Biohub leads $1.8B initiative to create AI-ready biology data with federal and tech partners
Biohub, a research nonprofit backed by Mark Zuckerberg and Priscilla Chan, is leading a $1.8 billion initiative to develop AI-ready biology datasets. The effort includes partnerships with federal agencies and tech giants like Meta, Google DeepMind, and Isomorphic Labs, aiming to train AI models to predict cellular behavior and accelerate scientific discovery.
Editor, Lazyfounder

30 SEC SUMMARY
- Biohub, a research nonprofit backed by Mark Zuckerberg and Priscilla Chan, is leading a $1.8 billion initiative to create AI-ready biology datasets.
- The initiative includes federal agencies like the US Department of Energy and the National Institutes of Health, alongside tech giants Meta, Google DeepMind, and Isomorphic Labs.
- The goal is to train AI models to predict cellular behavior, accelerating drug development and scientific research.
- Commercial funders will receive one year of exclusive data access before datasets become public.
- The first dataset is expected within a year, with accurate predictive models targeted within five years.
TABLE OF CONTENTS
KEY HIGHLIGHTS
- Biohub is leading a $1.8 billion initiative to develop AI-ready biology datasets for training predictive AI models.
- The US Department of Energy is committing over $500 million for lab measurements, modeling, and computing.
- The National Institutes of Health (NIH) will contribute datasets built with over $500 million in prior federal funding.
- Meta, Google DeepMind, and Isomorphic Labs are collectively contributing $300 million to the initiative.
- Biohub’s Virtual Biology Initiative will receive $500 million, with $400 million allocated for in-house research and $100 million for external projects.
- Datasets will be publicly released after a one-year exclusive access period for commercial funders.
A $1.8 billion commitment to AI-ready biology data
Biohub, a research nonprofit backed by Mark Zuckerberg and Priscilla Chan, is leading a $1.8 billion initiative to create AI-ready biology datasets. The effort includes contributions from federal agencies such as the US Department of Energy and the National Institutes of Health (NIH), alongside tech companies Meta, Google DeepMind, and Isomorphic Labs. According to SiliconANGLE and The Next Web, this is the largest coordinated effort to date for developing biology data optimized for AI training.
The initiative aims to train AI models to predict cellular behavior, which could accelerate drug development and scientific research. Datasets will be standardized to ensure compatibility with AI systems, with a first release expected in about a year.
Who is contributing and how funds will be used
The US Department of Energy is the largest contributor, pledging over $500 million over five years for lab measurements, modeling, and computing. According to The Next Web, these funds will support advanced tools like exascale supercomputers, X-ray and neutron scattering, cryo-electron microscopy, and automated labs.
The NIH will provide access to datasets developed with over $500 million in prior federal funding. Biohub will standardize these datasets for AI training, ensuring consistency and usability.
Biohub itself is committing $500 million to its Virtual Biology Initiative. Of this, $400 million will fund in-house research, including cryo-electron microscopy (cryo-ET) to collect detailed cell data. The remaining $100 million will support external research projects.
Meta, Google DeepMind, and Isomorphic Labs are collectively contributing $300 million. Nvidia will also support the initiative by providing accelerators and specialized software for research workloads. Additional partners include the Allen Institute, Broad Institute, Gladstone Institutes, Wellcome Sanger Institute, Human Cell Atlas, and Human Protein Atlas.
Renaissance Philanthropy is assisting in raising additional funds, and Biohub plans to engage drug companies and philanthropies for further support.
Data access and timeline
Commercial funders, such as Meta and Google DeepMind, will receive one year of exclusive access to the datasets they help fund. After this period, the data will become publicly available. Government-funded work, however, will be accessible without restrictions from the outset, according to The Next Web.
The consortium aims to release its first dataset within a year. The goal is to develop accurate predictive models of cellular behavior within five years, though current datasets are still far from the billions or trillions of cells required for highly reliable simulations.
Competing efforts in AI-driven biology
Biohub’s initiative is not the only effort in this space. Anthropic has built its own biology lab to develop AI-driven drug discovery tools, while the OpenAI Foundation has launched a $125 million grant program for biology datasets. Additionally, French startup Rivercell recently raised $25 million to build an AI virtual cell, as reported by The Next Web.
Why this initiative matters for biology and AI
Traditional biological research often requires years of work and millions of dollars to study how cells respond to new therapies. AI-driven simulations could dramatically reduce this time and cost, enabling faster and more efficient drug development.
Creating accurate virtual cells requires advanced AI models trained on vast, high-quality datasets. This initiative addresses that need by pooling resources from public and private sectors to generate standardized, AI-ready biology data.
What this means
Lazyfounder analysis — our interpretation, not reported fact.
This initiative is a significant step toward bridging the gap between biology and AI. For founders and operators in biotech, pharmaceuticals, or AI, the datasets generated could become foundational resources for training proprietary models or developing new therapies. The one-year exclusive access period for commercial funders also creates a potential advantage for early partners, who may gain a head start in leveraging the data for drug discovery or other applications.
That said, the scale of the challenge is enormous. Current datasets are still far from the volume and granularity needed for highly accurate simulations. The success of this initiative will depend on sustained collaboration, additional funding, and advances in AI modeling. Competitors like Anthropic and OpenAI are also investing heavily in this space, so differentiation will be key for startups looking to carve out a niche.
For policymakers and investors, this initiative signals strong government and industry interest in AI-driven biology, which could lead to more funding opportunities or regulatory support for related projects.
Key takeaways
- This initiative represents a major public-private effort to standardize biology data for AI, potentially reducing the time and cost of drug development.
- Founders in biotech and AI should monitor the datasets’ release timeline, as they could become critical resources for training proprietary models.
- The one-year exclusive access period for commercial funders may create early opportunities for partnerships or licensing deals.
- Federal involvement signals strong government interest in AI-driven biology, which could lead to additional funding or policy support for related projects.
- Competing efforts, such as those by Anthropic and OpenAI, highlight the growing importance of AI in biology and the need for startups to differentiate their approaches.
FAQ
What is the goal of the $1.8 billion initiative led by Biohub?
The initiative aims to create AI-ready biology datasets to train AI models that can predict cellular behavior. This could accelerate drug development and scientific research by enabling faster and more cost-efficient simulations of cell responses to therapies.
Who are the key contributors to this initiative?
The initiative includes federal agencies like the US Department of Energy and the National Institutes of Health, as well as private sector partners such as Biohub, Meta, Google DeepMind, Isomorphic Labs, and Nvidia. Biohub is committing $500 million, while Meta, Google DeepMind, and Isomorphic Labs are collectively contributing $300 million.
When will the datasets be available to the public?
Commercial funders will have exclusive access to the datasets for one year after their creation. After this period, the data will be released publicly. However, government-funded work will be available without restrictions from the start.
How does this initiative compare to other efforts in AI-driven biology?
Other efforts include Anthropic’s biology lab, OpenAI’s $125 million grant program for biology datasets, and Rivercell’s $25 million funding to build an AI virtual cell. Biohub’s initiative is notable for its scale and public-private collaboration, making it the largest coordinated effort of its kind so far.
Related on Lazyfounder
Sources
- The Next Web · 2026-10-07
Biohub, Meta, Google DeepMind and the US pool $1.8bn for AI biology data - SiliconANGLE · 2026-10-07
US government, tech giants and Biohub commit $1.8B to AI biology initiative - The Decoder · 2026-10-07
Zuckerberg's Biohub leads a $1.8 billion push to build AI models that predict cell behavior
This story is an original summary drafted with AI by Lazyfounder from the reporting listed above and checked by automated validation. Facts are attributed to their original publishers; sections marked as analysis are Lazyfounder's. Where a source is in another language, facts were machine-translated and quotations are reported, not reproduced. Read the original coverage via the links, and see our AI policy and corrections policy.
About the author
Editor, Lazyfounder
Tarun Mottlia edits LazyFounders, covering Indian startups, funding rounds, AI and product launches. Every story on the site is AI-assisted and checked against its cited sources before publication.
More stories by Tarun MottliaGet the LazyFounder Brief
Startup, funding and AI news in a five-minute read. Join the early-access list.


