The U.S. government, tech behemoths, and leading scientific institutions have joined forces in a $1.8 billion initiative to build massive, open biological datasets designed to train predictive artificial intelligence (AI) models.

The Virtual Biology Initiative consortium, announced Wednesday, brings together the Chan Zuckerberg Biohub, Department of Energy (DOE), the National Institutes of Health (NIH), and corporate titans Meta Platforms Inc., Alphabet Inc.’s Google DeepMind, and Isomorphic Labs.

The effort, which seeks to bridge a critical data gap in modern medicine, aims to digitize cellular research and drastically compress drug development timelines from decades to just a few years.

While current cellular datasets encompass hundreds of millions of cells, experts estimate that reliable AI models require datasets spanning billions or trillions of cells to accurately map biological systems.

“We need to capture the language of biology, we need to capture the language of the cell. And that doesn’t exist today,” said Alex Rives, head of science at Biohub. “An accurate predictive model of biology could dramatically accelerate scientific discovery by enabling scientists to perform experiments digitally.”

The $1.8 billion commitment draws from both public coffers and private industry capital.

The Department of Energy is investing more than $500 million over five years through its Genesis Mission. Funding will deploy exascale supercomputers, X-ray and neutron scattering, cryo-electron microscopy, and self-running laboratories to generate high-resolution cellular data.

The National Institutes of Health is contributing access to existing datasets and repositories built through more than $500 million in prior federal investments. Biohub will standardize these records for AI training.

Chan Zuckerberg Biohub is contributing $500 million pledged earlier this year, allocating $400 million to advanced cellular measurement tools and $100 million to external research grants.

Meta, Google DeepMind, and Isomorphic Labs are jointly investing $300 million to support data generation and model development.

Additionally, NVIDIA Corp. will provide compute infrastructure and specialized software, while academic partners — including the Broad Institute, the Allen Institute, and the UK’s Wellcome Sanger Institute — will assist in research execution.

To incentivize private investment, the initiative utilizes a hybrid open-science model. While all government-funded data will be made public immediately without restrictions, datasets funded by commercial partners will subject private sponsors to an exclusive access period (reported as one year) before becoming available as a public scientific resource.

“We have always held this as a community asset, not just for one group, so that it can build upon itself over time,” said Dr. Priscilla Chan, co-founder of the Biohub alongside Meta CEO Mark Zuckerberg.

The alliance expects its first dataset within a year, targeting fully operational predictive models within five years.

The launch comes amid an escalating race across the tech sector to apply AI to biological sciences. Anthropic recently expanded its footprint with a dedicated wet lab, while the OpenAI Foundation launched a $125 million grant program for medical datasets.

Internationally, Paris-based startup Rivercell secured $25 million on Wednesday to build a virtual cell model, underscoring a swift global transition toward AI-driven digital biology.