Biohub, together with the U.S. Department of Energy (DOE), the National Institutes of Health (NIH), Google DeepMind, Isomorphic Labs, and Meta, has announced a major expansion of the Virtual Biology Initiative, bringing the total commitment to about $1.8 billion in funding, data, computing power, and measurement technology. The goal is ambitious: to build the open, AI-ready datasets that scientists will need to create predictive models of how human cells behave, and how they respond when something goes wrong.
The expansion was announced on October 7, 2026, just over five months after Biohub first launched the initiative in April with a founding commitment of $500 million.
Why This Matters
AI has already transformed parts of biology. Protein structure prediction, a challenge that resisted solution for decades, became computationally tractable in just a few years. The key to this success is often overlooked: the Protein Data Bank, a comprehensive, freely available database of experimentally solved protein structures.
For the cell as a whole, no equivalent resource exists yet. To predict how a cell will respond to a drug, a mutation, or a disease, a model needs to learn from enormous amounts of consistent data showing cells across many types and conditions, both at rest and when deliberately disturbed. Today, that data is scattered across labs, generated with different methods, and often too inconsistent to train AI models reliably. The Virtual Biology Initiative is designed to close that gap.
Where the $1.8 Billion Comes From
A large share of the commitment is not cash, but data, computing resources, and lab capacity. According to the announcement, the main contributions are:
- Biohub – $500 million (founding commitment): about $400 million for new measurement technologies and engineering tools, and about $100 million to fund research outside Biohub.
- U.S. Department of Energy – more than $500 million over five years: lab measurement, modeling, and computation through the Genesis Mission, a DOE-led cross-agency effort. This draws on national laboratory resources spanning imaging, AI analytics, exascale supercomputing, and autonomous labs.
- National Institutes of Health – existing data: NIH will coordinate datasets, repositories, and knowledge bases built from more than $500 million in prior federal investment, including resources from the National Library of Medicine and NCBI. Biohub will work with NIH to standardize this data for AI training.
- Google DeepMind, Isomorphic Labs, and Meta – $300 million combined: to build technologies and multimodal datasets for predictive models of biology.
NVIDIA is supporting the effort with accelerated computing and technical expertise, and Renaissance Philanthropy is helping expand funding for data generation. Major research organizations, including the Allen Institute, Broad Institute, Gladstone Institutes, Human Cell Atlas, Human Protein Atlas, and Wellcome Sanger Institute, are also part of the effort.
What Kind of Data Will Be Generated?
The initiative focuses on multimodal data, meaning different types of measurements of the same biology, captured at several scales. According to Biohub, this includes genomic, transcriptomic, proteomic, cellular, and tissue-level data, generated with techniques such as spatial omics, cryo-electron tomography, and large-scale microscopy of living tissues and organisms.
A key emphasis is on cell response data: measuring how cells change when they are perturbed, across far more cell types and conditions than have been studied so far. This kind of cause-and-effect data is exactly what predictive models need, and exactly what is hardest to find in today’s public databases.
The effort also builds on Biohub’s existing protein AI work, including the ESM family of models and the ESM Atlas of protein sequences and structures.
Open by Design
Biohub says the data it generates will be made openly and freely available to the global research community. The partners describe a shared resource built on common standards and identifiers, with a single point of access. Detailed licensing and access terms have not yet been published, so it remains to be seen exactly how researchers will be able to use each dataset.
This focus on openness follows a broader trend in the field. Last month, Google DeepMind released AlphaGenome Atlas, a free resource covering the predicted effects of all 9 billion possible single-letter DNA variants. Large, openly shared resources like these are increasingly becoming the foundation that the rest of the research community builds on.
What It Means for Bioinformatics
For bioinformaticians and computational biologists, initiatives like this shift the bottleneck. As standardized, AI-ready data becomes available at scale, the most valuable skills will include:
- Working with multimodal data: integrating transcriptomics, proteomics, imaging, and spatial data rather than analyzing each in isolation.
- Single-cell and spatial analysis: the core data types behind cell-level models.
- Data standards and metadata: understanding ontologies, identifiers, and reproducible pipelines, which decide whether data can actually be used to train models.
- Using and evaluating foundation models: knowing what predictive models can and cannot do, and how to test them against real biology.
A Long Road Ahead
It is worth being clear about what has and has not been achieved. This announcement is about building the data and infrastructure for future models, not about a working “virtual cell” today. Predicting how any human cell will respond to any intervention is one of the hardest problems in biology, and the DOE commitment alone runs over five years. Some details also remain to be worked out, including the specific cell types and conditions to be prioritized and the exact terms under which data will be shared.
Still, the scale and breadth of the coalition are notable. Bringing together government agencies, leading AI companies, and major research institutes around a shared, open data resource could do for cell biology what the Protein Data Bank did for protein structure: give AI the foundation it needs to make real predictions about life.
Virtual Biology Initiative: Frequently Asked Questions
The Virtual Biology Initiative is a Biohub-led effort, launched in April 2026, to generate open, AI-ready biological data for building predictive models of human cells. In October 2026, it expanded to about $1.8 billion in funding, data, computing power, and measurement technology with partners including the U.S. Department of Energy, NIH, Google DeepMind, Isomorphic Labs, and Meta.
Biohub provides a $500 million founding commitment, the U.S. Department of Energy contributes more than $500 million over five years through the Genesis Mission, NIH coordinates datasets built from more than $500 million in prior federal investment, and Google DeepMind, Isomorphic Labs, and Meta are investing $300 million combined. NVIDIA supports the effort with computing resources.
A virtual cell is a computational model that aims to predict how a real cell will behave, for example, how it responds to a drug, a genetic mutation, or a disease. Building one requires large amounts of consistent data on many cell types under many conditions, which is what the Virtual Biology Initiative sets out to generate.
Biohub says the data it generates will be openly and freely available to the global research community, built on shared standards and a single point of access. Detailed licensing and access terms have not yet been published.
The Genesis Mission is a cross-agency initiative led by the U.S. Department of Energy. Through it, DOE is contributing lab measurement, modeling, and computation from its national laboratories, including imaging, AI analytics, and exascale supercomputing, to the Virtual Biology Initiative.
AlphaFold predicts the 3D structure of individual proteins and was trained on decades of openly shared data in the Protein Data Bank. The Virtual Biology Initiative works one level up: it aims to create the open data needed to model whole cells, so that AI can predict how cells behave and respond, not just what their proteins look like.
Sources
Disclaimer:
The research discussed in this article was conducted and published by the authors of the referenced sources. CBIRT has no involvement in the research itself. This article is intended solely to raise awareness about recent developments and does not claim authorship or endorsement of the research.
Important Note: This blog is based on an organizational announcement rather than a peer-reviewed publication. It reports on a funding and data initiative, not on established research findings.
Follow Us!
Learn More:
Dr. Tamanna Anwar is a Scientist and Co-founder of the Centre of Bioinformatics Research and Technology (CBIRT). She is a passionate bioinformatics scientist and a visionary entrepreneur. Dr. Tamanna has worked as a Young Scientist at Jawaharlal Nehru University, New Delhi. She has also worked as a Postdoctoral Fellow at the University of Saskatchewan, Canada. She has several scientific research publications in high-impact research journals. Her latest endeavor is the development of a platform that acts as a one-stop solution for all bioinformatics related information as well as developing a bioinformatics news portal to report cutting-edge bioinformatics breakthroughs.











