Teacher showing a holographic DNA image in the classroom
Credit: FG Trade / Getty Images / E+

Within each cell, billions of biomolecules—nucleic acids, proteins, carbohydrates, lipids, and metabolites—dance to the rhythm of countless chemical reactions per second in a sea of trillions of water molecules. Imagine being able to zoom in, slow down, predict, and rewrite the story of this microscopic universe. This vision, once limited to science fiction, is now on its way to becoming a reality thanks to the concept of the “virtual cell.”

A collaboration between the Arc Institute, 10x Genomics, and Ultima Genomics, with echoes of the Human Genome Project, is catalyzing this potential paradigm shift in biology and medicine. Together, the blossoming research organization and two biotech industry leaders are driving the development of the Arc Virtual Cell Atlas—a repository of computation-ready, single-cell measurements designed to model cellular behavior.

“This is not just about individual research efforts, but also is about building a broader scientific community,” Patrick Hsu, co-founder of the Arc Institute, told Inside Precision Medicine. “Solving these challenges requires partnerships across academia, industry, and technology sectors. It’s inspiring to be part of a collective effort to tackle questions that once seemed insurmountable.”

Just two months after its launch with over 300 million cells, the Arc Virtual Cell Atlas is expanding rapidly. With the Arc Virtual Cell Atlas now exceeding 400 million cells and growing, this initiative aims to redefine biomedical research and drug discovery, paving the way for scientific innovations previously thought impossible.

“Talking in tangible ways about building a virtual cell, simulating biology in a computer, is a little mind-bending,” Serge Saxonov, PhD, co-founder and CEO of 10x Genomics, told Inside Precision Medicine. “Growing up, that notion has felt very fantastical. It’s amazing to be here, discussing concrete plans with researchers and scientists who will build and install them.”

The collaboration is integrating cutting-edge technologies from 10x Genomics and Ultima Genomics, making data collection faster, more detailed, and more affordable than ever. By enabling researchers to explore how cells respond to various conditions, the Virtual Cell Atlas is intended to be a critical resource for predictive biological models—models that could accelerate insights into disease mechanisms and revolutionize drug development.

“I don’t think people understand quite how complex a cell is, the scale that will be needed, and how incredibly impactful it will be when it’s solved,” Gilad Almogy, PhD, co-founder and CEO of Ultima Genomics, told Inside Precision Medicine. “All the forces—the work by Serge, 10x, and others in making tagging so many cells possible, the science being done by Patrick and others by perturbing cells, and what has happened in the world in AI over the past few years—came together for this magical moment. But it’s clear that this program is going to make a big impact, but it’s far from the end of the road—it’s ground zero for revolutionary biology.”

What in the world is a “virtual cell”

The idea of the virtual cell is not a novel concept but rather a shared mission that is gaining momentum in scientific and technological circles. Hsu elaborated, “It’s in the zeitgeist, and people are asking, ‘What is the virtual cell? What does it mean? Sounds interesting, but what can it actually do?’” The virtual cell aims to address a fundamental biological challenge: the slow, resource-intensive process of running real-world experiments on cells, tissues, and animals.

Hsu envisions the virtual cell as predictive models that act as a scientific co-pilot—an AI system capable of understanding cells and representing their dynamic states. Such a tool, Hsu believes, could revolutionize biology by accelerating research exponentially. “The virtual cell could be a co-pilot that understands cells…to accelerate scientific research by orders of magnitude, potentially by parallelizing things that would normally take real-world experiments into GPUs,” said Hsu.

The concept of the virtual cell isn’t just about throwing AI against a wall of biological data and seeing what sticks. Hsu emphasized that it is not just about “cramming AI into different domains” but solving deeply biological problems in a strategic, highly orchestrated endeavor that combines advanced AI and large-scale cellular data to unlock new scientific frontiers. 

The challenge of cell complexity

Biological systems are enchantingly labyrinthine. Even within a single organism, such as a human, cells exhibit remarkable diversity despite sharing the same genome. “I think it’s vastly underrated how complex biology is and how little of it we actually understand,” said Saxonov. “The cell is really the fundamental unit of biology, and every one of us is about 37 trillion cells. The sheer number of states, tissues, and conditions in which cells can exist is truly astounding.”

Indeed, the malleable map of human cell identities is becoming simultaneously sharper while undergoing constant revision as new dimensions are woven in by the hands of innovation. Early research focused on basic cell types identified through microscopy, while advancements in molecular biology and sequencing have revealed a more complex and diverse landscape of human cells. The number of known human cell types has evolved significantly over time, from initial estimates of around 200 to more recent estimations reaching 400 or even more.

This complexity presents both a challenge and an opportunity. Understanding cellular behavior requires data that spans multiple modalities, such as DNA, RNA, proteins, and spatial phenotypes. Hsu explained the importance of integrating these modalities much like tokens—the fundamental building blocks of text that large language models (LLMs) process to understand and generate human language. “Biology has its own ‘tokens,’ much like the language tokens in AI models,” said Hsu, “and by unifying these layers in AI models, we can tackle the combinatorial diversity of biology.”

Prepare to be perturbed

However, creating such models, as is the goal with the virtual cell, is no small feat. One of the keys to unlocking the virtual cell is perturbation data. Almogy explained, “All my cells have roughly the same genome, but their states and functions are vastly different. When you overlay the natural diversity of these cells with the variety of perturbations we can apply, the universe of possible states becomes mind-boggling.”

Biology has long been defined by the interplay between observation and perturbation. “If you look at the history of biotechnology, it has been about breaking down biological complexity,” said Hsu. “Starting with the invention of the microscope in the Dutch glassblowing tradition, we learned to magnify cells and began exploring the individual units of life.” From these humble beginnings, the field evolved rapidly. Researchers learned to culture cells in aseptic conditions, extract natural products, and develop small molecules and antibodies, gradually building a toolkit that could perturb cellular states.

The Arc Virtual Cell Atlas represents a shift from observational to perturbational data. Hsu elaborated, “Creating high-quality perturbational data is critical. This involves not only measuring pathways but also ensuring the data meets rigorous quality-control metrics. By combining these datasets, we can train ML models to predict cellular responses and design better drugs.”

Hsu added, “Observational data is valuable, but to train machine learning (ML) models capable of understanding causality, we need high-quality perturbational data. At Arc, we’ve spent years building functional genomics capabilities to create the world’s largest dataset of perturbed single cells. This data forms the foundation for training models that can predict and manipulate cell states.”

Scaling data acquisition

Data acquisition lies at the heart of creating a virtual cell. “Biology has a source code—DNA, RNA, and proteins—and the biological sciences are incredibly complicated and information-starved,” said Almogy. “The founding premise of Ultima is the general notion that biological sciences are not progressing as fast as they should because they don’t have enough information. Our instruments are designed to read massive amounts of DNA with unparalleled efficiency. Whether it’s sequencing an inherited genome, analyzing cell-free DNA in the bloodstream, or studying single-cell perturbations, our goal is to digitize biology at scale.”

Almogy’s background in the semiconductor industry shaped his perspective, likening the demand for biological data to Moore’s Law. “In semiconductors, you never knew what would drive the next need, but you always knew it was coming,” said Almogy. “Similarly, we’ve always believed in the growing need for biological data. When we started, we didn’t anticipate the convergence of single-cell perturbation and AI. We invested in scalability—whether it’s the three billion bases of DNA in the inherited human genome, the nearly ten billion people, or the 37 trillion cells in a human, or the size of the experiments that Patrick is planning. This is a form where you need massive amounts of data that we did not see coming, and we’re thrilled to be part of this transformative moment.”

Saxonov’s 10x Genomics was similarly established with Moore’s Law in mind, thinking about what kind of data could be generated with the expected advancements in sequencing. “Our premise was that DNA reading capabilities would continue to scale, effectively removing the bottleneck for data acquisition,” said Saxonov. “This assumption allowed us to focus on converting biological samples into multi-omic measurements. The rise of companies like Ultima complements this vision, enabling us to analyze diverse phenomena with unprecedented depth.”

Leveraging 10x Genomics’ high-resolution single-cell analysis platforms, particularly its Chromium Flex technology, Arc researchers can generate perturbational data at an unparalleled scale and quality. This technology enables the examination of millions of perturbed individual cells simultaneously, at the lowest per-cell cost.

Additionally, Ultima Genomics’ UG 100 sequencing system and recently launched Solaris chemistry are pivotal to the Virtual Cell Atlas’s next phase. Ultima’s innovative wafer-based sequencing architecture and high-throughput Solaris Boost mode allow Arc to produce significantly more data at a fraction of the cost of alternative methods. Arc researchers have extensively validated UG 100, finding that it integrates seamlessly with 10x’s technology for high-quality data generation.

AI as the integrative force

The Arc Virtual Cell Atlas is being developed as an open-source resource, the world’s largest repository of multi-omic single-cell data, uniformly processed to minimize technical artifacts. Explaining its significance, Hsu said, “Over the years, single-cell data has been generated using various computational tools, leading to inconsistencies. By reprocessing this data with modern methods, we’ve created a standardized dataset that anyone can use to build virtual cell models.”

The virtual cell initiative relies heavily on AI to make sense of the massive datasets generated by 10x Genomics and Ultima Genomics. Saxonov described this convergence as an “exciting inflection point” and said, “We’re on an exponential path where data acquisition, AI, and biology are coming together. While a generalized model for all cellular behavior is still years away, the progress we’re making now is setting the stage for rapid breakthroughs in drug discovery and precision medicine.”

Hsu emphasized the role of ML in uncovering higher-order patterns: “Machine learning provides a universal framework for processing unstructured data. By integrating multi-omic measurements, we can create models that not only predict cellular behavior but also simulate novel cell states.”

The road to a new paradigm for drug discovery

The potential of a virtual cell is undeniable. “This is more than a technical achievement; it’s a vision for the future of science,” said Saxonov. “By integrating biology, technology, and AI, we can accelerate the progress of medicine by orders of magnitude.”

One of the most exciting applications of the virtual cell is drug discovery. “It’s not just about understanding basic mechanisms but also about transforming how we discover and develop drugs,” Hsu said. By simulating experiments computationally and harnessing the power of GPUs, the virtual cell has the potential to parallelize and drastically shorten drug development processes that would otherwise take years.

Hsu, highlighting the shift from descriptive to predictive models, said, “Traditionally, drug discovery has relied on trial and error. Machine learning has given us a universal paradigm for how to take massive amounts of unstructured, unlabeled data and identify emergent, higher-order patterns. We hope that virtual cell models will enable us to do that at the level of cell architecture and function to better understand biology. By using AI-driven models, we can predict how drugs will interact with cellular systems, dramatically accelerating the development process. Fundamentally, drug matter is a way of trying to perturb cell states, so going from this paradigm of simply reading things to knowing how to understand and react is something that we’re excited to explore.”

This capability is especially critical for complex diseases like cancer and autoimmunity, where understanding T cells’ behavior is paramount. “T cells are not only important for basic science but also for therapeutic applications,” Hsu explained. “An accurate T cell model would be transformative for cancer immunotherapy and aging research.”

Whose virtual cell?

The high-quality, curated, open datasets of the Arc Virtual Cell were assembled using both observational and perturbational data, namely Tahoe Bio’s Tahoe-100M and the Arc’s AI agent-curated scBaseCamp dataset. The scBaseCamp boasts over 230 million single cells, making it the largest public repository of single-cell data, surpassing other prominent resources, such as CZ CELLxGENE (107 million cells) and the Human Cell Atlas (65 million cells). It also covers a broader range of biological contexts, spanning 21 organisms and 72 distinct tissues, which raises a question and possible concern: Who does the virtual cell represent? From this description, it sounds more like a universal protocell than a purely human cell.

Along these lines, one thing I asked in my conversation with Hsu, Saxonov, and Almogy that got skirted had to do with whether there was a strategy for choosing the cells that will be used for constructing the virtual cell going forward. I would speculate that a virtual cell would benefit from using a mix of the following approaches: (1) a comprehensive map of primary cells from at least one individual, (2) several cell matching types from a cohort of individuals whose genomic diversity covers a majority of the global population, and (3) starting from embryonic or pluripotent stem cells that are then being differentiated into tissues.

Nevertheless, while listening to Hsu, Saxonov, and Almogy talk about the virtual cell, I couldn’t help but begin daydreaming of a holographic projection of a cell akin to a planet, star system, or Death Star in Star Wars that can be zoomed in and out and spun around with one’s hands (or even directly from the mind). I began to think of how this wobbling, floating, spherical biological universe, by clicking on a particular gene and inserting a mutation or modulating its expression, could then transform from one cell state or type to another. I could see a cardiomyocyte contracting rhythmically, then causing it to hiccup due to arrhythmia or dedifferentiating into a more progenitor state, only to then enucleate and become a mature erythrocyte.

I don’t think it’s unreasonably optimistic to think that I’ll live to see it.

Also of Interest