[BidClub_]
No Priors · · 58 min

No Priors Ep. 103 | With Vevo Therapeutics and the Arc Institute

Sarah GuoNima AlidoustJohnny YuPatrick HsuDave BurkeHani Goodarzi

YouTube
TL;DR
  • Tahoe 100 shifts single-cell perturbation data toward machine-learning scale. Vivo reports 100 million single-cell measurements across 50 cancer models from different patients, 1,200 drug treatments, and 60,000 drug–cell or drug–cell-line interactions, versus roughly 1–2 million publicly available perturbational data points before it. Dave Burke compares the hoped-for impact to an “ImageNet moment” for cellular biology.

  • The differentiated asset is causal, diverse, consistently generated data—not another pile of healthy-cell observations. Perturbations provide a before-and-after signal, while disease models expose states absent from normal tissue. Four people performed consistent work across 60,000 experiments with effectively no batch effects. The team argues that scale and breadth can make the data itself generate hypotheses and surprises.

  • Virtual cells address the systems-level question that protein models leave open. Protein language models can learn folding, binding, and structural biology, but a drug acts on a protein embedded in a cell, tumor, and broader system. The proposed virtual cell is a “notional CPU” that predicts how a genetic edit or chemical changes the transcriptome—and eventually solves the inverse problem of moving a diseased state toward health.

  • Current virtual-cell models remain commercially immature, with predictive performance on differentially expressed genes on the order of 10% and no accepted benchmark. Tahoe 100 represents roughly 200–300 billion training tokens under current encodings. Dave describes roughly 1 trillion tokens as a comfortable comparison point from other domains, while stressing that biology’s scaling laws remain uncertain. Hani Goodarzi places protein models beyond GPT-3 maturity but cellular models “closer to GPT-1 than 2.”

  • Open sourcing is an operating strategy, not philanthropy detached from Vivo’s economics. Tahoe 100 joins Arc’s roughly 230 million-cell scBaseCamp collection to form a 330 million-cell Virtual Cell Atlas; the community can train and critique models while Vivo continues building its Mosaic data-generation platform. Nima’s framing is “a new stake in the ground”: a small internal team can recruit an external research ecosystem without building a large organization.

  • If virtual cells work, the potential prize is simultaneous compression of target risk, chemical search, and experimental time. Roughly 90% of drugs fail in clinical trials; the guests argue failures reflect both poor drug matter and selection of the wrong target. A sufficiently accurate model could search tens of millions of compounds, generalize across patient contexts, and help teams “measure twice and cut once”—but at 10% accuracy, Patrick Hsu cautions, “you’re just simulating noise.”

  • The technical inflection may be now, but therapeutic proof will remain slow. Chinese biotechs’ cost, pace, antibody manufacturing, and IND-enabling packages are pushing US companies toward leaner teams and stronger vendors; Nima calls it “morning in bio” and urges building now rather than promising results in three to five years. Sarah Guo stresses that treatments commonly require 11-plus years, and even moving success from 10% to 30% leaves a “law of small numbers” that may take a decade-long window to validate.

Digest · the substance, structured for research

1. Tahoe 100 makes perturbational biology a machine-learning-scale resource

  • Johnny Yu defines Tahoe 100 as the world’s largest single-cell RNA-sequencing dataset: 100 million cells spanning 50 cancer models from different patients, 1,200 drug treatments, and 60,000 drug–cell or drug–cell-line interactions. His strongest claim is that it may be “the first data set that’s going to enable machine learning in this space.”

  • The relevant comparison is not merely total cells. Publicly available perturbational data previously amounted to roughly 1–2 million single-cell points, while most historical datasets were observational, fragmented, poorly labeled, and drawn from healthy tissue rather than disease states.

  • Dave Burke’s analogy is ImageNet in 2009: a sufficiently large, purposeful dataset can produce a nonlinear capability jump. The hope—explicitly still a hope—is that cellular perturbations do for systems biology what foundational protein-structure datasets, including PDB and CASP, helped enable for models such as AlphaFold.

2. A virtual cell models causality beyond protein structure

  • Nima Alidoust distinguishes two languages: protein models learn “the language of structural biology”—folding, molecular binding, and antibody–protein interaction—while virtual cells aim at “the language of systems biology,” where a target sits inside a cell, tumor, immune environment, and broader organism.

  • Dave’s computing analogy makes the abstraction concrete: DNA is ROM, encoding the cell; RNA is RAM, its changing working memory; and a virtual-cell model infers a “notional CPU” that maps an edited gene or applied drug into a transcriptomic response.

  • The inverse problem is the therapeutic one: given a diseased cell’s expression profile, which genetic or chemical perturbation might return it toward a healthy state—or, for cancer, kill it while sparing healthy cells? That capability remains an intended destination, not a demonstrated product.

  • Hani Goodarzi’s case for perturbational data is causality. Observing a liver sample yields associations; intervening genetically or chemically creates a defined before and after, allowing a model to learn changes that drive cell state rather than merely accompany it.

3. Information content matters more than another redundant million cells

  • Nima reports that early single-cell foundation models could discard 99% of a roughly 60-million-cell collection with little performance loss. His inference is that much of the available data repeated similar biological contexts, so raw scale overstated how much the model could learn.

  • Diversity supplies the missing information. A model intended to reason about heart, brain, liver, bone, or cancer needs normal and diseased states across cell and tissue types; perturbations help explore the high-dimensional “manifold” rather than repeatedly sampling one neighborhood.

  • Each cell contributes roughly 2,000–5,000 gene-and-expression tokens, making Tahoe 100 approximately 200–300 billion tokens under the team’s estimate. Dave cites about half a trillion tokens for GPT-3 and 700 billion for ESM-3, with roughly 1 trillion as a comfortable comparison point, but says the field will not know the relevant scaling law until it gets there.

  • The bottleneck also varies by domain. Decades of genome sequencing make DNA models increasingly compute- and context-length-limited; cellular models remain data-limited because scalable single-cell profiling emerged only recently.

4. Mosaic replaces one-patient-at-a-time screening with pooled experiments

  • Vivo’s Mosaic platform pools cells from many patient-derived models—including multiple cancer types and distinct genetics—into one “mosaic tumor,” then screens hundreds or thousands of drugs while resolving each model’s response. Sarah’s reaction captures the jump from serial experiments to multiplexed biology.

  • Initial perturbations target cancer-relevant genes, growth, DNA regulation, and pan-cancer pathways. Hani argues that these conserved pathways may broadly apply to neuroscience and immune-cell development, while future datasets could add rare diseases and use model feedback to fill gaps.

  • Scale changes experimental philosophy as well as throughput. Nima says Vivo can generate about 50 times more perturbational data in five weeks than was publicly available from the previous decade, reducing the need to preselect a narrow hypothesis: teams can go large on chemical and patient-sample space and become more unbiased.

5. Open data lets a tiny company borrow the field’s intelligence

  • Nima says Vivo chose to open-source the data within hours of discussing the opportunity for two reasons: to put “a new stake in the ground” for expected dataset scale, and to keep an internal team of roughly three or four people focused while recruiting outside researchers to expose strengths, defects, and model opportunities.

  • Arc’s Virtual Cell Atlas combines Tahoe 100 with scBaseCamp, an approximately 230 million-cell observational collection, for roughly 330 million cells. The proposed workflow is complementary: one could potentially pretrain on broad observations, then add perturbations to teach dynamics and improve prediction.

  • Arc built an agent likened to “a Google crawler and index” to find, annotate, and uniformly reprocess public sequencing data. That matters because changing tools, tool versions, genome builds, and reagent chemistries otherwise introduce analytical and experimental batch effects into merged datasets.

  • Operational consistency is itself part of the asset: Dave says four people performed exactly consistent work across 60,000 experiments. Fewer hands reduce the familiar biological caveat that a result works only “in my hands.”

6. Accurate simulation could compress biological time and chemical search

  • Patrick’s aging-project anecdote illustrates the constraint: one experimental round could require aging animals for two years. In-silico parallelism would be transformative, but he draws a hard line—if a virtual cell is only 10% accurate, “you’re just simulating noise.”

  • Vivo’s intended application is to predict how a new chemical entity interacts with cells from different patient models, asking whether it can move a diseased cell toward health—or, in cancer, kill it without killing healthy cells. Johnny’s future vision is a drug generated from a virtual-cell model, not a demonstrated capability today.

  • Nima separates two generalization problems. Across cells, models must transfer from observed patients to new individual variation; across chemistry, they must navigate tens of millions of candidate compounds and effectively unbounded biologics, identifying the small region worth synthesizing rather than screening familiar libraries incrementally.

  • With about 90% of drugs failing in clinical trials, Patrick argues the industry may have both bad drug matter—potency, toxicity, pH, kinetic profiles, and related properties—and the wrong targets. Virtual cells could narrow the target search before expensive chemistry and development begin.

7. The transcriptome is the abstraction, while cellular context remains visible

  • Dave describes RNA expression as a 1980s graphic equalizer with roughly 20,000 moving bars, responding to environment, stress, aging, health, and disease. Patrick says the group views the transcriptomic layer as rich enough to capture consequential state without simulating every molecular detail of the cell.

  • Multicellular modeling can then ladder upward through spheroids, organoids, in-vivo models, immune environments, and spatial data. Hani’s nuance is that environmental information is “filtered through the cell”; observing enough contexts may let a nominally single-cell model infer effects produced by surrounding tissue.

8. Platform companies are designed to abandon bad hypotheses

  • Nima contrasts Vivo with a single-hypothesis biotech, whose organization and incentives become committed to making one thesis work—sometimes advancing a drug after testing it on only three patient samples. A platform can generate competing hypotheses and remain “a lot more scientific” about what reaches the clinic.

  • Chinese biotech sharpens the operating challenge. Patrick points to lower cost bases, faster pipelines, strong safety and toxicology packages, IND-enabling work, and efficient antibody manufacturing; he sees that competition as potentially beneficial because patients, investors, and companies all want working molecules faster and cheaper.

  • Neither a vendor-chained “virtual biotech” model nor full vertical integration has solved the problem cleanly: the former proved slow, while owning everything proved expensive and bureaucratic. The proposed middle is capable vendors plus lean companies, with Nima arguing that the industry should incorporate these capabilities rather than rely solely on regulatory limits.

  • Nima’s “morning in bio” manifesto has three parts: build now rather than announce delivery in three to five years; organize small, concentrated teams of “superstars”; and challenge domain experts’ long lists of reasons cross-domain machine-learning techniques cannot work. His hedge remains intact: “Maybe it works, maybe it doesn’t work—but if you don’t try, you’ll never know.”

9. The technical inflection may be now, but therapeutic proof will lag

  • Dave reaches back to neural networks before ImageNet and AlexNet: ideas can appear unproductive until compute, data, and model architecture cross an inflection. Single-cell resolution, scalable perturbations, and modern training methods are the biological equivalents the group believes may now be converging.

  • Evo 2 is offered as evidence of emergent biological learning: Arc trained it on 9.3 trillion nucleotides without explicit DNA instruction, yet it learned ribosome-binding sites, codon degeneracy, and zero-shot BRCA1 variant pathogenicity with an area under the ROC curve of around 94%.

  • Maturity differs sharply by domain. Hani places protein language models beyond GPT-3, while Patrick and Dave describe cellular biology as roughly GPT-2-level or still developing toward it; Hani ultimately places virtual-cell models “closer to GPT-1 than GPT-2.” Existing best models predict differentially expressed genes with performance on the order of 10%, and no accepted benchmark yet settles comparisons.

  • Sarah notes her decade of skepticism about AI biotech and stresses that GPT-4-like cellular capability would not be instantly obvious: treatments generally require 11 years or more, and moving success from 10% to 30% would be extraordinary but still subject to the “law of small numbers.” Dave says better models could help “point the cannon in the right direction,” but proof would accumulate slowly over a roughly 10-year development window.

Sarah Guo

Today, we’re here with the CEO, CTO, and core investigator of the Arc Institute, as well as the co-founders of Tahoe Therapeutics, to talk about their release of Tahoe 100, the largest single-cell drug-perturbation data set ever created. We’ll also discuss where we are in AI for biology, why we need a virtual-cell model and not just protein-structure-prediction models, and when we should finally expect to see treatments from the growth of machine learning in biology.

Johnny Yu

Hi, I’m Johnny, and I work on single-cell sequencing at Tahoe Therapeutics.

Nima Alidoust

I’m Nima. I’m one of the founders, together with Johnny Yu. I’m a quantum chemist by background, but I’ve converted to being a computational chemist who loves playing with biological data. We’re building Vivo to really do that: to predict how chemicals interact with cells in different biological contexts. Some people call it the virtual cell. That’s basically what we’re working on.

Patrick Hsu

I’m Patrick Hsu, one of the founders of the Arc Institute, which is working at the interface of biology and machine learning to try to understand and, one day, treat complex human diseases—which are most of the major killers.

Dave Burke

I’m Dave, CTO at Arc Institute, focused on computational biology and building novel AI models for biology.

Hani Goodarzi

I’m Hani, a core investigator at Arc. I work very closely with Dave and Patrick to push our virtual-cell initiative.

Sarah Guo

Congratulations, everyone. It’s a big day. Let’s jump right into it. What is Tahoe 100, and what is its significance?

Johnny Yu

Tahoe 100 is the world’s biggest single-cell RNA-sequencing data set, and it enables a ton of machine-learning applications, including things like the virtual cell. It also enables a lot of drug-discovery applications. Broadly, in the context of where I think we are as a field, it’s the beginning of a different way of doing drug discovery: understanding how to build medicines and bringing AI and machine-learning people into the mix.

Nima Alidoust

Over the last 20 years or so, people have accumulated massive amounts of data when it comes to protein structures, protein function, and how drug molecules interact with proteins. One thing we haven’t had as much of is information about how different cells behave in different contexts, and how different genes in each of those cells function in the presence of the other genes in these different biological contexts.

We believe this is the era for that. We’ve seen the emergence of protein language models built on the data sets that have been accumulated over the last 2 decades, but now is the era for having data on cells: how they function, how they interact with drug molecules, and exactly what Johnny was describing. Tahoe is really a landmark data set that allows us to measure how drugs interact with different cells from different patient models. That gives us the ability to build models similar to the protein language models, but in the cellular context.

Dave Burke

If you think about the history of AI, it’s punctuated by these data sets that come about. If you think about ImageNet in 2009, which Fei-Fei Li put together, and what that did to drive a nonlinear jump in machine vision, I think the hope here is that, by producing perturbational data sets that allow us to elicit cellular responses, we’ll be able to drive forward the ability to model at the cellular level—not just at the protein level.

Hani Goodarzi

Lots of people have been talking about what those foundational data sets look like for biology. Those data sets have been really useful for training protein-structure-prediction models, like AlphaFold, which was built on CASP, the competition built on top of PDB data. But how do you do this for cells and cellular dynamics, which is really what tells us about biology and how it responds in health and disease?

I think those are the core steps forward where we want to bring up our ability to study higher levels of abstraction in biology—not just the individual molecular machines, but how they operate in the context of an entire cell.

Sarah Guo

Congratulations also to the entire Arc team. Given that you’re working on both virtual-cell models and protein-structure-prediction and protein-language models, can you contextualize why we need both and where we are in the progress of each?

Dave Burke

I think we’re learning that as we look at these emerging properties of biology by training large-scale foundation models on nucleic acids and these virtual-cell models, which we’ll talk more about today. We have this debate often internally. I have an engineering and computer background, so the way I think about it is this: if you think about the cell, the DNA lives in the ROM—the read-only memory. It’s coding for the cell. The RNA lives in the RAM, so it’s like the working memory.

The RNA is constantly changing its expression level. It’s almost like one of those 1980s graphic equalizers, where you have 20,000 bars for each gene, and it’s constantly adjusting its expression level depending on what the cell is experiencing—whether that’s the environment, stress, aging, disease state, or healthy state.

What we’re trying to do with this data, as a field, is create these virtual-cell models. In a way, that’s inferring a notional CPU for the cell: how does the cell respond to an input? That input could be an edited gene or the application of a drug. Then how does that reflect in the transcriptomic profile? That CPU is an analogy to the AI model that you want to build.

Once you have an AI model, what’s really interesting is that you can start posing the inverse question. Given a cell in a certain disease state that’s exhibiting a certain transcriptomic profile, how do I perturb that cell—whether with a gene edit or a drug—to perturb it back into a healthy state?

I think that’s what’s really exciting about this data. It enables these models, which then enable these tools, and hopefully can accelerate drug discovery.

Nima Alidoust

One thing I’ll quickly add is that, when we think about different domains in biology and building AI models of those domains, there are parts where we are data-poor and parts where we are compute-limited.

When it comes to DNA language models, thanks to the field and decades of sequencing a ton of genomes, we’re not as data-limited. Compute—specifically, context and how long we can consume DNA, what size of inputs we can use, and all of that—is a big limitation that we’ve tried to solve.

When it comes to cell-state models, that is an area where we’re absolutely very much data-limited. Being able to profile cells at single-cell resolution is basically a new technology. It emerged over the past decade, but the explosion really happened over the past 5 or 6 years. We’re just getting to the point where we can generate that kind of data at scale.

It’s not just the scale, though. Before scBaseCamp, which is the data set being released together with Tahoe as part of the Arc Virtual Cell Atlas, the data set that the Arc folks created by collating all publicly available data, I think the number of human cells that we had collated together was on the order of 45 or 50 million. If you generate 60 million single-cell data points, scale is one thing. The question is how much information content there is in that data as well.

We built early versions of some of those virtual-cell models. We called them single-cell foundation models, or whatever name you want to use for them. What we saw is that if you reduce those 60 million cells by downsampling them by 99%—so you use just 1% of that data to train your models—the model’s performance doesn’t reduce that much.

That means the information content of the data you’re using to train those models isn’t amazing. Having data that comes from very different biological contexts is key in providing information content for the model so it can learn. That goes back to what Dave was saying: perturbational data sets allow you to create new contexts and new cell states that the model can learn from, and therefore use for different kinds of applications.

I’ll let Johnny talk later about the challenge of creating those perturbational data sets.

Sarah Guo

Before we go there, can we zoom out for a second and have you describe, in layman’s terms, what the data actually tells you and where the prior data came from—even if it was information-poor?

Johnny Yu

If you look at the data that’s been generated over the past decade, it’s basically all kinds of academic groups, like us, or people in industry generating all these little data sets. There are a ton of problems with this.

First, there are batch effects. Even if one person runs an experiment on 2 different days, the data can look different, even if it’s the same cells. When you think about trying to build the internet of biology—which is what you need to build this ChatGPT moment in terms of scale—you need big data. Machine learning isn’t going to do anything for us if we don’t have big data, or if you have a data set that’s poorly labeled and full of batch effects.

This data set is basically doubling the size of all the data that’s out there cumulatively from the past decade. It covers 50 different cancer models from different patients, so it has cells from 50 different patients and 1,200 drug treatments. It’s a really deep and rich data set that effectively has no batch effects.

We think this is not only an additional data set for machine learning; we actually think it’s the first data set that’s going to enable machine learning in this space.

Hani Goodarzi

One thing that might be worth touching on is why perturbational data is important. The key is that we’re going from correlation—which is what a lot of biological research is. It’s descriptive: you stare at things and try to see, when you poke this way, what else is changing—to causation.

Going from associative changes to causation is where genetic or chemical perturbations allow you to have a very clear before-and-after. You have a set of causal changes that can actually drive a particular cell state.

The key is being able to do this in a generalizable way. You need to be able to look across many different cell types and tissue types. An ML model would need to train on all of that diversity in order to learn a general sense of the possibilities of cell states.

Dave Burke

In a topological sense, what the model is trying to do is create a manifold in a high-dimensional latent space. To explore that manifold, the model needs to see lots of different perturbations and responses. Once you do that, you have this generalized manifold that allows the model to make predictions for data that it hasn’t seen in its sample, but that still fits the manifold.

Hani Goodarzi

To make it more tangible, almost the entirety of the data that was publicly available before this came from healthy tissue. Very little came from diseased cells. Almost all of it was observational: you take cells from a liver sample, for example, and perform single-cell RNA sequencing on them.

That has the limitation Patrick was talking about: does it capture the causality of the gene interactions you’re trying to model? The second question is whether it allows you to model how a new perturbation will impact the cells, whether it’s a genetic perturbation or a drug perturbation. That is really the focus for Tahoe in this situation: perturbational data sets.

Nima Alidoust

If you put all of the perturbational data sets in the world together, then, if you’re generous, it’s 1 or 2 million single-cell data points of publicly available data. Tahoe is 100 million. We’ve massively increased that amount.

When you couple that with the huge amount of observational data sets from different species that are already in the world—which is what the Arc team put together—it turns out to be 230 million single-cell data points. They’ve tried to reduce the variations between these data sets as much as possible so they’re consistent with each other and can be used to train machine-learning models.

That’s the significance of this data.

Hani Goodarzi

I want to make a finer point on this. If you want a model that can learn about changes going on in the heart, the brain, the liver, or the bones, you need to be able to train across all those different cell types.

If you just look at normal, healthy cells, you wouldn’t necessarily learn how the manifold in latent space changes in disease. Being able to look at many different tissue types across different cancers is one way to get at those critical disease states that both basic science and drug discovery really care about.

Sarah Guo

How should we think about 100 million data points or 230 million data points, and the scale of this release, in terms of where we are? Is that enough to be useful? What do we know about scaling laws now?

Dave Burke

The short answer is that it’s a very hard question. We won’t know until we get there. What we can draw inspiration from is large language models in human language, as well as things like DNA language models, where we do have enough data to understand scaling laws.

In those cases, around 1 trillion training tokens is where you want to get to. GPT-3 was, I think, half a trillion tokens. ESM-3 was 700 billion tokens, so close to 1 trillion. A trillion sounds like a comfortable mark to hit.

The question then becomes how you count tokens, because cells aren’t exactly sentences. But if you count genes and their expression as tokens, I think this collection gets us close to where we want to be to start asking and answering those questions.

Nima Alidoust

A cell collection for these data sets has 2,000 to 5,000 genes, and each gene and its expression are basically a token in what we’re doing. So 100 million single-cell data points is equivalent to around 200 to 300 billion tokens.

There’s a finer point, which is how many of those tokens are actually informative to the model. But I think you understand the gist of it.

Sarah Guo

How do you decide where in the genetic landscape to start? How do you choose perturbations?

Hani Goodarzi

You want to match your perturbation toolkit—the arrows you’re throwing at the biology—against the biology you have. For cancer, that means going after cancer-relevant genes: genes that impact the growth of cells, genes that impact DNA regulation, and drugs that target key pan-cancer pathways.

For cancer-relevant questions, this data set is heavily based around chemical perturbations and cancer. But these pathways are so conserved and fundamental that they broadly apply to neuroscience and immune-cell development in general.

I think it’s really the foundation model that’s going to be able to take this data, ingest it, build a model, and then understand how to translate that data to a completely different context.

Johnny Yu

This is one of the really special things we have at Vivo: the Mosaic platform. It allows us to take cells from many different patients and, in cancer, that means all kinds of cancer—lung cancer, pancreatic cancer, and so on—from different patients with their own special genetics.

We pool them together into a single mosaic tumor, which we can then reproducibly screen against hundreds or thousands of drugs. This key innovation allows us to test tens or hundreds of cancer models at a time instead of testing one cancer model at a time. It makes this a really scalable data-generation platform.

That’s what we used to generate Tahoe-100M.

Nima Alidoust

When we think about how we build these pools in terms of information content, we want to maximize it by covering a lot of cancer patients. For this data set, we covered the biggest cancer types by how frequently they occur annually.

As we continue to grow the data set, we want to bring in rare diseases and perhaps more coverage of different parts of the cancer space. We also want to let the machine inform us by helping us fill in the gaps in the foundation models.

Another direction is chemical space. Frankly, when we generate data, we generate 50 times more perturbational data sets in 5 weeks than are publicly available from the past 10 years. You don’t have to prioritize as much, and that’s the beauty of it, in my opinion.

You can go large on chemical space and on patient-sample space. That way, you don’t have to come up with a hypothesis a priori about what you have to feed the models. You can just generate as much as you want and be more unbiased.

Hani Goodarzi

As scale increases, you can generate hypothesis-free, unbiased data. That’s really the beauty here.

Sarah Guo

Can the data surprise you?

Nima Alidoust

Exactly. This is one of the things I like to talk about. I hope the people we have here are representatives of the new generation of biologists.

One thing that has been slowing progress in biology is that we’ve always been super hypothesis-driven. There are reasons for that: a lot of these experiments are expensive, and they take a lot of time and resources. But now sequencing costs have gone down, the per-sample cost of single-cell sequencing has gone down, and compute costs have gone down.

I think it’s time to change that mentality in biology as well—to be more courageous and more freewheeling in terms of data generation and the kinds of samples you put together.

Sarah Guo

I want to talk about being more ambitious in biology and the open-sourcing of this in a second, but first I think we should zoom out and talk about, in layman’s terms, what the platform does. You can correct me if any of this is wrong.

You have these tumors that are a mosaic of cells from different patients, representing a huge amount of patient genetic variation. Each mouse can then be treated with different drugs, and the signal you extract afterward is the interaction of those drugs with each of these different patient types.

Johnny Yu

That’s right.

Sarah Guo

Nobody else thinks this is crazy?

Patrick Hsu

No, it’s not crazy, because it’s happening every day in our lives. But it’s really science fiction, honestly.

Sarah Guo

I’m just trying to boil it down to a very simple, non-biologist understanding. When you say it’s a platform with this super-tumor where you can pull all of this data out, it’s wild to think about how efficient that is compared with observing one patient type at a time.

Patrick Hsu

I think this is a really interesting point. If you map the number of tokens per experiment across the last 50 years of biomedical research, it will look like the hockey stick that all investors and founders know and love: just going up into the right.

The way we think about doing science is changing based on this. There’s a roiling discussion today about hypothesis-driven versus hypothesis-free research, and whether we should be doing mechanism-focused work versus large-scale profiling.

Honestly, I think this stuff is going to wash out with scale.

Nima Alidoust

Exactly. You don’t have to choose between those two.

Patrick Hsu

That’s my hot take about this era of machine learning and biology. The vast majority of mechanistic data that’s been generated today was really made to ask very specific, well-scoped questions.

Going from a few tokens per experiment to many more tokens per experiment is just going to be the way to do it.

Nima Alidoust

I can say it another way. In biology, what we’ve done is treat humans as the foundation models that ingest information and come up with hypotheses. But now we want to go beyond that, because humans come with their own intuitions and biases.

At UCSF, for example, we often say that we use some of our medicinal chemists, like Kevan Shokat, as the last layer of a neural network. They’ve built this intuition about whether a chemical that was generated by an AI model actually looks like something real.

They can’t even verbalize why they think it might be good.

Patrick Hsu

People criticize these models for hallucinating, but if you think about it, the process of scientific research involves hallucination. That’s what creativity is.

Sarah Guo

You’re all adherents to that lesson in this field as well: that intuition being baked into the models or the process is not necessarily the right thing. We just need to scale the data.

Dave Burke

At least, we hope you don’t have to make that choice. We’re seeing evidence of scaling laws in biology across proteins. That’s been shown in protein language models and across DNA, which is what we’ve shown in our Evo series of models at Arc.

We’re also seeing inference-time scaling laws. In our most recent study, there are early signs of promise. We’ll need good benchmarks, and we’ll have to look at this across different data types over time.

Nima Alidoust

The funny thing for me is that, if you’ve been in this field long enough, you know that every time you take a success from field A and try to translate it to field B, a lot of people—including people in our own organizations—come up with a list of 100 reasons why the learnings from field A aren’t applicable to field B.

But then you’re surprised every time. When you try to do the same thing from B to C, the same kinds of lists start emerging.

Something that’s underappreciated is that the same models that learned human language are learning the language of structural biology. With the Evo work, they’re learning the language of DNA. This is incredible. I don’t think it’s trivial.

If you’ve been in the field long enough, you know there were a lot of people saying that protein language models would never work and that you needed domain-specific models to model these phenomena.

The ethos we have to bring here is that, when Hani says we should use the learnings from those models to translate them here, that’s exactly what we should be doing. We should think about what worked and at least try it in these new domains. Maybe it works, maybe it doesn’t, but if you don’t try, you’ll never know.

This is the domain we’re talking about, and it’s the domain Vivo is excited about. The virtual-cell part of Arc is excited about it as well: the language of systems biology. The first thing you should do is try the things that worked in other domains in this domain.

Sarah Guo

That’s music to my ears. This is one of the only things we have really strong conviction about at the fund we invest out of: that many of these techniques work and scale in domains where people aren’t sure yet. We wouldn’t have expertise in the traditional types of discovery and company-building, but they seem to apply very generally.

I think this is a great segue to a question about why you’re open-sourcing the data.

Nima Alidoust

We generated the data at Vivo, and Vivo is a private, venture-backed company—a startup. When the idea of Tahoe originally came up and Johnny told me, “Nima, there’s this opportunity. We can generate 100 million single-cell data points,” I said, “Can we?”

He said, “Yes, we can.”

I said, “Let’s go and do it.”

Within hours of that conversation—we’re transparent about this—Johnny, Hani, and I are co-founders of Tahoe Therapeutics, and we said, “Let’s do it, and let’s open-source it.”

Why do we want to do that? First, we want to put a new stake in the ground. We want to show that there’s a new game in town and that it’s possible to up our game as a community and as a field.

We want to show people that they should move on from generating 100,000 or 1 million observational single-cell data points. They should up their game and go to a much more massive scale. That’s the first reason.

Second, we wanted to support the DNA of our company, which is to be very small: a small team of superstars rather than hiring a large number of people. Paradoxically, open-sourcing allows us to do that.

We talked to Dave about Tahoe around the night before the new year, sometime between Christmas and New Year’s, and the entire Arc team got really excited about it. If there hadn’t been an open-source aspect to it, it wouldn’t have been as exciting.

The whole community is getting excited about playing with this data and telling us what’s good about it and what’s not good about it. That allows a team of 3 or 4 people in-house to keep it that way and bring in the entire community of like-minded people who have the same vision of building virtual cells to help us in this quest.

For us, the idea was to remove the main bottleneck in doing that. Everyone has been saying that the bottleneck is data.

Dave Burke

The serendipity of this was that Arc is all about mission-driven science and pushing science forward. We were conceiving of creating and launching this week what we’re calling the Arc Virtual Cell Atlas.

The idea was: can we find high-quality, curated data sets and put them out into the world to accelerate virtual-cell modeling? Then we started chatting, and it was, “You’ve got what?”

What we’re assembling this week is this new atlas. The star of the show, in some ways, is the Vivo Tahoe 100 data set. We’re also augmenting it with observational data.

We created something called scBaseCamp. You can almost think of it like a Google crawler and index. We built an agent that goes onto the internet and mines public single-cell RNA data, then curates it in a very uniform way. The result is a very nice observational data set of about 230 million cells.

You add that to the 100 million cells from Tahoe 100, and you now have 330 million cells. This is a really exciting resource for scientists around the world who are interested in modeling at the cell level.

It’s very complementary to have an observational data set that you can potentially pretrain a model on, and then the perturbational data set from Tahoe 100 allows you to bring in those dynamics and make the model richer and more predictive.

We’re super excited about AI agents for science at Arc and across the community. The capabilities are still very early today, but we wanted to show an example of how an agent can do something really useful.

It’s very clear now that basically all dry-lab workflows are going to get automated with agents or copilots. This would ordinarily be the type of thing that a team of computational biologists would be slaving over.

Our core insight was that the Sequence Read Archive is the largest repository of biological data from next-generation sequencing. If you get an NIH grant, for example, you post all of that data online. If you publish in a journal, you put all of the data online as part of the publication.

But it’s extremely fragmented, poorly annotated, and sprawling. There are no requirements for the submission data to be uniform.

We built this agent to crawl all of that data, collect it, organize it, and process it. In doing so, it isolates and removes many of the batch effects and data biases of previous methods.

Hani Goodarzi

These are foundational data sets for the entire field. People work with them, interpret them, and write papers on top of all this data.

Nima Alidoust

Exactly. These data sets have been generated over time, going back a decade. Tools have changed, versions of tools have changed, and genome builds have changed. By simply taking processed data sets and collating them, you’re infecting and contaminating the data with analytical effects and batch effects.

Our idea was to at least remove that. There are a lot of technical and experimental batch effects, and over that span of time, the chemistries of reagents have changed. But at least we can do our part and remove the analytical component.

We were surprised by the extent to which those effects were observable in the data, and by how helpful it was to remove them.

Dave Burke

One thing I would add is that we’re really excited about the leverage of this approach. Sometimes I ask Hani and Johnny, “What does drug A do to cell line X?” There’s this phrase that biologists use: “In my hands, it does this and this.”

A computer scientist wouldn’t say, “In my hands, this model worked.” They might say, “In my computer, in my environment, it worked.”

The genius of what Johnny has built is that this was done by very few hands. Automation is going to scale you to a certain level, but you haven’t even done much automation.

The beauty of Tahoe is that a few people, with a few hands, did exactly consistent work across 60,000 experiments. It’s 100 million single-cell data points, but it’s actually 60,000 drug–cell interactions, or drug–cell-line interactions. Having that done by 4 people reduces the data-set infection that Hani was talking about.

Sarah Guo

This is the first time in history that there’s an opportunity for scientists and entrepreneurs to work on this data set and create these virtual-cell models. How do you tell the quality of one of these models?

Dave Burke

The core idea is its predictive ability. You take a cell and perturb it. You can do that genetically by suppressing or upregulating genes, or you can apply drugs and look at the response.

The measure of the model is how well it predicts what we call the differentially expressed genes, or DEGs. The reality is that today the best models are very poor at this. Their predictive ability for the DEGs is on the order of 10%.

Sarah Guo

Is there an accepted benchmark for this today?

Dave Burke

No, but that’s something else the industry would benefit from. If you think about where we want to go, one of our conjectures is that one reason the models aren’t doing well isn’t simply model structure.

We have a lot of rich structures that we understand in the machine-learning space. The issue is data quality. The hope is that, with this new Arc Virtual Cell Atlas and Tahoe-100M, we finally have a starting point where we can build rich models and get high predictive value from these virtual-cell models.

That’s why this is such an exciting moment in time.

Patrick Hsu

It might be worth speaking plainly about why we even care about virtual-cell models. We have real cells, so why not just do experiments on those?

Biology is very slow. Many of us have tried to pick up pipettes, move clear liquids from one tube to another, grow cells, and make animals. Biology happens in real time.

In the last year of my PhD, my adviser tried to convince me to start an aging project, which would have involved aging animals for 2 years. That’s one experimental round. As you can imagine, I declined. I said, “May I please graduate, sir?”

That’s actually what happens. You’re constrained by biological time, which is completely crazy to me coming from an engineering background. It’s also important to many fields, like neurodegeneration and anything else that takes time to progress.

The idea of massively parallelized, in-silico simulations sounds great, but they need to be accurate. If they’re 10% accurate, you’re just simulating noise.

How do we go from a discipline that primarily respects experiments today to something more like physics, where theory drives a lot of progress? I think these virtual-cell models are a core wedge in making that happen.

Sarah Guo

Can you make that more concrete? If these virtual-cell models work—and we don’t even know how to measure them yet because they don’t exist in a productive way today—what would scientists, the biotech field, or patients expect to gain?

Nima Alidoust

From a drug-discovery perspective, what we’re focused on at Vivo is predicting how a new chemical entity interacts with cells from different patients or patient models. That really is the core of it.

Patrick was talking about in-silico simulation. Can I predict in a computer whether this new chemical structure—drugs are chemical structures, by the way, in case that surprises you—will take a diseased cell, such as a cancer cell, from a diseased state to a healthy state? In the case of cancer, can it kill the cell?

If I can predict that, then my ability to design new chemicals that do that effectively—killing the cancer cell without killing healthy cells—improves massively. That’s what we want to do, and that’s the kind of data we’re generating to train those models.

Johnny Yu

A big part of our future vision and road map is that we think there will be a moment when a drug is generated from a virtual-cell model. The drug will cause a diseased cell to become a healthy cell again.

I think that’s the goal, and it will reshape how we do drug discovery.

Nima Alidoust

There are 2 dimensions of generalizability to think about. One is the cell dimension, and the other is the chemical dimension.

On the cell side, every disease is unique. There are similarities—there are truncal cancer mutations and other things that drive the disease—but there are also very significant individual variations.

You can observe cells from patients, but you can’t do that for every tumor that arises. That’s what these folks do with Mosaic. The idea is that, using a virtual-cell model, you can take those learnings and apply them to all of these new observations that you can make in patients. That’s one dimension.

The other dimension is chemicals. In silico, you have libraries with tens of millions of compounds, and biologics are effectively infinite if you really put your mind to it. Most of these molecules have never existed and will never exist because there’s no use for them.

A model that can traverse that massive space of chemistry and find which parts you need to pay attention to, synthesize, and test would be massively enabling.

Everyone else has libraries that are well behaved—perhaps a couple hundred thousand compounds—and uses fragments and tries to put them together. The process of designing drugs is this slow screening process. This would allow us to leapfrog that entire pipeline.

Patrick Hsu

Ninety percent of drugs fail in clinical trials, so we’re pretty bad at making drugs. I think that implies 2 things.

First, our chemical matter may not be very good in terms of potency, ability to bind the target, toxicity, pH, kinetic profiles, and all of those other properties.

The other possibility is that we’re drugging the wrong target. The idea of virtual-cell models is that they’ll allow you to significantly cut down the search space for the right target. Then you can focus your time on making the right chemical matter or drug composition to make the right kinds of changes in the right kinds of cells.

That’s why mechanism and drug discovery are so tightly interwoven. These models need to help accelerate both.

Nima Alidoust

This is the core reason we need virtual cells in addition to protein language models. Protein language models speak the language of structural biology: what a protein structure looks like, how it folds, how it interacts with a ligand, how a small molecule binds to a protein, or how an antibody binds to another protein.

These are binding questions. You’re trying to see whether one chemical binds to another chemical. But biology is more complex, and there is a context to the protein target we’re trying to hit.

It’s part of a cell. The cell is part of a tumor in the case of cancer. The tumor is part of a broader biological system. Virtual cells will allow us to go beyond the language of structural biology and venture into the language of systems biology.

We can understand how a drug interacts with the broader biological system rather than simply with one target whose binding we’ve cracked.

Sarah Guo

I have a higher-level systems question. At the single-cell level, what about multicellular aggregates, organelles, and organs? Is all of that going to be possible in the future?

Patrick Hsu

I think the first question for virtual cells—or any modeling—is what the right level of abstraction is. Our belief is that the right level of abstraction is the transcriptomic level, because you have these very complex gene pathways.

Whenever a cell is changing in response to its environment, that change will be reflected in the transcriptome. That’s the first question: even within a cell, what’s the right abstraction?

A cell is an exquisite piece of machinery. You could make an arbitrarily complex model of it, but we believe the genetic level is the right level to model.

Going beyond that, you can create very advanced models. You see people doing this with spheroids and organoids. You take mixtures of cells and run them together, trying to simulate cardiac tissue or brain tissue.

What’s really interesting is that, if you have an organoid with 20,000 cells, you can still apply these techniques: take these drug perturbations or genetic perturbations, apply them to the cells, and look at the responses.

You’re going beyond a single cell, but you’re also capturing the intercellular dynamics in the models.

Hani Goodarzi

I’ll make one small comment. It is a single cell that we’re modeling, but that context dependency also captures many of the effects that arise from the environment.

The models we have are actually ex vivo models in this specific experiment for Tahoe, but we also have in vivo models. We have humanized mice that capture some of the immune system of the mouse.

In a way, yes, you’re simulating and building an in-silico model of a cell. But if a model is any good, it can simulate cells in different biological contexts: in the presence of this kind of immune environment, in the presence of one kind of tumor versus another kind of tumor, or in the presence of one mutation versus another mutation.

We call it single-cell modeling, but the whole idea of having so many single-cell data points is that you have those cells in different contexts.

Sarah Guo

That seems like a really important nuance.

Hani Goodarzi

The information about the environment is filtered through the cell. If you’re observing the cell with enough resolution, you can predict it. You can also add spatial data.

Sarah Guo

Definitely. I have a few hot-take questions to end with. Nima, I’ll start with you, because we were having a passionate discussion about why it was important to you that Vivo be a platform company rather than a single-hypothesis company, like 99.9% of biotechs out there. What’s the difference?

Nima Alidoust

The difference is the kind of team you build and the ambition you have.

A single-hypothesis company is based on the idea that the human being—the foundation model Hani was talking about—comes up with hypotheses, and then we test those hypotheses in different experiments.

Companies built around a hypothesis are heavily incentivized to make that hypothesis work. What you see in biotech is that a company will take a drug to the clinic after testing it on 3 different patient samples.

If you’re a platform company, you’re trying to have enough hypotheses and a hypothesis-free way of generating new hypotheses. That means you aren’t wedded to a single hypothesis, and it allows you to be much more scientific in your quest for new drugs or new targets to treat disease.

That’s why we decided to make Vivo a platform company. It allows us to be much more rigorous about what we decide to take to the clinic.

Sarah Guo

There’s been a lot of news recently about the rise of Chinese biotechs. For the core members of the research community here, is that a threat? How do you think about it?

Patrick Hsu

Their cost basis is definitely more competitive. A lot of the discussion in biotech and pharma is how they’re able to do things at this pace and cost, and why their data packages look so good.

They have safety data, toxicology, and all these IND-enabling studies. It’s really competitive. People were surprised by the efficiency of the pipelining and the ability to manufacture all these different antibodies, primarily.

I think that’s great for the industry. Everybody—including patients, investors, and the biotech companies themselves—wants lower cost bases. We want the ability to make molecules that work faster.

All of these things will compete in the system to reduce the currently high cost basis of doing these things here in the United States.

One of the core challenges right now is that we have a wide array of services, CROs, and contract research collaborators that you can try to chain together.

The virtual biotech was previously a concept that was very much in fashion. In reality, when you try to do this, even though it looks good on paper, it’s incredibly slow.

Then people tried the other way: let’s fully vertically integrate and own everything. But that was incredibly expensive. The answer is probably more of a Goldilocks solution in the middle.

We need really competent vendors and CROs that understand the drug-discovery and development process. Then individual companies need to be able to run in a capital-efficient and lean way.

I think the industry is trying to reshape itself around these changes right now to figure out the right way to build startups and the right way to build drugs.

Dave Burke

I totally agree. It’s an important moment. One thing I haven’t seen is an acknowledgment of it. It just hit us in the face.

The United States is the innovation hub, but I think we need to be more intentional about that in biotech. You see innovation in tech, and you see that as the mantra. Innovation in biotech has actually been viewed as one of the things Chinese CROs and companies are good at.

What we’re finding out is that’s not actually innovation. The kinds of things we’re working on—putting big data and AI into the first layer of how we do biology—are what innovation should look like in our space.

If we don’t push that forward as a community, we’re not going to have that innovation in the industry.

Nima Alidoust

This actually slapped us in the face. It caught us by surprise, but one of the first conversations Johnny and I had 3 years ago, when we were thinking about starting Vivo, was about this happening in China and the whole thesis around the commoditization of things we think are massively important, such as molecular design.

There are 2 ways to respond. One is regulatory capture: lobby the government and put limits on how much we can interact with Chinese companies.

The other is to make it part of our ecosystem and change our thinking about business models and how we build teams. To Patrick’s point, do we build a fully integrated team with $100 million in the bank, or a small 14-person team like we have at Vivo?

These are the kinds of things we should be thinking about.

I want to make this into a bigger statement: I think it’s morning in biology. There’s a different game we should be playing here. If you want to stick to the same old-school way of doing things, it isn’t going to work.

The old-school way involves a lot of planning. I was texting with Dave about this a couple of days ago. If I had a penny for every time some massive organization announced an extraordinarily impressive thing and said, “We’re going to give it to you in 3 to 5 years,” I’d be extremely rich right now.

That’s the ethos in biology. You announce some massive thing and say you’re going to do it in 3 to 5 years. But now we have the tools. It’s time to build, and it’s time to do it right now.

That’s how Vivo gets created in a matter of months. That’s how Evo 2 gets created in a matter of months, from the first Evo paper to what happened. That’s how Tahoe gets created.

The second piece is small, super-focused teams of superstars. Massive, vertically integrated organizations aren’t just capital-intensive; they’re inefficient. They move slowly and get bogged down in bureaucracy.

The third piece is associated with this naysaying. In everything you want to do in biology, there are strong biologists who will tell you why it isn’t going to work.

That has to change. We have to think differently, try things out, and use the tools we now have.

Dave Burke

When I talk to pharma companies, they’ll say, “AI and drug discovery are very interesting, but I don’t spend that much of my top-line budget on drug discovery. Most of it is wrapped up in clinical development.”

A lot of them are much more excited about things like natural-language workflows to summarize clinical-trial documents, which are massive regulatory filings, and make it easier to write and read them. They’re also excited about AI-assisted cohort analysis and reducing costs in that part of the cycle.

What they’re going to see as these models get better is that virtual-cell models can help you find the right target, so you can point the cannon in the right direction. You can measure twice and cut once.

The cost basis for the industry should go down, and accuracy should go up.

Sarah Guo

I’m glad both of you brought up the naysayers, because if you weren’t going to, I was going to. I’ve been pitched AI for biotech companies for at least a decade. We haven’t seen a lot of results, and there’s also the natural life cycle of bringing treatments to market. You generally need 11 years or more.

If you were going to leave a broader audience with a single claim about why this is true—there have been different approaches over the last decade, such as computer vision and consumer-scale sequencing data—why should this work now? When should we actually begin to see treatments from these machine-learning approaches?

Dave Burke

I’d go back to analogies. In the machine-learning space, we had artificial neural networks for a long time. People got wrapped up in whether a perceptron could model an XOR gate or whatever it was—around 1990.

Things bounced around for a while. Then we had increased compute, increased data, and more sophisticated models, and we hit nonlinear inflection points.

I mentioned the ImageNet moment in 2009. It drove the development of convolutional neural networks. AlexNet was the model that really showed the way.

Before that, people thought only humans could recognize images at high quality and that computers would never do it. Now we know computers can do it better than humans.

I think it’s the same thing in AI and biology. When I look at single-cell sequencing, the capability is mind-blowing if you’re not a biologist. The idea that, at single-cell resolution, we can look at how expression changes over time is incredible.

You take that ability to generate lots of data, add much more sophisticated models and model training, and suddenly things are happening.

If you look at the Evo 2 model, we trained it on 9.3 trillion nucleotides. We didn’t tell it anything about DNA. We just gave it a lot of DNA from across the planet—every piece of DNA we could get hold of.

What did the model learn? It started learning all sorts of things. It knows where ribosome-binding sites are. It knows about codon degeneracy. One of the things we showed is that it can predict the pathogenicity of BRCA1 variants, which are known to drive breast cancer, with an area under the ROC curve of around 94%.

We never taught it anything. It learned this stuff zero-shot.

I think we’re at that point of inflection now. We’re all in agreement that we’re at a point in time where we’re going to see that inflection. The difference between where we were yesterday and where we are starting this week is going to be the data.

Sarah Guo

Are we somewhere between GPT-1 and GPT-4 in biology?

Patrick Hsu

I’d say we’re closer to GPT-2.

Dave Burke

We’re developing GPT-2, but we don’t have enough data. We need more data.

Hani Goodarzi

If you go a little deeper and talk about different domains, I think protein models are past GPT-3. When it comes to single-cell models and virtual-cell models, we’re closer to GPT-1 than GPT-2 right now.

Sarah Guo

That’s still a pretty exciting timeline if you take the progress in other domains and apply it here. But the difficulty is exactly what you said: with GPT-4, you immediately knew what you had.

If we hit GPT-4 for cell-state models, for example, it will take some time to prove that point. Drug discovery is also subject to the law of small numbers. A platform that takes your success rate from 10% to 30% is amazing, but it’s still 30%, and you need to get lucky.

You still have a drug-development cycle on the order of 10 years, so you have to wait for the model to prove itself. The success rate can slowly go up in a 10-year rolling window.

Nima Alidoust

If we’re all optimists here, I’d say we’re going to treat it as a system. If this was a terribly debilitating bottleneck at the beginning, then hopefully this is a breakthrough.

Sarah Guo

I think that’s a great note to end on. Thank you all for joining me.