[BidClub_]
Latent Space · · 85 min

🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik

Ron AlfaDaniel Bear

YouTube
TL;DR
  • Noetik’s core thesis is that oncology’s 90%-95% clinical failure rate is primarily a patient-selection problem, not a molecule-making problem. Ron Alfa argues that trials often show no placebo effect in cancer, so responders indicate that relevant biology is active; conventional development cannot identify the right subgroup. Noetik wants to discover therapeutically meaningful patient subtypes directly from human tumors, then use them both for reverse-translated target discovery and trial design.

  • The company’s claimed moat is a purpose-built, intentionally controlled human-tumor dataset rather than public biological data assembled after the fact. It has generated spatially resolved transcriptomics for more than 100 million cells, paired with H&E histology, protein imaging, and genotype data—“at least an order of magnitude larger” than comparable datasets Noetik has seen. Training on only 40% or 10% materially worsened its models, particularly when generalizing into unseen cancer types.

  • Noetik trains on expensive multimodal assays but can run clinical inference from the standard H&E image already collected for almost every oncology patient. Travis McKie calls H&E the “lingua franca of pathology”: the model can separate responders from nonresponders, predict locally expressed genes, and expose biology beyond single-mutation or single-protein biomarkers. That creates a plausible path from archived trials to prospective diagnostics using the same inexpensive input across many drugs.

  • Its “virtual cell” is deliberately practical and top-down, not an attempt to simulate every biochemical reaction inside a cell. OctoVC asks what T cells, macrophages, tumors, or gene expression would do in a particular patient context; PerturbMap tests model predictions through roughly 100 barcoded genetic perturbations within mouse tumors. The ambition is to predict which drug fits which patient without first constructing a mechanistically perfect cell.

  • The modeling work is moving from masked reconstruction toward autoregressive spatial prediction, with tissue context emerging as the scaling variable. OctoVC uses extreme masking—roughly 99% in the hosts’ recollection—to prevent the model from merely completing local edges. TARiO’s next-token objective showed larger models outperforming mainly at longer context lengths, suggesting that more surrounding tissue, not parameter count alone, unlocks complex patient-level biology.

  • The $50 million GSK agreement is the clearest commercial validation, but its structure matters: that figure includes upfront payment and milestones, alongside a separate annual model-license fee. GSK receives OctoVC models already trained on lung and colon cancer and can fine-tune them on its own translational data. Akash Tiwari’s key framing is that this resembles a meaningful biopharma business-development deal, except “the substrate is actually a model,” not a molecule.

  • The investment risk is inseparable from the moat: Noetik spent roughly four years building infrastructure and closer to $10 million than $50 million before knowing the approach would work. It lacked enough data to train a model for at least 18 months, individual transcriptomics runs took two weeks for two slides, and the company has not yet disclosed the trial-reanalysis results it says are coming. Dan Barr says several hundred patients across each major and selected minor cancer indication might generalize across oncology, but explicitly hedges that all disease biology could require another order of magnitude of data.

Digest · the substance, structured for research

1. Patient selection—not pharmacology—is the proposed failure point

  • Alfa opens with Noetik’s contrarian premise: “90%, 95% of cancer drugs fail in the clinic,” even though industry is better than ever at pharmacology, target selection, and making molecules. He locates the dominant failure downstream, in identifying which patients carry the biology a drug can exploit.

  • His evidence is the responder hidden inside the failed aggregate. Alfa says that cancer trials often have no placebo effect; when a patient responds, that indicates real activity. The trial may have failed because that biology was diluted across an indiscriminately enrolled population.

  • The platform therefore works in both directions. It can begin with human tumors and reverse-translate toward new targets, or analyze phase two and phase three biopsies to identify the biology predicting response and redesign the next study. Noetik says it is doing substantial work on the latter, although those results remain undisclosed.

  • Alfa pushes the subtype claim further: “Nobody actually knows what the subtypes are.” A category long treated as one lung-cancer subtype might contain three functionally distinct diseases, and those latent divisions may matter more therapeutically than the classifications pathologists have used for more than a century.

2. “Frankensteinian” preclinical models leave clinical teams guessing

  • Bear’s critique begins with immortalized cell lines that have persisted for 40 or 50 years, often carrying abnormal chromosome counts and expression programs unlike recognizable human cells. Researchers still label them colon or lung cancer, but he calls them “Frankensteinian cells” whose experimental response often fails to map back to patients.

  • Moving those cells into animals does not close the gap: oncology commonly implants them under the skin “in weird places,” tests hundreds through a contract research organization, and infers an indication from which nominal colon or ovarian lines respond. Even lines derived from colon cancer may lack mutations characteristic of human colon tumors.

  • The resulting clinical chain is punishing. With no preclinical guidance about patients, a team may enroll all eligible tumors into an open-label study of roughly 50 people while simultaneously learning dose, safety margin, and efficacy signal. If lung cancer alone hypothetically contains 10 relevant subtypes, few responders are statistically unsurprising—and the molecule may nevertheless be canceled.

  • The hosts’ restatement captures the proposed fix: stop assuming today’s indication labels are homogeneous, and find the subpopulation that responds. Bear adds that existing biomarkers—one mutation, one stained protein, or one gene signature—are “biased towards simplicity” and usually correlate only weakly with clinical success.

3. Biological training data must be designed before it can be scaled

  • Tony Bui rejects the idea that a useful biology corpus can simply be scraped together. The Protein Data Bank was intentionally accumulated over decades; Shawn Wang compares ImageNet’s carefully curated, labeled 1.2 million-image scale. The lesson is that “you really need to be intentional about the data that you generate” and anticipate the models it must support.

  • Bui treats scale as “necessary, if not sufficient.” Language displays extraordinary scaling partly because of its corpus size, but thousands of hours of video have not automatically produced equivalent behavior. Biology adds tens of thousands of genes and proteins, their spatial organization, healthy tissue, other diseases, and even other species.

  • Shawn Wang offers the counter-thesis: because biology arises from underlying physical processes, perhaps enough of the space becomes in-distribution earlier than language does. Bui’s answer stays hedged—his “hunch is that biology is pretty complex,” Noetik remains far from covering all of it, and “I don’t know.”

  • Another useful check comes from protein folding: the discussion notes that good models may use only a small fraction of PDB, and some researchers argued that its 1990s coverage was already sufficient for a sufficiently capable algorithm. Dan Barr’s narrower estimate is that several hundred patients across major and selected minor cancers might be enough to generalize broadly across oncology.

4. Noetik captures tissue, cells, molecules, and batches together

  • Lessons from Tony Bui’s six years at Recursion shaped the experimental design. Images are information-dense, allow many patients on one slide, and reduce marginal data-generation cost relative to sequencing, where each additional run can correspond to another patient set.

  • The clinical anchor is H&E: hematoxylin and eosin create the familiar purple-pink tissue contrast used by pathologists to classify nearly every resected tumor. It captures tissue architecture but cannot reliably identify all relevant cell types, so Noetik adds multiplex immunofluorescence for B cells and other components of the tumor microenvironment.

  • Spatial transcriptomics supplies the molecular layer, detecting roughly 1,000 to 19,000 genes at resolved locations in the same cells. A probe binds each RNA species and the instrument cycles through detection for weeks; the discussion likens the output to an image with “20,000 color channels” instead of RGB’s three. Genotyping adds the underlying DNA alterations.

  • Batch control is built into the physical samples. Noetik samples each tumor dozens of times, randomizes hundreds of patient specimens across arrays, and represents every patient on multiple slides and processing runs. That lets researchers ask whether a patient embedding reflects immunotherapy response or merely “staining batches.”

5. The virtual cell is a drug-making heuristic, not a complete cell replica

  • Noetik separates two definitions. A comprehensive virtual cell would simulate millions of intracellular chemical reactions after any external signal; the discussion considers that intellectually interesting but says today’s data modalities cannot solve it. Noetik instead wants “some heuristic that’s useful for making drugs.”

  • Most current virtual-cell work, as described in the discussion, predicts transcriptomic changes after a small-molecule or CRISPR perturbation in cultured cells. The objection is translational: if the endpoint is what happens in a patient, “modeling data that comes from a patient” is more likely to succeed than modeling an in-vitro abstraction.

  • Noetik’s model therefore learns genes, proteins, cells, tissue, and patient context through self-supervision. It intentionally has not leaned heavily on electronic health records: the team does not want the representation constrained by what a doctor, operating with the knowledge of that moment, happened to record.

6. One H&E image can support discovery, trial rescue, and diagnostics

  • Before a molecule reaches patients, Noetik can simulate its target across lung, colon, ovarian, or broader oncology cohorts. The result might be a negative indication choice: a target initially intended for lung cancer could appear biologically unimportant there but relevant in ovarian cancer.

  • OctoVC can ask what a T cell would express inside a particular tumor microenvironment, or what happens if a target gene or protein is removed. Useful outputs include increased immune function, reduced tumor growth, or another response believed to correlate with clinical success.

  • The cleanest retrospective use case starts with a treated cohort: if responders all occupy one self-supervised patient cluster and not the other nine, that yields a direct enrollment hypothesis. Although training uses the full multimodal stack, inference needs only H&E—including a digitized image from a trial conducted years earlier.

  • The model can also predict where genes are expressed from that H&E. Seeing a drug’s protein target enriched in responder tissue provides an interpretability check, while the surrounding multigene pattern explains why one-protein biomarkers miss response. Noetik is applying the same input across collaborations, including one announced with Agena, with an eventual diagnostic as the natural endpoint.

7. Scale, architecture, and interpretability compound the data moat

  • Academic paired datasets may contain only a hundred or a few hundred patients, with spatial transcriptomics often below even that. Noetik reports more than 100 million spatially resolved cells, all paired with H&E and protein imaging—“at least an order of magnitude larger” than alternatives it has examined.

  • Its internal ablations support the scale claim: reducing training data to 40% or 10% made models “a lot worse,” especially when generalizing into cancers excluded from training. A sudden early doubling of the dataset similarly produced an immediate performance jump.

  • The team does not present data volume as sufficient. Noetik builds custom multimodal architectures and self-supervised objectives aimed at patient differences, gene or protein counterfactuals, and readable biological outputs. They call these “world models” because the task is to predict what follows from an action, not merely classify an image.

8. PerturbMap reconnects human predictions to causal animal experiments

  • The discussion addresses the tension in Noetik’s human-first approach after the company acknowledges that it still uses mice and injected cells. Human data should train the primary models, but the FDA may still want evidence that a new mechanism works in an animal system, leaving developers to “back into this system” unless they deliberately build a bridge.

  • PerturbMap creates roughly 100 distinct CRISPR knockouts, each carrying a combinatorial protein barcode, and injects them together into mouse lungs. The result can be hundreds of tumors whose identity and spatial biology remain readable—not one cell line implanted beneath the skin as a stand-in for human diversity.

  • Noetik can map genes causally to immune-cold or immune-hot phenotypes, then layer pharmacology on top; the team describes panels such as 50 knockouts across 50 drugs. Human tumors without immune infiltration can be recreated genetically in the mouse, where they likewise lack immune cells and fail to respond to immunotherapy.

  • The more surprising step is “in silico humanizing the mouse”: a model trained on human H&E and spatial transcriptomics is run directly on mouse H&E and emits human-gene predictions. Known antigen-presentation knockouts correctly appear cold, and several genes from one signaling pathway produce similar inferred phenotypes. The team acknowledges that genuinely novel regimes remain uncertain.

9. Autoregression scales only when the model sees enough tissue

  • OctoVC used masked autoencoding: divide each modality into tokens, hide most of them, and reconstruct gene expression, protein patches, or histology patches from what remains. The host recalls masking roughly 99%, and the response confirms that this is consistent with their approach.

  • Extreme masking is purposeful. At only 10%, a model can succeed through “boring behaviors” such as extending a nearby edge; hiding much more forces it to learn correlations between proteins and the holistic structure of tissue rather than local continuity.

  • TARiO changes the objective to autoregressive next-token prediction, effectively a structured form of masking that mirrors the loss behind LLM scaling. It was not Noetik’s first attempt, but it produced clearer gains from larger models and longer contexts on spatial-transcriptomic data.

  • The subtle result is that bigger models helped mainly when context was longer—meaning they could see more tissue at once. Low-expression but predictive genes might partly explain it, but controlled comparisons of the same context over smaller versus larger physical regions also favored larger regions.

10. GSK licenses a model as the asset, not a molecule

  • Noetik’s announced GSK agreement licenses OctoVC models trained on lung and colon cancer. The $50 million headline includes upfront payment and milestones, while an annual model-license fee sits separately; Tiwari also describes significant multimillion-dollar upfront and near-term economics.

  • GSK can use the models for simulation and therapeutic discovery, then fine-tune them on its own “mountains and mountains” of translational pathology and trial data. That turns a shared foundation model into something closer to GSK’s proprietary version without requiring the pharma to unify every silo and build the base model itself.

  • Tiwari sees pharma demand moving from bespoke, single-program collaborations toward access across dozens of pipeline programs. The deal resembles a traditional industry business-development transaction, but its breakthrough is structural: “The substrate of the deal is not a molecule. The substrate is actually a model.”

11. The moat required committing before any learning curve appeared

  • Tiwari says Noetik opened a laboratory, bought instruments, sourced human tumors, and ran two-week transcriptomics runs processing two slides at a time with “no prior indication that any of this would work.” His summary is blunter: “Big zero. Big crazy bet.”

  • The company lacked enough data to train a model for at least 18 months; individual transcriptomics runs took two weeks for two slides. Tiwari puts the investment “closer to the 10” million dollars than $50 million. There was no obvious off-the-shelf approach for spatial data, so the AI team had to explore an “alien landscape of data” from first principles.

  • Their advice to other startups is to begin with the machine-learning problem and design the dataset backward from it. “Any dataset” does not automatically have a valuable ML use case, and small pilots can fail below a critical scale even when the full experiment would work: “There’s no shortcut to it.”

  • Noetik’s closing bet is top-down abstraction. As simplified neural networks predicted real brain responses better than painstakingly stitched biophysical neuron models, functional tissue models may predict patient response sooner than bottom-up biochemical simulations. This may be the “first inkling of the ChatGPT moment for bio,” but literature-reading agents alone will not replace new data, new ML, and clinical translation.

Ron Alfa

We basically opened the lab, hired a team, got all the instruments, and started sourcing tumor samples. There was no prior art here indicating that any of this would work.

Brandon Allgood

Big zero.

Ron Alfa

We just started generating data and sourcing human tumors. We built this whole processing pipeline to get the tumors into these arrays and formats. You have two-week runs where you're processing 2 slides, and we're just churning data for months.

We couldn't even train a model. We sort of just built all this, and then, let's say, 18 months later: “Hey, Alondra, can we train a model off of it?” It wasn't obvious.

Dan Bear

There wasn't really anything major to go off of. There were transformers developed for single-cell data, but there just weren't really datasets out there that people had been able to develop on. We do a lot of custom model building.

Hi there, I'm R.J. Honnold and this is Brandon Anderson. We're the co-hosts of the Latent Space Science podcast and today we're really happy to be in the studio with some of the people from Noetic.

Ron Alfa

I'm Ron Alfa, co-founder and CEO of Noetic, and a physician-scientist by training. My hobbies are making hot takes about AI curing cancer.

Hi, I'm Dan Bear. I'm VP of AI at Noetic. I'm a biologist by training and did PhD work in neuroscience and then moved into comp neuro, computer vision, self-supervised learning and have, you know, been doing AI research at Noetic for the past few years.

Brandon Anderson

Maybe we should start with what Noetic is, why you founded it, and what the difference is between Noetic and the other virtual-cell companies.

Ron Alfa

Maybe just start with a little bit of a contrarian thesis, which is really the reason for founding Noetic. We all know the numbers: 90% to 95% of cancer drugs fail in the clinic. Why do they fail?

Our thesis is that they fail not because we're bad at pharmacology, not because we're bad at target selection, and not because we're bad at making the drug. We're actually better at that process than we have ever been in the history of drug development.

Most of those drugs fail, we'd argue, because we're bad at selecting which patients those drugs are going to work in. Oftentimes, you see trials where there is no placebo effect in cancer. Some patients respond to these drugs, and if you have a patient that responds, that tells you that there's some active biology there.

But you have a problem in patient selection. That's really the thesis behind why Noetic exists: Can we build models that can fundamentally understand patient biology from the very beginning and help you position molecules in the right patient population?

R.J. Honnold

So you're actually using the models, at least partly, to select the patient cohort, not just so you can imagine it working either way. You could design it one way: “I think that this molecule will do well because I know something about the patient population.” But you could also say, “I think that this patient population is the match for this molecule.”

Ron Alfa

That's where the power of the models is. Once you've trained these models on patient data, you can use them on both sides of the equation.

You can use them for discovering new targets directly from the patient data, which people often refer to as reverse translation: starting from humans and then trying to understand which targets to go after. Then you can use that to develop molecules.

But you can also use them directly on patient data. If you have, let's say, a phase 2 or phase 3 trial, you can use these models to understand which patients, or what underlying biology of the patients in the trial, is a predictor of response. We've been doing a ton of that recently.

R.J. Honnold

Are you doing a lot of restudying trials that had a bad effect?

Ron Alfa

We are doing a lot of looking at data from phase 2 and phase 3 trials, and then using the models essentially to run inference on patient biopsies and understand whether there's underlying biology that would help us design the next trial. We haven't shared any of that yet, but you'll see this soon.

Brandon Allgood

Cancer is kind of infamous in that there are many, many different types of cancer. Whenever someone says “cure cancer,” that's almost a meaningless, vacuous statement.

Your point is that, even among cancers—or if you pick a specific type of cancer, and then a subtype within a subtype—there are a bunch of different patient populations, and each one of them will respond differently to drugs. Your point is that you can figure this out right now: Some subpopulation will do well and respond to this drug, while generally speaking, the rest of the population will not, even though historically we've classified this as one type of cancer, one indication, or so on.

Ron Alfa

Yeah, that's exactly right. I would maybe even go further and say that nobody actually knows what the subtypes are. There are cancers that originate in a certain tissue, like the lung, that have been classified into subtypes based on pathologists looking at them for more than a century. Those subtypes certainly have some connection to the real, like, carving nature at its joints—what are the actual functional subtypes of disease there?

But our thesis is that if you look at much richer data—the multimodal data that we're generating in our lab—we're going to see that what people thought was one subtype of lung cancer is really 3 distinct subtypes of cancer. That is going to be critical for figuring out which patients should get which drugs.

Dan Bear

Maybe I'll just go back to one of your first questions. You were asking why drugs fail in patients, and I was saying that many drugs fail in patients because we don't understand which patients they will work in, in oncology. Why do we end up in that situation?

Whenever you make a new drug, you do a set of experiments in cell culture—cells in a dish. Those cells are often cell lines. These cell lines have existed for 40 or 50 years, and they're immortalized. They have genomes that allow them to persist, with abnormal numbers of chromosomes. They have gene-expression patterns that don't represent any known cell in the human body. These are sort of Frankensteinian cells. They can't be derived or analyzed in that way. They're mostly cancer.

You can do your experiments in these cell lines in a dish, or you can move them into animal models. In oncology, you often have a panel of different animal models, with different cancer types that you'll test these in.

In doing these experiments, we sort of convince ourselves that some of these cell lines are, let's say, lung cancer cell lines or colon cancer cell lines. Then, even in the mouse context, we say that some of them are colon cancer cell lines and some are lung cancer cell lines. In the mouse, we implant them under the skin in weird places, treat the mice with drugs, and see how they respond.

Ultimately, there's a big gap because they don't translate to patient biology most of the time. Even if these cell lines were derived from a colon cancer, most of them don't have the mutations that human colon cancers have in many cases.

Pharma has done this for 20 or 30 years. You develop a drug, test it against hundreds of these cell lines, and it's not a hard experiment. You can send this out to any CRO. They'll test your drug against hundreds of different cancer cell lines, and you can sit back and say, “Okay, which of the 50 colon lines responded to my drug, and which of the 50 ovarian cancer lines?”

You could try to map that to human biology, but the problem is that cell lines, as an abstraction, do not relate in any way to your patients. Ultimately, no matter what you do preclinically, the molecule gets into the clinic and the clinical team says, “Look, we don't really know how to design this trial because none of the data that you've produced gives us any insight into which patients to run the trial on.”

So we're going to run an open-label study. We're going to enroll all tumors—all patients who are enrollable in the trial—and see where we get signal.

Imagine doing that in an early-phase trial where, let's say, you have 50 patients and you're trying to test different doses. You don't really know the dose of the drug, you don't know what the safety margins are, and you're also trying to figure out where your signal is.

What if I told you that, let's say, in just lung cancer, hypothetically, there are only 10 different subtypes of lung cancer, and you don't even know if it's lung. It could be any. This is what happens. Oftentimes, you get to the end of these early-stage trials and don't see very many responders, as you would expect statistically, and then these molecules get canceled.

R.J. Honnold

So you're imagining that, with your Noetic system, you help the pharmaceutical company characterize that they expect people with a certain genetic profile, or even transcriptomic profile, will respond to this drug. Then you sequence the patient and say, “Yes, this is a match,” or no. Is that the sort of grand vision?

Ron Alfa

I would say we're even less biased than that. We're saying, “Okay, we want the model to learn, let's say, from lung cancers. We want the model to learn how many different therapeutically relevant subtypes of lung cancer there are, just from self-supervised learning on the data.”

Dan Bear

Those subtypes could be driven by large genetic changes, immune changes, or really any biology that the model is learning in the process of training. We do see different types. Feel free to contradict this as the actual doctor here, but the biomarkers that people have been using are biased toward simplicity: does the patient have this particular mutation, sometimes a stain for this single protein, or transcriptomics to look for a particular gene signature? There’s no reason to think that biology, or the biology of cancer, is so simple that you’re going to capture most of the meaningful variation with such simple biomarkers.

Most of them have weak correlations with clinical success, but the hypothesis here is that, again, if you were to carve nature at its joints and figure out what’s really going on, there are these 5 subtypes, and the correlation between which patients you give a particular drug to and whether you have success is much, much stronger than if you’re forcing yourself to go with these very simple biomarkers.

Alessio Fanelli

You mentioned the lab. You do a lot of data generation in the lab. Why do you think that, versus using existing public repositories or whatever, is appropriate?

Tony Bui

We generate all our data in the lab, from sourcing tumor samples themselves to processing them and generating the data. Maybe another hot take I have, just in AI and bio, is that you’re not at the order of magnitude of data that you are in other spaces for building training models. It becomes really hard to brute-force these problems just by collecting data.

We have a couple of pretty good examples of where someone has designed a dataset. PDB was designed and has been built over the past 50 years or so, and it’s not an accident that that dataset exists. Someone decided that we were going to design this dataset and collect this data over decades and decades, with the intuition that potentially this would help solve protein folding down the road—and it did. It’s not just that PDB is a bunch of random data that people have organized from the web.

I think that in bio, you really need to be intentional about the data that you generate and how you generate it, and have some foresight around what models we’re going to want to train and what modalities we need to learn from the very beginning. So that’s why we’ve taken this approach.

Shawn Wang

A good comparison is the ImageNet dataset, which kicked off the deep-learning revolution in computer vision with convolutional neural networks, actually demonstrating that neural networks can do better than other methods on object categorization. ImageNet is, at least the part of it that people were developing models on, 1.2 million images, very carefully curated. These are high-quality images, not random images from the internet or multiple datasets cobbled together.

Tony Bui

And labeled.

Shawn Wang

Yeah, and labeled. I think with the data that we’re generating, we’re around that scale right now, but, of course, people have gone much, much larger in image datasets and language datasets—text datasets, obviously, for LLMs.

Tony Bui

We think that we need to get the data up to that scale before we can really see meaningful progress on the algorithm side.

Alessio Fanelli

The scale of language data.

Tony Bui

Language is really the only modality where people are seeing these very impressive scaling results. Part of that has to be just the scale of data that’s there and that the models are trained on. That can’t be the only thing, because there’s a lot of video data as well. People are training on thousands of hours of video data and haven’t seen any of the scaling results that you have in language modeling, but having the right scale of data is necessary, if not sufficient, to really make progress here.

Shawn Wang

Can I offer a contrarian take to that? There’s this whole concept of the giant frontier of LLMs, and in certain regions it can be really good at solving some problems and then remarkably stupid at solving nearby problems. Maybe the argument is that a lot of these frontier models are just becoming massive; everything is becoming in-distribution.

If everything starts out OOD, and you just get more data, that becomes in-distribution. Is it possible that for biological systems, because there are underlying physical processes here, you can basically make things more in-distribution earlier and that you can’t actually cover the space? I kind of have some thoughts with PDB, but maybe I’m just curious at this point.

Tony Bui

I think it’s a good question: how much data and what kind of diversity do you need in biology to solve, say, the drug translation problem—figuring out which drugs are going to work in which patients? My intuition from working in biology for a while is that we’re still pretty far from that, because we’re building datasets that are focused right now on cancer, and I’ve generated data from thousands of patients in a few major cancer subtypes. But there’s every other disease, healthy tissue, and even other species.

There’s a lot of biology to learn, especially if you think about it as having to learn the spatial and functional patterns of tens of thousands of genes, tens of thousands of proteins, how their spatial arrangement contributes to the function of organs, and so forth. My hunch is that biology is pretty complex and that we still need to generate a lot more data, but I don’t know.

Alessio Fanelli

Yeah. But as a cancer company, do you think you could actually do this hypothetically for cancer? I mean, for at least some subclasses of cancer?

Tony Bui

Definitely. I think we’ve done experiments that suggest that if we can generate data from several hundred patients in all of the major cancer indications and some of the less major indications, that will result in a model that can generalize pretty well to any type of cancer we would throw at it.

Alessio Fanelli

Backing up, what is the data you’re collecting? My understanding is that you use some pretty specialized instruments and gather very specific datasets. How did you come to that decision about how much data to collect, how much to spend on it, and what types of data?

Tony Bui

I’ll give a hat tip to my previous employer, Recursion. I spent 6 years at Recursion, from the very beginning, and a lot of what we were doing in the early days was figuring out the things we didn’t understand about the datasets and what the problems would be in the dataset: batch effects, control design, orientation of samples on plates, things like that.

Flash forward to founding Noetik: I started the company already with some principles around how we should think about building the dataset. What are some things that we know mattered? For example, over many years we learned that images were actually a really powerful dataset for machine learning, for many reasons. One is their scale. We can put patient samples on slides, and on a single slide we can capture many patients’ worth of data.

The images themselves are very rich sources of biological information. Beyond that, we have a very information-dense modality, and we can decrease the cost of data generation, so then we can increase the amount of data generation over the whole dataset. That’s always been a really big benefit to image-based modalities over, let’s say, sequencing, where every time you run a sequencing run, your run is a patient set, for example. That was one way to think about it.

The other was: how do we design these datasets so we can control for things that we know are going to be important, such as batch effects? For example, if I have a slide and we do, let’s say, a spatial transcriptomic run on that slide, you stain the slide, do a bunch of wet-lab processing, put it into a machine, and get data out. If you do that on 2 different days, there are going to be different variables that impact the data quality. That’s going to be a large source of variation in datasets.

You want to be able to control for things like batch effects. Really, you want more patients represented on multiple different slides so you can process them in different batches. You want to be able to control for things like this so you can go downstream and look at the data and say, “Okay, once we have, let’s say, patient-level embeddings, we can ask: Do the patient-level embeddings represent patient response to immunotherapy, or do they represent staining batches?”

Alessio Fanelli

So you’re actually taking 1 patient and spreading it across multiple slides so that you can get a—it’s sort of a calibration across the slides?

Tony Bui

Yes. Our data looks very different from anyone else in the space of generating data on histology or digital pathology types of specimens. We receive a sample, sample those samples dozens of times to build these arrays, and each array has hundreds of patient samples randomized. Every patient is represented on multiple different arrays, and so we’re getting a lot of different representations of each patient that we’re sending through the data-processing pipeline.

And then that lets you downstream answer some of these questions and control for some of these variables.

Shawn Wang

You mentioned some terms I just wanted to define for people. What is spatial transcriptomics?

Tony Bui

Yeah. This was your first question: What are the data types? If you sit back—and this is not my background in terms of spatial biology—again, everything we did previously was cell biology in a dish. If you said, “Okay, I want to train a foundation model that understands human biology,” what does that mean? How would you go after that problem? That was really the starting point for the company: “From first principles, how would we do this?”

You probably want tissue-level biology. You want to understand tissue: cells are organized into tissues. You probably want some modality that is relevant in clinical use, so you can relate clinical data to what your models are learning. That’s why we generate pathology H&E. That’s what every patient gets: a tumor removed, and then it gets stained with H&E.

I can explain what H&E is: it’s basically 2 different dyes, hematoxylin and eosin, and it really just creates a contrast over the tissue. You’ve probably seen these purplish pathology specimens. Pathologists can look at those and identify different cellular structures, and they use those to classify tumors based on the classical classifications—adenocarcinoma, small-cell carcinoma, things like that—based on cellular structures.

Shawn Wang

Okay, so there are specific patterns that show up when you add these 2 stains, and it is well established that you classify tumors based on those patterns?

Tony Bui

Based on pathology, yeah. Basically every tumor that gets processed in the hospital will get this H&E stain, and it’s how the pathologist typically classifies a tumor. So that’s the first level. You probably also want to understand cell types.

It’s really hard to understand cell types from just that stain because it doesn’t reveal that much that a human can use to classify cell types, at least. You can say, “Well, I want to know whether there are immune cells and different subtypes of immune cells.” We want to have some layer of cell biology.

swyx

Okay, so you want to know about immune cells because you have these cancer cells, and oftentimes the immune response dictates whether or not you’ll have an effective treatment?

Aviv Regev

The immune environment of the tumor will be a core—we know it’s a core constituent of whether a patient is going to respond or not. So you want to give the model this tissue-level information. There’s not enough cell-level information in there for the model to learn enough cell biology about different subtypes, so we also want to present it with some cell-level information.

We use protein stains, so standard immunofluorescence. You basically use antibodies against a small set of cell markers to label your different key cells—B cells and the standard subtypes of cells in the tumor microenvironment.

swyx

So, in this stain, just for those who are familiar with it, the antibody has a fluorescent protein. When you hit it with a certain frequency of light, it fluoresces, so you can tell the antibody bound to a certain protein, and now it has a fluorescent protein attached to it.

Yep. And in terms of the data, from the tissue layer you have an RGB image. From the next layer, you have a multichannel image, with each channel representing, let’s say, 1 color. For example, certain immune cells are each in a different channel, so you have this multichannel image.

That’s great: we’ve got tissue, we’ve got cells. But if we actually want to make drugs, we need some type of molecular information. We need to tie all of this down to what’s happening in the genome. What is the cell doing? What are the mechanistic principles of the biology?

So then we get spatial transcriptomics. That’s spatially resolved RNA: DNA is transcribed into RNA, which is translated into proteins. We get basically the RNA in a spatially resolved pattern for the same cells that we’re seeing in all of these other layers. Now you have between 1,000 and 19,000 different genes. Again, these are all image layers, with spots showing where those RNAs are and in which cells.

swyx

And this one works a little bit similarly to how we talked about protein, where you have a segment of RNA and then you have a fluorescent protein. Usually there’s some sort of combinatorial thing, so if you see these 4 colors at this amplitude, that means this gene because they’re right next to each other or something like that.

For the detection, you’re basically binding a probe to each one of those RNAs, and then you’re cycling it. It takes weeks to run 1 of those assays. The machine will cycle across each species, they’ll amplify, and you’ll get a signal for each RNA species.

At this point, you now have this very rich data layer where you have the tissue, the cells, and the molecular information. You can use all of that to train the model. We think of it as, essentially, the central dogma, if you will. We also have DNA; we genotype just so we understand the genomic alterations in these tumors.

swyx

All right. So you have this stack of images, basically, that you can train models on, with an understanding of the expression of genes and the proteins that are being expressed at the time that the sample is taken, all in the image information. Then you can train your models with that.

Yeah. Spatial transcriptomics is particularly dense because if you think, let’s say, there are 20,000 genes in the genome, we’re running assays that are detecting nearly all of them in a single sample. You can think of 1 of those data points as an image, except instead of being an RGB image that has 3 color channels, now all of a sudden it has, like, 20,000 color channels.

It’s a very meaty computer-vision problem to try to look at those data and figure out what makes patient A different from patient B, and then go from that to which drug is going to work in which one.

swyx

So you have a hot take about virtual cells. I want to understand how—okay, so you have this big pile of data: every single sample has a massive data set with it, and then you have many, many samples. How do you turn that into useful knowledge?

Maybe just: What is a virtual cell? Everyone is always asking that question. I think there are really 2 ways to think about it. One is that we want to be able to simulate all the biochemical processes in a cell.

We want to have this comprehensive foundation model where we understand that if some signal from outside the cell interacts with the cell, then here are the millions of intracellular chemical reactions that are going to happen, and you could predict them from the model. That’s 1 view. I think that’s interesting; it’s an interesting intellectual pursuit. I don’t think we have all the modalities of data that you would need to solve that problem.

I tend to see the virtual cell problem as something more practical. We’re trying to make drugs that work in patients. From a virtual cell perspective, what we really want to do is understand cell biology in some heuristic that’s useful for making drugs. The heuristic could be either a way to understand drug targets or a way to map your cell-level biology up to patient-level biology.

The way we design these first virtual cell models is really just to simulate the biology of a cell in some context. The biology of that cell is, let’s say, the cell being in some context, with the output being the transcriptome in that context, or the protein in that context. These types of input-output relationships allow us to design experiments.

The very simplistic thing that we’re doing is that the model can simulate the biology of a cell—or many cells—in different contexts and allow you to run some simulations in that regime.

swyx

Yeah. I think most of the things that people are calling virtual-cell models right now are focused on single-cell gene expression, so transcriptomics data, RNA data, and they’re largely geared toward predicting what’s going to happen to the transcriptome—the set of genes expressed—when you hit cells with either a small molecule or drug, or a genetic perturbation. Typically, these are cells grown in vitro, either in cell culture or as primary cells, something like that.

The genetic perturbation is where I knock out a gene or add a gene and see how that impacts the expression of the various RNA.

Exactly. I think my view, and I think Ron shares it, too, is that that may be of interest in some cases, but the problem we’re really trying to solve is predicting what’s going to happen in a patient. Just modeling data that comes from a patient is, in my mind, much more likely to translate to what happens when you give a patient a drug than something that’s happening in cell culture.

swyx

Is there other clinical data that you’re pulling into the model besides the actual—so you’re calling it the context of the cell, just the surrounding cells—but is there other “this drug caused a bad reaction” kind of stuff?

Yeah, we’re pulling in data from the entire patient, not just the very local neighborhood of the cell.

So far, we haven't done much integration of electronic health records or other information that one could get about the patient, and that's pretty intentional. We really want these models to learn basic biology—again, the central dogma, but not just the central dogma: the basic biology of genes, proteins, cells, and tissue—in a self-supervised way.

We want to do that purely from the data that we're generating and not be biased by what the doctor wrote about that patient. Our thesis is a bit like this: most of the therapeutically predictive and important information is not contained in the very small number of patients who've been treated with a given drug, or in whatever the doctors thought was important to write down given the state of knowledge at that time. It's much more about trying to discover what's really there in patient biology than going based on the text that people have written about it.

Alessio Fanelli

So, you have this self-supervised model. You use a lot of data. You have essentially some clusters of patients now. How do you translate those clusters of patients into making decisions? You go to a pharma company and say, “We can repurpose this drug,” or, “We can suggest this subtype should be the focus of your phase 2 trials.” What is the process for that? What data do they need to provide you, and how do you translate your models?

Aviv Regev

It depends on what the problem is. I think it's important to note that one of the more interesting aspects of these models is that they're useful for a broad array of use cases, as we were talking about from the very beginning.

You, as the pharma company, could say, “Okay, well, I have this molecule, and the target of the molecule is X, and I want to design my clinical trial. The molecule has seen 0 patients so far. All I know is the target and some biology around the target.” We can run simulations using the models and our cohorts of patients.

Let's say, if we were to look in lung cancer, we can run simulations around the target and ask, “Which sets of patients here would this target be important in?”—across a cohort of lung cancers and colon cancers, or across all of oncology. You might see—and we see this sometimes—that your target probably isn't something you want to put in lung cancer. Maybe you want to put it in ovarian cancer because it's not really important in lung cancer.

Alessio Fanelli

What are you simulating here? Are you saying that this drug is expected to knock down this gene, and therefore you want to look for clusters where knocking down this gene inhibits tumor growth rather than enhancing tumor growth?

Aviv Regev

That's certainly 1 way we could do it. There are other types of simulation where you might just want to ask: if there were an immune cell here, like a T cell, which is responsible for actually killing tumor cells, what would happen to it? What genes would it express, or what proteins would it express, in this particular patient's tumor microenvironment?

Travis McKie

That's what we've called these virtual cell simulations. We have a model called Octo Virtual Cell that does this, and that can give quite powerful answers to the question of whether these drugs are going to work in these patients. You might find, actually, as Ron was saying, that the thing this drug targets is just not important in this particular patient's tumor, and it's not going to have any effect on the T cells, the macrophages, or some other cell type there.

Then there's the type of simulation you alluded to, where you can ask the model, “What would happen to this patient's tumor if you were to knock down this particular target gene or its protein product?” You might be looking for cases where the model predicts that removing that gene or protein is going to have a large effect—either increase the immune system's function, its ability to fight that tumor, or decrease the tumor's ability to grow—or some other readout that you think is correlated with clinical success.

I just want to call out that perhaps the simplest use case is the one where there's a company that has a drug, they've given it to some patients, and we know some of those patients responded. It becomes a question of whether the space of patients that the model has learned via self-supervision tells us that all of the responsive patients are in 1 of these clusters and not the other 9 clusters, or something. If we know that, then there's a pretty straightforward hypothesis that this is the right cluster.

Alessio Fanelli

So, that's the scenario where you would sequence something. What would you collect about those? You have a cohort that responded and 1 that didn't.

Travis McKie

This is getting back to something Ron mentioned earlier, which is this type of data called H&E. It's the standard pathology stain that makes these pinkish- and purplish-looking images.

Right now, we've built models that are trained on all of the multimodal data we generate, but once they're trained, at inference time, all they need is an H&E image. That could be something that we generate in our lab, or it could just be a digital image that they have from a trial that was run years ago.

The reason that's so powerful and flexible is, again, because H&E is the lingua franca of pathology, and especially oncology. Almost every patient who's been given a clinical-stage drug is going to have that.

Alessio Fanelli

You can look at the 2 cohorts—the responders and the nonresponders—and say, “These H&Es live in this part of the latent space, and these H&Es do not.”

Travis McKie

Exactly. One way we've gone further than that is that, given the H&E, the model can say, “I predict that these genes are expressed at this location in this patient.”

Not only do we have these clusters, these embeddings that say all of the responders to this drug are over here and all of the nonresponders are over there, but we can actually see, “Okay, for the responders, these are the genes that are expressed much more highly—or predicted to be expressed much more highly—in the responder cluster versus the nonresponder cluster.”

That adds a major level of interpretability, because we can see things like, “Good, the responders are actually expressing the protein target of this drug.” We would be worried if that weren't the case, but we can see that it is. On the other hand, we also see that the biology is very complicated. That's part of explaining why these simple biomarkers, like looking at a single gene or a single protein, just really don't capture what is predictive of therapeutic response.

Alessio Fanelli

Yeah, so I have a million directions I want to go here. The H&E actually gives you a pathway to a diagnostic then as well?

Travis McKie

Exactly.

Alessio Fanelli

Yeah, right.

Travis McKie

Yeah. You can imagine that after the drug hopefully makes it to the market, the doctor says, “Oh, you have cancer. I'm very sorry. We're going to do an H&E stain of your tumor, and then we're going to put it in the model.” It says, “This one won't work, but this one will.”

Alessio Fanelli

That's right.

Travis McKie

We're using the same approach today, where we're looking at many different mechanisms from different collaborations that we have in place. One of them we've announced with a company called Agena. These are all different mechanisms. The input is still H&E, using some of the same indications.

Using H&E, we're looking at whether drug A works in some sets of patients and whether drug B works in other sets of patients. You can take that to its natural progression and say, “Well, okay, if you can use that same input—just H&E—for experimental drugs, why not use it also for drugs that are already on the market?” In a sense, the same assay can be very predictive across many different cancers and many different potential therapeutics.

Alessio Fanelli

There are lots of models that take H&Es and go to gene expression out there, open source and otherwise. They do so-so. I've read on Twitter, in your Twitter feed and whatever, that you feel you have a data moat, right? So, why is Noetik's model better?

Travis McKie

Sure. I think the scale of data that we've trained these models on is pretty different from a lot of what's out there. The reality is there's just not that much of this kind of paired H&E plus other data modalities.

Typically, there are some datasets generated by academic labs, and others where they might have maybe 100 or a few hundred patients' worth of data. With paired spatial transcriptomics, that might even be an overestimate.

In comparison, we're generating these data with multiple patients per slide and individual patients distributed across multiple slides. We've generated more than 100 million cells with spatially resolved transcriptomics. That's all paired with H&E and protein as well. It's at least an order of magnitude larger than any of the other datasets that we've seen out there, and I think that makes an enormous difference.

We've seen with our own models that if you drop down to 40% or 10% of the data used in training, the models get a lot worse. They especially get worse at generalizing to other types of cancer from the ones that they've been trained on.

So, I think that's a big piece of it. I also think that the algorithmic side of it is important. We've developed custom architectures specifically for training on this multimodal data. Again, my background is in computer vision, and specifically in self-supervised learning there.

We've tried to develop self-supervised learning approaches for these data that are really adapted to solving the problem of figuring out what is different in one patient versus another, and then simulating what would happen if you were to knock down a particular gene or protein or something. This is why we call these world models: we're trying to build models that can simulate what's going to happen if you take a particular action. I think that's another big differentiator for these models. And then, again, interpretability is probably a third one.

Alessio Fanelli

But Travis, you were just talking about how one of the other strategies people take for this is to do perturbations on cells and then watch the response. Your experience, plus your strategy here, is that you can simulate this sort of counterfactual perturbation idea without even having to collect the data to do that, and you can see this—

Travis McKie

Well, there's—yeah, there's a big piece that we haven't talked about yet, which is that we are actually running perturbation experiments, except they're in vivo perturbations using a platform based in mice. We have another platform called PerturbMap. Ron, if you want to describe any of it—basically, this is a platform for generating highly multiplexed knockouts of individual genes.

It's the same kind of CRISPR knockouts that people are doing for individual cells in vitro, except when we knock out a gene in a cancer cell, that cancer cell gets injected into a mouse. It's barcoded, so we know which gene was knocked out, and it's being injected alongside roughly 100 other cell types with different genes knocked out. You end up with mice that have tumors that are barcoded and have 100 different genetic perturbations in them.

We can actually use that to validate our models and ask whether what the models are predicting in humans via simulation is actually borne out when you do these perturbations in a mouse system.

Shawn Wang

Sorry, there's a lot to unpack there.

Ron Alfa

Yeah, barcode. Yeah, so, sorry—barcoding. This is a technology in which an individual gene is knocked out with CRISPR, but this also introduces a set of protein tags in that cell that get expressed. It's a combinatorial code, so gene X might have proteins A, B, and C, while gene Y, when it's knocked out, has proteins D, E, and F.

We can tag those proteins, or label them with antibodies, so that when we go and look in the mouse, we know exactly which gene was knocked out based on which of those protein tags were expressed.

Alessio Fanelli

So, you knock out a gene, but you also add a gene that has the barcode proteins encoded on it?

Ron Alfa

Yeah, exactly.

Travis McKie

And I mean, the system is designed so everything that we're doing here is tissue-level, or could be in vivo—you know, tumors that are in vivo and in the form of the tumor in the whole tissue. Here in this mouse system, you have hundreds of tumors in the lungs of a mouse. If you look at these images, it's a mouse lung with literally hundreds of tumors in it.

Each tumor has a distinct biology that's driven by the biology of the knockout of the gene that's being perturbed, and we can capture the biology of each tumor in a spatially resolved way. What you can see is that we have certain tumors in humans that don't have immune cells in them. Those tumors are very aggressive, and they don't respond to immunotherapies.

You can generate those same tumors in this mouse system, and again, they don't have immune cells in them. You can do it genetically, so you can start to map the causative gene relationships between these different immune—or, more broadly, tumor—genotypes or biological profiles, if you will, to what you see in the human.

You can then treat those mice with drugs and see how hundreds of tumors in a single mouse respond to treatment with one drug. Or you can treat, let's say, 50 different knockouts across a panel of mice with 50 different drugs, and you can start to build this intersectional pharmacology and genetic experiment.

swyx

On Twitter and various places, I've heard you say, “no attic is no no cell lines, no war bottles.” Maybe you even said that, you know, if you—what's a cow?

Travis

And then we just said we have a mouse model.

swyx

Yeah, that's it. And we're injecting cell lines into the lungs, not under the skin.

So, yes. Fundamentally, you would think it's really important to build models that are trained on human data, and we're sourcing all these human tumors to build human-centric models. So, that is also true.

From the very beginning, we have asked this question: let's say we want to develop a drug from the very beginning, and let's say the FDA—and I know things have changed a little bit with the FDA—wants you to have some data in an animal that says your new mechanism works in some animal system. What do you do?

You're kind of stuck because you've now generated arguably the best data that you can in the human system, and then the FDA says, “Well, cool, but does it work in the mouse? How does it work in the mouse?” And so you have to back into this system that doesn't translate.

From the very beginning of the company, this has been a question. We started, probably at the same time we started generating the mouse-to-human data, building this mouse platform with the aim of drawing connectivity between these 2 systems.

We wanted a platform that would allow us to map the diversity of human tumors, because we know that if we just run a mouse model with 1 tumor, that tumor has no connectivity. In the mouse system, we want to have diversity of tumors, and we want to see a mapping of diverse tumor biology to the tumor biology that we're seeing in humans across many different locations.

We've been building this system so you can see many different perturbations that produce a lot of the tumor biologies—the plural—that you see in humans. We also want to be able to get from this mouse system to biologically relevant targets or genes in humans.

One of the fundamental problems in mouse systems is that we share many genes with mice, but there are a lot of genes and biological processes that we don't share with mice, as is obvious. Oftentimes, when you're developing drugs, you run into this situation: you have a target, and you have some biology that works really well in mice, but maybe that doesn't even exist in humans, or maybe that pathway is useless in humans.

One of the things we've started to develop, which I'll share more about soon, is a way to use one of these models to essentially infer human biology from the mouse directly. We're in silico humanizing the mouse. All the outputs in terms of the transcriptome from the mouse are in the form of human genes. When we read out this mouse system, we're reading it out in the form of human organ health.

swyx

How do you validate that? I mean, that's a pretty impressive claim if you can do it, but, man, it seems like a tricky validation task.

In my experience, both here at Owkin and at my previous employer, I could say a version of that. A lot of the approaches you're looking for when you're building these types of models involve asking whether the models are recognizing biology that you know to be true.

For example, in the human context, we know that 12% of patients with lung cancer respond to immune checkpoint inhibitors. Do the models recognize those patients? Can they recover those patients without training?

swyx

I mean, cold?

Yeah. And we see that. When you go look at those patients, we see that the underlying features of those patients map to what we know about those patients in the clinic.

In the mouse system, we have control genes. We ask: if you look at the mouse tumor embedding space, do the tumors that should be really cold look really cold from the human inference?

swyx

Cold in the sense of—

Yeah, like they don't have immune cells. And then hot in the sense of lots of immune cells. We try to build systems where you have these handholds, and the more of these examples that you know to be true that work, that you see, the more confidence you have.

Obviously, when you're in the regime of something very new, it's still uncertain for some of these.

Alessio Fanelli

So, the bridge between the mouse and the human is that you build a world model on the human, you build a world model on the mouse, and then you say, “What are the parallel structures in the 2 latent spaces?” Is that kind of the intuition here?

Travis

That's one thing that we're doing, but actually, this is even simpler. We've trained models on human H&E, spatial transcriptomics, et cetera, and then we're just running inference on mouse H&E, which is easy to generate.

Apparently, mouse H&E looks enough like human H&E that the models think it's perfectly valid H&E. They make predictions about whether it's immune-hot and immune-infiltrated, cold, fibrotic, or some other tumor phenotype.

And those predictions are accurate. These are some of the controls that Ron mentioned. We know that in mice and humans and everything, if you knock down tumor cells' ability to present antigens to immune cells, those are very cold. Immune cells are nowhere near those tumors. That's exactly what we see in the mouse, and that's exactly what the models—the in silico humanized models—predict.

There are other examples where, again, we're recovering the biology that we expect to see there. Then there are findings that are novel but also make total biological sense. For instance, we have done knockouts in the mouse of half a dozen genes that are all in the same pathway. You might predict that knocking down those genes is going to produce the same phenotype because they're all on the same pathway.

swyx

And that was a pathway?

Yeah, a pathway is like protein A signals to protein B, which signals to protein C, and there's a chain of events that leads to the cell having some behavior—changes in its metabolism, its growth, et cetera. I don't know if you've ever seen these crazy-looking protein-signaling diagrams that make you want to stay away from biology. People have worked out a lot, and they know that these 2 proteins interact physically and signal to each other and so forth. And so, some chain of those—

swyx

Interactions

this protein binds this protein, and that causes it to upregulate a gene that causes another protein to be formed, blah blah blah, until you get to some phenotype, meaning the cell changed the way it looks or the—

Exactly. Based on decades of biological literature and experiments on these, there's a very strong biological prior that if you hit gene A, gene B, and gene C, and they're all in the same pathway, you should get similar phenotypes. This is kind of how old-school genetics was done, and we see that with these in silico humanized mouse models, which is amazing to me as a biologist: you have a model that's trained on human data, then you show it some mouse histology, and it's able to say these 5 different tumor genotypes all look like they have the same phenotype—and lo and behold, they're 5 genes that are in the same pathway.

swyx

So, you guys are switching gears a little bit because we want to talk about models on the Latent Space podcast. You guys recently—there was an interesting blog post about the TARiO model. It's a transformer-based model. Do you want to talk about that?

Sure. This is a new model architecture that we developed after the first virtual cell model, OctoVC, that we developed. TARiO is just a different transformer architecture. One major difference between it and our prior models—if this is a model podcast, this gets into the self-supervised learning objective—is that for a while, including with OctoVC, we were training models on what's called the masked autoencoding loss function, or objective.

You have a piece of data, chunk it up into small chunks, mask out some of those chunks, and the training task is for the model to predict the masked-out chunks from the revealed chunks.

swyx

Like BERT.

Travis

Yeah, exactly like BERT.

swyx

What are the chunks? Because this is multimodal, and I would imagine the different channels contain wildly different levels of information. I remember seeing something like 99% masking in OctoVC, if I'm—

Yeah, yeah. So—

swyx

And that was kind of surprising because when you have 20,000 channels, and maybe some of the channels are fairly—most of the signals are fairly sparse—

Yeah.

swyx

Then it seems like it'd be either there's a huge redundancy here in your data, or you really risk just throwing the baby out with the bathwater.

Travis

Yeah. What are the chunks? That totally depends on which modalities we're talking about. In spatial transcriptomics, one chunk, or one token, might be the level of expression for a particular gene at a particular spatial location. For protein images—multiplex protein images—it might be the image patch for that particular protein at a particular location, and so on. For histology images, those are usually just patches of the image, so pretty standard, like vision-transformer style.

The masking—and the maybe surprising result that you can, and actually need to, mask out large amounts of the data to get the model to learn anything interesting—is important. If you ran the hypothetical where you only mask out 10% of the image, more like BERT, for instance, in language modeling, what do the models learn? They learn these boring behaviors, like how to continue an edge a little bit between 2 regions of an object or something. They can learn that task very well, but they don't end up learning anything about the holistic structure of the image data.

Akash Tiwari

We found pretty early on at Noetik that the same thing was true with these multimodal transformers: if you mask out a lot of it, there are actually pretty strong correlations between where protein A is expressed and where protein B is expressed, and forcing the models to learn them is really what gives them this predictive power.

Shawn Wang

And so, TARiO, though, is an autoregressive model?

Akash Tiwari

Yeah, exactly. That was going to be the plan. Prior models, including OctoVC, were of this masked autoencoding-style training objective. TARiO is an autoregressive model, which, if you think about it, is kind of a particular choice of masked autoencoding, except instead of randomly masking out part of the data, you're always asking the model to predict the next token in a sequence.

We know that this is something that scales very well with LLMs, like training on the next-token prediction task, and it's still an open question: how do you get models of other data modalities to scale the way that LLMs have scaled? TARiO was not actually our first attempt, but one of our subsequent attempts to bring that autoregressive, next-token prediction task into modeling spatial transcriptomics data.

We found that when we use this architecture and this task, we started to see much better scaling behavior, where bigger models, and especially at longer context lengths, were really outperforming smaller models at shorter context lengths.

Shawn Wang

Because they can see further in the image?

Akash Tiwari

Yeah, that's probably a big part of it. I think there's actually a pretty subtle but very interesting result in that blog post with TARiO, which is that you only really see the benefits of using larger models when you're looking at longer context lengths. Here, longer context really means, again, that you're seeing more tissue at once, more area at once.

I'm not super deep into the language-modeling literature, but I don't know if there's an analogous thing with language models, where you only see these scaling behaviors at longer context. So, what we could be finding here is that, with patient data, you really do need to incorporate more of the patient's spatial context to get the models to learn these more complicated nonlinear patterns in the spatial transcriptomics and take advantage of it.

Shawn Wang

Is it possible that part of this is because you have some number of low-expression genes, and that the behavior is driven entirely by some of those low-expression genes?

Akash Tiwari

Yeah, definitely possible that the more context you have, the more likely you are to catch these low-expression but highly predictive genes, et cetera. I would guess it's a combination of that and larger area. We've done some experiments just comparing a model with the same amount of context, but in smaller or larger areas, and there definitely seems to be an advantage to looking at larger regions of tissue as well.

Shawn Wang

I want to hear about this—you did a big deal recently. You got a lot of press, and I think you have the distinction of being one of the only AI-for-bio tooling companies that's making money.

Akash Tiwari

Accidental. Nope.

Shawn Wang

So, could you tell us whatever you can disclose about that? We love hearing about this.

Akash Tiwari

We were really excited to announce a deal with GSK where we licensed them OctoVC, which is our virtual cell foundation model. We announced that back in January. It's a $50 million deal and includes an upfront payment, milestones, and, separately, an annual model licensing fee.

I think this was an attractive deal for both parties, for us and for GSK, because the deal focuses on models that we've already trained on lung cancer and colon cancer. It allows us to provide them with access to the models. GSK is one of the top AI teams in biopharma, so they know how to use these types of capabilities. They can use them for their internal use, and they can also use them to fine-tune on their data.

That was a really big sell for GSK as well, because GSK—and every pharma company—is sitting on mountains and mountains of so-called translational data. The types of data that we're training the models on come from clinical trials and pathology specimens across many different therapeutics. Everyone's sitting on a lot of this data, and it's been very hard to unlock.

All of a sudden, GSK can use our models both to do simulations and to do therapeutic discovery.

But they can also fine-tune models on their data, and in a way, the model then becomes GSK's version of the model. This was super exciting. It was the first announced foundation-model licensing deal in the space, and frankly, it was one we'd been trying to do for a long time, even before Noetik.

I think a lot of companies have been trying to do these types of deals, and it's historically been slow for adoption on the pharma side. It's also been slow to demonstrate a very clear value proposition for different types of capabilities. What's unique about this deal is that it doesn't look exactly like a software licensing framework—for, let's say, a small amount of money with a number of seats, where you license a model.

It looks like a real business-development deal in the industry, where there's a very significant, multimillion-dollar cash upfront, near-term payment. But then the substrate of the deal is not a molecule. The substrate is actually a model, which is what really made this break.

Shawn Wang

Why do you think there's appetite for this suddenly? It seems like almost whiplash. It seems like only maybe a year or 2 ago that bio was dying and whatever, and now suddenly there's this deal. Boltz is getting a ton of attention. There's so much attention on Isomorphic Labs.

Akash Tiwari

I mean, maybe not totally, but increasingly more people in pharma, across the industry, are seeing the value of different capabilities. They're able to use some of the open-source capabilities and demonstrate the value to themselves internally.

If you look at a pharma company, these companies are working on dozens and dozens of programs. My opinion is that pharma increasingly wants to be able to access models not just for one collaboration, where you and I are working together on this one program. They want to be able to access the technology across the whole pipeline.

I think that's going to create a driving force for not just bespoke, project-driven licensing, but actual broad licensing, where a pharma can access the technology in many different therapeutic programs.

Dan Fu

With the structure of prediction models—protein-structure-prediction and binding-prediction models—there's a massive public data set. There are increasing amounts of data people can generate to augment that. There's enough data to the point where people can train very good models, but maybe not just on the data that any one biopharma company has.

I think the same is true, but even more so, for the types of models that we're building, which are foundation models at the patient-biology level. No one company—I mean, these companies may have a lot of data, but it's scattered, it's siloed—and pulling everything together to train an actual foundation model may not be as easy as it sounds within a single company.

That's the nice thing about being a startup here: we can make that bet that you actually do benefit from generating all of this data in a uniform, unbiased way, at very high quality, and then use that to develop and train the models. My opinion is that you need to have data at that scale before you can even think about developing models that actually work.

You can't do the AI R&D or build the algorithms until you have a good enough data set to tell you whether your favorite algorithmic idea is actually working or not. That's a major advantage for us: we have enough data to see whether my idea or someone else's idea about how to build the model is actually leading to improvements there.

Akash Tiwari

Yeah, I mean, this is a good point. Sometimes people ask me, "Why don't you just generate your data?" We just started generating data 4 years ago. There was no model.

Shawn Wang

How many years? Like 2 years, maybe? A year and a half at least before you had the first trained models working?

Akash Tiwari

3 or 4 years still. So this is year 4 again. We basically opened the lab, hired a team, got all the instruments, and started sourcing tumor samples. There was no prior indication that any of this would work.

We started generating data and sourcing human tumors, processing them. We built this whole processing pipeline to get the tumors into these arrays and formats. It takes weeks—it takes literally 2 weeks for a machine to run a couple of slides on transcriptomics. You've got these 2-week runs where you're processing 2 slides, and we're just churning data for months.

We didn't even have enough data to train a model for at least a year and a half. Then you're building processing pipelines, you have to align all the data, and you've got to post-process it off the machine. We built all this, and then, let's say, 18 months later: "Hey, I wonder if this stuff—"

It wasn't obvious. It wasn't like, "Oh, we're going to train this off-the-shelf on some open-source architecture." Dan and the team have done a ton of work.

Shawn Wang

Big zero. Big crazy bet.

Akash Tiwari

I just went for it. We started generating data and sourcing human tumors, processing them. We built this whole processing pipeline to get the tumors into these arrays and formats. It takes weeks—you know, it takes literally 2 weeks for a machine to run a couple of slides on transcriptomics. You've got these 2-week runs where you're processing 2 slides, and we're just churning data for months.

Dan Fu

Yeah, there wasn't really anything major to go off of. There were transformers developed for single-cell data, but incorporating spatial data into that was—again, there just weren't really data sets out there that people had been able to develop on.

We do a lot of custom model building, and I enjoy that. I think people enjoy that.

Shawn Wang

Yeah, you're really unique, innovative, and bold. Sorry, who are you looking for? What kind of people?

Anybody excited about doing ML research on this kind of alien landscape of data, where you really have to figure out what's working from first principles, and obviously the work we do should have very, very large impact.

We're definitely not restricted to people who have a biology background. People who just like tackling very challenging machine-learning problems and are open to learning the minimum amount of biology necessary to make progress would be great candidates.

Shawn Wang

Talking to you guys reminds me a lot of the bio lads.

Dan Barr

Yeah.

Shawn Wang

I know that both of you are part of the Recursion mafia.

Dan Barr

You know, I'm not.

Shawn Wang

Well, yeah, yeah, yeah, but—but you—this, yeah.

Dan Barr

Yeah, yeah, yeah. We're going to be on the show in the future, too. We're looking forward to that.

Shawn Wang

It's interesting because both of you seem to have really similar philosophies. You have deep convictions that you're just going to start collecting data before you know this is going to work. You're going to brute-force it—go, go, go—and eventually it will work. You have signs of progress. I don't know, I think that's really impressive.

I wonder if there's something about Recursion that's in the water, which has led to this sort of thinking of, "We're going to commit to doing things at scale, and it may not work at first. You have to hit a certain point before it will."

Dan Barr

I mean, we failed a lot at the beginning.

Shawn Wang

Yeah. You mean at Recursion?

Dan Barr

At Recursion, yeah, yeah. We had to build it from first principles, and we really did. We spent many years trying to figure out what the data should look like. Ian and I were both involved in platform development: how to design these data sets, how to design the experiments, and iterative cycles over the years saying, "These things did work; these things didn't work."

At the end of coming out of Recursion, I think what a lot of folks there had was an understanding of what we need to think about. Even if I wanted to design a different data set today, and it's totally different, what are the things we learned over mistakes—or not mistakes, but trial and error, basically—over that many months that we would try to insert in our new approach?

I don't know that everything that I've predicted at Noetik in terms of how to generate the data set has been important, necessarily. I know that we could start at the very beginning and say, "Okay, let's make sure we do these 10 things." I know every one of these 10 things was important before, so let's at least make sure we do these 10 things.

I don't know that all 10 things are important for us today, but I would presume that many of them are, and that lets you leapfrog that process of trial and error a little bit. Certainly, we still have trial and error, but hopefully we're not having to solve 15 problems. Maybe we're only solving 3 or 4 problems over time.

Shawn Wang

So, for small biotech startups, which are probably in the AI space and are collecting their own data—their own data moat—do you have any advice or suggestions for how to be more successful there?

Dan Barr

I think you need to think ahead to, okay, what am I trying to do on the machine-learning side, and what is the right data for solving this problem? Oftentimes, I see a lot of companies say, “I want to generate X data set. I’m just going to generate X data set, and then I’m going to do machine learning on that.” That might not be the right data set; you might not have designed it the right way. It doesn’t follow that any data set has a machine-learning use case, or that that data set is going to solve the problem you’re trying to solve.

For me—even in life—it was: What problem are we trying to solve, and what data are going to help solve that problem? Rather than going from the data directly to trying to solve the problem.

I also had a quick piece of advice: Pay attention to where the technology is and where it’s changing rapidly. I finished my PhD in 2016. I did a lot of looking at spatial RNA via a technique called in situ hybridization, the same technique that is at the base of what we’re doing. I could look at maybe 2 genes at a time on a single sample, and that took me a full week of manual work.

I came to Noetik 5 or 6 years later, and all of a sudden there were platforms where you could look at 1,000 genes or 20,000 genes at once with a single machine that could run this assay. It’s expensive, but it’s data beyond the wildest dreams of Dan Barr in 2016, and that is only improving rapidly. I think it’s important to see what the technology of today allows and also where it’s going in terms of what data to generate.

Shawn Wang

And what does that pitch look like? “I’m going to generate data for a year and a half, and then I’ll spend $50 million, and then—”

Dan Barr

It wasn’t $50 million. It was maybe closer to $10 million. If you’re going into a regime where there are no data and you want to do something different, there’s no shortcut to it, right? You’re going to have to generate the data set, and you’re not going to know the answer until it’s there. That’s why a lot of companies are not going into that space where there are no data sets, because it can be challenging to do that.

Shawn Wang

I think a lot of smaller biotech AI startups will try this pattern. They’ll either start with a public, open-source data set, or they’ll try a pilot—incrementally collect a small amount of data and see if something works or doesn’t. Oftentimes, there’s almost a critical point where, below that, you’re just not going to get any signal. You have to have conviction that you need to collect up to a certain point before you start really driving something fundamentally valuable.

Dan Barr

Yeah, I mean, imagine trying to train a foundation model on not enough data.

Shawn Wang

Yeah. Yeah. But with your claim about AlphaFold, right? You have your GPT—GPT-3, GPT—you know, with GPT-1, 2, and 3, there was a clear progression there. With each one of them, you could see there was something which worked with scale, and there was this insight: Oh, we’re going to scale this up.

Sometimes with biological data, the process of collecting lots of data is just very expensive to begin with. You can’t just take something off the shelf and expect that you’re going to hit the threshold of GPT-3-like fullness.

Dan Barr

Fullness.

Shawn Wang

Yeah, yeah.

Dan Barr

So, yeah, totally.

Shawn Wang

It takes some conviction.

Dan Barr

It definitely takes conviction. I think it also takes a scientific belief that there’s a lot out there that we just don’t know yet, and that you’re not going to capture the biology you need to by having, right now, an agent that reads all of the biological literature, because again, that’s just a tiny slice of what’s out there.

I don’t know if it’s a great analogy or if I’m going to botch the history here, but in astronomy, it required Tycho Brahe collecting this enormous amount of astronomical data at his observatory. That was the substrate for Kepler figuring out the first laws of motion of the planets, and then that was superseded by Newton’s laws and so forth.

I sometimes don’t know how you even get started without this large repository of really high-quality data. Maybe there’s a tragedy-of-the-commons problem here of who’s going to generate that data and who’s going to capture the value of it, but I’m very glad that we’re taking that bet and seeing it pay off.

Shawn Wang

Yeah, this is not my expertise, but hypothetically speaking, how much of the PDB do you need to train?

Dan Barr

Some people argued that—yeah, and then you can get some pretty good models with, I think, 1% of 1%. There are people going back to the 1990s who argued that the PDB was already complete, in the sense that if you had a sufficiently smart algorithm, you could have done a pretty reasonable job at protein folding even back then.

Shawn Wang

Interesting.

Dan Barr

You don’t need a lot to get a pretty big boost, but the community was sort of independently collecting PDB data for quite some time without necessarily being convinced that this was going to lead to solving protein folding. Most of those structures were quite useful in and of themselves. Maybe that’s the counterpoint: Oftentimes, just knowing a protein was very helpful for us, anyway, with data.

We did see a transition from the early data. How many samples did we have? I’m guessing probably on the order of a few hundred before there was a huge bolus. There was definitely a moment very soon after I joined where the data set just kind of doubled in size overnight because there was a huge bolus, and the models immediately got a lot better at that point.

Now we run more controlled experiments: What happens if you train on 10% of the data versus 40% versus 100%? What happens if you hold out all of the pancreatic cancer or all of the breast cancer? We have a much better idea of what kind of diversity and scale we need.

If we were sticking to cancer, maybe we’re not that far off. If we end up generating a few hundred patients in a bunch of major and some minor indications, which we’re going to do this year, maybe that’s enough to generalize to kind of all cancer, because there is a lot of shared biology in cancer and immune cells across different tissues, mutations, and so forth. But if you think about all of the disease biology there is for a model to learn, maybe that’s another order of magnitude.

Shawn Wang

But even being able to solve all cancer biology would be pretty impressive.

Dan Barr

Yeah, to cure a cancer would be great.

Shawn Wang

Well, if it’s all cancer biology, it doesn’t take your cancer away. It also takes a different face.

Dan Barr

But yeah, at least if you go down and just take 1 drug, if you could look at 1 drug mechanism across the whole of oncology, that’s incredibly powerful. Imagine what Merck has done with Keytruda. Merck has run hundreds of trials with Keytruda—possibly even more than 1,000 trials—with different populations to find all these different indications: the subset of ovarian cancers, the subset of lung cancers, the subset of colon cancers.

That’s all been done by enrolling trials. If you can look at that biology from model embeddings and at least have a very well-defined starting point—if I’m going to run a trial, it doesn’t have to be as broad as it would need to be if I didn’t have any answer—then that can be a really powerful tool for a diversity of mechanisms.

Yeah, maybe just as a last point, going back to the virtual-cell hot takes: If your goal is to build an actual mechanistic model of an individual cell and then build up from 1 cell to an entire tissue, and then from tissue to patient and so forth, you might need a lot more data and a lot more data modalities than just gene expression or something like that.

We’re taking much more of a top-down approach. We’re trying to first solve the problem of what is determining heterogeneity among actual patients, and which of that variability is predictive of drug response. My intuition is that you don’t need to model the mechanism at the subcellular level necessarily to solve the problem of which patient should get which drug, or which targets are important in which patients.

I saw a similar debate play out in neuroscience and computational neuroscience, where for a long time people were really trying to build these biophysical models of individual neurons, and then they were going to stitch them together into models of the brain and so forth.

What actually ended up working in terms of building computational models of the brain and behavior is this abstraction: we're just going to treat individual neurons as linear-nonlinear units and put them together in neural networks connected by linear weight matrices. We stack a bunch of layers together and build neural network models of the brain that abstract away all of the biophysical details of what a neuron is doing.

Those are now by far the most predictive models of how a given neuron is going to respond to real-world stimuli in a real brain. I think my bet is that the same is going to be true for these models, too. Modeling at the level of functional tissue, where you have a bunch of cells interacting in a disease context, is going to get you to the problem of predicting patient-level behavior much faster than trying to first model a cell and then stitch a bunch of those cells together.

swyx

Yeah, that makes sense to me. It's a good analogy. I like analogies.

Do you have any call to action for the listeners?

Yeah. I would say, first, everyone should be excited about biology. Sometimes a lot of my hot takes on X recently are just that I feel like there's a huge amount of enthusiasm in the mainstream tech ecosystem, and people aren't really following a lot of what's happening in the biology space.

But at the same time, you're hearing, you know, an AI lab saying we're going to cure cancer. People should actually look at the folks working on curing cancer, working on aging, or working on other areas of biology. These are really exciting problems. There are real, significant ML problems in the space.

One call to action is that I would love for people to be more stoked about learning about applications of machine learning in the biological sciences and solving some of these hard problems, because I think these are the problems that are going to massively impact humanity in the next 10 years. This is really the very beginning. Maybe we're in the first inkling of the ChatGPT moment for bio, but it's very much just the beginning.

Yeah, in line with that, really dig in and learn more about the details. A lot of the time, it's presented as: We have these protein-folding models, we have these binding models, and we have AI-for-science agents that are reading all of the literature and automating these computational-biology workflows.

I think it's important to realize that there are a lot of problems in AI for biology, AI for biochemistry, and so on. Some of them are very important, but solving any one of those is not going to solve the problem of how we develop better therapeutics. We're focused on a particular slice of that process, which is translating things that we know work well in some patients into successful drug trials where we know exactly which patients to give them to.

That requires building foundation models at a particular level—the patient level—but people should not be under the impression that this is all going to be solved immediately because AI agents like LLMs are going to read the literature and figure out what the right drug is. There's a lot more data to generate, a lot more ML problems to solve, and a need to translate those methods into actual successful drugs. There are a lot of different places to contribute.

swyx

Lots to do.

Yeah, indeed.

swyx

Great. Thank you very much.

🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik | BidClub