🔬训练 Transformers,解决癌症试验95%的失败率——Ron Alfa & Daniel Bear,Noetik
Noetik 的核心判断是,肿瘤学90%-95%的临床失败率,主要是患者筛选问题,而不是分子设计问题。 Ron Alfa 认为,癌症试验往往不存在安慰剂效应,因此出现应答就说明相关生物学确实处于活跃状态;传统开发流程无法识别正确的患者亚群。Noetik 希望直接从人体肿瘤中发现具有治疗意义的患者亚型,再将其同时用于反向转化靶点发现和试验设计。
公司声称的护城河,不是事后拼装的公开生物学数据,而是一套专门构建、受控生成的人体肿瘤数据集。 Noetik 已为超过1亿个细胞生成空间分辨转录组数据,并与H&E组织学、蛋白成像和基因型数据配对;按其观察,这一数据集“至少大一个数量级”。仅使用40%或10%的数据训练,模型表现就明显恶化,尤其是在泛化到未参与训练的癌种时。
Noetik 用昂贵的多模态检测进行训练,但临床推断可以依赖几乎每位肿瘤患者都已经采集的标准H&E图像。 Travis McKie 将H&E称为病理学的“通用语言”:模型可以区分应答者和非应答者,预测局部表达的基因,并揭示超越单突变或单蛋白生物标志物的生物学。这为从历史试验走向前瞻性诊断提供了一条可信路径,并可用同一种低成本输入覆盖多种药物。
它的“virtual cell”刻意追求实用和自上而下,而不是试图模拟细胞内的每一条生化反应。 OctoVC 关注在特定患者背景下,T细胞、巨噬细胞、肿瘤或基因表达会发生什么变化;PerturbMap 则在小鼠肿瘤内通过约100种带条形码的基因扰动检验模型预测。其目标是在不先构建机制上完美的细胞模型的情况下,预测哪种药物适合哪位患者。
模型工作正从掩码重建转向自回归空间预测,组织上下文开始成为真正的扩展变量。 OctoVC 采用极高比例的掩码——按主持人回忆约为99%——以避免模型只是补全局部边缘。TARiO 的下一个token预测目标显示,更大的模型主要在更长上下文中胜出,暗示真正解锁复杂患者级生物学的,可能是更多周边组织,而不只是参数量。
5,000万美元的 GSK 协议是最明确的商业验证,但交易结构同样关键:这笔金额包括首付款和里程碑付款,另有独立的年度模型许可费。 GSK 获得了已在肺癌和结肠癌数据上训练的 OctoVC 模型,并可用自身的转化数据进行微调。Akash Tiwari 的关键表述是,这看起来像一笔有实质意义的生物医药业务开发交易,只是“交易底层实际上是模型”,而不是分子。
投资风险与护城河不可分割:Noetik 花了约4年搭建基础设施,在确认方法有效之前,投入更接近1,000万美元而非5,000万美元。 公司至少有18个月没有足够数据训练模型,单次转录组实验处理2张切片就要2周;它所说即将公布的试验重分析结果目前也尚未披露。Dan Barr 认为,每个主要癌种以及部分精选的小癌种各有数百名患者,可能就能泛化到整个肿瘤学,但他也明确保留意见:覆盖全部疾病生物学,可能还需要再增加一个数量级的数据。
1. 患者筛选,而非药理学,是被指认的失败环节
Alfa 以 Noetik 的逆向判断开场:“90%、95%的癌症药物最终都会在临床失败。”但与此同时,行业在药理学、靶点筛选和分子设计方面都已达到前所未有的水平。他认为,主要失败点出现在更靠后的环节:没能识别出哪些患者拥有药物可以利用的生物学特征。
他的证据,是藏在失败总体结果中的应答者。Alfa 表示,癌症试验往往不存在安慰剂效应;当患者出现应答时,说明药物确实产生了活性。试验之所以失败,可能只是因为这种生物学特征被稀释在一个不加区分、广泛入组的患者群体中。
因此,这个平台可以双向工作:既可以从人体肿瘤出发,反向转化寻找新靶点;也可以分析2期和3期活检样本,识别能够预测应答的生物学特征,再据此重新设计下一项研究。Noetik 表示,后一个方向已经投入了大量工作,但相关结果仍未披露。
Alfa 进一步推进了“亚型”这一判断:“实际上没人知道这些亚型究竟是什么。”一个长期被视为单一肺癌亚型的类别,可能包含3种功能上截然不同的疾病;这些潜在分层在治疗上的意义,可能超过病理学家沿用超过1个世纪的分类体系。
2. “Frankenstein 式”临床前模型让临床团队只能猜
Bear 的批评从永生化细胞系开始:许多细胞系已经存续40年或50年,往往带有异常的染色体数量和表达程序,与任何可识别的人体细胞都不相似。研究人员仍然把它们标记为结肠癌或肺癌细胞,但 Bear 称其为“Frankenstein 式细胞”(“Frankensteinian cells”),因为它们在实验中的反应往往无法映射回患者。
把这些细胞移植到动物体内,也没有弥合这一鸿沟。肿瘤学研究通常会把它们植入皮下某些“奇怪的位置”,再通过合同研究组织测试数百个细胞系,并根据哪些名义上的结肠癌或卵巢癌细胞系出现应答来推断适应症。即使源自结肠癌的细胞系,也可能缺少人体结肠肿瘤中的典型突变。
由此形成的临床链条极其残酷:在缺乏患者层面临床前指引的情况下,团队可能把所有符合条件的肿瘤患者都纳入一项约50人的开放标签研究,同时摸索剂量、安全边际和疗效信号。假设肺癌包含10个相关亚型,那么最终只有少数患者应答在统计上并不意外,但分子仍可能因此被终止开发。
主持人的总结指向了拟议中的解决方案:不要再假设今天的适应症标签内部同质,而要找出真正会应答的患者亚群。Bear 补充称,现有生物标志物——单一突变、单一染色蛋白或单一基因特征——都“偏向简单化”,与临床成功之间通常只有很弱的相关性。
3. 生物学训练数据必须先被设计,之后才能扩展
Tony Bui 不认同有用的生物学语料库可以事后简单搜集。Protein Data Bank 是经过数十年有意积累形成的;Shawn Wang 则以 ImageNet 精心整理并标注的120万张图像为例。核心教训是,“必须有意识地设计要生成的数据”,并提前考虑它需要支持哪些模型。
Bui 将规模视为“必要条件,但未必充分”。语言之所以展现出惊人的规模效应,部分原因在于语料库足够大;但数千小时的视频并没有自动产生同等水平的能力。生物学还要面对数万种基因和蛋白质、它们的空间组织、健康组织、其他疾病,甚至其他物种。
Shawn Wang 提出相反的判断:既然生物学源于底层物理过程,也许这个空间会比语言更早进入模型的分布之内。Bui 的回答仍然保留余地——他的“直觉是生物学相当复杂”,Noetik 距离覆盖全部生物学还很远,而他的结论是“我不知道”。
蛋白质折叠提供了另一重校验:讨论指出,优秀模型可能只需要使用PDB中很小一部分数据;也有研究者认为,只要算法足够强,PDB在1990年代覆盖的数据就已经足够。Dan Barr 的范围更窄:主要癌种以及部分精选的小癌种各有数百名患者,可能就足以广泛泛化到整个肿瘤学。
4. Noetik 把组织、细胞、分子和批次放进同一个体系
Tony Bui 在 Recursion 工作6年的经验影响了实验设计。图像信息密度高,可以在一张切片上容纳许多患者;与测序相比,图像还能降低边际数据生成成本,因为测序中每增加一次实验,往往就对应另一组患者样本。
临床锚点是H&E:苏木精和伊红染色形成病理学家熟悉的紫粉色组织对比,用于对几乎所有切除肿瘤进行分类。H&E 能捕捉组织结构,但无法可靠识别所有相关细胞类型,因此 Noetik 又加入多重免疫荧光,以识别B细胞和肿瘤微环境中的其他组成部分。
空间转录组学提供分子层数据:在同一批细胞的已解析位置上检测约1,000至19,000个基因。探针会结合各类RNA,仪器再持续数周循环完成检测;讨论将最终输出比作一张拥有“20,000个颜色通道”的图像,而不是RGB的3个通道。基因分型则补充底层DNA变异信息。
批次控制直接嵌入物理样本设计。Noetik 会对每个肿瘤采样数十次,将数百份患者样本随机分配到不同阵列,并让每位患者出现在多张切片和多次处理流程中。这样,研究人员可以判断患者嵌入反映的是免疫疗法应答,还是仅仅反映“染色批次”。
5. Virtual cell 是制药启发式工具,不是完整细胞复制品
Noetik 区分了两种定义。完整的 virtual cell 需要在受到任何外部信号后,模拟细胞内数百万条化学反应;讨论认为这在理论上有意思,但现有数据模态无法解决。Noetik 追求的则是“一种对制药有用的启发式工具”。
按讨论中的描述,目前多数 virtual cell 工作是在培养细胞中,预测小分子或CRISPR扰动后的转录组变化。问题在于转化价值:如果终点是患者体内会发生什么,那么“对来自患者的数据进行建模”更有可能成功,而不是去模拟一种体外抽象。
因此,Noetik 的模型通过自监督学习基因、蛋白质、细胞、组织和患者背景。团队有意没有大量依赖电子健康记录,因为他们不希望模型表征被医生在当时掌握的信息以及最终记录下来的内容所限定。
6. 一张H&E图像可以同时支持药物发现、试验挽救和诊断
在分子进入患者之前,Noetik 可以在肺癌、结肠癌、卵巢癌或更广泛的肿瘤学患者队列中模拟其靶点。结果可能是否定最初的适应症选择:一个原本为肺癌设计的靶点,可能在肺癌中生物学意义不大,却与卵巢癌相关。
OctoVC 可以询问,在特定肿瘤微环境内,T细胞会表达什么;也可以模拟移除某个靶基因或靶蛋白后的结果。有效输出包括免疫功能增强、肿瘤生长减缓,或其他被认为与临床成功相关的反应。
最清晰的回溯性用例从已接受治疗的患者队列开始:如果应答者全部位于某个自监督患者聚类,而不是另外9个聚类,那么这就直接给出了入组假设。训练需要完整的多模态数据栈,但推断只需要H&E,包括多年以前完成的试验所保存的数字化图像。
模型还可以仅凭H&E预测基因在何处表达。如果应答者组织中富集了某种药物的蛋白靶点,就能提供一重可解释性校验;周围的多基因模式则解释了为什么单蛋白生物标志物会漏掉应答者。Noetik 正在合作项目中使用同一输入,包括已宣布与 Agena 的合作,而最终诊断产品是顺理成章的终点。
7. 规模、架构和可解释性共同强化数据护城河
学术界的配对数据集可能只有100名或几百名患者,空间转录组学数据规模往往还低于这一水平。Noetik 表示,其拥有超过1亿个空间解析细胞,全部与H&E和蛋白成像配对,按其考察结果,规模至少比替代数据集大一个数量级。
内部消融实验支持规模效应:把训练数据缩减至40%或10%,模型表现都会“差很多”,尤其是在泛化到训练中排除的癌种时。数据集在早期突然翻倍,也带来了即时的性能跃升。
团队并不认为数据量本身就足够。Noetik 正在构建定制的多模态架构和自监督目标,针对患者差异、基因或蛋白质反事实以及可读的生物学输出进行训练。他们称之为“世界模型”,因为任务是预测某个动作之后会发生什么,而不只是对图像进行分类。
8. PerturbMap 把人体预测重新接回具因果性的动物实验
讨论触及了 Noetik 以人体为先的方法所面临的张力:公司承认仍在使用小鼠和注射细胞。人体数据应当用于训练主模型,但FDA 可能仍要求证明新机制在动物系统中有效;如果没有有意搭建桥梁,开发者最终只能“倒推回这个系统”。
PerturbMap 构建约100种不同的CRISPR敲除,每种都携带组合式蛋白条形码,再将它们共同注入小鼠肺部。最终得到的可以是数百个肿瘤,同时保留可读取的肿瘤身份和空间生物学,而不是把单一细胞系植入皮下、充当人体多样性的替代品。
Noetik 可以将基因与免疫冷或免疫热表型建立因果映射,再叠加药理学;团队描述过类似“50种敲除对应50种药物”的面板。缺乏免疫浸润的人体肿瘤,也可以通过基因方式在小鼠中重建;这些小鼠肿瘤同样缺少免疫细胞,也不会对免疫疗法产生应答。
更令人意外的一步是“在计算机中将小鼠人源化”:在人体H&E和空间转录组学上训练的模型,直接运行于小鼠H&E,并输出人类基因预测。已知的抗原呈递敲除会正确呈现为免疫冷表型,同一信号通路中的多个基因也会产生相似的推断表型。团队承认,面对真正全新的生物学机制,结果仍然存在不确定性。
9. 只有看见足够多的组织,自回归才会扩展
OctoVC 采用掩码自编码:把每种模态切分成token,隐藏其中大部分,再根据剩余内容重建基因表达、蛋白质区块或组织学区块。主持人回忆掩码比例约为99%,团队回应称这与其方法一致。
极高比例的掩码是有意设计的。如果只隐藏10%,模型可以通过延伸附近边缘等“无聊行为”完成任务;隐藏更多内容,则迫使模型学习蛋白质之间的相关性以及组织的整体结构,而不是依赖局部连续性。
TARiO 将目标改为自回归的下一个token预测,本质上是一种结构化掩码,与LLM扩展背后的损失函数相似。这并不是 Noetik 的第一次尝试,但在空间转录组学数据上,它让更大模型和更长上下文带来的收益变得更加清晰。
微妙之处在于,更大的模型主要在上下文更长时才有帮助——也就是说,它们一次可以看到更多组织。低表达但具有预测力的基因可能是部分原因;不过,在相同上下文、对比更小与更大物理区域的受控实验中,更大的区域同样更有利于大模型。
10. GSK 获得的是模型许可,而不是分子资产
Noetik 已宣布与 GSK 达成协议,授权其在肺癌和结肠癌上训练的 OctoVC 模型。5,000万美元的 headline 金额包括首付款和里程碑付款,年度模型许可费另计;Tiwari 还提到,首付款和近期经济收益达到数百万美元级别。
GSK 可以用这些模型进行模拟和治疗发现,再用自身“堆积如山”的转化病理学数据和临床试验数据进行微调。这样,共享基础模型就能变成更接近 GSK 专有版本的工具,而药企不必先统一所有数据孤岛,也不必自行构建基础模型。
Tiwari 认为,药企的需求正从围绕单个项目定制的合作,转向覆盖数十个管线项目的访问权限。这笔交易看起来像传统的行业业务开发交易,但其结构性突破在于:“交易底层不是分子,交易底层实际上是模型。”
11. 这条护城河要求公司在任何学习曲线出现前就下注
Tiwari 表示,Noetik 开设实验室、购买仪器、获取人体肿瘤样本,并以每次2张切片、耗时2周的速度运行转录组实验,而当时“没有任何迹象表明这一切会奏效”。他的总结更加直接:“一个大大的零。一个大得离谱的疯狂下注。”
公司至少有18个月没有足够数据训练模型;单次转录组实验处理2张切片就要2周。Tiwari 表示,投入“更接近1,000万美元”,而不是5,000万美元。当时不存在显而易见的空间数据现成方案,AI团队只能从第一性原理出发,摸索这一片“陌生的数据大陆”。
他们给其他初创公司的建议是:从机器学习问题出发,再反向设计数据集。“任何数据集”都不会自动对应一个有价值的机器学习用例;小规模试点可能因为低于临界规模而失败,即使完整实验最终会成功。“这件事没有捷径。”
Noetik 最后的下注是自上而下的抽象化。正如简化的神经网络比 painstakingly 拼接的生物物理神经元模型更早预测真实大脑反应,功能性组织模型也可能比自下而上的生化模拟更早预测患者应答。这或许是“生物领域首次隐约接近 ChatGPT 时刻”,但仅靠阅读文献的智能体,不会取代新数据、新机器学习方法和临床转化。
We basically opened the lab, hired a team, got all the instruments, and started sourcing tumor samples. There was no prior art here indicating that any of this would work.
Brandon Allgood
Big zero.
We just started generating data and sourcing human tumors. We built this whole processing pipeline to get the tumors into these arrays and formats. You have two-week runs where you're processing 2 slides, and we're just churning data for months.
We couldn't even train a model. We sort of just built all this, and then, let's say, 18 months later: “Hey, Alondra, can we train a model off of it?” It wasn't obvious.
Dan Bear
There wasn't really anything major to go off of. There were transformers developed for single-cell data, but there just weren't really datasets out there that people had been able to develop on. We do a lot of custom model building.
Hi there, I'm R.J. Honnold and this is Brandon Anderson. We're the co-hosts of the Latent Space Science podcast and today we're really happy to be in the studio with some of the people from Noetic.
I'm Ron Alfa, co-founder and CEO of Noetic, and a physician-scientist by training. My hobbies are making hot takes about AI curing cancer.
Hi, I'm Dan Bear. I'm VP of AI at Noetic. I'm a biologist by training and did PhD work in neuroscience and then moved into comp neuro, computer vision, self-supervised learning and have, you know, been doing AI research at Noetic for the past few years.
Brandon Anderson
Maybe we should start with what Noetic is, why you founded it, and what the difference is between Noetic and the other virtual-cell companies.
Maybe just start with a little bit of a contrarian thesis, which is really the reason for founding Noetic. We all know the numbers: 90% to 95% of cancer drugs fail in the clinic. Why do they fail?
Our thesis is that they fail not because we're bad at pharmacology, not because we're bad at target selection, and not because we're bad at making the drug. We're actually better at that process than we have ever been in the history of drug development.
Most of those drugs fail, we'd argue, because we're bad at selecting which patients those drugs are going to work in. Oftentimes, you see trials where there is no placebo effect in cancer. Some patients respond to these drugs, and if you have a patient that responds, that tells you that there's some active biology there.
But you have a problem in patient selection. That's really the thesis behind why Noetic exists: Can we build models that can fundamentally understand patient biology from the very beginning and help you position molecules in the right patient population?
R.J. Honnold
So you're actually using the models, at least partly, to select the patient cohort, not just so you can imagine it working either way. You could design it one way: “I think that this molecule will do well because I know something about the patient population.” But you could also say, “I think that this patient population is the match for this molecule.”
That's where the power of the models is. Once you've trained these models on patient data, you can use them on both sides of the equation.
You can use them for discovering new targets directly from the patient data, which people often refer to as reverse translation: starting from humans and then trying to understand which targets to go after. Then you can use that to develop molecules.
But you can also use them directly on patient data. If you have, let's say, a phase 2 or phase 3 trial, you can use these models to understand which patients, or what underlying biology of the patients in the trial, is a predictor of response. We've been doing a ton of that recently.
R.J. Honnold
Are you doing a lot of restudying trials that had a bad effect?
We are doing a lot of looking at data from phase 2 and phase 3 trials, and then using the models essentially to run inference on patient biopsies and understand whether there's underlying biology that would help us design the next trial. We haven't shared any of that yet, but you'll see this soon.
Brandon Allgood
Cancer is kind of infamous in that there are many, many different types of cancer. Whenever someone says “cure cancer,” that's almost a meaningless, vacuous statement.
Your point is that, even among cancers—or if you pick a specific type of cancer, and then a subtype within a subtype—there are a bunch of different patient populations, and each one of them will respond differently to drugs. Your point is that you can figure this out right now: Some subpopulation will do well and respond to this drug, while generally speaking, the rest of the population will not, even though historically we've classified this as one type of cancer, one indication, or so on.
Yeah, that's exactly right. I would maybe even go further and say that nobody actually knows what the subtypes are. There are cancers that originate in a certain tissue, like the lung, that have been classified into subtypes based on pathologists looking at them for more than a century. Those subtypes certainly have some connection to the real, like, carving nature at its joints—what are the actual functional subtypes of disease there?
But our thesis is that if you look at much richer data—the multimodal data that we're generating in our lab—we're going to see that what people thought was one subtype of lung cancer is really 3 distinct subtypes of cancer. That is going to be critical for figuring out which patients should get which drugs.
Dan Bear
Maybe I'll just go back to one of your first questions. You were asking why drugs fail in patients, and I was saying that many drugs fail in patients because we don't understand which patients they will work in, in oncology. Why do we end up in that situation?
Whenever you make a new drug, you do a set of experiments in cell culture—cells in a dish. Those cells are often cell lines. These cell lines have existed for 40 or 50 years, and they're immortalized. They have genomes that allow them to persist, with abnormal numbers of chromosomes. They have gene-expression patterns that don't represent any known cell in the human body. These are sort of Frankensteinian cells. They can't be derived or analyzed in that way. They're mostly cancer.
You can do your experiments in these cell lines in a dish, or you can move them into animal models. In oncology, you often have a panel of different animal models, with different cancer types that you'll test these in.
In doing these experiments, we sort of convince ourselves that some of these cell lines are, let's say, lung cancer cell lines or colon cancer cell lines. Then, even in the mouse context, we say that some of them are colon cancer cell lines and some are lung cancer cell lines. In the mouse, we implant them under the skin in weird places, treat the mice with drugs, and see how they respond.
Ultimately, there's a big gap because they don't translate to patient biology most of the time. Even if these cell lines were derived from a colon cancer, most of them don't have the mutations that human colon cancers have in many cases.
Pharma has done this for 20 or 30 years. You develop a drug, test it against hundreds of these cell lines, and it's not a hard experiment. You can send this out to any CRO. They'll test your drug against hundreds of different cancer cell lines, and you can sit back and say, “Okay, which of the 50 colon lines responded to my drug, and which of the 50 ovarian cancer lines?”
You could try to map that to human biology, but the problem is that cell lines, as an abstraction, do not relate in any way to your patients. Ultimately, no matter what you do preclinically, the molecule gets into the clinic and the clinical team says, “Look, we don't really know how to design this trial because none of the data that you've produced gives us any insight into which patients to run the trial on.”
So we're going to run an open-label study. We're going to enroll all tumors—all patients who are enrollable in the trial—and see where we get signal.
Imagine doing that in an early-phase trial where, let's say, you have 50 patients and you're trying to test different doses. You don't really know the dose of the drug, you don't know what the safety margins are, and you're also trying to figure out where your signal is.
What if I told you that, let's say, in just lung cancer, hypothetically, there are only 10 different subtypes of lung cancer, and you don't even know if it's lung. It could be any. This is what happens. Oftentimes, you get to the end of these early-stage trials and don't see very many responders, as you would expect statistically, and then these molecules get canceled.
R.J. Honnold
So you're imagining that, with your Noetic system, you help the pharmaceutical company characterize that they expect people with a certain genetic profile, or even transcriptomic profile, will respond to this drug. Then you sequence the patient and say, “Yes, this is a match,” or no. Is that the sort of grand vision?
I would say we're even less biased than that. We're saying, “Okay, we want the model to learn, let's say, from lung cancers. We want the model to learn how many different therapeutically relevant subtypes of lung cancer there are, just from self-supervised learning on the data.”
Dan Bear
Those subtypes could be driven by large genetic changes, immune changes, or really any biology that the model is learning in the process of training. We do see different types. Feel free to contradict this as the actual doctor here, but the biomarkers that people have been using are biased toward simplicity: does the patient have this particular mutation, sometimes a stain for this single protein, or transcriptomics to look for a particular gene signature? There’s no reason to think that biology, or the biology of cancer, is so simple that you’re going to capture most of the meaningful variation with such simple biomarkers.
Most of them have weak correlations with clinical success, but the hypothesis here is that, again, if you were to carve nature at its joints and figure out what’s really going on, there are these 5 subtypes, and the correlation between which patients you give a particular drug to and whether you have success is much, much stronger than if you’re forcing yourself to go with these very simple biomarkers.
Alessio Fanelli
You mentioned the lab. You do a lot of data generation in the lab. Why do you think that, versus using existing public repositories or whatever, is appropriate?
Tony Bui
We generate all our data in the lab, from sourcing tumor samples themselves to processing them and generating the data. Maybe another hot take I have, just in AI and bio, is that you’re not at the order of magnitude of data that you are in other spaces for building training models. It becomes really hard to brute-force these problems just by collecting data.
We have a couple of pretty good examples of where someone has designed a dataset. PDB was designed and has been built over the past 50 years or so, and it’s not an accident that that dataset exists. Someone decided that we were going to design this dataset and collect this data over decades and decades, with the intuition that potentially this would help solve protein folding down the road—and it did. It’s not just that PDB is a bunch of random data that people have organized from the web.
I think that in bio, you really need to be intentional about the data that you generate and how you generate it, and have some foresight around what models we’re going to want to train and what modalities we need to learn from the very beginning. So that’s why we’ve taken this approach.
Shawn Wang
A good comparison is the ImageNet dataset, which kicked off the deep-learning revolution in computer vision with convolutional neural networks, actually demonstrating that neural networks can do better than other methods on object categorization. ImageNet is, at least the part of it that people were developing models on, 1.2 million images, very carefully curated. These are high-quality images, not random images from the internet or multiple datasets cobbled together.
Tony Bui
And labeled.
Shawn Wang
Yeah, and labeled. I think with the data that we’re generating, we’re around that scale right now, but, of course, people have gone much, much larger in image datasets and language datasets—text datasets, obviously, for LLMs.
Tony Bui
We think that we need to get the data up to that scale before we can really see meaningful progress on the algorithm side.
Alessio Fanelli
The scale of language data.
Tony Bui
Language is really the only modality where people are seeing these very impressive scaling results. Part of that has to be just the scale of data that’s there and that the models are trained on. That can’t be the only thing, because there’s a lot of video data as well. People are training on thousands of hours of video data and haven’t seen any of the scaling results that you have in language modeling, but having the right scale of data is necessary, if not sufficient, to really make progress here.
Shawn Wang
Can I offer a contrarian take to that? There’s this whole concept of the giant frontier of LLMs, and in certain regions it can be really good at solving some problems and then remarkably stupid at solving nearby problems. Maybe the argument is that a lot of these frontier models are just becoming massive; everything is becoming in-distribution.
If everything starts out OOD, and you just get more data, that becomes in-distribution. Is it possible that for biological systems, because there are underlying physical processes here, you can basically make things more in-distribution earlier and that you can’t actually cover the space? I kind of have some thoughts with PDB, but maybe I’m just curious at this point.
Tony Bui
I think it’s a good question: how much data and what kind of diversity do you need in biology to solve, say, the drug translation problem—figuring out which drugs are going to work in which patients? My intuition from working in biology for a while is that we’re still pretty far from that, because we’re building datasets that are focused right now on cancer, and I’ve generated data from thousands of patients in a few major cancer subtypes. But there’s every other disease, healthy tissue, and even other species.
There’s a lot of biology to learn, especially if you think about it as having to learn the spatial and functional patterns of tens of thousands of genes, tens of thousands of proteins, how their spatial arrangement contributes to the function of organs, and so forth. My hunch is that biology is pretty complex and that we still need to generate a lot more data, but I don’t know.
Alessio Fanelli
Yeah. But as a cancer company, do you think you could actually do this hypothetically for cancer? I mean, for at least some subclasses of cancer?
Tony Bui
Definitely. I think we’ve done experiments that suggest that if we can generate data from several hundred patients in all of the major cancer indications and some of the less major indications, that will result in a model that can generalize pretty well to any type of cancer we would throw at it.
Alessio Fanelli
Backing up, what is the data you’re collecting? My understanding is that you use some pretty specialized instruments and gather very specific datasets. How did you come to that decision about how much data to collect, how much to spend on it, and what types of data?
Tony Bui
I’ll give a hat tip to my previous employer, Recursion. I spent 6 years at Recursion, from the very beginning, and a lot of what we were doing in the early days was figuring out the things we didn’t understand about the datasets and what the problems would be in the dataset: batch effects, control design, orientation of samples on plates, things like that.
Flash forward to founding Noetik: I started the company already with some principles around how we should think about building the dataset. What are some things that we know mattered? For example, over many years we learned that images were actually a really powerful dataset for machine learning, for many reasons. One is their scale. We can put patient samples on slides, and on a single slide we can capture many patients’ worth of data.
The images themselves are very rich sources of biological information. Beyond that, we have a very information-dense modality, and we can decrease the cost of data generation, so then we can increase the amount of data generation over the whole dataset. That’s always been a really big benefit to image-based modalities over, let’s say, sequencing, where every time you run a sequencing run, your run is a patient set, for example. That was one way to think about it.
The other was: how do we design these datasets so we can control for things that we know are going to be important, such as batch effects? For example, if I have a slide and we do, let’s say, a spatial transcriptomic run on that slide, you stain the slide, do a bunch of wet-lab processing, put it into a machine, and get data out. If you do that on 2 different days, there are going to be different variables that impact the data quality. That’s going to be a large source of variation in datasets.
You want to be able to control for things like batch effects. Really, you want more patients represented on multiple different slides so you can process them in different batches. You want to be able to control for things like this so you can go downstream and look at the data and say, “Okay, once we have, let’s say, patient-level embeddings, we can ask: Do the patient-level embeddings represent patient response to immunotherapy, or do they represent staining batches?”
Alessio Fanelli
So you’re actually taking 1 patient and spreading it across multiple slides so that you can get a—it’s sort of a calibration across the slides?
Tony Bui
Yes. Our data looks very different from anyone else in the space of generating data on histology or digital pathology types of specimens. We receive a sample, sample those samples dozens of times to build these arrays, and each array has hundreds of patient samples randomized. Every patient is represented on multiple different arrays, and so we’re getting a lot of different representations of each patient that we’re sending through the data-processing pipeline.
And then that lets you downstream answer some of these questions and control for some of these variables.
Shawn Wang
You mentioned some terms I just wanted to define for people. What is spatial transcriptomics?
Tony Bui
Yeah. This was your first question: What are the data types? If you sit back—and this is not my background in terms of spatial biology—again, everything we did previously was cell biology in a dish. If you said, “Okay, I want to train a foundation model that understands human biology,” what does that mean? How would you go after that problem? That was really the starting point for the company: “From first principles, how would we do this?”
You probably want tissue-level biology. You want to understand tissue: cells are organized into tissues. You probably want some modality that is relevant in clinical use, so you can relate clinical data to what your models are learning. That’s why we generate pathology H&E. That’s what every patient gets: a tumor removed, and then it gets stained with H&E.
I can explain what H&E is: it’s basically 2 different dyes, hematoxylin and eosin, and it really just creates a contrast over the tissue. You’ve probably seen these purplish pathology specimens. Pathologists can look at those and identify different cellular structures, and they use those to classify tumors based on the classical classifications—adenocarcinoma, small-cell carcinoma, things like that—based on cellular structures.
Shawn Wang
Okay, so there are specific patterns that show up when you add these 2 stains, and it is well established that you classify tumors based on those patterns?
Tony Bui
Based on pathology, yeah. Basically every tumor that gets processed in the hospital will get this H&E stain, and it’s how the pathologist typically classifies a tumor. So that’s the first level. You probably also want to understand cell types.
It’s really hard to understand cell types from just that stain because it doesn’t reveal that much that a human can use to classify cell types, at least. You can say, “Well, I want to know whether there are immune cells and different subtypes of immune cells.” We want to have some layer of cell biology.
swyx
Okay, so you want to know about immune cells because you have these cancer cells, and oftentimes the immune response dictates whether or not you’ll have an effective treatment?
Aviv Regev
The immune environment of the tumor will be a core—we know it’s a core constituent of whether a patient is going to respond or not. So you want to give the model this tissue-level information. There’s not enough cell-level information in there for the model to learn enough cell biology about different subtypes, so we also want to present it with some cell-level information.
We use protein stains, so standard immunofluorescence. You basically use antibodies against a small set of cell markers to label your different key cells—B cells and the standard subtypes of cells in the tumor microenvironment.
swyx
So, in this stain, just for those who are familiar with it, the antibody has a fluorescent protein. When you hit it with a certain frequency of light, it fluoresces, so you can tell the antibody bound to a certain protein, and now it has a fluorescent protein attached to it.
Yep. And in terms of the data, from the tissue layer you have an RGB image. From the next layer, you have a multichannel image, with each channel representing, let’s say, 1 color. For example, certain immune cells are each in a different channel, so you have this multichannel image.
That’s great: we’ve got tissue, we’ve got cells. But if we actually want to make drugs, we need some type of molecular information. We need to tie all of this down to what’s happening in the genome. What is the cell doing? What are the mechanistic principles of the biology?
So then we get spatial transcriptomics. That’s spatially resolved RNA: DNA is transcribed into RNA, which is translated into proteins. We get basically the RNA in a spatially resolved pattern for the same cells that we’re seeing in all of these other layers. Now you have between 1,000 and 19,000 different genes. Again, these are all image layers, with spots showing where those RNAs are and in which cells.
swyx
And this one works a little bit similarly to how we talked about protein, where you have a segment of RNA and then you have a fluorescent protein. Usually there’s some sort of combinatorial thing, so if you see these 4 colors at this amplitude, that means this gene because they’re right next to each other or something like that.
For the detection, you’re basically binding a probe to each one of those RNAs, and then you’re cycling it. It takes weeks to run 1 of those assays. The machine will cycle across each species, they’ll amplify, and you’ll get a signal for each RNA species.
At this point, you now have this very rich data layer where you have the tissue, the cells, and the molecular information. You can use all of that to train the model. We think of it as, essentially, the central dogma, if you will. We also have DNA; we genotype just so we understand the genomic alterations in these tumors.
swyx
All right. So you have this stack of images, basically, that you can train models on, with an understanding of the expression of genes and the proteins that are being expressed at the time that the sample is taken, all in the image information. Then you can train your models with that.
Yeah. Spatial transcriptomics is particularly dense because if you think, let’s say, there are 20,000 genes in the genome, we’re running assays that are detecting nearly all of them in a single sample. You can think of 1 of those data points as an image, except instead of being an RGB image that has 3 color channels, now all of a sudden it has, like, 20,000 color channels.
It’s a very meaty computer-vision problem to try to look at those data and figure out what makes patient A different from patient B, and then go from that to which drug is going to work in which one.
swyx
So you have a hot take about virtual cells. I want to understand how—okay, so you have this big pile of data: every single sample has a massive data set with it, and then you have many, many samples. How do you turn that into useful knowledge?
Maybe just: What is a virtual cell? Everyone is always asking that question. I think there are really 2 ways to think about it. One is that we want to be able to simulate all the biochemical processes in a cell.
We want to have this comprehensive foundation model where we understand that if some signal from outside the cell interacts with the cell, then here are the millions of intracellular chemical reactions that are going to happen, and you could predict them from the model. That’s 1 view. I think that’s interesting; it’s an interesting intellectual pursuit. I don’t think we have all the modalities of data that you would need to solve that problem.
I tend to see the virtual cell problem as something more practical. We’re trying to make drugs that work in patients. From a virtual cell perspective, what we really want to do is understand cell biology in some heuristic that’s useful for making drugs. The heuristic could be either a way to understand drug targets or a way to map your cell-level biology up to patient-level biology.
The way we design these first virtual cell models is really just to simulate the biology of a cell in some context. The biology of that cell is, let’s say, the cell being in some context, with the output being the transcriptome in that context, or the protein in that context. These types of input-output relationships allow us to design experiments.
The very simplistic thing that we’re doing is that the model can simulate the biology of a cell—or many cells—in different contexts and allow you to run some simulations in that regime.
swyx
Yeah. I think most of the things that people are calling virtual-cell models right now are focused on single-cell gene expression, so transcriptomics data, RNA data, and they’re largely geared toward predicting what’s going to happen to the transcriptome—the set of genes expressed—when you hit cells with either a small molecule or drug, or a genetic perturbation. Typically, these are cells grown in vitro, either in cell culture or as primary cells, something like that.
The genetic perturbation is where I knock out a gene or add a gene and see how that impacts the expression of the various RNA.
Exactly. I think my view, and I think Ron shares it, too, is that that may be of interest in some cases, but the problem we’re really trying to solve is predicting what’s going to happen in a patient. Just modeling data that comes from a patient is, in my mind, much more likely to translate to what happens when you give a patient a drug than something that’s happening in cell culture.
swyx
Is there other clinical data that you’re pulling into the model besides the actual—so you’re calling it the context of the cell, just the surrounding cells—but is there other “this drug caused a bad reaction” kind of stuff?
Yeah, we’re pulling in data from the entire patient, not just the very local neighborhood of the cell.
So far, we haven't done much integration of electronic health records or other information that one could get about the patient, and that's pretty intentional. We really want these models to learn basic biology—again, the central dogma, but not just the central dogma: the basic biology of genes, proteins, cells, and tissue—in a self-supervised way.
We want to do that purely from the data that we're generating and not be biased by what the doctor wrote about that patient. Our thesis is a bit like this: most of the therapeutically predictive and important information is not contained in the very small number of patients who've been treated with a given drug, or in whatever the doctors thought was important to write down given the state of knowledge at that time. It's much more about trying to discover what's really there in patient biology than going based on the text that people have written about it.
Alessio Fanelli
So, you have this self-supervised model. You use a lot of data. You have essentially some clusters of patients now. How do you translate those clusters of patients into making decisions? You go to a pharma company and say, “We can repurpose this drug,” or, “We can suggest this subtype should be the focus of your phase 2 trials.” What is the process for that? What data do they need to provide you, and how do you translate your models?
Aviv Regev
It depends on what the problem is. I think it's important to note that one of the more interesting aspects of these models is that they're useful for a broad array of use cases, as we were talking about from the very beginning.
You, as the pharma company, could say, “Okay, well, I have this molecule, and the target of the molecule is X, and I want to design my clinical trial. The molecule has seen 0 patients so far. All I know is the target and some biology around the target.” We can run simulations using the models and our cohorts of patients.
Let's say, if we were to look in lung cancer, we can run simulations around the target and ask, “Which sets of patients here would this target be important in?”—across a cohort of lung cancers and colon cancers, or across all of oncology. You might see—and we see this sometimes—that your target probably isn't something you want to put in lung cancer. Maybe you want to put it in ovarian cancer because it's not really important in lung cancer.
Alessio Fanelli
What are you simulating here? Are you saying that this drug is expected to knock down this gene, and therefore you want to look for clusters where knocking down this gene inhibits tumor growth rather than enhancing tumor growth?
Aviv Regev
That's certainly 1 way we could do it. There are other types of simulation where you might just want to ask: if there were an immune cell here, like a T cell, which is responsible for actually killing tumor cells, what would happen to it? What genes would it express, or what proteins would it express, in this particular patient's tumor microenvironment?
Travis McKie
That's what we've called these virtual cell simulations. We have a model called Octo Virtual Cell that does this, and that can give quite powerful answers to the question of whether these drugs are going to work in these patients. You might find, actually, as Ron was saying, that the thing this drug targets is just not important in this particular patient's tumor, and it's not going to have any effect on the T cells, the macrophages, or some other cell type there.
Then there's the type of simulation you alluded to, where you can ask the model, “What would happen to this patient's tumor if you were to knock down this particular target gene or its protein product?” You might be looking for cases where the model predicts that removing that gene or protein is going to have a large effect—either increase the immune system's function, its ability to fight that tumor, or decrease the tumor's ability to grow—or some other readout that you think is correlated with clinical success.
I just want to call out that perhaps the simplest use case is the one where there's a company that has a drug, they've given it to some patients, and we know some of those patients responded. It becomes a question of whether the space of patients that the model has learned via self-supervision tells us that all of the responsive patients are in 1 of these clusters and not the other 9 clusters, or something. If we know that, then there's a pretty straightforward hypothesis that this is the right cluster.
Alessio Fanelli
So, that's the scenario where you would sequence something. What would you collect about those? You have a cohort that responded and 1 that didn't.
Travis McKie
This is getting back to something Ron mentioned earlier, which is this type of data called H&E. It's the standard pathology stain that makes these pinkish- and purplish-looking images.
Right now, we've built models that are trained on all of the multimodal data we generate, but once they're trained, at inference time, all they need is an H&E image. That could be something that we generate in our lab, or it could just be a digital image that they have from a trial that was run years ago.
The reason that's so powerful and flexible is, again, because H&E is the lingua franca of pathology, and especially oncology. Almost every patient who's been given a clinical-stage drug is going to have that.
Alessio Fanelli
You can look at the 2 cohorts—the responders and the nonresponders—and say, “These H&Es live in this part of the latent space, and these H&Es do not.”
Travis McKie
Exactly. One way we've gone further than that is that, given the H&E, the model can say, “I predict that these genes are expressed at this location in this patient.”
Not only do we have these clusters, these embeddings that say all of the responders to this drug are over here and all of the nonresponders are over there, but we can actually see, “Okay, for the responders, these are the genes that are expressed much more highly—or predicted to be expressed much more highly—in the responder cluster versus the nonresponder cluster.”
That adds a major level of interpretability, because we can see things like, “Good, the responders are actually expressing the protein target of this drug.” We would be worried if that weren't the case, but we can see that it is. On the other hand, we also see that the biology is very complicated. That's part of explaining why these simple biomarkers, like looking at a single gene or a single protein, just really don't capture what is predictive of therapeutic response.
Alessio Fanelli
Yeah, so I have a million directions I want to go here. The H&E actually gives you a pathway to a diagnostic then as well?
Travis McKie
Exactly.
Alessio Fanelli
Yeah, right.
Travis McKie
Yeah. You can imagine that after the drug hopefully makes it to the market, the doctor says, “Oh, you have cancer. I'm very sorry. We're going to do an H&E stain of your tumor, and then we're going to put it in the model.” It says, “This one won't work, but this one will.”
Alessio Fanelli
That's right.
Travis McKie
We're using the same approach today, where we're looking at many different mechanisms from different collaborations that we have in place. One of them we've announced with a company called Agena. These are all different mechanisms. The input is still H&E, using some of the same indications.
Using H&E, we're looking at whether drug A works in some sets of patients and whether drug B works in other sets of patients. You can take that to its natural progression and say, “Well, okay, if you can use that same input—just H&E—for experimental drugs, why not use it also for drugs that are already on the market?” In a sense, the same assay can be very predictive across many different cancers and many different potential therapeutics.
Alessio Fanelli
There are lots of models that take H&Es and go to gene expression out there, open source and otherwise. They do so-so. I've read on Twitter, in your Twitter feed and whatever, that you feel you have a data moat, right? So, why is Noetik's model better?
Travis McKie
Sure. I think the scale of data that we've trained these models on is pretty different from a lot of what's out there. The reality is there's just not that much of this kind of paired H&E plus other data modalities.
Typically, there are some datasets generated by academic labs, and others where they might have maybe 100 or a few hundred patients' worth of data. With paired spatial transcriptomics, that might even be an overestimate.
In comparison, we're generating these data with multiple patients per slide and individual patients distributed across multiple slides. We've generated more than 100 million cells with spatially resolved transcriptomics. That's all paired with H&E and protein as well. It's at least an order of magnitude larger than any of the other datasets that we've seen out there, and I think that makes an enormous difference.
We've seen with our own models that if you drop down to 40% or 10% of the data used in training, the models get a lot worse. They especially get worse at generalizing to other types of cancer from the ones that they've been trained on.
So, I think that's a big piece of it. I also think that the algorithmic side of it is important. We've developed custom architectures specifically for training on this multimodal data. Again, my background is in computer vision, and specifically in self-supervised learning there.
We've tried to develop self-supervised learning approaches for these data that are really adapted to solving the problem of figuring out what is different in one patient versus another, and then simulating what would happen if you were to knock down a particular gene or protein or something. This is why we call these world models: we're trying to build models that can simulate what's going to happen if you take a particular action. I think that's another big differentiator for these models. And then, again, interpretability is probably a third one.
Alessio Fanelli
But Travis, you were just talking about how one of the other strategies people take for this is to do perturbations on cells and then watch the response. Your experience, plus your strategy here, is that you can simulate this sort of counterfactual perturbation idea without even having to collect the data to do that, and you can see this—
Travis McKie
Well, there's—yeah, there's a big piece that we haven't talked about yet, which is that we are actually running perturbation experiments, except they're in vivo perturbations using a platform based in mice. We have another platform called PerturbMap. Ron, if you want to describe any of it—basically, this is a platform for generating highly multiplexed knockouts of individual genes.
It's the same kind of CRISPR knockouts that people are doing for individual cells in vitro, except when we knock out a gene in a cancer cell, that cancer cell gets injected into a mouse. It's barcoded, so we know which gene was knocked out, and it's being injected alongside roughly 100 other cell types with different genes knocked out. You end up with mice that have tumors that are barcoded and have 100 different genetic perturbations in them.
We can actually use that to validate our models and ask whether what the models are predicting in humans via simulation is actually borne out when you do these perturbations in a mouse system.
Shawn Wang
Sorry, there's a lot to unpack there.
Yeah, barcode. Yeah, so, sorry—barcoding. This is a technology in which an individual gene is knocked out with CRISPR, but this also introduces a set of protein tags in that cell that get expressed. It's a combinatorial code, so gene X might have proteins A, B, and C, while gene Y, when it's knocked out, has proteins D, E, and F.
We can tag those proteins, or label them with antibodies, so that when we go and look in the mouse, we know exactly which gene was knocked out based on which of those protein tags were expressed.
Alessio Fanelli
So, you knock out a gene, but you also add a gene that has the barcode proteins encoded on it?
Yeah, exactly.
Travis McKie
And I mean, the system is designed so everything that we're doing here is tissue-level, or could be in vivo—you know, tumors that are in vivo and in the form of the tumor in the whole tissue. Here in this mouse system, you have hundreds of tumors in the lungs of a mouse. If you look at these images, it's a mouse lung with literally hundreds of tumors in it.
Each tumor has a distinct biology that's driven by the biology of the knockout of the gene that's being perturbed, and we can capture the biology of each tumor in a spatially resolved way. What you can see is that we have certain tumors in humans that don't have immune cells in them. Those tumors are very aggressive, and they don't respond to immunotherapies.
You can generate those same tumors in this mouse system, and again, they don't have immune cells in them. You can do it genetically, so you can start to map the causative gene relationships between these different immune—or, more broadly, tumor—genotypes or biological profiles, if you will, to what you see in the human.
You can then treat those mice with drugs and see how hundreds of tumors in a single mouse respond to treatment with one drug. Or you can treat, let's say, 50 different knockouts across a panel of mice with 50 different drugs, and you can start to build this intersectional pharmacology and genetic experiment.
swyx
On Twitter and various places, I've heard you say, “no attic is no no cell lines, no war bottles.” Maybe you even said that, you know, if you—what's a cow?
Travis
And then we just said we have a mouse model.
swyx
Yeah, that's it. And we're injecting cell lines into the lungs, not under the skin.
So, yes. Fundamentally, you would think it's really important to build models that are trained on human data, and we're sourcing all these human tumors to build human-centric models. So, that is also true.
From the very beginning, we have asked this question: let's say we want to develop a drug from the very beginning, and let's say the FDA—and I know things have changed a little bit with the FDA—wants you to have some data in an animal that says your new mechanism works in some animal system. What do you do?
You're kind of stuck because you've now generated arguably the best data that you can in the human system, and then the FDA says, “Well, cool, but does it work in the mouse? How does it work in the mouse?” And so you have to back into this system that doesn't translate.
From the very beginning of the company, this has been a question. We started, probably at the same time we started generating the mouse-to-human data, building this mouse platform with the aim of drawing connectivity between these 2 systems.
We wanted a platform that would allow us to map the diversity of human tumors, because we know that if we just run a mouse model with 1 tumor, that tumor has no connectivity. In the mouse system, we want to have diversity of tumors, and we want to see a mapping of diverse tumor biology to the tumor biology that we're seeing in humans across many different locations.
We've been building this system so you can see many different perturbations that produce a lot of the tumor biologies—the plural—that you see in humans. We also want to be able to get from this mouse system to biologically relevant targets or genes in humans.
One of the fundamental problems in mouse systems is that we share many genes with mice, but there are a lot of genes and biological processes that we don't share with mice, as is obvious. Oftentimes, when you're developing drugs, you run into this situation: you have a target, and you have some biology that works really well in mice, but maybe that doesn't even exist in humans, or maybe that pathway is useless in humans.
One of the things we've started to develop, which I'll share more about soon, is a way to use one of these models to essentially infer human biology from the mouse directly. We're in silico humanizing the mouse. All the outputs in terms of the transcriptome from the mouse are in the form of human genes. When we read out this mouse system, we're reading it out in the form of human organ health.
swyx
How do you validate that? I mean, that's a pretty impressive claim if you can do it, but, man, it seems like a tricky validation task.
In my experience, both here at Owkin and at my previous employer, I could say a version of that. A lot of the approaches you're looking for when you're building these types of models involve asking whether the models are recognizing biology that you know to be true.
For example, in the human context, we know that 12% of patients with lung cancer respond to immune checkpoint inhibitors. Do the models recognize those patients? Can they recover those patients without training?
swyx
I mean, cold?
Yeah. And we see that. When you go look at those patients, we see that the underlying features of those patients map to what we know about those patients in the clinic.
In the mouse system, we have control genes. We ask: if you look at the mouse tumor embedding space, do the tumors that should be really cold look really cold from the human inference?
swyx
Cold in the sense of—
Yeah, like they don't have immune cells. And then hot in the sense of lots of immune cells. We try to build systems where you have these handholds, and the more of these examples that you know to be true that work, that you see, the more confidence you have.
Obviously, when you're in the regime of something very new, it's still uncertain for some of these.
Alessio Fanelli
So, the bridge between the mouse and the human is that you build a world model on the human, you build a world model on the mouse, and then you say, “What are the parallel structures in the 2 latent spaces?” Is that kind of the intuition here?
Travis
That's one thing that we're doing, but actually, this is even simpler. We've trained models on human H&E, spatial transcriptomics, et cetera, and then we're just running inference on mouse H&E, which is easy to generate.
Apparently, mouse H&E looks enough like human H&E that the models think it's perfectly valid H&E. They make predictions about whether it's immune-hot and immune-infiltrated, cold, fibrotic, or some other tumor phenotype.
And those predictions are accurate. These are some of the controls that Ron mentioned. We know that in mice and humans and everything, if you knock down tumor cells' ability to present antigens to immune cells, those are very cold. Immune cells are nowhere near those tumors. That's exactly what we see in the mouse, and that's exactly what the models—the in silico humanized models—predict.
There are other examples where, again, we're recovering the biology that we expect to see there. Then there are findings that are novel but also make total biological sense. For instance, we have done knockouts in the mouse of half a dozen genes that are all in the same pathway. You might predict that knocking down those genes is going to produce the same phenotype because they're all on the same pathway.
swyx
And that was a pathway?
Yeah, a pathway is like protein A signals to protein B, which signals to protein C, and there's a chain of events that leads to the cell having some behavior—changes in its metabolism, its growth, et cetera. I don't know if you've ever seen these crazy-looking protein-signaling diagrams that make you want to stay away from biology. People have worked out a lot, and they know that these 2 proteins interact physically and signal to each other and so forth. And so, some chain of those—
swyx
Interactions
this protein binds this protein, and that causes it to upregulate a gene that causes another protein to be formed, blah blah blah, until you get to some phenotype, meaning the cell changed the way it looks or the—
Exactly. Based on decades of biological literature and experiments on these, there's a very strong biological prior that if you hit gene A, gene B, and gene C, and they're all in the same pathway, you should get similar phenotypes. This is kind of how old-school genetics was done, and we see that with these in silico humanized mouse models, which is amazing to me as a biologist: you have a model that's trained on human data, then you show it some mouse histology, and it's able to say these 5 different tumor genotypes all look like they have the same phenotype—and lo and behold, they're 5 genes that are in the same pathway.
swyx
So, you guys are switching gears a little bit because we want to talk about models on the Latent Space podcast. You guys recently—there was an interesting blog post about the TARiO model. It's a transformer-based model. Do you want to talk about that?
Sure. This is a new model architecture that we developed after the first virtual cell model, OctoVC, that we developed. TARiO is just a different transformer architecture. One major difference between it and our prior models—if this is a model podcast, this gets into the self-supervised learning objective—is that for a while, including with OctoVC, we were training models on what's called the masked autoencoding loss function, or objective.
You have a piece of data, chunk it up into small chunks, mask out some of those chunks, and the training task is for the model to predict the masked-out chunks from the revealed chunks.
swyx
Like BERT.
Travis
Yeah, exactly like BERT.
swyx
What are the chunks? Because this is multimodal, and I would imagine the different channels contain wildly different levels of information. I remember seeing something like 99% masking in OctoVC, if I'm—
Yeah, yeah. So—
swyx
And that was kind of surprising because when you have 20,000 channels, and maybe some of the channels are fairly—most of the signals are fairly sparse—
Yeah.
swyx
Then it seems like it'd be either there's a huge redundancy here in your data, or you really risk just throwing the baby out with the bathwater.
Travis
Yeah. What are the chunks? That totally depends on which modalities we're talking about. In spatial transcriptomics, one chunk, or one token, might be the level of expression for a particular gene at a particular spatial location. For protein images—multiplex protein images—it might be the image patch for that particular protein at a particular location, and so on. For histology images, those are usually just patches of the image, so pretty standard, like vision-transformer style.
The masking—and the maybe surprising result that you can, and actually need to, mask out large amounts of the data to get the model to learn anything interesting—is important. If you ran the hypothetical where you only mask out 10% of the image, more like BERT, for instance, in language modeling, what do the models learn? They learn these boring behaviors, like how to continue an edge a little bit between 2 regions of an object or something. They can learn that task very well, but they don't end up learning anything about the holistic structure of the image data.
Akash Tiwari
We found pretty early on at Noetik that the same thing was true with these multimodal transformers: if you mask out a lot of it, there are actually pretty strong correlations between where protein A is expressed and where protein B is expressed, and forcing the models to learn them is really what gives them this predictive power.
Shawn Wang
And so, TARiO, though, is an autoregressive model?
Akash Tiwari
Yeah, exactly. That was going to be the plan. Prior models, including OctoVC, were of this masked autoencoding-style training objective. TARiO is an autoregressive model, which, if you think about it, is kind of a particular choice of masked autoencoding, except instead of randomly masking out part of the data, you're always asking the model to predict the next token in a sequence.
We know that this is something that scales very well with LLMs, like training on the next-token prediction task, and it's still an open question: how do you get models of other data modalities to scale the way that LLMs have scaled? TARiO was not actually our first attempt, but one of our subsequent attempts to bring that autoregressive, next-token prediction task into modeling spatial transcriptomics data.
We found that when we use this architecture and this task, we started to see much better scaling behavior, where bigger models, and especially at longer context lengths, were really outperforming smaller models at shorter context lengths.
Shawn Wang
Because they can see further in the image?
Akash Tiwari
Yeah, that's probably a big part of it. I think there's actually a pretty subtle but very interesting result in that blog post with TARiO, which is that you only really see the benefits of using larger models when you're looking at longer context lengths. Here, longer context really means, again, that you're seeing more tissue at once, more area at once.
I'm not super deep into the language-modeling literature, but I don't know if there's an analogous thing with language models, where you only see these scaling behaviors at longer context. So, what we could be finding here is that, with patient data, you really do need to incorporate more of the patient's spatial context to get the models to learn these more complicated nonlinear patterns in the spatial transcriptomics and take advantage of it.
Shawn Wang
Is it possible that part of this is because you have some number of low-expression genes, and that the behavior is driven entirely by some of those low-expression genes?
Akash Tiwari
Yeah, definitely possible that the more context you have, the more likely you are to catch these low-expression but highly predictive genes, et cetera. I would guess it's a combination of that and larger area. We've done some experiments just comparing a model with the same amount of context, but in smaller or larger areas, and there definitely seems to be an advantage to looking at larger regions of tissue as well.
Shawn Wang
I want to hear about this—you did a big deal recently. You got a lot of press, and I think you have the distinction of being one of the only AI-for-bio tooling companies that's making money.
Akash Tiwari
Accidental. Nope.
Shawn Wang
So, could you tell us whatever you can disclose about that? We love hearing about this.
Akash Tiwari
We were really excited to announce a deal with GSK where we licensed them OctoVC, which is our virtual cell foundation model. We announced that back in January. It's a $50 million deal and includes an upfront payment, milestones, and, separately, an annual model licensing fee.
I think this was an attractive deal for both parties, for us and for GSK, because the deal focuses on models that we've already trained on lung cancer and colon cancer. It allows us to provide them with access to the models. GSK is one of the top AI teams in biopharma, so they know how to use these types of capabilities. They can use them for their internal use, and they can also use them to fine-tune on their data.
That was a really big sell for GSK as well, because GSK—and every pharma company—is sitting on mountains and mountains of so-called translational data. The types of data that we're training the models on come from clinical trials and pathology specimens across many different therapeutics. Everyone's sitting on a lot of this data, and it's been very hard to unlock.
All of a sudden, GSK can use our models both to do simulations and to do therapeutic discovery.
But they can also fine-tune models on their data, and in a way, the model then becomes GSK's version of the model. This was super exciting. It was the first announced foundation-model licensing deal in the space, and frankly, it was one we'd been trying to do for a long time, even before Noetik.
I think a lot of companies have been trying to do these types of deals, and it's historically been slow for adoption on the pharma side. It's also been slow to demonstrate a very clear value proposition for different types of capabilities. What's unique about this deal is that it doesn't look exactly like a software licensing framework—for, let's say, a small amount of money with a number of seats, where you license a model.
It looks like a real business-development deal in the industry, where there's a very significant, multimillion-dollar cash upfront, near-term payment. But then the substrate of the deal is not a molecule. The substrate is actually a model, which is what really made this break.
Shawn Wang
Why do you think there's appetite for this suddenly? It seems like almost whiplash. It seems like only maybe a year or 2 ago that bio was dying and whatever, and now suddenly there's this deal. Boltz is getting a ton of attention. There's so much attention on Isomorphic Labs.
Akash Tiwari
I mean, maybe not totally, but increasingly more people in pharma, across the industry, are seeing the value of different capabilities. They're able to use some of the open-source capabilities and demonstrate the value to themselves internally.
If you look at a pharma company, these companies are working on dozens and dozens of programs. My opinion is that pharma increasingly wants to be able to access models not just for one collaboration, where you and I are working together on this one program. They want to be able to access the technology across the whole pipeline.
I think that's going to create a driving force for not just bespoke, project-driven licensing, but actual broad licensing, where a pharma can access the technology in many different therapeutic programs.
Dan Fu
With the structure of prediction models—protein-structure-prediction and binding-prediction models—there's a massive public data set. There are increasing amounts of data people can generate to augment that. There's enough data to the point where people can train very good models, but maybe not just on the data that any one biopharma company has.
I think the same is true, but even more so, for the types of models that we're building, which are foundation models at the patient-biology level. No one company—I mean, these companies may have a lot of data, but it's scattered, it's siloed—and pulling everything together to train an actual foundation model may not be as easy as it sounds within a single company.
That's the nice thing about being a startup here: we can make that bet that you actually do benefit from generating all of this data in a uniform, unbiased way, at very high quality, and then use that to develop and train the models. My opinion is that you need to have data at that scale before you can even think about developing models that actually work.
You can't do the AI R&D or build the algorithms until you have a good enough data set to tell you whether your favorite algorithmic idea is actually working or not. That's a major advantage for us: we have enough data to see whether my idea or someone else's idea about how to build the model is actually leading to improvements there.
Akash Tiwari
Yeah, I mean, this is a good point. Sometimes people ask me, "Why don't you just generate your data?" We just started generating data 4 years ago. There was no model.
Shawn Wang
How many years? Like 2 years, maybe? A year and a half at least before you had the first trained models working?
Akash Tiwari
3 or 4 years still. So this is year 4 again. We basically opened the lab, hired a team, got all the instruments, and started sourcing tumor samples. There was no prior indication that any of this would work.
We started generating data and sourcing human tumors, processing them. We built this whole processing pipeline to get the tumors into these arrays and formats. It takes weeks—it takes literally 2 weeks for a machine to run a couple of slides on transcriptomics. You've got these 2-week runs where you're processing 2 slides, and we're just churning data for months.
We didn't even have enough data to train a model for at least a year and a half. Then you're building processing pipelines, you have to align all the data, and you've got to post-process it off the machine. We built all this, and then, let's say, 18 months later: "Hey, I wonder if this stuff—"
It wasn't obvious. It wasn't like, "Oh, we're going to train this off-the-shelf on some open-source architecture." Dan and the team have done a ton of work.
Shawn Wang
Big zero. Big crazy bet.
Akash Tiwari
I just went for it. We started generating data and sourcing human tumors, processing them. We built this whole processing pipeline to get the tumors into these arrays and formats. It takes weeks—you know, it takes literally 2 weeks for a machine to run a couple of slides on transcriptomics. You've got these 2-week runs where you're processing 2 slides, and we're just churning data for months.
Dan Fu
Yeah, there wasn't really anything major to go off of. There were transformers developed for single-cell data, but incorporating spatial data into that was—again, there just weren't really data sets out there that people had been able to develop on.
We do a lot of custom model building, and I enjoy that. I think people enjoy that.
Shawn Wang
Yeah, you're really unique, innovative, and bold. Sorry, who are you looking for? What kind of people?
Anybody excited about doing ML research on this kind of alien landscape of data, where you really have to figure out what's working from first principles, and obviously the work we do should have very, very large impact.
We're definitely not restricted to people who have a biology background. People who just like tackling very challenging machine-learning problems and are open to learning the minimum amount of biology necessary to make progress would be great candidates.
Shawn Wang
Talking to you guys reminds me a lot of the bio lads.
Dan Barr
Yeah.
Shawn Wang
I know that both of you are part of the Recursion mafia.
Dan Barr
You know, I'm not.
Shawn Wang
Well, yeah, yeah, yeah, but—but you—this, yeah.
Dan Barr
Yeah, yeah, yeah. We're going to be on the show in the future, too. We're looking forward to that.
Shawn Wang
It's interesting because both of you seem to have really similar philosophies. You have deep convictions that you're just going to start collecting data before you know this is going to work. You're going to brute-force it—go, go, go—and eventually it will work. You have signs of progress. I don't know, I think that's really impressive.
I wonder if there's something about Recursion that's in the water, which has led to this sort of thinking of, "We're going to commit to doing things at scale, and it may not work at first. You have to hit a certain point before it will."
Dan Barr
I mean, we failed a lot at the beginning.
Shawn Wang
Yeah. You mean at Recursion?
Dan Barr
At Recursion, yeah, yeah. We had to build it from first principles, and we really did. We spent many years trying to figure out what the data should look like. Ian and I were both involved in platform development: how to design these data sets, how to design the experiments, and iterative cycles over the years saying, "These things did work; these things didn't work."
At the end of coming out of Recursion, I think what a lot of folks there had was an understanding of what we need to think about. Even if I wanted to design a different data set today, and it's totally different, what are the things we learned over mistakes—or not mistakes, but trial and error, basically—over that many months that we would try to insert in our new approach?
I don't know that everything that I've predicted at Noetik in terms of how to generate the data set has been important, necessarily. I know that we could start at the very beginning and say, "Okay, let's make sure we do these 10 things." I know every one of these 10 things was important before, so let's at least make sure we do these 10 things.
I don't know that all 10 things are important for us today, but I would presume that many of them are, and that lets you leapfrog that process of trial and error a little bit. Certainly, we still have trial and error, but hopefully we're not having to solve 15 problems. Maybe we're only solving 3 or 4 problems over time.
Shawn Wang
So, for small biotech startups, which are probably in the AI space and are collecting their own data—their own data moat—do you have any advice or suggestions for how to be more successful there?
Dan Barr
I think you need to think ahead to, okay, what am I trying to do on the machine-learning side, and what is the right data for solving this problem? Oftentimes, I see a lot of companies say, “I want to generate X data set. I’m just going to generate X data set, and then I’m going to do machine learning on that.” That might not be the right data set; you might not have designed it the right way. It doesn’t follow that any data set has a machine-learning use case, or that that data set is going to solve the problem you’re trying to solve.
For me—even in life—it was: What problem are we trying to solve, and what data are going to help solve that problem? Rather than going from the data directly to trying to solve the problem.
I also had a quick piece of advice: Pay attention to where the technology is and where it’s changing rapidly. I finished my PhD in 2016. I did a lot of looking at spatial RNA via a technique called in situ hybridization, the same technique that is at the base of what we’re doing. I could look at maybe 2 genes at a time on a single sample, and that took me a full week of manual work.
I came to Noetik 5 or 6 years later, and all of a sudden there were platforms where you could look at 1,000 genes or 20,000 genes at once with a single machine that could run this assay. It’s expensive, but it’s data beyond the wildest dreams of Dan Barr in 2016, and that is only improving rapidly. I think it’s important to see what the technology of today allows and also where it’s going in terms of what data to generate.
Shawn Wang
And what does that pitch look like? “I’m going to generate data for a year and a half, and then I’ll spend $50 million, and then—”
Dan Barr
It wasn’t $50 million. It was maybe closer to $10 million. If you’re going into a regime where there are no data and you want to do something different, there’s no shortcut to it, right? You’re going to have to generate the data set, and you’re not going to know the answer until it’s there. That’s why a lot of companies are not going into that space where there are no data sets, because it can be challenging to do that.
Shawn Wang
I think a lot of smaller biotech AI startups will try this pattern. They’ll either start with a public, open-source data set, or they’ll try a pilot—incrementally collect a small amount of data and see if something works or doesn’t. Oftentimes, there’s almost a critical point where, below that, you’re just not going to get any signal. You have to have conviction that you need to collect up to a certain point before you start really driving something fundamentally valuable.
Dan Barr
Yeah, I mean, imagine trying to train a foundation model on not enough data.
Shawn Wang
Yeah. Yeah. But with your claim about AlphaFold, right? You have your GPT—GPT-3, GPT—you know, with GPT-1, 2, and 3, there was a clear progression there. With each one of them, you could see there was something which worked with scale, and there was this insight: Oh, we’re going to scale this up.
Sometimes with biological data, the process of collecting lots of data is just very expensive to begin with. You can’t just take something off the shelf and expect that you’re going to hit the threshold of GPT-3-like fullness.
Dan Barr
Fullness.
Shawn Wang
Yeah, yeah.
Dan Barr
So, yeah, totally.
Shawn Wang
It takes some conviction.
Dan Barr
It definitely takes conviction. I think it also takes a scientific belief that there’s a lot out there that we just don’t know yet, and that you’re not going to capture the biology you need to by having, right now, an agent that reads all of the biological literature, because again, that’s just a tiny slice of what’s out there.
I don’t know if it’s a great analogy or if I’m going to botch the history here, but in astronomy, it required Tycho Brahe collecting this enormous amount of astronomical data at his observatory. That was the substrate for Kepler figuring out the first laws of motion of the planets, and then that was superseded by Newton’s laws and so forth.
I sometimes don’t know how you even get started without this large repository of really high-quality data. Maybe there’s a tragedy-of-the-commons problem here of who’s going to generate that data and who’s going to capture the value of it, but I’m very glad that we’re taking that bet and seeing it pay off.
Shawn Wang
Yeah, this is not my expertise, but hypothetically speaking, how much of the PDB do you need to train?
Dan Barr
Some people argued that—yeah, and then you can get some pretty good models with, I think, 1% of 1%. There are people going back to the 1990s who argued that the PDB was already complete, in the sense that if you had a sufficiently smart algorithm, you could have done a pretty reasonable job at protein folding even back then.
Shawn Wang
Interesting.
Dan Barr
You don’t need a lot to get a pretty big boost, but the community was sort of independently collecting PDB data for quite some time without necessarily being convinced that this was going to lead to solving protein folding. Most of those structures were quite useful in and of themselves. Maybe that’s the counterpoint: Oftentimes, just knowing a protein was very helpful for us, anyway, with data.
We did see a transition from the early data. How many samples did we have? I’m guessing probably on the order of a few hundred before there was a huge bolus. There was definitely a moment very soon after I joined where the data set just kind of doubled in size overnight because there was a huge bolus, and the models immediately got a lot better at that point.
Now we run more controlled experiments: What happens if you train on 10% of the data versus 40% versus 100%? What happens if you hold out all of the pancreatic cancer or all of the breast cancer? We have a much better idea of what kind of diversity and scale we need.
If we were sticking to cancer, maybe we’re not that far off. If we end up generating a few hundred patients in a bunch of major and some minor indications, which we’re going to do this year, maybe that’s enough to generalize to kind of all cancer, because there is a lot of shared biology in cancer and immune cells across different tissues, mutations, and so forth. But if you think about all of the disease biology there is for a model to learn, maybe that’s another order of magnitude.
Shawn Wang
But even being able to solve all cancer biology would be pretty impressive.
Dan Barr
Yeah, to cure a cancer would be great.
Shawn Wang
Well, if it’s all cancer biology, it doesn’t take your cancer away. It also takes a different face.
Dan Barr
But yeah, at least if you go down and just take 1 drug, if you could look at 1 drug mechanism across the whole of oncology, that’s incredibly powerful. Imagine what Merck has done with Keytruda. Merck has run hundreds of trials with Keytruda—possibly even more than 1,000 trials—with different populations to find all these different indications: the subset of ovarian cancers, the subset of lung cancers, the subset of colon cancers.
That’s all been done by enrolling trials. If you can look at that biology from model embeddings and at least have a very well-defined starting point—if I’m going to run a trial, it doesn’t have to be as broad as it would need to be if I didn’t have any answer—then that can be a really powerful tool for a diversity of mechanisms.
Yeah, maybe just as a last point, going back to the virtual-cell hot takes: If your goal is to build an actual mechanistic model of an individual cell and then build up from 1 cell to an entire tissue, and then from tissue to patient and so forth, you might need a lot more data and a lot more data modalities than just gene expression or something like that.
We’re taking much more of a top-down approach. We’re trying to first solve the problem of what is determining heterogeneity among actual patients, and which of that variability is predictive of drug response. My intuition is that you don’t need to model the mechanism at the subcellular level necessarily to solve the problem of which patient should get which drug, or which targets are important in which patients.
I saw a similar debate play out in neuroscience and computational neuroscience, where for a long time people were really trying to build these biophysical models of individual neurons, and then they were going to stitch them together into models of the brain and so forth.
What actually ended up working in terms of building computational models of the brain and behavior is this abstraction: we're just going to treat individual neurons as linear-nonlinear units and put them together in neural networks connected by linear weight matrices. We stack a bunch of layers together and build neural network models of the brain that abstract away all of the biophysical details of what a neuron is doing.
Those are now by far the most predictive models of how a given neuron is going to respond to real-world stimuli in a real brain. I think my bet is that the same is going to be true for these models, too. Modeling at the level of functional tissue, where you have a bunch of cells interacting in a disease context, is going to get you to the problem of predicting patient-level behavior much faster than trying to first model a cell and then stitch a bunch of those cells together.
swyx
Yeah, that makes sense to me. It's a good analogy. I like analogies.
Do you have any call to action for the listeners?
Yeah. I would say, first, everyone should be excited about biology. Sometimes a lot of my hot takes on X recently are just that I feel like there's a huge amount of enthusiasm in the mainstream tech ecosystem, and people aren't really following a lot of what's happening in the biology space.
But at the same time, you're hearing, you know, an AI lab saying we're going to cure cancer. People should actually look at the folks working on curing cancer, working on aging, or working on other areas of biology. These are really exciting problems. There are real, significant ML problems in the space.
One call to action is that I would love for people to be more stoked about learning about applications of machine learning in the biological sciences and solving some of these hard problems, because I think these are the problems that are going to massively impact humanity in the next 10 years. This is really the very beginning. Maybe we're in the first inkling of the ChatGPT moment for bio, but it's very much just the beginning.
Yeah, in line with that, really dig in and learn more about the details. A lot of the time, it's presented as: We have these protein-folding models, we have these binding models, and we have AI-for-science agents that are reading all of the literature and automating these computational-biology workflows.
I think it's important to realize that there are a lot of problems in AI for biology, AI for biochemistry, and so on. Some of them are very important, but solving any one of those is not going to solve the problem of how we develop better therapeutics. We're focused on a particular slice of that process, which is translating things that we know work well in some patients into successful drug trials where we know exactly which patients to give them to.
That requires building foundation models at a particular level—the patient level—but people should not be under the impression that this is all going to be solved immediately because AI agents like LLMs are going to read the literature and figure out what the right drug is. There's a lot more data to generate, a lot more ML problems to solve, and a need to translate those methods into actual successful drugs. There are a lot of different places to contribute.
swyx
Lots to do.
Yeah, indeed.
swyx
Great. Thank you very much.