[BidClub_]
No Priors · · 58 分钟

No Priors 第103期|与 Vevo Therapeutics 和 Arc Institute 对谈

Sarah GuoNima AlidoustJohnny YuPatrick HsuDave BurkeHani Goodarzi

YouTube
TL;DR
  • Tahoe 100 将单细胞扰动数据推向机器学习规模。 Vivo 报告称,该数据集覆盖来自不同患者的50个癌症模型、1,200种药物处理,以及60,000种药物–细胞或药物–细胞系交互,共计1亿次单细胞测量;在此之前,公开可用的扰动数据约为100万—200万条。Dave Burke 将其潜在影响类比为细胞生物学领域的“ImageNet时刻”。

  • 真正形成差异化的资产,是具备因果性、多样性且生成一致的数据,而不是又一堆健康细胞观测。 扰动提供前后对照信号,疾病模型则暴露出正常组织中不存在的状态。4个人以高度一致的方式完成了60,000次实验,几乎没有批次效应。团队认为,规模和广度本身就可能让数据生成新的假设与意外发现。

  • 虚拟细胞回答了蛋白质模型留下的系统层面问题。 蛋白质语言模型可以学习折叠、结合和结构生物学,但药物作用于嵌入细胞、肿瘤及更大系统中的蛋白质。虚拟细胞的设想是一个“概念性CPU”,预测基因编辑或化学物质如何改变转录组,并最终解决如何将疾病状态推向健康状态的逆向问题。

  • 当前虚拟细胞模型仍未达到商业成熟度,对差异表达基因的预测表现约为10%,且尚无公认基准。 按当前编码方式,Tahoe 100约对应2,000亿—3,000亿个训练tokens。Dave表示,其他领域约1万亿tokens可以作为较为宽裕的比较点,但也强调生物学的扩展规律仍不确定。Hani Goodarzi认为,蛋白质模型的成熟度已超过GPT-3,而细胞模型“更接近GPT-1,而不是GPT-2”。

  • 开放源代码是一种运营策略,并非脱离 Vivo 经济模型的慈善行为。 Tahoe 100与 Arc 约2.3亿细胞的 scBaseCamp 数据集合并,形成一个3.3亿细胞的 Virtual Cell Atlas;社区可以训练和批评模型,Vivo则继续建设 Mosaic 数据生成平台。Nima将其定义为“重新立一根标杆”:一支小型内部团队可以借助外部研究生态,而无需组建庞大组织。

  • 如果虚拟细胞能够奏效,最大的潜在收益是同时压缩靶点风险、化学搜索空间和实验时间。 约90%的药物会在临床试验中失败;嘉宾认为,失败既可能源于药物本身不理想,也可能因为选错了靶点。足够准确的模型可以搜索数千万种化合物,跨患者情境泛化,帮助团队“量两遍、再下刀”(measure twice and cut once);但 Patrick Hsu警告,准确率只有10%时,“你只是在模拟噪声”。

  • 技术拐点可能已经到来,但治疗验证仍将缓慢推进。 中国生物科技公司的成本、速度、抗体生产能力和IND申报支持材料,正迫使美国公司转向更精简的团队和更强的供应商;Nima称这是“生物医药的清晨”,并呼吁现在就开始建设,而不是承诺3—5年后交付成果。Sarah Guo强调,治疗通常需要11年以上,即使把成功率从10%提高到30%,仍会受“少数样本定律”制约,可能需要一个长达10年的窗口才能验证。

摘要 · 为研究而整理的核心内容

1. Tahoe 100 将扰动生物学变成机器学习规模的数据资源

  • Johnny Yu将 Tahoe 100定义为全球最大的单细胞RNA测序数据集:覆盖来自不同患者的50个癌症模型、1,200种药物处理和60,000种药物–细胞或药物–细胞系交互,共计1亿个细胞。他最强的判断是,这可能是“第一个能真正推动该领域机器学习的数据集”。

  • 关键不只是细胞总数。此前公开可用的扰动数据约为100万—200万条单细胞数据点,而历史上大多数数据集属于观察性数据,彼此割裂、标签质量差,且来自健康组织而非疾病状态。

  • Dave Burke将其类比为2009年的 ImageNet:足够大且目标明确的数据集,可能带来非线性的能力跃升。这仍明确只是希望——细胞扰动数据或许能为系统生物学带来类似基础蛋白质结构数据集之于 AlphaFold 的作用,包括 PDB 和 CASP 在内的数据集曾帮助后者实现突破。

2. 虚拟细胞在蛋白质结构之外建模因果关系

  • Nima Alidoust区分了两种语言:蛋白质模型学习的是“结构生物学的语言”,包括折叠、分子结合以及抗体–蛋白质交互;虚拟细胞要学习的则是“系统生物学的语言”,即一个靶点如何处于细胞、肿瘤、免疫环境乃至更大生物体系统之中。

  • Dave用计算机作类比,让这一抽象概念变得具体:DNA是编码细胞的ROM,RNA是持续变化的工作内存RAM,而虚拟细胞模型则要推断一个“概念性CPU”,将被编辑的基因或施加的药物映射为转录组响应。

  • 逆向问题才是治疗真正要解决的问题:给定一个疾病细胞的表达谱,哪种基因或化学扰动可以将其推回健康状态;对于癌症,则是在不伤害健康细胞的情况下杀死癌细胞?这仍是预定目标,而非已经验证的产品能力。

  • Hani Goodarzi认为,扰动数据的核心价值在于因果性。观察肝脏样本只能得到相关性;对基因或化学物质进行干预,则能形成明确的前后对照,让模型学习哪些变化驱动了细胞状态,而不仅仅是伴随细胞状态出现。

3. 信息含量比再增加100万个重复细胞更重要

  • Nima表示,早期单细胞基础模型可以舍弃一个约6,000万细胞集合中的99%,而性能几乎不受影响。他的推断是,大量可用数据重复了相似的生物学情境,原始规模夸大了模型实际能够学到的信息量。

  • 多样性提供了缺失的信息。一个要理解心脏、大脑、肝脏、骨骼或癌症的模型,需要覆盖不同细胞和组织类型中的正常与疾病状态;扰动有助于探索高维“流形”,而不是反复采样同一个局部邻域。

  • 每个细胞大约贡献2,000—5,000个基因与表达tokens,因此按团队估算,Tahoe 100约对应2,000亿—3,000亿个tokens。Dave提到,GPT-3约使用5,000亿个tokens,ESM-3约使用7,000亿个tokens,约1万亿可以作为较为宽裕的比较点,但他说只有真正达到这个规模后,行业才会知道相关的扩展规律是什么。

  • 不同领域的瓶颈也不一样。经过数十年的基因组测序,DNA模型越来越受制于算力和上下文长度;细胞模型仍受数据限制,因为可规模化的单细胞分析直到最近才出现。

4. Mosaic 用混池实验取代逐个患者筛选

  • Vivo的 Mosaic 平台将多个患者来源模型——包括多种癌症类型和不同遗传背景——混合到一个“马赛克肿瘤”中,然后同时筛选数百乃至数千种药物,并解析每个模型的响应。Sarah的反应体现了实验模式从串行走向多路复用的跃迁。

  • 初期扰动针对癌症相关基因、细胞生长、DNA调控和泛癌通路。Hani认为,这些保守通路也可能广泛适用于神经科学和免疫细胞发育;未来数据集还可以加入罕见病,并利用模型反馈填补数据缺口。

  • 规模改变的不只是实验吞吐量,也改变了实验哲学。Nima表示,Vivo在5周内生成的扰动数据约为此前10年公开数据的50倍,因此不必再预先锁定狭窄假设:团队可以在化学物质和患者样本空间上大规模探索,减少偏见。

5. 开放数据让小公司借用整个领域的智慧

  • Nima表示,Vivo在讨论这一机会后的数小时内就决定开放数据,原因有二:一是在预期数据集规模上“重新立一根标杆”,二是让约3—4人的内部团队保持专注,同时借助外部研究者暴露数据的优势、缺陷和模型机会。

  • Arc的 Virtual Cell Atlas 将 Tahoe 100与约2.3亿细胞的观察性集合 scBaseCamp 合并,总规模约为3.3亿细胞。设想中的工作流是互补的:模型可以先在广泛的观察性数据上预训练,再加入扰动数据来学习动态过程并提升预测能力。

  • Arc构建了一个类似“Google爬虫和索引”的智能体,用于寻找、标注并统一重处理公开测序数据。统一处理很重要,因为工具、工具版本、基因组构建版本和试剂化学体系的变化,否则会在合并数据集中引入分析和实验批次效应。

  • 运营一致性本身也是资产的一部分。Dave表示,4个人以完全一致的方式完成了60,000次实验。减少操作人员数量,有助于降低“只有在我手里才有效”这一常见的生物学疑虑。

6. 准确模拟可以压缩生物学时间和化学搜索空间

  • Patrick讲述的衰老项目说明了现实约束:一轮实验可能需要让动物衰老2年。计算机模拟的并行能力将带来颠覆性变化,但他的底线很明确——如果虚拟细胞只有10%的准确率,“你只是在模拟噪声”。

  • Vivo计划预测一种新化学实体如何与来自不同患者模型的细胞互动,判断它能否将疾病细胞推向健康状态;对于癌症,则是在不杀伤健康细胞的情况下杀死癌细胞。Johnny设想的未来,是由虚拟细胞模型生成药物,但这目前还不是已经实现的能力。

  • Nima将泛化问题分为两类。对于细胞,模型必须从已观察到的患者迁移到新的个体差异;对于化学,它们必须在数千万种候选化合物以及几乎无穷多的生物制剂中进行搜索,找到值得合成的小范围区域,而不是继续在熟悉的化合物库上逐步筛选。

  • 约90%的药物会在临床试验中失败。Patrick认为,行业面对的可能既有糟糕的药物分子——效力、毒性、pH、动力学特征等属性不达标——也有错误的靶点。虚拟细胞可以在昂贵的化学研发和药物开发开始前,先缩小靶点搜索范围。

7. 转录组是抽象层,细胞环境仍然可见

  • Dave将RNA表达描述为一台1980年代的图形均衡器,大约有20,000根不断移动的柱条,响应环境、压力、衰老、健康和疾病。Patrick表示,团队认为转录组层足够丰富,可以捕捉具有决定性意义的细胞状态,而不必模拟细胞中的每一个分子细节。

  • 多细胞建模随后可以逐级扩展到球状体、类器官、体内模型、免疫环境和空间数据。Hani的细化观点是,环境信息会“经过细胞过滤”;如果观察到足够多的情境,一个名义上的单细胞模型或许就能推断周围组织产生的影响。

8. 平台型公司生来就是为了放弃错误假设

  • Nima将 Vivo 与单一假设型生物科技公司作对比:后者的组织和激励机制都会围绕一个论点展开,甚至可能只在3个患者样本上测试过药物,就继续推进开发。平台型公司可以生成相互竞争的假设,并对哪些项目进入临床保持“更加科学”的判断。

  • 中国生物科技公司进一步加剧了运营挑战。Patrick提到更低的成本基础、更快的研发管线、扎实的安全性和毒理学材料、IND申报支持工作,以及高效的抗体生产能力。他认为这种竞争可能是有益的,因为患者、投资者和公司都希望更快、更便宜地获得真正有效的分子。

  • 无论是由供应商串联起来的“虚拟生物科技公司”模式,还是完全垂直整合,都没有干净地解决这一问题:前者已被证明速度缓慢,后者则昂贵且官僚化。拟议中的中间道路是能力更强的供应商加精简的公司;Nima认为,行业应当吸收这些能力,而不是只依赖监管壁垒。

  • Nima关于“生物医药的清晨”(morning in bio)的宣言包含3部分:现在就开始建设,而不是宣布3—5年后交付;组织由“超级明星”组成的小型集中团队;挑战领域专家列出的、解释跨领域机器学习技术为何无法奏效的长串理由。他仍保留了不确定性:“也许能行,也许不行——但如果不试,永远不会知道。”

9. 技术拐点可能已到来,但治疗验证仍将滞后

  • Dave回顾了 ImageNet 和 AlexNet 之前的神经网络历史:当算力、数据和模型架构尚未跨过拐点时,一些想法可能看起来毫无产出。该团队认为,单细胞分辨率、可规模化扰动和现代训练方法,可能正是生物学领域正在汇聚的对应要素。

  • Evo 2被视为生物学学习出现涌现能力的证据:Arc在没有明确教授DNA知识的情况下,用9.3万亿个核苷酸训练了它,但模型学会了核糖体结合位点、密码子简并性,并在零样本 BRCA1 变异致病性判断上取得约94%的ROC曲线下面积。

  • 不同领域的成熟度差异明显。Hani认为蛋白质语言模型已经超过GPT-3阶段;Patrick和Dave将细胞生物学描述为大约GPT-2水平,或仍在向这一水平发展;Hani最终认为,虚拟细胞模型“更接近GPT-1,而不是GPT-2”。现有最佳模型对差异表达基因的预测表现约为10%,而且尚无公认基准可以确定不同模型的比较结果。

  • Sarah提到自己对AI生物科技长达10年的怀疑,并强调,即便细胞模型达到GPT-4式能力,效果也不会立即显现:治疗通常需要11年或更久,把成功率从10%提高到30%将是非凡进步,但仍受“少数样本定律”制约。Dave表示,更好的模型或许能“把大炮指向正确方向”,但验证仍会在约10年的开发窗口中缓慢积累。

Sarah Guo

Today, we’re here with the CEO, CTO, and core investigator of the Arc Institute, as well as the co-founders of Tahoe Therapeutics, to talk about their release of Tahoe 100, the largest single-cell drug-perturbation data set ever created. We’ll also discuss where we are in AI for biology, why we need a virtual-cell model and not just protein-structure-prediction models, and when we should finally expect to see treatments from the growth of machine learning in biology.

Johnny Yu

Hi, I’m Johnny, and I work on single-cell sequencing at Tahoe Therapeutics.

Nima Alidoust

I’m Nima. I’m one of the founders, together with Johnny Yu. I’m a quantum chemist by background, but I’ve converted to being a computational chemist who loves playing with biological data. We’re building Vivo to really do that: to predict how chemicals interact with cells in different biological contexts. Some people call it the virtual cell. That’s basically what we’re working on.

Patrick Hsu

I’m Patrick Hsu, one of the founders of the Arc Institute, which is working at the interface of biology and machine learning to try to understand and, one day, treat complex human diseases—which are most of the major killers.

Dave Burke

I’m Dave, CTO at Arc Institute, focused on computational biology and building novel AI models for biology.

Hani Goodarzi

I’m Hani, a core investigator at Arc. I work very closely with Dave and Patrick to push our virtual-cell initiative.

Sarah Guo

Congratulations, everyone. It’s a big day. Let’s jump right into it. What is Tahoe 100, and what is its significance?

Johnny Yu

Tahoe 100 is the world’s biggest single-cell RNA-sequencing data set, and it enables a ton of machine-learning applications, including things like the virtual cell. It also enables a lot of drug-discovery applications. Broadly, in the context of where I think we are as a field, it’s the beginning of a different way of doing drug discovery: understanding how to build medicines and bringing AI and machine-learning people into the mix.

Nima Alidoust

Over the last 20 years or so, people have accumulated massive amounts of data when it comes to protein structures, protein function, and how drug molecules interact with proteins. One thing we haven’t had as much of is information about how different cells behave in different contexts, and how different genes in each of those cells function in the presence of the other genes in these different biological contexts.

We believe this is the era for that. We’ve seen the emergence of protein language models built on the data sets that have been accumulated over the last 2 decades, but now is the era for having data on cells: how they function, how they interact with drug molecules, and exactly what Johnny was describing. Tahoe is really a landmark data set that allows us to measure how drugs interact with different cells from different patient models. That gives us the ability to build models similar to the protein language models, but in the cellular context.

Dave Burke

If you think about the history of AI, it’s punctuated by these data sets that come about. If you think about ImageNet in 2009, which Fei-Fei Li put together, and what that did to drive a nonlinear jump in machine vision, I think the hope here is that, by producing perturbational data sets that allow us to elicit cellular responses, we’ll be able to drive forward the ability to model at the cellular level—not just at the protein level.

Hani Goodarzi

Lots of people have been talking about what those foundational data sets look like for biology. Those data sets have been really useful for training protein-structure-prediction models, like AlphaFold, which was built on CASP, the competition built on top of PDB data. But how do you do this for cells and cellular dynamics, which is really what tells us about biology and how it responds in health and disease?

I think those are the core steps forward where we want to bring up our ability to study higher levels of abstraction in biology—not just the individual molecular machines, but how they operate in the context of an entire cell.

Sarah Guo

Congratulations also to the entire Arc team. Given that you’re working on both virtual-cell models and protein-structure-prediction and protein-language models, can you contextualize why we need both and where we are in the progress of each?

Dave Burke

I think we’re learning that as we look at these emerging properties of biology by training large-scale foundation models on nucleic acids and these virtual-cell models, which we’ll talk more about today. We have this debate often internally. I have an engineering and computer background, so the way I think about it is this: if you think about the cell, the DNA lives in the ROM—the read-only memory. It’s coding for the cell. The RNA lives in the RAM, so it’s like the working memory.

The RNA is constantly changing its expression level. It’s almost like one of those 1980s graphic equalizers, where you have 20,000 bars for each gene, and it’s constantly adjusting its expression level depending on what the cell is experiencing—whether that’s the environment, stress, aging, disease state, or healthy state.

What we’re trying to do with this data, as a field, is create these virtual-cell models. In a way, that’s inferring a notional CPU for the cell: how does the cell respond to an input? That input could be an edited gene or the application of a drug. Then how does that reflect in the transcriptomic profile? That CPU is an analogy to the AI model that you want to build.

Once you have an AI model, what’s really interesting is that you can start posing the inverse question. Given a cell in a certain disease state that’s exhibiting a certain transcriptomic profile, how do I perturb that cell—whether with a gene edit or a drug—to perturb it back into a healthy state?

I think that’s what’s really exciting about this data. It enables these models, which then enable these tools, and hopefully can accelerate drug discovery.

Nima Alidoust

One thing I’ll quickly add is that, when we think about different domains in biology and building AI models of those domains, there are parts where we are data-poor and parts where we are compute-limited.

When it comes to DNA language models, thanks to the field and decades of sequencing a ton of genomes, we’re not as data-limited. Compute—specifically, context and how long we can consume DNA, what size of inputs we can use, and all of that—is a big limitation that we’ve tried to solve.

When it comes to cell-state models, that is an area where we’re absolutely very much data-limited. Being able to profile cells at single-cell resolution is basically a new technology. It emerged over the past decade, but the explosion really happened over the past 5 or 6 years. We’re just getting to the point where we can generate that kind of data at scale.

It’s not just the scale, though. Before scBaseCamp, which is the data set being released together with Tahoe as part of the Arc Virtual Cell Atlas, the data set that the Arc folks created by collating all publicly available data, I think the number of human cells that we had collated together was on the order of 45 or 50 million. If you generate 60 million single-cell data points, scale is one thing. The question is how much information content there is in that data as well.

We built early versions of some of those virtual-cell models. We called them single-cell foundation models, or whatever name you want to use for them. What we saw is that if you reduce those 60 million cells by downsampling them by 99%—so you use just 1% of that data to train your models—the model’s performance doesn’t reduce that much.

That means the information content of the data you’re using to train those models isn’t amazing. Having data that comes from very different biological contexts is key in providing information content for the model so it can learn. That goes back to what Dave was saying: perturbational data sets allow you to create new contexts and new cell states that the model can learn from, and therefore use for different kinds of applications.

I’ll let Johnny talk later about the challenge of creating those perturbational data sets.

Sarah Guo

Before we go there, can we zoom out for a second and have you describe, in layman’s terms, what the data actually tells you and where the prior data came from—even if it was information-poor?

Johnny Yu

If you look at the data that’s been generated over the past decade, it’s basically all kinds of academic groups, like us, or people in industry generating all these little data sets. There are a ton of problems with this.

First, there are batch effects. Even if one person runs an experiment on 2 different days, the data can look different, even if it’s the same cells. When you think about trying to build the internet of biology—which is what you need to build this ChatGPT moment in terms of scale—you need big data. Machine learning isn’t going to do anything for us if we don’t have big data, or if you have a data set that’s poorly labeled and full of batch effects.

This data set is basically doubling the size of all the data that’s out there cumulatively from the past decade. It covers 50 different cancer models from different patients, so it has cells from 50 different patients and 1,200 drug treatments. It’s a really deep and rich data set that effectively has no batch effects.

We think this is not only an additional data set for machine learning; we actually think it’s the first data set that’s going to enable machine learning in this space.

Hani Goodarzi

One thing that might be worth touching on is why perturbational data is important. The key is that we’re going from correlation—which is what a lot of biological research is. It’s descriptive: you stare at things and try to see, when you poke this way, what else is changing—to causation.

Going from associative changes to causation is where genetic or chemical perturbations allow you to have a very clear before-and-after. You have a set of causal changes that can actually drive a particular cell state.

The key is being able to do this in a generalizable way. You need to be able to look across many different cell types and tissue types. An ML model would need to train on all of that diversity in order to learn a general sense of the possibilities of cell states.

Dave Burke

In a topological sense, what the model is trying to do is create a manifold in a high-dimensional latent space. To explore that manifold, the model needs to see lots of different perturbations and responses. Once you do that, you have this generalized manifold that allows the model to make predictions for data that it hasn’t seen in its sample, but that still fits the manifold.

Hani Goodarzi

To make it more tangible, almost the entirety of the data that was publicly available before this came from healthy tissue. Very little came from diseased cells. Almost all of it was observational: you take cells from a liver sample, for example, and perform single-cell RNA sequencing on them.

That has the limitation Patrick was talking about: does it capture the causality of the gene interactions you’re trying to model? The second question is whether it allows you to model how a new perturbation will impact the cells, whether it’s a genetic perturbation or a drug perturbation. That is really the focus for Tahoe in this situation: perturbational data sets.

Nima Alidoust

If you put all of the perturbational data sets in the world together, then, if you’re generous, it’s 1 or 2 million single-cell data points of publicly available data. Tahoe is 100 million. We’ve massively increased that amount.

When you couple that with the huge amount of observational data sets from different species that are already in the world—which is what the Arc team put together—it turns out to be 230 million single-cell data points. They’ve tried to reduce the variations between these data sets as much as possible so they’re consistent with each other and can be used to train machine-learning models.

That’s the significance of this data.

Hani Goodarzi

I want to make a finer point on this. If you want a model that can learn about changes going on in the heart, the brain, the liver, or the bones, you need to be able to train across all those different cell types.

If you just look at normal, healthy cells, you wouldn’t necessarily learn how the manifold in latent space changes in disease. Being able to look at many different tissue types across different cancers is one way to get at those critical disease states that both basic science and drug discovery really care about.

Sarah Guo

How should we think about 100 million data points or 230 million data points, and the scale of this release, in terms of where we are? Is that enough to be useful? What do we know about scaling laws now?

Dave Burke

The short answer is that it’s a very hard question. We won’t know until we get there. What we can draw inspiration from is large language models in human language, as well as things like DNA language models, where we do have enough data to understand scaling laws.

In those cases, around 1 trillion training tokens is where you want to get to. GPT-3 was, I think, half a trillion tokens. ESM-3 was 700 billion tokens, so close to 1 trillion. A trillion sounds like a comfortable mark to hit.

The question then becomes how you count tokens, because cells aren’t exactly sentences. But if you count genes and their expression as tokens, I think this collection gets us close to where we want to be to start asking and answering those questions.

Nima Alidoust

A cell collection for these data sets has 2,000 to 5,000 genes, and each gene and its expression are basically a token in what we’re doing. So 100 million single-cell data points is equivalent to around 200 to 300 billion tokens.

There’s a finer point, which is how many of those tokens are actually informative to the model. But I think you understand the gist of it.

Sarah Guo

How do you decide where in the genetic landscape to start? How do you choose perturbations?

Hani Goodarzi

You want to match your perturbation toolkit—the arrows you’re throwing at the biology—against the biology you have. For cancer, that means going after cancer-relevant genes: genes that impact the growth of cells, genes that impact DNA regulation, and drugs that target key pan-cancer pathways.

For cancer-relevant questions, this data set is heavily based around chemical perturbations and cancer. But these pathways are so conserved and fundamental that they broadly apply to neuroscience and immune-cell development in general.

I think it’s really the foundation model that’s going to be able to take this data, ingest it, build a model, and then understand how to translate that data to a completely different context.

Johnny Yu

This is one of the really special things we have at Vivo: the Mosaic platform. It allows us to take cells from many different patients and, in cancer, that means all kinds of cancer—lung cancer, pancreatic cancer, and so on—from different patients with their own special genetics.

We pool them together into a single mosaic tumor, which we can then reproducibly screen against hundreds or thousands of drugs. This key innovation allows us to test tens or hundreds of cancer models at a time instead of testing one cancer model at a time. It makes this a really scalable data-generation platform.

That’s what we used to generate Tahoe-100M.

Nima Alidoust

When we think about how we build these pools in terms of information content, we want to maximize it by covering a lot of cancer patients. For this data set, we covered the biggest cancer types by how frequently they occur annually.

As we continue to grow the data set, we want to bring in rare diseases and perhaps more coverage of different parts of the cancer space. We also want to let the machine inform us by helping us fill in the gaps in the foundation models.

Another direction is chemical space. Frankly, when we generate data, we generate 50 times more perturbational data sets in 5 weeks than are publicly available from the past 10 years. You don’t have to prioritize as much, and that’s the beauty of it, in my opinion.

You can go large on chemical space and on patient-sample space. That way, you don’t have to come up with a hypothesis a priori about what you have to feed the models. You can just generate as much as you want and be more unbiased.

Hani Goodarzi

As scale increases, you can generate hypothesis-free, unbiased data. That’s really the beauty here.

Sarah Guo

Can the data surprise you?

Nima Alidoust

Exactly. This is one of the things I like to talk about. I hope the people we have here are representatives of the new generation of biologists.

One thing that has been slowing progress in biology is that we’ve always been super hypothesis-driven. There are reasons for that: a lot of these experiments are expensive, and they take a lot of time and resources. But now sequencing costs have gone down, the per-sample cost of single-cell sequencing has gone down, and compute costs have gone down.

I think it’s time to change that mentality in biology as well—to be more courageous and more freewheeling in terms of data generation and the kinds of samples you put together.

Sarah Guo

I want to talk about being more ambitious in biology and the open-sourcing of this in a second, but first I think we should zoom out and talk about, in layman’s terms, what the platform does. You can correct me if any of this is wrong.

You have these tumors that are a mosaic of cells from different patients, representing a huge amount of patient genetic variation. Each mouse can then be treated with different drugs, and the signal you extract afterward is the interaction of those drugs with each of these different patient types.

Johnny Yu

That’s right.

Sarah Guo

Nobody else thinks this is crazy?

Patrick Hsu

No, it’s not crazy, because it’s happening every day in our lives. But it’s really science fiction, honestly.

Sarah Guo

I’m just trying to boil it down to a very simple, non-biologist understanding. When you say it’s a platform with this super-tumor where you can pull all of this data out, it’s wild to think about how efficient that is compared with observing one patient type at a time.

Patrick Hsu

I think this is a really interesting point. If you map the number of tokens per experiment across the last 50 years of biomedical research, it will look like the hockey stick that all investors and founders know and love: just going up into the right.

The way we think about doing science is changing based on this. There’s a roiling discussion today about hypothesis-driven versus hypothesis-free research, and whether we should be doing mechanism-focused work versus large-scale profiling.

Honestly, I think this stuff is going to wash out with scale.

Nima Alidoust

Exactly. You don’t have to choose between those two.

Patrick Hsu

That’s my hot take about this era of machine learning and biology. The vast majority of mechanistic data that’s been generated today was really made to ask very specific, well-scoped questions.

Going from a few tokens per experiment to many more tokens per experiment is just going to be the way to do it.

Nima Alidoust

I can say it another way. In biology, what we’ve done is treat humans as the foundation models that ingest information and come up with hypotheses. But now we want to go beyond that, because humans come with their own intuitions and biases.

At UCSF, for example, we often say that we use some of our medicinal chemists, like Kevan Shokat, as the last layer of a neural network. They’ve built this intuition about whether a chemical that was generated by an AI model actually looks like something real.

They can’t even verbalize why they think it might be good.

Patrick Hsu

People criticize these models for hallucinating, but if you think about it, the process of scientific research involves hallucination. That’s what creativity is.

Sarah Guo

You’re all adherents to that lesson in this field as well: that intuition being baked into the models or the process is not necessarily the right thing. We just need to scale the data.

Dave Burke

At least, we hope you don’t have to make that choice. We’re seeing evidence of scaling laws in biology across proteins. That’s been shown in protein language models and across DNA, which is what we’ve shown in our Evo series of models at Arc.

We’re also seeing inference-time scaling laws. In our most recent study, there are early signs of promise. We’ll need good benchmarks, and we’ll have to look at this across different data types over time.

Nima Alidoust

The funny thing for me is that, if you’ve been in this field long enough, you know that every time you take a success from field A and try to translate it to field B, a lot of people—including people in our own organizations—come up with a list of 100 reasons why the learnings from field A aren’t applicable to field B.

But then you’re surprised every time. When you try to do the same thing from B to C, the same kinds of lists start emerging.

Something that’s underappreciated is that the same models that learned human language are learning the language of structural biology. With the Evo work, they’re learning the language of DNA. This is incredible. I don’t think it’s trivial.

If you’ve been in the field long enough, you know there were a lot of people saying that protein language models would never work and that you needed domain-specific models to model these phenomena.

The ethos we have to bring here is that, when Hani says we should use the learnings from those models to translate them here, that’s exactly what we should be doing. We should think about what worked and at least try it in these new domains. Maybe it works, maybe it doesn’t, but if you don’t try, you’ll never know.

This is the domain we’re talking about, and it’s the domain Vivo is excited about. The virtual-cell part of Arc is excited about it as well: the language of systems biology. The first thing you should do is try the things that worked in other domains in this domain.

Sarah Guo

That’s music to my ears. This is one of the only things we have really strong conviction about at the fund we invest out of: that many of these techniques work and scale in domains where people aren’t sure yet. We wouldn’t have expertise in the traditional types of discovery and company-building, but they seem to apply very generally.

I think this is a great segue to a question about why you’re open-sourcing the data.

Nima Alidoust

We generated the data at Vivo, and Vivo is a private, venture-backed company—a startup. When the idea of Tahoe originally came up and Johnny told me, “Nima, there’s this opportunity. We can generate 100 million single-cell data points,” I said, “Can we?”

He said, “Yes, we can.”

I said, “Let’s go and do it.”

Within hours of that conversation—we’re transparent about this—Johnny, Hani, and I are co-founders of Tahoe Therapeutics, and we said, “Let’s do it, and let’s open-source it.”

Why do we want to do that? First, we want to put a new stake in the ground. We want to show that there’s a new game in town and that it’s possible to up our game as a community and as a field.

We want to show people that they should move on from generating 100,000 or 1 million observational single-cell data points. They should up their game and go to a much more massive scale. That’s the first reason.

Second, we wanted to support the DNA of our company, which is to be very small: a small team of superstars rather than hiring a large number of people. Paradoxically, open-sourcing allows us to do that.

We talked to Dave about Tahoe around the night before the new year, sometime between Christmas and New Year’s, and the entire Arc team got really excited about it. If there hadn’t been an open-source aspect to it, it wouldn’t have been as exciting.

The whole community is getting excited about playing with this data and telling us what’s good about it and what’s not good about it. That allows a team of 3 or 4 people in-house to keep it that way and bring in the entire community of like-minded people who have the same vision of building virtual cells to help us in this quest.

For us, the idea was to remove the main bottleneck in doing that. Everyone has been saying that the bottleneck is data.

Dave Burke

The serendipity of this was that Arc is all about mission-driven science and pushing science forward. We were conceiving of creating and launching this week what we’re calling the Arc Virtual Cell Atlas.

The idea was: can we find high-quality, curated data sets and put them out into the world to accelerate virtual-cell modeling? Then we started chatting, and it was, “You’ve got what?”

What we’re assembling this week is this new atlas. The star of the show, in some ways, is the Vivo Tahoe 100 data set. We’re also augmenting it with observational data.

We created something called scBaseCamp. You can almost think of it like a Google crawler and index. We built an agent that goes onto the internet and mines public single-cell RNA data, then curates it in a very uniform way. The result is a very nice observational data set of about 230 million cells.

You add that to the 100 million cells from Tahoe 100, and you now have 330 million cells. This is a really exciting resource for scientists around the world who are interested in modeling at the cell level.

It’s very complementary to have an observational data set that you can potentially pretrain a model on, and then the perturbational data set from Tahoe 100 allows you to bring in those dynamics and make the model richer and more predictive.

We’re super excited about AI agents for science at Arc and across the community. The capabilities are still very early today, but we wanted to show an example of how an agent can do something really useful.

It’s very clear now that basically all dry-lab workflows are going to get automated with agents or copilots. This would ordinarily be the type of thing that a team of computational biologists would be slaving over.

Our core insight was that the Sequence Read Archive is the largest repository of biological data from next-generation sequencing. If you get an NIH grant, for example, you post all of that data online. If you publish in a journal, you put all of the data online as part of the publication.

But it’s extremely fragmented, poorly annotated, and sprawling. There are no requirements for the submission data to be uniform.

We built this agent to crawl all of that data, collect it, organize it, and process it. In doing so, it isolates and removes many of the batch effects and data biases of previous methods.

Hani Goodarzi

These are foundational data sets for the entire field. People work with them, interpret them, and write papers on top of all this data.

Nima Alidoust

Exactly. These data sets have been generated over time, going back a decade. Tools have changed, versions of tools have changed, and genome builds have changed. By simply taking processed data sets and collating them, you’re infecting and contaminating the data with analytical effects and batch effects.

Our idea was to at least remove that. There are a lot of technical and experimental batch effects, and over that span of time, the chemistries of reagents have changed. But at least we can do our part and remove the analytical component.

We were surprised by the extent to which those effects were observable in the data, and by how helpful it was to remove them.

Dave Burke

One thing I would add is that we’re really excited about the leverage of this approach. Sometimes I ask Hani and Johnny, “What does drug A do to cell line X?” There’s this phrase that biologists use: “In my hands, it does this and this.”

A computer scientist wouldn’t say, “In my hands, this model worked.” They might say, “In my computer, in my environment, it worked.”

The genius of what Johnny has built is that this was done by very few hands. Automation is going to scale you to a certain level, but you haven’t even done much automation.

The beauty of Tahoe is that a few people, with a few hands, did exactly consistent work across 60,000 experiments. It’s 100 million single-cell data points, but it’s actually 60,000 drug–cell interactions, or drug–cell-line interactions. Having that done by 4 people reduces the data-set infection that Hani was talking about.

Sarah Guo

This is the first time in history that there’s an opportunity for scientists and entrepreneurs to work on this data set and create these virtual-cell models. How do you tell the quality of one of these models?

Dave Burke

The core idea is its predictive ability. You take a cell and perturb it. You can do that genetically by suppressing or upregulating genes, or you can apply drugs and look at the response.

The measure of the model is how well it predicts what we call the differentially expressed genes, or DEGs. The reality is that today the best models are very poor at this. Their predictive ability for the DEGs is on the order of 10%.

Sarah Guo

Is there an accepted benchmark for this today?

Dave Burke

No, but that’s something else the industry would benefit from. If you think about where we want to go, one of our conjectures is that one reason the models aren’t doing well isn’t simply model structure.

We have a lot of rich structures that we understand in the machine-learning space. The issue is data quality. The hope is that, with this new Arc Virtual Cell Atlas and Tahoe-100M, we finally have a starting point where we can build rich models and get high predictive value from these virtual-cell models.

That’s why this is such an exciting moment in time.

Patrick Hsu

It might be worth speaking plainly about why we even care about virtual-cell models. We have real cells, so why not just do experiments on those?

Biology is very slow. Many of us have tried to pick up pipettes, move clear liquids from one tube to another, grow cells, and make animals. Biology happens in real time.

In the last year of my PhD, my adviser tried to convince me to start an aging project, which would have involved aging animals for 2 years. That’s one experimental round. As you can imagine, I declined. I said, “May I please graduate, sir?”

That’s actually what happens. You’re constrained by biological time, which is completely crazy to me coming from an engineering background. It’s also important to many fields, like neurodegeneration and anything else that takes time to progress.

The idea of massively parallelized, in-silico simulations sounds great, but they need to be accurate. If they’re 10% accurate, you’re just simulating noise.

How do we go from a discipline that primarily respects experiments today to something more like physics, where theory drives a lot of progress? I think these virtual-cell models are a core wedge in making that happen.

Sarah Guo

Can you make that more concrete? If these virtual-cell models work—and we don’t even know how to measure them yet because they don’t exist in a productive way today—what would scientists, the biotech field, or patients expect to gain?

Nima Alidoust

From a drug-discovery perspective, what we’re focused on at Vivo is predicting how a new chemical entity interacts with cells from different patients or patient models. That really is the core of it.

Patrick was talking about in-silico simulation. Can I predict in a computer whether this new chemical structure—drugs are chemical structures, by the way, in case that surprises you—will take a diseased cell, such as a cancer cell, from a diseased state to a healthy state? In the case of cancer, can it kill the cell?

If I can predict that, then my ability to design new chemicals that do that effectively—killing the cancer cell without killing healthy cells—improves massively. That’s what we want to do, and that’s the kind of data we’re generating to train those models.

Johnny Yu

A big part of our future vision and road map is that we think there will be a moment when a drug is generated from a virtual-cell model. The drug will cause a diseased cell to become a healthy cell again.

I think that’s the goal, and it will reshape how we do drug discovery.

Nima Alidoust

There are 2 dimensions of generalizability to think about. One is the cell dimension, and the other is the chemical dimension.

On the cell side, every disease is unique. There are similarities—there are truncal cancer mutations and other things that drive the disease—but there are also very significant individual variations.

You can observe cells from patients, but you can’t do that for every tumor that arises. That’s what these folks do with Mosaic. The idea is that, using a virtual-cell model, you can take those learnings and apply them to all of these new observations that you can make in patients. That’s one dimension.

The other dimension is chemicals. In silico, you have libraries with tens of millions of compounds, and biologics are effectively infinite if you really put your mind to it. Most of these molecules have never existed and will never exist because there’s no use for them.

A model that can traverse that massive space of chemistry and find which parts you need to pay attention to, synthesize, and test would be massively enabling.

Everyone else has libraries that are well behaved—perhaps a couple hundred thousand compounds—and uses fragments and tries to put them together. The process of designing drugs is this slow screening process. This would allow us to leapfrog that entire pipeline.

Patrick Hsu

Ninety percent of drugs fail in clinical trials, so we’re pretty bad at making drugs. I think that implies 2 things.

First, our chemical matter may not be very good in terms of potency, ability to bind the target, toxicity, pH, kinetic profiles, and all of those other properties.

The other possibility is that we’re drugging the wrong target. The idea of virtual-cell models is that they’ll allow you to significantly cut down the search space for the right target. Then you can focus your time on making the right chemical matter or drug composition to make the right kinds of changes in the right kinds of cells.

That’s why mechanism and drug discovery are so tightly interwoven. These models need to help accelerate both.

Nima Alidoust

This is the core reason we need virtual cells in addition to protein language models. Protein language models speak the language of structural biology: what a protein structure looks like, how it folds, how it interacts with a ligand, how a small molecule binds to a protein, or how an antibody binds to another protein.

These are binding questions. You’re trying to see whether one chemical binds to another chemical. But biology is more complex, and there is a context to the protein target we’re trying to hit.

It’s part of a cell. The cell is part of a tumor in the case of cancer. The tumor is part of a broader biological system. Virtual cells will allow us to go beyond the language of structural biology and venture into the language of systems biology.

We can understand how a drug interacts with the broader biological system rather than simply with one target whose binding we’ve cracked.

Sarah Guo

I have a higher-level systems question. At the single-cell level, what about multicellular aggregates, organelles, and organs? Is all of that going to be possible in the future?

Patrick Hsu

I think the first question for virtual cells—or any modeling—is what the right level of abstraction is. Our belief is that the right level of abstraction is the transcriptomic level, because you have these very complex gene pathways.

Whenever a cell is changing in response to its environment, that change will be reflected in the transcriptome. That’s the first question: even within a cell, what’s the right abstraction?

A cell is an exquisite piece of machinery. You could make an arbitrarily complex model of it, but we believe the genetic level is the right level to model.

Going beyond that, you can create very advanced models. You see people doing this with spheroids and organoids. You take mixtures of cells and run them together, trying to simulate cardiac tissue or brain tissue.

What’s really interesting is that, if you have an organoid with 20,000 cells, you can still apply these techniques: take these drug perturbations or genetic perturbations, apply them to the cells, and look at the responses.

You’re going beyond a single cell, but you’re also capturing the intercellular dynamics in the models.

Hani Goodarzi

I’ll make one small comment. It is a single cell that we’re modeling, but that context dependency also captures many of the effects that arise from the environment.

The models we have are actually ex vivo models in this specific experiment for Tahoe, but we also have in vivo models. We have humanized mice that capture some of the immune system of the mouse.

In a way, yes, you’re simulating and building an in-silico model of a cell. But if a model is any good, it can simulate cells in different biological contexts: in the presence of this kind of immune environment, in the presence of one kind of tumor versus another kind of tumor, or in the presence of one mutation versus another mutation.

We call it single-cell modeling, but the whole idea of having so many single-cell data points is that you have those cells in different contexts.

Sarah Guo

That seems like a really important nuance.

Hani Goodarzi

The information about the environment is filtered through the cell. If you’re observing the cell with enough resolution, you can predict it. You can also add spatial data.

Sarah Guo

Definitely. I have a few hot-take questions to end with. Nima, I’ll start with you, because we were having a passionate discussion about why it was important to you that Vivo be a platform company rather than a single-hypothesis company, like 99.9% of biotechs out there. What’s the difference?

Nima Alidoust

The difference is the kind of team you build and the ambition you have.

A single-hypothesis company is based on the idea that the human being—the foundation model Hani was talking about—comes up with hypotheses, and then we test those hypotheses in different experiments.

Companies built around a hypothesis are heavily incentivized to make that hypothesis work. What you see in biotech is that a company will take a drug to the clinic after testing it on 3 different patient samples.

If you’re a platform company, you’re trying to have enough hypotheses and a hypothesis-free way of generating new hypotheses. That means you aren’t wedded to a single hypothesis, and it allows you to be much more scientific in your quest for new drugs or new targets to treat disease.

That’s why we decided to make Vivo a platform company. It allows us to be much more rigorous about what we decide to take to the clinic.

Sarah Guo

There’s been a lot of news recently about the rise of Chinese biotechs. For the core members of the research community here, is that a threat? How do you think about it?

Patrick Hsu

Their cost basis is definitely more competitive. A lot of the discussion in biotech and pharma is how they’re able to do things at this pace and cost, and why their data packages look so good.

They have safety data, toxicology, and all these IND-enabling studies. It’s really competitive. People were surprised by the efficiency of the pipelining and the ability to manufacture all these different antibodies, primarily.

I think that’s great for the industry. Everybody—including patients, investors, and the biotech companies themselves—wants lower cost bases. We want the ability to make molecules that work faster.

All of these things will compete in the system to reduce the currently high cost basis of doing these things here in the United States.

One of the core challenges right now is that we have a wide array of services, CROs, and contract research collaborators that you can try to chain together.

The virtual biotech was previously a concept that was very much in fashion. In reality, when you try to do this, even though it looks good on paper, it’s incredibly slow.

Then people tried the other way: let’s fully vertically integrate and own everything. But that was incredibly expensive. The answer is probably more of a Goldilocks solution in the middle.

We need really competent vendors and CROs that understand the drug-discovery and development process. Then individual companies need to be able to run in a capital-efficient and lean way.

I think the industry is trying to reshape itself around these changes right now to figure out the right way to build startups and the right way to build drugs.

Dave Burke

I totally agree. It’s an important moment. One thing I haven’t seen is an acknowledgment of it. It just hit us in the face.

The United States is the innovation hub, but I think we need to be more intentional about that in biotech. You see innovation in tech, and you see that as the mantra. Innovation in biotech has actually been viewed as one of the things Chinese CROs and companies are good at.

What we’re finding out is that’s not actually innovation. The kinds of things we’re working on—putting big data and AI into the first layer of how we do biology—are what innovation should look like in our space.

If we don’t push that forward as a community, we’re not going to have that innovation in the industry.

Nima Alidoust

This actually slapped us in the face. It caught us by surprise, but one of the first conversations Johnny and I had 3 years ago, when we were thinking about starting Vivo, was about this happening in China and the whole thesis around the commoditization of things we think are massively important, such as molecular design.

There are 2 ways to respond. One is regulatory capture: lobby the government and put limits on how much we can interact with Chinese companies.

The other is to make it part of our ecosystem and change our thinking about business models and how we build teams. To Patrick’s point, do we build a fully integrated team with $100 million in the bank, or a small 14-person team like we have at Vivo?

These are the kinds of things we should be thinking about.

I want to make this into a bigger statement: I think it’s morning in biology. There’s a different game we should be playing here. If you want to stick to the same old-school way of doing things, it isn’t going to work.

The old-school way involves a lot of planning. I was texting with Dave about this a couple of days ago. If I had a penny for every time some massive organization announced an extraordinarily impressive thing and said, “We’re going to give it to you in 3 to 5 years,” I’d be extremely rich right now.

That’s the ethos in biology. You announce some massive thing and say you’re going to do it in 3 to 5 years. But now we have the tools. It’s time to build, and it’s time to do it right now.

That’s how Vivo gets created in a matter of months. That’s how Evo 2 gets created in a matter of months, from the first Evo paper to what happened. That’s how Tahoe gets created.

The second piece is small, super-focused teams of superstars. Massive, vertically integrated organizations aren’t just capital-intensive; they’re inefficient. They move slowly and get bogged down in bureaucracy.

The third piece is associated with this naysaying. In everything you want to do in biology, there are strong biologists who will tell you why it isn’t going to work.

That has to change. We have to think differently, try things out, and use the tools we now have.

Dave Burke

When I talk to pharma companies, they’ll say, “AI and drug discovery are very interesting, but I don’t spend that much of my top-line budget on drug discovery. Most of it is wrapped up in clinical development.”

A lot of them are much more excited about things like natural-language workflows to summarize clinical-trial documents, which are massive regulatory filings, and make it easier to write and read them. They’re also excited about AI-assisted cohort analysis and reducing costs in that part of the cycle.

What they’re going to see as these models get better is that virtual-cell models can help you find the right target, so you can point the cannon in the right direction. You can measure twice and cut once.

The cost basis for the industry should go down, and accuracy should go up.

Sarah Guo

I’m glad both of you brought up the naysayers, because if you weren’t going to, I was going to. I’ve been pitched AI for biotech companies for at least a decade. We haven’t seen a lot of results, and there’s also the natural life cycle of bringing treatments to market. You generally need 11 years or more.

If you were going to leave a broader audience with a single claim about why this is true—there have been different approaches over the last decade, such as computer vision and consumer-scale sequencing data—why should this work now? When should we actually begin to see treatments from these machine-learning approaches?

Dave Burke

I’d go back to analogies. In the machine-learning space, we had artificial neural networks for a long time. People got wrapped up in whether a perceptron could model an XOR gate or whatever it was—around 1990.

Things bounced around for a while. Then we had increased compute, increased data, and more sophisticated models, and we hit nonlinear inflection points.

I mentioned the ImageNet moment in 2009. It drove the development of convolutional neural networks. AlexNet was the model that really showed the way.

Before that, people thought only humans could recognize images at high quality and that computers would never do it. Now we know computers can do it better than humans.

I think it’s the same thing in AI and biology. When I look at single-cell sequencing, the capability is mind-blowing if you’re not a biologist. The idea that, at single-cell resolution, we can look at how expression changes over time is incredible.

You take that ability to generate lots of data, add much more sophisticated models and model training, and suddenly things are happening.

If you look at the Evo 2 model, we trained it on 9.3 trillion nucleotides. We didn’t tell it anything about DNA. We just gave it a lot of DNA from across the planet—every piece of DNA we could get hold of.

What did the model learn? It started learning all sorts of things. It knows where ribosome-binding sites are. It knows about codon degeneracy. One of the things we showed is that it can predict the pathogenicity of BRCA1 variants, which are known to drive breast cancer, with an area under the ROC curve of around 94%.

We never taught it anything. It learned this stuff zero-shot.

I think we’re at that point of inflection now. We’re all in agreement that we’re at a point in time where we’re going to see that inflection. The difference between where we were yesterday and where we are starting this week is going to be the data.

Sarah Guo

Are we somewhere between GPT-1 and GPT-4 in biology?

Patrick Hsu

I’d say we’re closer to GPT-2.

Dave Burke

We’re developing GPT-2, but we don’t have enough data. We need more data.

Hani Goodarzi

If you go a little deeper and talk about different domains, I think protein models are past GPT-3. When it comes to single-cell models and virtual-cell models, we’re closer to GPT-1 than GPT-2 right now.

Sarah Guo

That’s still a pretty exciting timeline if you take the progress in other domains and apply it here. But the difficulty is exactly what you said: with GPT-4, you immediately knew what you had.

If we hit GPT-4 for cell-state models, for example, it will take some time to prove that point. Drug discovery is also subject to the law of small numbers. A platform that takes your success rate from 10% to 30% is amazing, but it’s still 30%, and you need to get lucky.

You still have a drug-development cycle on the order of 10 years, so you have to wait for the model to prove itself. The success rate can slowly go up in a 10-year rolling window.

Nima Alidoust

If we’re all optimists here, I’d say we’re going to treat it as a system. If this was a terribly debilitating bottleneck at the beginning, then hopefully this is a breakthrough.

Sarah Guo

I think that’s a great note to end on. Thank you all for joining me.