🔬生物学正在变成软件——Matt McPartland 与 Neil Patel,Chai Discovery
- Chai Discovery 有意不自己开发药物,而是向药企出售建模和产品层。 Matt McPartlon 表示,Chai 刚成立时,这一判断颇具争议。对话中提到的4家合作方是 Eli Lilly、Pfizer、Novartis 和 Genentech。Neil Patil 称 Chai“几乎是一家中立的软件工厂,专门帮助制造药物”。
- 商业化真正打开局面的是 Chai-2:针对50个靶点设计抗体,其中约一半找到了结合体,平均结合命中率约为20%。 团队最初选择的是有趣的靶点,早期许多靶点没有成功后,转而选择 CRO 目录中已经验证过的靶点。对话中“没有已知抗体结合体”的说法,适用于引用的冷冻电镜/数据泄漏检验所用靶点,并不一定适用于全部50个靶点。
- Chai-1 论文中的一次冷冻电镜验证产生了0.33埃的误差,大约是原子宽度的三分之一。 由于预测结构与电子密度点云几乎完全重合、肉眼看不出差异,团队最初认为结果一定有误。
- 价值不只是加快发现速度,还在于触达传统免疫方法难以实现的分子类型和作用机制。 Matt 提到精准 GPCR 激动剂、多特异性形式和双特异性抗体;在传统筛选中寻找两个彼此独立的结合体,会把难度叠加起来。
- 表位预测——判断治疗分子应该在哪里结合——被认为更难,而且基本仍未解决。 Matt 在不确定的情况下提到,AlphaFold 2/其 multimer 版本在抗体—抗原预测案例中的准确率大约为11%,意味着大多数案例都是错的。他说 Virtual Cell 可能是最接近当前前沿的方向,但距离成熟仍有一段路。
- 算力是生物 AI 面临的结构性逆风。 Neil 表示,在他的例子中,超大规模云厂商和头部 AI 实验室买下了约10,000块 B300 中超过95%,初创公司只能争夺剩余算力。模型中的 pair representation 和 L 的立方次方批处理,对算力和内存的需求不同于 LLM。Chai 还融资了另外4,000万美元,不是4亿美元。
- Chai 的产品刻意做成类似 CAD、而非聊天机器人的形态:它是分子领域的 Autodesk、SolidWorks 或 Figma,提供表位涂色工具,以及类似内容识别填充的结合体生成工作流。 单租户部署帮助解决了药企对 IP 的担忧,公司也会与合作方共同开发专用或微调后的模型版本。
- 这个领域希望从靶点发现、命中发现、优化层层推进的瀑布式流程,转向模型辅助的循环。 Matt 的研究北极星是直接由模型生成类药分子;产品北极星则是迭代式研发活动,最终可以从表位上升到通路等更高抽象层级。Matt 认为可以靠“意志力”消除的瓶颈是验证延迟,Neil 认为则是人才稀缺。关于扭转 Eroom 定律的讨论来自主持人的框架,并非 Chai 声称已经取得的结果。
1. Chai 卖的是建模层,不是药物——这本身就是公司判断
- Matt 表示,Chai 从一开始就认定自己应该成为软件和建模层,这一判断“在当时非常有争议”。团队押注模型最终会强大到足以设计出有用分子。Multimer 结构预测在大约2021年才出现,而 Matt 认为它是打开设计能力的关键;随后逆折叠开始在真实实验中奏效,这得益于 Baker 实验室的验证工作。
- Neil 明确将 Chai 定位为不与客户竞争:Chai 不认为自己是一家要亲自开发药物的 AI 生物科技公司。他称其“几乎是一家中立的软件工厂,专门帮助制造药物”,因此可以支持药企推进药物发现。Matt 表示,Chai 的成功依赖合作伙伴成功。
- 对话中提到的4家合作方是 Eli Lilly、Pfizer、Novartis 和 Genentech。Chai 的核心价值在于规模化的通用能力:Chai-2 之后,团队没有只在1—2个靶点上证明成功就停下来,而是决定“全押上去”。
2. 为什么是抗体——以及现有最大竞争者是一只小鼠
- 对话使用了锁和钥匙的类比:与疾病相关的蛋白质是锁,抗体则是被设计来牢牢贴合它的钥匙。嘉宾将抗体描述为灵活、通用的蛋白质,其框架区域相对统一,结合末端则具有可变性。
- 抗体是 Y 形蛋白,结合发生在两端,框架区域相对恒定,在很大程度上可以从免疫系统已经识别的框架中选择。这使抗体成为有吸引力的设计对象:模型可以专注于“指尖”部分,同时保留熟悉的支架。
- Matt 表示,Chai CEO Josh 喜欢把公司最大的竞争对手称为“小鼠——或者说,在某些意义上,是自然本身”。传统发现流程可能包括免疫接种活动,或大规模酵母展示筛选。后者针对单个靶点至少要搜索数十亿个候选分子,最终可能得到1个、2个或十几个命中体。
- 局限在于,传统方法得到的命中体可能只证明了它会黏在靶点上。研究人员未必知道它在哪里结合、是否具备类药性,或是否具有目标治疗效果。
- Matt 还指出一个不对称性:预测已有抗体如何与靶点结合,出了名地困难;但设计可以更有选择地决定尝试哪些结构。如果模型拥有选择空间,它可能会主动挑选更容易解决的案例。
3. 精准性不只是速度——表位、选择性与交叉反应
- 主持人认为,Chai 的一个差异化优势是“有意图”:用户可以指定结合体应当接触的区域,随后检查经验证的抗体是否以预期姿态结合。这既有助于治疗机制推理,也有助于设计选择性。
- Matt 说,目标可以是一个特定表位,也就是结合位点,甚至可以精确到一组特定原子。历史上,暴力筛选可能只是在分子某处找到一个结合体,用户却无法控制它究竟从哪里“戳”入靶点。
- GPCR 是最典型的案例。Matt 将其描述为细胞膜上的“门铃蛋白”。如果抗体经过工程化设计,能够以特定方式刺激 GPCR,就可能触发下游级联反应,产生激动剂效果,而不只是阻断靶点。
- Neil 用药物开发举例定义交叉反应:一种药物可能需要同时结合人类和猴子的对应变体,才能在猴子身上进行测试。产品和模型可以识别保守区域并对其进行靶向。选择性则是反向问题——结合目标蛋白,同时避开一个相似的人体蛋白,因为意外阻断后者可能导致毒性或副作用。
- 当主持人将其描述成大规模反筛选时,Matt 收窄了说法:重点在于用户可以明确希望结合什么、避开什么。他表示,更大的模型最终可能同时纳入更多这类约束。
4. Chai-1 开源,是一次逼着团队补齐基础设施的工程
- Neil 表示,早期团队在做蛋白质设计时意识到需要 MSA 流水线和大量基础设施。AlphaFold 3 发布后,团队决定开源一个模型,并搭建支撑它所需的基础设施。对于主持人所说、这一项目也是学习如何建设生产级基础设施的方式,Neil 表示认同。
- 当时 Chai 只有5个人。开源 Chai-1 给了团队一个清晰目标,迫使他们以公司级而不仅是研究级的标准建设基础设施。
- MSA,即多序列比对,会向结构预测模型提供许多相关蛋白序列。保守氨基酸和协同突变可以暗示哪些位置在三维空间中彼此接近。主持人将其概括为:从进化中学习哪些变化能够保留或破坏蛋白质;Neil 也认为这种方法能够奏效非常不可思议。
- 早期发布故事体现了资源约束:OpenAI 联合领投了 Chai 的种子轮,5人团队曾在 Mission 区一间几乎空置的 OpenAI 办公室里工作。他们连续48小时赶完论文、网页服务器和技术报告。随后 Josh 在早上大约7:00接受 Bloomberg TV 或类似媒体采访,其余团队成员则躲了起来,导致采访者评论说这家公司看起来根本没有员工。
5. 模型内部:分词器、类语言模型主干和扩散组件
- Neil 对 Chai-1 的概括是:大致由一个分词器、一个 transformer 或类语言模型主干,以及一个类似图像扩散模型的组件拼接而成。这里的分词器并非传统的词语分词器:分子中的原子会被组合成 token,同时电荷、元素类型等属性也会被编码。
- 模型最终必须回到三维原子坐标。扩散组件负责将内部表示转换为预测结构。
- Matt 将 Chai-2 描述为全原子扩散模型。它不只是针对固定序列预测位置,还能设计原子、放置原子、决定哪些原子存在,再将这些原子选择映射回氨基酸。
- Chai-1 是折叠模型:给定序列,预测结构。Chai-2 是设计模型:给定目标结构,生成计划与之结合的候选分子。Matt 表示,大约在录制前1年,这一能力跨过了抗体设计的实用门槛。
- Matt 认为,拥有生物学学位并不是从事这一领域计算工作的前提。他将 AI 生物学比作视频模型:做底层机器学习问题,不需要先成为电影导演。对话提到,Matt 的背景是理论计算机科学;随后又说,除了 Matt 和 Kevin,研究团队大部分人都没有正式的生物学背景。
6. Chai-2 跨过实用门槛:50个靶点、一半找到结合体、命中率约20%
- 主持人用一个类比描述 Chai-1 与 Chai-2:Chai-1 只是识别出一张图片里有一只猫,Chai-2 则是在指定场景中生成一只猫。更深层的区别是,Chai-2 必须同时生成序列和与之兼容的结构。
- Matt 表示,这一过程是迭代式的:模型可以调整结构,询问什么序列能够支撑该结构,再次调整,最终收敛到自洽的序列—结构组合。Neil 将这一过程比作类似 EM 的算法。
- 这次50个靶点的活动部分出于现实考量。团队最初选择了有趣的靶点,但湿实验流程尚在搭建,许多靶点没有成功。Matt 表示,大约一半靶点完全没有奏效,因此团队转向 CRO 已经验证过的靶点,并从这些目录中挑选了一组有趣的对象。
- 结果是:50个靶点中约一半找到了结合体,平均结合命中率约为20%。对话称,正是在这个阶段,药企开始看到该方法可能在其项目中奏效的迹象。
- “没有已知抗体结合体”这一说法属于后续验证案例。对于引用的 Chai-1 冷冻电镜工作,Matt 表示,相关靶点特意选择为没有已知抗体结合体,因此任何命中都将是该靶点首次已知的抗体命中。对话并未证明 Chai-2 的全部50个靶点都是按这一标准选出的。
- 商业层面的结果是,药企和生物科技公司开始主动询问能否使用该模型,Chai 因此搭建了产品,并筹集了服务这些需求所需的算力。
7. 没有直接真值时如何验证设计——以及0.33埃的意外
- 蛋白质设计缺少一个显而易见的 ground truth 指标。常见替代方法是,将设计出的序列输入一个独立的结构预测模型,检查它预测的结构是否与设计结构相似。
- Matt 直言这种方法存在失败模式:研究人员可能把自洽性“刷出来”,尤其是在模型反复生成几乎相同结构的情况下。因此 Chai 还会考察模型置信度和生成多样性。一个既自洽又高置信、但每次都生成同一序列和结构的模型,并没有证明自己具备广泛解决问题的能力。
- 物理验证速度很慢。冷冻电镜可以提供结构证据,但单次验证可能需要数月。Chai 也在开发可能预测实验室成功率的计算机内指标。湿实验反馈周期已经从数月缩短到数周,虽然仍慢于扩展 LLM 评估,但已经足以启动递归式自我改进。
- 在 Chai-1 论文中,团队将预测结构叠加到电子密度点云上,最初几乎看不到任何差异。Neil 给出的误差是0.33埃,大约相当于原子宽度的三分之一。团队认为结果不可能正确,实验室一定是把错误的设计发回来了。
- 被问到是否存在数据泄漏时,Matt 表示,该案例中的靶点特意选择为没有已知抗体结合体。因此,结果不能简单解释为模型复现了一个已知的抗体—靶点复合物。
8. 从 Chai-2.5 到 Chai-3:团队押注模型,而不是解剖失败案例
- Chai-2 之后,团队讨论过如何处理25个没有产生命中体的靶点。一种选择是详细研究它们的共同属性;另一种选择是押注扩大规模、改进模型,最终解决这些问题。Chai 选择了后者。
- Matt 将发布路径描述为渐进式迭代:Chai-2、Chai-2.5、Chai-2.7,最终到 Chai-3,性能逐步提升。
- 结合亲和力是一个主要维度。弱结合体并不够,Chai 希望生成达到或接近治疗级别的分子。可开发性同样重要,包括分子是否安全、稳定、可制造,以及是否不易发生自聚集。
- Chai-2.5 纳入了一项可开发性研究。Neil 表示,团队对这些性质得到的改善幅度感到相当惊喜。
- 主持人认为,与原始结构预测准确率相比,这些辅助性质可能更决定产品的实际效用。嘉宾同意这些性质很重要,但 Matt 表示,结构预测仍是有用基准,因为它拥有相对清晰的 ground truth。
9. 产品是分子 CAD,单租户部署解决了 IP 顾虑
- Neil 表示,最直观的界面本来会是聊天机器人,但团队最终选择了视觉化产品。他将 Chai 的产品比作 Autodesk、SolidWorks 或 Figma,而不是 ChatGPT。
- 这套设计工具类似 Photoshop 环境:用户可以载入分子、涂出表位,再通过类似内容识别填充的工作流生成结合体。产品还包含科学分析和绘图功能,帮助用户检查结果,避免以错误方式对模型进行条件设定。
- 药企对 IP 极其敏感。Neil 表示,早期人们告诉他,客户绝不会把专有数据放入共享平台,并在那里生成新药。他的安全背景帮助其设计了严格的数据分隔和单租户部署,使每个客户都拥有独立版本或账户。
- Eli Lilly 是早期合作方之一,曾与 Chai 密切合作开发第一版设计工具。对话称,在许多交易中,Chai 会与合作方共同训练或微调一个针对该合作方的模型版本;但对话并未说明这些细节对每一笔交易都是公开信息。
- 关于科学家的产品采用,Matt 表示,结果会创造“启动能量”。药企科学家很务实,可能先拿一个此前长期未解决的靶点交给 Chai。只要结果足够有说服力,他们就愿意试用产品。
- Matt 将此与自己销售安全产品时的经历作对比:当时他觉得客户在技术上没有那么成熟。而在 Chai,他表示合作方往往包括已经研究某个靶点5年、10年或20年的科学家。
- Matt 讲述了一位药企科学家的故事:在一次评审中,她因为 Chai 帮助为一个自己研究了10年的靶点生成了第一个初始结合体而落泪。
- 内部严谨性来自 Nathan Rollins 等人的加入。RJ 描述 Rollins 14岁时就开始在 Baker 实验室工作,18岁从 Harvard 毕业,大约21岁时在 Marks 实验室拿到 PhD。他起初持怀疑态度,但看到 Chai-2 的结果后表示,自己需要“把这件事做得无懈可击”,现在还没人应该庆祝。
10. 从瀑布到循环:两个北极星,以及一款注定被替代的产品
- RJ 将靶点发现、命中发现和优化描述为瀑布式流程,每个阶段都有一道闸门,耗时可能从数月到数年。早期尝试的成本很高。如果模型能够生成有潜力的候选物,流程就可能变成更像敏捷软件开发的循环。
- Matt 表示,Chai 不希望仅仅因为初始分子通常很差,就继续保留这些阶段区分。公司的北极星是直接由模型生成类药分子。他承认这会遇到许多障碍,包括更好的条件控制和强化学习系统,但认为目标可以实现。
- RJ 描述了两个表面上相互冲突的目标。研究侧希望生成越来越好的从头设计候选药物,可能先从拮抗剂等更容易的类别开始;产品侧则希望支持迭代式工作流,让实验室结果成为下一轮模型运行的条件输入。
- 两个目标可以共存,因为更难的分子类型需要更高层级的产品工作流。RJ 举例称,激动剂、双特异性抗体和 ADC 都属于这类对象。Neil 追问,模型如何可靠地一次性“拨动”细胞上的开关;这些正是产品未来需要支持的下一层抽象。
- Neil 表示,Chai-4 可能淘汰当前分子可视化产品的可能性让他“一度非常存在主义”。他接受产品可能只存在大约1年,将其作为交付价值、推动更好研究的桥梁。
- 今天的产品更像分子领域的 Cursor:用户检查化学键和分子性质。后续版本可能围绕表位选择编排研发活动,再围绕一条通路中的多个靶点编排活动,最终抵达“科学的外循环”。
- Chai 不运行自己的药物开发管线。内部科学团队维护基准集,其中既有已知治疗药物对应的靶点,也有专门用来推动模型能力的靶点。他们开展实验是为了验证和改进模型,而不是开发这些药物。
11. 表位预测是更难的问题,两位嘉宾都认为它仍然棘手
- 主持人认为,表位预测可能比寻找结合体更难。Matt 表示同意,称这是“荒谬地困难的问题”,因为它需要全局生物学背景,才能判断什么在发生相互作用、以何种方式相互作用。
- Matt 区分了几个层级:找出哪些蛋白质导致疾病,理解这些蛋白质如何相互作用,以及确定应该打断哪一种相互作用。需要阻断的具体位置就是表位。
- 当被问及假设出现 SARS-CoV-3、新型流感病毒或类似事件时,Matt 给出了一套带有保留的流程:运行结构预测模型,检查模型预测的结合位置;如果模型置信度很高,就使用该位点。他随即强调,结构预测总体上还没有解决。
- Matt 用“I think”和“like”等保留表达提到,AlphaFold 2/其 multimer 版本在抗体—抗原预测案例中的准确率大约为11%,意味着大多数案例都是错误的。对话并未确定这一11%数据的确切适用范围。
- 主持人解释说,抗体通过重组不同组成部分来识别陌生病原体,因此不像其他蛋白质那样拥有可借鉴的保守进化模板。Matt 认同这一框架。
- Matt 表示,Virtual Cell 可能是解决更广泛背景问题时最接近当前 state of the art 的方向,但距离成熟仍有一段路。
12. 经济学上的反问:重点是新分子类型,而不只是节省200万美元
- RJ 提出了经济学上的疑问:在一项成功的药物开发项目中,抗体发现可能只需在内部花费几百万美元,而整个项目可能耗资约5亿美元;通常被引用的26亿美元则包含了失败项目。如果 Chai 只是节省发现阶段的成本,产品为什么会如此有价值?
- Matt 首先质疑了这个前提。Neil 强调,精准设计抗体本身非常困难。Matt 更广泛的回答是,Chai 能够打开传统免疫方法难以触达的靶点、机制和分子类型。
- Matt 提到了 Chai-2 设计 GPCR 激动剂的案例。抗体可以被工程化设计为精准命中 GPCR 上的开关,而不是只在蛋白质某处结合。
- 他还提到多特异性和四头形式,这些分子必须从第一性原理出发进行设计。在双特异性抗体中,两条臂都必须分别结合不同靶点。如果每条臂产生结合体的概率都是十亿分之一,传统方法就会面临乘法级难题。
- RJ 补充了组合层面的价值:Chai 帮助的不是单一药物,而可能是合作方的一整个靶点组合。Neil 表示,平台让合作方能够在某个子领域集中学习,并将这些学习成果应用到多个项目中。
- 公司的使命是将药物发现“从科学实验变成工程学科”:用户应该能够声明式地定义希望获得的治疗性质,再由模型填补缺口。
13. 算力面向 LLM,工程基础组件反而成了瓶颈
- Matt 认为,阻碍科学走向工程化的首要问题是数据基础设施。生物学文件格式包含多份结构副本、未解析区域、替代位置、结构解析方式等大量边界情况。工程问题在于:是编码大量特殊情况,还是采用更简单、更有原则的方法。
- 当主持人提出是否可以让 LLM 处理解析时,回应是 LLM 未必知道所有相关边界情况。Chai 更偏好人类能够理解和审计的代码。
- Neil 的答案是算力。在现货和按需市场经历容量紧张后,Chai 决定自己购买硬件。他以市场上约10,000块 B300 广泛出货为例,称超大规模云厂商和头部 AI 实验室买走了超过95%,初创公司只能争夺剩余部分。
- Neil 表示,B300 和 Vera Rubin 等新系统是按 LLM 需求设计的,配有大型 KV cache,以及72块 GPU 彼此通信的配置。这些系统可能提供不错的性能,但算力和软件市场尚未针对生物模型优化。
- Matt 解释说,AlphaFold 2/3 类模型处理的是 pair representation,其规模大致为序列长度 L 的平方,而不是 L。批处理因此可能达到 L 的立方次方复杂度。结果是每个 token 的算力消耗很高,内存带宽开销也很大;即使层归一化也可能成为实质性瓶颈。
- Neil 另外指出,triangle layers 在现代 GPU 上成本高、效率低,因为它将较小的隐藏维度与较大的序列维度结合起来,形状正好与 GPU 最擅长处理的方向相反。他将其描述为用参数换算力。
- Neil 的基础设施解决方案是通过 Temporal 实现持久化执行。分布式模型调用可能因为不稳定的存储桶、数据库或 GPU 而失败。没有持久化执行,团队大部分时间都会耗在队列、重试和编排上。
- Chai 用 Temporal 处理数据库调用、副作用、模型调用和长时间运行的数据流水线,使失败可以自动重试,而不必为每种情况单独编写队列逻辑。
- 临近结尾时,Neil 表示 Chai 又融资了4,000万美元,需要购买另一座算力集群,以支持更大规模的推理和训练。他明确说不是4亿美元。他的结论是,可靠的工程基础组件将为更大胆的生物学工作提供支撑。
14. 简单性与 Bitter Lesson 的冲突——以及对数据效率的反驳
- Matt 表示,系统复杂化与“被 Bitter Lesson 彻底说服”在根本上存在张力。他以 AlphaFold 3 为例,谨慎地说其可能有23个子模块。复杂到这个程度后,很难理解修改一个模块会如何影响整个系统。
- Chai 办公室墙上挂着 SpaceX 的 Raptor 1 和 Raptor 2 发动机照片,提醒团队删除不必要的组件、简化系统。
- Neil 反驳说,AlphaFold 2 和 AlphaFold 3 的成功部分来自它们相对小型、数据高效、计算密集,并围绕大量归纳偏置构建。若移除这些偏置,就需要新的数据来源,或对数据采取根本不同的处理方式。
- Neil 提到一篇 Apple 论文,该研究在一个非常大的、源自 ESMFold 的数据集上进行蒸馏。他表示,模型确实获得了有用信号,但无法泛化,因为它是在做模式匹配,而不是推理。
- Matt 同意蛋白质结构数据很难处理:序列数据非常丰富,但实验结构数据少得多。序列数据帮助了 ESM 类模型,而将同样的方法直接应用于实验结构数据则困难得多。
- 他的解决方案是提炼 AlphaFold 中有用的思想,而不是复制整个架构。Triangle layers 就是一个例子:即使代价高昂,它仍然代表了一种有价值的归纳偏置,可以被修改并进一步构建。
- Matt 将此与自己的理论计算机科学背景联系起来。他的导师告诉他,永远不要低估多项式时间;Matt 的第一篇论文使用了一个 n 的20次方时间复杂度算法。更广泛的启示是,可以先允许解空间足够宽,再在后续阶段进行简化。
15. 商品化、数据护城河问题,以及药企作为资本配置者
- 主持人将蛋白质设计描述为拥挤赛道,可能已有10家或15家初创公司。他们还回忆了 RFdiffusion 发布后不久的一个案例:有人生成了皮摩尔级结合体,并用冷冻电镜完成验证。对话没有证明这些具体是 mini-binder。
- Neil 预计,一些分子类型会逐渐商品化,但更具挑战性的任务仍将由前沿模型服务。他将这一局面与 LLM 类比:开源模型可以覆盖部分使用场景,但闭源前沿模型凭借解决更难任务的能力,以及与之配套的更好产品,仍然捕获了大量价值。
- 他表示,产品层与模型层同样重要。如果缺少周边工作流和工具,用户未必会选择开源模型。即便 AGI 最终能够一次性完成所有工作,他仍预计产品在相当长一段时间内都很重要。
- Matt 认为,生物学缓慢的测量循环能够抵御即时商品化。公开序列数据可以构建基础模型,但有用的测量需要时间,也必须不断迭代。他表示,AGI 未必能立刻解决这些技术瓶颈。
- Neil 否认 Chai 不可能拥有数据护城河。他将公司与 Anthropic 和企业合作的情况作比较:企业数据未必可以被用于训练通用模型。Chai 正在投资于把算力转化为数据的方式,并利用合作方的需求来指导研究。
- 在许多交易中,Chai 会与合作方共同训练或微调专用模型版本。合作方可能拥有实验数据或偏好的性质,帮助模型在特定类别的靶点上表现更好。有时提升可以非常具体,例如确保生成的设计具备合作方偏好的某种性质。
- Neil 表示,Chai 没有建立自己的药物开发管线的计划。他看重的是激励一致性:改进模型、帮助合作方成功,再利用这些结果改进产品。
- Neil 认为,药企 token 的下游价值异常高,因为成功药物可以成为价值数十亿美元的资产。他表示,两款 GLP-1 药物合计可能代表一项价值1万亿美元的资产,并保留了相应的不确定性。
- 主持人而非 Neil 表示,直到大约3个月前,GLP-1 收入还超过 AI 实验室收入总和。另一位发言者指出,Genentech 曾是硅谷早期最重要的创业成果之一。
16. Eroom 定律、愿望清单与核心结论
- 对话将 Eroom 定律定义为倒过来的 Moore 定律:药物开发成本不是下降,而是在上升。主持人表示,这最终会使新增药物开发的边际回报转为负值。这是主持人的框架,并非经过测量的 Chai 结果。
- 另一位发言者推测,Chai 这样的公司可能帮助翻转或弯曲这条曲线。对话没有证明 Chai 已经做到这一点。
- 团队反复使用资本配置作为运营隐喻。Chai 将约10人的研究团队描述为把想法配置给算力;更广泛的公司在对话中约有30人,被描述为借助 AI 工具配置注意力和算力。
- Matt 认为可以用意志力消除的瓶颈是验证循环:能否即时知道一个蛋白质设计假设是否有效。他表示,这个领域目前仍然需要在黑暗中摸索。
- Neil 的答案是人才稀缺。许多技术能力很强的人进入 LLM、软件或 SaaS,相比之下进入计算生物学的人很少。他认为,部分原因在于这个领域不为人熟知,也缺乏视觉可访问性,而 Chai 正试图通过网站和产品改善这一点。
- Neil 的收尾判断是,生物学正在开始转向声明式的精密工程。折叠模型的误差已经在1埃以内,设计模型的命中率超过50%,这意味着一个96孔板里可以有48个有趣的结合体。他将这一转变类比为软件、Cadence 中的电路设计,以及机械 CAD。
- Matt 的收尾判断是,这个领域确实正在奏效:生命迹象已经出现,商业牵引力已经存在,同时仍有大量研究问题和容易摘取的机会。他表示,三维几何、扩散模型、类语言模型主干与核心机器学习的结合,使这一领域异常宽广,也可能产生重大影响。
It looks a lot less like a ChatGPT and a lot more like Autodesk, SolidWorks, or Figma, if you’ve used those things where you can load up your molecule. There’s almost a Photoshop-esque design suite. You have the equivalent of a paint tool to paint your epitope and the equivalent of a Content-Aware Fill tool to get your binders generated from Chai.
To add to what Matt’s saying, this notion of target discovery, hit discovery, and optimization—where each of these has a gate and takes a few months to a few years—is a very waterfall model. The cost of trying things and getting things early is very expensive. But if you start to get into a regime where models can give you really promising candidates, you can start to make that look a lot more like a loop. It’s akin to becoming more agile in software development.
1. Intros: Matt and Neil
But now the next problem is agonists. How do you reliably one-shot hitting a switch on a cell? Or bispecifics or ADCs? I think these are the levels of abstraction that we’re going to have to climb with the product as the models get better. If you have really good primitives for structure prediction, binding, and design, and you can compose them, then you can start to grow into the outer loop of science.
Welcome to Latent Space AI for science. I’m Brandon. I build RNA therapeutics at Atomic AI. I’m joined by my co-host, RJ Honicky, CTO and co-founder of MiraOmics. It’s a pleasure to have with us in the studio today Matt McPartlon and Neil Patil of Chai Discovery. Chai is a protein-design startup that’s about 2½ years old and has made quite a splash in those few years. You have several very exciting announcements that I think you’ll tell us about today. To get started, could you two give us a bit about your backgrounds and what you do at Chai?
Thank you very much for having us. We’re super excited to talk about Chai today. I’m Matt McPartlon, one of the co-founders of Chai. My background is in AI- and biology-related work during my PhD. I actually started my PhD in theoretical computer science and then transitioned to this later.
I’ve been doing this stuff for about 8 years, and I came into the field at an interesting time, when protein structure prediction was just starting to see signs of life. This was during the AlphaFold 1 days. I was in the field during AlphaFold 2 and got to see a lot of the interesting developments at that time. I’d always been interested in applying this stuff in the real world, and Chai was a perfect opportunity to do that.
I’m Neil Patil. I help lead platform and product here at Chai, so a lot of the work around the infrastructure to train models, serve them, and then the productization piece—the design suite that lets you use the models.
I have a more meandering path. I got into programming about 15 years ago, making apps in the App Store, and got really addicted to the dopamine hits you get from that. Then I got nerd-sniped by robotics and worked on that for a bit, including self-driving cars in 2018 and 2019. I got really jaded and decided I didn’t want to touch hardware for a while.
I ended up switching and joining a SaaS company called Vanta as one of the first employees there and grew with it. I started my own security company afterward. A few years into that, I thought, “You know what? Atoms are kind of cool. I want to work on something a little more meaningful.” I joined Chai about a year ago, right after Chai-2 was announced, to help with a lot of the platform and commercialization pieces.
2. Four pharma partnerships
Awesome. It’s like the 5 stages of grief or something.
Yeah. We’re at acceptance.
Awesome. You have, I think, 4 big partnerships now and have raised a whole bunch of money. Can you tell us a little bit about those partnerships? What I really want to know is, what are you telling investors and customers that’s so compelling that they’re willing to do these big deals?
We’ve been very fortunate to partner first with Eli Lilly and then with Pfizer, Novartis, and Genentech. It’s been a really interesting ride, and I think our business model is also very compelling to a lot of people. We really care about the partners succeeding. Chai as a company really depends on how the partners succeed.
Neil probably has some interesting takes on what we actually offer and what makes that so compelling, so I’ll hand it over to him.
As you all know, drug discovery is a very lengthy process. A lot of these pharma companies are spending years and years and billions of dollars trying to find initial therapeutic candidates. At Chai, we train models that can help accelerate that process and find those initial binders and then some.
There are a lot of bio companies and AI-for-bio companies that are making their own drugs. We really don’t see ourselves that way. We see ourselves as almost a neutral software factory for making medicines. That’s what lets us work with and support all of these other pharma companies in their drug-discovery journeys.
3. The software-layer bet
A lot of this capital is just another proof point that we can start to accelerate that software factory, go after harder modalities, train bigger models, and ultimately build what our partners and customers ask us for.
But what is it? Why you and not other structural companies? Why are they compelled to buy from you?
The thesis of Chai has always been to be the software and modeling layer, which I think was very controversial at the time.
That was 2 years ago, and it’s already a completely different world. Yes.
Yeah, it’s pretty crazy. People tried this play for a while, and I think the models just weren’t there yet. Even for us, we were taking a risk in the very beginning. We were banking on the models getting there.
I had seen early signs of life in my work, and our CEO, Josh, who was on the original ESM papers team at Meta, was seeing pretty early signs of life that there might be scaling laws here. We thought we’d actually be able to start designing things. Structure prediction is getting really good.
One crazy thought is that we didn’t have a multimer structure-prediction model until 2021. That was 5 years ago, when we could start using deep learning to actually predict the shape of 2 proteins at once. AlphaFold 2 was this huge breakthrough, but then AlphaFold-Multimer came out a year later. You really needed that to unlock design in the first place. We weren’t even trying to predict multiple proteins at once.
Then inverse folding started working, and we thought, “Oh, protein design—this actually works in the lab.” Credit to the Baker lab for doing all this excellent lab validation on their models, but we’re starting to see them do interesting things and actually work in real-world experiments.
Now is probably the time to start betting on this. Before then, maybe you could take some experimental data from a campaign on 1 target that you had and cared about, and you might be able to make some progress on that and keep hill-climbing in this 1 very specific case. General models weren’t really a thing back then.
4. The 50-target challenge
We took that bet seriously and decided to push as hard as possible and shoot for generality in our approach. When Chai-2 came out, our second paper after Chai-1, we showed the world that this was actually possible and possible at scale. We didn’t show this for 1 or 2 targets. It kind of worked, and we thought, “Let’s just go all in.”
Josh likes to say we set a bold companywide challenge to design antibodies for 50 targets, and we actually saw some signs of life. We thought, “All right, let’s do this with real statistics and see if this actually works.”
It’s an interesting story how we chose these targets. We asked, “What targets are we going to choose? We should choose some interesting targets.” At that point, we were ramping up with CROs and figuring out what our wet-lab process looked like. After trying some things with many proteins, we said, “Here are the interesting targets. This is what we should look at.”
Half the time, the targets just didn’t work. We were still learning, so we thought, “Maybe we should just go with targets that the CROs have actually validated. Let’s get the CRO catalog and see what they’ve already worked on, and restrict that to an interesting set.”
5. What is an antibody, and why target it
From that, we chose 50 targets, designed antibodies against them, and got hits for half. At that point, I think pharma started to realize, “Okay, there are actually signs of life here, and this might work in some of our programs.” So antibodies are maybe a more challenging domain than other structure-prediction problems.
So why tackle antibodies? Maybe back up: what is an antibody?
Yeah, and what do you do with it, and why is it an attractive target? The analogy that everyone gives is this lock-and-key kind of problem, where your target—this protein that you're trying to bind to—might be some disease protein that's kind of like your lock. Then you want to design this key that fits into it and, in our case, just sticks there.
The interesting thing with antibodies is that they're really flexible, general proteins. In a lot of ways, they're very general; in a lot of ways, they're actually pretty uniform. But at least the way they bind to a target is very general, so you have a lot of optionality in how you design this kind of binding interface.
The structure-prediction problem for antibodies—predicting how this antibody actually binds to the target, how the key fits into the lock—has been a notoriously difficult problem. The nice thing is that we've made a lot of progress on structure prediction. The field as a whole has come a long way in getting structure prediction to where it is.
But in the design setting, you can be a lot more selective about the types of designs you want to make and the types of structures you actually want to focus on. In some cases, it might actually be even easier to design a protein binder that is an antibody than to actually predict how it might bind that target in general. If you have the freedom to choose, you can just pick the easy cases, if that makes sense.
With antibodies, there's a whole machinery in the body that works with them. What does the body do with them naturally, and what can you do with them that is sort of not natural but useful for therapeutics?
This is coming from a non-biologist here, but I think of antibodies as these Y-shaped proteins. They kind of look like a peace sign with your fingers. Each of these fingers is kind of like an arm of the antibody, and it's actually only the tips of your fingers—the tips of the antibody—that engage in binding.
So this makes them really nice therapeutic design targets for that particular reason. The nice part is that the rest, apart from the tips, is actually relatively constant. This is called the framework region of an antibody. In the design problem, you're typically just designing the very fingertips, and you can choose, for the most part, these framework regions that your immune system already recognizes.
6. ADCs, bispecifics, and pressing the switch
Antibodies are these Y-shaped proteins that your immune system recognizes. It knows them really well. They're kind of one of the lines of defense in your body against pathogens and other types of diseases. I guess antibodies can, on one end, connect to proteins on the surface of a cell, typically, or other things, and then the other end helps the immune system identify a pathogen, typically. But you can also do things like you mentioned: ADCs, antibody–drug conjugates. That means putting a drug on the other side or something like that, which causes the drug to be released into the cell when you bind to something.
Right. They're this very general framework where, on the ends, you have these CDR loops, and you can design them to bind to arbitrary things. Maybe on one end you bind to a cancer cell, and on the other end you bind to a toxic molecule. You're now precision-delivering that toxic molecule to a cancer cell, right?
Or you just have two ends bind to things and force induced proximity to have some effect in the body. A lot of drugs historically are really just about blocking things—antagonist behavior. But maybe you can have agonist behavior. You can really precisely press a switch.
There's a GPCR, or G protein-coupled receptor, which is one of these doorbell proteins that sits in your cell membrane. You can have an antibody very precisely engineered to poke it in a certain way that causes a downstream chain reaction. One of the things that's really exciting about where we're getting to with some of these models is that we can start to get that precise.
7. How this was done before: mice and yeast display
We can really target a very specific epitope, meaning a binding spot, or a very specific set of atoms for the antibody to go after. Historically, with a lot of drugs, you're just brute-forcing a lot of antibodies and trying to come up with a bunch of things to see what sticks. Maybe that gets you a binder to some spot on your target molecule, but that doesn't let you precisely engineer where you're poking it.
I know you're not biologists, but do you have any idea how they used to design these before these models came up?
Or what is actually still the state of the art in terms of drugs that have made it to the clinic?
Yeah. Josh, our CEO, likes to say that our biggest competitor is the mouse—or nature, in certain ways. Traditionally, these types of drug-like molecules were either discovered in immunization campaigns. You would literally just infect a mouse with a disease and see what antibodies it makes to try to combat that.
Other ways of doing this are super-large yeast display, and so on. You might start with, “Hey, I really like this framework. How am I going to figure out the right loops to design to bind this target?” Then you try as much as you possibly can and literally search for a needle in a haystack. This would be on the order of at least billions of potential molecules that you're screening against this 1 target.
In that case, you might end up with 1, 2, maybe a dozen potential hits to this target. You don't know much about those hits. All you know is that they kind of stick to the target. You don't necessarily know where, or whether they're even drug-like.
I think one big separator of Chai—and a thing that our partners definitely like to see—is that you can be really intentional with how you want to do this design process. You can say, “I want to bind this target in this particular area.”
You can even go back and look at the designs after we've validated them. You can go back and look and say, “Is this antibody engaging the target in the way that I expect? Do I think this will actually have the therapeutic effect that I'm going after?”
8. Selectivity and cross-reactivity
One of the cool things about knowing that you have the right binding pose is that you can now also design selectivity into it. Does your platform have some technique for selectivity?
Yeah, there's a nice mix of ideas that went both into the modeling side and especially into the product side for dealing with selectivity and cross-reactivity. In some cases, you want your molecule to bind 1 target and avoid another one. So you might have a healthy variant of a protein and a disease variant of a protein, and you want to avoid the disease variant. Or you might have some other similar protein that's not actually harmful in your body that you don't want to artificially block.
On the modeling side, we've come up with ways of doing that, but I think it's even more interesting on the product side. How do you enable customers or partners to actually intentionally design for these things?
Yeah. And maybe to back up and define cross-reactivity: it turns out that when you're developing a drug, you're not necessarily going straight to injecting that into a human. You might want to put it in monkeys first, for example. The monkey might have a mostly similar but slightly different variant of it, so your drug not only needs to bind to the human variant but also the monkey variant.
The way we've tried to approach this in the models and the product is to let you account for those very general cases where you say, “Hey, I'm trying to design something that can bind to both of these things so that I can actually go and develop the drug.” You can identify the region that's conserved and then target that. Conserved means it doesn't change much between the 2, and you target that exact region.
Similarly, with selectivity, there might be a very similar protein in the human that, if you accidentally bind that one, is very bad. You only want to bind the target protein. That's why a lot of drugs fail, are toxic, or have really bad side effects. You're kind of having this combinatorial problem of binding only these things and avoiding only these.
What's been really exciting with some of the progress recently is the improvements we've been able to make in the level of specificity we can get to with those models.
So you're not only designing the binder, but you're also making sure that it doesn't bind to another thing. Are there other ways that CAR-Ts have tried to tackle this by having some molecular or some sort of signaling pathway that says, “If I bind, I only fire if I bind this one and this one doesn't bind”? But you're saying you just design an antibody that will actually bind only to the thing that you care about.
We're getting to the point where, in some cases—it's nuanced, right?—you can actually try that.
Okay, that's amazing. So you're saying you essentially call it a counter-screen, or you have, as part of your platform, a way to reliably counter-screen against a large, diverse set of proteins that might be issues downstream?
9. Chai-1 and the MSA detour
Yeah. I would say the framing is more that you can be very specific about what you care about binding versus what you care about avoiding. But I think, for example, a lot of the money we're raising now will let us train bigger models that can maybe be even more general and start to account for even more things at the same time, right?
Maybe we should back up. Let's talk about the history of the Chai series of models.
Well, why don't you tell the story?
We started Chai around 2.5 years ago. For the first couple of months, we were like, “All right, we're going to work on protein design.” We were working on this and making some progress. We were like, “Oh, this is pretty interesting.” We had some ideas and models.
Then that was right when AlphaFold 3 came out. We'd been talking about, “Man, we really need an MSA pipeline. We need all this infrastructure.”
MSA is multiple sequence alignment.
Why is this? We've covered this before, but what is an MSA in 2 sentences, and why is it important?
If you want to predict the structure of a protein, it might be really useful to see a bunch of very similar protein sequences. What those really similar protein sequences tell you is which positions—or which amino acids—end up being conserved across many variants of this protein. If you see high levels of conservation, or high levels of mutation like correlated mutations, it typically gives you some indication that these amino acids are close in 3D space. So you kind of have this 2D view of a protein, which can then be used to help you predict this 3D structure.
So you're learning from evolution what was conserved because the things that weren't conserved probably broke the protein, and something died or didn't make it.
Exactly. It's pretty remarkable that this works, honestly. One of my favorite bio facts here.
We were thinking, “Oh man, it'd be nice to have a lot of infrastructure and whatever.” Then AlphaFold 3 came out. We were like, “Hey, we should open-source this model. We should just bunker down and build all the infrastructure that we need.” I think this will pay back in the long term, for sure, both as a forcing function to get us where we are and to contribute to the community as a whole.
So it's interesting that you chose—okay, what we're actually doing here is building a model, but what we're really doing is learning how to build the infrastructure. Is that kind of what you're saying?
10. Sitting in the OpenAI offices
Yeah, that's exactly right. I had built a lot of similar infrastructure in my PhD, but not at a production level for a company. At that point, I think there were 5 of us at Chai, and we were like, “All right, this is our forcing function. We have a clear goal to work toward. It's very direct. Let's get this thing going and see how fast we can do it.”
You guys were, at this time, sitting in the OpenAI office?
We were sitting in the OpenAI office, yeah—in the Mission.
Right. So what's the backstory on that? It's really interesting.
Two of our other co-founders, Josh and Jack, had a relationship with some of the OpenAI people, actually. OpenAI co-led our seed round. So we were thinking, “All right, should we get an office while we're only 5 people?” It turned out that office was mostly vacant, so we got to sit in the OpenAI offices for a while. We built Chai 1 open source and learned about infrastructure.
Yeah. So then after that, we really set our sights on protein design.
And it's worth pointing out: Chai 1 was a structure-prediction model, right? You have the sequence; what is the structure that it folds to? And then that was Chai 2.
Yeah, Chai 1. Chai 1's finished. One other crazy story there—let's see if we can actually share this. This is a hilarious one.
We were like, “Oh man, we really want to be the first to put this out.” We were like, “Okay, we're 1 week out. The model's almost done training. Should we build a web server?” Then we were like, “Oh yeah, maybe not.” But we ended up spinning up this whole web server so people could use it rather than just download the Git repo. It's kind of annoying, especially for biologists, and we actually wanted people to use this, so we thought, “Let's spin up a web server. Let's get the technical report out,” and all this stuff.
We ended up being up for 48 hours straight, just getting the paper over the line and getting the last things done on the web server. Then Josh was interviewing with Bloomberg TV or something that morning. We'd been up for 48 hours straight, so Josh ran into a room to do this interview on Bloomberg TV. I think it was 7:00 in the morning. Everyone was in the office, and we didn't want to be seen. The interviewer was like, “Interesting company. It doesn't look like there are any employees here.” [laughter]
Yeah, it was a really fun time. The early startup days were just super fun. After that, we set our sights on design, and what we were thinking is, we always had antibodies in mind. We thought of this as the most tractable problem. The nice thing with proteins is you have this beautiful sequence representation, and there's already a lot of research on how to autoregressively generate sequences. This sequence-generation problem is well studied.
So we were thinking, what's a nice area to apply sequence generation to in the bio space? It's pretty natural to do linear sequences of amino acids. So we started working on design. A unique thing about Chai is that we're not just designing antibodies—we're not an antibody company. We don't really pigeonhole ourselves into 1 therapeutic area.
We tried to tackle this problem very generally. We were thinking: Can we design many proteins? Can we design antibodies? Can we scaffold protein complexes? We wanted to take a holistic view of how you design proteins in general. That eventually led to Chai 2. That was our first flagship design model, and that's where the Chai 2 paper and our BOLD target discovery project came in.
We designed antibodies to 50 targets for that paper, got binders to about half of them, with, I think, on average, around a 20% hit rate for binding. Afterwards, we started working on Chai 3. That's our latest series of models. But let's take a break there.
11. Anatomy of a structure prediction model
Before we talk about Chai 3, can you tell us, especially for listeners who may not be familiar with structure-prediction models, what the model looks like? How does it work in general?
Let's take a look at Chai 1. Chai 1 has, roughly, a tokenizer, a transformer, something that looks like a language model, and then something that kind of looks like an image-diffusion model, and they're all stitched together. The tokenizer is not your typical WordPiece-style tokenizer. This is like: I have a bunch of atoms in a molecule, and now I want to pull those into what I would call tokens for my LM-looking trunk. Then that conditions this big diffusion model, which will emit the image, which is some 3D structure.
So is it atoms or is it amino acids that are the input?
That's an interesting question as well. We have all these different input tracks. One thing about biology is that the data is inherently multimodal in a sense. You have this token-sequence representation. Each of these tokens has a set of atoms that kind of dangles off it. You also have some properties of the different atoms: an atom might have a different charge, or it might have a different element type—the periodic table of elements.
These all get bunched together into tokens. Once tokenized, you can process this in very standard ways. But ultimately, you have to get back to these 3D coordinates. In order to predict the structure, this is just some 3D object, and to emit that object, you go through what looks like an image-diffusion model, where you go back from tokens to the atom representation.
I see. So the tokens go in, the transformer establishes the relationship between the different tokens, and then the diffusion model turns that latent representation into a 3D structure.
That's exactly right.
Okay, great. So that's Chai 2. That was Chai 1.
Okay, okay. So Chai 1 was a folding model. Yeah. All this biology stuff sounds kind of scary: atoms, tokens, amino acids. My background, personally, is theoretical computer science. That's what I spent all my earlier years doing. I transitioned to this pretty late in my PhD.
12. Chai-2: all-atom diffusion and crossing into design
But I think the background that you need is really similar to the background that you need for any other field of machine learning. There are all these domain-specific things that you learn about. One analogy, or anecdote, I like to say is that people think you can't work on AI bio unless you're a biologist. But it's kind of like saying you can't work on video models unless you're a director or something. There are all these super domain-specific things, like understanding lighting in a video, but at the end of the day, these are just machine learning problems, and they're all solved the same way.
Okay. So then Chai-2—there's a jump in capability as well as an architectural change, right?
Yeah. What we've disclosed about Chai-2 is that it is an all-atom diffusion model. We're trying to predict atoms in 3D space still, but we're doing it in such a way that the model actually has the ability to design atoms, place them, and decide which atoms are actually there. One way to represent an amino acid, like a protein token, is by which atoms are present. So in the Chai-2 case, we were just predicting, all right, let the model pick what atoms it wants to keep, and then map that back to what amino acids there are.
What are you able to do with Chai-2 that you can't do with Chai-1? Is it just better, or are there new capabilities it brings?
It's design. Chai-1 lets you say, “Hey, I know the sequence of amino acids”—that text string—“and I know the structure.”
That you would get from the genome, exactly?
Chai-2 says, “Okay, I have a target structure that I want to design a binder to.” Chai-2 will then generate candidate molecules, candidate medicines, that bind to that target. This is a design model, or a design family of models. I think that's where you really cross the threshold of usefulness. Chai-1 is very useful because you can at least intuit and reason about the structure and see what you're looking at. But the ultimate goal here is to design medicines and new molecules, and I think Chai-2 really crossed the threshold of performance for doing that with antibodies a year ago.
One analogy here would be, back to the image domain, Chai-1 would be like, “There is a cat in this image.” Thanks, Chai-1. Chai-2 is, “I'll show you a background. Maybe I'll prompt you with some image information, like, hey, put a cat in a field,” and Chai-2 actually just gives you back an image of a cat in a field, and you're like, “That's a good-looking image,” or it's not. You might have some other model that ranks the image, but fundamentally, it's the generative problem.
13. Co-designing sequence and structure
Yeah. Taking that analogy a step further, it's maybe more like you show it a background, and then it generates that there is a cat. It generates an image of the cat at the same time, and it makes sense both that there is a cat in this field and that the cat works in the image. It's an interesting problem because you have to generate 2 things at the same time, both the sequence and the structure. I don't know if you can, but could you talk a bit about how that works? How do you do that? So you co-design the sequence—
—in a way that the structure also fits and makes sense. One way to think about it is the classic way of doing this. Let's talk about both in structure prediction: I know the sequence, and I can roughly figure out the 3D shape from that. Then there's the inverse-folding problem, which is: given a 3D shape, give me back a sequence that would fold into this. Now you need to do both things at the same time.
I think similar principles apply. You can have the model think a little bit about what the structure should look like, then have some other part of the model think about what sequence would support this. A nice thing with diffusion is that you can do this pretty slowly and iteratively. You can give the model a lot of time to think: if I change the structure like this, how should the sequence change? You can play this back and forth and eventually it converges on something that's self-consistent.
14. How do you know a design is any good?
It's almost like an EM algorithm. Yeah, exactly. So you have this model now, Chai-2, which is able to predict, or sample, a structure and a sequence that generates that structure. Just because you can generate a structure doesn't necessarily mean it's accurate enough to do something. Do you have other scaffolding on top of that? Are there additional problems? Are you one-shotting these things, or do you need to generate thousands of them and then have a ranking or scoring? Having a candidate is maybe not enough. What do you do once you sample a structure, or—
Traditionally, what's done—and when co-design in protein structure design started to become a thing, we were kind of at a loss for metrics—is ask: how do you know that your protein is legitimate after you design some sequence and structure? How do I know that this is legit? Even biologists are like, “I have no idea if this thing actually folds. Maybe some of it looks right.” Even our biologists are surprised by some of our designs that do end up working.
What was done at the time is that we came up with a bunch of metrics, and AlphaFold really enabled this. You'd take the sequence that you predicted and run that through a totally distinct structure-prediction method. This is completely independent of your model, and you say, if an independent model thinks that this sequence folds to a similar structure, then it has a higher likelihood of being correct than whatever the prior likelihood would be. You can take your sequence and measure how consistent this structure-prediction method is with the structure that you actually predicted for that sequence. You can now compare your design to an independent model's structure prediction, and that became a really good way of gaining conviction that your design model was correct. People kind of gamed these benchmarks for a while and kept pushing and pushing. It turns out it’s easy to get self-consistency, consistent design of structures, if all of your proteins look identical. There are a lot of problems that this creates, but then people started adding more and more on top of this.
It's totally out of domain now, right? By definition.
Yeah. As a human, you can look at this thing and be like, “I don't know. It checks out.” Even our biologists are surprised by some of our designs that do end up working.
That's an interesting point that I think some people have acknowledged in the community. How did you solve that? You can see that if you use your oracle and your sampler at the same time, you eventually converge. What do you do to stop that, or to convince yourselves that you're doing something valuable?
One of the nice things about structure-prediction methods is that usually you have some calibration in how confident the model is in its prediction. It turns out these models can give you a pretty well-calibrated confidence prediction. Rather than just saying, “This is what I think the structure looks like,” it'll say, “This is what I think the structure looks like, and here are the parts that I'm not really certain about.” You can aggregate this down to a single scalar.
15. Validation, cryo-EM, and the 0.33 Å result
Typically, people will look at not only how self-consistent they are, but also how much this independent folding model even likes the structure that it output. That was one way, early on, to gain confidence. Another thing that people often do is look at the diversity of their generations. Again, you could have a model that's perfectly consistent and gives you great confidence predictions back, but it might be the same structure every time, the same sequence every time. You also want to see how diverse the solutions are. How many of these new problems can I solve, in a sense?
If I had a whole lot of money to validate, how would you do that? Can I go and do cryo-EM or something like that and try to figure out the structure, to get some ground truth on that?
It's more that the feedback loop is really slow. You can validate a few structures like this, but it might take months, and it's just not a very scalable direction. I think that's a problem for the field as a whole. People are spending a lot of time—especially at Chai, I think—thinking about how to validate these problems at a bigger scale. How do we increase the throughput of our validation or decrease the cycle time? If you're waiting months to figure out, “Hey, was my model correct?” it's hard to iterate in a research environment that way.
The good news is that this is getting a lot better. There's a whole network now of wet labs that you can work with that will run these assays and experiments and tell you things like whether your protein binds to its target. Thankfully, we're not at years; we're down to weeks, which isn't as fast as LLM land, where you can scale up an eval, throw more compute at it, and get results back in hours, but it's fast enough that you can start to recursively self-improve.
And I think we also spend a lot of time figuring out what metrics we can compute in silico—on the computer—that are predictive, perhaps, of lab success. But, to your question about cryo-EM: yeah, you also have to measure the structure, and as you know, that's expensive because you have to freeze the protein and shoot these electron beams at it and see how they bounce off.
I remember there's this really funny anecdote. We'll see if I can share it. In the Chai-1 paper, we actually did that: we took some of the proteins that the model predicted and ran cryo-EM, and we got the results back and were like, “Wait, the results look wrong,” because we had overlaid the prediction over the electron density point cloud, and we didn't see any difference. The point being, we're getting to the point now where these structure-prediction models are within a few angstroms or less of the actual atomic positions.
And in this case, it was a 0.33-angstrom error, which is 1/3 the width of an atom.
And we were like, “This can't even be right. Clearly, they just sent us back the wrong design.”
They just sent us back our design.
Yeah, exactly. Yeah.
Did you check for data leakage?
Yeah, in this case, there were no known antibody binders. We actually chose these targets specifically to have no known antibody binder, so if we did get a hit, it was definitely the first antibody hit to this target.
16. Chai-2.5, Chai-3, and betting on the models
I think one of the things I didn't realize about biology was just how much of it is literally feeling around in the dark—and that's not even a metaphor. You literally can't see how these things look, right? So structure models are so huge, because now you can actually predict within an angstrom how these things look, and that enables you to do things like Chai-2 with the design models.
This, to me, is one of the cornerstone problems in AI for science, right? You don't know—
You fundamentally don't even know how to measure your problem in a lot of cases, so it's very difficult to validate.
Yeah. So you're getting these sub-angstrom predictions with Chai-2. Chai-3—why Chai-3? What's better, or what?
Yeah, I think with Chai-3—honestly, there was Chai-2, there was Chai-2.5, there was Chai-2.7, and eventually there was Chai-3, and each time we saw better and better performance. I think the main thing with Chai-3 is that we looked at Chai-2 and the targets it could solve. There was a lot of internal discussion after Chai-2: “Hey, we made successful binders for half of these 50 targets. What about the other 25? What can we do to make those better?”
Then we were split. We thought, “All right, should we study these targets that we missed and figure out exactly whether there are properties of these that we can look at, or should we just bet on the models? Will the models just get there if we put more time into them? Should we be a little less empirical in that sense and really bet on the models getting better?” We definitely took the latter approach: we bet on the models getting better, and we pushed as hard as we could on that front.
Scaling up the model, the data, whatever, to just build more accurate models.
Yeah.
Is it accuracy? Is that the main thing? Is it binding affinity?
So I think binding affinity is a big one. You can't just bind weakly. For this to be a useful tool, especially for our partners, we need to start producing molecules that are at or very close to therapeutic grade, which means they have to bind really tightly. They also have to be developable. They have to have all these nice therapeutic—
Properties. And developability—I think we talked about it; he mentioned Chai-2.5, right?—which we released a few months after Chai-2. There was a study we did on the developability of the molecule.
For the audience, obviously, the molecule has to stick well and stick tightly, but there are these other properties you care about. To use the non-biological terms: is it safe, is it stable, is it easy to manufacture, and does it self-aggregate? We've been pleasantly surprised at how much we've been able to climb and push the performance in those areas.
It seems like one of the reasons that you want to do antibodies is the developability.
Yeah, you get a lot for free there, right, with that antibody framework.
Yeah, it's interesting. To me, there are many structure-prediction models out there. I feel like these other ancillary factors are actually going to be the most impactful in the usefulness of a product, let's say.
Yeah. Right.
Yeah, absolutely. The nice thing about structure prediction is that there is a ground truth you can compare against. For design, you don't really have that. You're like, “Here's some new disease molecule. Give me a binder for that.” If you want to know if this thing really binds, you have to send it off to the lab and wait a while.
With structure prediction, you can be like, “All right, the model hasn't seen this sequence before. It's never seen anything close. Does it actually fold up into the correct shape?” We can just hold that out of the data set and check. I've always thought of structure prediction as this really nice speedrun benchmark to validate ideas on.
Right. Sorry, I didn't mean to say—I meant structural models in general. Yeah. Maybe we can talk a little bit more about getting into the product side of things. Thank you for coming. [laughter]
17. Building the product under pharma IP constraints
Actually, I really think this goes throughout, not only for structural models like this but also virtual cell and whatever. It's really all the other stuff around the drug-development process that is going to have the biggest impact. So can you talk a little bit about that?
Yeah, I think that's actually a good thing to talk about after Chai-2, because I think Chai-2 is where it started to get really fun from a product perspective. With Chai-2, we crossed the threshold of usefulness where, after we released that paper, a lot of pharmas and biotechs approached us and said, “Hey, this model might be able to do some stuff for us. Can we use it?” Then we thought, “Oh, man, we should build a product. We should build something to let you use that model.”
That's right around when I joined, and there was this mad buildout to both build the product—which we can talk about the shape of—and secure the compute so we could serve those models to our partners.
Another third piece there that was really interesting is security and IP. We want to be a very neutral platform that anyone can design medicines on, but, as you guys know, pharma is a notoriously IP-sensitive industry. When I joined, a lot of people told me this couldn't be done: they're not going to put their data in a platform and have all their new medicines generated out of it.
Having a bit of a background in security helped a bit. Actually, if you're really aggressive about how you segment data and set up single tenancy—where you're almost deploying a separate version or separate account in the product per customer—you can actually build a platform and ship it to them.
Through the summer of last year, we started doing that. We'd been working with—or talking to—Eli Lilly, and they were one of the first partners to really work with us closely on that. That kind of made the V1 of the design suite that you can use to engineer some of those molecules. Maybe it's worth talking a bit about that design suite.
I think we have these really powerful models now that can do all of these crazy things if you condition them in the right way. If you give them the right context about the structure you're going after or maybe the constraints around the model—like, “Hey, I want to design an antibody that hits this GPCR protein, but doesn't collide with the cell membrane and also targets the specific epitope on that as well”—
And we looked at it and thought, “I guess we could put a chatbot around it. That'd be really easy to talk to.” But you're trying to build something almost very visual, right? You can finally build something really visual with some of these structure-prediction models.
If you look at the Chai product, it looks a lot less like ChatGPT and a lot more like Autodesk, SolidWorks, or Figma, if you've used those things, where you can load up your molecule. There's this almost Photoshop-esque design suite. You have the equivalent of a paint tool to paint your epitope. There's an equivalent of a content-aware fill tool to get your binders generated from Chai.
18. Convincing scientists who hate AI tools
You, of course, have a lot of the scientific analysis and plotting and whatever to understand the results of the models. But we've just been surprised at how much complexity is actually in doing that, so that you don't shoot yourself in the foot when you're prompting these models to give you advice. So, are you sitting with people who are designing these antibodies and then complaining to you or whatever?
Yeah. How do you convince med chemists to use your tools? Med chemists hate AI tools—it's notorious. They're like, “I don't want to touch this thing,” or, “I don't understand it.” They will not touch things they do not understand.
Well, it helps a lot to have the models working really well, right? When we had the results of Chai 2 and Chai 2.5, I think that's enough of an activation energy where pharma companies and the scientists within these companies are like, “Oh, let's try it. Actually, Chai, can you guys just try running the model against a few of these targets and let's look at the results?” And then we do that, and the results are good, and they're like, “Okay, let me try to get on that product and let me try to use it.”
No, I think pharma is incredibly pragmatic, actually. I've been very impressed with everyone that we've worked with so far. They're very pragmatic about this, and they're willing to be proven wrong. I don't blame them for not trusting the models. I have used these models, and rightly so, I'm pretty skeptical when I see a new release. I always have been.
You really just need to show them the proof. They can give you a target that they're interested in, or maybe it's more something they've worked on in the past. They probably don't want to share IP right out of the gate, but they can be like, “Hey, I've had trouble with this particular target in the past. Let's see how you guys can do on this.” And then once you show them the proof, they're almost overwhelmingly willing to accept that.
I come from a cybersecurity background and have worked on security products before, and those were dark, dark years because you spend a lot of your time selling to people who are surprisingly not that technical. You think cybersecurity people are very technical; in many cases, they're not. It is this kind of uphill enterprise slog to a very unsophisticated customer.
I think I've just been pleasantly surprised by how much I enjoy working with our partners and our customers. These are scientists who have been spending 5, 10, 20 years of their lives working on one target, often in some cases, and they've studied everything about it. They're very sophisticated. They're very smart. Getting to collaborate with them is just a gold mine, and we learn a lot about how to make the product better.
19. The ten-year target
There's this anecdote. A few months ago, we were showing some of the results from a target through a pharma partnership, and one of the scientists in the room started tearing up and crying.
Oh, wow. [laughter]
“You really hit the nail on the head with that one.” And she was like—we were like, “What's wrong?” She's like, “No, I've just literally spent 10 years trying to get an initial binder to this thing, and you guys were able to help me do it.”
Oh, that's—
And that feels really special. To answer your question, there's, of course, the teams of scientists and computational biologists that we're working with within each of our partnerships. There's also the people we have within the building, right?
20. Battle-testing without your own pipeline
I think one of the things that I really appreciate about Chai is how cross-disciplinary it is. We have people who are maybe engineering experts and less bio-experts, like myself. We have great AI scientists, or ML scientists, but we also have a bunch of scientists that we work with and have joined Chai to help us test the limits of the models—see what Chai 2 is actually capable of, what targets it can do, what it can't do—and inform some of the research direction there.
I want to add to that. In the Chai 2 days, we started with a bunch of engineers and people who had AI-bio experience. We didn't have a hardcore lab scientist, and one of our first hires in that realm was Nathan Rollins. He started working in the Baker lab at 14, graduated from Harvard at 18, and got his PhD by 21 or something like this, in the Marks lab.
He was super skeptical about Chai at first. Then the results started to come in, and he was like, “Okay, this is kind of interesting. This could work.” And then once the Chai 2 results came back, he was like, “I need to bulletproof this. Nobody celebrate yet. All this—” [laughter]
I think it's been really nice to have that level of rigor: to have people who have spent the time in the lab, who have designed proteins themselves, who have literally, in the case of Andy, led several therapeutic programs and brought drugs to the clinic themselves. We have all these people internally at Chai using the product and really battle-testing that.
So if you don't have your own platforms—right? I mean, you don't have your own programs, right? You're a pure platform or partnership model, right?
Yeah.
How do you battle-test something if you basically don't have a use case where you have to continuously push it forward? Or if you are just pushing things forward, you end up with your own candidates if you're successful, and then what do you do about that?
We have benchmarks of our own internal cases, right? There's a set of targets that have known therapeutics against them. There's a set of targets that we pick to push ourselves, and we're constantly refining that set and adding to it.
21. Waterfall to loop
That's what the internal science team helps with: expanding that and almost running the experiments to try to get initial binders there. We don't care about developing those drugs; we just do that in service of validating and making our models better.
And then, of course, there's a loop with our partners, too.
Would you consider yourself hit discovery, or are you using some jargon like hit-to-lead or lead optimization? Where do you live in this? Hit discovery might be one part of it, which you can do, but the later parts are often much more bespoke and special. How do you balance that? It seems much more difficult to me to be general—to solve general lead optimization—than it does to solve discovery.
Ideally, we really want to be able to, rather than think of this as a bunch of stages, produce drug-like molecules straight out of the models. I think part of the reason why we think of it that way is because the initial molecules are usually not good enough to be drugs. We're kind of—
At the inflection point now. We're really seeing this internally at Chai, where the models are getting pretty close to producing molecules that could eventually be drugs, or are very close to being drugs.
We try not to make too much of a distinction between hit discovery, lead optimization, and all the different parts of this preclinical pipeline. Our north star is to just really produce drug-like molecules straight out of the models.
22. Levels of abstraction, and throwing the product away
Of course, this is going to be hard, and there are going to be tons of roadblocks. You need to be able to actually prompt the model to do this. You need the whole RL stack to learn different properties and things along those lines. But I think it's very achievable.
And I think, to add to that, this notion of target discovery, hit discovery, and optimization—where each of these has a gate and takes a few months to a few years—is this very waterfall model, where the cost of trying things and getting things early is very expensive.
But I think, to what Matt's saying, if you start to get into a regime where you can have models give you really promising candidates, you can start to make that look a lot more like a loop. It's akin to becoming more agile in software development.
Internally, we have two north stars. At first pass, they almost sound contradictory, but the north star in research is to start to de novo, one-shot, better and better medicinal candidates that are as close to being ready for the next phase as possible.
But within product, we do want to expand into whatever these iterative workflows look like, where maybe I get a binder, I get some results from the lab, and I'm using that to condition my next run of the model. I think they sound contradictory, but they're actually not, because I think what's going to happen is that the research is going to get better at identifying a de novo candidate for a specific class of drugs. Say, antagonists—blocking things—might be a little bit easier. Maybe we can get to a state where we can one-shot pretty good drugs there.
But now the next problem is agonists, right? How do you reliably one-shot a switch on a cell? Or bispecifics or ADCs? I think there are these levels of abstraction that we're going to have to climb with the product as the models get better.
One of the things I got very existential about a few months ago was that I thought, “Man, all this stuff we're building in the product to visualize molecules and do this—maybe I'm just going to have to throw it all away when Matt ships Chai-4.” But I think that's the reality of building products now. You're using them less as an end in and of themselves. Maybe you'd have built software that was supposed to last 20 years; now it's supposed to last maybe 1 year, but it is the bridge to deliver value and enable the research that gets you to the next thing.
I'd imagine we're probably going to rewrite our product at higher and higher levels of abstraction. Right now, we have something a little more akin to Cursor, where you're inspecting the molecule in the same way you're inspecting the code, because you really need to verify the bonds that are forming and the properties of the things that you're getting. But then you get to a point where that stuff is solved enough that the product is actually just helping you orchestrate these campaigns of hypotheses. Maybe you have 1 target and you're orchestrating a bunch of different epitope choices or whatever against that. Then maybe you're going up 1 level of abstraction, where you're now doing a whole campaign against all of the targets within a pathway.
23. Epitope prediction: the harder problem
I think what's really exciting about that is that if you have really good primitives for structure prediction, binding, and design, and you can compose them, then you can start to grow into the outer loop of science. Maybe the thing runs itself, and you start to get to some really, really cool drugs at the end of it. I actually want to push on what you just said about epitope prediction, because I think a lot of people in the field would argue that this might be the much harder problem than finding antibodies and binders. Where do you think the state of the art is in general, and also with regard to Chai, in terms of epitope prediction? Is this a problem that has a reasonably solvable time horizon? Also, can you define epitope prediction?
I'll think of this at a few different levels. The most basic level is: I have some disease that I want to target, and what proteins are actually responsible there? Actually figuring out biologically what's going on—what should I be targeting in the first place with the drug? I guess once you figure that out, it's kind of a structural biology problem at that point. You're like, “All right, this set of proteins is responsible, and what's going on there?”
Well, this is interacting with some other protein that it shouldn't be interacting with. Conventionally, you'd just want to block that interaction or something with an antibody. But that's where these proteins interact, and the type of interactions that you want to disrupt—that's typically the epitope. It's the actual site on the protein that you want to block. This is a ridiculously hard problem. I'm with you on this. This is the harder problem—the amount of context and global understanding you need in order to figure out what's interacting and how.
But maybe let's take a few specific cases. What about when SARS-CoV-3 comes around, or the new flu or whatever? What would you do there? Is that something you think you could reasonably tackle?
In that case, yeah, you could just run a structure-prediction model and see where the model thinks this thing will bind. If it's highly confident in that, you might say, “Okay, here's the site that we want to block.” I think in general it's still very hard. Even structure prediction is getting really good, and a lot of people think AlphaFold 2 solves structure prediction. Not really. AlphaFold 2 got, I think, 11%; the multimer version of this got, like, 11% of antibody–antigen prediction cases correct. That means 90% of the time it's wrong.
Yeah, but AlphaFold 2 solved a certain class of monomeric proteins with MSAs.
Right. And just to clarify—I had to understand this myself, so maybe I can help listeners who aren't familiar. An antibody—the whole point of an antibody—is that it can identify new things that the body hasn't encountered before. The design of antibodies, as opposed to other types of proteins, is such that the system is designed so that you can quickly recombine different components in order to match proteins from unknown pathogens, more or less. This is why it's not conserved in evolution the way that other proteins are.
24. The economics question
Yeah. So, back to the epitope prediction problem: I think it's still hard. There are a lot of cases that are maybe tractable, but in general, if you want to discover this for a new target, it's still a really difficult problem. Maybe Virtual Cell would be the closest thing to the state of the art there, but that's still a ways out.
I wanted to dig in a little bit on the product, because there's something I don't understand about the economics of basically all the structural stuff that's happening right now. Obviously, a lot of people think it's very valuable, so I'm not grokking something. When you look at the cost of developing an antibody, it maybe is a couple million dollars, right? You identify a target somehow and then say, “Okay, I need an antibody to match this.” Then you optimize it in various ways and maybe try it in animals—with antibodies, you typically get to animals faster.
If you look at how much it costs to bring a drug all the way to market, if you're prescient and pick the right target and the right technology, it might be half a billion dollars. Typically, that $2.6 billion number is advertised over all the failures as well. If you look at just the cost of that 1 success, depending on the disease, maybe it's less, but half a billion might be a good median number or something. You're saving a couple million dollars in a half-billion-dollar campaign. Why is this so attractive?
I would maybe challenge the premise in a few ways. Sure, if you're trying to get an antibody for a very simple target, maybe. But what we've been most excited by is our partners using antibodies in more sophisticated ways. For example, in Chai-2, we showed GPCR agonist activity, where you can really hit the switch on a cell—a doorbell protein, so to speak—in a very precise way.
That's very, very, very hard to do with antibodies if you can't be that precise, right?
25. Modalities you can't get from immunization
So you're unlocking a new capability. I would think about it as less like, “I'm taking the existing drugs I can make and making them faster.” There is some of that too, right? But it's more like, “No, there are better targets to go after that are maybe more precise and more effective.”
There are also drug modalities that you just can't discover with immunization. You're not going to design your crazy multispecific, four-headed, super-intense formats. These are things where you kind of have to design them from first principles. Even with bispecifics in particular, both arms now need to bind different targets, and you have this multiplicative effect on your binding rate. If you have a 1-in-a-billion chance of finding a binder in arm 1 and a 1-in-a-billion chance in arm 2, you're not—this just isn't going to work with the traditional approach.
Exactly. I think the other thing I'd think about is that you're not just helping your partner with maybe 1 drug. There might be a portfolio of targets that they're going after, or a portfolio of drugs that they're trying to make. The nice thing about the platform approach, rather than developing individual drugs, is that we can scale with them as they pursue more targets, in addition to more ambitious targets.
Right. So it lets you concentrate your learning in a subdomain, and everybody benefits from that.
Okay. So what are some of these capabilities? You mentioned a few. Are there more that are really interesting that you guys are chasing?
Yes. We talked about cross-reactivity and selectivity, and some of these really interesting additional modalities with bispecifics. There's a set of things that our partners have been asking us for that we've been working on, but I can't get too into them because that starts to reveal some of the targets that they're going after.
But I think the point being, once you get precise, you can start to do some really, really cool drugs.
It's a new technology, right? Technology in pharma means, how do you deliver your therapeutic? CAR-T is a technology, right? And so this is maybe a new technology in the sense that you can have these highly, highly engineered therapeutics.
26. What's really blocking science-to-engineering
Right, and that comes from the mission of the company, which is to really turn drug discovery from a scientific experiment into an engineering discipline, right? How do you get to the precision-engineering phase for biology, where you can start with, almost declaratively, defining the thing you're trying to get and have the model fill in the gaps and get you that?
So what is the biggest blocker to going from science to engineering?
Oh, man, there are so many things. That's the thing about Chai.
Yeah. [Laughter] Yeah.
I don't even want to talk about the amount of headaches.
Too late. You already know.
Okay, just when you're actually parsing—first of all, file formats for biologists. They just don't care. There are standardized file formats. Are they the best? I don't really know. But there's also a lot of information that you want to pack in: I have this structure; here are the people who solved it; this is the method I used to solve it. There's a lot of stuff going on.
Depending on the method that you use to actually figure out what this 3D structure is, you might have multiple copies of that structure. Part of it might not have really been resolved, or you're thinking, “It could be here; it could be there. I'm just going to give you both options.” So the actual parsing problem on the engineering side of working with this type of data is really difficult.
This seems like something that LLMs can excel at, though.
They don't know all the edge cases often, right?
This is more back to just a simplicity approach. LLMs are very good; I will absolutely give you that. Then you're thinking: Should this function have 20 special cases, or should we be really principled in how we approach this? How opinionated should we be in how we do this?
We want a strategy that's easy enough for humans to understand. When we're reading through the codebase, we really need to know what's going on here and what the potential problems are. Sometimes that just comes down to looking at examples.
Once you've figured out all the infrastructure work and how you get data into the models, there's then scaling the model, and then scaling the infrastructure around the model to train bigger and bigger versions of this. That's a lot of work that Neil and the product team actually do.
27. Buying compute is a terrible job
Yeah, I mean, that would have been my answer: the infrastructure part. I mean, not to beat a dead horse, but compute—getting the compute and using it in the right way—is such a challenge, especially for startups.
This has been such a—yeah. Anthropic is single-handedly holding back science.
Totally. One of the things that I help a lot with at Chai is buying compute for the company.
Worst job, man. I would not recommend it. It is very stressful.
Back to you—the hardware job.
Yeah, yeah. I know. Exactly. In the wrong way. But in September of last year, we started to really notice that things were getting tight, right? We were doing a lot of our inference on spot and on-demand markets, and we'd have these days where you would get these capacity crunches. We thought, “Okay, we should probably start to get ahead of buying some compute for ourselves.”
I think everyone probably says this, but, man, it was hard. I didn't realize how much of a power law this is, where there are, say, 10,000 B300 units shipping everywhere, and the hyperscalers and the biggest AI labs are buying 95%+ of it. Then you have the startups fighting over the scraps.
The other thing that's really interesting, especially if you look at these later compute versions, the Vera Rubin systems or the B300s, is that a lot of this stuff has been built very LLM-forward. You have systems with huge KV caches where you have 72 GPUs that are all required to talk to each other. Obviously, some performance gains there help us, but it's interesting just how much the compute market has gotten LLM-pilled.
I think there's probably a whole set of compute-stack and inference optimizations and things that need to be made for this class of models. I think this class of models is going to be just as big and just as impactful as LLMs, but the compute market doesn't realize that yet, both in the capacity sense and in the software-stack sense. We actually spend a lot of our time, even just doing basic optimizations of compute, to get it to work better for the types of models that we have.
28. Triangle attention and why GPUs hate it
Yeah, I know that some structure models are more recursive than LLMs, for example, and that changes maybe the compute-to-memory ratio that you need and things like that. What are some of the cool or interesting optimizations that you've done there, depending on the type of model?
So we can go back to a Chai-1-type model. In that case, we're following the AlphaFold 2/3 architecture, and rather than doing attention over a normal sequence representation, you're loosely doing attention over this pair representation. You can think of this as a sequence of length L² rather than, typically, length L. If you're doing attention over that, the way that you actually batch this up ends up being L³. Now you're in a pretty heavy compute regime.
The amount of FLOPs that you're putting into every token is pretty high. The memory-bandwidth overhead of just transferring that from SRAM to whatever is a real bottleneck in these architectures. Even something as simple as layer normalization can take a long time; that can be a significant amount of the compute that you're using.
29. Durable execution and Temporal
On our side, we've spent a lot of time optimizing it and taking engineering very seriously so that these operations are at least better. We're always looking at how new chips perform compared to the older versions. Sometimes that's even different for training versus inference, and of course Neil knows this really well.
Well, there's what you're doing on the individual GPU, and then there's how you orchestrate fleets of GPUs, right? You basically shard your computation. When you're designing a molecule on Chai, it's not necessarily 1 call; it's a lot of GPUs being thrown at the problem across a lot of compute. Actually, I would say that one of the hardest things to get right in software engineering is durable execution. Are you all familiar with that term? Can I go on a little?
Ultimately, if you're computing a lot of data—model calls across a very wide set of infrastructure—you always run into problems where some part of the infrastructure is flaky. Maybe the bucket you're grabbing your data from goes down, your database has a blip because there are too many transactions against it, or your GPU errors out.
I've been at companies before where you spend so much of your time just dealing with this. You're basically putting all these queues and retries in place, duct-taping things together, and it becomes this mess where what used to be a pretty simple computation that's just distributed turns into spending 95%+ of your time on all of this queuing and retry stuff.
We're huge fans of a company called Temporal. There's this idea: if you're trying to get a really long-running job to run, at the end of the day, what do you need? You need a queue, your flaky thing pulling off the queue, retry logic to put things back on the queue if they fail, and a whole orchestration system to tie all the queues together and monitor them.
What's really cool about Temporal is that it's a tech company that's invented a framework for doing this. One technical decision we made early on that was very helpful was to run as much stuff as we can on Temporal. That includes calls out to the database from the app, so the database transaction goes through without failing; side effects can sit on Temporal and get retried smartly without us having to write our own queue logic; model calls; and the orchestration of really long data pipelines.
Point being, one of those primitives is that you need to get durable execution right so that you're not stuck in retry hell. A really deep engineering thing that you wouldn't realize unless, like me and Jack, you've been burned by this many, many times before.
30. Complexity vs. the bitter lesson
I think we're at this state now, right, where we've raised another $40 million. I have to go buy another compute cluster. We're going to have really, really, really large runs, inference, and training sets. Getting those foundations right is what's actually going to let us do more ambitious things. To kind of answer your question, I actually think that's a lot of the bottleneck to making biology more like engineering: just having the right engineering primitives.
I have an analogous tangent on the model side. Actually, one of the things that's nice about those problems is that they're super visible. At least, this crashed, this failed; we just see that the loss curve didn't go down, or we see weird gradient behavior, or whatever. I think a lot of these same principles—engineering first—also apply on the research team.
One thing that I like to say is that complexity and being Bitter Lesson-pilled are fundamentally at odds. For example, I think AlphaFold 3—I might get this number wrong—had 23 submodules, and at that point, that's a really difficult system to optimize and study. You're like, "All right, what happens if I tweak this thing in submodule 30 or 21? What happens to the whole system?"
You can always think, "Hey, we can make this better by adding submodule 24," but should you? Or should you think about just removing things and lowering that complexity down? I think that's a pretty fundamental thing at Chai: the engineering culture and being very simplicity-biased.
31. Inductive bias and data efficiency
Have you all seen the picture of the SpaceX engines? Raptor 1 has a bunch of pipes, and Raptor 2… We have a picture of that on our office wall because it's just true, right? How do you delete, delete, delete more things?
Yeah.
But the only way you can accomplish that is—I mean, the reason AlphaFold 2 and AlphaFold 3 worked is that they were small models, relatively speaking. They were very compute-intensive, but they were very data-efficient.
Yes.
And there was inductive bias after inductive bias—
It was brought in by human intuition and probably hard-fought experience. Mm-hm.
They were incredibly efficient. If you try to knock down those things, they're not like a house of cards. Everything is an incremental improvement on top of it. In order to get beyond that, it seems to me like you really just need new sources of data. You need to at least treat data fundamentally differently, in a way that is much more efficient.
I'm actually kind of surprised to hear that you have scale to that degree, because I suggest that you're doing something very different from the way the community is thinking about it. I don't know if you can comment about that, but—
We're pretty first-principles people. The whole research team at Chai—except for me and Kevin, really—we're the only people with a bio background. Even still, we're pretty far removed.
We try to look at every problem as a core ML problem. We try to think of what's the analog in other spaces. Even for image models, CNNs were built to process images, so images should be looked at in patches. That was the nice inductive bias there.
Then people were like, "Well, you can just tokenize this thing, throw it into a transformer, and it's going to work," and it did end up working, even on a relatively small data set. But for proteins in particular, it is really hard. There's not as much structural data. There's a ton of sequence data, and that's one of the unlocks for ESM working. You can get that to just run on a transformer. If you try to do the same thing with experimental structure data, good luck. You need—
I mean, there was that Apple paper where they distilled on ESMFold, which was actually really cool: you could distill on a very large data set and get good signal, but it didn't generalize at all because it wasn't reasoning. It was really pattern matching.
One of the things these triangle layers you were talking about, for example, is that they do have a very nice inductive bias. Maybe it's not the triangle inequality, like the paper originally proposed, but it's a clean inductive bias, and it unambiguously is one of the things that made it work. It just comes at a huge cost.
Yeah.
Yeah. No, I think that's definitely true. These layers are pretty costly, and that kind of limits what you can do with the architectures. They're not only costly in terms of compute; they're just not efficient on modern GPUs either. You have small hidden dimensions and large sequence dimensions—it's exactly the opposite of what GPUs are designed to process.
One takeaway from triangle layers is that you're just trading off parameters for compute in that sense. That's one mental model for thinking about this. I might want to throw more compute at the problem and trade that off for parameters, because I won't be able to hold as many parameters. I can't literally store these large pair representations and still do normal attention.
I think there are fundamental things you can abstract from ideas like AlphaFold, but you can tweak these and start building off of them in your own way. It sounds like you have quite a bit of fundamental research going into this direction. For, I guess, an audience looking for a nerd snipe in ML engineering for new problems, it's a very different research direction than a lot of the community is going in.
Yeah. Yeah. I think what we built at Chai is very unique in a lot of ways, but also very tied to what core ML is good at, kind of what I was saying before. We try to map every problem into a core ML problem. We think, how would you approach this if it were an LLM or something like that?
At the end of the day, we really, really value simplicity. We really encourage people who don't have a bio background not to be scared of this stuff.
I think that extends into the product too, where there's a balance to be had here, right, between how general you make the product. Do you build a cross-reactivity workflow, a selectivity workflow, and a bispecifics workflow? Or do you say, "No, let's make the model general enough to say, 'I'm going to condition on arbitrarily binding or avoiding something,'" and then you just have a very general screen in your CAD suite where you can say, "Hey, I just want to avoid or bind to these parts of these different structures"?
I think, kind of like the ML team, I don't have a formal bio background. Most of the product and platform team doesn't have a formal background either. There is some amount of regretting my words that I'm going to have, right, because I'm sure there are a million nuances, and I don't want to come off as too brash or naive there.
But I think it's sometimes helpful not to be burdened by all of the nuance and nonsense, and you can get to be maximally general because that's kind of what we're seeing in the research. The models are very general; that lets the product be very general.
I'm thinking back to my CS theory days. My first adviser was like—we were working on some problem, and we needed a polynomial-time algorithm for something. He would always tell me, "Never underestimate the power of polynomial time. This is basically like you're allowed to choose whatever exponent you want."
My first paper was an n^20-time algorithm for this problem, and I was like, "Andy, I did exactly what you said."
He's like, "Wait a minute, I didn't mean it like that."
32. Will protein design get commoditized?
Yeah, but I think you can really help yourself—you can free yourself a lot when you're like, "All right, I can kind of do whatever I want and then simplify it later." I think that's really a pretty fundamental way of thinking about things that we leverage a lot at Chai.
The space of binders in protein design, and binders in general, is actually a fairly crowded space. I'm curious about your general outlook of the field and the industry. I mean, I can go back to some anecdote. I was at maybe NeurIPS 3 or 4 years ago, right? The one right after RFdiffusion came out.
I was talking to someone in the Baker lab, and they were like, "Man, I just one-shotted." I don't think they even used "one-shot"—one-shot wasn't even a term back then—but they were like, "I just got picomolar binders out of RFdiffusion and just threw in the cryo." Great, right? It didn't seem like that just solved the problem. It's not like, "Oh man, now every—"
Yeah. But there are lots of people who I think have seen that you can actually do protein design, at least in some categories, quite well.
Is it mini proteins or mini binders? Ironically, nanobinders are actually smaller than mini proteins, or maybe a little bit harder. Antibodies are typically considered even harder. But is this something that can and will be commoditized, at least in some part? How do you compete? Where does the field go from here?
I think the answer is that it's kind of all of the above. There probably will be some commodity layer for certain types of modalities or drugs, right? At the same time, we're going to be able to do more and more ambitious drugs, and it's just like what's happened in LLM land. You have your open-source models that are maybe general and helpful for some things, but people are still buying frontier models, right?
Actually, if you look at the amount of value captured, it's the closed-source frontier models. The whole pie is growing, but it's growing so fast that even as the share of open-source models expands, the frontier models are still able to capture the majority of the value.
Right. And what are the reasons for that?
One, if you have more intelligence, you're going to go after harder tasks, right? If we have more intelligent biomodels, we're going to go after more crazy bio tasks. But also, a lot of the reason I don't use the open-source model is because I don't get Claude Code, right? I don't get Claude. I think there's a product layer to be built that's just as important as the model layer.
We learn a lot from our partners and the people in the building as well. What are the really tough things that they get stuck on using the models? Some of them are the dumbest things, right? I want to be able to better visualize this piece and focus on that. Some of them are actually very sophisticated things that we then have to build a pretty vertical product for.
Look, maybe in the fullness of time, AGI one-shots everything and it doesn't matter, but I think there's quite a bit of time until we get there. The product makes a huge difference for that. That'd be my answer. You probably have a more model-forward answer.
No, I think biology is slow, which is one nice thing, and there's not that much labeled data. You could take all the publicly available sequence information out there, and that might give you a good base model, but you still need some measurements on that data. That's still pretty time-consuming, and then you need to iterate on it.
I think there are even data blockers to unlocking this. If we really want to do zero-shot design and start generating candidate molecules that are almost ready to go into the clinic, I think there's more to that than just AGI. AGI might not solve that right away. There are definitely some technical blockers there.
33. No pipeline, no data moat?
But even in the space of specialist companies, there are probably 10 or 15 protein design startups. The 2 things that it sounds like Chai has gone in on are, 1, an all-in-one product, and 2, you're not trying to do your own platform if you don't have your own data moat. Is that going to help you win out in the end, or is that going to be a blocker? I'm just curious.
Yeah, that's a great question. Chai definitely has no plans to start a pipeline. We take the partnership model pretty seriously. From a personal standpoint, I love the incentive alignment: We make the models better, the partners succeed more, and that iterates on itself.
I think that's a pretty unique part of Chai. One, we're able to partner with a lot of people. Two, we get feedback on the product, so we know that it's very real. This is in the hands of legitimate big pharma companies, and they're actually running campaigns on this stuff.
We really have to be model-forward and model-focused. We need to keep delivering value, and that puts a lot of pressure on the research and product teams. First of all, the product team has to serve these needs, while the research team is always shooting for better and better versions.
The way I think about this is, if you're a bit more model-forward kind of thinker or company, then there comes a certain point where there's a lot to do on both the model and data side, but I don't think either is exhausted. It would be stupid to say we don't need any more data, but it would also be stupid to say the models are stuck and we can only use data to solve these problems. There's tons of room to grow on both sides. We're taking both very seriously.
I would also push back on the no-data-moat premise. That'd be kind of like saying, “Hey, all the enterprises that work with Anthropic, you're not letting Anthropic train on your data, so they can't build models that are good at enterprise workflows,” right?
We are investing in this. There are ways to turn compute into data and get more, and we're doing those. But also, what is the kind of data that you're trying to get? What's cool about working so closely and supporting so many of these partners is that we get to really learn what would be helpful in research.
34. Per-partner fine-tuning
Rather than doing research in a vacuum based on what would hypothetically be cool, we're able to do informed research based on what our partners have been very organically asking us for help with.
I see. I assume that you aren't allowed to train general models based on your partners' data. Do you train specialized models? Is there an AstraZeneca model and a Pfizer model?
Yeah. In a lot of these deals, we're working with them to train or fine-tune a version of our model for them, and I think there's probably so much more we can do there over time.
My brother started a company called Applied Compute. Great company. They're kind of doing this thing for LLM design and helping enterprises really understand the value of their language data and do that for specialized tasks. I think there's a whole world where we could potentially do that for biological data.
What is the value there? What is the lift that you get from using their data? Is it just that it's more data, or is it more that it's specialized to a problem?
They have a lot of scientific data that they're getting from experiments that can maybe help our models do better in particular classes of candidates or targets that they care about.
35. Are all AI companies consulting companies?
Yeah. Even something as simple as they might just have some preferred way of doing things that's not native to the Chai model, and they can ask the product team, in a sense, “Hey, our designs have property X. Can you make sure that they have those?” Even things as simple as that actually have a pretty big impact for them.
This goes along with a pet hypothesis I have that all AI companies, and especially bio and scientific ones, are actually consulting companies. Pharma is particularly the case because you're developing a new drug. It's almost by definition new, so the existing stuff has to be customized in many cases, unless you're doing something that's just a reiteration of old stuff. A lot of the big pharma companies are pushing the boundaries of science.
Certainly, we aim to make the models very general, and we aim to make the product very general and powerful. But there is integration work with every customer. To answer your question, you do get some defensibility just by doing that.
36. The highest-value token in any domain
What's nice about building trusted relationships with these partners is that, hopefully, if we execute really well over the first year, they'll continue working with Chai to do more ambitious and more drugs beyond that.
I mean, it's going to be hard to switch, right?
I hope so. Yeah.
Just getting the security.
Yeah, yeah. Maybe one other interesting point is that, if you think of this on a per-token basis, I don't know if there's another domain where the downstream value of a token is as valuable as it is for pharma. Think about the actual drugs that come out: These can be multibillion-dollar assets. In the case of GLP-1s, I think the 2 GLP-1 drugs combined are maybe a trillion-dollar asset.
37. Pharma as VC: Genentech and Eroom's law
Yeah. Until, I think, 3 months ago, GLP-1s' total revenue was more than all of the AI labs put together.
I don't think people realize that. I didn't realize that. It's crazy, right?
But yet the market cap is way lower.
It’s crazy how relatively speaking the market is.
And I didn’t realize how much of a VC business pharma is. They’re, in some sense, taking really ambitious bets. If you study the history of Silicon Valley, obviously people think of Silicon Valley as software, but in the 1980s, one of the biggest venture outcomes—and one of the first ones—was Genentech. It’s such a VC model: you get this string of tokens that can then give you so much value downstream.
Just a general shout-out to Outpost’s blog series about finance and funding. Really fantastic.
Before that, I knew a lot of those points, but I didn’t realize just how deep that rabbit hole went.
Yeah. Maybe the single biggest problem in biopharma is actually just the funding model.
Have you heard of Eroom’s law?
Yeah. Oh, yeah.
Yeah, yeah.
Moore’s law backwards.
Yeah, Moore’s law backwards. In compute, it kind of scales, so you have this nice exponential, log-linear scaling of compute. You have almost the exact opposite in pharma: the cost of actually making a drug is increasing exponentially. The amount of money put into each drug is growing at an exponential rate, which is pretty interesting to see.
Which guarantees that, at some point, the marginal return on new drug development will be negative.
Exactly. So unless someone—maybe Chai—figures out how to fix this, I think we might be on the verge of flipping some of these—
Bending the S-curve.
Yeah. Just to double down on the point, pharma and VC fundamentally are both optimizing a portfolio.
38. Everyone is a capital allocator
Yeah. Thinking of pharma as sophisticated capital allocators, where they have this portfolio of targets and they’re allocating between them, was a big reframe for me. I think we’ll just see more of that in the future. Hopefully, pharma can take riskier bets and pursue really, really cool drug targets.
That analogy—the VC-type investor model—is actually how we think a lot about research at Chai as well. Our research team is relatively small, definitely compared to a lot of the Isomorphic Labs and DeepMind. Our research team is around 10 people, so we’re a relatively small team, but we think of it almost like an investing job, where you’re investing ideas into compute. In the same sense, you’re really just capital allocators in that respect.
I actually think—maybe this is too cute—but I would make the broader point that we kind of think of everyone at Chai as a bit of a capital allocator. One of the things that surprises people is that we’re pretty small. We’re only 30 people, and that’s because everyone we hire onto the research or engineering teams, especially now that they’re, in some ways, very empowered with AI, is mostly allocating their attention to the right ideas and allocating their compute.
This is, I think, a characteristic to some extent of machine-learning and AI projects, and also of science. If you’re building an API for some B2B SaaS company that’s not building foundation models, whatever your limit is, it’s mostly people. The resource you’re allocating is almost entirely people, whereas if you’re building hardware, AI models, or something scientific, your constraint is those resources. The bottleneck is the lab, the compute, or other things. You have to really be in the mentality of, “I have these limited allocations. I have some shots on goal. How do I allocate those shots?”
Well, I would say yes and no. Going back to the example of building an API for a B2B company, that API has incremental costs. You have to support it. It adds complexity to the product. It’s another thing you have to take to market and sell. Maybe you should actually be allocating that to a different bet—a different thing on your product roadmap that you should be prioritizing instead of the other thing.
In a world where building things just gets really cheap and increasingly free, the scarce thing is your attention—both what you can put into it to keep your product simple and grokkable, and what your customer can put into it to really understand how to use it. I see it less as a binary thing and more as all of us, as engineers, becoming a little bit more like allocators of attention.
Yeah, which is what executives are. We’re all just becoming—
39. One bottleneck removed by fiat
Everyone—well, I mean, I was listening to a podcast with Satya Nadella. He says Microsoft wants to make everyone a manager of infinite minds. If you really take that to its extreme, everyone’s going to be an executive. I certainly feel like an executive, and I talk to Claude every day.
A little suite of interns who are all going out and eagerly solving problems you may or may not have actually wanted, but they’re solving the problems.
So, we have 2 typical questions that we ask. We’ve already kind of asked one, but I’m going to ask it again, maybe more directly: If you could remove a bottleneck from your problem space by fiat, what would that be?
That’s an interesting question. One thing that would be really nice—I’m always in research land; it’s very hard to turn off—is probably just the validation loop of protein design in general. Just being able to say instantly, “Hey, this thing works; this thing doesn’t.” There’s still a bit of walking around in the dark that you’re doing. At Chai, we’ve taken this very seriously, but it’s probably along the lines of just validating hypotheses and knowing for certain that things work.
Yeah, that’s an unsolved problem for sure.
Unsolved problem.
Yeah, and it would be hugely valuable.
Hugely valuable. I’m going to take a much more abstract answer to that, which is actually talent scarcity. There are a lot of smart people going and working on LLMs. There are a lot of people working and becoming software engineers for SaaS, but not that many smart people go and work on bio.
I didn’t work on bio in high school because I thought, “Oh, I could pick up my computer and program apps,” but if I wanted to work on bio, I had to study, get good grades in school, and maybe get a PhD or whatever. Maybe that’s one reason for it. I think another reason is that a lot of this stuff is really obscure. We threw around a lot of big words during this podcast, and you can’t really visualize the things.
It’s one of the things we care a lot about at Chai: how do we make the whole thing feel visual on our website and in the product? Part of the reason we’re here is that more people should realize you don’t need a super, super, super-specialist bio background to contribute to this computationally.
I think a lot about talent flows and where talent goes in the economy. In the 1990s, everyone was flowing to finance, and since the 2000s, people have been flowing to tech. Big tech ate up a lot of the talent until a few years ago, and now maybe LLMs and the big AI labs are eating up a lot of the good talent. At the meta level, how do you allocate talent better? Selfishly, I want more talent going into bio. We probably want more talent going into manufacturing and physical-world things, and these other problems that the US has. If I had a megaphone to talk to everyone, that’s what I would try to do.
40. Takeaways
Okay, so that leads to the second question. Maybe the answer is the same, but what is the 1 takeaway that you would like people to have from the episode?
Yeah, I think biology has been this somewhat obscure-feeling field where you’re stumbling around in the dark. You don’t know what you’re looking at. You’re dealing with non-determinism in your experiments, and you’re having to do a very long and iterative trial-and-error loop across a very long amount of time.
At some point, you cross the threshold of what you can do computationally—when you can get folding models down to being within an angstrom, and when you can get design models to give you hit rates north of 50%. Now you can put them in a 96-well plate and actually have 48 interesting binders. You start to get to the point where you can declaratively precision-engineer what you want, rather than betting on nature or trial and error to get you there.
I think we had the same thing happen in software, where you can write code and deterministically get an outcome, or in electrical engineering, where instead of your schematic being drawn out, you can put it in Cadence Design Systems and get it made in silicon, or CAD for mechanical engineering, where you can precision-engineer your part and get it printed or manufactured. The same thing is happening in bio, and it's happening very quickly.
Yeah. And that really opens the door for a lot of really interesting people for whom it maybe wasn't as scrutable or accessible before, right? Software engineers like myself, researchers like Matt. Obviously, we're still going to want the specialists, but the generalists can often really accelerate the precision engineering happening in the domain.
Yeah, I think for me, the biggest takeaway is that the field is actually working. Not only does it have commercial traction, but the research is actually showing signs of life. It's not even just showing signs of life—the signs of life have been shown. We're actually in a place where the models work. They're delivering value, and there are still tons of really interesting research problems to solve.
So I think there's a lot more low-hanging fruit in this field than there would be in other fields. And I think the amount of impact that you can have, especially as a researcher, is just unmatched in this field. For us, we're all very mission-driven. But even if you're not, there are a lot of fun puzzles to solve.
There's this kind of 3D geometry angle. If you like diffusion models, there's a million problems to solve in that regard. We have this LLM-like trunk in Chai-1. There's just so much of core machine learning that's touched by these problems. Although we've made a ton of progress, there's still a lot to be done. I think it's just one of the most interesting fields to be working in, while also having some of the largest impact on humanity.
Thank you so much for making the long journey. It's been a great—
22-minute walk.
Yeah. And we look forward to tracking Chai's progress. Awesome. Thank you guys. Thank you very much.