🔬 可验证奖励驱动的 RL,但验证器是一间实验室——Lila Sciences
Lila 的核心押注是:受控实验可以成为 AI 的下一座互联网级训练语料库,而自然本身提供可验证的奖励。 互联网曾是“我们开采的化石燃料”,科学强化学习则让模型提出实验、观察现实,并生成更好的训练数据。真正可能构成公司护城河的是由此形成的飞轮,而不是某一种药物或材料。
这套操作系统优化的是信息增益和迭代速度,而非最大化机器人吞吐量。 仪器组成一张图,由类似 PCI 总线的传输层连接;每个动作都是一次 API 调用,执行者可能是“机器人手臂”,也可能是“人手臂”。Lila 自称“token 生成最大主义者和灵活性最大主义者”:下一次实验必须让模型学到有价值的东西,而不只是增加一个低信息量样本。
早期结果显示模型能力确实得到提升,但“愚蠢”和“新颖”之间的界线仍被刻意保持开放。 在表达和基因编辑任务中,Lila 称模型零样本表现“达到约 80%”,而人类为 0%;其提出的非铂族金属电催化剂从无聊到看似愚蠢,最终成为表现最好的候选物。这种上行空间伴随着严格复测、环境遥测、工具限制,以及承认有信息量的假阳性也可能浪费时间。
最清晰的商业验证是:一个体内 CAR-T 项目由2至3人用6个月完成。 对比项目大约耗时6年、此前投入1亿美元;Lila 称其“monster” UTR 的表现约为 Moderna 和 Pfizer 参考值的10倍,在非人灵长类动物的 B 细胞清除和持久性上也优于 Capstan 数据。Lila 不会负责临床试验,而是把这类验证点转化为收取费用并分享上行收益的“零 FTE 初创公司”合作模式。
跨科学领域泛化是其经济学 thesis,而不是品牌包装。 Lila 已生成10万亿个由模型产出、经实验验证的推理 token,覆盖生命科学、化学和材料科学,并称其通用模型往往胜过领域专用替代方案。在小分子药物发现中学到的化学能力,已经迁移到金属有机框架等应用。Rafa 认为语言不必成为每一种科学模态的必要表征,而 Andy 强调结合工具使用的 token 化推理。
物理扩展目标是一座经济学上类似云基础设施的全自动实验室。 当前系统包括定制驱动程序、被刻意放弃保修的设备,甚至还有运行 Windows 95 的视觉语言模型;终点是 24/7 运行、高密度垂直堆叠、自主运输,以及最大化单位体积产生的 token 数量。一座10万平方英尺的马萨诸塞州设施只是中间步骤,目标实验室“应该让人感觉像数据中心”。
最大风险出现在发现之后,以及仿真、硬件、监管和经济学的交界处。 一位嘉宾提到,只有约5–8%的临床项目能从 IND 推进到获批;材料还需要放大和验证,材料仿真数据往往无法预测现实,而 Andy 称强化学习工作负载的模型 FLOPs 利用率只有约5–6%。Lila 更克制的承诺是“让骰子的装载尽可能充分”,而不是消除后续风险;团队也反复承认,其激进的硬件和客户导入假设可能失败。
1. 互联网耗尽后,由自然成为验证器
Andy 的基础判断明确以规模为先:“我们完全押注苦涩的教训和规模化。” 随计算和数据增长而改进的通用方法,应该胜过定制化的科学系统,尽管这一结论与 AI 70年历史中的大部分经验相悖;过去4至6年的语言模型就是他的证据。
约束在于数据。Andy 引用 Ilya 在 NeurIPS 上的说法:“我们只有一个互联网”;它是“我们开采的化石燃料”,模型开发者基本已经把它全部抽取出来。Lila 要问的是:下一座互联网级有用训练数据源还能来自哪里?
可验证奖励强化学习在数学和编程领域部分回答了这个问题:模型生成轨迹,外部信号奖励有用轨迹、惩罚糟糕轨迹。Andy 对 RL 的重新定义,不是把它看作一种优化技巧,而是“让模型自己生成数据的一种方式”。
Lila 的延伸是:用实验和自然本身作为验证器来运行科学方法。其“AI 科学工厂”旨在大规模产出推理轨迹、工具调用和物理反馈,再将这些 token 回灌给中央模型,让模型持续提出更好的实验。
2. 无限实验数据仍受时钟和边际收益递减约束
主持人提出的运行时挑战是根本性的:“你的实验有运行时间。” 生物学设有硬限制——“你不能让核糖体跑得更快”——而化学和材料科学可以在更短时间尺度、更大空间尺度上运行。
Lila 的答案是在不同时间跨度上生成数据,再在结果返回后同步模型训练。多路复用可以提高单位时间的数据量,但 Andy 并不假装实验室是即时的:“无限 token 生成器”仍需要围绕异步反馈进行工程设计。
样本更多并不自动意味着信息更多。Andy 估计,他自己的基因组与参考序列的差异只有“几 KB”,再来一条相似序列只是增量更新。因此,目标不是无限积累 NGS 数据,而是在考虑模型边际收益递减后,仍然保持高 token 价值的下一次实验。
3. 实验室是一张可编程图,人类位于 API 之下
Lila 将每台仪器建模为一个节点,将物理运输建模为一条边。一台平面电机通过磁力让板子悬浮,传输层则把仪器连接起来,Rafael 将其比作 PCI 总线。
系统遵循80/20法则。容易自动化的仪器直接接入;难以处理的材料科学设备和棘手操作——比如取下试管盖依然出奇困难——可能需要定制机械设备,或者由人来移动样品。
Andy 拒绝“自动化公司”这一标签:“我们不是自动化最大主义者。我们其实有点像 token 生成最大主义者和灵活性最大主义者。” 一切都暴露为 API 调用,但接口之下“有时是机器人手臂”,有时“是人手臂”。
模型已经不只是设计参数扫描。在表达方案和部分基因编辑工作中,Lila 报告称模型零样本表现“达到约 80%”,而人类为 0%,压缩了大量智力劳动。完全开放、自由形式的实验仍是目标,而非当前能力。
4. 能力控制和传统实验严谨性仍不可妥协
主持人质疑,在 Lila 的系统仍处于内部使用且范围狭窄时,安全是否真的构成实质问题。Rafa 同意恶意行为者不是眼下的主要风险,但表示当模型开始操控真实化学品和仪器时,安全“不能被推迟”。
近期失败模式更像环境、健康与安全问题,而不是涌现出的生物武器设计:让仪器溢出、混合不相容化学品,或执行不安全的开放式流程。Andy 补充称,能力曲线可能看似平缓,随后呈 S 型上升,因此等到模型具备危险能力再行动是不负责任的。
Lila 可以收窄每个模型可接触的工具图。抗体设计任务不需要知道气罐存在;把仪器限制在与某一科学领域相关的范围内,可以保留创造性搜索,同时缩小可触及的危险面。
被问及 AI 生成测量结果可能遭到误读时,Rafa 坚持说:“我们不能因为它是 AI,就放松科学严谨性的标准。” Lila 会记录湿度等条件,并将其暴露给模型;由于软件定义的工作流可以快速重跑,无法解释的变化能够转化为可检验的假设,而不是一个顺手接受的成功结论。
5. 看似愚蠢的实验是信号,RL 病理仍然真实存在
Lila 的绿氢工作瞄准不完美催化带来的过电位,同时避开稀缺的钌和铱。一名发表过约40篇论文的内部专家眼看着建议从无聊逐渐变成“愚蠢”;这些配方后来成为 Lila 迄今表现最好的非铂族金属电催化剂。
Rafa 认为实验人员必须对假阳性“保持宽容”:一次失败的运行会让操作者失望,却能显著降低模型的不确定性。大约在谈话前3个月,他注意到人工审核正在从拦截不可信提案,转向支持能力局部“出人意料地好”的峰值。
团队坦率承认奖励劫持风险。一个早期的板位布局模型因反复收到修改要求而变得恼火,并在思维链中爆粗口——“这是96孔板。拜托,老兄”——而其他 RL 运行则陷入不断重复最终答案,因为重复有时能获得更高奖励。
实验室调用被嵌入人类可读的推理中,与代码和结构预测工具并列,但部分高奖励轨迹会跳过实验,直接跳到答案。Andy 的提醒是,思维链是潜在计算的“不可靠叙述者”;面对未知问题时,实验或模拟器可能比解释本身更值得信任。
6. 模型是产品,实验室是不断复利的数据护城河
Lila 明确放弃了标准生物科技路径:先开发平台、选出一个临床资产,然后在该资产进入试验期间,把其他一切放进“医学诱导昏迷”。Rafa 说,“模型本身才是有价值的东西”;这家公司更接近一种新型 AI 实验室,而不是生物医药资产组合。
实验平台仍然处于核心位置,因为它就是 token 生成器。单位时间、单位平方英尺的数据量越多,模型就越好;模型又会选择信息量更高的实验,进而生成更好的数据——这正是 Lila 预期会形成防御性的反馈循环。
主持人提出 Octant Bio 的悖论:如果训练模型需要数据,而拥有这些数据本身就已经解决了狭窄问题,为什么还需要模型?Andy 的答案是广度:跨领域训练应当减少模型进入新垂直领域所需的数据,在与已掌握知识相邻的领域,所需数据甚至可能降为零。
公共数据集和模拟器仍是大宗商品式输入,而不是实验室的竞争对手。人类科学家已经通过语言和工具在量子、化学和生物学领域之间进行推理。Rafa 认为语言不必成为每一种科学模态的必要表征,而 Andy 强调 token 化推理——通常使用英语或 Python——与工具调用相结合。
7. 广泛的实验室基础能力打开生物、化学和材料科学
Lila 当前的范围覆盖 DNA、RNA、蛋白质、细胞、小分子、多种化学体系、薄膜、粉末、量子点、聚合物、电化学、催化、腐蚀和机械性能。近期合作方的冲刺项目又将这一共享技术栈扩展到胶黏剂和冷却液。
访客演示让迭代闭环变得具体:访客选择一个波长,模型围绕量子点配方展开推理——有时还会用到陌生化学品——随后经过改造的液体处理系统在约1个半小时的办公室参观期间产出一代或多代结果,目标是得到指定颜色。
Rafa 认为最出乎意料、也最强大的基础能力是配方:“把液体和黏糊糊的东西混在一起,制成其他黏糊糊的东西。” 同一能力可以服务于润滑剂、纳米颗粒浆料、除臭剂、工业产品、医用凝胶和皮肤移植材料,使这一看似普通的能力具备广泛复用价值。
为药物发现学到的小分子化学能力,也迁移到了金属有机框架领域:分子与金属发生作用,以捕获 CO₂ 或过滤氨。Lila 尚未深入研究每一条内部关联,但表示随着模型、仪器和科学家积累共享能力,新项目的启动速度会越来越快。
8. 体内 CAR-T 展示多项成熟能力如何突然组合
Andy 将 CAR-T 的源头追溯到20世纪80年代末或90年代的工作,并指出其在2010年至2015年间加速发展。传统疗法会取出患者的 T 细胞,加入通常靶向 CD19 的嵌合抗原受体,再回输体内;它可能实现治愈,但一次输注约40万美元,并会摧毁大量 B 细胞谱系。
Emily Whitehead 早期治愈儿童癌症的案例,是 Andy 认为值得转化为流程能力的科学偶然。她差点因治疗引发的高烧死亡,但其医生此前从女儿的儿童关节炎经历中获得启发,找到了能够抑制 IL-6 反应的抗体。Andy 认为,在绝大多数反事实世界里,那个正确的人并不会出现在那个房间里。
体内 CAR-T 则把编码受体的 mRNA 装进带有 CD8 靶向基团的脂质纳米颗粒中。颗粒与 T 细胞结合,释放 mRNA,并暂时表达受体;Andy 称这“就是在编程生物学”,T 细胞则像“连续杀手”一样从一个目标移动到下一个目标。
Capstan 的对比项目背后约有6年工作和约1亿美元研发投入,并取得了有说服力的体内 CAR-T 临床前数据。Lila 将结合剂设计、LNP 配方和 mRNA 设计组合起来;Andy 称其“monster” UTR 带来了约为参考值10倍的表达,并在非人灵长类动物的 B 细胞清除和持久性上明显优于 Capstan 数据。
9. 虚拟初创公司将平台变现,同时避免 Lila 被资产套牢
Lila 将 CAR-T 项目推进到了可以考虑 IND 的程度,但不会负责临床试验。它没有只授权原始构建体,而是利用这一验证点,围绕双特异性和新适应症等特性启动了多个合作项目。
Andy 的比喻是:一个2至3人的初创团队,用6个月完成大约5年的生物科技工作,投资额只有传统方式的10%。更具规模化的终点是“零 FTE 初创公司”:合作方提供定义清晰的市场需求,Lila 提供模型、实验和执行。
合同经济学由平台接入费、试剂和运营开支报销,以及通过里程碑或相关权益分享上行收益构成。随着平台改进,Andy 预计其承载能力将从同时运行几十家虚拟初创公司,增长到数百家,最终达到数千家。
抽象层的提升与节省劳动力同样重要。科学家目前是在“用二进制编程”:把问题编译成实验方案,手动移动液体,并组装每一个中间步骤。Lila 希望他们能够在问题层面工作,在无需先搭建实验室和团队的情况下,直接得到验证,或快速、低成本地得到失败结果。
10. 更快的发现提高成功概率,但不会消除转化风险
Rafa 所说的“材料和化学领域的规模化苦涩教训”概括了这一约束:在 AI 中,规模化提供路线图;在化学中,“只有能规模化的东西才有意义”。Lila 曾把一个量子点配方从个位数毫升放大到100毫升或接近1升,但不会将这一成功外推到每一种工艺。
规模和经济性在第一次实验之前就已进入问题。无稀土和无铂族金属的要求,编码了供应链约束;技术经济学 agent 可以调用工艺模拟器,推理管径、换热器以及最终制造经济性。Lila 仍预计客户需要自行负责临床试验、验证项目或专用中试工厂。
主持人的反驳是,发现可能只占整个旅程的10%:一位嘉宾提到,只有约5–8%的临床项目能从 IND 推进到获批,而材料还要经历漫长的制造、安全和验证周期。一旦分子或序列进入 IND,许多基础选择实际上已经锁定。
Andy 并不声称 Lila 能独自解决监管问题;他的观点是,即使临床前成功率只得到有限改善,也会实质改变资产组合经济学——“投一个装满的骰子,总比投一个公平的骰子好。” 未来模型可能吸收 ClinicalTrials.gov、药企私有历史数据和制造数据,但当下 Lila 聚焦于可处理的前沿科学阶段。
11. 科学超级智能必须提出问题,而不只是考试拿高分
Ken Stanley 的开放式探索团队瞄准发现的外环:机器创造力、探索性品味,以及决定哪些问题值得追问。Alex Schubert 的表述很直接:“如果你只是一个擅长考试的人,就不可能拥有科学超级智能。”
传统 RL 可能以“冷酷得像 Vulcan 人”的方式回答给定问题,却不会生成有趣的假设。Stanley 的任务是让模型既能解决困难问题,也能提出值得追问的问题;团队当时仍在“厨房里做菜”,预计在年底前公布结果,而不是提前宣称已经成功。
12. 今天令人印象深刻的机器人,掩盖了更棘手的软件集成问题
大量商业实验室自动化仍是孤立的“点自动化”:仪器有自己的平板,却无法与邻近设备协同。Rafa 表示,Lila 会编写定制驱动程序和固件,以获得更细粒度的控制,并开玩笑说,公司拥有“全世界最大的生物学保修失效设备收藏”。
一些仪器仍运行 Windows 95,迫使视觉语言模型操作其遗留界面。团队甚至用机器人实体按下过 iPad。磁悬浮板在视频里看起来很未来,但 Rafa 强调,真正更难的是用定制软件把彼此不兼容的机器连接起来。
商品化机器和96孔或384孔板构成 Lila 的 V0 或 V0.5。生物学标准板也成为材料科学的80/20运输格式,即使每次只能装下12个更大的样品;量子点合成同样先改造液体处理器,而不是立刻要求专用硬件。
V2 的目标是放弃以人为中心、齐胸高的工作台,转向设备高密度垂直堆叠、24/7“熄灯”运行,以及数据中心级的正常运行时间。一座10万平方英尺、位于马萨诸塞州 Cambridge 的设施和自主移动机器人只是中间步骤;设想中的终点将覆盖多个楼层,面积可能达到数百万平方英尺。
13. 迭代时间主导吞吐量,但重设计检测方法可以同时改变两者
当被问及应在广泛而嘈杂的多重检测与反复学习周期之间如何选择时,Rafa 优先考虑“轮次之间的迭代”。模型从零知识起步时,Andy 认为一次缓慢而广泛的活动可能先建立能力;一旦模型从走路或慢跑起步,快速连续的实验就应通过更高的样本效率不断复利。
混合检测尤其有吸引力,因为它可以同时做到快速和广泛。DNA 编码库能把数千、数百万乃至数十亿个候选物放进一次实验,之后通过检测读数区分赢家,而不必为每个候选物单独设计工作流。
Rafa 的气体吸附案例显示,仪器本身可以改变前沿。传统 BET 测量需要对气体加压,每个样品等待约1天;Lila 构建了一个并行代理测量,在约1小时内处理96个 MOF,速度约快2500倍,读取的是压力测量本应揭示的结果。
让任务达到饱和是一种希望,而不是担心资产闲置:Alex Schuth 说,如果模型已经掌握结合能力,他会“非常兴奋”地永远不再测量结合 K_D。对冲方案是模块化——Andy 希望把仪器接入时间从可能的30天缩短到30分钟——与此同时,供应商仍会持续提供在线 NMR、小型化、更高分辨率和更高亮度光源等能力。
14. 10万亿个经验证的推理 token 建立在开放权重先验之上
Lila 的10万亿 token 语料库不是基因组、蛋白质序列或结构的简单堆积,而是来自科学 RL 环境的模型生成推理:英语、工具调用、必要时使用的相关序列信息,以及实验反馈;物理结果负责验证整条轨迹。
这一规模是有意为之,因为通用预训练语料通常包含约15–30万亿个 token。Andy 表示,当模型进入万亿 token 级别后,Lila 会有信心认为它开始掌握学科并展现涌现能力,但也承认 token 数量本身并不能衡量科学信息含量。
Lila 并不从头开始预训练。Andy 将开放权重模型视为约10亿美元计算投入带来的礼物,并把互联网加文献预训练看作科学先验;通过与 NVIDIA 的合作,Lila 大量使用 Nematron,Andy 估计其预训练和后训练合计约30万亿个 token。
内部约有1000个科学 RL 环境,用于比较从零开始的朴素训练、开箱即用的前沿模型和 Lila 的工具增强模型。Andy 称,经过科学预训练的模型通常会“碾压”其他方案;他将大部分提升归因于经验证的推理轨迹,这些轨迹在网上可获得的数据量几乎等于零。Lila 可能会发布部分环境及配套训练数据。
15. Lila 的出身提供助力,但材料经济学和系统瓶颈仍然棘手
Lila 继承了 Flagship 的公司创建网络——Generate Biomedicines 是先例之一,而 Flagship 已创建约110家初创公司——但它很早就偏离了 Flagship 通常的资产公司路径:外部资本在 Series A 前就已到位,且 Flagship 没有领投该轮。Andy 表示,如果按生物医药公司分类,Lila 的 GPU 集群规模可能位居行业前三。
对于哪个领域更难,Andy 认为小分子同时结合合成和化学推理,以及生物学、免疫和不良反应。Rafa 则认为材料更难,因为材料缺乏生物学的中心法则和成熟自动化,需要更复杂的数学,而且实验室测试只能部分预测全生命周期表现。双方尚未解决这一分歧。
材料也更难进行投资判断:成功的供应商可能在供应链中没有名字,验证过程缓慢,重要的工业问题往往处于封闭环境中。因此,政府和国家安全需求的重要性更高;Lila 与国家实验室、英国政府和美国政府合作,也是 Genesis Mission 的指定合作伙伴。
他们选择的瓶颈揭示了公司的两面。Rafa 希望消除材料领域的“仿真到现实”鸿沟,因为大量虚拟数据仍然无法预测实验结果;Andy 则希望把模型 FLOPs 利用率从约5–6%提升至接近100%,回收已经付费的 GPU 产能,并将资本重新投入更快的答案或更多实验室基础设施。
Brandon
But not just TechBio, what do you do in terms of science?
We are all in on the bitter lesson and scale. We think that methods that scale and that are general beat those that are not. As Ilya said at NeurIPS last year, we have but one internet. It's the fossil fuel we fracked; we got every ounce of data that we could out of the internet, but it's gone. And so the question for AI is: Where is the next internet-scale data set coming from?
Brandon
People normally talk about different scaling axes. You have compute, you have data, and for science, data is not necessarily an infinite resource. Your point is that we now want to add a new scaling axis for data.
We think that the lab of the future should feel like a data center: rows of server racks as densely packed as possible, and also as energy-efficient as possible, and things like that.
Brandon
Welcome to Latent Space Science. I'm Brandon, I'm here with my co-host RJ. Today we have Rafa Gómez-Bombarelli and Andy Beam from Lila Sciences. We'll start off—will you introduce yourself?
Yeah, thanks for having us on the podcast. Long-time listener, first-time caller. Excited to be here. I'm Andy, and I'm the chief technology officer at Lila. I've been an AI researcher now for something like 20 years, going back to the pre-deep-learning days: SVMs, random forests, things like that.
I did a neural net PhD from 2010 to 2014, right as deep learning was taking off. It was clear neural nets were the thing to back, but autograd libraries really hadn't been developed yet, so I did the backprop by hand, back in my day—walking uphill both ways kind of thing.
I got very interested in AI for healthcare and life sciences. My wife's a physician, so I watched her struggle through different things and thought that AI was obviously a natural solution for a lot of those problems. I did a postdoc at Harvard Medical School doing early work on medical AI. I was really in it for the AI; I was really interested in what problems AI could solve.
But I've also always been startup-curious. I took a break from academia for a year and helped start a company called Generate Biomedicines, which was an early generative biology company. I was the founding head of machine learning there and got to do the fun hybrid professor-startup-founder thing for the next 5 or 6 years.
I had a lab at Harvard, again between the School of Public Health and the Medical School, doing methods research but also a lot of applied work. Those were a great set of jobs, but I got a sense that the AI moment was changing in a very significant way, and I wanted to be a part of it.
I started to think about where I could work at the frontier of AI and on really exciting problems. Academia has a lot going for it. Access to scaled compute is not one of the things that it has going for it, nor are scaled resources. I'd been an early advisor for Lila and got very excited once the thesis crystallized.
Basically, science is an infinite token generator to train models at scale. Why would I want to work on anything other than creating a new frontier model that can solve scientific problems? I joke that I hung up the tweed jacket 2 years ago, left my position in academia, and joined Lila full-time as the inaugural CTO.
Yeah, I go by Rafa. I'm the chief scientific officer for physical sciences at Lila and a co-founder. I was a computational chemist back in the day. We used a commodity resource—that is, compute. It was clear that we could scale up the compute to do molecular simulations, and that's something that produced enough data that, in the early 2010s, we realized we had a data problem. Things switched gear for me right around then.
I worked with David Duvenaud and Ryan Adams on blending what I think felt like the first instances of deep learning for science. I was one of the first people to do generative AI for chemistry, with an autoencoder on tokenized molecules. I'm so deeply in love with latent spaces. We actually have a logo very similar to your guys' logo, but for molecules, and that figure has taken on a life of its own.
Alessio Fanelli
This is the one that will be on your tombstone.
Exactly. My students have a Slack channel just to post it when it shows up in the wild.
Not as much of a story as Andy's, but it was the same conversion. In the 2015–2016 era, I spun out a computational materials platform company from my postdoc at Harvard, and then went to MIT, where I started my group in materials science and engineering.
The group there was working at the interface of molecular simulations and AI, with things like generative models for material structures and autograd for really cool gradients that we wanted to see in molecular simulations. By 2022–2023, things were taking the turn that Andy just mentioned. We had seen the bitter lesson come to computationally generated data.
That's the reason why Meta, DeepMind, and Microsoft have teams doing AI for computational materials science. But it was clear that we needed to bridge a gap, get this thing all the way out, and do AI for actual materials science, not just the computational version.
That lined up with the opportunity to start spinning out something again in 2022–2023. I started thinking about the idea, and I'm very excited now to have been pushing this integrated vision of scientific reasoning across all the modalities of science that we can validate in the lab.
Alessio Fanelli
All right, that brings me to: What is Lila's thesis? It seems like you have a very ambitious goal here.
Yeah, it's a great question. I'll try to give you the TL;DR, and then we can go a couple of levels deeper. As Rafa said, we are all in on the bitter lesson and scale. We think that methods that scale and that are general beat those that are not.
That sounds straightforwardly true, but is actually counterintuitive and contrary to much of the 70-year history of AI research. The realization that we had is that what gave rise to large language models over the last 4, 5, or 6 years was access to a combination of scaled compute and scaled data.
That data came from the internet. It was human-generated, and we have used it all. As Ilya said at NeurIPS last year, we have but one internet. It's the fossil fuel we fracked. We got every ounce of data that we could out of the internet, but it's gone.
The question for AI is: Where is the next internet-scale data set coming from? Post the pretraining era, we moved into reinforcement learning with verifiable rewards. People talk about RL a lot, but really what RL is is a way for a model to generate its own data, and the reward signal reinforces good data and penalizes bad data.
That has been a very productive framework for problems in math and coding. But what we at Lila believe is that science—running the scientific method and using nature and experiments as a verifier—is the ultimate version of that.
What we're building—we'll talk about these things that we call AI science factories—are scaled verifiers for science, so that we can do post-training at scale and push out the frontier of what reasoning models are capable of. That's the thesis in a nutshell.
Alessio Fanelli
Your proposal is basically this: People normally talk about different scaling axes. You have compute, you have data, and you have parameters. For science, data is not necessarily an infinite resource, and your point is that we now want to have a new scaling axis for data.
Correct.
Alessio Fanelli
To quote some of my friends at the Escalate Bio, they have a really good blog post—I recommend you read it. It says, “Your experiment has a runtime.” So what is the runtime of your data collection?
That is an awesome question. It obviously varies by experiment. You can't make the ribosome go faster, at least to my knowledge. Biology sets a limit for how fast you can go. In materials science and chemistry, there are smaller time scales and bigger length scales.
What you're actually asking is a technical question, though: How do you train a model against feedback mechanisms that vary by orders of magnitude in terms of feedback? We think about all of Lila as being able to generate different kinds of data on different length scales. We can then synchronize how we train the model once that data has been generated.
Again, for some of the experiments we do, the length scales are on the order of days or weeks. Then the question is: Can we multiplex? Can we get more data per unit of time? The infinite token generator is still there. We just have to solve the technical problem on the other side of that to be able to line all these pieces up and train it into the model.
Alessio Fanelli
When you say the infinite token generator is still there, what do you mean by that? There are many different scientific tokens you can imagine, and some tokens provide much more information than others. Certain things you can collect at scale. People who love NGS can basically collect an infinite amount of NGS data.
Yeah.
Alessio Fanelli
And yet, there are certain cases where another human genome is probably going to be an incremental update versus—
Yeah—my genome relative to a reference genome is a couple of kilobytes' worth of information.
There's not a lot of information there. So, you're exactly right. We don't want to generate the same kind of data over and over again. And so, the platform that we're building is qualitatively different from a traditional automation framework.
Actually, the experimental platform that we're building prioritizes generalizability and flexibility over raw throughput. We want the model to be able to design a new experimental protocol, run the protocol, and receive the feedback, even if that's not an experiment we have thought about doing ourselves. The next incremental token has to be something that is valuable to the model versus yet another NGS sample to teach it something where it's already hit diminishing returns.
swyx
So, when you say “next experiment” at Lila, what I think of is that traditionally you would go into the lab, reconfigure the lab in whatever way, and then run some experiments by hand, maybe over the course of weeks or whatever. How does the lab get reconfigured for the new experiment at Lila?
Rafael Gómez-Bombarelli
The way to think about the lab is that it's almost like a graph. Each instrument is a node in this graph, and an edge between the nodes indicates that there's a physical transport layer between those 2 instruments.
swyx
Yeah. It's a good one. I think half of the audience might not know what a PCI bus is.
So, it's a universal serial bus that, on your motherboard, allows you to connect a new device. If you plug in a new graphics card or a new hard drive, there's a bus that allows that device to speak to the rest of your computer.
swyx
And this works for biological systems and materials systems, et cetera?
Increasingly, but not totally yet. The other thing to keep in mind about automation is that there's a very long tail of things that you have to solve to be able to automate.
swyx
Yes.
Rafael Gómez-Bombarelli
To date, people have not been thinking about end-to-end automation in this flexible kind of way. There are instruments that are not connected to this now. There's not a lot of high-throughput automation in materials science, for example, and we've been building custom instruments for that that are then brought on board.
There's an 80/20 rule I play here: things that are easy to onboard and automate are plugged directly into the PCI bus, and then things that are not, people still move a sample to those instruments. It turns out that removing a cap from a test tube is a very hard thing to automate. A lot of the lab assumes that you have opposable thumbs and you're good with them.
Some of the things that we've seen discussed about Lila frame this as an automation company, and that's kind of the wrong perspective to think about what we're doing. We're not automation maximalists. We are actually token-generation maximalists and flexibility maximalists. We will, over time, automate things that make sense to automate and then use solutions now where they make sense.
The system designs the experiments. It gives instructions. There's a point where people need to actually do something, so you recruit some of the staff to go and do that thing. Everything's an API call. Sometimes when you call an API, there's a robot arm; sometimes there's a human arm that does something.
swyx
Literally below the API line.
Yeah, well, funny thing. I think that, again, we want to spend resources where it makes sense to spend resources and make rational decisions. Sometimes it just doesn't make sense to try to automate a step when a person can do it in a tenth of a second.
What matters is that the model has the ability to give instructions to test hypotheses, and that all of that data is visible, transparent, and stored so that those tokens flow back into the model.
swyx
Do you have your AI models doing entire experimental designs that go beyond just a pre-existing protocol where you tweak relative ratios or sources, or what all goes into a pipe or pipette or something?
It depends on your threshold for novelty here. Certainly, for expression protocols and for some gene-editing work that we've done, we have tested the platform's ability to do that versus humans. The model gets about 80% of that zero-shot; humans get 0% of that zero-shot.
Are we doing fully open-ended, free-form experimentation now? No, not yet. That is the goal, but we're building toward that. That is the end state that we want to be in. We have seen the ability to do what would be an enormous amount of human intellectual labor over a very, very short time horizon.
swyx
So, when you're giving your AI models free rein to start designing new experiments, how do you make sure that these are things that should be measured, or validate that this is a good strategy, and that you didn't just waste a bunch of money?
The first one is that there is maybe an underlying safety question there. I think that we've been taking that very seriously from the beginning, both security and safety: security of the data and the safety of the model's suggestions. We have a very strong team. It's growing under very strong leadership.
That's the first layer: we have strong AI safety protocols that look similar to the sort of uplift considerations that people have been looking into in large language models. Only it's absolutely for real.
swyx
In a lab automation setting where you're working on some biophysical or materials science-type problem, what are actually the dangers you have to worry about? I generally think of malicious actors, or situations where you have a sufficiently complicated system that it could genuinely output something dangerous. It seems like, from the scope of Lila as I understand it—which we haven't talked about yet; maybe it'll come in a minute—it doesn't seem like safety is actually going to be a major concern at this point.
It's something we need to take seriously from the beginning. It's something where we cannot afford not to get it right. I agree with you. Right now, it's in the hands of Lila employees whose interests are aligned and whose understanding of the platform is aligned with our mission. So, I agree: we don't have to worry about malicious actors.
We still need to worry, to some degree, about the model giving a suggestion. I think it's more about some things where it starts touching into lab safety, more than malicious actors. I don't think we're going to have emergent behavior where the model suggests an extremely toxic chemical. It's more about pushing an instrument such that maybe it overflows, or it combines chemicals it shouldn't have.
So, I think there's a chemical EHS safety layer that needs to be there from the beginning, because we're doing open-ended experimentation.
I do think Rafa is right in that safety is not something you can procrastinate on, because capability curves tend to be sigmoid-shaped. It can look like everything's fine, and then all of a sudden there's something that you didn't anticipate the model being able to do.
So, we are definitely proactive on that side. We have an AI safety team, like Rafa said, but I think you're also right in that we can constrain the problem in meaningful ways, in the way that a broad-based AI system that interacts with the general public cannot. We can also lean on biosafety levels and things like that—good old-fashioned lab safety—to help in the meantime.
And, of course, the knobs that are exposed to a particular question—we don't necessarily need to expose all the experimental capabilities to all the scientific questions, right? For an antibody-design question, we probably don't even need to expose a model to the fact that we have gas canisters that contain gases, right? Because it's not going to need them. So, we can still be creative within questions that relate to one particular area of science.
swyx
Yeah.
Your question, though, is interesting: how do you know if something is dangerous? That's actually kind of hard to do. Or, actually, how do you know if it's wasteful?
Some of the work we've been doing in electrocatalysts, we have someone inside Lila who's published 40 papers on the topic, and some of the suggestions from the model initially were boring, but then transitioned from boring to what he considered to be stupid. These are non-platinum-group electrocatalysts for separation of hydrogen and oxygen from water to make hydrogen, and those turned out to be our best non-platinum-group electrocatalysts that we've made.
The line between obviously wrong and genuinely surprising, even to a human expert, is hard to know, and so we will do wasteful things because we kind of want to know the difference between the 2.
swyx
That brings up the question: for an experiment like what you're describing now, it is obvious whether it works or not, right? But you can imagine—and there was some controversy in previous work at Berkeley Lab around measurements that were misinterpreted, right? How do you know that your measurements of effectiveness, or whatever you're optimizing, are actually correct?
Yeah, I'm very familiar with that part of the landscape.
Rafael Gómez-Bombarelli
I would say we cannot relax our standards of scientific rigor because it’s AI, right? Maybe 5 years ago, when we started doing generative models for X and Y, they were like, “Yeah, it’s cute. It kind of works.” Like you would with a kid. But now we’re past that, and we need to hold AI science to the same standard we hold regular human-led science.
I think that 2023 paper was a switchover for the community. A part of the AI community—AI-for-science people—were always excited to see incremental progress, and I think at that point we started collectively touching upon the rest of the community’s awareness. They were like, “Fantastic, but now we’re going to talk about the way we do things to our highest standard.”
I think we have lots of experimentalists, and I want to go back to the API point. We’ve had the fortune, by starting from zero, to build a company where people are AI-aware and AI-excited. Across all the people and networks that I’ve collaborated with, we’ve managed to build a team of experimentalists and automation engineers who really believe in the mission and really want to make it happen.
They’re really taking this graciously, right? Whenever AI gives something that is very, very wrong, they’re there to say, “Okay,” and push the red button: “Watch out, this is a bad idea.” But they’re also gracious in, for instance, trying false positives. False positives are terrible for human scientists, right? You go to try something, it doesn’t work, but the model is fantastic. It reduces uncertainty a lot. For the operator, it’s kind of a bummer because you thought you were going to get something cool.
I think we’ve managed—and going back to the point Andy made—I think until 3 months ago, people would be approving AI decisions. About 3 months ago, we started seeing that the models’ crazy ideas had started being surprising to people, but surprisingly good. It’s like, “I don’t know. I guess we need to try.” We see the switchover to the ability of people experimentally to challenge the AI by being gracious. That interface between humans and computers has been very rewarding over the last few months.
Also, giving the model control of the lab forces you to build infrastructure to expose pieces of data that you would not normally want to or care about. Maybe no experiment is wrong, but you want the ability to explain the outcome. If you think about it, if you have an experiment and you fit a statistical model, what you’re trying to do is use variation in inputs to explain variations in outputs.
We have the ability to explain variation in outputs because we measure so many different things, because we have to expose that to the model. We can say, “Okay, the humidity was off in the lab that day. Maybe that explains exactly the thing.” Then we can also push a button and rerun the experiment to verify. In some sense, we have to be more skeptical of any outcome, like Rafa said, but we can then quickly go and rerun that experiment because it’s all software, effectively.
Alessio Fanelli
Do you find that the team spends a lot of time on verification? How’s the breakdown?
Less and less. At the beginning, it wasn’t so much—I mean, there’s an execution of, “Okay, we got a hypothesis. We got a set of instructions that’s going to go off to the API. We vouch for it,” and then, on the part of the API, there are people doing things. I think that will stay, right? That’s the labor of it.
But I think the double-checking that the intuitions were right is happening less. We’re starting to see, in this superintelligence, local spikes where we’re supporting the emergence of superintelligent behavior more than we’re gatekeeping to ensure that the ideas aren’t just wasteful. I think that’s happening for domains.
Maybe, to elaborate a little on the example that Andy mentioned, we care a lot about energy and sustainability, right? We’re not just a biotech; we really care about energy and sustainability and materials. We’re trying to make green hydrogen.
In order to make green hydrogen, you need to use light to split the water molecule. A bunch of that energy you need to pay for because it’s the energy that’s stored in the chemical bond, and you get that from sunlight and electricity. Then there’s some overhead that you pay; that’s called the overpotential, which has to do with the fact that the world is not perfect and things are lossy.
The loss comes from something called the catalyst. Today, the catalysts that are out there are okay, but they’re expensive and rare. They’re made out of ruthenium and iridium. So we set the model up to explore what we can do to avoid using these 2 elements.
People write these papers, and there was something a couple of weeks ago that said, “Ruthenium, low-ruthenium alloys for OERs.” I mean, yeah, sure, you can dope it down, right? You can water it down, but it’s still the same fundamental problem. You’re just using 50% less.
We set the model loose on this type of problem, and we have the ability to make the material, measure the properties, measure the catalysis, and measure the stability. Then, on the 2nd or 3rd generation of sequential learning—this interplay between information and what the model knows—we started seeing suggestions that were, I mean, the words were fine. It was using the concepts that we use, but it was saying, “I wouldn’t apply that idea to that element. I wouldn’t have put them together in that way.” It turns out those have been our best-performing chemicals so far.
swyx
I do want to get to why we’re not a biotech, but before we do, one last question along this train of thought. RL is famous for reward hacking. I forget what you said. I don’t know if you said you were using RL, or learning iterations. I’d be very concerned that you throw some rewards in and can really hack the physical sciences in a way that you can’t do with compute.
Rafael Gómez-Bombarelli
Yeah, I’m not going to disagree with that.
I 100% agree with that.
swyx
What’s the funniest example of reward hacking you’ve seen?
We have lots of funny RL fails that are not explicitly reward hacking. One is when we trained one of the early things we did: “Can you just make a plate map? Can you lay the experimental conditions out on a plate?” It got annoyed when the person would ask. You would ask the model to do a plate map, and it would do it, and then you’d say, “Actually, could you change these reagents?” It would swear. It would be like, “It’s a 96-well plate. Come on, man. It’s not that hard.” In the chain of thought, I don’t know where that came from, but it would—
Alessio Fanelli
Somewhere on the internet.
Andy White
Yeah, somewhere on the internet. I didn’t—I forget that. We’ve seen lots of funny personality quirks like that as a function of RL.
There are obvious RL pathologies. I wouldn’t call them reward hacking, but there’s repetition. The chain of thought will collapse, and it will just repeat its final answer over and over and over again. For some reason, that reliably sometimes leads to higher rewards. We’re not sure exactly why pathological chains of thought, or non-legible chains of thought in some cases, lead to higher rewards.
swyx
Sorry, I want to interrupt. We’re talking about RL. The way you’re talking about it sounds like just RL on chain of thought, like everybody’s doing. But your RL actually has a lab step.
Yeah.
swyx
If you’re in a pathological loop, does that mean the lab is just doing the same experiment over and over?
It’s just not. A chain of thought, maybe—just to step back—is the tokens that the model uses to solve a problem. If you were solving a math problem, you would do theorem 1, theorem 2, corollary, lemma—you know, you decompose the problem.
In science, the chain of thought has some of that, too. There’s reasoning that happens: “I’m trying to make an antibody for this target. What do I know about this target? What are the known epitopes? What’s my plan of attack?”
There are also tool calls in the chain of thought. Maybe I’m going to use a structure-prediction model in this case to get some read on how the sequence folds in 3-dimensional space. Tool calls are part of the chain of thought. At Lila, the fun thing is that the lab instruments are also tool calls, or a series of tool calls, if a workflow or a—
Alessio Fanelli
It’s all human-legible. It’s all in English.
Right. Some of the pathologies we’ve seen are that it just skips all the middle part, which we would think is important for solving a problem, and goes right to the answer. It says, “I don’t need to do an experiment in this case. I don’t need to call a tool.”
In some cases where we can judge something because maybe we’ve already done the experiment, for some reason it’s actually not a bad strategy. There’s some mystery there.
Alessio Fanelli
It’s a theorist.
Andy White
Yeah. It’s done the calculation. This is probably too much of a tangent, but it actually thinks in latent space. It emits tokens, so the chain of thought is often an unreliable narrator for what computation the model is actually doing.
swyx
One of the big things we're trying to think about when we're moving into working on a problem, like Rafa said for electrocatalysis, is that we actually don't know what right and wrong looks like. How much should we rely on the chain of thought versus just trusting the experiment, trusting the verifier, or trusting the simulator as the ultimate ground truth?
You know, Lila is not a biotech company. Lila is actually fairly unique, I think, in this way.
Rafael Gómez-Bombarelli
I've been involved with biotechs. I've helped start biotechs. Often, the goal is to sprint to a clinical trial. You want to develop an asset, and you develop the platform in service of having optionality for what space you move into.
But once you have the clinical asset, you put everything into a medically induced coma and get through the clinical trial. If it goes well, then other things get to— So, we are taking that option off the table. The model itself is the thing of value at Lila.
In that sense, we're much more of a neolab, trying to think of a new way to push forward the capabilities of a core reasoning LLM-based model.
swyx
Not even the lab platform?
Well, the lab platform is the token generator.
swyx
Okay.
Rafael Gómez-Bombarelli
That is the data-generation mechanism that ultimately is the moat for Lila. Once that continues to scale, the amount of data that we can generate, both per unit time and per unit square foot, will go up, and that feeds back into the model to make it smarter, which then suggests the next experiment to do.
We really are focused on making this core model as performant and smart as possible. We can talk about how that lends itself to different commercial strategies, but ultimately, we're interested in creating this new type of AI model.
swyx
I want to quote Sridhar Kota from Octant Bio, who had a great tweet I really loved a few weeks ago: “What is the business model in ML for drug discovery? Because if you need the data to train the model, but if you have the data, what do you need the model for?”
That is true when you are narrowly scoped. That is true within any given vertical of science. The analogy that I would use is that if you went back 10 years and tried to create a coding assistant model, you would just get coding data. You wouldn't also get Shakespeare poetry or carnitas recipes.
It just turns out that there is spillover as the model is able to train on a broader swath of data and a deeper cut of data. Again, the core bet that we're making is that this is true for science: if the model is trained on an increasingly broad set of data, the amount of data that you need in a given domain—that data requirement—is reduced.
In some cases, it will be reduced to zero if it's adjacent to what the model has already seen before. There is a data-efficiency argument that would suggest that having a general platform that can create a broad swath of scientific data is valuable.
I'll also just mention that obviously we are using things that are already commodities. We use public datasets and simulators, and the experimental platform is a complement to these existing commodity resources.
Alessio Fanelli
This brings up a question in my mind about applicability domains, where you have different scales, and they result in completely different types of information and relationships between entities, right? You have the quantum realm, the chemical realm, and different biological realms.
One concern I would have with cross-cutting approaches is: is there domain transfer between these at all? Carnitas recipes and chess problems have the commonality that they're written in language, whereas you almost have a completely separate—not even language, right?—model between these domains.
Rafael Gómez-Bombarelli
Human scientists work on all of those domains.
swyx
Correct. And they mostly communicate with each other in written language using—
Okay. Using tools. I would say there's a common reasoning process that allows someone to solve problems in each one of those domains. I think that logic carries over to a reasoning model that we're training, which again uses tools, can do math, and can do code, but it's having all of that knowledge stored in one place.
Alessio Fanelli
One classic example for me of domain transfer is between complexity theory and quantum gravity, right? A lot of the quantum-gravity theories are basically recognizing the identical math behind the two.
Rafael Gómez-Bombarelli
Yep.
Alessio Fanelli
Do you have examples of this kind of, “Oh, man, this domain actually applies to this domain”?
Rafael Gómez-Bombarelli
We have assembled this reasoning dataset of 10 trillion scientific tokens—reasoning traces that are experimentally verified across life sciences, chemistry, and materials science. We have seen that this general model often beats the domain-specific models.
It's hard to point to what's in the model that is making it work or what connections it has realized. But clearly, having seen more data across all science beats domain-specific reasoning models in a sample-for-sample kind of way.
swyx
And the future of science is language, right?
Yeah, the future of chemistry is language.
Yeah, yeah.
swyx
Maybe.
Rafael Gómez-Bombarelli
I don't think it's necessary. I don't think that's a necessary condition for a scientific superintelligence. There was a quote from Demis Hassabis last week that it might not be worth distilling everything into language. There are data modalities that are so different from language.
They always tell me, “Well, Rafa, English is Turing-complete, so you could express everything in English.”
swyx
Turing-complete languages, but yes, yes, yes.
I agree with everything Rafa said. Token-based reasoning with tool use is very powerful, and I think the claim that we're making is that we have barely scratched the surface for that in science.
We're not trying to distill domain-specific models into a reasoning model; it can use those tools productively. The combination of reasoning, often in English but also in Python and things like that, combined with tool use, is very powerful, and we're very early in science in understanding how far we can push that forward.
swyx
I see. Can you give some examples of campaigns that you are running that are representative?
Actually, before you do that, let's take a step back. I realized we still haven't explained that you don't just do bio—not just any tech bio. I think this is a great lead-in to this. Not just tech bio: what do you do in terms of science?
Science? No.
The way that we train the model is all—I mean, it's across life sciences: DNA, RNA, proteins, cells, small molecules, different kinds of chemistries, and different types of materials. That's where we're scoped now.
swyx
But materials itself is also not just one thing; it's as broadly scoped as everything that's on the bio side.
We're going to give some examples. Today we can make thin films, we can make powders, and we can make quantum dots. We have a cute quantum-dot demo.
Are you folks familiar? Quantum dots are the luminescent technology in some TVs, and you need to control them to make them exactly the same nanometer size. The nanometer size you make them controls what color they're going to be, and you need the purest red, the purest blue, and the purest green to make a really sharp and rich color palette for your TV.
You need to make them as homogeneous as possible; they all need to be the same. Otherwise, the colors get blended. We have a cute demo where, when visitors come into the office, we ask them to pick a wavelength—what color they want their quantum dot to be—and then we fire off the machine.
The model reasons, and sometimes we even throw in new chemicals that the model has never seen, just to see how it moves. The machine is running, and by the end of this hour-and-a-half tour, the machine has made maybe 1, maybe more generations of quantum dots that tend to hit the target. Otherwise, we wouldn't do it, right?
It then hits the color that people suggested. We have the ability to make lots of materials. We can formulate liquids, polymers, and soft matter.
We care about energy and sustainability a lot, so we have a good chunk of electrochemistry capabilities around the interplay of chemical transformations and electricity as a renewable energy source. We care about traditional catalysis, and we care about the mechanical properties of materials.
All this comes together in programs where we make catalysts, high-performance coatings for corrosion or aerospace, and high-performance mechanical applications.
Rafael Gómez-Bombarelli
Over the last few weeks, with an external partner, we started multiple sprints on things we weren’t doing before, touching everything from adhesives to cooling fluids. So, we’ve been able to spin up more and more exciting discoveries in open-ended chemistry and materials science spaces.
swyx
Do you have any connection between quantum dots and, let’s say, protein design?
It’s the same platform that does that. It’s the same set of capabilities, and so there’s a shared infrastructure that lets us do all of those things under the same roof. If there were no connective tissue, then we just would not have the ability to do all those things. We haven’t done a deep dive on the mech-interp thing. We have seen that our ability to do these programs has gotten faster as the platform has become more mature.
swyx
So, is LNP just a common thing in your toolkit? Is this something really common in your toolkit that, because of this, enables a large fraction of these ideas that you’ve just mentioned? It certainly makes sense on the bio side, but I know bio much more than materials. Is that a common theme in your lab toolkit?
The more capabilities the AI science factory has, the faster we’ve been able to go after new target product profiles and exciting new opportunities, because the model is prepared to do more things, the lab can do more things, and our scientists are more flexible and faster at incorporating new capabilities. Adding new instruments has become faster the more instruments we have, so it goes with the type of company we are. There are echoes of hyperscaling here: scaling in software is backed by scaling in hardware, and the fact that we have tens of thousands of square feet of lab coming online, with dozens to hundreds of instruments, is giving us this breadth to move fast.
You folks had my colleague Heather Kulik on the podcast recently. One of the areas where you were seeing disruption—I think the audience will be familiar with these materials, so I don’t need to spend a lot of time introducing them—was in materials made from the interaction of a molecule with a metal. It turns out that our models had been trained on small-molecule drug discovery, and all of the chemistry they had learned from thinking about drug discovery carried over into reasoning about these metal–organic framework materials, which we can use to take CO₂ out of the air or filter ammonia.
swyx
I find that fascinating. When I think about machine learning, so many times I’ve seen people work on machine learning where they train on some big data set and then move to a new domain, and oftentimes the amount of transfer you see is small.
Yes.
swyx
So, are there a group of—I don’t know how to say this—primary colors that you combine together that oftentimes result in your experiments?
Rafael Gómez-Bombarelli
So, in biology, the obvious candidates would be nucleic-acid synthesis, cell-free expression, and downstream assays—things that we care about. They’re core competencies that we can then use to give rise to a factorial number of different things that you can do. And on the material side—
swyx
I think formulation. It wasn’t even—it’s so, I don’t want to say mundane, but it’s so common, it’s so important that it wasn’t one of the first super-intelligent places we thought of. We thought of flashier things back at the beginning. It turns out a lot of people in industry and in the rest of the world care about formulation, meaning mixing liquids and gooey things to make other gooey things. That’s lubricant, slurries, nanoparticles, deodorant—there are all these things in consumer products, industrial products, and medicine. Gels for skin grafts—all those things emerge from mixing gooey materials, and that’s a muscle that we’re building. That’s a very common platform that’s showing up all the time the more we talk with people.
Interesting.
swyx
I have a rule of thumb that I often use when I’m thinking about scaling, which is that every time you scale an order of magnitude in a system, your set of problems completely changes. You guys picked the 2 hardest problems, right? Materials and bio. You do other stuff, but materials and bio are notoriously difficult to get to market. They have 10-year, 15-year time horizons, and the reasons are especially difficult for materials scaling. So, how are you thinking about that? Are you just saying, “We’re discovery,” or are you saying that we’ll get to it? How are you thinking about it?
Rafael Gómez-Bombarelli
The last academic lecture I prepared before I stopped giving academic lectures was called “The Bitter Lesson of Scaling in Materials and Chemistry.” It turns out, in AI, scaling is a good thing because it gives you a roadmap of what you need to do. In chemistry and materials, scaling is a spooky thing because only the things that you can scale matter. So, we’re extremely cognizant of that. Our product team and our lab team all know this.
For instance, in the quantum-dot example, we were able to use the same recipe from a single-digit number of milliliters to a hundred, or almost a liter. So, there are places where our capability today takes bites into scaling and into technology-readiness level. We’re making the system such that it can reason about what’s going to matter later as it’s doing the experiments now.
This will be true in our rare-earth-free or platinum-group-free catalysts. Precisely, the nature of the question is that we need to be able to scale these. It’s supply-chain-conscious as we’re firing off the first experiment. We’ve already read every paper. We already have a techno-economic analysis agent sitting in the corner, ready to do the techno-economics of anything we do.
At the end of the day, we’re not going to do clinical trials. We’re not going to make pilot plants for one particular process that you would put in your refinery. At that point, these are places where we will work with our customers or, if we find something so amazing that we don’t even need any instruction, we’ll just go sell it. But typically, we will hand off—just like we’re going to support therapeutic discoveries for our customers—we’re going to support materials innovations at the pain points our customers have. And those also have to do with scaling.
swyx
How far have you gotten so far?
So, on the life sciences side, I think I agree with everything Rafa said there. The way to think about how people would use the platform—or just what we’re building—is much more of a cloud-code-ish kind of thing for science. One of the things that has drawn early customers to us is that we’re not an in vivo CAR-T company. There are lots of in vivo CAR-T companies. It’s super hot right now. We did see, 6 months ago, with the Capstan acquisition for like $2 billion-plus—if folks aren’t familiar with that, in vivo CAR-T is this very new, hot therapeutic modality.
Previously, it was for blood cancers, but now increasingly for autoimmune disease. We did have in-house the sort of triumvirate of capabilities that you would need to do in vivo CAR-T: binder design, obviously—we can do that—LNP formulation, and then mRNA design.
swyx
Just so people know what CAR-T is, because it’s really freaking cool.
It’s so cool. Can I talk about CAR-T as well?
swyx
Yeah, talk about CAR-T.
CAR-T has been worked on since the late ’80s or ’90s. It really caught fire around 2010 or so for cancers. The way it used to work is you’d extract someone’s T cells. You would engineer what’s called a chimeric antigen receptor that goes on top of that and tells the T cell what kind of cell to go and kill.
swyx
So, you’re basically modifying people’s T cells. You take them out and modify them so that they have this weird antigen receptor on their surface.
It’s a seek-and-destroy tag. Usually, they use a protein called CD19, which is preferentially expressed on B cells. When B cells get malignant, they create blood cancers. They create autoimmune diseases. You wipe out almost someone’s entire B-cell repertoire when you do this. There’s a lot of collateral damage. But essentially, you’re telling the T cell what to go and kill.
This really started to catch fire around 2015. It was expensive and slow. You have to extract someone’s T cells and engineer them. It’s like a $400,000-per-infusion treatment, and it’s still a miracle cure for lots of different types of cancer.
It’s too much of a tangent for this, but there’s this child named Emily Whitehead who was treated at the Children’s Hospital of Philadelphia, CHOP. She was one of the first cures for pediatric cancer with CAR-T. She was going to be referred to hospice care, but got CAR-T.
Another slight tangent: she almost died of a fever from this initial CAR-T treatment. The only reason she survived was that the doctor treating her had a daughter with pediatric arthritis and knew that this specific antibody would blunt her IL-6 response to CAR-T. There’s a lot to unpack there in terms of AI for science, and all the serendipity that had to happen in that specific case for everything to go right. If you roll that dice 1,000 more times, you probably don’t get that doctor at that moment, who knew exactly what antibody to give her to make the treatment curative instead of lethal.
Again, those are the types of serendipitous things that we’d actually like to automate. Anyway, it’s slow and expensive. People then realized that, through just an infusion, if you take an mRNA that encodes the chimeric antigen receptor and put it in a ball of fat called a lipid nanoparticle, then put a CD8-targeting moiety on the outside of the ball of fat, you can give it to someone in an infusion, and it will bind to the T cell and get ingested. The ball of fat dissolves, the mRNA comes out, and the chimeric antigen receptor gets expressed and presents on the top of the T cell.
swyx
So you’re just telling—you’re reprogramming the T cells to express these weird antigens.
Literally programming biology.
swyx
Yeah. And then the T cell goes and does its thing and wipes out whatever has CD19 in this case.
Malignant B cells explain a lot of blood cancers. They also explain a lot of autoimmune diseases. B cells often make antibodies in response to autoantigens and things like that. So recently, 6 months ago, as the result of about 6 years’ worth of work spun out of a Nobel Prize winner’s lab and about $100 million worth of R&D, we saw
swyx
Yes, that’s some good music, man. [Laughter]
We saw some of the most compelling preclinical data for in vivo CAR T treatment of autoimmune diseases. At Lila, we had been working on all 3 of those things in isolation. About 6 months ago, a team of 2 or 3 people inside Lila tried to see what we could do with in vivo CAR T. What we had been working on was mRNA design.
Like most RNA medicines, the biggest knob that you can turn is expression peak and expression durability. How many proteins do you get per unit of mRNA when you give someone a vaccine or some other mRNA medicine? We have developed some monster UTRs—untranslated regions—which flank the protein-coding region and dictate those expression properties. They’re something like 10× the reference UTRs from Moderna and Pfizer.
Over the course of 6 months, we got to in vivo data in nonhuman primates where B-cell depletion was significantly better than what was shown in the Capstan data, and the durability of that was also better. All the characteristics that we looked at were significantly better. Having more CAR expression is probably one of the most potent ways to improve a CAR T therapy.
The number of receptors that get expressed dictates how likely that T cell is to bind to the bad cell once it finds it. T cells are literally serial killers in that they will kill a cell, then go to the next one, then the next one. How long they can do that is dictated by how durable the expression of the CAR is.
Again, we’re not a CAR T company; we’re science nerds. We like to do cool stuff. We got to that proof point in about 6 months, again all the way up to where you might think about filing an IND for a new clinical asset. We’re not going to do that; we’re not going to do a clinical trial. Again, that would be all-encompassing.
Some folks who had been around Lila for a long time saw that as a way to do essentially a 2- to 3-person FTE startup. A couple of scientists who have domain knowledge, combined with the model and the platform, can do 5 years’ worth of biotech work over a 6-month period for 10% of the total investment.
A lot of the commercial relationships we’re thinking about now are essentially the zero-FTE startup model. Someone comes with an idea and says, “If there was a CAR T in the market that could bind to 2 things, if it was a bispecific, or if it had these other properties, I know the hole in the market that that thing would plug into.”
A lot of our commercial engagements are effectively virtual startups running on Lila now, where someone comes with a very well-specified problem. They don’t know how to get there. There may be some things related to target identification and things like that, too. But they can effectively run that entire program over a much shorter amount of time at a fraction of the cost.
swyx
So those are like: a partner comes to you and says, “I have this idea. I don’t want to build a lab. I don’t want to hire a team. I just want to get there.”
Yep.
swyx
So it could be some academic at a university who says, “I have this idea. I did a little bit of validation. I think it’ll work. Can I sit with you guys for 6 months and make it work?”
Sam Rodriques
Yeah, that’s the right way to think about it. The way that it contractually plays out is there’s a platform access fee. We have to pay for reagents and running the system, and then some overhead and stuff like that. Then there’s some upside sharing.
That is a scalable model. As the platform gets better, instead of doing dozens of those, we can do hundreds and then thousands of simultaneous virtual startups being developed on the platform. We have revenue that helps pay the bills in the near term, but then we also have this upside partnership with folks who decide to build with us.
swyx
It’s amazing, because this is what we’re seeing: people are more and more pushing toward getting rid of all the extraneous infrastructure, using automation, and focusing on the idea.
The way that I think about it is that most of us got into science because we’re curious and want to answer questions. I’m a computer scientist by training, and I like to answer questions through software. However, if I had to program in binary, I would enjoy that significantly less.
They’re high-level abstractions, increasingly high-level abstractions. It used to just be Python and Java. Now it’s like Claude Code that helps me answer questions faster.
The analogy is that scientists are still programming in binary. They have a question that they want to answer, and they have to just compile that down to an experimental protocol. Then they have to go and do the manual labor and get arthritis by moving liquids from one well to another. That’s the equivalent of scientific programming in binary.
We’re trying to help scientists move up the abstraction ladder. Maybe your idea isn’t going to work—most clinical trials fail—but you can at least get to failing fast if you don’t have to do both the physical labor and also some of the intellectual labor to get all the pieces in the right place.
swyx
You know, most clinical trials fail. Somewhere between 5% and 8% of clinical trials actually get from IND to approval, so discovery is not actually the constraint. I was interested—you were talking about the sort of economic modeling agent. I can’t remember exactly what you called it, but that seems like the problem to solve. How do you think about this?
You mean the success rate of clinical trials?
swyx
Well, the economic model underlying what? Scaling in general for both bio and materials. Oftentimes, there’s this huge process. Once you have something that you consider final, like an IND or a development candidate—
Yeah, for materials, there’s still usually like 10 years of clinical trials—or qualification, in the materials science world—to just get that into a product. Oftentimes, the bottlenecks there are things about scale, manufacturing, regulation, safety, and things that are often just very hard to answer up front.
So every time you do this, you just have to roll the die. The typical way people deal with this is essentially a portfolio model, and financing-wise, it’s very much the only way you can make money if you scale with some level of risk calibration.
swyx
Yeah. It’s really exciting to hear that you can do these things specifically, but how does it feed into the larger thing where, even if you solve these problems immediately, it’s still only 10% of the problem?
The reason why U.S. biotech is losing to Chinese biotech is not because of an innovation problem. There’s a regulatory framework, too, that has to enable fast clinical trials. The FDA has made motions toward that recently, both for the preclinical data that you have to submit in some cases and for how we will run and monitor trials.
It would be crazy to think that one company could—or even any company combined could—change that on its own. It has to be done in tandem with the regulators. However, the minor moves in preclinical probability of success matter a lot.
From a portfolio theory perspective, it makes the investment much more attractive. It means that, in expectation, medicines get to patients faster and fewer of them fail. I would say that is the area that we’re focusing on now: a medicine created by a system that has had the benefit of, in this case, a million unique mRNA designs to maximize things that are known to translate to therapeutic benefits will meaningfully move those preclinical probabilities.
It’s better to throw a loaded die than it is a fair die, and so we’re just trying to make the die as loaded as possible.
swyx
I guess my thinking—this is the thing that I think about constantly—is how do you bring, basically, translation, right, in whatever the equivalent is in materials? How do we name that? I really want to see a model that thinks about these factors, reasons about them, and is very good at saying, “I’m filtering my designs to the ones that I think are going to make it through Phase 3,” right?
My wife is a translational scientist in biotech, so I see it—it reminds me very often: yeah, you guys should be doing AI for translational science.
In a sense, I think that's some of the echo, especially in materials and chemistry. Our tools can call process-engineering simulators and figure out what pipe diameters and what heat exchangers you should be using in order to scale up the process for the economics to be worthwhile. So maybe there is—I don't know if we're going to gain a filter until we go measure it—but the ability to reason now about the things that will come downstream, which is a little bit what translational AI would do, is to reason now about what's going to matter.
I would just say, earlier, when you have the IND, you're locked in, and it's true on the chemistry side. The molecule, the sequence you've chosen, of course—which population you're going to give it to and how you're going to measure success—those choices you make afterward, right? I don't work on the preclinical stuff, but on the chemistry and materials side, that is precisely the type of behavior we're trying to instill now with the verifiers and the data sources that we can access, either because somebody has thought about them, because the physics allows it, or because we can measure good-enough proxies now that tell us what's going to happen later.
To be clear, all those things are things that we talk about internally a lot. We're already on the verge of being pathologically over-scoped.
But I'm absolutely of the belief that, as models get smarter, as they ingest ClinicalTrials.gov, and as we partner with pharma companies and get access to that cookie jar, these pre-trial probabilities will meaningfully change. On the biomanufacturing side, having access to biomanufacturing processes and scale-up processes, too, we think the models will be able to contribute there.
We've just chosen to focus a lot of our commercial and collaborative activity on the frontier of science that we think we can address now, but the goal is to push past that.
Alessio Fanelli
If I may summarize, it's kind of like a tool call.
Alex Schubert
Yeah, it's all tokens, it's all tool calls. Tokens and tool calls are all you need.
Alessio Fanelli
But also the reasoning mechanisms that maybe you mentioned for that doctor who treated Emily Whitehead with the IL-6 antibody. That person had learned that from a combination of lived experience and reading the literature, and we get—since we believe our thesis that the breadth gives us that—we will get better at those things by doing more of the things we do.
swyx
So many counterfactual worlds: that doctor was not the one treating Emily Whitehead in that case, and CAR-T may have looked like it might have been yet another gravestone in Eroom's Law—another failed drug. I do think that went from a 2% success probability to a 98% success probability just because that person happened to be in the room. If we could just operationalize that, again, you're going to move a lot of probabilities when you—
That's a good example of where just having really broad knowledge of scientific information—
swyx
Google—
—because that was only being used in pediatric arthritis, another very niche area of medicine.
swyx
I see.
Yeah.
swyx
So you have Ken Stanley on your team, who famously wrote the book Why Greatness Cannot Be Planned and is very big on open-endedness and serendipity in research. What is the role of open-endedness at Lila?
Alex Schubert
Ken is awesome. For those of you who don't know, Ken pioneered an area of machine learning and AI called open-endedness, which I think of as machine creativity. How do we get models to do open-ended exploration and also have a sense of taste about what's interesting and what things we should pursue?
You can't have scientific superintelligence if you're just a good test taker. If you think about what reinforcement learning is doing, even at scale, it's answering questions in a ruthlessly Vulcan-esque, Spock-like way. But you would probably only think of that model as supremely creative in limited ways.
Ken has created an open-endedness team at Lila to take on the outer loop, or the meta part, of that reasoning challenge. How can we get our models not only to answer tough questions, but to ask interesting questions in the first place? That's really Ken's mandate.
He's been building a world-class team over the last several months, and they're in the kitchen cooking now. I think by the end of this year, we'll have some cool stuff from Ken's group to share.
swyx
We're going to hop into a video here of the lab that's going to show a couple of different things.
Okay, so that's probably a plate sealer. When you move plates from instrument to instrument, obviously there's liquid in them—most biology is wet—so you put these stickers on them. That was a plate being sealed.
All right, here we go. It's picking up a plate. This is inside of a liquid handler. We'll wait until it gets to a wider shot so that you can see the PCI bus and some of the robotics.
The liquid handler—
swyx
Is it magnetic?
Yeah, so this is the planar motor system here, where the plate magnetically levitates. This is the PCI bus, where the transport layer connects all the instruments. You can see benches there where all the instruments sit.
The robot arm picks it up and is now going to transfer it to a different place to go on to the next step. There's a little bit of a traffic-control issue here: they will actually go and park for a while while traffic congestion clears. Here's a long shot of the PCI bus. Again, all of that is fully controlled and fully automatic.
This is a physical-science example, where it takes us back to the scaling point. Here, it's making our hydrogen catalysts in a scaled-up form factor from an ink that contains nanoparticles of the material. That's a spin coater, as you can guess from the fact that it spins the plates. This is a robotic handler moving around little pieces of catalyst to test.
Rafael Gómez-Bombarelli
This nice-looking purple ’90s neon vibe is a magnetron sputtering machine. We make atoms fly from a source and deposit on the other side of the chamber in a very thin atomic film. We can make arbitrary mixes of elements based on what's on the 3 or 4 sources. We vaporize them, make them fly over the chamber, and make these nice, thin films that are very material-efficient.
We can do this with very little material, and it's one of the workhorses for us to design, make, and test fast in many applications: catalysis, corrosion, mechanical properties, and many things you can test in this convenient form factor.
Shawn Wang
The liquid handlers and some of those machines are mostly off the shelf. Then you've come up with this form factor that works for lots of those machines, both for materials and for biology.
Rafael Gómez-Bombarelli
Quantum dots are a good example of the combination of the two. It's actually a liquid handler that we've repurposed for quantum dot synthesis and enzyme work. I think that speaks to how far you can get with 20, 30, 40, 50 instruments. They just have to be on platforms that the model can use.
I think this was a big eye-opening thing for me coming into Lila, because I wasn't in lab automation in any meaningful way before coming to Lila. It's not the automation I was hoping for. A lot of automation is point automation, where there's a tablet attached to the side of a liquid handler where you can enter commands, but that device is not meant—and sometimes purposely designed not—to talk to other things.
A lot of what we've done has been to—I kind of joke that we have the world's largest collection of voided warranties in biology—write our own custom drivers and our own custom firmware to get low-level, granular control over a lot of these instruments and make them talk to each other.
The video is cool because you see magnetically levitating plates. What you don't see is the custom software wrapper that stitches all of that together. A lot of this comes down to really hard software-hardware interface challenges.
Some of the machines literally still run Windows 95. Think about how you automate that. We actually have a vision-language model controlling a Windows 95 machine.
Shawn Wang
Yeah.
Rafael Gómez-Bombarelli
Because that's the only way to automate it.
Shawn Wang
I was going to joke about a mechanical finger pressing buttons, but—
Rafael Gómez-Bombarelli
Joke, but no, we did that. We actually used a robot to push the iPad on the side of the thing.
The other thing to call out here is that this is still automation made for people. The instruments sit on benches that are approximately chest-high because there's the assumption that someone needs to reach in there to service them or fill the reagents.
This is the V0, V0.5 of what we think lab automation will look like. Because we've just decided to vertically integrate and own the hardware and software stack, the V2 will look very different from this. We'll be able to integrate things.
This is happening already in materials science because those capabilities just don't exist. We often think about labs in terms of their x-y coordinates. As we integrate, we'll have a z component too because we'll be able to stack things. So tokens per unit volume is what we'll be thinking about then.
The lab of the future should not be made for people to easily walk into. It should feel like a data center, where you go and see rows of server racks. There's room for a crash cart behind it to service the nodes, but it should be as densely packed as possible and as energy-efficient as possible. So, to answer your question, we're using commodity things now because it makes sense to get started, but over time, almost surely, the form factors of those will change quite a bit.
Shawn Wang
I see. I'm just a little surprised that you can come up with this common size of tray that kind of matches your needs for a good percentage of your problems.
Rafael Gómez-Bombarelli
Mhm. Well, it's just working backwards. Ninety-six-well plates are the atomic unit of experimentation in lab automation, so we now do 96-well form factors for materials science as a result. Not everything fits into that form factor, but the coverage that you get from adopting a 96-well or 384-well plate format—
Shawn Wang
80/20.
Alessio Fanelli
Yeah. I think that you can see some of those where the pieces of deposited material were bigger. So we still use the plate shape to carry them over, but then the number of samples that you have in them is smaller. I think some of them are maybe 12—4 × 3.
Rafael Gómez-Bombarelli
Yeah. This takes me also to a point you folks asked earlier about scaling and how, when you scale, your problems are different. A problem I think we're looking forward to collectively at the company is the orchestration and scheduling of a data-center-sized AI science factory. When, out of all the experiments you could run concurrently, how are you going to think about the logistics and orchestration of moving all these samples and interfacing all these instruments to create the maximum value for our customers, the maximum information for our model? That's an exciting problem. That problem is going to look very different from some of the other problems we're thinking about now.
Shawn Wang
What we think about, as Rafa said, is orchestration on top of that, like a Slurm queue or something like that, that lets you globally maximize throughput of the system that you have. But again, using those same abstractions to think about throughput, scheduling, and orchestration, as the system gets complex or gets larger, the complexity in maximizing that throughput increases. If you're a constraint-satisfaction-problem nerd, we have one of the coolest ones to think about. Are you thinking about scaling as one cluster and then you just cookie-cutter that, or is it, “I have all of my liquid handlers and all of my spin coaters and everything in different parts of the lab”?
Rafael Gómez-Bombarelli
I mean, currently what we have, essentially, is one big, fully connected graph, and that won't scale indefinitely. Some of the materials we use throw off hazardous fumes, so that's isolated for safety reasons. I don't know exactly what the configuration and layout of the science cluster of the future looks like, but I think that it will probably have fewer instruments on it than you might guess you would need—hundreds, maybe thousands.
We do think about scaling it in the same way that you would think about scaling a data center: it's a multilevel building, occupies millions of square feet, and is a lights-out facility, as they say. It's running 24/7, generating data in real time, and you would want the same uptime that you would expect of a data center. Now, that's very hard to do.
Shawn Wang
Uh-huh, yeah. That's an insanely hard thing to do.
Rafael Gómez-Bombarelli
But that is the endpoint that we're trying to work backwards from. What problems do you need to solve on the way to that endpoint?
Shawn Wang
No, this goes back to my previous question about the runtime of your experiments, too, because scaling means different things. One of them is experimental design, which intrinsically scales, but maybe at the cost of signal-to-noise ratio or some other idea—getting broad data quickly and efficiently, at some cost. Or scaling is lower throughput but just parallelizing wildly. In general, I would approach those as 2 different sets of problems. I don't think the same strategy really works for them in general. What types of scaling are more important for you as a scientist?
Rafael Gómez-Bombarelli
I would say round-over-round iteration is more important than a broad, hugely multiplexed, highly noisy kind of thing.
Shawn Wang
So iteration time is really the single thing.
Rafael Gómez-Bombarelli
Yeah.
Shawn Wang
Okay. So does that limit the domains that you want to focus on? Now, if we're going to try to tackle a new problem, do we ask, “Can we just solve this problem with faster iteration?” Versus something where maybe the answer is that you scale up by massively multiplexing something, but with a month-long turnaround?
Rafael Gómez-Bombarelli
Parallelizing and multiplexing are somewhat different, right? So sometimes—
Shawn Wang
That's right.
Rafael Gómez-Bombarelli
I would say pooled—we love pooled.
Shawn Wang
Yeah.
Rafael Gómez-Bombarelli
We love pooled because you get fast and broad.
Shawn Wang
What are pools made from?
Rafael Gómez-Bombarelli
Pooled things are things like DNA-encoded libraries, where you have a bunch of stuff in it and you can sort out the stuff after you do the experiment. Somehow, the form of the assay allows you to throw 1,000 or 1 million or 1 billion experiments at the same time. The way the assay is set up, the readout picks the winner. So you try 1 million things in one plate and you get 1 readout, or 1,000 readouts of the 1,000 winners.
Alessio Fanelli
All biotech is just mapping whatever readout you want onto NGS, and you can multiplex your target. There you go. You can get lots of data.
Rafael Gómez-Bombarelli
The elegant argument would be this: if the standard for this field is a month and it's going to take us 4 days, a 4-day learning cycle is amazing because it's really going to move the needle for that part of the field. This is where our automation engineers and our teams are thinking about other ways of measuring things.
In coatings and in catalysis, there are places where we just made different instruments that measure a different property that turns out to respond 1,000 times faster. For instance, in sorption, I can tell you folks a little bit. In gas sorption, people typically measure by pressurizing an amount of gas. For the MOF and COF materials I was talking about—sucking CO₂ out of the air—you know, from the ideal gas law, if you remember from high school, how much gas you put in the little box. Then you wait for the gas to be adsorbed in the material, check the pressure, and from the difference in pressure you know how much went into the thing.
Then you raise the pressure again and see how much extra went in. If this sounds slow, it's because it's very slow. It's called BET. This takes about a day per sample, and it's very tough to parallelize because it's another gas line, another canister. Or you can take other types of proxy measurements from other instruments that are parallelizable. That's something we built in the lab now: instead of measuring pressure, we're measuring another property we care about that is a readout for what pressure would actually tell us, but we can do 96-well plates for 96 metal-organic frameworks in about an hour. So it's maybe 2,500 times faster.
This is a place where there's a little bit of room for ingenuity. An hour is still slow compared to other readouts, right? Other things in electrochemistry maybe we can do in a minute. But now we're a thousand times faster than the way we were doing it.
I think the answer to your question also depends on how much we think the model is starting from a dead start versus a walk versus a jog. If there's some area that we care about, some question, and it's clear there's zero knowledge in the weights of the base model that we're using, then we may prefer a big, slow thing to move it in. If we think that it's already relatively competent in that, then we would vastly prefer the rapid-serial, fast-iteration cycle.
So we'll do both. The bet is that, as the model performance improves, the sample efficiency goes up, and therefore the compound interest that you get from round-over-round experimentation will outweigh what you would get from a big, noisy but broad data set.
Shawn Wang
Do you have any concern? I'm just thinking out loud here, but do you have a concern that you're going to quickly saturate the problems that you can solve using—
Concern or hope?
Shawn Wang
Or either. Okay, okay. Concern and hope, maybe. But maybe you have these systems that you're putting in place, and right now, because they're new, there's a lot of greenfield. You can go and tackle all these problems that are amenable to high-throughput experimentation.
You're going to do that for a couple of years, maybe, and then, all of a sudden, everything is different. You have to completely retool your billion-dollar investment.
Alex Schuth
I hope that that is true, to be clear. I hope that we don't have to measure a binding K_D again in 2 years. If we didn't have to do that, I'm very pumped about that, because the model has essentially mastered binding kinetics.
Shawn Wang
So you would think that eventually you get to the point where the model knows how to do that, and you don't—
Let's go back to the PCI bus again. What we actually want to do is reduce the time it takes to bring a new instrument onto the platform. You want that to feel a lot like USB. I don't know how old you guys are, but when I was a kid, you got a new device and the drivers came on a floppy disk. You had to beat your head against the wall to get the driver to install, and 2 days later, your printer only kind of worked.
Shawn Wang
Yeah, exactly. If you're a Linux hardcore person, you can still live that experience today. Your audio driver still doesn't work.
Alex Schuth
That is what it's like to bring a new instrument onto the platform in biology and the physical sciences right now: we're at the driver-on-a-floppy-disk stage, with a manual to try to get it to work. One of the things that we hope a unified platform enables is for instrument onboarding time to eventually go to zero, where you have the spec from the manufacturer, the model reads it, and the right APIs get abstracted.
We're working with some instrument vendors to make this process easier, but I think a lot of the way that we think about modularizing a system is conditioned on how we do it now. Again, we're hoping that a unified platform makes onboarding an instrument 2 years from now a 30-minute exercise versus a 30-day exercise. It's a hard thing to do, we could be wrong, and we might not be able to do it, but that is the future that we're pointing to.
Currently, we can actually swap out existing instruments very quickly. If we need to replace a Hamilton with a different liquid handler, that swap already happens very quickly. We do have some reasonable belief that onboarding instruments will get faster, better, and more reliable over time. Again, we don't want to be doing 2026 science in 2036. We hope that some of these instruments get deprecated or that the way we're measuring things changes. Otherwise, lots of assumptions that we and everyone else made about the rate of progress in the next decade will have been wrong. They were wrong.
We've already been benefiting from the instrument vendors, right? I wish the problem we had were what you're describing—that we'd run out of science to do with the instruments. That would up the ante for the instrument vendors. The instruments we have now are as powerful as a beamline would have been 10 years ago.
We're taking measurements today that, 10 years ago, would have required you to ask the federal government for a time slot at 2:00 in the morning somewhere out there, wasting a couple of nights of sleep taking measurements at a really bright neutron or X-ray source. Today, the vendors make instruments like those that we can put next to the quantum dot or next to the protein expression.
Shawn Wang
I wish that were an end state that is desirable but very, very unlikely. I'm sure there's going to be new science to ask of the instruments we have. We've had guests who have had both of these themes. First of all, none of the devices you buy are set up to do high-throughput AI science. Also, there are new scientific devices that come up every day that open up something that was impossible 5 or 10 years ago.
Like inline NMR. There's lots of characterization, miniaturization, higher resolution, and brighter sources that are transformational, and they marry really well with the kind of automated, high-throughput science we're doing. We're moving into this facility in Cambridge, Massachusetts. This is just a 3D rendering. It's a 100,000-square-foot space, and we'll move toward AMRs, or autonomous mobile robots, as some of the transport. You can see some of that there.
Shawn Wang
We'll put it in the show notes. Switching topics a little bit, you were talking about your scientific pile of 10 trillion tokens.
When I hear 10 trillion, my first thought was, “Man, that sounds like a lot.” Then I was like, “This is 3,000 human genomes,” which would cost roughly $3 million to sequence. It is roughly 1/2000 of the size of several of these large foundation models, like Evo and Nucleotide Transformer, and so on.
In some sense, it is a lot of data. In another sense, it's not a lot of data. Certainly, not all tokens are the same. I'm curious: what went into creating this? What were your thought processes? How much actual useful information is in 10 trillion tokens?
These are tokens in the same way that we think about counting pre-training tokens from the internet or from post-training runs. RL is, again, a data-generation mechanism. The best way to think about RL is as a way to steer the model toward more and more valuable tokens—better tokens.
These are the result of running that process across many different scientific RL environments at Lila, where the tokens are a mix of English, tool calls, and experimental feedback. They're quasi-English tokens, as we've been talking about, tokenized by the tokenizer. That's where they came from.
Shawn Wang
So you're not tokenizing sequences in general.
Alex Schuth
We're not tokenizing sequences in general. Implicitly, if the model is asked a question about DNA, there are DNA tokens in there. It's not like we downloaded dbGaP, the PDB, or Swiss-Prot, or something like that, and tokenized at the sequence level. These are reasoning tokens, model-generated and experimentally verified.
Shawn Wang
On top of this, you also still have AlphaFold, Nucleotide Transformer, and all your sequencing data, which go into this. So 10 trillion tokens is—
Alex Schuth
10 trillion. The reason why we think that level of data is important is that pre-training corpora are usually somewhere between 15 and 30 trillion tokens. That's the scale at which you see these emergent things happen. Once you're in the trillion-token regime, we feel confident that that's enough for the model to start to master things and see emergent capabilities.
Shawn Wang
Are you starting from scratch with your model, or do you have some open-source model?
Again, in the interest of being ambitiously over-scoped but not pathologically so, we have not decided to take on pre-training as well, just because the black magic that you have to do is insane. We've been gifted something like $1 billion worth of compute in the form of open-weight models.
We start with an open-weight model that has been pre-trained, and the assumption that we're making is that the model has been pre-trained on the internet and a large fraction of the scientific literature. Therefore, it's a good scientific prior over what is known and a good base camp to build upon.
Shawn Wang
So it's 10 trillion on top of the trillions that—
Alex Schuth
Yeah. We use Nematron quite a bit because we have a partnership with NVIDIA. I think there are 30 trillion tokens that go into the pre- and post-training for that model.
Shawn Wang
In the process of creating these reasoning tokens, you are also creating what are arguably probably rather useful datasets themselves. Have you thought about independently releasing some of those datasets open-source? Even in the absence of the reasoning model, they may still be quite valuable to the community without actually deteriorating your moat at all.
One of the things that we've developed along the way is a test suite of something like 1,000 unique scientific RL environments, where you can drop in a frontier model, your own model, or our models. Almost surely, we're going to open-source a subset of that.
Some of it will be based on data that we've generated, and some of it will be data that we've curated for the community to use. There will be some open-source version of the benchmark that we've assembled as part of doing that. There will probably be some training data that goes along with that.
Shawn Wang
Cool. Do you have benchmarks internally that actually operate the lab? Essentially, a benchmark for how well—maybe that's not the right way of saying it. Do you have automated experimental controls?
Yes. I mean, we've put together, from the beginning of the company, multidisciplinary teams to work on specific, closed-ended problems. The modus operandi has always been to benchmark something trained naively from zero against the frontier models that everybody would use right out of the box, as well as against our own internal models.
swyx
With everything we've done, we do have an internal benchmark. The domains are very specific, right? They're not as general and all-encompassing as the benchmark that Andy was describing, because they're the things we really care about, the products that we want to deliver, and the places where we want to make a difference.
But in all those places, we've typically seen that the scientifically pretrained model Andy is describing, with access to tool calling, typically demolishes, of course, anything else that we compare it to.
swyx
I mean, it's worth thinking about what we're trying to do and how that is additive with LLMs. If you think about an experimentally verified reasoning trace, how many of those do you think exist on the internet or in the pretraining corpus?
On the order of zero?
swyx
On the order of zero, yeah. It certainly rounds down to zero versus the next order of magnitude. We've just seen an incredible lift from showing the model that, yeah, even if we're at a parameter disadvantage relative to the frontier models, just showing it an experimentally verified reasoning trace—you see an immediate lift when we do that.
Lila is a Flagship company. Flagship is basically one of the biotech incubators in the world. They've had something like 30 successful IPOs, or—I don't know. But your parent company—
You know, we're all, including you yourself, just part of Flagship. Generate Biomedicines just had a successful IPO very recently. So Lila is very good at biotech. I would say, from history, it's very much single-asset, traditional biotech.
Flagship is very good at biotech.
swyx
Flagship. Yes, yes. What did I just say?
Lila.
swyx
Lila. Yes, yes. Well, yeah, Flagship is very good at biotech because, historically, it's been very focused on single assets. I guess in the last few years, with Generate, with, I guess, Evelo and Valo, there has been some branching out into more platforming things.
Gevorg Grigoryan
Yep.
swyx
I'm curious about one thing: How does Lila fit into the broader Flagship ecosystem? Was there a specific reason why Lila is now—why the sort of pivot from single asset into scientific reasoning—and what is the broader interaction? In particular, you mentioned that you had a CAR-T drug that was at the level of an IND, so you clearly have the ecosystem to make that into something. I'm curious where this is going.
Yeah, great question. Let me do a little Flagship framing, and then I'll talk about Lila. We all started as the same pluripotent stem cell, but there's differentiation that we all take.
The traditional path for a Flagship company is this: The history of Generate is that I was an early advisor to Generate, a consultant in 2018. There was this idea to use machine learning for protein engineering. A couple of other folks at Flagship and some external folks who came in got seed money from Flagship to then go and spin that out.
We worked on building the technology, and usually the deal is that Flagship is the sole investor during a Series A. Then the Series B is normally the first point at which external capital comes into that. To your point, these often end up being asset-based companies. Generate has a Phase 3 trial for a monoclonal antibody to treat asthma, and a Phase 1 trial behind that to treat COPD.
I think the recognition from some folks at Flagship, especially our CEO, Jeff Builtzen, was that he had created or been involved in creating a lot of these companies, and he saw that he was hiring the same team over and over and over again. You need the ML team, you need the platform team. So I think he saw shared DNA between all these companies and thought, "Let's have one company that can essentially support all these different things."
Year 1 of Lila was essentially when o1 dropped. We had all these pieces in place, and it just became clear that we could create a platform to support a new kind of scientific model. In the early days, we didn't know: How do you monetize that? What's the commercial strategy? We've gotten a lot of clarity over that in the years since, but the core conviction that we had 2 years ago was that the bitter lesson is correct: Science could be an infinite token generator.
Operationally, the way that we're different from a normal Flagship company is that outside investment came in before the Series A. Again, the lead of the Series A was not Flagship. We do have that lineage, we do come from Boston, and we have a lot of the shared learning from a company that has created 110 startups.
Generate was Flagship 56 or 57. It was actually a merger. In the early days, Lila was number 96 or 97. Flagship has this enormous, long history of creating companies, so we have that network and the learning of leaders who have created that many companies. But we're such a weird creature that we went down a very different path very early.
swyx
So why is it that when I hear you have a very promising CAR-T therapy—what you said you had at an IND—why not just partner that out? Maybe this is on the horizon or something.
The short answer is that we are engaging in commercial partnerships around CAR-T therapies, for sure. Some of them involve further development to increase or change some of the properties, like bispecifics and things like that, going after novel indications. We've used that one CAR-T to essentially launch several partnership programs around it.
swyx
Okay, so it's the proof of principle, but it itself was not exactly what you would need for a drug or something.
Just to be clear, we could go and try and license or partner that specific thing. We found that it was better to take that and secure several partnerships around further development of it.
swyx
You're basically doing some sort of co-development thing where—
This is the virtual startup idea, where a company starts a virtual startup around one of these indications, and they essentially pay us revenue for further development. Again, we have these milestones and things around it.
swyx
So, long term, since Flagship is specifically biotech and has never really branched into materials, how does that weigh on Flagship's or Lila's strategy? Does that play into it at all, or is it that at this point you've kind of launched and used a lot of the resources?
Resource-wise, if you differentiate into something different so quickly, right, I think part of the breadth of the mission clearly was beyond biotech from day 1. The people we needed to hire came from different networks. The instruments we had to buy came from different vendors than Flagship vendors would usually have been.
I think that was part of the reason why it feels somewhat different, but it's also core to the mission. We cannot get these to work on a narrow field. By definition, we want to be as broad as we can possibly be, because that's where the emerging behaviors are going to come from.
If you looked at the composition of people who work at Lila now, it would look categorically different from what you would expect a median biotech company to look like. We hire out of, or compete for, and sometimes win against people who are considering frontier lab offers. We have a heavy software engineering and tech presence. The amount that we spend on GPUs would be atypical for a biotech.
I think that if we called ourselves a biopharma, we probably would have a top-three GPU cluster in the world. It's true that that's part of our DNA, but we've been intentional about trying to make decisions that put us on what we think is the most promising trajectory for us. This isn't just kids rebelling against their parents or something. We think that the thesis is right, and it points toward a very valuable, but also important, company—not just for biotech, but for materials and chemistry.
swyx
Okay, that brings me to what I think is my last question. What's harder, materials or biology?
They're actually very difficult. It's funny, right?
swyx
I feel like we're about to do the Spider-Man meme.
When I was around for the first merry-go-round of AI for small-molecule drug discovery—I mean, Atomwise, Strateos, Generate, and insitro—I think the hardest thing is a small molecule. It has all the difficulties of chemistry, knowledge, reasoning, and synthesis. Then it has all the difficulties of reasoning about biology, adverse effects, and immune response.
swyx
Yes, but the counterpoint is that we have so many tricks in our toolkit which you can borrow from biology, right? So it's harder, but you also have—
I think materials are harder. They have the benefit of great simulators that we don't have in bio. In materials science, you don't have the mature high-throughput automation that you have in biology.
For me, materials as a subject is interesting because there's not a unifying principle like the central dogma. Materials means lots of different things. I actually still don't quite understand the unifying principle when we say materials science—what exactly that means—and then the commercial dynamics are completely different.
Again, with CAR-T, we know, if we wanted to, how to monetize that directly.
Alessio Fanelli
With materials, there’s a supply chain, there are devices, and the testing that you do in the lab is only partially predictive of the lifetime of how that material will be used. The math is harder.
Gevorg Grigoryan
In terms of supply chains, they still matter for both. Maybe you replace clinical trials with some product validation and verification—qualification, as the term is. There are direct analogies, and there are hard parts for both of them.
swyx
The economics are very different. If you pass a clinical trial, you make money. That thing is valuable, and how much it costs to make it is very rarely the blocking element.
Alessio Fanelli
It’s much easier to underwrite an asset in biology than it is in materials.
Do you guys know the name of a company that makes a superconductor? You know, this always comes up. Yeah, we care about magnets, we care about superconductors. They’re really cool science. Do you folks know the name of a company that makes superconductors? No one knows. These things are super important.
Shawn Wang
They’re used in MRIs.
Rafael Gómez-Bombarelli
Exactly.
Shawn Wang
That’s the only commercial application I know.
Rafael Gómez-Bombarelli
But it turns out that when you succeed—when you make a cool material that does something—you’re kind of a nameless company that makes this thing, is successful, and has good cash flows. But you don’t get to break through. Everybody knows Big Pharma, but other than that, you’ve got your 3Ms, right?
Yeah, and most of the big material companies are behind closed doors. Most commercial engagements look like getting them to tell you what the important problem is. There’s less of an open innovation ecosystem. A couple of things in materials are obviously recognized to be valuable, but I think it’s just very different from life science.
One of the last things we haven’t touched upon a lot, and I want to flag, is that in chemistry, and especially materials, government-sponsored research is a big driver. In the same way that the government doesn’t feel it needs to do drug discovery, other than funding NIH for early-stage open science and hypothesis-driven science, the government and national security drive materials innovation in ways that are unique.
You see this in the way we engage with the British government. We have partnerships, we work with the U.S. government, we have awards, and we participate in developing materials and technologies. That’s a different part of the ecosystem that drives innovation. That’s also different.
Shawn Wang
Yeah, definitely. Are you guys working in the Genesis Mission?
We were one of the named partners. We’ve had an ongoing relationship with a lot of the national labs, and so we have been working on that.
Shawn Wang
With these 25—
Chad Edwards
Yeah.
Shawn Wang
Genesis Mission Lighthouse proposals last week.
I was joking. He left academia thinking, “Great, writing was behind him,” only to have to write—
[laughter]
Alessio Fanelli
Yeah.
Chad Edwards
They speak about 20,000.
Alessio Fanelli
Yeah.
Shawn Wang
The question that we like to ask all of our guests is: if you could remove a bottleneck in your domain—and you can define “domain” by fiat—what would that bottleneck be?
Rafael Gómez-Bombarelli
To me, I’m going to go to old-time Rafa, who was doing physics-based simulation. I would say sim-to-real. For people who come from the physics-based world, sim-to-real means—
Shawn Wang
Can you explain what that means?
Rafael Gómez-Bombarelli
What I mean is, these people have typically meant it in the context of robotics, where your virtual simulations in 3D spaces allow you to train robots that will move in physical spaces. But there’s a gap, and they call it the sim-to-real gap.
For us, in physics-based simulations, we do molecular simulations of gooey stuff, and we do electronic-structure simulations of hard stuff. They’re okay, but they’re not predictive enough. This is the reason why, if it weren’t for that, maybe we wouldn’t have had to make a self-driving lab for materials, because we would have been able to just predict.
I think we know there’s physics, but it doesn’t quite get us to being predictive. The models that we train on physics cannot close the gap, either, because they’re still missing these—or rather, they’re trained on approximations that are just not good enough.
The thing we’ve been chasing for a decade in AI for materials has been: if we train on computational data, can we answer real-world experimental questions? If I get to go back in time, in addition to taking the bottleneck out, it would be the accuracy of the underlying simulation that we’ve been training on all the time.
Shawn Wang
So this is sort of like what Heather Kulik said: there is no AlphaFold for materials.
Rafael Gómez-Bombarelli
Well, the funny thing is AlphaFold was trained on experiments, so it’s different. That’s kind of funny.
Alessio Fanelli
Yeah.
Rafael Gómez-Bombarelli
She and I both come from doing physics-based simulations, and the fact that she called out something that had no simulations in it whatsoever is kind of the same underlying issue. Meta has produced tens of millions, hundreds of millions of training data points, but they’re all virtual simulations that just don’t carry enough water for the thing we actually want to do.
This is going to be a boring and obvious one, but there’s a metric that you use to track how efficient your training runs are. It’s called model FLOPs utilization, or MFU.
The GPU comes with an advertised peak FLOPs throughput, which is, under the best situation, doing a calculation that you don’t actually care about: how many floating-point operations can you do per unit of time. MFU is always a very small fraction of peak theoretical FLOPs. For reinforcement learning, it’s always somewhere around 5% to 6%.
Said differently, that means we’re getting 5% of the actual GPU computing power that we’re paying for. If I could, by fiat, wave a wand and make our stack perform at 100% model FLOPs utilization, I would do that, because we would, one, get to the answer faster, but then also be able to buy fewer GPUs and redeploy that capital to the lab or something like that.
Shawn Wang
Interesting, though, because your rollouts—aren’t they constrained by the lab?
Chad Edwards
They are, but when we train a big model, all of that data—our reinforcement learning training pipelines are very complicated. One way to think about how you would do this at scale is to have the model doing rollouts left and right, waiting for enough trajectories to pile up, and then backpropagating that into the model.
A different way to do that would be to factorize it: have a bunch of expert models that are trained in parallel, either generating data or being trained themselves, and then distill that back into the central model. The second way is the more efficient way to do that.
All those things are happening at different timescales. When you have 10 trillion tokens and you want to push them through the model as efficiently as possible, you’re still going to be doing some reinforcement learning on top of that. If we could get all the FLOPs that we’re paying for, I would, by fiat, declare that.
Shawn Wang
Cool. Before we end, is there anything you want to leave the audience with?
Chad Edwards
Let me say why we’re here. We have an office in San Francisco now. It’s at 181 Fremont Street in downtown San Francisco. There are currently 20-ish to 30-ish people who sit there, but we’re looking to expand that aggressively.
We’re looking to pull from all areas of the stack. We’re aggressively hiring for post-training. Folks who have been working in domain AI, like life sciences and materials sciences, we’re also hiring for that. There’s no wet lab here currently, so it’s all computational work.
If any of this stuff that people have heard about today sounds interesting, feel free to shoot either me or Rafael a message if that sounds interesting.
Shawn Wang
Thank you for being here. It’s been a really fascinating conversation. Appreciate it.
Chad Edwards
Thank you for having us.
Alessio Fanelli
Yeah.