🔬「最具创新性的 Diffusion 研究正在药物发现领域发生,而不是图像生成领域」
Genesis Molecular AI 认为,基础 AI 研究的前沿已从熟悉的 LLM 架构转向用于生成3D分子结构的扩散模型。 GANs 在蛋白质和蛋白质–配体系统上失败,而扩散模型提供了迭代生成物理结构所需的“正确原语”。主持人更尖锐的招聘话术是:LLM 实验室很大程度上仍在重新排列2017年发布的 transformer 层,而“最具创新性的 Diffusion 研究”如今发生在结构预测领域。
Genesis 将约1 Å的准确度视为蛋白质–配体预测的实用门槛,因为传统的2 Å尺度可能掩盖化学上致命的错误。 在2 Å精度下,芳香环可能发生翻转,但输出看起来依然合理;氢键通常只落在2.7–3.3 Å的距离窗口内。Evan Feinberg 的表述很直接:“药物发现本质上是一门分辨率科学”,而1.8–1.9 RMSD的模型可能生产出由 agent 放大的“slop”。
Genesis 已将 LLM 的部分 scaling 技术栈迁移到分子领域:合成数据预训练、推理时迭代计算,以及最终的强化学习。 公开结构数据库只有约200,000条记录,而估计的类药小分子数量达到10^60,因此物理模拟可以提供额外训练数据。推理时,模型“以晶体结构的方式思考”,反复完善内部表征,同时由基于物理的引导来控制扩散过程。
Genesis 认为,AI 杠杆最大的一层,是已知疾病生物学与临床测试之间缺失的设计层。 知道致病靶点,与能否成功将其药物化是两回事,有时甚至负相关;Feinberg 认为,对于具备强遗传学证据、药代动力学、安全性和转化模型的候选药物,所谓临床成功率10%的笼统统计“其实是明显偏低的”。这一机会既包括从零到一发现结合剂,也包括为现有药物寻找更好的后继者,后续代 ALK 抑制剂就是例证。
一款可行药物需要同时优化30多个 ADMET 相关终点,而不只是获得准确的结合构象。 效力、选择性、溶解度、膜通透性、细胞色素 P450 抑制、hERG 风险、组织暴露等属性中的任何一项,都可能单独淘汰一个分子。更棘手的是,这些目标彼此反相关:让化合物更“油”可能改善结合,却损害溶解度;随后增加极性,又可能阻止其进入细胞——这相当于在分子尺度上“打地鼠”。
Genesis 的前瞻性数据优势,来自将模型接入真实药物项目和实验室反馈:Incyte 提供已披露的合作项目,Insitro 则提供快速化合物生产和测量能力。 已披露的 Incyte 工作,既包括推动已有化学先导物进入开发候选物阶段,也包括为一个没有专利、论文或共晶结构的靶点找到首批已知结合剂。与 Insitro 的合作形成反复的设计–合成–测试–分析循环,并积累覆盖结构、效力和 ADMET 的训练数据,但合成与高保真验证仍然难以自动化。
Sapphire 是 Genesis 将专业模型转化为全天候药物发现劳动力的尝试,同时不把科学家排除在战略控制之外。 一个 LLM 负责编排构象、效力、ADMET 和化学工具,让药物化学家不必掌握每个参数;设想中的结果是“数百人的虚拟科学家团队”全天候运行。Feinberg 不认同彻底替代人类:专家负责设定方向和评估结果,agents 则承担执行与重复性工具调用。
OpenBind 提供了 Genesis 此前因私有合作数据而无法展示的外部泛化测试。 在未见过的 EV-A71 3C protease 靶点上,柔性环必须绕配体移动,Pearl 据称取得了远大于公开分布内基准的性能差距,并且“几乎每一个构象都基本正确”。剩下的约束是算力:两位嘉宾都把 GPU 视为瓶颈,而 Edunov 认为,相比生命科学领域,“纯 LLM 领域剩下的 alpha,已经变得有些可疑”。
1. 扩散模型成为分子 AI 缺失的原语
Edunov 加入 Genesis 前的路径,始于物理学,随后进入软件工程、FAIR 研究,并负责 Llama 2 和 Llama 3 的预训练。加入公司担任 CTO 后,他得以“回到物理学的根源”,同时将大规模 AI 方法应用于分子系统。
Feinberg 则从相反方向走来:物理与计算机科学背景、医学世家,以及在 Stanford VJ Pond 实验室从事图机器学习。Edunov 研究的是“大量大图”,Feinberg 研究的则是由原子、化学键和空间相互作用构成的“许多小图”——分子。
Feinberg 记得自己在2017–2018年前后曾断言 GANs 是图像生成的未来,随后模式崩溃让 GANs 在蛋白质和蛋白质–配体复合物上失去效果。这个领域只能“等待正确的原语”出现,而扩散模型证明更适合生成3D结构。
令人意外的是,基础扩散研究的重心已不再集中于消费级图像生成。Feinberg 的判断是:“最具创新性的 Diffusion 研究正在我们的领域发生”,使3D结构预测成为10年前无人预料会出现的核心支柱。
2. 药物发现 AI 正在复利式演进,而不是等待一次 iPhone 时刻
Genesis 约7年前成立时,创始人担心自己已经来晚:老牌公司筹集的资本多出几个数量级,投资人也质疑市场是否还需要另一家 AI 药物发现公司。Feinberg 如今回看,认为这种担忧近乎荒谬——“当时以为已经晚了,结果发现仍处于比赛前半段”。
他的原始判断依然成立:约20,000个蛋白质编码基因都可能参与疾病,没有任何单一突破能够解决整个空间。“并不存在一个 iPhone 时刻”;即使是智能手机、ChatGPT 和自动驾驶汽车,也是通过持续迭代逐步变得有用,而不是靠一次干净利落的从零到一。
可解决靶点的范围预计将继续扩大,而“巨大跃迁”会嵌入累积式迭代之中。Feinberg 说,当前系统已经比10年前有用得多,并预计未来10年还会出现一次指数级改善。
主持人的核心挑战是泛化:分子模型过去往往只是模式匹配器,告诉研究人员他们已经知道的事情。Feinberg 认同,这正是 AI 进入物理世界时最紧迫的问题——模型天然会在训练数据附近插值,而药物发现要求可靠外推。
3. 分子设计是连接生物学与临床试验的缺失中间层
Feinberg 的工作类比是:蛋白质像一把锁,药物像一把旨在改变其功能的钥匙。结合只是“必要但不充分”的条件;分子还必须避开脱靶蛋白、抵达正确组织、保持安全,并满足约30项额外的成药性指标。
10年前,研究人员就提出,准确的3D蛋白质–配体复合物能够改善亲和力和效力预测,但几乎无法验证:计算方法预测结合构象的能力很差,而实验晶体学或冷冻电镜可能耗资数万美元,并耗费数月、数年,甚至整篇博士论文的时间。
Genesis 认为,近期突破不只是生成看起来漂亮的结构,更在于证明系统性提升构象准确度能够传导至效力预测。更好的复合物还可能揭示此前被认为“不可成药”的蛋白质中存在可成药构型。
靶点识别、分子设计、监管申报准备、患者分层和临床试验需要不同模型,尽管它们共享部分基础设施。Feinberg 认为设计环节的杠杆最大:患者常被告知医生知道疾病由什么引起,却仍然缺少能够选择性作用于该原因的疗法。
4. Pearl 用稀缺结构数据攻克10^60的搜索空间
Pearl 接收蛋白质序列和配体表征,然后预测二者共同的3D复合物,这是一项共同折叠任务。Genesis 有意聚焦小分子和中等大小分子,包括口服药物、大环分子和肽,而不是把所有生物相互作用都视为同一个问题。
“小”并不代表计算上容易。Edunov 估计,类药小分子约有10^60种,每个分子还存在旋转和多种构象;在主持人提出“在干草堆里找一根针”后,Feinberg 将其反转为“在针堆里找干草”,因为大多数候选物要么无效,要么具有危险性。
分子领域最接近互联网规模预训练的数据源,是 RCSB Protein Data Bank,但其中也只有二十多万条结构;尽管每条结构都包含大量潜在信息。新的实验结构以“冰川般的速度”增加,因为生产它们既昂贵又困难。
Genesis 通过物理模拟小分子来补充这一数据集。蛋白质–蛋白质系统规模大得多、建模成本也更高,但分子动力学及相关计算可以为小分子预训练生成成本更低的结构数据,前提是谨慎处理其中的偏差。
5. Pearl 引入 scaling laws,但没有照搬语言模型
Edunov 将 LLM 配方拆成3个阶段:预训练扩展、通过微调或强化学习进行后训练,以及推理时扩展。Genesis 先创建合成预训练数据,再让模型投入额外推理计算,而不是立即输出一个结构。
这种类比对应的是功能,而不是语言。模型是在“以晶体结构的方式思考”——内部表征被部分物化,并随着 Diffusion head 反复迭代、不断修正预测坐标。
由于扩散过程本身就包含多个去噪步骤,Genesis 可以在生成过程中注入基于物理的引导。主持人将其描述为在学习得到的力场与物理力场之间取得平衡;Edunov 接受这种引导直觉,但不愿声称研究人员确切知道网络内部究竟表征了什么。
Feinberg 将物理先验视为一种经过约束的表征选择,类似于把图像编码为像素网格、把语言编码为 token 序列。Genesis 在输入、架构和输出验证中使用物理信息,同时尽量不强迫模型继承“我们这些渺小人类”持有的每一个假设。
6. 1 Å 将合理图像与有用化学区分开来
讨论指出,在2 Å精度下,结构就像一张细节错误的生成图像;但与肉眼可见的模糊不同,分子错误可能看起来完全可信。整个芳香环或杂环都可能翻转,却仍然通过粗粒度结构指标。
Feinberg 用氢键解释为何目标是1 Å:供体到受体重原子的距离通常落在2.7–3.3 Å之间,窗口只有0.6 Å。距离太短会发生碰撞,太长则会迅速削弱相互作用。“药物发现本质上是一门分辨率科学。”
主持人提出了一个严肃的反方观点:单一构象只是对概率分布的抽象,而亲和力还包含焓、熵和动力学贡献。Feinberg 接受这种抽象,但坚持认为它必须可检查——如果没有结构让科学家判断效力预测是否合理,那么一个标量效力输出“完全可以被视为幻觉”。
Feinberg 区分了高效力配体中高度确定的结合核心与暴露在溶剂中的区域,后者可能确实会“到处晃动”。目标是在蛋白质相互作用发生的区域实现亚埃级准确度,同时接受其他区域存在动态性;这个核心随后可支持自由能预测,并回答实际问题:“下一步该合成什么分子?”
7. 领域的2 Å基准制造了一场评估危机
当被问到 Genesis 如何跨过这一门槛时,Feinberg 给出的答案“极其无聊”:“数据、基础设施和评估。”因为“只有测得出来的东西才能被改进”,优化亚埃级准确度会改变筛选、课程学习、架构、损失函数以及大量相互叠加的小工程决策。
真实合作项目和内部项目让2 Å的失败变得明显,而学术排行榜没有做到这一点。Feinberg 认为,RMSD 低于2 Å源自为发表论文而设计的早期 docking 研究,后来被 AI 继承,并不是药物化学家证明它足以支持前瞻性设计。
他有意用 SWE-bench 作类比:模型得分很高,并不意味着它会成为任何人的首选编程工具。分子评估同样正在从 RMSD 转向物理有效性、PoseBusters 和 lDDT;Feinberg 称,自己的领域正处于“一场正在转型中的评估危机”。
8. 结构预测只是30属性问题中的一根支柱
Feinberg 反驳了 AlphaFold 时代结构预测已经解决药物发现的流行说法。一个静态且分辨率相对较低的结构,无法涵盖动力学、选择性、暴露以及将结合剂转化为药物所需的其他终点。
ADMET 在实际操作中由30多个检测项目构成。Feinberg 提到溶解度、口服生物利用度、不同细胞色素 P450 变体,以及 hERG 抑制;后者作用过强可能造成心脏毒性。
这些终点的可学习程度各不相同。有些终点直接对应与特定蛋白质的相互作用,可能通过3D建模解决;另一些终点则汇总了多条通路,而公开数据集可能“小得可笑”。
Genesis 的广度早于 Pearl:其 Stanford 系谱曾产出 MoleculeNet 和面向制药规模 ADMET 预测的多任务图网络。Feinberg 强调的是连续性——公司一直聚焦药物发现所需的全部模型,包括分子生成,而没有扩展到无关的靶点识别或临床试验问题。
9. 已知生物学同时留下 first-in-class 与 best-in-class 机会
Shawn Wang 的商业挑战是:那些生物学机制已知、结构也具备可操作性的靶点,是否已经被挑剩了。Feinberg 将变量拆开:生物学验证与药物化难度是正交关系,甚至可能呈反相关,因为最有吸引力的疾病靶点往往极难实现选择性结合。
Feinberg 还反对不加区分地引用10%的临床成功率。对于遗传关联紧密、生物学机制清楚、动物模型能够转化、预测药代动力学足够好且安全性强的候选药物,从 Phase 1 到 Phase 3 的获批率“相当高”,但他没有给出一个普适数字。
因此,机会既包括没有任何已知结合剂的真正从零到一项目,也包括改进不完美化学先导物的从一到十项目。他公开举出的例子是 ALK 抑制剂:后续代产品带来了质量更高的生存曲线,说明“我们已经攻克 ALK”并不意味着开发应该停止。
Genesis 聚焦小分子和中等大小分子;Feinberg 表示,仅小分子就约占 FDA 批准药物的65%。因此,目标市场既包括打开新靶点,也包括替代临床或已上市药物中表现欠佳的产品。
10. 合作数据将模型变成前瞻性药物项目
Genesis 披露了与包括 Gilead 在内的公司的合作,以及扩大的 Incyte 合作。一个 Incyte 项目从困难靶点和已有化学先导物开始;Genesis 在合作方数据上微调基础模型,推动项目接近选择开发候选物这一二元里程碑。
另一个端点则始于强生物学关联,却没有任何已知化学先导物——没有专利、论文或配体共晶结构。团队找到初始 hits,随后将其推进为在生化检测和活细胞检测中均有活性的抑制剂。
商业模式将 Genesis 的 AI 专长与药企在生物学、开发和商业化方面的优势结合起来。将 Genesis Therapeutics 更名为 Genesis Molecular AI,体现的是这种身份,而不是放弃药物;内部项目仍然会反复使用这些模型,产生坦率的药物化学家反馈,并让平台经受“实战检验”。
Feinberg 将组织形容为一条“双螺旋”:一侧是深度 AI 研究,另一侧是经验丰富的药物发现人员,其中一些人曾参与研发多款获批药物。目标是让模型直接服务于众多药物开发者,同时保留足够多的内部发现工作,从而真正理解这项工作,而不是变成“只会敲键盘的人”。
11. 湿实验室反馈的价值,恰恰在于自动化仍然混乱
Feinberg 的近期强化学习路径从基于物理的奖励开始,未来可能延伸到实验室闭环 rollout:预测分子、合成分子、测量下游属性,再将结果反馈到训练中。他说,Genesis 已经看到强化学习在其模型上有效的早期迹象。
Insitro 快速生产和测量化合物的能力,支持围绕结构、效力和 ADMET 持续进行设计–合成–测试–分析循环。Feinberg 称,这一合作格外重要,因为它将历史数据与前瞻性数据结合起来,用于联合训练基础模型。
物理工作并不适合简单套用“机器人实验室”的叙事。合成需要匹配的试剂、催化剂、溶剂、温度和实验方案;化合物还需要纯化,并通过 NMR、质谱或相关方法确认,证明试管中确实是研究人员想要的东西。
高通量筛选和 DNA 编码化合物库可以测试数百万甚至数十亿个化合物,但它们与重新合成以及高保真检测之间的相关性,可能“R平方低得令人震惊”。自动化也更偏好受约束的化学空间,而药物发现要在彼此反相关的目标之间寻找新颖的 Pareto 异常值——速度可能要以分子质量为代价。
12. 只有底层模型奏效后,agents 才能扩张药物发现团队
Feinberg 将分子 agents 与编程 agents 相比:两者都会放大正面和负面价值,因此“agents 的有用程度,取决于它们所编排的底层模型”。如果构象系统只能达到1.8–1.9 RMSD,那么它们只会自动化生产药物化学家会拒绝的“slop”。
Sapphire 是 Genesis 为 agentic 平台设定的代号,设想是让“数百人的团队”由药物化学家和 CADD 科学家组成,全天候工作。其 LLM orchestrator 可以调用构象、效力、ADME 和化学工具,同时处理一整套软件栈中没有任何人能够全部掌握的参数选择。
晶体结构可以成为另一种模型模态:agent 可能检查一张图像,调用专业几何工具,或读取原生 token 化的3D表征。这样,它就能基于结构预测进行推理,而不是把每个预测器当作不透明的标量 oracle。
Feinberg 预计会沿着 Cursor 的轨迹发展——先是自动补全,随后逐步走向自主执行;但“我不相信完全自动化或替代人类”。科学家仍然负责战略指挥,agents 则吸收日常执行工作;在他的设想里,药物发现人员会成为拥有数百名计算工人的“大战略家”。
13. OpenBind 暴露泛化能力,GPU 如今限制上行空间
公开基准往往将模型表现压缩在一个狭窄区间,因为每个团队都在围绕这些基准优化。未见过的 EV-A71 3C protease 靶点提供了更干净的测试:Pearl 没有在该靶点上训练或围绕它开发,而该靶点的柔性环必须移动以容纳配体。
Feinberg 表示,Pearl 的数字“高出很多”公开模型,并且在模拟柔性环移动方面“几乎每一个构象都基本正确”。Genesis 将其视为对更大差距的公开确认,而公司称这种差距也存在于保密合作靶点上。
两位嘉宾都将 GPU 视为瓶颈。Edunov 认为,LLM 公司正在消耗本应用于药物发现的算力;Feinberg 理想中的干预是拥有一座巨大的 H100 集群,同时指出 NVIDIA 已两次投资 Genesis,并与其合作优化 kernels 和 Pearl。
Edunov 的算力配置逻辑是,药物在经济周期中始终具有价值,而“纯 LLM 领域剩下的 alpha,已经变得有些可疑”。主持人将主流 LLM 架构与分子扩散进行对比:前者仍与2017年发布的 transformer 高度相关,后者则是一个架构差异更大、且直接关系临床结果的领域。
swyx
I remember very clearly, in 2017 and 2018, talking about GANs, or generative adversarial networks, and how they were clearly the future of image generation. Obviously, they didn't work very well for proteins or protein–ligand systems, and we had to wait for the right primitive to be created. That turned out to be diffusion, which turned out to be a much more useful primitive for the space.
What's kind of cool is that, right now, for people who are interested in really core, fundamental AI research, some of the most innovative diffusion research is happening in our field—in 3D structure prediction. No one would have predicted that then, but now that's a pillar of diffusion, I'd say.
Hi there. Welcome to the Latent Space AI for Science podcast. I'm swyx, CTO of Latent Space. I'm joined by my co-host, Alessio Fanelli, and we're privileged to have Evan Feinberg, founder and CEO of Genesis Therapeutics, and Sergey Edunov, who led Llama 2 and Llama 3 pretraining before he joined Genesis as CTO.
Hi, I'm Sergey. I studied physics at school. This was a long time ago. After graduation, I happened to work in software engineering. I thought I would never need physics again. When machine learning came up and became a thing, it turned out that a lot of the things you do in machine learning were actually very similar to what you would do in physics.
So I jumped on the machine-learning bandwagon and did a lot of AI research as part of the FAIR team at Facebook for quite a while. I later led the Llama team—Llama 2 and Llama 3—and then recently decided to pivot my career again and rediscover my roots in physics a little bit. I joined Genesis as its CTO.
Hey, I'm Evan. I'm the founder and CEO of Genesis Molecular AI, and, like Sergey, I'm also a physics major. I was a bit different from everyone in my family growing up. Almost everyone was a medical professional of some kind, and my sister became an accomplished TV writer, playwright, and novelist. My love was physics and computer science.
My mental model of an adult was, “You should help people and help patients, ideally.” So I was always searching for the right way to do that. After arriving at Stanford and doing my PhD in VJ Pond's lab, we were excited in the mid-2010s by everything amazing going on in machine learning—for images and for language, using a dated term.
Around the same time that Sergey was at FAIR working on a lot of big graphs, I was at Stanford working on many small graphs. Molecules are really networks of atoms and bonds and spatial interactions. If you're at the right place at the right time to bring our backgrounds in physics to bear on improving AI algorithms for looking at molecules, you can make progress. We published a few papers in the area of graph machine learning.
Like Sergey, I thought, “Well, I won't need this physics again, because there's machine learning.” But it turns out the massive amounts of protein simulations we ran on GPUs also came in quite handy as Genesis evolved. We've been really excited over the past few years to figure out how to build foundation models for a totally new domain and make them useful for patients.
Alessio Fanelli
Yeah, that's a really nice lead-in to my first question, which is: you and I have both been in this domain of machine learning for molecules and biology for roughly 10 years. It's kind of like an entire generation of tech bio has come and gone since then.
Yeah.
Alessio Fanelli
While a lot of machine learning for molecules has been, I think, quite effective, one domain where it's been—or has historically really resisted machine-learning modeling—has been the world of protein–small-molecule interactions. With some of the recent advances that Genesis has put out, it seems like you might have actually started to make real improvement on this in a way that we haven't seen for a long time.
Can you talk about what you have done, what Genesis has done, the developments that led to this improvement, and why you think this is actually a real improvement over some of the traditional machine-learning strategies, which were ambiguously helpful?
I totally agree, Alessio. The amazing thing is that when we were founding the company initially as a spinout of the research we were doing in AI at Stanford, we were afraid that we could be too late. You and I first met around that time, right? Seven-plus years ago, as you rightly point out, it was almost a different generation in the area.
At the time, when we were a little seed company, there were already incumbents in the space who had raised orders of magnitude more money than us. There were already some really exciting academic papers out, and we had to constantly answer the question, “Is there room for another company doing AI in drug discovery?” Which is a crazy thing to say, looking backward. Hindsight's 20/20, I suppose.
The claim that we made then is the same one we're making now: there are 20,000 protein-coding genes, and each of them can cause a disease. In the same way that, in other domains of AI, things didn't go from zero to one and then suddenly become solved, the reality has been that there's been no single iPhone moment. Even the phone required iterative development over time to become what it is today.
People love to talk retrospectively about the ChatGPT moment, but the first ChatGPT models that became widely used were vastly less useful than they are now. I think we should look at drug discovery from an AI context in the same way. As we go, we should expect to solve more and more problems. There will be large leaps, there will be iterative improvements, but they'll compound over time.
That's why I think, in the same way that 10 years ago was very early but clearly value was starting to be created for drug development, from an AI perspective, fast-forward 10 years and our technologies—which we'll talk about in this discussion—are vastly more useful than they were then. In the next 10 years, we'll see another exponential improvement as well.
In the same way that your car, in terms of autonomy, can do things today—we're recording this in San Francisco; you look outside and see Waymos—we'd have been blown away a few years ago by those technologies. I think we should look at it as every few months, every year, the technology gets more and more useful for discovering new medicines with artificial intelligence. That will continue to be true for years to come.
Alessio Fanelli
Maybe we can go back in time a little bit and explain what the state of protein–small-molecule drug discovery was, let's say, about a decade ago. What could we actually expect to do with machine-learning models, and what were some of the failure modes that people, I think, didn't really see coming from a practical standpoint back then?
It's funny—just breaking the fourth wall a little bit. I came in not really knowing any of these questions, but Alessio really knows how to speak my love language. Because, like the other people in this room, my background is much more on the physics and computer science side, I've been very, very focused, in a deeply passionate way, on the area of how we can make medicines with AI for over a decade.
I've seen so many tectonic shifts. To give you one concrete example—and I'll give a little background for those who don't have as much biological background—drug discovery is akin to finding a key for a lock, where the lock is usually a protein, or in some cases a nucleic acid, and the drug is the key. It's usually some small molecule, a peptide, an antibody, or some other modality that wants to bind to that lock—that protein, the receptor—and change something about it.
If it's an enzyme, you usually want to stop it from functioning. Sometimes you want to activate that endogenous protein. The goal is to introduce an external molecule, an exogenous molecule, to change something about that biological pathway.
A necessary but not sufficient part of that process is finding a molecule that binds well to that receptor or protein. Of course, that's not sufficient. We also need to make sure it does not bind to certain antitargets, certain proteins you want to avoid. That often mediates toxicity. You need ADME and toxicity, which is typically around 30 properties, to make sure your molecule is safe in addition to being effective and gets to the right tissue. These are all critical, all are necessary, and none are sufficient.
One concrete example to answer your question about protein–ligand interactions is that we had a hypothesis for a long time: if we could predict the 3D structure—the 3D coordinates, which, for the computer-science people, you can imagine as a point cloud—the 3D positions of your drug and the 3D positions of the protein, and if you could model that with high accuracy, that would naturally lead to measuring binding affinity or predicting potency more accurately.
That was a hypothesis that fundamentally could not be tested because models for predicting those 3D poses, the 3D structures of complexes, were so bad.
Alessio Fanelli
Or, if you can predict this, it requires so many computational resources that you might as well just solve the problem.
Right? You might as well solve the 3D crystal structure or cryo-EM structure, is what you're saying, which can cost tens of thousands of dollars. It can take months or years; entire PhD theses can sometimes be devoted to it, and postdocs can sometimes work literally 24/7 trying to solve one 3D structure.
The hypothesis was that if we could make this orders of magnitude faster and sometimes more accurate—which we can talk about with artificial intelligence, or before that with molecular dynamics—we would thereby not only accelerate drug discovery but enable the discovery of medicines for targets that were previously thought to be undruggable.
And that was just a hypothesis. It took until the last few years to show that if we systematically improve the accuracy of those predictions—if we can improve the accuracy of a protein-ligand 3D complex—we can make potency prediction more accurate as well. That's something that I am personally extremely gratified to feel like we've been showing, because that was just a hypothesis when we started a company a few years ago, and now it's become a reality thanks to the convergence of a lot of different ideas that were able to enable Pearl and other foundation models like it.
Shawn Wang
Yeah. Can you talk about Pearl a bit? I think one of the interesting things about Pearl, in my mind, is that it was an attempt to scale what seems unscalable to me: RNA and small molecules. Normally, if you scale them, models just come out as pure pattern matchers. These things love to tell you what you already know, and sometimes they're maybe not so great at telling you something you don't know. It seems like there are hints that Pearl has actually moved beyond this. What were the key insights that went into this?
Evan Fineberg
Sure. And by the way, Brandon, what you describe shouldn't surprise any pure ML practitioners in the audience, just because machine learning is initially constructed for pattern recognition. It's the same when you fire up Claude or ChatGPT or what have you: it's going to do best when it's closest to the training data, and it's going to pattern-match off it. I think that has been the big sense of urgency in AI meeting the physical world: how to extrapolate and how to make generalizable models. That's what we've been working really hard on. Maybe Sergey wants to give some more color on that, maybe to step back a little bit and discuss what Pearl is.
It's a structure-prediction model, which basically means it takes as input a protein sequence and a ligand representation that you try to attach to this protein, and it predicts how this protein and ligand are going to look together as a structure in 3D space.
Shawn Wang
A term for this that people sometimes use is co-folding. It's multiple folding—AlphaFold 3, and then some of the other models, like Boltz and OpenFold 3, have also implemented some of these ideas in their own way.
Yes, there are several models. Some are open source, and some are closed source, that are pursuing a similar direction. Our fundamental difference is that we're focusing on small-molecule space. A lot of models are doing protein-protein interactions, and it works well for small-molecule space. At first glance, it may sound like it's an easier task because, well, it's a small molecule, right? Why would it be challenging? But the reality is the search space is so vast. There are 10^60 drug-like small molecules in the universe, so good luck searching that space. There are also all sorts of variations: you can rotate a molecule, and you can have different conformations of those molecules.
Shawn Wang
Maybe the intuition there is that when you type a Google search, if you just use a few words, you're going to get a huge list of matches. But if you have a very specific query, you're going to have a few matches, and they're more likely to match. Similarly, with a small molecule, if you have a very small query, which is your molecule, then it's likely that you're going to have to search a huge space of possible matches. Whereas if you have something very complex, it's easy to rule out a match. Is that a good way to think about it?
That's one way to think about it. Definitely. For me, it's just the computational complexity of figuring out which small molecule will attach to this protein.
Shawn Wang
It's such a vast problem to solve that it's impossible to do without good models. It's like finding a needle in a haystack.
Mhm.
Shawn Wang
Where everything except your needle is very, very dangerous or just doesn't bind.
Or they are needles.
Shawn Wang
Yeah. Right, right. Finding hay in a needle stack might be a more apt analogy.
Yeah. So now we can go and think about how we train those models, right? A lot of training data that people use is so-called PGB, which is a public database of all historical crystal structures, and it's not that big. It's like 200,000 crystal structures, and it's very hard to expand. It takes a lot of time, energy, and money to create new crystal structures. Although there are some projects pursuing this, it's still expanding at a glacial pace. Expanding this database is pretty hard.
But what we figured out is possible to do, and it's very relevant to our small-molecule space, is that in small-molecule space you can actually model your small molecules with physics. You can model their behavior, and that allows you to create more data. You can train a model on something which is not necessarily possible in protein-protein space.
Shawn Wang
Because they're too complex for a protein.
Yeah, those are very big molecules. It's very hard to model them with physics—not impossible, it's just computationally very, very hard.
Shawn Wang
Right. So I guess what you're saying is you do a really good job of modeling these small molecules using MD or something like that. You can put together those structures in your training set at a much lower cost, so you can use that as your synthetic training set.
Maybe to step back a little bit and talk about our roadmap. It's not fundamentally different from the roadmap for LLM scaling. Remember, in LLMs we have all of the stages: you have pre-training scaling, then you have post-training scaling, where you do either fine-tuning or RL. Now everybody's doing RL. Then you do inference-time scaling. All 3 of those concepts connect, and that's what led us to state-of-the-art LLMs these days.
It's not fundamentally different on our side. We also have pre-training scaling, where we create a lot of synthetic data to train better models. We do that. Then the second thing we started doing is actually the third step in LLM scaling: we started doing inference-time scaling. It's fundamentally very similar. In an LLM, when we talk about inference-time scaling, we're talking about thinking tokens, where the LLM, instead of giving you the answer right away, goes and thinks for a while and then comes up with a response.
We're doing a very similar thing with our models, where a model is forced to think, except it's not thinking in language tokens; it's thinking in terms of crystal structures—not fully materialized crystal structures, but some sort of crystal-structure representation in memory. The model kind of goes back and forth with those. We use physics-based guidance during this process to steer the model output in the right direction. What we found is that it improves model performance by a lot.
Shawn Wang
It's easy to understand how thinking tokens work because they're basically parts of the transcript that you don't see. But I think that, at least for other models, do you have some sort of loop in there? Is that similar to what you're doing—a looped transformer or something like that? I know you mentioned that there's some physics-based verification as part of that in the loop. Is there more to it than that?
One fundamental block of our models is a diffusion-based head. It's like the same diffusion models people are using for image and video generation. We use them for crystal-structure generation. A diffusion head is iterative by nature: it has multiple steps, and you're refining your predicted structure. As you're doing this process, you can steer the model in the right direction.
Shawn Wang
So you have a steering simulation or something in the loop? You go through a pass of diffusion, look at what comes out, run some verifier, and use that to say, “I like this” or “I don't like this,” or give it a vector to move toward. Is that kind of the idea?
Yeah.
Shawn Wang
So your diffusion head is basically learning something like a force field, and you're balancing between a diffusion-based force field and a physics-based force field. Is that a way of thinking about it?
I struggle to understand what those models are really learning underneath, and I think the whole idea of explainability is a big research challenge. You can identify what a neuron in a transformer is potentially thinking about, but it's a really difficult problem. We have some internal representation, but the output comes out as a crystal structure, and it's really hard to tell what exactly is happening inside.
Shawn Wang
Is it possible to do something along the lines of the standard mechanistic-interpretability techniques, like sparse autoencoders and the more complicated things, in this context? Or is there a problem with doing that kind of thing?
I think it would be interesting as a research direction. That's not something we are pursuing. I guess potentially we could explore this.
Shawn Wang
Okay. Yeah.
Brandon
Sergey is a lot more credible than I am talking about the transferability of LLMs to our space. Sergey is being humble, but Sergey led the Llama 2 research team at Meta when he was still there. He's trained one of the most widely used language models on the planet. But I guess what I can add, on the more physical and interpretability side, is I like to look at AI models in terms of the inputs, the models themselves, and the output.
And we find it really improves outcomes, as well as interpretability, when you can include as many physical priors as possible without, of course, biasing the model toward what we puny humans believe.
But I've always had the view that AI fundamentally is representation learning. To me, that's actually not that different from what's happened in the language and vision domains. When we started using convolutional neural networks, we were enforcing a real human prior on images. We're saying, well, pictures are grids of pixels. There's something inherent about that. That's a human construct, and we constructed ConvNets on top of that.
And language as sequences of tokens in 1 dimension: we first built RNNs, then built transformers, but it's still baking in a pretty serious prior on how to look at that data. We don't view it very differently here.
In our case, from the input perspective, the disadvantage that we have in our field over the more traditional domains of AI is that we don't have the internet to work with. We can't just download Reddit posts, buy some subscriptions to The Wall Street Journal, train a model, and voilà—the pretrained model works fairly well. The closest equivalent that Sergey mentioned is the RCSB Protein Data Bank, which has more like a couple hundred thousand crystal structures.
Although there's a lot more—there are a lot more tokens per structure—so there's a lot more latent information than that figure would describe. We have to be clever on the input side, which means generating more pretraining data, as I was mentioning, and using physics as much as possible.
In terms of the model architecture itself, the idea is to have the model relearn as little physics as possible, so it's less likely to overfit. Then, at the output stage, we're enforcing physicality.
I think that also goes to one of our main focuses as a company, which is that we've been focused from the beginning on small- and medium-sized molecule discovery. What does that mean? Small molecules tend to be drugs that you can take as a pill, so they're orally bioavailable. It's how most people think of medicine.
And medium-sized molecules—think about macrocycles, peptides, modalities that break the traditional Rule of 5, but are still growing modalities in medicine. That said, of course, small molecules are still 65% of FDA-approved drugs, so we're talking about the biggest part of the pie.
We're building this with a judicious focus from day 1 on what drug hunters, medicinal chemists, and CADD scientists need to discover drugs faster and better than they could before. We wanted to build in that sort of human interpretability and usability from the beginning, and also physical usability. We want outputs from these models to be used by physical methods that require 3D coordinates to actually make sense.
We use the term force field. We want the outputs of our model to be useful as inputs to a force field. There were a few papers that came out. One was in Cell, I think, last year, which showed that, for all the claims about AlphaFold solving drug discovery, people tried to take AlphaFold-produced protein structures, use them for traditional docking, and found no value in it.
Those pockets just weren't high-resolution enough. They couldn't be used by physical screening methods. And so we wanted to be able to build our systems so they'd be useful to humans, and useful with all the many adjacent and powerful tools from the physics and computational chemistry communities from day 1, for that interoperability. Does that make sense?
Shawn Wang
Yeah. So this might be a good time to just step back for a minute. When you develop a drug, there are many steps to that, from discovery to toxicity to bioavailability. There are names for all these things, which I'll let you state, and then there's the clinical stuff and, finally, eventually, approval, hopefully.
Where do your models sit in that, or which parts of that process do they serve? What are the steps, just briefly? I've read that there are 12 steps—maybe, I don't know, 12 steps in drug discovery. Acceptance is the last one.
Except that your Phase 3 trial failed.
Shawn Wang
Exactly. Yeah. So where do your models help people do their jobs—in the discovery or in the development process?
As you rightly point out, if we go from the very beginning to the very end, first we need to identify what target is causing the disease, what signaling cascade is involved, what's going wrong, and what's causing the phenotype of the patient. That's not a trivial exercise, but there are many—not all, but many—diseases where the cause is actually very clear.
The clearest are the monogenic diseases, but even in others, from various panels, it's clear that this overexpressed, overactive, or deleted gene or protein is causing the disease. So we know many, but not all, of the universe of causes of diseases.
Then, once a target is identified, we have the process of finding a drug for that target that's ideally selective and very potent. It gets to the right tissue. Some diseases are multisystem; some are very specific to a certain organ or tissue. You want to make sure that drug can get to that tissue.
And then there is the process of GLP tox and IND enablement. IND—investigational new drug—you have—
Shawn Wang
GLP in this case is not good—not—
Not good laboratory practices.
Shawn Wang
Thank you. We're not talking about the metabolic space right now.
There are clinical trials, usually of ascending size. Phase 1 starts typically more with safety, then Phase 2 and Phase 3, and, as you point out, ideally approval at whatever regulatory body—the FDA or EMA or what have you.
Our contention is that the highest-leverage application of artificial intelligence is the drug discovery and drug design process. The reasons are that, even though they're all valuable—and I cheer on all the peers, many of whom I know, who are working at all the different parts of the stack—the first thing to make clear is that, just like a vision model is going to be very different from a coding model, they might share some similarities under the hood, but it requires real focus to do any one area correctly.
There's a similar, if not greater, difference between the sorts of models you need for target identification—that is, biology—drug discovery, preparing for regulatory filings, and helping segment patients for clinical trials. These are all very important, complementary, but ultimately distinct problems.
I think about all of the times when a patient goes to a doctor and gets told that they have a very clear diagnosis. In addition, the physician will tell them there have been breakthroughs where we understand why this disease happens. We've sequenced your genome, even. We know—or we've sequenced your tumor, and we know what's causing your condition—but we do not have a selective therapy. There is no precision medicine for your condition.
The hope is that at least they know what's wrong with them, but they don't really know how to treat it. Our aim is to have as many moments as possible where patients are told they have a very specific condition and there's a specific treatment for them.
That's going to require bending the cost curve of discovering and developing new medicine, but also, importantly, solving those cases where there are certain “undruggable” targets—undruggable proteins that we need to figure out how to drug—where those targets have proven resistant to traditional methods or even intractable in some cases.
Shawn Wang
Just to clarify my understanding, there's the discovery part, and then there's the clinical part, where you're answering, “Does this thing actually work?” There can be effort there where, okay, I want to find the target that's impacting this disease the most, or maybe what's the thing that I want to go after in order to solve a medical problem. Then there's the clinic, where you're answering whether this thing actually works. Maybe there's a feedback loop between them, even.
But if you don't have the thing in the middle that's like, “Okay, here's how we actually build this drug so that it can hit my target,” it's not selective enough, so it kills cells or does things to cells that I don't want it to do, then it doesn't matter, right? I can have the answer, “This drug should go after this target,” but if I don't know how to do that effectively and potently, then the answer doesn't matter. Is that kind of what you're saying?
Yeah. Because people love to debate about success rates in clinical trials, but the reality is that if one focuses on those drug candidates that are aimed at proteins that have a close genetic linkage to a disease and/or where the biology is well understood, where the animal models translate well to the disease, and where the molecule is predicted to have good pharmacokinetics and the levels in patients' serum are going to be high enough, those molecules have fairly high FDA approval rates.
The success rate from Phase 1 to the end of Phase 3 is substantially higher than average. People love to cite the 10% success rate, but that's really a lowball, because we often know what genes, proteins, and targets are causing the disease. They're just really hard to drug, or the molecules we put into patients are inferior to the target product profile that they deserve.
In the cases where, if you just look at the preclinical data, you focus on good targets with good biology that are predicted to be distributed well and have good safety profiles, those molecules are very likely to get approved and create tremendous value for patients. And so our view is—I think about the physicians that run trials who are clamoring for those kinds of molecules. I think about the patients who are looking for more selective therapies.
And so our view is that that is the highest-leverage application of AI in health care and medicine more broadly.
Shawn Wang
Just a business question here. What is the space of things where the target and the biology are understood, the structure is understood, and the only thing you really needed to do is find the right lock? It seems like you probably would have already picked over that class of targets in some sense. Those are easy targets, right? So how much opportunity is there for that?
There are 2 orthogonal concepts. Known biology is orthogonal to the ease with which one can drug that target. Sometimes, unfortunately, it seems they anticorrelate, in that often the most appealing targets from a validation perspective seem to be really hard to drug.
Shawn Wang
No, but do they anticorrelate, or is it just that we have picked all the low-hanging fruit that solves both of those? Those are easy targets, right?
For a variety of reasons, many of them have been drugged—not all. But even in those cases, people love to have this false dichotomy of easy versus hard targets: this one is drugged versus it's picked over, therefore it's not. At Genesis, we do most of our work through large pharma. Most of what we do is providing AI services, basically, to major pharma companies like Gilead, and we recently announced the expansion of our collaboration with Incyte, which we're excited about.
I can't give details of the specific targets that they're working on, but what I can say is that there's a real range, from first-in-class chemical matter. What that means is we believe this target causes a disease, but there are no known molecules that bind to that target. We need to find the first binders ever—the true zero-to-one cases. But there's a wide variety of targets we work on that are more, I'd say, 1-to-10 cases.
Sometimes there's an approved agent, or sometimes there are molecules that are only preclinical but aren't optimal. If you can improve upon those preclinical molecules and get them to development candidates, they're ready to get into patients. Or, if you can improve upon the existing clinical or approved agents, you can create a lot of value for patients.
I'll give a public example. We're not working on this target at all, but you can just look at the progression of ALK inhibitors. To make it very concrete, look at charts of patient survival. You might have said, "When the first ALK inhibitor came out, why do we need another? We've drugged ALK. We're done." But it is a qualitative, not just a quantitative, difference when you compare patient-survival curves from first-generation ALK inhibitors to later-generation ALK inhibitors: how much longer those patients live and how many more get benefit from them.
I think there's not only value in first-in-class; there's enormous value for patients in best-in-class, too. I think both are areas that we focus on, if that makes sense.
Shawn Wang
Okay, yeah. So far, we've talked about Genesis modeling in terms of, let's say, this one specific problem of protein-ligand binding. How does your overall workflow work in terms of developing this? I assume—and I do know—that you have other things that you've worked on in solving the broader early-phase drug-discovery problem.
I keep joking that when the Nobel Prize was given for AlphaFold 3, a lot of people thought that drug discovery was solved.
And it's very far from the truth. Very far.
Shawn Wang
Yeah, so maybe you can predict a crystal structure, right?
At low resolution.
Shawn Wang
At low resolution. Well, we're doing better, I think.
Yes.
Shawn Wang
Assume you can predict a crystal structure. Does it mean drug discovery is solved? No, obviously not.
The broader community—the protein-folding community and the drug-discovery community—immediately said there are other things you care about, like dynamics. Just a single static structure isn't enough to understand what's going on with interactions. There are lots of things there in addition to just having a static structure. I think many people thought those were very useful; they've really accelerated science, but it was very clearly a starting point and not an end condition.
Shawn Wang
Right, but if you look outside of people working in the field, in the general popular community, it was a pretty widespread belief that the problem was solved. I just have to clarify that; otherwise, there's an entire ML structural-biology community that will just jump on us and say, "No, no, it's not a solved problem." Got to throw those caveats in.
Absolutely not a solved problem.
Shawn Wang
Yeah.
In reality, you need to predict so many other properties, like ADMET properties that I haven't mentioned. Basically, you're designing a key for a lock—that small molecule that sticks to your protein in the body. But you don't want this small molecule to stick to everything in your body, because that's probably going to cause issues.
You also want to make sure that the small molecule can be soluble, so that you can ingest it as a pill. Those are important properties, too. You want to make sure it doesn't have any other side effects. Predicting all of those properties is just as important as, or maybe even more important than, predicting the crystal structure itself.
And of course, as a company, we're very proud of it. We're not just building models for predicting crystal structure; we're building models for predicting all of those properties and basically enabling drug hunters to become 100× more effective in their daily job.
Shawn Wang
I'm curious about the pipeline that you have. You mentioned you have different pharma partners, GSK or whoever you're working with. Do they have targets or drugs that were designed by you that are in clinical trials or have been approved? Where do we stand with all that?
We're limited in what we can say about pharma companies. They're famously secretive, which makes sense. What I can say, which was recently publicly disclosed as an example, is that we just expanded our partnership with Incyte, and that started a little over 1 year ago with the initial work together. That spanned the 2 sort of bookends of the drug-discovery process in some way.
One of the programs we worked on together was a case with a very challenging target where chemical matter existed, but there was a key binary event, which is getting to a DC, or development candidate. That is the molecule, or ideally a set of molecules, all of which are possible to be the agent for a Phase 1 clinical trial.
We had to do some work together where we take our foundational models, our base models, fine-tune them on Incyte's data, and use them in close collaboration to work together to get to a DC faster. We're getting substantially closer in that case, which is one of the reasons we're excited to expand our work together.
On the flip side of that, one of the other areas we worked together was on a protein with a very nice linkage to a very severe disease, where there was no known chemical matter—no patents, no papers, and, in this case, no co-crystal structure of a ligand, another synonym for a small molecule, that bound to that protein.
So we had to find the first-ever known chemical matter to bind to it and then progress those—what are called hits, the first molecules that bind to your protein—into inhibitors that are active in biochemical assays, which are more enzyme-based, and cellular assays, which means that in an actual living-cell model, your molecule is active as well. It was really based on a concrete set of accomplishments together that we want to expand the collaboration.
Unfortunately, outside of that, there's really very little we can say. The objective of Genesis is to create medicines that patients wish they had. The way we'll be able to do that is by working with as many pharma and biotech partners as possible, for whom their comparative advantage is discovery, clinical development, and commercialization, while our comparative advantage, of course, is AI.
We can put those 2 areas of expertise together in a synergistic, not just additive, way and make medicines together that otherwise would not be possible. That's the name of the game that we're in. This is what we're in for: those clinical outcomes that you're pointing out, and I'm very excited to be able to share those as they arise in the coming years.
Shawn Wang
One thing that you guys talk about on your website is this 1-angstrom threshold, at which a protein-structure prediction and the binding between that structure and a small molecule become useful. Can you talk a little bit about why that's the case? How did you do it when others have failed? You've spoken a little bit about that with the biases in the model and the synthetic data. Maybe there's more. How does that impact downstream, the actual things we're talking about here with ADMET and all that stuff?
When we're talking about molecular interaction, the 2-angstrom scale that people typically measure is just too big. You can think about it like an image-generation model: you would generate a picture and it's fuzzy.
Shawn Wang
So, yeah, it's like in general, right? But you can't discern the details, and the details really matter here. With 2-angstrom accuracy, your entire aromatic ring can be flipped and it will still be a valid output.
The worst part is that, unlike an image that's blurry, you don't even know it's blurry, right? You flip around an aromatic ring, and it looks just fine.
Alessio Fanelli
Yeah, or even worse, a heterocyclic aromatic ring, and then you're really in trouble.
Yeah.
Shawn Wang
And it really matters in this case because individual atoms need to establish connections here. So that level of accuracy—what we're pushing for, 1 angstrom—is really important. So maybe it's almost like with a lot of generative image models: they will fill in a detail that just doesn't exist. So if you were to try to do forensics and tell your vision transformer to enhance an image, and suddenly it just pops up a face, but that's actually not the person's face, right?
But you don't know that, and it feels just fine. And so, if you're trying to use this as a structural hypothesis as a med chemist, you're kind of screwed. If you're looking at the wrong face as an investigator, it's not going to take you anywhere.
Yeah, I really like your framing. And then, as Evan said, downstream, you run a bunch of other models, like physics-based models, and all of those little issues compound. Then, obviously, if you made the wrong prediction—fundamentally wrong—the downstream predictions are also going to be wrong.
I just wanted to say, for your question, I really appreciate your mid-2000s TV reference about enhancing forensic images. I just want to say I appreciate that.
Shawn Wang
Before we get back into details here, we talked about poses. Maybe the contrarian take is that a pose isn't even a well-defined concept, and that the best way of thinking about these things is as a probability distribution over things. There's a small molecule that probably lives in a binding pocket, something like this.
I mean, is this pose an abstraction that humans use and not really, in some sense, ground truth? So I'm actually almost kind of interested to hear that we have this warning threshold, and that this is really important for med chemists. I would also talk to some chemists who would say, I mean, this pose isn't even real, versus just maybe a most probable configuration.
How do you think about that abstraction in general? Do you explore conformational space and provide that as tools? How do you interact with just a single structure versus an ensemble? Sorry, I don't know if that question even makes sense.
Yes, it's an abstraction, but it's a very useful abstraction. It helps us to build up confidence that a particular model output is actually valid. You're not just straight-up hallucinating something, because, yes, ultimately what matters is binding affinity or potency, and you can straight-up predict that with your model and skip the entire pose-generation step.
But then you only have a single number, and that number might as well be completely hallucinated. You have no means to validate whether that number even makes any sense. So, as much as poses are not perfect, they're still a very useful tool for the entire process.
Shawn Wang
Sorry, going into that correction. I'm afraid I'm jumping into a technical rabbit hole here, but just because you have a pose, there are also entropic and enthalpic contributions. Predicting binding affinity is not just about getting the energy right. Is this even likely to make it into the binding pocket, and is it likely to live there long-term?
So it's much more than just affinity or potency prediction. It's much more than just, is this the right pose? Does this have the broader properties it needs to be a long-lived molecule in this state? How do you deal with those things, too?
A note for the wider AI audience that probably resonates with practitioners as well as users: obviously, something that's blossomed in the past 6 months is agents. And we love agents.
Alessio Fanelli
Who doesn't?
There are a few necessary but not sufficient conditions. I'd say, in response to your really deep question there, we all remember what agents were like, let's say, in the middle of last year. Let's just say there is positive value and there's negative value, and both can be amplified by agents. And why is that? Well, agents are only as useful as the underlying models that they're orchestrating.
Let's think about coding. If your coding model even makes subtle but real bugs, your agents are just going to amplify those issues, and you're going to end up with not only slop but something that may be anti-useful, that might give the user incorrect information. We all remember what that era was like in the middle of last year, which made a lot of people lose patience, I would say, with claims about LLMs for agent engineering.
Something changed. Clearly, a threshold was met. Even though these models are still not perfect, the utility of LLMs for software engineering is so obvious now. Obviously, it's been a huge tailwind and very useful for us, for replacing a lot of the drudgery of coding and getting it focused a little bit more on some of the more strategic issues that matter.
You could draw a direct analogy from that to what we're talking about here. It will be no surprise that we're working on an agentic platform for 24/7 drug discovery. You can imagine just fleets of hundreds of med chemists and CADD scientists working nights and weekends all the time for your drug targets. The codename for that agent is Sapphire.
The prerequisite for that was that we needed the underlying models for pose, 3D complex prediction, potency, and ADME to all be good enough for an agent using these models 24/7 to create molecules that medicinal chemists would actually want to make and not laugh at. If your model is sitting at 1.8–1.9 RMSD, that's slop, most likely. Let me be really direct about it: what you're hearing is a bimodal distribution from people about the utility of a 3D pose and how real it is.
The reality is that for a highly potent ligand, almost certainly there is a large portion of that molecule with a very well-defined 3D position, down to even half an angstrom. If you don't believe me, you can open up the electron density diagram in the PDB. It's all available online, and you can literally see, in some cases, aromatic rings with missing density in the middle.
So it's not just a construct on a blackboard that your organic chemistry professor showed you. You can literally see a donut, a torus of electron density for an aromatic ring. However, there are going to be solvent-exposed areas that are often less important for binding affinity but maybe are important for solubility or other properties of your molecule. Those sometimes are less defined because, in reality, to your point, they're more dynamic. They're flopping around in solvent.
But the critical piece, as an upstream indicator that's valuable for predicting the free energy of binding, is to get the core of your molecule that specifically interacts with the protein correct to sub-angstrom resolution. And why does this matter? Just to use it intuitively, everyone in the audience has probably heard about a hydrogen bond. They're among the most critical forms of noncovalent interactions. It's how nucleic acids are held together, how most ligands and proteins interact, and how proteins form secondary structures.
Hydrogen bonds have a very specific angle and distance. The distance is from the donor to the acceptor heavy atom. It's 2.7 angstroms to 3.3 angstroms. If I do my math right, that's a 0.6-angstrom gap. Outside of that, it's not a hydrogen bond. If it's less than that, it's a clash. If it's more than that, the interaction is much, much weaker pretty quickly. An angstrom, for those who don't know, is one-tenth of a nanometer.
Drug discovery really is a science of resolution. If your accuracy is not sufficient, it will therefore not be useful for the downstream things that you care about, which are both potency prediction and prospective design: what molecule do I make next? That's why our view is that the history of innovation in startups is that the ones that do best focus on one well-defined but very important problem. We think our ability to get higher-resolution predictions stems from our judicious focus on small- to medium-sized molecule design rather than boiling the ocean.
Shawn Wang
Yeah. So that's the what and the why. What about the how? How did you get to 1 angstrom and sub-1-angstrom?
I'm going to give you an extremely boring answer.
Shawn Wang
Okay.
Which I think is actually true for the entire AI field: three things matter in AI. It's data, infrastructure, and evals.
Shawn Wang
Okay.
Right? So you can only improve what you measure, and once you are very careful about measuring what matters and you have really talented people on the team, we're going to figure out how to hill-climb that measure, right? From the start, we actually focused on sub-1-angstrom precision, and that led to a bunch of small decisions in the process that compound.
If your team never looks into this metric at all, then you will never train a model that's good at it. And if you're constantly looking into that, then you're going to achieve that.
Shawn Wang
So it's the right objective plus good science—
And it propagates everywhere, right, through the whole stack. It propagates to how you look at the data—maybe, how do we filter out data? Some data is noisier, so maybe we don't need to see it, or maybe you don't need to see it later in the training.
It propagates through your modeling architecture and propagates through your loss.
Shawn Wang
How long have you been working toward this specific goal? I’m curious because this isn’t something you broadly hear the community talk about wanting. People do say that 2 Å is sort of the canonical benchmark, and I’ve never heard someone say 1 Å is the cutoff until reading things coming out of Genesis. How long have you decided this is the number we need to hill-climb on? How direct has this focus been in the evolution of the company? I’m just wondering: how did this come about?
I think one of our most important secrets is that we’re working on real drug programs, either with partners or in-house. When you actually work on real drug programs, you see the failure modes, and you see what works and what doesn’t work. It’s pretty obvious, when you look at the outputs of real programs, what kinds of failure modes are happening. It becomes very obvious that 2 Å is just not working out for those setups.
Shawn Wang
Right? But I mean, there are a lot of really smart medicinal chemists who also think very carefully about benchmarks, people I respect very much, and who have also prosecuted actual drug discovery programs. I’m just curious: why hasn’t this become part of the community? Is it literally just that the community has never been able to succeed at a winning benchmark, and settled on 2 Å as something we can aspire to? I don’t know.
To your point, it sounds like, in my experience with pharma, there are plenty of problems for which there are known things among the technical experts in a subdomain, but that information doesn’t get out of pharma or doesn’t get the attention it deserves. Some of the information is proprietary and gets passed from company to company, but never really released to the public. Is this your estimation of what’s going on?
Have you ever heard of SWE-bench? Gemini does pretty well in SWE-bench. Sometimes Gemini publishes models that win on some of those software benchmarks. Raise your hand if you’re using Gemini to write code right now instead of the obvious competitors.
Shawn Wang
Yeah, not me.
No one. Why would you do that? It’s obviously worse in practice.
I’d say, actually, that in this case, if you look at the provenance of how that happened, RMSD less than 2 Å came originally from docking studies, long before AI models for pose prediction. When physics-based docking studies came from academic institutions—because usually the large proprietary software makers didn’t want to benchmark their methods against other methods—academics had to get licenses and try them. Then they introduced RMSD less than 2 Å, which isn’t surprising. They’re academics; they’re not using these things to make drugs. They’re using them to write papers.
The first big innovation there was PoseBusters, put out by a lab at the University of Oxford, which pointed out that RMSD itself is insufficient and that we need to look at physical validity as well. I’m talking about PoseBusters as the metric, not the benchmark. In the latest release from OpenBind, we just published our benchmarks at Genesis on the OpenBind set.
Shawn Wang
We’ll talk about that in a bit.
If you look at their original publication, their default metrics are RMSD, PoseBusters validity, and lDDT. There’s clearly an acknowledgment that RMSD less than 2 Å is insufficient, and the field is now rapidly evolving to acknowledge that.
Shawn Wang
Do you think the academic literature is going to establish some benchmark—maybe it’s 1 Å, maybe it’s some lDDT or another metric—that will converge and then hopefully, as a community, drive forward these more accurate modeling strategies?
I would say there was an eval crisis in our field that is now in transition. Our field was previously a lot quieter, but if you go to NeurIPS and ICML, year after year the workshops get a lot more crowded. Those evals are in transition now that we realize the flaws of what came before.
Shawn Wang
That was a really interesting point about picking the right evals and then just doing good machine learning like you normally would. But I picked up on something Evan said earlier, which is that you started out with, I guess, what we call PotentialNet. I don’t think we defined that earlier, but it’s this graph-based network, and then you did a lot of computational simulations after that.
Now you’re sort of going back in, once generative modeling took off. I’m curious about how that evolution worked. How did you find computational techniques to be really crucial to building on the ML for a while? Did that computation lead back into generative modeling, or did you start collecting data to build a generative model? Is computational data a way of data augmentation for a generative model? What’s the history of that, and what led to those decisions, to the extent you can talk about it?
I will say that the last line of PyTorch I’ve written is much further back in history than the most recent line of PyTorch that Sergey has committed. I want to make sure that he gives his opinion as well.
I want to make sure that I directly answer your question. I want to make sure that I address something you said a little bit ago that I think kind of got lost, which is that you asked about how we do other things, including ADME prediction. I can spend all day, and I’m happy to, diving into details about Pearl and its evolution, and we will. I just want to give a shout-out to the fact that it is 1 important pillar, but not the only pillar that matters in the trajectory of a drug discovery campaign.
You mentioned that PotentialNet paper, which we’re happy to see had become influential. You and I were working in the space when it was very much in the future, let’s say, but now it is the future, so it’s great. Another paper we published around the same time was on neural networks for ADME prediction, and we published 2, actually, on ADMET.
Shawn Wang
These are all the properties you talked about before that you just have to get right in order to make a drug successful.
Correct. I would say there are over 30 or so assays, each of which you can imagine, if you’re a neural net person, as a multitask neural network or multi-head model. It’s got to predict over 30—three dozen-ish—properties, each of which, if it’s in the wrong range, means your molecule isn’t a drug; it’s just a tool.
These are things like solubility, which Sergey mentioned is important for formulation: can it be made into a pill that you can take orally? There’s oral bioavailability, whether or not you’re inhibiting certain enzymes called the cytochrome P450s and their different variants, and the hERG channel, which, if you inhibit it too much, can cause cardiotoxicity. So these are an alphabet soup of things that most people here haven’t heard of.
A lot of these are extremely hard specifically because it’s not a single causal effect. There are often many processes and many pathways involved in defining a single endpoint.
Shawn Wang
And datasets are often comically small.
Unless you have pharma to help you out. At least in open source, it’s a hard problem. It’s sparse in the public domain. I’d say there’s actually a range, from really directly predictable properties to ones that, as you point out, are actually amalgams of other signaling events. Whether or not you’re inhibiting CYP3A4, for example, is really specific. It’s a certain protein that you’re inhibiting.
Shawn Wang
So that’s something that Pearl, for example, could actually model and where you’d expect to see some performance—a useful prediction.
We’ve been doing 3D work on ADME prediction before anyone else was. There was a bunch of work that ended up launching the company. Back in the day, when AI could still be in peer-reviewed journals and not just random white papers on arXiv, we published MoleculeNet, which has been cited a few thousand times at this point.
We also published this paper showing that multitask graph neural networks were the best at the time for doing ADMET prediction on large pharma datasets, and that paper is now one of the most cited papers on AI for ADME ever. A lot of papers, as you know, are like—you write them and kind of move on—but both of those works have become quite influential as well.
Our history from the beginning has been to work on not just 1 problem, but to focus on drug discovery. In doing so, we’ve been able to have the bandwidth to focus on building all of the ML models that are needed for drug discovery and not get distracted by the biological discovery processes that are needed—the target identification side or the clinical trial side—but really focus on all those tasks that are required for drug discovery, as well as molecular generation.
These are all important. I just wanted to make the point at the beginning, but as you rightly point out, Brandon, Pearl is a 3D structure prediction model, and many of these ADMET properties can be posed that way.
So it is therefore useful for not just on-target potency prediction. All of them matter, and we've had to tackle all of them.
Shawn Wang
Do you use Pearl for all these tests, or is it focused on some of them?
We'll be sharing some things publicly in the coming period. But I think we've had enough news of results lately. Every time you publish something, it's a lot of work for the team to put together. So we'll do that in the coming period, but the most recent thing we published was obviously the OpenBind results, which was a standard 3D prediction task.
Shawn Wang
So, one thing I want to ask about—I keep hitting at this. You talked about there's a lot of synthetic data, and then there's ML modeling, which is very generative, modern generative-model flavor. I want to talk about the feedback between how your computation, your ML, and wet-lab data went into developing this model.
We talked about priors. I think maybe one thing that's typically been hard is scaling protein-ligand models well in a way that meaningfully generalizes. This recent Pearl paper and some of your other recent results, which we'll talk about shortly, have shown true generalization. What I'm curious about is how these different aspects feed together to give you this power.
Specifically, just saying computation is interesting, but you have to be very careful about computation. MD has all sorts of biases; if you're not right, it can give you poor results. I'm curious about how all of this worked together, and especially in the history of the development of the company, how did you get to this point?
There's the classic concept of the narrative fallacy, where in retrospect everything seems obvious and you can draw really linear processes of how we got from point A to point Z. But the reality is always more interesting, much messier, and much more nonlinear, I think.
For Sergey and me, who've been training or taming neural nets for over a decade, we remember a lot of different eras in the space. We would have been so excited if we could tell ourselves the capabilities of our technologies 10 years ago. If we told ourselves how we got there, it might have seemed somewhat obvious, but still there were some things that would have been really difficult to foresee.
One example is that the concept of using generative AI for the molecular space is not new, but the reduction to practice would have been very difficult a decade ago. As one concrete example of that, I used to co-run the Stanford AI Salon. We had fun organizing it at Stanford, in the Gates Computer Science Building, and we'd get some interesting speakers and about 25 people having wine and cheese. They're all now some very straight-up famous people, right? But back then it was a much smaller field.
I remember very clearly, in 2017 or 2018, talking about GANs and how generative adversarial networks were clearly the future of image generation, obviously. There was a lot of work then on whether we could just apply GANs to produce conformations of proteins, or try to use them to produce protein-ligand poses. For all the same reasons that those models were really tricky to train for images—mode collapse was the most famous problem—they didn't work very well for proteins or protein-ligand systems.
We had to wait for the right primitive to be created, and that turned out to be diffusion, which was a much more useful primitive for the space. Interestingly, a lot of image and video models—some are using diffusion, but some of those have actually gone autoregressive. What's kind of cool is that right now, for people interested in really core fundamental AI research, some of the most innovative diffusion research is happening in our field, in 3D structure prediction. No one would have predicted that then, but now that's kind of a pillar of diffusion, I'd say. Everything I just said was unpredictable 10 years ago.
In parallel to that, long before the advent of 3D diffusion models for generative tasks in chemistry, we were building a variety of tools for the problems of drug discovery at hand. Some of them were using physics-based methods for predicting potency or even for helping predict certain ADMET properties. The same was true with molecular generation—that is, using different techniques to generate new molecular ideas.
As Sergey pointed out, there's 10^60 drug-like molecules. Searching that space efficiently is hard. So we've been working on that problem, and those things happened to be available when we wanted to take what was then the very nascent area of co-folding, which was clearly exciting but not at all ready for prime time—not at all ready for what a medicinal chemist would want to use in their day-to-day work—and take it to the realm of useful and, in some cases, irreplaceable, clearly superior to a non-co-folding method.
It just happened that we'd been building those other primitives, and so they were available to us. We could put them together to build out, for example, the synthetic data pipeline or to use some of the inference-time techniques. We were ready to do those rapidly because we'd been working on orthogonal techniques and ways of approaching problems.
There was something here that was serendipitous, but almost any discovery—like the discovery of penicillin—has this element. I'm not saying what we're doing is on the same order of magnitude or benefit to humanity as penicillin, but there's always this element of, if you're laser-focused on a problem for long enough and spend enough time in the lab—or, in our case, on the computer, banging your head against these problems—you can create the luck, as it were, that enables some of the developments.
Shawn Wang
How does the lab interact with the development process?
As I mentioned before, we train other models, not just structure-prediction models, and for those specific models, lab outputs are extremely useful right away. Potency prediction, right? Those outputs can be used directly for model training. But what I'm most excited about going forward is reinforcement learning. I think that's coming in our field, and we've seen early signs of it working already with our models.
You basically put, initially, maybe physics-based feedback into your models to improve them through RL—the typical RL loop. But eventually, you can go all the way down to lab-in-the-loop setups, where your model is not only producing predictions, but you synthesize based on those predictions, measure downstream properties, and then feed those results back into the model.
Shawn Wang
So how do you get enough volume? I mean, do you have automation to do that, or do you do large campaigns that aren't necessarily automated but generate a lot of data?
One thing I'm super proud of is our partnership with Insitro. They're such an amazing company. They're extremely good at producing data—basically taking the compounds, generating them, creating them, measuring the downstream properties, and then sending those results back.
This is such a match made in heaven for us between Genesis and Insitro, where we're able to train models, give predictions, and have results back from Insitro super quickly.
Shawn Wang
So your rollouts include a lab iteration, essentially. Yeah, that's amazing.
Yeah. And diffusion is slow, so it's about the same speed as that sometimes. Give it a week.
It is true. I'm a big believer that companies are typically really good at 1 or 2 things and do best when they can focus on them. In the same way, we've been really doggedly focused on developing the best AI models for drug discovery, and Insitro has that level of maniacal focus on optimizing drug discovery and development.
Another trend people like to talk about is, of course, the rise of China in biotech. It's the elephant in the room; there's no point avoiding it. It's so in the zeitgeist right now, and so many Western companies have become more reliant on CROs to do a lot of their wet-lab work. Insitro became so state-of-the-art in terms of the experimental capabilities they had in-house. Their productivity is extremely high, and as Sergey said, it's a match made in heaven for us.
What we thrive on is continuous learning of the models. We want to have design, make, test, analyze cycles that are as rapid as possible and continuously fine-tune, and in some cases retrain, the models based on what we see in the lab. That partnership is one of, if not the first ever, that enables that sort of true joint foundation-model training on historical and also prospective data.
As ML practitioners, we all know that data is such a critical input—the whole critical ingredient. It's just extremely exciting for us, and that's a range reflecting our conversation so far here, from structure and potency to a variety of ADMET properties as well. I think it will immediately improve the strength of our models and be really powerful for generally accelerating drug discovery.
Shawn Wang
We were joking around about the time that it takes to do diffusion, but I don't know if you can disclose this: how long does it take? You email them with, "Here's some—I don't know how many—10, 100 compounds," and then they get back to you with measurements of those synthesized compounds in what, a day, a week, a month, an hour?
It depends. Some compounds are really easy to knock out, such as a very well-characterized reaction where the usual conditions just work. Sometimes it’s, “Oh, this coupling actually didn’t work like the literature said it would,” and we have to try different conditions. So it depends.
I’m saying this so the audience understands that, for all the claims out there about robotic labs automating synthesis, the reality is a lot more complicated than that.
Shawn Wang
Sorry. What are some of the complications there? They all do some sort of automation. Again, I’m not trying to get you into some fight with them, but what was the contrary take here? What goes wrong with an automated lab?
Okay. So, this is also one of the reasons, as Sergey mentioned, we’re really excited about reinforcement learning, because it circumvents the problems of having to do very fast design–make–test–analyze cycles from the lab to feed back into the model. Ideally, we can throw more and more GPUs at the problem and have the model self-train in the same way that we once did for board games in the late 2010s or, most recently, for coding. That’s a key north star for us.
However, before anyone tells you that there will be a cadre of geniuses in a data room just solving drug discovery, there are some problems in our space that don’t lend themselves to RL as naturally and that will require experimental intervention.
To drive into the specifics and give some examples: coding was always my north star, what I wanted to do with my life. But I spent a fair bit of time myself in the wet lab as a very mediocre experimental scientist before I found my true joy back when I got to graduate school and could just focus on coding all day.
The reality is quite messy. In order to make a new molecule, you have to synthesize it. What that means is bringing different chemical reagents and catalysts, temperatures, and solvents together in the right way and in the right protocol to try to make your new molecule.
Once you do that, you need to purify it. If you don’t do that, your positive or negative assay read could be an artifact. You need to characterize it, so you need to use NMR and other techniques like mass spectrometry to indicate that what you made in your vial is what you actually sought to make in the first place, which is not as trivial as it sounds either.
That’s not only true for us, by the way. For anyone doing materials science or AI, it’s the same sort of challenge: Did I make what I actually made? That’s not a trivial question at all. It’s easier in small-molecule drug discovery than it is in materials, but it’s still nontrivial and takes time.
We’ve seen so many cases where screening large libraries sounds great on paper because you can test millions or billions of compounds in one go, so you have more shots on goal. As a bonus, you have data for your model, right? You just got millions of data points.
The reality is that the translation of high-throughput screens, whether it’s DELs—DNA-encoded libraries—or more traditional screens, the R² of those predictions compared to the actual business of resynthesizing a molecule de novo and doing a low-throughput, high-fidelity experiment is shockingly low. That false-positive rate is enormous for a variety of reasons.
That’s one of the many reasons why people are so excited about AI making those predictions, because in theory they can be even cleaner than a lot of the wet-lab work, especially the high-throughput work, and yield higher enrichment and higher true-positive rates.
Shawn Wang
So why can it be more? Is it just because the experiments go wrong a lot? Even if you have some sort of very precise robot doing it, there are still too many variables. There’s too much slop in the system.
It’s a few things. One is that drug discovery frequently involves finding the outliers.
Shawn Wang
Yeah.
One of the many aspects that’s so wild about our space is that the properties one is seeking in a molecule very often anticorrelate. Binding tends to improve the more hydrophobic—the more greasy—your compound is, because protein-binding pockets are greasy. Guess what? That makes your solubility worse.
You want to improve your solubility, often by adding polarity to your compound. Suddenly, for those who remember high school biology, cells are surrounded by lipid bilayers. Your molecule might not get through the cell membrane because you made your molecule too polar.
The properties you want often anticorrelate, and that’s where multiparameter optimization of molecules ends up feeling like playing whack-a-mole. You’re often searching for real outliers, and that often requires making molecules that are fundamentally novel.
The kinds of chemistries that can be automated today are fairly constrained. Your ability to search chemical space broadly for those really top Pareto-optimal compounds is actually very limited. The benefits you get in speed have a very harsh trade-off with the actual quality and novelty of the molecules that you can make.
Shawn Wang
So you can do a lot of boring stuff very quickly. Is that—
If you’re lucky, yeah. That would be—yeah, if you’re lucky.
Shawn Wang
But to get to the answer to the cutting-edge questions, it’s hard to get a robotic system to do those kinds of things. Not where we’re at today, right? And so, I mean, I know—again, I’m not asking you to disclose anything that you can’t—but what is Insilico’s approach that makes them so fast and effective?
I think if one is willing to focus on one problem and say no to distractions, it’s amazing what can happen. One, their talent density is very high, and they’re willing to do what is not universally done now in large- or medium-sized pharma, which is build so many experimental capabilities in-house.
Shawn Wang
Mhm. And I actually think that it’s possible for more companies to try to do that, given the global dynamic of pharmaceutical development that’s changed so dramatically.
Right. CROs are taking over everything. Now suddenly it has become commoditized, and your capability to do the cutting-edge experiments goes down. It can.
It’s been a debate for hundreds of years about the benefits of vertical integration. It comes with all the risks and challenges and upfront costs, but it has all the benefits of controlling the process end to end.
Shawn Wang
When it works, you can do amazing things when you do things end to end.
Yeah. Of course, it comes with greater challenges, and I think an appetite is needed to do it.
Shawn Wang
This kind of brings me to another question I have, again more industry-focused. Maybe now is a good time to talk about your pivot. There’s been a name change from Genesis Therapeutics to Genesis Molecular AI, and I think that goes hand in hand with perhaps a pivot in the strategy of the company.
Is it that you’re going to focus on the AI portion and then sell that to pharma, rather than be a sort of drug company that also sells the tool? Is that sort of accurate?
When we started the company, the origin of the company was the fundamental deep-learning research methods that we had developed at Stanford.
Shawn Wang
Right?
The objective we were trying to solve for was: How can we have the greatest impact for patients? At the time, the founding of the company was purely AI research. There was no real precedent at that time for a pure model company really making it in the realm of just creating value for partners in pharma, or really most legacy industries.
The other component was the company being founded on AI research. I think we really wanted to show people that we were serious about the biotech domain. So, I think that’s part of the origin of, let’s call it, Genesis Therapeutics.
Shawn Wang
Yeah.
We seemed a lot more serious to all the tried-and-true folks in pharma, the dyed-in-the-wool medicinal chemists and drug developers who we really wanted to work with. I think there was an element of an attempt at nominative determinism there, a bit, but also an acknowledgment that I don’t think anyone had really shown at that time that you could be a pure AI model company and really make it in the space.
Shawn Wang
And really make it in the space.
That was obviously 2019. So we had the benefit of being real forebears, pioneers in the space. We had the downside of being pretty early.
Shawn Wang
Yeah. We thought we were late, but it turns out it was still early innings.
As the company grew and evolved, we really were structured kind of like a double helix. We had extremely strong, focused AI research, and we were fortunate to be able to recruit some of the most experienced and accomplished drug hunters—people on our management team and board who have discovered, co-invented, and developed many FDA-approved drugs.
You talk to most medicinal chemists, and most med chemists feel very fortunate if they’re able to work on 1 molecule that gets FDA approval in their whole career. We’re really fortunate to have convinced these accomplished drug hunters to join our management team and board at the beginning and steer the show.
The outcome has always been the same: Everything we’re talking about only matters if we end up helping patients and their families at the end of the day.
And we have felt from the beginning that the way to do that at scale is to primarily partner with large and medium-sized pharma companies and biotech companies. They can do what they're best at, which is novel biology and clinical development, while we can do what we're best at, which is AI, and bring those 2 together.
In addition to that, as Sergey has said, there are numerous advantages to having in-house wet-lab data generation: clearly dogfooding the models and getting really direct human feedback from real med chemists about what we need to build, in an extremely candid way. And, of course, the third thing is that the single most valuable thing in the universe—which I like to joke is the mere result of an assay test—if you ask people what is the 1 thing they most want, and they really thought about it, most people would say a medicine for themselves or a family member. So there is just immense intrinsic value to having one's own pipeline programs.
When we interact with pharma partners, on average, they love that we're working on molecules ourselves because it means that we're not just a bunch of keyboard jockeys, but we live the dream that we're pitching, and we need to use these models in practice for ourselves, too. So they're battle-tested in an additional way, and we're sort of domesticated in that we know what's really difficult about the space. We're not just selling some pipe dream.
Shawn Wang
It seems to me like there's been a shift in the last 1–2 years where suddenly pharma has become interested in buying tools off the shelf, whereas previously there were a lot of companies that started trying to build AI and then had trouble, so they became pharma companies because that's the only way they could make money, right? You raise a bunch of money and hope that in 10 years you have a drug, and if that drug is successful, then you've made it; otherwise, you fold or sell off the assets or whatever.
It seems like that's changing now. You have a lot of recent AI acquisitions in drug discovery, pathology, and lots of different related areas. Have you seen this shift, and is that part of it? Are you shifting toward selling models, or is this just rebranding because of that shift? What's going on?
To finish that narrative arc, the name change was just to reflect what we were really doing in practice, which is that we are an AI company, as we have been from day 1, when we were coding in PyTorch 1.7. Sergey was scaling transformers on PyTorch, the first to do that, and we were scaling graph neural nets in PyTorch, one of the first to do that. So it's reflecting what our identity has been from day 1. But also alluding to the fact that we are very serious molecule makers ourselves, and we are full-stack in that way.
We are also, to be clear, on a mission to create as many medicines as possible that otherwise would not have been created. I think the way to get there is to put our technology directly into the hands of as many drug developers as possible—not just our own, but in big pharma, mid-cap pharma, and biotech. That's partially why we've been building our agents, so that our models, which have become so much more powerful, can be easily used and orchestrated by med chemists and CADD scientists who might never have written a line of code in their life.
Shawn Wang
I find it interesting to hear that you've now started bringing agents into your development process. It sounds like, from the get-go, your philosophy was more that these are tools for scientists: you make the best possible tool, and this lets them optimize whatever pipeline they're going to use. Ultimately, this is a tool for med chem, and now you're switching and saying, "We're actually going to have automated systems which make decisions and then advance."
First of all, you did allude to having hit maybe the threshold—the November threshold—where it just magically became useful. Maybe that's part of it, but are there other reasons for going in this direction? Why now?
Well, let me ask you a question in return. How many tools do you think a human med chemist or CADD scientist can realistically learn and be proficient in?
Shawn Wang
I was going to say, how many can they learn, and how many do they use? I've seen that they like to use lots of them, but maybe they're not necessarily good with them. But it's really hard, right? A lot of the tools are extremely complex. They have a lot of parameters that you need to set properly and configure, right?
And this is where agents are extremely good: they actually know how to use all of the tools and how to orchestrate them well.
Shawn Wang
I wanted to interrupt you when you say "agent." Some people like using the term agent to just mean something automated, but other people use agent to specifically mean that you have an LLM responsible for orchestrating decisions. Outside of biotech, it's almost universally the latter, but sometimes in bio, people are still using agents to refer to any sort of decision-based system that's automated rather than run by humans. Maybe some of them are even more like traditional rule-based systems or something, right?
In our case, I mean, our definition of agents is an LLM that is able to use all of the tools that we have been painstakingly building over the many years of this company's existence. That's a very powerful concept, because now, as a CADD scientist or med chemist, you don't need to deeply understand all of the details and all of the hyperparameters that you need to set. You can just let this thing go and figure things out and solve the problems.
Shawn Wang
You're saying that you can take a crystal structure, which I guess in this case is framed as a graph or a 3D image that Molstar is interpreting in some way, and then it can actually make decisions about what is binding to what and make hypotheses? You've gotten to the point where you've seen it making effective decisions, right?
There are definitely multiple layers to this question. You can start with basic image understanding here, or you can use additional tools to understand what is happening in the crystal structure. You can even train a model that is able to natively understand crystal-structure representation. Image modality is a modality for LLMs these days; crystal structure can be a modality for an LLM as well.
Shawn Wang
I can fine-tune it on sequences that look like they're somehow capturing tokens that are 3D crystal structures, right?
Yeah.
Shawn Wang
So, if you historically treated humans as the people you were making tools for, and now you have agents, how do you reimagine the interaction between humans and agents and the decision-making process going forward? Do you hope to essentially automate humans out? Do you want to scale up and do more hypotheses?
And then the skeptic would ask: how do you make sure that your agents aren't making weird decisions, where they tweak the wrong hyperparameter, and now you give the human a bunch of outputs without understanding the parameters well enough to know that that happened?
Yeah, it's a great question, and I think the answer here can be seen in how coding tools have evolved, right? You remember Cursor when it just appeared? It was like tab autocomplete. It was still useful, but not to the extent that it is useful right now, where I can fire up Cursor and it just goes and does everything for me.
I think we're going to see a very similar trajectory here. These agents are already useful, but now humans need to provide consistent feedback and basically steer and guide the agent along the way. Over time, the agents are going to become more independent still. I don't believe in full automation or replacing humans; I believe humans are going to become way more efficient in the job that they're doing using the tools. A human will be providing strategic direction for what this agent needs to do, and the agent is going to be able to go and execute on behalf of a human.
Shawn Wang
I think this is the best time ever to be a creative human being, or really to be any human, because the long arc of history, starting from when we decided, "Hunting and gathering? I don't know about that. I think I'm going to settle down and do this whole agriculture thing," started thousands of years of drudgery. Sometimes you wonder why we actually adopted agriculture, because everything I've seen seems like agriculture was probably a net loss in the short term.
The only benefit I see in the long run is that if you were a hunter-gatherer and got sick, were born with a genetic illness, or got injured, you were in trouble. There wasn't much more we could do for you; medicine just didn't really exist. But ultimately, I think the crowning achievement of civilization is that the sick people among us in society are not seen as burdens, but as people we want to help. And that isn't true in the history of the animal kingdom, including humans.
It's a sad reality, but it's true. Whereas now, because of medicine, we see people who are sick as people we want to help, who we can help. We can help them, and they could become even more productive members of society if that's the thing that you care about.
If you fast-forward that to now, I think that so much of creative work and day-to-day work, if you look at it as a pie chart, time is our most precious resource. So much of it was being taken up by repetitive tasks that actually didn't require much thought or creativity. And now it's like we've created this additional headspace for deep thinking. For us, you're solving some of the most technically challenging but important problems. We can use these tools to become wildly more productive and creative.
That goes to the med chemists, drug hunters, and CADD scientists who are our direct users: they can be grand strategists for a drug discovery campaign. And when you have hundreds of drug discovery scientists working 24/7 on large data centers, just the output you get in terms of discovery is going to be greater, especially if it's used in the hands of expert scientists who know what they're doing.
Shawn Wang
We probably should have talked about this earlier, but you have some really exciting results from your Pearl model based upon a recent open dataset challenge that you alluded to earlier. Sorry we didn't talk about it before, but I would like to give you a chance to talk about this. There was this OpenBind—was it EV-A71 3C protease?—which was a very, very hard target. It had a lot of things that made it hard for the community, or for co-folding methods, to actually work, and you ran this internally. I think you had some fun results.
We can also go back to our conversation on evals and what's wrong with evals in general. If you look at the open-source models and their performance on public benchmarks, it's kind of very similar across the board. What was surprising when OpenBind came up with this new benchmark is that you see a wide variety of performance from the models on that. In part, it's because it's just a single target. Okay, I get it. But it's also because this is a target that those models haven't seen during training, so nobody has optimized for that specific target before. Nobody trained on it.
It was very interesting for us to see how our models were going to perform out of the box on this particular target, which we also had never seen before during training or when we developed these models. When it came up, we obviously wanted to look into it. We ran the Pearl system on that, and we were genuinely surprised and excited about the results. Our numbers are way higher than other published, openly available models.
To me, it reflects how Pearl actually behaves in real drug discovery programs, because we have seen these kinds of differences and shifts in internal programs or partnership programs before, but we were unable to talk about them because a lot of this data is either partnership data or internal data that we cannot disclose. Here, you have a perfect example where it's an external target that we didn't know about, and we just ran it. We see this staggering difference. That's what makes me excited.
There are some specifics of this target: it has this flexible loop that needs to move when you insert your ligand into the right location, and a lot of methods are just unable to handle it. Where Pearl was exceptionally good was figuring out how to move the loop, and we were basically correct for every single pose. We're predicting this movement really well. That's the exciting part.
Alessio Fanelli
I'm curious about this. What is it about Pearl that allows it to model a dynamic target like that?
The Pearl model, the way it's trained, is trained to predict the protein structure and ligand together. The protein structure and ligand together aren't necessarily in the same shape or form as the protein in isolation, so the whole training process nudges Pearl in that direction.
Alessio Fanelli
So are there lots of examples in the training set like that, where there's this dynamic nature, where during the binding event the conformation changes?
There is induced fit throughout the training set.
Alessio Fanelli
This is probably inherited from the physics-based simulations that you're using.
It certainly helps. I think the other component that Sergey alluded to was that we haven't just been benchmark-maxing on the same benchmarks. We obviously published state-of-the-art numbers on RMSD and PoseBusters last year, when we published the Pearl technical report with NVIDIA at GTC at the end of last year. We use RMSD and PoseBusters. Pearl does the best on those benchmarks. We did that.
Sergey Levine
However, as Evan is alluding to, the gap from Pearl to the next-best model is even larger on the partner drug targets that we work on. What that's in part a reflection of is that the objective of our company is not maxing benchmarks. The objective of the company is creating value for biotechs and pharma companies that are usually working on hard drug targets that are different from what's in the PDB. We've had to build the models in a way that allows them to extrapolate. The OpenBind challenge was just the latest example of that.
Shawn Wang
Before we wrap up, we always ask 2 questions. One is: if you could remove a bottleneck from your industry by fiat—and this is not, like, for inputs, outputs, or business—what would that be?
Sergey Levine
I can say where my main bottleneck right now is: GPUs.
Shawn Wang
Okay. Everyone says—
Sergey Levine
Yeah, look, GPU prices are going up and up and up, and LLM companies are kind of sucking up all of the GPU capacity out there. I firmly believe that what we're doing is genuinely important for humanity. Discovering new medicines for yourself and your loved ones is something that we all need at some point in our lives. I think it's very important to unblock that GPU bottleneck for drug discovery.
Shawn Wang
Is Anthropic actually throttling science because of GPU purchases?
Sergey Levine
Are you getting spicy?
Shawn Wang
I don't know if you can ask this, but are you looking at non-NVIDIA? I mean, I guess you guys are NVIDIA partners, whatever, but I've seen that some companies are pivoting toward a multi-architecture model so that they can take advantage of other suppliers because of the shortage.
Sergey Levine
NVIDIA's been a great supporter of us. They've invested twice in Genesis, and we collaborate closely—not just in terms of the capital, but we optimize kernels together. Pearl was co-authored by Genesis and NVIDIA. They publicized the OpenBind results with us. We've been very grateful for that. They're a very, very, very big company, and we're not at the scale of a customer like Meta or something, so we're grateful that they've given us time and attention nonetheless.
But I think part of that is a self-serving aspect here, which is that obviously no one can time the market, but the amount of hype and investment in the traditional LLM space for chatbots and coding, that sort of thing, at some point the amount of alpha will not be what's desired compared to the amount of investment. And I think, regardless of the vicissitudes of the economy, medicines will always be high in demand.
Whether it's a desire to help humanity or a self-serving interest in looking for the next big thing, I do think that chipmakers, including NVIDIA, are going to want to get a lot more invested in life sciences, because they will always be high in demand and the amount of alpha left in pure LLM space is just getting a little questionable.
Shawn Wang
I mean, all the LLM companies are starting to look into life sciences right now anyway, right? They GTB Roslin, there's a cloud initiative, and there are lots of startups and whatever.
Sergey Kiselev
Yep. We think those are all great tailwinds for us.
swyx
And then what about you? What's the bottleneck besides GPUs?
Oh, I had the same answer as Sergey, for sure. If I could wave a magic wand, I'd love to have just an enormous H100 GPU cluster. That would be amazing.
swyx
And then the second question we ask is: do you have a call to action for our audience? Meaning AI engineers plus scientists—get involved with something, come apply for jobs. What's the call to action for them?
We are definitely hiring. We're definitely looking for top talent in the space. I personally transitioned from the LLM space to drug discovery. I've never done drug discovery before coming to Genesis. I find this field fascinating and extremely interesting.
In LLM companies, you often find research scientists excited about working on architectures but forced to work on boring stuff. And honestly, LLM architectures are relatively boring. I don't know. I probably alienate half your audience.
swyx
I think they probably think the same thing. But it's like, it's a transformer layer in the end. The paper was published in 2017, and you go to any LLM lab today, you will see very, very similar pieces in their models.
Yes, there is a little bit of flavor of MoEs, but otherwise architectures are fundamentally very similar to what was discovered back in 2017. Our architectures and our models are actually very different and very, very interesting to work with, so if somebody is excited about working on architectures, this space is actually a very, very interesting place to work.
swyx
Yeah, I'd like to comment: not only are your architectures very different from the LLM space, but you're also very different from what a lot of the other bio-ML space is doing as well.
Alessio Fanelli
It’s sort of unique both in their subdomain and more broadly.
swyx
We saw, as a recent guest, you had one of the progenitors of ESMFold.
Yeah, they cite our model architecture paper that we published and open-sourced a few months ago. So, the work that you do here at Genesis as an AI researcher, there are lots of opportunities for publication, open source, and influencing the field beyond the obvious ways to influence the field by working with top customers that are actually making medicines that change people’s lives.
It’s both incredibly intellectually engaging and there are a lot of opportunities to influence the field and get your name out there. But if you want, you could also just stay being a cog in the machine of an LLM company.
Alessio Fanelli
No shade, no shade. Yeah.
swyx
All right. Well, thank you so much for coming and speaking with us. I’ve personally really enjoyed this. Hopefully the audience has as well. Evan, Sergey, thank you. We look forward to seeing how things develop, and hopefully you get some good job applications out of the—
Yeah, we love the show, so we really appreciate you taking the time to chat.
Sergey Kiselev
Yeah, thank you for having us.