到2050年治愈所有疾病|Sam Rodriques,Edison Scientific
Rodriques 的核心论点是,生物学受制于稀缺的科研人才,而“资本可以规模化,物流可以规模化,人才就是无法规模化”。 AI 可以通过高吞吐推理缓解这一瓶颈:阅读比任何研究人员都更多的证据,检验更多假设。但它无法取消物理验证,尤其当决定性实验是一项持续多年的人体临床试验时。
FutureHouse 的首个完整闭环智能体为干性年龄相关性黄斑变性生成了治疗假设,随后从湿实验验证推进到动物实验,并在访谈前2天发表于 Nature。 Rodriques 称2025年5月的结果是“未来已来”的时刻。其后继系统 Cosmos 可能已被用于生成20,000-30,000项新发现,但这个数字指的是提出的发现,而不是已经交付的获批药物。
Biewald 将商业张力概括为“进步来自发现,但商业价值在开发”。 Rodriques 表示,发现成功的频率很低,可能10年都不会显现价值;更快的实验则能推动药物走向患者。AI 可以协助准备试验方案和监管文件、协调临床试验中心,并推动化合物通过研发管线;分子设计专家也能同步改进候选药物。Rodriques 认为,未来的制药公司会精简得多,并用同样规模的人手推进更多项目,但人体试验仍是物理和运营层面的硬约束。
Edison Scientific 提出的护城河是专门化科学推理、面向客户的训练和最后一公里部署,而不是试图在所有领域击败前沿实验室。 Rodriques 表示,在合成化学等细分领域,少量任务数据就可能让模型相对通用模型取得“巨大提升”,而专有制药数据则能保留客户差异化。他也保留了判断:如果智能达到任务饱和,专门化可能失去优势;而在编程领域,专用模型还没有证明自己能击败最强通用模型。
这套平台有意利用前沿模型彼此不重叠的优势,而不是押注单一供应商。 Rodriques 表示,在复现科学论文分析的评测中,Anthropic 领先 Edison;而在 Cosmos 早期的一项世界模型任务中,只有 Gemini 2.5 能够完成。多模型架构也回应了制药公司不愿被锁定在 OpenAI、Anthropic 或 Google 之中的顾虑,毕竟领先者仍在不断变化。
Rodriques 一边热衷于加快实验,一边对不受控制的生物黑客行为持深度怀疑,并提出了颇具争议的临床试验改革方案。 未经研究的肽可能无效,也可能只是产生安慰剂效应,甚至造成无人追踪的伤害;在自行用药的人群中,“没有人在寻找”不良事件。但繁重的试验要求也在助推这种行为,因此他支持其所描述的澳大利亚和中国式分散化早期审批,并有条件地考虑在安全性得到证明后放宽疗效要求。
尽管 Rodriques 宣称这一轮 AI 周期有所不同,他仍把关键性 Phase 3 试验和 FDA 批准设为行业的硬指标。 自大约2012年以来,AI 药物发现一直过度承诺,AlphaFold 是显著例外;制药业内人士不会因为某个模型“提出了一种药”就买账。“答案在布丁里,而布丁就是获批药物。”
1. 生物学的人才供给上限,推动了从物理转向 AI
Rodriques 开玩笑说,物理学“实际上已经没有未解决的问题了”,随后给出严肃区分:大多数可见现象都能追溯到亚原子粒子,但人为什么会生病、老鼠为何会有某种行为、大脑如何运作,仍未被解释。“除了生物学,大多数事情都已经被很好地理解了。”
他在 MIT 的经历让他把科学产出归结为3种投入:资本、物流和人才。资金和实验基础设施都可以扩张,但“人才就是无法规模化”;治愈疾病、理解大脑和应对衰老,都需要一种能够放大研究人员推理能力的方式。
离开前景不错的神经科学和生物工程职业,感觉像是把一切都扔掉,“纵身跳下悬崖”。AI 吸引他的原因不是 AI 本身,而是它看起来最有可能把人才从科学瓶颈中移除。
2. 在需求提前数年到来之前,非营利机构是合适的孵化器
2022年,彼时已有 GPT-3 和 InstructGPT,但还没有推理模型范式,Rodriques 预计 AI 科学家需要5或10年才能出现,也看不到明确的商业模式。因此,他与智能体先驱 Andrew White 创办 FutureHouse,将其设为服务基础研究和广泛科学利益的非营利机构,并包括开源发布。
这一架构沿用了 Rodriques 早先提出的聚焦型研究组织思路:为那些“大到无法由学术界承担、又不具备风险投资经济性”的项目设立目标导向型非营利机构。比如绘制大脑中微观神经纤维,需要公司级工程能力,却可能要20年才能产生药物价值;FRO 可以围绕一个明确目标推进,完成后再结束运作。
FutureHouse 于2023年11月宣布成立,但强大的智能体在大约2年内就出现了,而不是需要10年。到2025年春季,制药高管已从需要别人解释智能体,转向要求立即部署;因此,Edison Scientific 的营利性分拆更像是扩大非营利研究成果的手段,而不是非营利模式失败的证据。
3. Robin 提供验证点,Cosmos 将搜索规模放大
Robin 是为药物再利用假设打造的、由人工编排的多智能体系统。它可以生成假设、规划实验、分析湿实验室返回的数据,并提出下一轮实验——构成“科学发现的完整闭环”,但物理工作仍由人执行。
2025年5月,Robin 为干性年龄相关性黄斑变性提出了一种新的治疗路径。Rodriques 称该疾病影响50岁以上人群中的约5%-10%。团队先通过湿实验验证,随后开展动物实验;论文在录制前2天发表于 Nature。
Cosmos 取代了 Robin 的轨道式工作流,改由一个能够协调自身子智能体的编排器运行。它的“世界模型”并不神秘,本质上是复杂的上下文管理:系统整合数十次或数百次运行积累的领域知识,并利用不断演化的表征进行发现。
Rodriques 估计,自 Cosmos 发布以来,用户已经生成了20,000-30,000项新科学发现。这个规模是其产品叙事的核心,但节目始终区分了提出潜在新结果与通过实验验证其有效性。
4. AI 擅长答案可核验或吞吐量决定价值的领域
Rodriques 毫不避讳地承认模型的智能极其“棱角分明”:有一次他让模型搜索 Notion 和 Slack,再为一项即将公布的交易起草新闻稿,结果糟糕透顶。模型真正的实用优势,集中在可验证任务,以及评估更多证据或假设本身就能创造价值的问题上。
编程和数学可以快速获得反馈,而科学只能在原则上被验证。针对药物的闭环强化学习可能需要证明安全性、让人类服用药物,再等待6个月,因此反馈循环或许长达3年;依赖海量数据和快速奖励的常见方法,无法顺畅迁移到科学领域。
巨型机器人实验室只有在自动化实验真正对应现实目标时才有帮助。移液机器人可以执行有价值的筛选,但在实验室代理指标上训练的系统无法判断药物是否对人有效:“模型的好坏,只取决于训练它的数据和目标。”
Rodriques 所说的“高吞吐推理”填补的是另一种空缺。智能体可以跨越单个研究者无法遍历的证据,理解生物学上下文,但研究人员仍需控制多重假设检验和 P-hacking,并最终开展必要实验。
5. 寄生虫和细菌操纵子展示了高吞吐推理的价值
为寻找自身免疫疾病治疗方法,FutureHouse 的一名研究员从自然界自身的免疫抑制剂入手:经过进化、能够抑制人体免疫反应的寄生虫。智能体会利用序列、结构、基因组上下文和生物体信息,检查寄生虫基因组中的每一种候选蛋白,再将数千种可能性缩减为实验室候选清单。
入选的寄生虫蛋白被合成到 DNA 芯片上,并筛查其对 T 细胞的影响。这个案例体现了分工:机器让原本不可能完成的搜索变得可行,物理实验则决定假设能否经受生物学现实的检验。
细菌蛋白提供了第二类推理任务。当一个功能未知的蛋白位于操纵子中,旁边是一个功能已知的蛋白时,智能体可以推断该操纵子的生化作用,并提出未知蛋白的可能功能——潜在目标是发现酶或生物工程工具,而不只是匹配序列。
Biewald 坚持认为测试仍需要“某种物理实体”,Rodriques 对此完全接受。他的主张不是科学可以仅靠推理完成,而是假设生成本身是瓶颈,如今可以被大幅扩张。
6. 即使发现创造进步,开发才承接价值
客户既用智能体生成早期假设,也用它们推进开发:订购实验材料、组装实验方案、协调临床试验资源和试验中心、准备监管文件,并推动药物走向患者。
Biewald 的商业框架非常直接:“进步来自发现,但商业价值在开发。”Rodriques 表示,发现经常失败,可能10年都不会显现价值;更快的实验,尤其是人体试验,才是把想法转化为科学和经济价值的环节。
在确定生物学机制后,另一种 AI 能力可以设计实际分子。对于口服候选药物,它需要具备合适的生物利用度、足够的血液半衰期、正确的受体活性、合适的组织分布、代谢和药代动力学、可耐受的毒性,以及尽可能少的脱靶效应——然后还要面对最初机制可能就是错的这一现实。
Rodriques 列举了 Chai Discovery、Isomorphic Labs、Boltz、Lambda Labs 和 Profluent Bio 等进军分子生成的公司。制药公司内部可能拥有很强的 AI 团队,但训练全球最好的抗体生成模型,“就是制药公司的赛道”,而不是专注型模型构建者的赛道。
7. 更精简的制药公司仍会撞上人体瓶颈
Rodriques 表示,未来的制药和生物科技公司会用同样的人手,并行推进更多药物项目。这正是移除人才瓶颈在运营层面的含义,也可能支持一个横跨多种疾病、规模类似人类基因组计划的医学项目,让此前经济性不足的疾病获得研发资源。
Biewald 提出一个表面上的矛盾:如果临床试验才是瓶颈,为什么还要强调候选药物生成?Rodriques 的回答是,药物开发存在多个瓶颈,但最终的运营约束仍然是人——执行试验的操作人员,以及必须在其中验证疗效的患者。
在大多数领域,除非受试患者稀缺,增加一倍资金和一倍操作人员,大致就能支持2倍数量的想法。
Biewald 指出,只要更多好想法具备可投资性,资本就可以规模化。Rodriques 的回应是,限制因素往往是把这些想法推进到临床试验所需的人和操作人员。
8. 对肽的怀疑与试验改革,是同一个论点的两面
被问到自己的肽类药物组合时,Rodriques 回答说自己没有,因为他“对生物学了解得太多”。研究药物开发会让人看到大量细微环节可能出错,也会看到安慰剂效应有多强,因此他对从未在严格对照试验中验证过的化合物持怀疑态度。
服用肽后感觉好转,几乎不能证明什么:这无法说明究竟是肽真正起效,还是安慰剂在发挥作用。更严重的是,即便心脏骤停的发生率系统性地达到10%,也可能无人察觉,因为在自行用药的人群中“没有人在追踪这些不良事件”。
Biewald 表示,他钦佩生物黑客愿意尝试新事物。Rodriques 仍认为,不受控制的自行用药通常是不明智的;问题在于,人们在无法清楚判断不确定性、潜在伤害或可靠测量方式的情况下进行实验。
另一面是,缓慢而繁重的临床试验会把人推向不受控制的测试。Rodriques 表示,澳大利亚和中国将部分早期研究审批分散到单个中心,使其能够在速度上展开受监管竞争;而在美国,即便是小型试验也需要 FDA 集中审批,因此美国生物科技公司会把研究带到海外。
9. 放宽疗效规则,既带来个体化也带来统计滥用
Rodriques 提议考虑一种审批路径:在证明安全性的基础上,不总是要求药物必须对某一种预先定义的疾病证明疗效。一个药物可能帮助真实存在的亚组,却在总体试验中失败;安全性得到确认后,现实世界的使用或许能揭示谁会受益,尤其是当标准治疗已经对患者无效时。
个体内比较可以提供证据:如果一种抗抑郁药让症状没有变化,第二种却带来明显改善,那么发生变化的药物就具有信息价值,也部分避开了从“不接受治疗”转向“接受任何治疗”所带来的安慰剂效应。与已上市药物进行比较,也可能揭示更优的结果。
Biewald 反驳说,回顾性真实世界分析会诱发 P-hacking 和事后亚组选择。Rodriques 同意风险存在,但将其框架化为统计功效问题:在大量亚组中搜索后发现的模式可能只是伪象;而如果某种疫苗效应在1,000万人的数据中始终只出现在18-35岁男性身上,即便仍需确认性实验,也说明存在值得调查的信号。
Biewald 还问,真实世界数据有多少次能在没有选择偏差的情况下达到这种规模。Rodriques 指向覆盖大规模人群的疫苗和常见癌症,同时承认数据质量问题极其棘手:患者会漏服药物、跳过随访,医生记录质量不一,医疗记录也分散在不同医疗系统中。他表示 Edison 已经与客户合作开展这类分析,但现有真实世界证据“永远达不到你希望的程度”。
10. 医疗决策支持,用明确不确定性替代令人安心的答案
Biewald 将医学描述为高风险却始终存在歧义的领域:研究彼此冲突,数据集带有不同偏差,而患者越来越多地要求语言模型为个人决策综合研究。Rodriques 表示,他的朋友和家人已经这样使用 Cosmos。
一位朋友遭遇乳腺癌疑云,询问推迟1或2个月治疗会如何影响康复或生存概率。Rodriques 表示,平台会“毫无偏见地”检索文献,依据论文而非“感觉”给出答案;这位朋友最终被证实没有问题。
Biewald 讲述了自己将女儿撞到头的同一段描述分别输入2个账户:他的模型说女儿没事,而妻子的模型建议就医,这可能反映了对话历史,或模型被强化到迎合用户想听的答案。Cosmos 更可能返回发生严重后果的明确概率,以及一切无恙的概率——更有证据基础,但有时也更令人沮丧。
更大的张力在于,智能增加并不会自动改善判断。“更多信息不一定总能让决策更好”;用户需要的是正确的信息,而不一定是最大量的信息。
11. 专业化、专有数据与模型多样性构成护城河
面对 OpenAI、Anthropic 和 DeepMind,Rodriques 认为专业模型应当击败通用模型,因为后者必须把所有知识编码进权重。Edison 针对特定科学任务训练推理模型,并称在合成路径等细分领域,少量数据就能带来巨大提升,同时投入更多测试时算力来压制幻觉。
Biewald 以编程发起挑战。Rodriques 表示,很多人都在尝试让专用编程模型击败前沿通用模型,但目前还没有做到;编程可能是例外,因为每一家主要实验室都将其列为优先事项。他还提出第二重保留:一旦某项任务达到“智能饱和”,就像人类已经掌握最优井字棋一样,继续增加专业化或智能可能不再有帮助。
制药公司的专有数据强化了 Edison 的位置,因为客户希望模型基于自己的证据训练,而不是让每个竞争对手都能获得同样的智能。Edison 将定制化与最后一公里的内部部署结合起来,而前沿模型供应商可能会回避这项工作,因为它们不想为每个客户分别打造一个模型。
客户也不愿依赖单一前沿供应商:某一年 OpenAI 似乎遥遥领先,录制时 Anthropic 又让人感觉领先。Rodriques 表示,这些模型的优势存在显著差异——Anthropic 在复现论文分析的基准测试中特别强,而在 Cosmos 早期的一项世界模型任务中,只有 Gemini 2.5 能够完成——因此组合多个供应商可以显著改善结果。
12. 长时运行智能体逐步转向可操控性,而非缩小目标
最初的 Cosmos 可以运行6至12小时,在一次运行中写出45,000行代码、阅读1,500篇论文。Rodriques 称其为现有最密集的智能体之一,但也意识到科学家很少愿意提交一个问题后消失半天。
最近的一次更新加入了交互式前端,但没有放弃长周期工作。科学家现在可以提供反馈、引导调查方向,同时底层系统继续编排持续时间很长的研究。
这一演进体现了更广泛的成熟过程:从展示智能体框架,走向交付可用的科学基础设施。Rodriques 仍称 Edison 的智能体在科学能力上领先前沿实验室产品,但越来越多地通过部署、专有训练、工作流适配和管线加速来描述其价值。
13. 科学生产率正在放缓,但人的品味仍有价值
Rodriques 提到 Eroom 定律——Moore 倒过来拼写——描述药物开发实际成本在历史上的上升,并谨慎回忆其大约每9年翻倍。他表示,这种恶化在2011或2012年左右出现转折,可能是因为人类基因组计划让靶点更容易被发现,但此后生产率大致持平,而不是继续提升。
一种解释是容易摘的果实已经被摘完;另一种是“必须比 Beatles 更好”的问题。新疗法必须击败现有的同类最佳治疗,但药物无法像半导体那样叠加此前机制和工程进步;首创新机制往往带来最大的跃升,而之后每一次发现都会更难。
Rodriques 仍建议攻读 PhD,因为其目的在于学习如何做研究,并预计人类科学家在相当长一段时间内仍不可或缺。科学并不能被廉价验证,因此人类提供的是品味,尤其是判断哪些问题有趣,即便模型越来越擅长提出可能有效的想法。
他的不确定性保持得很明确:在5或10年内,模型是否会超越顶尖研究人员的判断,他并不知道。但在录制时,他认为“毫无疑问”,诺贝尔奖得主 Frances Arnold 提出的定向进化研究会优于 Claude 3.7。
14. 7000万美元买来部署,审批决定可信度
Edison 的营利性公司融资7000万美元,独立于非营利机构的另行融资。Rodriques 看重能够帮助其打开制药行业大门的投资人:Spark 的 Yasmin Razavi 带来前沿 AI 视角,前 Grail CEO Jeff Huber 带来广泛的行业关系,一家未具名的大型机构生物科技投资者则帮助其接触高管并促成部署。
两类投资人的尽调差异很能说明问题:科技投资人反复追问 OpenAI、Anthropic 或 Google 为什么不会吞并 Edison 所处的市场,而制药业内人士很少这样问。Rodriques 的解释是,真正了解药物开发的运营者“对价值有切肤之感”,能直接看到专业化工作流和部署问题。
他承认,AI 药物发现已经过度承诺了10多年,AlphaFold 是蛋白质结构驱动分子工作中显著的成功案例。他那句绝对化表态——“我向你保证,这次不一样”——随即被一个硬终点约束:在关键性 Phase 3 试验成功或收到 FDA 批准函之前,提出的发现不会让制药公司留下深刻印象。
1. Using AI to synthesize medical research
If we want to go and cure all diseases, understand how the brain works, solve aging, and so on, AI seemed like the right way to do that. When it comes to the world as we experience it, most things are pretty well understood except for biology. To this day, we don't have the wiring diagram for the human brain. We kind of have the wiring diagram for the fly, and that's the best that we have.
When we got started in 2022, remember, we didn't have any notion of reasoning models or whatever, right? We showed the first multi-agent system that was capable of doing the full loop of scientific discovery, and it came up with a new hypothesis about a way to treat a form of blindness called age-related macular degeneration. It actually just got published in Nature 2 days ago. Since we launched Cosmos, people have probably used Cosmos to make 20 or 30,000 novel scientific findings, which is wild.
The pharma companies of the future will be much leaner. You're going to be able to pursue many more drug programs in parallel than you can today with the same number of people.
There's a long history of AI overpromising in drug discovery. Is this time different?
Oh, yeah. I promise you, this time it is different.
2. Introduction
You're listening to Gradient Descent, a show about making machine learning work in the real world, and I'm your host, Lukas Biewald. All right, I'm here with Sam Rodriques, the founder and CEO of Edison Scientific and Future House. He's right at the forefront of using agents for scientific innovation, and he's got a lot to say on how to use agents well and science itself. I really hope you enjoy this podcast.
All right, Sam, thanks for doing this. I've really been looking forward to it.
Likewise.
3. Origins and the role of nonprofits in technology
All right. So, let's get started. You had this incredibly promising, successful career in neuroscience and bioengineering. I think you studied theoretical physics originally, and then you moved into AI.
Yeah, right. What went wrong?
What were you thinking? Why did you do it, Sam?
Why would I give up such a promising career? Yeah, great question. I started doing theoretical physics and did quantum information theory. The problem with physics is there are actually no unsolved problems left in physics. That's a slight exaggeration. We need to build quantum computers, and there are some interesting, important materials science problems, and we still don't know how the universe works, right?
But actually, I think one of the key insights is that if you look around at any phenomenon that you can see in the room today—in this room that we're sitting in—I could probably explain to you how that phenomenon works down to the level of subatomic particles, right? Except for you, or the mouse. Why does the mouse do what it wants to do? Why do you wake up at the time you wake up? How do you get sick, and so on, right? When it comes to the world as we experience it, most things are pretty well understood except for biology.
And so that was originally what got me into biology. Then I did my PhD at MIT. I'm an inventor at heart. I invented a bunch of different technologies. The key thing that I learned at MIT about doing biology is that there are 3 things that you need to do science, right?
You need capital—you need money. You need logistics, which is: do you have the things that you need to run the experiments you need to run, in the place and when you need them? And you need talent. Capital scales, and logistics scales, but talent just does not scale. Fundamentally, in biology, we're limited by talent.
The thing I got to thinking about is: how do we remove talent as a bottleneck in science? If we want to go and cure all diseases, understand how the brain works, solve aging, and so on, we need to figure out how to scale talent, and AI seemed like the right way to do that. I had basically built up this career, and I kind of just threw it away and jumped off a cliff. That was pretty wild, but it's been working out well so far.
Originally, the company that you started—which you should explain what it does—I think it started as a new kind of nonprofit, right?
Yeah. So, this was 2022. I had figured out that the most important thing that seemed like it was going to happen in science in the next 10 years was figuring out how to build an AI scientist, precisely because it was going to unblock talent.
At that time, in biology, you're used to things going really, really slowly. We were at this point where GPT-3 was out, and it could kind of say things. It was very impressive. You could kind of see where things were going. But it seemed like it was going to take a really long time to get there, and I didn't know how to commercialize it. I didn't know what it was.
We were like, “Well, we should do this as a nonprofit. Just go and do the basic research,” which was great. I teamed up with Andrew White, my co-founder, who is really a pioneer in the space of AI agents for science. He was working with OpenAI on GPT-4 at the time and was a professor at the University of Rochester.
We had the same vision for building an AI scientist. I knew the biology side, and Andrew actually knew how to do it technically. It felt like the right thing to do was to be a nonprofit because we thought it was going to take 5 or 10 years, which was complete nonsense. That was totally not right. As we all know now, within 2 years of launching FutureHouse—we announced it in November 2023—we already had extremely powerful AI scientists.
The timing wasn't something that ended up being very important. I will say that the fact that we started FutureHouse as a nonprofit was not the only reason we started FutureHouse as a nonprofit. We also really wanted, and continue to want, those technologies to benefit the entire scientific community. At FutureHouse, we open-sourced a lot of stuff, which I think is also super important.
By spring 2025, one day Andrej Karpathy tweeted or something, and all of a sudden the entire world knew what AI agents were. We started getting phone calls. We literally went from having to explain to people in our decks what an agent was and how it was different from a language model to having senior executives at pharma companies call us to say, “Oh my God, how do we use your agents? How do we use your thing?”
It became evident pretty quickly that we were going to have to have a for-profit spinout in order to satisfy that demand.
Do you feel like there's a role for nonprofits in technology, then? We now have 2 examples of nonprofits developing interesting technology and mainly turning it into for-profits to get that technology out into the world. Is there something broken about the nonprofit model?
No. First of all, is there a role for nonprofits in technology development? Yes, absolutely. Is something broken about nonprofits? No, definitely not. This just feels like for-profits doing what for-profits do well, which is scaling.
Let's back up. When I was in my PhD, I was working on a connectomics project. Connectomics is figuring out how to map all the connections between neurons in the brain. A core piece of understanding how the brain works is figuring out how to get the wiring diagram, and to this day, we don't have the wiring diagram for the human brain. We don't have the wiring diagram for the mouse brain. We kind of have the wiring diagram for the fly, and that's the best that we have. It's really, really difficult to understand how the brain works without the wiring diagram.
We wanted to map the wiring diagram in the brain. It's really, really challenging to do that, basically because it involves tracing these tiny, tiny fibers—tiny meaning 1/100th the width of a human hair—through this gigantic tangle of the brain. You have to make no errors, because if you make errors, then you're connecting neurons that aren't actually connected to each other. It's like trying to imagine mapping all the roots in a field of grass, but way, way harder.
That was the problem I was interested in, and I tried to do it in an academic lab. I was like, “Wow, there's no way I'm going to be able to gather the engineering resources, the capital, and the talent that I'm going to need in order to do this.”
And you can't do that for profit, right? You actually can't go into a for-profit setting and get money to map the brain, because what are you going to do with a map of the brain? You're going to develop drugs, and that will take 20 years or something. That's not an attractive proposition for investors.
4. Is this time different for AI in drug discovery
So I was like, "Okay, I can't do this in academia. I can't do it for profit." Basically, there's no way for me to do this right now, right? So I proposed these things called focused research organizations, which is the idea that for some problems, like mapping the brain, that are too big for academia but can't be done for profit, we need a third way. We need some alternative kind of structure.
The idea of a focused research organization is that it's a nonprofit that operates a lot like a company, not like the slow-moving foundation that you think of when you think of normal nonprofits. It's a fast-moving, hard-charging research organization that has a specific goal. It pursues that goal, and then when it's done with that goal, it's done. You can actually spin down the nonprofit.
Since then, we've gotten philanthropists to fund a bunch of these focused research organizations, and they're doing projects that you just couldn't imagine happening in a for-profit setting. Sometimes, if they really work out—if they hit and go really well—then what is the right step afterward? It might just be spinning out a for-profit. That has always been the idea.
With FutureHouse, we started with the goal of building an AI scientist. We didn't know how to commercialize it, so a nonprofit made sense. Plus, we wanted to be able to share the basic research with the world. Then it went really well, and we made a bunch of progress. What is the next most sensible step? Spin out a for-profit and allow it to scale.
This was actually not the plan. At the beginning, we thought maybe we'd spin out a for-profit in 5 or 10 years. We certainly didn't think it would take us 2 or 3 years. But I think this is just a consequence of success, as it was also in OpenAI's case.
5. Where AI is strong and weak in science
What did you see that made you feel like the technology was even more promising than you thought?
When we got started in 2022, remember, we had GPT-3 and InstructGPT. The language models then knew how to respond to you, right? But we were very much in a world where these models just couldn't do very much. We didn't have any notion of reasoning models or anything like that. They just seemed very rudimentary.
It was 18 or 24 months before they started to make discoveries. The first discovery they made was described in a paper about a system called Robin. It was the first multi-agent system that we showed was capable of doing the full loop of scientific discovery: hypothesis generation, experiment planning, running the experiments, analyzing the data, and coming up with new experiments.
It came up with a new hypothesis about a way to treat a form of blindness called age-related macular degeneration, specifically dry age-related macular degeneration, which affects 5 or 10% of people over the age of 50. This was in May 2025. Our agent came up with a new way of treating it and proposed a new treatment that we were able to validate in some wet-lab experiments. Subsequently, we've been able to validate it in animals. It was just published in Nature 2 days ago, and that was the thing where we looked at it and thought, "Oh, the future is here."
That makes sense. Yeah. So where's it gone since then?
I mean, that was discovery number 1, back in May. We have since released an updated version of our agent called Cosmos. This is a much more powerful version of Robin.
Robin was kind of on rails. It could do this, but we manually orchestrated a bunch of different agents to get it to do 1 specific thing, which was identify new ways of treating diseases, particularly through drug repurposing. With Cosmos, we built in an orchestrator so that it could orchestrate itself.
We also built in this notion of world models, which is basically a very sophisticated context-management tool. It allows Kosmos to build up an integrated notion of the knowledge in a field over the course of dozens or hundreds of subagent runs. That is core to the way it makes discoveries.
Since we launched Cosmos, people have probably used it to come up with 20 or 30,000 novel scientific findings, which is wild.
Wow. Yeah. I think anyone who's used AI for any technical application, and maybe even for nontechnical applications, notices that the intelligence is really spiky in surprising ways. There are some things it does so much better than a human, and some things it does shockingly worse.
Yeah.
It wrote me a press release the other day.
Oh, my God. I asked it to go into our Notion and our Slack, look up all the details for a deal that we're going to announce shortly, and draft a press release. It was terrible. I was like, "I cannot—I don't know." Yes, extremely.
I appreciate you saying that because I've had a few guests on this show lately who kind of refuse to acknowledge any weakness in their algorithms, and it gets really boring and weird because these models are going to struggle with some things. In practical fields, where do you feel like it's really strong, and where do your customers get surprised that it can't do something?
Okay, so it's strong in 2 areas. It's strong on things that are verifiable, and it's strong on things where throughput matters a lot.
Okay, so it's strong in 2 areas: things that are verifiable and things where throughput matters a lot.
Verifiable means that you can tell whether or not an answer is correct. This has always historically been where AI has been strongest, because it allows you to get feedback very quickly, and then you can use reinforcement learning. AI is strongest where there's a lot of data and where the problems are verifiable. Coding is verifiable. Math is verifiable. Science is very much nonverifiable. It's verifiable in principle because you can go and run experiments.
Right, but the loop is expensive.
The loop is expensive. It takes forever, right? People say, "Why don't you just do closed-loop reinforcement learning to teach it how to find new drugs?" I'm like, "Because that loop means I need to prove that this drug is safe, dose some humans, and then wait for 6 months." The loop is going to be 3 years long.
The other thing I hear a lot is, "Why don't you build a gigantic science warehouse where you just have pipetting robots doing experiments over and over again?" That's a better idea, and there are a lot of things for which that is probably great and very useful. But in general, that requires the experiments you're doing in a lab to be reflective of what you want.
In our case, we want medicines for humans. I can't test whether a drug works in a human using a pipetting robot in a lab. I have to test it in a human. Your model is only as good as what you train it on.
Basically, AI is good at 2 things. It's good at tasks that are verifiable, and it's good at tasks that require high throughput. On the high-throughput side, I think this is where it has really shined so far in science, because it's able to consider so much more evidence than any human is able to consider and test so many more hypotheses than any human is able to test.
You have to control for p-hacking and multiple-hypothesis testing and so on, but in biology, we have a problem with throughput. This is sort of about statistical synthesis of data. I'll give you an example: one of our FutureHouse postdoctoral fellows is interested in figuring out how to cure autoimmune diseases.
Okay, in order to cure all of them?
Ideally, yeah. We'll start with 1, but actually we don't need to be picky. In order to treat autoimmune diseases, what you need are ways to manipulate the immune system.
If we go to nature for inspiration, where has nature figured out how to manipulate the human immune system? Parasites. Before modern hygiene, humans just lived with parasites. The way that worked was that the parasites had figured out how to manipulate the immune system to turn down immune reactions.
Wow. If we could figure out how they do it, maybe we could do it on ourselves in order to cure autoimmune diseases, right? This is the idea.
But there are many, many parasites, and each parasite has a genome with thousands or tens of thousands of proteins in it. What is the mechanism? How are you going to figure out what mechanisms the parasites use to regulate the immune system?
We're actually now using our agents on every single protein in any parasite genome. We're running our agents to look at that protein, look at its structure, look at its sequence, look at the context in which it appears in the genome, look at the biology of that organism, and figure out: Could this be a candidate for how this parasite moderates the immune system?
From that, we've come up with a short list of proteins that we need to test, and we're now testing them in the lab. That is something that previously there was no way to do. It would have been completely impossible.
Right. But presumably, the hard part of that is building tools to—when you say “look at a genome,” it's obviously not feeding the genome into the context window. It's finding tools to—
Correct.
—to look at what's going on.
But the biology—but I call this high-throughput reasoning. Yes, you can just run some homology. You can come up with an algorithm to look at the sequence, or you can use a model of protein structure and just look at the structure of all the proteins. That doesn't tell you what role that protein might play in the biology of the organism, which is fundamentally what we're interested in.
Another good example is bacteria. There are many proteins that have no known function, particularly in bacteria. If we knew what their functions were, we might be able to figure out new ways to create bioengineering tools, like finding new enzymes that do functions we can't do today.
One of the ways that you can do this is that bacteria organize their genomes into units called operons that are all functionally linked. If you have one protein with a known function inside an operon, and you also have a protein with an unknown function in that operon, you can reason about the function of the latter protein based on the former protein, which has the known function.
That's something where you really just need intelligence in order to think about the biochemical role of this operon. What role does it play in the bacterial life cycle? Based on that, can we come up with a hypothesis for what this protein of unknown function might do?
Right. But when you test that hypothesis, presumably you have to do something physical, right?
Absolutely. In the case of the autoimmune disease parasite proteins, we've now gotten a bunch of them synthesized on a DNA chip, and we're screening them to see what effect they have on T cells.
You absolutely can't just reason your way to solving science. You have to do experiments. But coming up with hypotheses is one of the steps that is limiting, and that is something where the models are able to help us.
Got it. So presumably, your customers are using your models mostly for drug discovery. Is that right?
Yeah. Customers use this in 2 ways. I just talked about hypothesis generation. The other place where you can imagine these models having a major impact is in the operational work of science.
How do we actually get all the materials together that we need to run an experiment? Once you figure out what experiment you want to do, you need to order and organize the materials and figure out what the protocol is. When you get into testing on humans, you need to coordinate all the resources you need in order to run the clinical trials. You need to figure out what the clinical trial sites are and prepare your regulatory documents.
All of that is part of the process of doing science, and we're used in both areas. We're used for early-stage hypothesis generation and for the actual operational tasks of development: How do we get this drug through the pipeline and to patients as quickly as possible?
That's really interesting, because Weights & Biases customers do both also. At first, we saw mostly drug discovery applications, and then, post-LLMs, we started to see all these operational applications. I started to think that maybe the operational applications are more important. Certainly, I think more dollars go into that, and it's more of the bottleneck.
Do you have a passion for the drug discovery side and want to focus on that, or do you think the bigger business here is on the operational side? There's this very funny situation in science where progress comes from discovery, but the commercial value is in development.
The reason is that discoveries pan out so infrequently, and you can't tell whether or not they're valuable until 10 years or something after you came up with your hypothesis. What matters commercially—and frankly, what matters practically—is being able to run experiments faster.
You can come up with as many hypotheses as you want, but the experiments that matter are human clinical trials, right? That's what tells us whether or not the hypotheses actually work in practice.
If you want to accelerate the process of coming up with medicines, you need to make those experiments faster. That's where the commercial value is, and that's where a lot of the scientific value is. This is not to say that early-stage discovery is not important. It's also critically important.
6. Changing drug discovery
We've had a whole slew of guests on this podcast doing different parts of the drug discovery pipeline. We've had CEOs and researchers. We've even had notorious pharma bro Martin Shkreli giving his take. He was actually down on the whole drug discovery-with-AI thing.
I'm curious how you think about the entire drug discovery market and how it's changing. What's working, what's not, what's changing, and what's static?
Yeah, great question. I think, obviously, there are 2 big areas where AI is having a major impact. The first one is what I talked about, which is all of the reasoning. That includes both hypothesis generation, which we do, and the operational aspect of development: How do you get these molecules through to patients faster? We also do that.
The second major important part is coming up with the molecule. If I have this hypothesis that agonizing the GLP-1 receptor is going to cause people to lose weight, then in order to test that, I need to have a molecule that, inside a human, is going to agonize that receptor, bind to the receptor, activate it, and so on.
That is way harder than it sounds. If we imagine that you want to do this orally, with a pill that you just swallow, odds are it's not going to go into your bloodstream. If it does go into your bloodstream, odds are it's just going to get filtered out immediately by the liver. Even if it doesn't get filtered out by the liver, it's probably not going to stay in your blood long enough to have an effect.
Even if it activates the receptor, maybe it activates the receptor in the wrong way. Even if it activates the receptor in the right way, maybe it activates 100 other receptors that you don't want to activate, leading to bad side effects. Then there's the possibility that your original hypothesis might not work.
Coming up with a molecule that satisfies all of those criteria—that it has the right bioavailability, meaning that it gets into the body; that it has the right distribution, metabolism, and pharmacokinetic characteristics; and that it has the right toxicity characteristics—is another extremely critical portion of the process of discovering and developing drugs.
There have been really exciting and revolutionary companies like Chai Discovery, Isomorphic Labs, Boltz, Lambda Labs, and Profluent Bio working on the problem of coming up with the actual molecule.
The upshot is that I think the biotech and pharma companies of the future will be much leaner.
Right.
You'll be able to pursue many more drug programs in parallel than you can today with the same number of people. I think it's time for us to start thinking about what that means. Given that we'll be able to remove the talent bottleneck, we're going to be able to pursue cures for so many more diseases.
What would the Human Genome Project look like for medicine?
But I feel like earlier you said the bottleneck was human clinical trials.
Absolutely.
And what you just said makes it sound more like the bottleneck is coming up with the candidates for the clinical trials. Is that really the bottleneck? Which is it?
No, no. First of all, let me just say that there are many bottlenecks that all need to be overcome. But I think that fundamentally, at the end of the day, there’s an operational bottleneck, which is that you need humans in order to run the clinical trials. If we had twice as much money and twice as many operators in drug discovery, then in most areas—this is not always true, because sometimes you’re limited by patients and so on—you would just be able to pursue twice as many ideas.
Right? So we’re actually limited. Like I said before, capital scales, right? If we had more good ideas that were investable, then there would be more money available for them. Do you think that the new upstarts that you mentioned, like Isomorphic Labs, Chai Discovery, and others, have a structural advantage over the incumbents in this process?
I mean, yes, in that mostly what those companies are doing is coming up with molecules that they then partner with pharma companies to develop. So they are mostly—not today, at least to my knowledge—developing their own drugs, and are instead trying to get this technology everywhere.
I think that the pharma companies don’t have the expertise in-house to do it themselves, and so that’s why they’re going out and partnering.
I mean, they do have teams that supposedly work on drug discovery, right?
Oh, they do. I mean, the pharma companies have strong AI teams, right? Actually, we have been very impressed with the quality of the AI people we find inside many pharma companies.
But when it comes to training the best model in the world for generating a new antibody, that’s just not the pharma company’s game, in the way that it is Chai’s game, Latent Labs’ game, and so on.
7. Peptides and clinical trial processes
Got it. Okay, so Sam, tell me about your peptide stack. [laughter]
I had to think about what you meant for a minute, because we’re actually working with one of our partners on training a model that is better at reasoning about peptides. So I was like, “Wait, how does he know about that?”
My peptide stack? Okay. I’m going to admit I’m not a biohacker.
Wow. No peptides?
No, no peptides.
Wow.
Yeah, I know. The reason is that I know too much about biology. And I think that you’ll find this—I won’t put this person on the record.
I was hanging out the other day with the head of research and development at a very large, top-20 pharma company. He was talking about the peptide fad, and he was just like, “Yeah, these people don’t understand what can go wrong.”
Oh, I see. I think that’s true. When you have studied biology and drug development, you get an appreciation for everything that can possibly go wrong, including things that are very, very hard to identify if you’re not looking carefully.
Interesting. You also gain a deep appreciation for things like placebo—the strength of placebos.
Those 2 things together just make you very intrinsically skeptical.
So, are you nervous even about something like Ozempic that tons of people take?
I mean, Ozempic has been used to treat diabetes for a long time, right? Or, at least, GLP-1 agonists have been. So I’m not that concerned about it. The GLP-1 agonists have now been through many, many robust, well-controlled trials, and so I don’t worry about that.
Well, these are separate concerns, right? Does it not do anything, or does it hurt me?
Yeah, those are different concerns, right? But I’m also just a scientist. I think the issue is that when you think about trying peptides on yourself that have never been studied in a really robust trial setting, you have no way to know whether they’re really working.
People can feel better. That’s great. Placebos also make you feel better, so that doesn’t tell you anything about what’s making you feel better, right?
Even if you say, “Oh, look, everyone who takes it feels better,” right? Yeah, that’s the point. Yeah, I know. That’s right.
If this group of people were systematically—even if a very large group of people taking some of these peptides systematically had, say, a 10% incidence of cardiac arrest or something—we wouldn’t know. No one is looking. No one is tracking those adverse events when people just go and dose themselves with peptides, right?
I don’t want to sound too much like a—my recommendation would always be: probably don’t. I feel like I have a professional, ethical, and moral obligation to say to people, “Generally, doing this is ill-advised.”
Right.
But this is totally off the record, so you can say anything you want on this podcast.
Yeah, exactly—completely off the record. [laughter] I do admire people who go out and want to just try things, because I’m a fan of just trying things, right?
We just need to make sure that you’re clear-eyed about what the risks are, because there are no guarantees with a lot of these things, and they can do damage. Would you change anything about the clinical trial process that we have in the United States?
So, the flip side of what I just said about peptides is that part of the reason why people feel compelled to go and test things on themselves is that the process for testing them in humans in a robust, rigorous, and well-controlled manner is so onerous and takes so long.
If it were possible to run really high-quality, well-controlled studies in a way that is well regulated and so on, quickly, then maybe people would be doing that instead, which would probably be better.
There are a huge number of things that we should be doing. The first one, which is just the most obvious, boneheaded thing—we absolutely should be doing it; it’s crazy that we’re not—is that in Australia and China, for early-stage studies, the process of getting a clinical trial approved is decentralized relative to the way it is in the US.
In the US, you need to get central approval from the FDA even just to do a small-scale initial trial. In China and Australia, you have individual centers that run trials that are capable of approving trials. That means those centers can compete over how easy they can make it to do trials while staying within the regulatory guidelines, right? That is a drive for efficiency, and that is great.
As a result, US biotechs are going to Australia and China to do their clinical trials. This is obvious, right? We should be fixing it.
Actually, if there’s any administration that is going to fix it, you would think that the bull-in-a-china-shop kind of approach that this administration takes would be a great candidate to do it. They need to be doing that.
Then I think there are other things that we should definitely be looking at. One of the more obvious ones is loosening the requirements for efficacy.
The FDA requires you to prove 2 things in order to get your drug approved: They require you to prove that your drug is safe, and they require you to prove that your drug is effective for treating a specific condition.
Usually, drugs fail on efficacy. The thing about this that’s a little bit perverse is that often a drug will work in a specific subpopulation—it will work in 1 population of patients, but it does not work in the entire trial population. Therefore, the trial fails, and therefore it can’t get approved, even though it works on some subset of patients.
That feels like a failure of the system. The alternative is to only require that drugs be proven to be safe, and then determine that they are effective in the course of using them on patients in the field.
Right now, you can imagine that at some point this would not be ethical, and this would not be something that you would want to sign up for in all cases. If there is a standard of care that is known to be effective and then there's a drug where we know it's safe but it's not necessarily effective—we don't know if it's effective—you might choose the effective one. On the flip side, if the effective one doesn't work for you and you have no other options, you might choose to go with this other one. I think that would probably reduce the barrier to doing clinical trials in many diseases—not always, but—
It's so hard to know what's effective, right? I feel like I have lots of smart friends who take lots of medicine or things like that that they think are effective. In my mind, I'm thinking, “Probably not,” because the science doesn't show that it's effective. I'm actually not even sure who's more likely to be right.
But what you can do, absolutely, is use randomized controlled trials. In the standard process, you have patients, and half of them get the actual medicine; half of them either don't get the medicine or, more often, get the standard of care, because you would usually consider it unethical to deny a patient a medication. Again, it depends on the condition and the disease.
The other way that you could do it, if we were to relax the efficacy requirement so that you only have to show a drug is safe, is to look at the extent to which a single patient improves—you could look at within-patient improvement. If you take drug A for depression—and, obviously, antidepressants are famously very variable in whether they work in any given individual—maybe you take antidepressant number 1 and it doesn't work for you and your symptoms persist, and then you take antidepressant number 2 and suddenly you get much better. That's pretty good as far as evidence goes, because the only variable there that has been changed is which medicine you're taking.
That avoids the placebo problem to some extent, right? The issue with the placebo is you go from not taking medicine to taking a medicine, and simply taking a medicine is often effective, regardless of whether the medicine works. The other way that you can think about doing it is patients who are on drug A, which is known to be effective, versus patients who are on drug B, which is not known to be effective. If drug B is more effective—or if the patients on drug B have a better outcome than the patients on drug A—that's very strong evidence in favor, right?
Although I imagine this all sounds good in theory, I would think it would be very vulnerable to p-hacking if you're just doing these natural experiments in the wild, tracking everybody, and looking for effects like this.
P-hacking is not especially something you can control after the fact. Picking your subpopulations post hoc can be a problem. Although, again, this really just becomes a question of power. P-hacking is always a question of power.
Except that you don't know all the possibilities that were considered, right?
Well, you need to control for that, or preregister what you're going to look at.
Let's imagine—let's just take the example of a vaccine, because vaccines' real-world evidence is very unambiguous: You either get the disease or you don't get the disease. If I go out and have a vaccine—I don't know if it's effective—and I just give it to a bunch of people, and then afterward, let's say at a population level it's not effective, which is to say people without the vaccine get the disease at a rate that's indistinguishable from people who get the vaccine, then I'm going to go in and look at, “What about just men? Women? Men above the age of 18 but below the age of 35?” I'll go and look at all the different subpopulations, and inevitably I'll find one where no one in that population got the disease. Therefore, I'm going to say it's 100% effective, and you're going to say, “No, it's just p-hacking,” and you'd be correct.
If I have 10 million patients and across 10 million patients I still find that somehow men between 18 and 35 who get my vaccine never get the disease, whereas all other populations get the disease at the same rate, no, that's not p-hacking anymore. That's obviously not p-hacking. Something is going on; we don't know what it is, but that's obviously not p-hacking.
That's what I mean: It's a question of power. If you get enough real-world data, you can robustly go back and identify these subpopulations in which a drug is effective without worrying about p-hacking. Now, you may still want to do more experiments because you may have no idea why, and you may want to do a confirmatory experiment. But you can imagine that it's better to be able to go back in and at least find those hypotheses than it is to just have the trial fail and the drug—
I mean, how often do we have real-world data at that scale where there's not sort of weird selection bias?
Yeah. Vaccines are often given at population scale. But you could imagine that antidepressants are a great example. You could imagine getting enough real-world data on antidepressants in order to be able to do this.
The more common cancers—I mean, cancer is a specific thing because the treatment is obviously life or death. You need to be more careful. You never want to deny a cancer patient the standard of care, right? But the common cancers, breast cancer, colon cancer, and so on, will have hundreds of thousands or millions of patients.
So I think that, obviously, for a rare disease—
Mm-hmm.
—you won't get a huge amount of real-world data, but also for a rare disease, you don't have that many subpopulations usually.
Mm-hmm.
Right.
Do you think that a model like yours could go through the existing real-world data and find new patterns you haven't seen before?
Absolutely. We have several clients that we're working with aimed at that kind of work. The challenge there is always the quality of the data, which is to say that real-world data is gathered in the real world, and real-world people—as probably many people watching this understand—do not always go to their follow-up appointments, do not always take the medication on the schedule they're supposed to take it on, and doctors do not always take high-quality notes.
Patients move between healthcare systems and then you lose the records; they end up disconnected. For all these reasons, existing real-world evidence is never as good as you want it to be.
There are a number of great companies trying to fix this. One of them is a company called Empower Medicine, which is focused on gathering extremely high-quality electronic health record data from patients so that you can design synthetic clinical trials. There are other companies like Tempus, Komodo, and so on that do this as well, but—
I guess in the single-treatment case, it makes sense that there would be a lot of public research and someone would want to decide once and for all: Is this effective or not?
Yeah.
I've found in my life that when I have medical issues, they feel high stakes, and when I go look at the research and the data, it never seems clear. It often seems like there's conflicting research, and every single data set seems biased in different ways. I've found myself more and more using LLMs to try to synthesize the research into making a sensible decision.
For example, fertility was a big issue for me and my family, and it's expensive, and there are upsides and downsides. I really wanted a model like yours, and I think I asked you to use your model for a recent thing that I was looking at. Do you expect that people might use your model in that way, to look at their individual, specific situation and try to dispassionately synthesize the research into a decision?
My mother does.
No way. Tell me about that. That's cool.
Well, I don't know that my mother wants me to air her medical history on a podcast. But this is totally off the record, man.
I know. Good. I forgot.
Exactly. No, look, I have a bunch of friends who use it. I have one friend who had a breast cancer scare, for example, who wanted to figure out, if she delayed treatment, what that would do to her recovery or survival odds. She had a reason why she wanted to delay treatment for a month or two. Luckily, she turned out to be fine, which is good.
But one of the things that our platform is very good at is going out and dispassionately surveying the evidence and giving you an answer that is directly grounded in hard evidence from papers and from the literature, as opposed to vibes.
All right. Actually, you're a recent father. Have you used these models, or your own model, for your child yet?
Yeah. We have been lucky enough so far that we've not needed to use them for medical reasons. But, 100%, I ask it things like, “When will my baby start to walk and crawl?” and so on. Yeah, yeah, yeah.
I'm sure you will. I actually had a recent, interesting experience with my wife where we both put in the same symptoms and incident involving our daughter. She hit her head, and we both described it identically. The model told me it was fine and told my wife to take the child to the doctor, which I think might have been a different kind of incentive in the model. I do feel like they kind of want to tell you what you want to hear. Maybe some RLHF gives it that. So I wonder if it sort of implicitly knew—
Or, I mean, it has the history. It probably—you know, it definitely does.
Yeah. I mean, like I said before, what we focus on is making sure that our answers are grounded in scientific fact, right? But it can be frustrating. You can ask ChatGPT or Claude whether you should take your child to the doctor, and they will say, “No, the child's fine.” If you ask Cosmos, our agent, whether you should take your child to the doctor, Kosmos will probably come back and say, “Children who hit their heads have an X-percentage chance of developing this serious problem and a Y-percentage chance of doing this.” Then it will give you something like, “On average, there's a 97% chance that if you don't take your child to the doctor, they're going to be fine.” And that can be more frustrating sometimes.
Honestly, that sounds fantastic. Yeah.
But it's a question of how much information you want. It's very interesting to think about. We're entering this era where information and intelligence are just so much more abundant than they were before, but it's not always good that we're going to want that.
Right. More information is not always better for making decisions, right?
Totally true.
Although I don't like to admit that.
I know. But often, you need the right information, but not the most information.
So OpenAI, Anthropic, and DeepMind are obviously working on similar things to you, at least in the sense of deep research, making the models more advanced, and making agents that think more. Do you feel like you need to keep some kind of structural advantage over what they're doing?
I mean, yes. If you're building a company in general, you want to try to have a structural advantage over what other people are doing.
Okay. So what is your structural advantage?
Right, great question. There are a couple of ways to think about this. The most—
By the way, were you yawning through my question?
Yeah.
Is that a bad question?
Yeah. Sorry, man. My whole life, I get asked, “OpenAI and Anthropic are doing this,” and I'm just like, “Oh—”
I'm giving you a softball, man. Let's go.
No, so—
Um, so—
Man, now I'm worried. Maybe we should move on.
No, no, no, no, no, no. It's a great question. I'm trying to tease you. It's a great question.
Fundamentally, at the end of the day, OpenAI, Anthropic, and DeepMind are focused on building AGI, or artificial superintelligence, or whatever, right? I think the key thing to understand is that having a specialized model for a particular task is always going to be better than having a generalist model for that task, even when you have superintelligence or whatever. The reason is that the generalist model has to do everything in the weights, and the specialist model only has to do some things.
Wait, so do you have a specialist model?
Absolutely. We train reasoning models on specific scientific tasks.
I see.
What we see is that with a very small amount of data, you can get enormous gains for those specific tasks over the frontier models. That's the first thing: having specialized agents and specialized models, which also, by the way, comes with specialized user experiences, like I was mentioning. We put way more test-time compute into minimizing hallucinations than the others would.
But presumably, you also use the bigger models, right?
Yeah, absolutely. We use the bigger models in some areas. We don't want to be better than Anthropic at coding, for example, right? But when it comes to very niche biological reasoning, we have our own models internally that are superior.
Doesn't the coding thing actually show an example of a specialized model? These coding models are not really specialized in coding, right? And yet—
They definitely—if your point is, could a specialized coding model beat GPT-5.5 or beat—
You think it seems unlikely?
Well, I think a lot of people are trying to do it, and they haven't yet. Coding is very special because I think all the labs have realized at this point that coding is the thing that they need to be really, really good at, right? Maybe it's not true in coding for that reason. But I can kind of guarantee you that if you want to train a specialized model for reasoning about synthetic chemistry and synthetic pathways, the specialized model is going to do better than the generalist model.
Right now, the place where this may fall down is when you get to intelligence saturation. The thing I tell people about saturation is that if you think about tic-tac-toe, any superintelligence will be exactly as good as humans at playing tic-tac-toe. No amount of intelligence above human intelligence will improve performance at tic-tac-toe, because humans are optimal at tic-tac-toe. You can always—
Some humans—
Some humans, not all humans, but there exist humans who are optimal at tic-tac-toe, right? Similarly, you should imagine that at some point, intelligence will get strong enough that we will have saturated, and more intelligence will not improve your ability to reason about chemistry. I don't know where that point is, but it's feasible that we would get there, and then my argument might fall down, right?
Yeah.
No, at that point the argument might fall down. But that's the first thing. The second thing—but I think we're very far away from that point, right? All of the pharma companies have their own internal datasets, and they need to have proprietary advantages over each other, right? They need models that are trained on their data, because if everyone has the same intelligence, then no one has any advantage in R&D.
Those are the areas where we really shine. The reason right now why all pharma companies—or it feels like all pharma companies—are fighting to do a deal with us, and we're just overwhelmed with demand, is because we have the best models for science, we do the last-mile integration to get those models deployed internally and actually accelerate the pipeline, and we train on their data. That provides them with their sustainable advantage, which is something that, at least today, OpenAI and Anthropic don't want to do because they don't want to have a different model for every customer.
It's kind of interesting. I feel like when I talked to you 6 months or a year ago, you talked more about the agent framework and the evals that you were doing. Have things changed, or is that less exciting to talk about?
No. I think that we have matured, or we are in the process of maturing. The way that we think about what we do and about the value proposition is maturing, but it remains the case that our agents are, I think, the best at science and are substantially further ahead of the offerings that Anthropic, OpenAI, DeepMind, and so on have. I think we'll be able to maintain that edge for a while.
But when you think about the macro dynamics of what the market will look like in 3 years, 4 years, or 5 years, there are a couple of key things happening. The first one is that these companies need models that are trained on their data to maintain their advantage.
The second one is that they don’t want to be locked into a single model provider. They don’t want to be locked into just Anthropic or just OpenAI or whatever, because last year OpenAI was way ahead, and today Anthropic feels like it’s way ahead, and so on. Those dynamics are pushing customers to work with us.
Do you think, when you look at the major model providers, that they’re different enough that using them for specialized use cases is a valuable thing to do, or are they essentially interchangeable?
This is a great question. I think that, at the moment, they’re mostly—they’re spiky, right?
But are the spikes overlapping, or are they—
I think there’s a lot of non-overlapping spikiness.
Oh, interesting.
So—
Give me one example.
Yeah, great. For example, we have a benchmark that we’re building that we’ll release shortly, which involves reproducing analyses that were previously done in scientific papers. On our evaluations as of when we’re recording this—it may change—but at least as of the last data I saw, Anthropic is particularly good at that, significantly more so than OpenAI and significantly more so than DeepMind.
Another good example, which maybe makes the point to an even greater extent, is that there was a time early in the development of Kosmos when there was a specific task that we needed the models to do in the process of updating Cosmos’s world model. At the time, the only model we could find that was able to do that task was Gemini 2.5. We could not figure out how to get any of the other models to do that task, which was wild.
There’s a lot of that spikiness, right? They have slightly different personalities, obviously, so that’s also interesting. We’ve been surprised by the extent to which you can get much better results on tasks by combining models across different providers.
I see. When I fire off Cosmos today, how long does it run for? How many calls is it making?
Yeah, great question. We released the original version of Kosmos back in November. It would run for 6 to 12 hours, write 45,000 lines of code, read 1,500 papers in a single run, and was way more powerful than any agent that anyone had seen. I think it’s still one of the most powerful, most intensive agents out there.
The issue with it was a UX issue. If you’re a scientist doing research, you don’t really want to ask a question, walk away, and come back 6 to 12 hours later. We recently announced a significant update to Cosmos that puts an interactive front end on it, so it can still go and do those extremely long-running tasks, but you can give it feedback throughout. You can steer it. It’s more interactive.
Let’s go to rapid-fire random questions here.
Yeah, let’s do it.
All right. You said—I saw that you said—scientific progress has slowed down a lot. That was different from how I think about scientific progress. What did you mean by that?
That depends on which time you’re talking about when I said this.
Do you still stand by that point that I pulled out of context?
The—okay, it is empirically the case that progress in medicine has slowed down.
What does “empirically” mean?
It means that we have, in medicine, what is called Eroom’s law.
Which is Moore spelled backward.
Exactly. It’s the observation that the amount of money, in real terms, that it costs to develop a new drug has doubled over the past 40 years. I forget what the doubling time is. It might be that it has doubled every 9 years or something like that.
Semiconductor prices have had this exponential drop. Drug prices have had this exponential increase. There are various reasons people have argued for this. That trend actually broke around 2011 or 2012. People think it’s largely because of the Human Genome Project.
Oh, actually, the Human Genome Project made it easier to find drugs.
But now it’s about flat, and certainly productivity is not yet increasing.
Is that because we found all the best drugs, or—
That’s one of the reasons. One of the reasons may be a lack of low-hanging fruit. One of the reasons may be what’s called the “better than the Beatles” problem. If you have a drug that treats condition X—if you have a drug that treats depression, maybe a bad example; whatever, you have a drug that treats colon cancer—you need to come out with a drug that is better than that drug—
Okay?
—in order to gain market share, right? In order for it to get approved.
Even just in order for it to get approved, and then in order for it to be used—to be useful.
Exactly.
Why do you call it “better than the Beatles”?
Just because there’s this observation that, I guess, the Beatles are still extremely popular.
I see.
Right. The Beatles have not been displaced as a band. They were, like, first. Any band that comes afterward—if the Beatles had not existed, maybe there would be some other band. There are still good bands, but they don’t get as much share as the Beatles, because the Beatles took that space in the consciousness.
Basically, all diseases have been cured. So—
No, but this is the thing. They haven’t been cured. It’s just that we can’t find better treatments for them, right? No one is going to go out and argue that depression has been cured or that colon cancer is cured. But if you want your drug to get approved, you have to be better than the best in class.
But, yeah, I think that’s a problem. Of course you need to be better than the best in class, right?
But that just means it’s getting harder, right? It’s getting harder to come up with new things, because you don’t build on previous innovations in medicine in the way you were able to build on previous innovations in semiconductors, right? Why not? Well, because if you have a drug that uses a particular mechanism to cure cancer or fight cancer, you need to come up with a different mechanism. There’s only so much that you can get from juicing that specific mechanism, right?
You can optimize the drug, make it marginally better, and so on, and there are a lot of gains that you get out of that. But it’s often the first-in-class drugs that lead to the dramatic improvements.
I see.
Right.
Interesting. So there’s some new mechanism that we’re using.
Yeah, and so I think science is moving more slowly for that reason. I also just think, as in the case of physics, that as you discover more things, it becomes harder to discover more things.
Right?
Interesting.
All right. Is a PhD or formal science degree still worth pursuing?
Right. Wow. Yeah, good question. I don’t know where to start my answer. Sorry, you can yawn. Your turn to yawn.
Oh my God. The—
Equivocating. I hate it.
The—
I think the answer is yes, because I don’t think human researchers are going anywhere anytime soon. Fundamentally, the point of a PhD is to learn how to do research, right? If you never learn how to do research, you’re definitely not going to be effective at doing research using the—
But aren’t you automating research?
We are, or certainly accelerating it. The more interesting question is whether we’re going to need human scientists.
Yeah, a related question for sure. What do you think?
I think the answer is yes for a substantial amount of time. The reason is that science is nonverifiable. Fundamentally, at the end of the day, we need human scientists for their taste. I’m not sure that it’s going to be—maybe in 20 years. On the 5- to 10-year time scale, I’m not sure that we’re going to get to a point—it’s not clear to me yet whether we’ll get to a point where the models will just have obviously better taste than humans.
8. Will AI ever replace Nobel laureates
Well, it’s interesting, because here it says you said language models will eventually be better than humans at coming up with ideas.
Yeah, but—
You take it that far out, though. I mean, eventually—and “better” exists along many different axes, right? Better can mean more likely to work. I think that’s definitely the case, right? Language models are going to do way better than humans at coming up with ideas that are better, that are more likely to work, right? But when it comes to which of these problems is the most interesting to pursue, it’s just harder because it’s not verifiable.
And so this is just uncertainty, right? Like I said, I’m a scientist. I admit when I don’t know things. I don’t know whether we’re going to get to a point within 5 years where we’re just like, “Oh, you know, there’s no point in asking Frances Arnold, Nobel laureate, what she thinks about evolving proteins because we could just ask the model.” I think it’s a pretty tall bar. When I look today at asking Claude 3.7 what it thinks we should be doing with directed evolution to improve chemistry or to open up new avenues in chemistry, versus asking Frances, there’s no question that Frances is going to have better ideas.
Interesting today, right? I mean, in 2 years, we’re going to be sitting here again and you’re going to be like, “Well, you said that, and now you know…”
In this off-the-record podcast.
All right. I wonder if it’s going to work going forward to tell guests it’s off the record. Chat, the house rules, people. Chat, the house rules.
I was literally just at an event where I was on a panel, and they were like, “It’s off the record.” I was looking around thinking, “What about all the cameras that are pointed at me?” They were just like, “Oh my God, nothing’s off the record. Are you kidding?”
9. Raising $70 million from pharma investors
You’ve raised quite a lot of money. I think I have $70 million in my notes. Is that even accurate?
Yeah, that’s right. $70 million.
And that was just the for-profit; the nonprofit raised more. I think, unlike most of the new labs and the kinds of companies that I come across, most of your investors are actually folks I don’t know. They’re coming more from the pharma world, I think. Why do you think that is?
Well, for us, success is getting inside the pharma companies. The pharma investors are the ones who have that capability. So we get very high value out of all of our investors.
Yasmin Razavi is an investor at Spark Capital who co-led our round, and she is absolutely incredible. She’s on the board of Anthropic. She led their first VC round and has extraordinary perspective and insight. She has really helped us with talent and so on.
When I think about the concrete business traction, the investors who have contributed the most value are Jeff Huber, who was the CEO of Grail and somehow seems to know literally every single person in pharma, and an unnamed institutional biotech investor—a very large, unnamed institutional biotech investor—who would be displeased if I said in this off-the-record forum who they are. But I think most people in biotech will know what that means because they’re very well known, and they have been extremely valuable in terms of setting us up with high-level connections into companies and getting us deployed.
It’s interesting. So the people who actually know your field better are more bullish on you than the maniacs doing AI investment in Silicon Valley.
Yeah. I think the maniacs are very excited about many things, right? They’re very excited about the concept of, “We’re going to automate science.” I mean, everyone is very excited about, “Let’s go automate science. Let’s go automate drug discovery,” right? But the insiders know the problems, so they really viscerally know where the value is. Let me put it this way: I don’t get any of the insiders asking me how we’re differentiated versus Anthropic, OpenAI, and Google.
I see.
Right. I get asked that by every single tech investor—
Right.
—who will ask me why Anthropic, OpenAI, and Google won’t do what we’re doing. But anyone who has operated inside a pharma company, who has large positions in pharma companies and biotechs—we never get that question from them.
Okay, there’s a long history of AI overpromising in drug discovery.
Oh, yeah.
Is this time different?
Yes. I promise you, this time it is different.
Yeah, man, back to 2012, people have been saying that pharma AI is going to revolutionize pharma, and it’s been one flop after another, with the notable exception of AlphaFold. AlphaFold really changed the way that a lot of things are done, but it changed the way a lot of things are done in 1 area of drug discovery and development: the actual process of coming up with the molecule because you have the structure of the proteins.
But I think there’s a ton of skepticism. That said, it is crazy to think that this time will not be different. We have so much evidence already that this time it’s really going to be different.
The thing I do want to emphasize is that the proof is in the pudding, and the pudding is approved drugs. You’re going to hear a lot of stuff in the next couple of years: “My model came up with this drug,” and “My model discovered this fundamental aspect of biology,” and so on. Anyone who has been in pharma is not going to care until they see the outcome of a pivotal Phase 3 clinical trial or an approval letter from the FDA.
So that seems like a great place to end.
Thanks so much for listening to this episode of Gradient Descent. Please stay tuned for future episodes.