AI科学家来了。但科研进展真的在加速吗?| EP 170
- Edison Scientific 的 Kosmos,将合作者所说的3至6个月博士级数据分析,压缩到约12小时的一次运行中。 这一结论来自这样的测试:把未发表研究中的相同目标和数据集交给 Kosmos,它在一夜之间复现了研究人员的发现;Sam Rodriques 起初认为「这不可能是真的」。他说,Kosmos 的结论约有80%的正确率,大致相当于让一名人类独立开展调查研究。
- Kosmos 的经济账看起来很不寻常:200美元的一条提示词,可能消耗1,500篇论文并生成42,000行代码。 Kosmos 通过一个能维持长任务连贯性的「结构化世界模型」,协调 OpenAI、Google、Anthropic 的数百个智能体,以及 Edison 内部训练的模型。Rodriques 称200美元只是推广期定价,并将其与科学家收集这些数据可能已经花掉的5,000至10,000美元相比较。
- Kosmos 已经在产出新发现,但「AI发现」目前仍意味着加速分析,而不是自主得出科学真理。 其论文中的7个结论里,3个复现了已知发现,另外4个被描述为新的科学贡献,其中包括一套拟议机制:将一个非编码2型糖尿病变异、结合蛋白、基因表达与参与胰腺胰岛素分泌的 SIRT1 联系起来。科学家仍必须理解、交叉核验并通过实验验证这些输出。
- 医学领域最大的瓶颈仍是临床试验,而不是缺少合理的研究假设。 Rodriques 认为,假设当下数亿美元规模的试验已经经过最优设计,「那你就是疯了」,因此 AI 可以改进实验和临床试验。但制造临床级材料、招募稀缺患者、给药并等待结果,仍是顽固的物理约束。
- Rodriques 称在10年内治愈所有疾病「很疯狂」,但认为30年的巨大跃进是可信的。 即便一种药物能够彻底阻止衰老,也可能需要5年或10年才能证明有效;而且「拥有无限智能」也不代表所有答案都能凭空获得,新的实验仍不可或缺。更快的生物标志物可能缩短这一时间,因此他承认,激进预测仍有可能超出自己的预期。
- 生成式生物学——而非单纯预测——是当前科学能力跃迁最大的方向。 模型如今可以从零开始提出具备目标特征的抗体、蛋白质或生物体;设想中的工作流是输入一种疾病蛋白,生成与之匹配的抗体,但后续仍要经过制造、验证和人体试验。Rodriques 认为 AlphaFold 3 受到的关注仍不够,实验室自动化的热度基本合理,而「虚拟细胞」、量子计算和脑机接口被过度炒作。
- 科学智能体正处于采用 S 曲线的起点,但普通实验室的实践方式会比模型前沿变化得更慢。 编程助手和文献检索能立即为保守的生物学家创造价值,而更深层的智能体式研究,需要靠已经展示出的结果逐步建立信任。Rodriques 曾认为2026年或2027年由智能体生成大多数高质量假设是过度炒作;如今他称2026年仍属激进预测,但2027年可能真的实现。
1. Kosmos 将数月分析压缩为12小时运行
Rodriques 最初听到「6个月」这一说法时的反应是「这不可能是真的」。Edison Scientific 的测试方式,是把学术研究中相同的目标和数据集交给 Kosmos,而这些研究当时尚未由合作者发表。Kosmos 在一夜之间重新发现了研究人员称耗时3个月、5个月或6个月才得到的结果——这不一定意味着连续6个月不间断工作,而是相当于投入了这么多研究精力。
Kosmos 看起来像一个提示词输入框,但它「不是聊天机器人」:用户提交研究目标,然后等待约12小时。每次运行都会同时调用 OpenAI、Google 和 Anthropic 的模型,以及 Edison 针对具体任务训练的内部模型,而不是押注某一个前沿系统。
真正的技术突破在于一个记录任务知识状态持续演变的「结构化世界模型」。它让数百个智能体可以并行工作、按顺序推进,同时不忘记目标,也不至于「脱轨」,最终产出一项连贯的调查,而不是一堆彼此割裂的模型输出。
工作量解释了定价:一次平均运行会阅读1,500篇论文、编写42,000行代码,而一次 Claude 会话可能只写几百行。当前200美元的收费是推广价,未来会涨价;但相较于收集实验数据需要花费的5,000至10,000美元,用户告诉 Rodriques,他们「简直不敢相信」这项服务只要200美元。
2. 新发现先于证明出现,但仍离不开科学家
Kosmos 团队在论文中给出了7个结论:其中3个是复现结果,另外4个被描述为对科学文献的新贡献。Rodriques 说,系统确实能给出很深的洞见,正确率约为80%,「有点像」让一名人类研究人员独立开展调查。
他最有力的例子,涉及数百万个与疾病相关、但作用机制仍未知的遗传变异。拿到2型糖尿病的原始数据后,Kosmos 将一个位于基因之外的变异,与在其附近结合的蛋白质联系起来,找出相关的基因表达,并进一步连接到 SIRT1 的作用机制;SIRT1 参与胰腺胰岛素分泌。这套机制此前没有任何人类研究人员从这些输入中完整拼接出来。
Casey Newton 的一个有用拆解是:科学研究先选择数据,再收集数据,最后得出结论;Kosmos 目前主要攻克的是第三步。Rodriques 回忆,自己攻读博士期间花了6个月分析数据集,当时年收入约40,000美元。智能体可以迅速浮现研究结果,但科学家首先必须理解这几个月的工作如何被压缩,再重新运行分析、交叉比对并进行实验验证;在他看来,这个具体结果不太可能最终成为药物靶点。
3. 临床现实,而非创意生成,决定医学进度
Casey Newton 的反驳值得保留:当临床试验、患者招募和 FDA 审批耗时远超药物发现时,药物发现可能并不是限制因素。Rodriques 基本同意这一点:科学家已经知道如何在小鼠身上治愈的疾病数量「多到天文数字」,因为在那里开展实验很容易;而人体实验天然更慢。
AI 仍然能在上游发挥作用,因为如果你认为每一项药物试验都已经基于全部可用知识进行了最优设计,「那你就是疯了」。临床试验成本高达数亿美元,而现有数据集里藏着大量此前无人有能力挖掘的洞见。因此,更好的分析应当带来更好的实验和临床试验,即使它无法取消临床试验本身。
Rodriques 对时间表的直白判断是:在10年内治愈所有或大多数疾病「很疯狂」;但30年内实现「巨大的跃进」是可信的,尽管终结衰老或治愈一切都不保证在物理上可行。假设一种药物能让25岁至65岁期间的衰老停止,研究人员可能仍需要5年或10年才能检测出效果。
监管只是延误的一部分。即使完全没有监管机构,团队仍需生产足够的人体临床级材料,通过与医生建立联系来找到稀缺患者,给患者用药,然后等待结果。「GPT-7」也不会简单地解释如何治愈 Alzheimer’s:即使「拥有无限智能」,关于现实世界的缺失事实仍然缺失,直到实验补齐它们。生物标志物可能加快迭代,Rodriques 也承认,这是一条乐观预测可能跑赢他自己判断的合理路径。
4. 更好的科学需要世界模型,也需要有生产力的噪声
Rodriques 将 AI 科学分成两类:一类是对自然世界建模,另一类是对科学研究如何开展建模。Kosmos 属于后者;蛋白质结构预测、抗体生成和生物体设计属于前者。能力变化最大的是生成式设计:根据要求从零产出蛋白质、抗体或生物体,这是他所说的「此前从未拥有过的新能力」。
Kevin Roose 用 Google 回答「2026年不是明年」这一例子,质疑科学分析的可靠性。Rodriques 表示,研究人员确实必须投入大量时间检查智能体,但发表论文本来就要求检查这类工作。无论有没有智能体,完美都无法实现;现实标准是达到与人类相当的可靠性,而「检查工作永远会比一开始就产出工作更快」。
尚未解决的问题是:优化是否会摧毁偶然发现。Rodriques 预计错误仍会存活,并提到孢子飘进一扇敞开的窗户,最终产生了青霉素这一观察结果。研究生一年级学生同样会通过做专家会否决的「最随机、最古怪的事情」推动进展。他的回应是,系统可能需要保留变异,或者「加入噪声」;生物进化也依靠随机性来发现有用的功能。
5. 采用从编程起步,智能体逐步逼近 S 曲线
大多数科学家的工作方式尚未发生根本变化。Rodriques 认为,生物学家偏保守,因为沿用下来的实验流程即使没人完全理解每一步,也往往确实有效,而逐一测试所有替代方案又不可能。因此,大多数实验室会继续使用熟悉的方法,直到看到同行取得可明确验证的更好结果。
编程和文献检索是眼下的例外。随着研究人员使用 Claude Code、OpenAI 模型和 Gemini,即使此前不会编程,生物学领域长期存在的编程瓶颈也正在下降;智能体则可以解析原本难以处理的科学文献。完整的科研智能体仍处于更前沿的位置,扩散速度可能会更慢。
在快速问答环节中,Rodriques 认为 vibe proving 作为终端用途可能被过度炒作,尽管它对推动 AI 很有价值;科学实验室机器人「热度合理」,但技术上仍不成熟;AlphaFold 3 可能反而没有得到足够重视,尽管已经受到大量关注。「虚拟细胞」这个名称被过度炒作,因为当前系统建模的只是狭窄的细胞功能,而不是完整细胞;量子计算和科幻式脑机接口同样被过度炒作,或者距离真正实现比人们想象得更远。
他定义2025年的3项进展是:科学智能体、包括 Chai 和 Nabla Bio 在内的团队开展的从头抗体设计,以及 Arc Institute 从零开始设计噬菌体——后者即便实际用途尚不明确,也「很棒」。他预计智能体将在2026年「渗透进一切」。当年由智能体生成大多数高质量假设仍是激进预测,但他曾认为自己是在过度炒作的那一判断,到2027年可能会成为现实。
I’m Kevin Roose, a tech columnist at the New York Times.
I’m Casey Newton from Platformer.
And this is Hard Fork. This week, FutureHouse CEO Sam Rodriques joins us in the studio to separate the hype from the reality of AI science.
Well, Casey, it’s time for some science.
Yeah, give me a second, Kevin. I’m just going to put on my lab coat here, get out my Bunsen burner, and see what you’ve got cooking for us today.
I have been obsessed with this question of what AI is and isn’t doing for science and scientific discovery. Obviously, this is something we hear a lot about from the leaders of the big AI companies. People like Dario Amodei, Sam Altman, and Demis Hassabis have all been saying things in recent months about how close they believe we are to solving new scientific problems, curing diseases, and fixing the climate with all of these new AI tools that they’re building.
Some of that is obviously hype, or at least has the sort of markings of hype. But there’s actually a lot of real stuff going on in AI and science that I just do not feel personally qualified to evaluate.
Yeah. I would also say that science has become one of the main ways that the leaders of these tech companies want us to evaluate them, because whenever one of their models does something horrible, the message we basically get back in response is, “Don’t worry, we’re about to cure cancer. Just hang on tight. I know that this chatbot might be driving you to madness, but if you could just give us a few more releases, we’re going to do some really good stuff.”
Yes. This is something that we’re also hearing now from the U.S. government. The Genesis Mission was announced by the White House just before Thanksgiving. That is what they’re calling a dedicated, coordinated national effort to unleash a new age of AI, accelerate innovation and discovery, and solve the most challenging problems of this century.
I thought the Genesis Mission was just them trying to get Phil Collins to play the White House Christmas party. [Laughter]
I guess not. So today we have brought in a bona fide scientist to help us understand which of the scientific discoveries and possibilities out there are real and which are not. We need an expert with a broad focus, someone tracking the impact of AI not just on biotech or drug discovery but across the different sciences. And, Casey, we have found the perfect person.
Let’s hear about him.
Sam Rodriques is the co-founder and CEO of FutureHouse and Edison Scientific, which are San Francisco-based organizations. I guess it’s both a nonprofit and a for-profit.
Have I heard that before? [Laughter]
Yes. Come back when he has his board coup.
FutureHouse is the nonprofit. Edison Scientific is the for-profit that spun out of it. I’ve been to their office in Dogpatch. It’s really fun. It feels like a wacky mad scientist lab. They’ve got all these lab machines that I don’t understand, and people running around in lab coats. They’re all talking about AI, and it just feels like a cool place to be.
They are building what Sam calls an AI scientist, which is an AI agent that can do parts of the process of scientific research. Sam is also himself a scientist. He has a Ph.D. in physics from MIT, and before he launched FutureHouse, he spent several years running an applied biotech lab. So he has seen this stuff happening from a couple of different angles.
Today we want to talk to him about what he is up to, but also get his vision of the entire landscape. Tell us what is working, what isn’t, where’s the hype, and where’s the real stuff. Sam has a lot to say about it.
I think it’s fair to say that Sam is on the more optimistic end of the spectrum of beliefs about what AI will do for science. But as you’ll hear in our conversation, he’s more skeptical than some of the most optimistic people who are claiming that we’ll cure all disease in 5 or 10 years.
Yeah, if you’ve been craving a little bit of cold water for the wildest projections, he has some of that to offer you.
So let’s bring him in. Sam Rodriques, welcome to Hard Fork.
Hello. Thank you.
We have brought you here today to be our science expert, our guide to the biggest recent AI-powered breakthroughs that are happening in science. This is an area that I understand in an ambient way is important and that there are big things happening, but neither of us are scientists, although I did make a killer baking-soda volcano in elementary school.
We have so much to talk about today, but before we get into some of the particulars, I want to ask you about the project that you’ve been working on. Last month, the commercial arm of your nonprofit, which is called Edison Scientific, launched a new AI scientist called Kosmos that you say can accomplish work equivalent to 6 months of a Ph.D. or postdoctoral scientist in a single run of the model. Tell us about how Kosmos works and where that 6-month number comes from.
Yeah, exactly. I’ll start out by saying that when I got that 6-month number, my reaction originally was, “There is no way that this is true,” right? We’ve now measured it in a bunch of different ways. I can walk you guys through that.
Just to take a step back: We’ve been working for 2 years on figuring out how to build an AI scientist. The concept here is that there’s so much more science that we can do than we have scientists, right? So how do we scale up science?
The thing that happened with Kosmos that is pretty cool is that Kosmos is the first thing that I think we’ve made that actually feels like an AI scientist when you’re working with it. You go in, give it a research objective, and it goes away and comes back with insights that are really deep and interesting and sometimes wrong, but about 80 percent of the time right. That’s similar to if you ask a human to go away and do something and come back: A similar percentage of the time, it’s right. It’s a new experience working with it, and that’s very exciting.
The 6-month number specifically: We measured this with a bunch of academic collaborators—scientists who had done science previously that they had not published yet. We gave the same research objective and the same data set to Kosmos and asked it to go away and make new discoveries. It came back having found the same things that the researchers had found overnight.
Then you ask the researchers how long it took them to find this in the first place, and they would say 3 months, 5 months, 6 months. That’s where the number comes from. It’s the amount of time that it took them to come up with the finding.
Let me ask you a couple of questions so I can ground myself here. Is this tool a box you type into like the other chatbots? If so, what is powering it? Did you build your own model from scratch? Did you fine-tune another company’s model?
Yeah. It is indeed a box that you type into. You give it a research objective. It’s not a chatbot; it runs for 12 hours or so before eventually coming back to you with its findings.
In terms of how it’s built, we build on top of a bunch of different language models from OpenAI, Google, and Anthropic. In any given run, we use models from all the different providers. We also have our own models for specific tasks that we’ve trained internally, where those models are much better for the specific tasks that we train them on than the models that the frontier providers make.
The key insight in Kosmos is this use of what we call a structured world model. One of the main limitations with AI systems today is that they’re limited in the length and sophistication of the task that they can carry out before they go off the rails—forget what they’re doing and are no longer on task.
What we figured out was a way to have them contribute to a world model that gets built up over time and basically describes the full state of knowledge about the task they’re working on. That means we can orchestrate hundreds of different agents running in parallel and in series, and have them all working toward a coherent goal. That was the real unlock.
Another thing that I found interesting about Kosmos is the cost. This model costs $200 per prompt.
Yeah.
Every time you give it a task, you’re paying $200. Why is it so expensive?
It uses a lot of compute. That’s the fundamental answer: It uses a lot of compute.
Give us a sense of how much.
An individual run from Kosmos will write 42,000 lines of code and read 1,500 research papers on average. If you run Claude, it might write a few hundred lines of code. That gives you some sense of how much compute is going into this.
Have you ever had a scientist whose cat walks across the keyboard and accidentally hits Enter and all of a sudden spends $600?
This is a problem. The thing that you have to understand is that if you are a scientist and you go and do an experiment, you get some data back.
You're going to spend $5,000 or $10,000 gathering that data. What scientists want is the absolute best performance that they can get. Scientists who have used Kosmos generally come back to me and are like, “They can't believe we're only charging $200 for it.”
Right?
And, you know, I will say: $200 right now is a promotional price. We actually have to eventually charge more.
It's going up. So get those prompts in before Christmas. [laughter]
Exactly. But really, if you have to spend thousands of dollars gathering the data, the cost at the end of the day is not the limitation. We do have to be very generous with refunds because people make mistakes. I made a typo, right? It sucks.
So what you just mentioned about the tests that you all ran to figure out how long this thing could run for and how much time it was saving scientists, that's about replicating existing research that's out there. But a lot of what we hear from the people who are running these big AI labs is the possibility that pretty soon AI will start making novel scientific discoveries. It will start doing things that existing scientific methods and processes can't do. How close are we to that?
That's already happening, actually. If you go and read the paper that we put out about Kosmos, we put out 7 conclusions that it had come to. Three of them were replications of existing findings; 4 of them are net-new contributions to the scientific literature, like new discoveries.
And of those, what's the most impressive?
One of the ones that we really like is about the human genome. It contains millions of genetic variants—differences between different people's DNA—that are associated with disease. For the most part, we know that a variant is associated with a disease, but we have no idea why.
We gave Kosmos a bunch of raw data about a huge number of different genetic factors: what the variants are, what proteins bind near the variants, all these kinds of things. We asked it, for type 2 diabetes, to identify a mechanism associated with one of these variants.
It came back and identified a variant that was not in a gene. Kosmos identified that this is actually somewhere where a different protein binds. It was able to identify what protein binds and what gene is being expressed, and connected that to the actual mechanism of the gene SIRT1, which is involved in the pancreas in secreting insulin.
Right.
So in this case, is what I'm hearing that your model was able to do some very fancy reasoning over some existing data and identify something that no other human scientist had gotten around to and might not have for a really long time?
Yeah, that's right.
Okay. I think science generally consists of deciding what data to gather, gathering that data, and then drawing conclusions. At this point, basically, it's step 3 that Kosmos is aimed at, and there's—
You left out step 0, which was getting the Trump administration to unfreeze your funding, [laughter] but everything else was right.
Yeah.
So what happens when you get a discovery like this from Kosmos? Do you have to then go validate it? Do you hand it to a team of researchers who then have to make sure it works? What happens next?
Yeah, absolutely. You have to go and validate it. In the paper, we describe how we went and validated that particular variant.
In general, when people are using it, literally, when you run a Kosmos run, the first thing you have to do is understand what it's telling you, because it has just done something that scientists think is 6 months' worth of work. You're going to sit there for a long time just reading and understanding it.
Once you've read and understood it, then yes, indeed, you're going to run various experiments, do your own analysis, cross-reference things, and try to convince yourself that this is true. Based on what your research objective is, you'll decide the next steps. In this case, I think it's probably unlikely that there's a new drug target from this particular finding, but you could run this on other findings. Eventually, maybe you find a new drug target and start a drug program.
One concern that I've heard people express about models like Kosmos is that this is just not where the roadblocks are. The reason that we don't have more AI-discovered drugs and designed drugs out there curing diseases is not actually because we don't have the research methods to discover them. It's because you have to go to trials, recruit human subjects, get FDA approval—all that stuff takes a lot longer than the actual discovery of the drug.
What problems are models like these helping to solve in our scientific process right now?
I really agree that the bottleneck at the end of the day in solving medicine is basically clinical trials. The easiest way to see this is if you look at the number of diseases that we know how to cure in mice. It's astronomical because, obviously, you can just run experiments, and in humans, things are slow.
That said, if you think that every experiment being run right now by pharmaceutical companies—every clinical trial—is optimally planned and optimally conceived, given the full state of knowledge, you are off your rocker. There's no way. Those experiments cost hundreds of millions of dollars.
The question is: We do have to run clinical trials, so how do we make sure that those experiments are the best experiments we could possibly be running, given all the knowledge and data we have? There's so much data that has insights in it waiting to be found, but we simply don't have people to go and find them. That's ultimately going to feed into better experiments and better trials.
Well, then I'm curious how you see your tool fitting into the workflow of today's scientist. Is it the sort of thing where I've completed my experiments and now I want some help doing some analysis? Is it that I have all these old experiments that I only did a little bit of analysis on, and I'm curious if I can squeeze any more juice out of them? What other ways are you seeing AI being really good right now for a working scientist?
Going back to me in 2019, which is when I was wrapping up my PhD, I had this gigantic data set, and I wanted to graduate because I was a PhD student, which meant that I was making $40,000 a year or something. There were a ton of great opportunities to go out and not be a PhD student anymore.
So I spent 6 months literally just sitting at my desk trying to analyze the data and draw conclusions, reading papers. That's where Kosmos fits in right now. You would just take that data set, give it to Kosmos, and it comes up with a lot of findings.
Right now, you need to go and do a bunch of manual work to validate those findings and so on. Pretty soon, it's going to come with findings and you're going to be like, “Great.”
Sam, I'm curious if you could help give us and our listeners a state of the world of AI science right now. Recently, the White House announced what it's calling the Genesis Mission, which is a federal effort to corral and harness all of these data sets that the federal government is sitting on and use them to do new scientific exploration.
We also have lots of efforts, including yours, and lots of things going on in and around the tech industry and the biotech industry, with people doing AI for materials science. Give us a sense of the lay of the land of what's hot right now in AI science. Where's the effort and money going?
In order to understand the landscape of AI and science, the first thing fundamentally that you have to understand is that AI is about building models. For example, what is a language model? A language model is fundamentally a model of human language.
It just so happens that when you build a model of human language, it learns how to think like a human in some sense, because humans encode their thoughts in language. This is one of the greatest discoveries, certainly in the 21st century, maybe of all time.
Similarly, when we talk about AI and science, what you have to think about is that you are modeling things. That is what AI does. There are 2 fundamental categories: modeling the natural world and modeling the process of doing science.
These things are fundamentally different. The reason to make this distinction is because we are modeling the process of doing science. The other side of the AI-for-science world is building models that can, for example, predict the structure of proteins, generate a new antibody, or create a new organism from scratch—all things that have happened in 2025, where there's a huge amount of momentum.
Of the things happening in the process of modeling the natural world, you mentioned protein folding and novel organisms. What has most excited you as a scientist that you've seen?
It's absolutely what's most exciting right now. I think, without a doubt, it's this trend toward what we call generative models. These are models that can produce examples of proteins or antibodies or whatever that have desired characteristics, basically from scratch. This is a new capability that we have never had before. And it's huge.
I'm curious about the reliability piece as you're running all of these experiments. I saw this going around on social media this week. I reproduced it myself. If you asked Google, “Is 2026 next year?” it said, “No, 2026 is not next year. It is the year after next.”
So, in such a world, Sam, some people might get concerned at the idea that we're now entrusting AI with all of our data analysis. How much time are scientists having to spend going back and essentially rechecking the work of the AIs? And what kind of tax does that place on their work?
Yeah, this is very funny. Look, you have to spend a lot of time going back and checking. But, to be clear, this is true regardless of whether an AI does it or whether you ask a friend to do it. If you're going to publish a paper, you damn well better go back and check it and be sure that you are confident.
And it's never going to be 100%, right? The best you're going to do is get to a place where it is similarly good to if you were doing it yourself, which is not 100% because you're not infallible. And checking the work is always going to be faster than producing it in the first place.
Yeah.
Right. By a lot.
A lot of our biggest scientific breakthroughs in history have come from these kinds of strange accidents, these moments of serendipity. Penicillin starts growing in a Petri dish, or we discover, “Oh, my God, this is great.”
Does AI preserve that kind of serendipity, those kinds of accidents, or does it optimize them away?
Yeah, this is a great question, and the fact of the matter is, we just really don't know yet. This is going to be a really important core question that a lot of people are asking. What's your intuition on that?
I think they probably will—
They probably will—
They probably will preserve it.
My understanding is that, basically, the window was left open on some agar with no antibiotic in it. Obviously, they didn't have antibiotics if this was the discovery of the first one, right? So the window was left open with some agar, and some spores flew onto it and began growing. They observed that the bacteria was inhibited, right?
That's a mistake. Someone screwed up, right? And that mistake led to something fantastic. You will have mistakes. I think that will be preserved.
But in the meantime, scientists should always leave their windows open. You never know what's going to happen.
Seriously, though, when you get first-year graduate students in academia, they have no idea what to do. And that is a huge source of scientific progress, because they just do the most random, kooky stuff that no one who knows anything would ever think to do. And it's actually really important.
You almost want your AI scientist model to hallucinate a little bit, so that it doesn't lose that—
Add noise, right? We talk about this as just adding noise. This is actually important for biological evolution also, right? The genome has a lot of noise, and that's how evolution randomly comes up with new stuff.
There's a protein that's just totally random and doesn't do anything. Then one day, all of a sudden, oops, it does something, and that's great, right?
What do you make of the leaders of the big AI labs, people like Demis and Dario and Sam Altman, who are saying AI is going to allow us to cure all diseases, or most diseases, within the next decade or two?
A decade is crazy. I think—and I'm happy to take a very strong stance on this, because if I'm wrong, it's a great thing, right? If I'm wrong, everyone wins. But a decade is crazy.
Why is it crazy?
Because of the reason we were talking about before: You have to run clinical trials, right? If we had a drug right now that prevented aging—completely halted aging in humans between the ages of 25 and 65 or something—you would not know for 10 years, because you can't detect in humans in that age range whether or not they're aging for at least 5 or 10 years. You don't detect from 1 year to the next that you're aging. So you won't know if the thing is working.
I don't know. Some people at my 10-year high school reunion were already looking pretty bad.
Hate to say it.
I did say 25. Twenty-five. Fair enough. Fair enough.
But, right, we have to conduct experiments. Those experiments will take time. Now, 30 years, I think, is very plausible. We don't know what is going to be possible. We don't know if it's possible to halt aging. We don't know if it's possible to cure all diseases or whatever.
But between now and 30 years from now, I think you should expect to see a humongous leap forward in terms of—
I want to drill in on that a bit, though, because I think some people might hear that and say that this is essentially a regulatory issue, that we just don't have the FDA set up to measure this. I'm curious about the experimental side of it, though, right? Because my understanding is we don't really have enough biologists to run all the experiments that we might. We might not have the funding to fund the experiments. And you did raise the point that some of these experiments actually take a long time to run, right?
So what are all of the factors that, in your mind, are just going to make it so hard to—
You have to go and, even supposing you have a molecule that you want to test in a human and you know which humans you want to test it in, you have to go and make it, right? Humans are big. They require a lot of it. You have to make sure it's high enough grade that you can actually put it into a human.
You have to find the patients, which means forming relationships with doctors and actually waiting until you have enough patients who are willing to do it. For many diseases, there just aren't that many patients, and so finding the patients is hard, right? Then you actually have to dose them. You have to wait and see what happens, right?
Even with no regulation, it would be slow.
There's no AI shortcut for almost any of that, at least not right now.
No. What AI will allow us to do is discover a lot of things where we already have the information to discover them. We just haven't figured that out yet.
The other thing that AI researchers sometimes talk about, which is probably not reasonable, is that you should not expect that you're one day going to get GPT-7 and just ask it how to cure Alzheimer's and it will just tell you.
My expectation is that there is not enough knowledge, right? We do not have enough knowledge to solve it in principle, even with infinite intelligence. With infinite intelligence, there would still be some things that are just not known about the world, where we have to conduct the experiments to see.
You'll be able to plan the best possible experiment given everything that's known. But you will not just be able to de novo figure it out, right?
Casey, I took Latin. That means “from new.”
Oh, thank you. Thank you. That's saved me a step. This isn't quite science per se, but I'm curious what you make of this, Sam. All of the big AI labs are obsessed with math.
With winning the International Math Olympiad, with putting up a gold-medal score, with solving these unproven math theorems. And I have a take about this, which is that I believe this is because these labs are filled with people who were themselves competitive math athletes in high school, took part in the IMO, and did pretty well.
And a lot of those people think that AGI will just sort of be a slightly smarter version of them. But I'm curious: Why are these places so obsessed with math as being one of the first places that they want to make a lot of progress?
There are 2 reasons. I think one of the reasons is exactly what you just said. It's just familiar, right? But the other reason is that you can measure progress, right?
Ultimately, what drives progress in machine learning—a big part of what drives progress—is benchmarks. With math, you can tell whether or not your proof is right, and there's kind of an infinite number of things to go improve. So it's just really easy to tell whether or not you're getting better. Things like the IMO just present great opportunities.
By contrast, if you look at some of the biggest breakthroughs recently—the biggest breakthroughs this year in AI for biology—things like Chai Discovery and Nabla Bio coming up with these extremely good models for producing antibodies de novo, right? Huge breakthrough, but ultimately the win for them is going to be when it's approved in a human, and that might be another 5 years or something, right?
Arc Institute putting out, like, the first time anyone has designed an organism from scratch.
They designed a bacteriophage. It’s a kind of virus that infects bacteria. Incredible, right? But it’s just harder to evaluate: How good is it? You’re not going to release it into the wild, and so on. It’s harder to evaluate, whereas the IMO is just super clean. And so I think that’s one thing that we think about a lot: How do we get really clear benchmarks that we can pursue to measure whether or not we’re doing a good job at science?
I have an answer here: International Cancer Curing Olympiad.
I like that.
Should we start this? We can give people a medal if they win.
Let’s get on it, labs.
So when the CEOs or the leaders of these companies make statements about how we’re going to cure all disease using AI in the next 10 or 15 years, or whatever timeline they give, are they doing that because they don’t understand the bottlenecks? These are very smart people. So what are they not seeing, or are they just doing this as a marketing exercise? Is this an attempt to get people excited about AI who might otherwise be freaked out about it? Why are they giving these projections?
No, look, I think reasonable people could disagree. There are lots of reasons why you could argue that the models will get super smart and figure out ways to measure whether or not we’re making progress before you run a clinical trial, and that will increase the iteration cycle. There are reasonable arguments to be made about that: that we’re just not going to do full clinical trials anymore. We’ll just use biomarkers. That’s not crazy, and that’s one way that I could be wrong and maybe in 10 years we do have cures for all diseases.
That’s part of it. Obviously, there’s part of it where they want to hype the thing. Part of it is, does Sam Altman really intimately understand what it takes to go and manufacture—scale up manufacturing for—a small molecule to put into the clinic? Probably not. So there’s a mixture. I don’t think any of it’s in bad faith. It’s just that people are very excited.
There will be a little bit of a collision with reality at some point. We’re going to see exactly where that is. But regardless, the future is going to be awesome, right?
At this moment in 2025, how much do you think AI tools have changed the life of a working scientist? And how different do you expect that will be a year from now?
I think you’d be shocked by the extent to which they have not yet. Scientists in general are extremely conservative people because you never know. If you’re running an experiment, you never actually fully know what is in biology, at least. You usually do not fully understand why the experiment works and why it doesn’t.
There are some things that you’ve inherited from protocols that you’ve run in the past, where it’s like, “We do it this way.” You could go and test it, but there are way too many things to test. So you’re just kind of locked in on your methods. It’s what works, and you just want to do what works. For that reason, biologists just adopt new methods slowly.
I think most labs around the world are still probably doing science the way they’ve done it before and probably will continue to do so for a while. And that’s okay. One place where I think a lot of people are already adopting it is coding, because historically, coding has been a big bottleneck in biology. It’s a huge unlock now that biologists who didn’t know how to code can do a lot of coding using Claude Code, OpenAI’s models, Gemini, and so on. That’s a huge unlock, and I think that’s going to see a lot of adoption quickly.
Literature search is another one. Being able to parse the immensity of the scientific literature is a huge unlock. That’s going to get adopted very quickly. The tools like what we’re building are a little more frontier. We also build literature-search agents that people use, but in general, the tools are a little more frontier. It might take a little longer to adopt, but that’s okay. Ultimately, people adopt them when they see other people using them and getting great results.
Sam, can we play a little lightning-round game here with you? We’re calling this one “Overhyped, Underhyped.” We’ll tell you something, and you tell us whether, in your scientific opinion, it’s overhyped or underhyped. Ready?
Yeah.
Vibe proving. This is when AI systems go out and write math proofs.
Probably, if I had to choose, overhyped. It’s great as a progress driver in AI, and being good at it will probably have implications elsewhere. But is it itself that useful? I’m not sure.
Robotics for AI lab automation.
Robotics for automating AI labs, or—
Yes, or for automating scientific labs.
Robotics for automating scientific labs: I think appropriately hyped. It is going to be totally transformative. The technology is not at all there yet. There’s a lot that we need to do, but, yeah, probably appropriately hyped.
AlphaFold 3.
I think that’s an interesting one. I would say probably underhyped, in that I think all of the protein-structure models—there’s a lot of hype around them—but they’re still probably going to be extremely transformative. So maybe I would say probably underhyped. It’s hard. There’s a lot of hype around it, though, so it’s a hard decision to make.
Virtual cells, like we heard from Patrick Collison this summer about what the Arc Institute has done in making a virtual cell.
This is overhyped, but for a specific reason. The models that they’re building at Arc are awesome. They’re doing similar things at NewLimit, Chan Zuckerberg, and many other great companies and organizations. I think calling it a virtual cell is a little overhyped, right? Ultimately, that kind of model models something very specific. Actually building a true virtual cell—being able to simulate a cell in a computer—is an amazing goal. We are very far away from that.
Quantum computing.
Overhyped.
Brain-computer interfaces.
I’m also—oh, man, this one’s really hard. I’m going to say overhyped. I’m a huge believer in BCIs. I think effective BCIs, the way that we imagine them in science fiction, are further out than people imagine. Even Neuralink is making amazing progress.
Yeah, Casey’s got one in his head right now.
It’s on the fritz.
So we’re nearing the end of the year. If we can put you in a bit of a reflective mode, what do you think were the top 3 AI-driven scientific advancements this year?
I think the first one is—this year has been the year of agents. This was the year when people discovered agents. In good faith, I have to put us on that list, also with Google’s AI co-scientist. We’re not the only people who are working on this. Google has been doing a great job, and there are a bunch of other people.
AI agents for science, definitely. And then generative design is just having a huge moment. The other ones would probably be the work that Chai has been doing, the work that Nabla Bio has been doing, and many others on de novo antibody design.
I’m really glad you defined de novo earlier in the broadcast, by the way. It’s come up a lot.
Yes. Sorry. When I say de novo, I just mean that it literally generates it from scratch. You don’t give it anything, or you give it a target that you want it to bind to, and it generates it from scratch. This is huge because the promise that companies like Chai, Nabla, and so on are going after is a world in which you can say, “We know to cure this disease, we have to target that protein.” You click a button, and you have an antibody that you can go and put in humans tomorrow, right? That’s huge.
Or, as I was mentioning before, manufacturing is a pain in the ass. You have to go and manufacture it. But it’s huge. It cuts out an enormous amount of what people had to do previously. So that’s a huge one.
And the third one, I just think, is what Brian Hie, Patrick Hsu, and so on at the Arc Institute have done with generating organisms from scratch.
We know what it means now.
This is our Pee-wee’s Playhouse word of the week this week.
The de novo design of organisms. Is it useful? I don’t know. Is it awesome? Absolutely. It’s such a big breakthrough.
And Sam, what should we be watching for next year? What are you excited about that may be coming down the pipe for 2026?
Honestly, you’re going to see an explosion in agents. Again, it’s going to be the agents that see an explosion. We are right now at the beginning of that S-curve, and that is going to continue, right? I was telling people back maybe a year ago that I thought in 2026, or maybe 2027, the majority of the high-quality hypotheses generated by the scientific community would be generated by us or by agents that are like the ones we’re building.
When I said it in 2024, I thought I was overhyping it, right? But I was just like, it needs some hype. At this point, it may be real.
I think 2026 would be ambitious for that. That's a huge leap—for the majority of the good hypotheses that come out to be made by agents, that's a huge leap—but 2027? Yeah, man. I mean, 2026 is going to be the year when we see these agents start to infiltrate everything: infiltrate labs, infiltrate people's normal lives. I mean, it's already happening.
Cool, yeah.
Well, I look forward to it. Sam, thank you so much for giving us the science education that we clearly didn't get in school.
Yeah, you've really given us some de novo things to think about. I appreciate that.
Good. Thank you guys.
Thank you.