METR 的 Joel Becker:指数级时间跨度评测、威胁模型与 AI 生产力的边界
METR 的主曲线显示,AI 的能力沿着一条异常笔直的趋势上升,但其“时间跨度”衡量的是人类任务难度,而不是 agent 的运行时长。 在 50% 可靠率下,评测套件覆盖从微小的 SWA 操作,到需要人类 20–30 小时完成的 HCAST 工作和 RE-Bench 研究工程任务。这里有一个重要限制:任务边界清晰,通常可以自动评分,大多不涉及视觉,也剥离了组织内部的默会语境,因此这张图不能作为所有现实世界工作的代理指标。
Opus 4.5 的跃升幅度足以改变资深工程师的行为,也对 Becker 更偏好的 7 个月翻倍趋势构成挑战。 他观察到,开发者从抵触代码助手,变成“几乎一行代码都不写”;这一版本的结果符合更快的 4 个月曲线,却“以某种方式”证伪了他的 7 个月趋势线。Becker 仍认为,单次发布可能反映的是任务分布噪声;更有信息量的信号,是 1 到 3 年的进展。
最初关于 AI 让开发者变慢的发现,不能直接套用到今天的工作流上。 如今开发者抵触被随机分配到禁用 AI 的工作,形成选择偏差;同时,他们越来越多地并行处理多个问题或工作流。表观生产力也可能被高估:AI 让一些项目得以启动,其反事实加速“可能是无限的”,但这些项目往往此前被搁置,是因为它们对 Becker 的价值更低;而组织也未必能吸收 10 倍的产出。
Becker 所担心的阶段跃迁,不是基准测试成绩本身有多亮眼,而是 AI 研发闭环是否完全闭合。 “自动化 90% 还不够”:剩下的尾部可能包括故障 GPU、数据中心运维、冷却基础设施、芯片设计或芯片生产。METR 关于 GPT-5 和 GPT-5.1 的报告认为,这些模型的能力还不足以造成灾难性伤害;但 Becker 表示,完整闭环的自动化可能创造出能力失控式爆发的条件。
算力既是显性的加速器,也可能成为能力增长的刹车。 如果算法发现本身需要昂贵实验,算力增速放缓就会产生双重冲击:原始扩展减少,算法进展也减少,从而可能推迟重大里程碑。这一判断仍有条件:有些想法几乎不需要算力,AI 劳动力可能加速研究,而且讨论中提到,基于 OpenAI 的数据几乎无法看清 Meta、xAI 和 DeepMind 的支出。
单一排行榜分数掩盖了那个可能决定自主性是否奏效的“秘密第 11 件事”。 Becker 希望看到更多开放式评测、agent 操作记录,以及代码是否真的会合并进主分支的测试,而不只是通过单元测试。主持人指出,测试 harness 可以让表现上下移动约 10 个百分点,因此脚手架在今天很有价值,尽管“跨模型代际来看,它就没那么有价值了”。
Becker 在 Manifold 的胜利说明,预测市场定价的可能是能动性乃至特权信息,而不只是纯粹的预测能力。 他向慈善机构捐出约 5,000 美元,推动一个捐赠市场越过结果边界,赢得虚拟货币,并成为其中最赚钱的交易者;与此同时,一个 AI 模型市场的交易量达到 2,800 万美元,尽管内部人士可能提前知道答案。他更新后的看法更为谨慎:经过校准的概率有价值,但“赌博式行为会给社会带来成本”。
1. METR 将能力测量连接到具体的灾难性风险案例
Becker 将 METR 定义为模型评测与威胁研究:衡量系统今天和明天可能达到什么水平、部署后实际会做什么,以及这些能力和行为倾向是否会连接到“对社会造成巨大乃至灾难性的风险”。
METR 关于 GPT-5 的报告,以及随后对 GPT-5.1 开展的类似工作,结论都是这些模型不构成此类大规模风险。Becker 的判断基于能力:尽管模型的基准测试成绩亮眼、日常用途广泛,METR 认为它们还没有能力执行所需的伤害行为。
该组织已将重点从自主复制——AI 获取资源并独立完成自我部署——转向实验室内部的研发加速。这一情景可能带来能力爆发,并造成系统性失稳。
Swyx 强调,METR 是独立筹资的组织,而不是由实验室出资的评测机构。Becker 表示,METR 脱胎于 ARC;如果没有独立的专业知识来源,他可以“永远敲这面鼓”,却无法改善社会获得的信息。
2. 时间跨度衡量任务难度,而不是 agent 能存活多久
这条曲线最初只是 2023 年一页零散的内部幻灯片:一条轴是能力,另一条轴是时间、算力或其他资源。METR 将能力操作化为模型在 50% 可靠率下能够完成的人类任务时长后,经验曲线变得“异常笔直”,远比最初的直觉草图规整。
Becker 强调,市场反复出现的误读是:5 小时的时间跨度不意味着模型能够连续 5 小时高效工作,而是意味着模型可以可靠解决一个人类大约需要 5 小时完成的任务,即便模型自己只花了“0 分钟或 5 分钟”。
约 170 项任务的分布,从 SWA 风格的原子操作——例如“名为 passwords.txt 的文件可能包含密码”——延伸到需要人类 20–30 小时完成的 HCAST 任务,最后是 RE-Bench 极具挑战性的全新机器学习研究工程任务。
任务选择限制了结论的外推范围。METR 偏好经济价值高、可扩展、通常能够自动评分,且熟练的“低语境人类”仅凭提供的信息就能完成的任务;视觉密集、开放式、需要外部交互或依赖默会语境的工作代表性不足。现实工作要“混乱”得多。
3. Opus 4.5 挑战了一条趋势线,却强化了更长期的趋势
Becker 称 Opus 4.5 无论在基准测试还是实际使用中,都是一次有意义的跃升。他认识的一些能力很强的工程师,已经从有选择地使用或抵触 AI 编程,转变为“几乎一行代码都不写”;这是一种不连续的行为变化,他自己也有切身体会。
Opus 4.5 符合 METR 发布相关研究时提出的较快的 4 个月时间跨度翻倍线,却不符合 Becker 更偏好的 7 个月线。他承认,这“以某种方式证伪了我的趋势线”,但仍无法确定偏离是源于潜在能力,还是因为这一版本恰好适配 METR 的任务分布。
主持人质疑 Claude Code 运行 5 小时或 30 小时的轶闻:它可能有很长时间都在“做彻头彻尾的蠢事”,一次成功运行也可能无法重复。Becker 的回答是把轶闻视为证据,但对单个版本的权重低于 1 年和 3 年趋势。
4. AI 生产力已经超出了原始随机对照试验的简洁设计
METR 正在重做开发者生产力研究,但 Becker 不愿透露结果。复制实验变得更难,因为预期 AI 会带来明显收益的开发者越来越不愿被随机分配到“禁用 AI”的组别,研究人员只能得到那些本来就怀疑 AI 帮助不大的任务。
自 2025 年 3 月前后以来,工作流也发生了变化。如今开发者会并行推进多个问题或工作线,而早期研究的参与者基本上一次提供所有问题,并行程度更低;随机化一个孤立任务,已经无法捕捉实际的编码方式。
Becker 区分了在同一任务上的提速与扩展到新工作的能力。他的一些副项目如果没有 AI 根本不会存在,因此表观加速“可能是无限的”;但这些项目的价值并非无限,它们对他的价值较低,这也解释了他此前为什么没有获得完成这些项目所需的专业知识。
Alessio 补充了组织层面的约束:即使 AWS 工程师的速度提高 10 倍,客户也无法吸收新增的 50,000 项服务。Becker 同意,乐观的自我估计很容易被高估,但强调 AI 公司的人可能仍然获得了显著加速。
5. 端到端研发自动化是 Becker 判断阶段跃迁的门槛
Becker 当前的安全判断结合了正式评测与实际观察。模型在操作记录中仍显得“有点笨拙”,会浪费资源,也会犯明显错误;而略弱一些的系统已经被大规模部署 6 个月,却没有造成非同寻常的危险,因此根据此前证据,认为一次小幅增量改进就带来灾难,会令人意外。
Swyx 反驳称,涌现能力可能发生不连续的融合,从而削弱“n−1 版本没事”的论证。Becker 承认,考虑到前沿版本数量很少,连续性判断“很脆弱”;但他表示,迄今为止进展呈现出的惊人规律性,仍让他相信进步可能继续保持连续。
真正令他担心的断点,是实验室内部实现完全自动化的 AI 研发。即使测得的时间跨度达到 1 年,含义仍然不明确,因为“自动化 90% 还不够”;某些尚未测量的最后 10%,可能足以阻止整个反馈回路闭合。
这个闭环可能是纯软件的:在硬件不变的情况下,更好的模型递归地产生更好的模型;也可能需要芯片设计,最终甚至需要芯片生产。Becker 无法排除闭环一旦形成就发生爆发的可能性:“谁知道那之后会发生什么。”
6. 研究基准只能捕捉自动化闭环的一部分
论文复现或原创研究基准确实衡量了 AI 研发的一项真实组成部分,但 Becker 的“反方观点”是,它们遗漏了漫长的运营尾部。GPU 会故障,需要有人维修数据中心;冷却系统出问题时,也需要有人联系供水公司——这些能力不会体现在论文分数中。
Swyx 主张建立一个公开的“车轮图”,列出约 10 项关键能力,而不是把一切压缩成一个数字。Becker 同意这样做会损失维度信息,但他预计,任何清单都会暴露出一个“秘密第 11 件事”:事前很难界定,只有在它成为瓶颈后才显而易见。
主持人建议像安全社区列出十大风险那样,每年更新这份清单。Becker 接受了这一挑战,但保留了更深层的不确定性:METR 目前测量的可能只是实现完全自动化所需能力中的一小部分,因此整个闭环的闭合时间,可能晚于基准外推所暗示的时间。
7. 算力放缓可能通过算法发现产生复合冲击
Becker 的模型首先假设时间跨度趋势应按字面理解,然后追问什么因素可能让它向下弯曲。算力是一个显而易见的候选因素:如果能力直接依赖算力,而算法进展也依赖高算力实验,那么增长放缓就会同时削弱这两个部分。
Transformer、RLHF、学习率调度等改进,并不是投入一阵劳动力就会偶然出现的成果;研究人员通常需要规模来显现算法收益,还需要“大量实验”才能发现算法。在算力是瓶颈这一强假设下,算力增速减半可能大致令时间跨度增速减半,并显著推迟里程碑。
他强调了其中的限制条件。有些创新几乎不需要算力,研究人员可能在前沿训练开始前就提出好想法,AI 劳动力也可能在没有完全能力爆发的情况下加速算法研究。结论取决于算法进展在多大程度上“基本由算力决定”。
METR 使用了 OpenAI 过去的纳税申报资料,以及 The Information 报道的未来研发算力预测,再将美元支出换算回 FLOPs。Swyx 提醒称,Meta、xAI 和 DeepMind 的支出大多不可见,而实验室失败、行业整合、蒸馏和重复使用算力,也让行业层面的图景远不够干净。
8. 预测市场的准确性可以由参与者制造出来
当被问及对前沿模型的朴素先验判断时,Becker 估计 2025 年时间跨度领先者的概率大致为:xAI 5%、OpenAI 50%、Anthropic 45%——“如果我完全错了,别朝我开枪”;他同时指出,换一项基准测试,分布就会不同。
他在 Manifold 排行榜上的第 1 名,主要来自一个慈善市场。他看到捐赠额按线性方式预测,于是用虚拟货币买入更高档位,之后进行真实捐赠,把结果推过边界;他重复了这一操作,随后尝试虚张声势但失败。最终捐出约 5,000 美元,仍然带来了足够的收益,让他登上排行榜首位。
主持人将其称为“高能动性的预测市场”。另一个最佳 AI 模型市场的交易量已达到 2,800 万美元,尽管员工可能知道尚未发布的基准结果。Becker 表示 Meta 员工不应参与;讨论还提出,内幕信息可能本身就是价格发现的来源。
Becker 对预测市场的社会价值已经不如过去确信。战争等重大事件的可靠概率判断确实有用,但散户亏损和“赌博式行为”会带来成本;如果市场由成熟交易者持续与散户对手盘交易,情况会比大型机构彼此交易更加令人担忧。
9. 更好的评测会观察 agent 在开放世界中如何失败
Becker 强调了 AI Village 风格任务的重要性,例如开一家商品商店、组织一次公园活动,或搭建一项人类受试者实验。新旧模型、对视觉的依赖和不受控条件都会增加推断难度,但开放式环境能揭示 agent 如何“当场摔个大跟头”。
他理想中的测试会提供广泛的操作权限,然后只下达一句指令:“自动化研发,开始。”他预计现有系统会因为资源使用不佳和长周期执行能力不足而失败,这正说明,在一个细致的软件问题上表现出色,与自动化一个组织的研发流程“完全是两回事”。
Agent 操作记录是另一个规模巨大但经过筛选的数据集。它们展现了迭代式行动、输出、令人印象深刻的行为,以及可能的偏好颠覆;但用户自然会选择有一定成功概率的任务,因此,一份醒目的不安全操作记录本身,并不能说明该行为的基础发生率。
基准测试成功,也应与代码是否真的会合并进主分支区分开来:它是否增加了测试、遵循代码库模式,并完成正确集成?METR 会在开发任务上调校 harness,再用留出任务进行评估,以限制过拟合;脚手架“还有很大的增益空间”,但在同一代模型内可能很有价值,到了下一代则可能被冲淡。
10. METR 的 2026 路线图将能力证据与安全保障结合起来
Becker 预计,未来会有更多时间跨度和生产力风格的证据,也会开展监测研究,判断安全保障措施能否成功应用于尝试执行危险任务的模型。当前工作总体上仍是黑盒评测,而非可解释性研究;到 2026 年,这些能力、行为倾向和监测证据将被纳入更完整的风险评估。
他没有给出具体的 2030 年成功指标,只将 METR 描述为一个资源精简但富有战斗力的环境,由有才华的人从事前沿科学研究。
So METR stands for M-E-T-R. The first 2 letters are model evaluation. That is, we think about what the capabilities of AI models might look like today and tomorrow, as well as their propensities—what they’ll actually do in the wild, given that they have some level of capability. Threat research is the final 2 letters. We try to connect those capabilities and propensities to particular threat models that we have in order to determine whether AI models pose enormous or catastrophic risks to society.
Alessio Fanelli
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, founder of Kernel Labs, and I’m joined by Swyx, editor of Latent Space.
swyx
Hello, hello. We’re back in the studio with Joel Becker from METR. Welcome.
Thank you very much, guys. It’s a great pleasure to be here.
So, Joel, your work has impacted the AI field a lot, especially over the last year. I invited you for the AI Engineer Summit. Thank you for speaking as well and doing the extra workshop. You have a lot of papers that have been very impactful, but upfront, a lot of people feel like METR just burst onto the scene. Could you explain and introduce METR?
1. METR Connects Capabilities To Risks
Yes. So METR stands for M-E-T-R. The first 2 letters are model evaluation. That is, we think about what the capabilities of AI models might look like today and tomorrow, as well as their propensities—what they’ll actually do in the wild, given that they have some level of capability. Threat research is the final 2 letters. We try to connect those capabilities and propensities to particular threat models that we have in order to determine whether AI models pose enormous or catastrophic risks to society.
Yeah. Would you say that you’ve done a lot more ME, and TR is the next phase, or is there a TR side of the work that I’ve missed?
I think there’s some TR. Some of the most publicized work does look more like the ME side: this time horizon stuff and the developer productivity RCTs, stuff like that. But there’s 1 full report on our website, the GPT-5 report, and an analogous one for GPT-5.1 as well, trying to make a more structured case that it doesn’t pose these really large-scale risks, eventually coming to the conclusion that it doesn’t.
But it’s worth thinking: Why exactly is that the case? If you and I work with GPT-5, it does seem very capable. That matches up to benchmark scores. Why is it not able to do something really enormously wrong? We go through the evidence. We think it’s not capable enough, on the basis of some of this capabilities evidence that you’ve alluded to, to commit these catastrophic harms, and it’s not going to be able to do this.
But perhaps in the future, we’ll think it’s capable of doing pretty extraordinary things—the kinds of things that would be necessary to provide really serious threats. Then maybe you’d lean more on the propensities part. Are the protections that we have against these dangerous capabilities sufficient for it not to pose an existential threat? That sort of thing.
So I think threat research very much is there, and very much is something that we’re aspiring towards. In some ways, you might see the capabilities evidence as a kind of input to that.
Yeah.
Alessio Fanelli
Have the threat models been updated a lot? Or do you feel like you’re still using the same threat models as GPT-2’s paperclip factory? How much are you increasing the bar?
Yeah. So I’m not an expert in the threat-modeling piece; I’m more focused on the capabilities piece. I do think they’ve been changing to some extent. Something like the autonomous replication threat model—that is, being able to set yourself up and control resources—has been deprioritized relative to R&D acceleration.
The possibility is that there could be some capabilities explosion inside a lab, and that could be destabilizing for all sorts of reasons that we could talk about. So, mainly, we’re focusing on that latter one, although we do think about a number of threat models.
Yeah. Let’s talk about the ME side. I would say the model time horizon chart is probably the most quoted, both in investment decks that I see and generally on Twitter. What was the origin story of it, and what other color would you give on it to introduce it to the audience?
2. The Time Horizon Chart
Yeah. So there are a couple of different ways to tell this story. One way is that there’s this internal METR PowerPoint from 2023. We were trying to lay out our ambitions for what METR research might look like in the future. There’s this graph. It has a Y-axis that’s some measure of autonomous capabilities or dangerous capabilities or something like that, and then an X-axis that’s labeled time or compute or whatever resources we want the Y-axis to vary over.
Then it has a bunch of scattered points that kind of go off to the right. We think capabilities are improving over time. Many of METR’s research bets have been trying to make this ever more concrete.
And then when we actually did the full thing, when we had something like this Y-axis—which turned out to be task difficulty, as measured by the length of time it takes humans to do the tasks that models can complete with 50% reliability—when we actually got that data and plotted it over time, it turned out to be remarkably straight, the straightest line you’re aware of from a familiar graph.
Part of what makes it so extraordinary is that this pattern does seem to be so regular. In fact, it’s just way straighter than this incredibly scattered graph that we had at the beginning, before I joined METR.
How did you pick the tasks? I would say that’s one question that people have. You have some labels, like “train classifier,” “fix bugs,” and “small Python library.” They all seem arbitrary. What’s the process of task selection?
3. How METR Picks Tasks
People are right to be worried about task selection; there are many finicky details in here. I would say the aspiration was to pick economically valuable tasks relevant especially to general autonomy and R&D, the threat models that we’re primarily interested in.
One misreading of the time horizon graph is that this is referring to the full distribution of any tasks that you might give AIs, and I think that’s clearly not right. In particular, for tasks that require vision capabilities, models are probably much less capable today, as measured by time horizon, than they are on the tasks we typically give them that do not require vision capabilities.
We try to sample these tasks by having people inside METR create them and by offering a bounty so that people from outside METR can provide us with tasks, stuff like this. That’s not a perfectly random selection process. In particular, it’s a process with a bunch of constraints in order to be able to run our evals scalably.
It’s helpful—not necessary, but helpful—for success on the tasks if they can be automatically graded. That means some types of tasks are included, while tasks that are harder to make automatically gradable are not included. But this is the aspiration.
Yeah. The computer vision point was interesting. Any other disqualifiers, so to speak? What are other things where you would expect the chart to be a lot worse?
One thing is fairness: we want tasks to be, in principle, completable by a model that has access to sufficient information. They shouldn’t be impossible given the information the model has.
The way we think about that is: could a low-context human who was sufficiently skilled at the general skills, but maybe not the particulars in the background, achieve success on this task? I think that rules out a lot of real work, because a lot of real work involves people having careful mental models of the situation that are not all fully listed in an issue description or the equivalent of that.
In some ways, you might think of us as not measuring things like that. Another thing is that our tasks tend not to be very open-ended, although they vary a little bit, or involve interacting with the outside world, or be this sort of messy thing, as we call it internally, which refers to a bunch of different things.
But you broadly get the picture from the descriptor “messy.” Relative to tasks that you might find in the real world, our tasks are somewhat nicely scoped. They’re quite neatly contained. Indeed, I think we’re going to talk about some of the developer productivity stuff later. Some of the interesting findings make more sense in light of the fact that those tasks are a lot messier than METR’s tasks.
swyx
Are there any that you would want to highlight in terms of task distribution? I think I’ve come across RE-Bench before. You have a particular affinity for RE-Bench. I don’t know if you want to introduce your side project, the RE-Bench Warmers.
Yeah.
And your side project, the RE-Bench Warmers.
I have a soccer team called the RE-Bench Warmers. We are the most enthusiastic and possibly least technically skilled soccer team in San Francisco. We made the playoffs last season.
Ooh.
Shout-out to the team for that. We’re certainly going to make the playoffs again this season, but possibly by the time this podcast is out, we’ll find out that we have not made the playoffs.
That’s the plan.
Is this the same league that you’re in?
Alessio Fanelli
Same organizer, but different field. We play Mission Bay. We play Palomar.
swyx
Okay. All right.
Alessio Fanelli
Okay.
swyx
HCAST was the first time you came across it and the others. SWA, are those the METR proprietary ones? And then, anything else that you’re considering adding?
Yeah. So there are private tasks in HCAST as well. But yeah, SWA is this list of atomic tasks, or these very small software actions. Maybe one example is: here’s a list of 4 files. One of them contains the passwords. One of them is called passwords.txt. Which file most likely contains the passwords?
I think GPT-2 can sometimes do that task and sometimes not. Claude Opus 4.5, I’m sure, can do that task 100% of the time. Then we go up to HCAST tasks, which span from only a little harder than those SWA tasks all the way up to something like 20 or 30 hours, which require more autonomy and more sequential actions. Many of them are much more challenging. Perhaps, in some sense, they’re built out of these atomic actions, although I’m not sure quite how clear that is.
And then these RE-Bench tasks are these very challenging, novel machine learning research engineering challenges.
So, totaling 170 tasks. I think this is very good. What’s really interesting is that people don’t understand that when people quote the number of hours, it is the human-equivalent hours, but machines will probably take a lot less time for that.
One thing I’ve always wondered was: why didn’t you publish a second chart where it was just, “Well, here’s the difference between what machines can do versus what humans can do”?
That’s a good question. I think you can think of time horizon, in some ways, as a summary statistic—a single number for how good models are plotted over time. We could have done how long the models can work productively. It’s not quite clear how to operationalize that. You do want some notion of success. Otherwise, how exactly do you threshold this—how long they can work for?
But in principle, we could do something like that. This is closer to the first thing we tried. This is the thing with the clear empirical trends. I do think it’s right that a common misconception about time horizon is that it’s about how long the models should work for.
The models are, as we all see, working for longer periods of time autonomously in the wild when we use them in Cursor, Claude Code, or Codex, but that’s not the primary thing going on. In some ways, I think it would be easier to explain time horizon if you assumed that the model solved all these challenges in 0 minutes or 5 minutes, just to emphasize that’s really not the thing that’s going on here. Instead, we’re just plotting the difficulty of tasks they can do over time, and that difficulty is measured in human time.
Alessio Fanelli
Yeah, I do think there’s some collision when people say, “I ran Claude Code for 5 hours,” which is the top of your chart right now. But that would mean 5 hours of a Claude Code run would be the equivalent of a—
30—
—30-hour task—
swyx
500—
Alessio Fanelli
—in your—
Yeah.
—thing, basically. And yeah, I think that’s interesting to—
swyx
Or it might not, because—
Alessio Fanelli
Yeah. Exactly. Yeah. Totally.
swyx
—it might spend 3 hours doing absolute bullshit.
Yeah. And a lot of these claims about Claude Code—
30 hours—
—was running for 30 hours or something.
Yeah.
I have a lot of questions about that. How good was that output really at the end?
Have we investigated those? Yeah.
We have and we haven’t, to some degree. I can talk about particulars. But there’s also the question of, if I attempted that again, how cherry-picked is this example? If it succeeded the first time, would it fail the second time? I think, in some ways, those anecdotes are interesting, but not so scientific.
Yeah. That is something for people serious about AI to understand. The state of people making claims on agent performance is very unscientific and much more anecdotal, and sometimes influenced by marketing desires. Let’s just put it kindly.
Yeah. I think METR’s out there trying to support civil society, trying to provide high-quality, independent information to the public. I couldn’t agree more that the information environment is less than perfect.
Let’s talk about Claude Opus 4.5. It’s a very big jump. This was the first time I called it out when you guys put it out, where I was like, “This is the first time—as far as I understand, you’re the first people to call out how much better Opus 4.5 was than the status quo.”
And I think this almost ties into your background as a superforecaster a little bit, because then, basically, over the entire holiday period, over New Year’s, people discovered what you had already discovered. What are your reflections on that? What were your reactions? Any stories to tell about that?
That’s very kind. I do want to attack you on 2 claims. Firstly, I have not been a superforecaster. I think there’s a particular group of people who work for Tetlock or something who are supposed to be called—
Okay. No, no. But you’re broader.
What I’m referencing is, you were number 1 on—
4. Opus 4.5 Breaks The Trend
You know, Opus 4.5 is a big jump on benchmarks as well. I think, in some ways, METR time horizon is highly correlated with a bunch of benchmark scores. It’s, in some ways, a more understandable way of thinking about what benchmark performance really means—slightly more interpretable.
I do feel intuitively like Opus 4.5 was a big bump. I’ve seen some of the most talented engineers I know go from being picky about not using AIs for coding to practically not writing a line of code. I’m sure many other people at previous model releases have seen similar things happen to them. I’m not sure what that implies. It’s so discontinuous.
In some ways, I think the story of time horizon is that progress has been remarkably continuous over so many years, so many orders of magnitude of compute and effective compute. But yeah, I think model capabilities are astonishing. It points to model capabilities being even more astonishing in the future.
It broke your trend line—the trend line that you were working so hard to build over multiple years—and it just did that.
Yeah, I’m not sure about the characterization. I think it’s—
So there was some speculation, even when the paper came out, that maybe the appropriate trend line to use is this faster 4-month doubling time—
Yeah.
—which Opus 4.5 would be perfectly in line for.
And picked 7 months. I was more of a believer in 7 months, and so it is falsifying my trend line in some way.
It’s slightly confusing to think about whether differences from the trend line represent differences in the difficulty of our task distribution at particular points versus something more fundamental, more like latent capability, and I don’t feel like I have a perfect handle on that. In general, I think the Twittersphere pays a lot of attention to particular model releases, and really the informative thing is what the trends look like over a period of 1 year or a period of 3 years.
It was a pretty significant update, I would say, for all of us. I would co-sign what you said there, with even very cynical, more senior developers being finally pilled into agentic coding. Now very serious people are telling me that they want to commit their organizations to full, “Don’t write a single line of code by human hands,” and just commit to 100% agentic coding, which is not something that you would have said a year ago.
That sounds right to me. I feel it. I feel it in my own case.
Alessio Fanelli
How do you validate previous research? So take the developer productivity study, right? Your AI slowed people down. If you were to redo it with Opus 4.5, would you expect the results to be dramatically different? And should we redo the study? Should we stop citing the study? How do you think about that?
5. Developer Productivity Gets Harder
We have been redoing it in the background. I think it’s—and I won’t comment on exact results—but I think it is much harder to do it today than it was in the past, for all sorts of reasons.
The first is, as AIs get better at coding, it’s harder and harder to find developers submitting tasks who are willing to be randomized to AI-disallowed. There’s a quote-unquote “selection issue,” where maybe we end up only observing the tasks that they thought AI wouldn’t greatly uplift them on ahead of time, because those are the tasks that they’re willing to be paid for to be flipped into AI-disallowed.
There are other issues. I think today a common workflow is to work on multiple issues or multiple lines of work at the same time, concurrently, and that wasn’t really true before. It’s difficult to know how to capture that in our study design. If you flip a single task to be AI-allowed or AI-disallowed, you’re supposed to work on that single task, but actually, that’s not how developers are working today.
I think basically these weren't threats to the previous study design in approximately March 2025. People weren't really working concurrently, or not nearly to the same degree. They basically were giving us all of their issues.
swyx
Yeah. We have Quentin Anthony, who was part of the study—
Yeah.
Alessio Fanelli
The only productive developer.
swyx
The only productive—
I have some questions about that. I think Quentin is very talented, as are all of the developers in the study, but we don't measure developer effects very precisely.
Yeah, no, I'm curious. I don't know if he's part of the new study. You don't have to share that.
Alessio Fanelli
Yeah.
swyx
But I think it'll be interesting to have people on again who've been in the study. I do feel like things are changing. Even 3 months ago, I was using Cursor a lot more, in pair with Claude Code.
I think today I do a lot of just async Claude Code and then review and iterate. I don't know, man. It's much better, and I don't know how to quantify it. I think that's part of some of your points before: people maybe overestimate. If you were to ask me how much did it speed you up, I'd say, “I don't know, 10x,” but it's probably not, right? I don't know how to calculate the actual percentage, so it's hard for everybody involved.
So here's some issues you might think about. If you took the tasks that you were completing personally in March 2025 and then submitted them to our uplift study now under the previous design, we might reason about how much faster those would go. You might expect them to go somewhat faster because AI capabilities have improved.
But you're doing a different and larger set of tasks now. I can think of a couple of side projects that I have that I simply wouldn't be doing were it not for AI existing. In some sense, the speed-up there is maybe infinite because these are things that I simply could not have done otherwise.
But if you were to equate speed-up with the additional value that these projects are providing, these wouldn't really line up. There's a reason I wasn't getting the expertise to do the other projects before. It's just less valuable to me.
Another problem is the concurrency thing that we just raised. I do think that very bullish estimates of speed-up today are, to some extent, inflated by what we document in that original paper: people's expectations of speed-up tend to be too optimistic, it seems. They also tend to be inflated, I think, by not quite grokking that the additional tasks they're able to complete are lower-value than you might think, and that there's a reason they weren't doing them previously. That said, I don't doubt that those tasks do have value, or that people are being sped up on even the tasks they would have done before. It's a complicated issue.
Yeah, I do think that a lot of companies have issues absorbing additional productivity, especially when you're a real product organization. If you think of the AWS console, right? If you gave AWS AI and everybody's 10x more productive, even if they ship 50,000 more services, customers can't really absorb 50,000 more services.
Right.
So I think there's some—you shouldn't really expect your engineers to do 10x more because your organization cannot push out 10x more product. And I agree. I spent a lot more time doing side projects and things, which have been fun, but not that valuable in an economic sense, but valuable to me, to my soul.
Yeah. I don't want to overstate that. I think probably people at AI companies today are being significantly sped up by access to AIs. I think you, for your non-side projects, are probably being—
Right.
—sped up by access to AIs. But yeah, it's tricky. It's easy to overstate.
Yeah. What's the Cognition internal tracking? How do you guys measure? How do you measure speed-up? How do you measure how much impact you have as—
How much do you think people—
Yeah.
Alessio Fanelli
What's your number? Oh, me personally, quite a bit, except that I am doing a lot of nontechnical stuff, like organizing a conference, which is mostly dealing with contracts and booking guests and all that other stuff that has nothing to do with code.
I would say that what I've seen internally in Cognition is a lot of just commit velocity, regardless of whether or not you had authored them. I do think, weirdly enough, the number of PRs, let's call it, is a pretty decent measure of how engaged you are in terms of shipping products and also debugging and maintaining things. I don't think that there's a good measurement of quality. There's no story points.
The other guest that we had at AI Engineer was talking about how they pay people by story points. You do more—you complete more story points, we'll pay you more, and there's no upper bound to that. I think that's a really interesting thing, except that you have to have a very confident relationship between the engineer and the person assigning story points. That's effectively what you're doing: your hour is a story point.
Yep.
We reward the models based on the story points that they complete.
Right. In some sense, ideally, you want to get Cognition and a bunch of other companies. You randomize the companies to use AI or not use AI, and then the outcome metric for your randomized controlled trial is how much profit they make or something, or their valuation after some period of time.
Yeah. I think basically no one is stopping to do science except for you guys. We know RCTs are the best, right? But sometimes human intuition is good enough that you're like, okay, when we lack data but enough humans agree, either it's mass psychosis and we're all wrong, or there's something here and we just cannot articulate it, but the benefits outweigh the cost of slowing down to do the science first.
This is not where we're introducing a new—I don't know—food to the general population, where we have to do a lot of safety testing. Here, it's just software, guys, so let's just ship it.
Totally. Thinking at METR about why models today aren't catastrophically dangerous, it's interesting to get the uplift numbers. It's interesting to get the time-horizon numbers. But really, why don't I believe they're dangerous? It's a mix. I watch the models do things in transcripts, and sometimes they're kind of derpy. They don't use resources well, or they just clearly have some of these obvious faults.
In broad deployment, only slightly worse models in the past 6 months have not been doing anything crazy or causing great danger. The next model is only a little bit better, and so it seems surprising on priors if it was so dangerous. I totally think that anecdotes and intuitions are real evidence. People should totally be taking that into account.
I do want to comment on this whole thing about how this threat-assessment site is in your name. Typically, I expect, let's say, EA-affiliated companies or organizations to be on the Eliezer side of the world, where they're banging the drum about danger. Whereas here, you're actually saying, “Actually, it's pretty balanced. We care about AI safety, but also we're not there yet, and we are actually the watchdogs looking out for it.”
I would say you stand out as someone not funded by the labs where—let's say ARC, is it ARC?—or some other groups that also do threat evaluations before model releases. They would typically be funded by OpenAI or some other big lab.
Yeah.
But now you're a separately funded organization. As far as I know, it's a big deal that you're not funded by a big lab.
METR came out of ARC, I think.
But now you're a separately funded organization. And as far as I know, it's a big deal that you're not funded by a big lab.
Yeah. I think if I don't have this independent source of expertise, I can bang that drum forever.
Yeah. The other thing also is just this concept of capability explosion, which is a word that you use. That's also something I wrestle with, right? If you believe in emergence, you believe in multiple capabilities fusing together to produce generalized capabilities that you may not be able to detect, it's hard to predict based on trend lines. It should be discontinuous in some sense, and I don't know that going, “Oh, the n − 1 model was fine, therefore the n model is probably fine,” is a good argument. It's really hard to tell.
The thing that gives me comfort is that yesterday I was at the OpenAI livestream, and even Sam Altman was like, “Yeah, I just let Codex YOLO dangerous permissions, whatever, on my computer, and I don't approve the model anymore. It just does whatever it wants to do on my laptop.” And I think, I guess the guard is every model lab leader dogfooding, and if it screws up their personal permissions, then they have skin in the game, is what I'm saying.
On the continuity argument, I'm not sure what I think. I agree that it's flimsy. There are only so many models, so many data points on this time-horizon trend.
How much should we expect it to be continuous, to keep going like this? I'm not sure. It may be an intuition that something might be discontinuous because models are providing so much effective labor in improving the next generation of models. Maybe that's a reasonable thing to think.
On the other hand, I've been pretty surprised so far about the degree to which it's continuous, and that gives me some faith that it might continue to be continuous in the future. It seems ambiguous to me.
We have breakpoints in physics, right? So I'm curious if—
Yep.
—it doesn't seem like—it's funny. When you think about water, water will boil—
Yeah, yeah.
—at this exact temperature. So maybe we do know, but I feel like we don't really know. I don't know if it seems like there's the same thing with models, because it's all just compounding of the same thing, if that makes sense. It's just scaling the same thing over and over.
Yep.
But yeah, maybe we will see it. I'm curious what you would need to see to feel that this is here. Because even if you look at Claude Opus 4.5, that's clearly out of trend, and so you were saying 4 months instead of 7 months. But if then the next model is, "Oh, maybe it shouldn't be 4 months; it should be 2 months," would that make you change your mind about whether or not the months thing even makes sense?
Or if we maybe pass some base level after which it accelerates and will keep going? I don't know. I feel like you must be having this discussion internally.
6. The Capability Explosion Threshold
In some sense, the thing that would really concern me is if AI R&D was fully automated inside of some lab. That would totally seem like the conditions are there for potentially a capabilities explosion. If I saw a time horizon of 1 year, I would still find it ambiguous, I think, at the moment, whether that was the case.
Because for things to be fully automated, 90% automated isn't enough. You need some full loop to be closed, and perhaps we're missing some sort of task that points to that missing 10%. So I think it's a tricky issue. I can't give a number. But yeah, my intuition for where water boils is at some point where this loop is fully closed.
There are interesting debates about what exactly that loop is. Some people talk about software-only intelligence explosions, which means that, even holding hardware fixed, we could get to the point where, just from models improving themselves, they would then be smarter and take the next step to create even better models with even fewer resources. This sort of thing could lead to some extreme takeoff.
Or maybe that fizzles out somewhat quickly, and instead you need, in addition to the software-only capabilities, chip design, or maybe you even need chip production, and that's this larger loop that can close. If you think that, I think you maybe should still think that closing the chip-production, software-only, and chip-design loop is potentially very destabilizing and concerning. But yeah, tricky issue.
swyx
I think that is the actual paperclip factory. If you incentivize a model to go build its own compute, it would just build whatever it needs, and then it would turn the planet into chips.
Alessio Fanelli
I don't think it can do it. We will stop it before that? Question mark.
swyx
I don't know if we have the power. There's no off button. Like, there's—
I think it's super hard to foresee. But a model that—
Yeah.
—had those kinds of capabilities, it's hard to rule out that there would be something like a capabilities explosion. Who knows what happens after that point?
Yeah. Okay. So there are a bunch of other benchmarks that actually directly track this, right? OpenAI has PaperBench, I think, which directly tracks its capability to reproduce papers, and I think there are a lot of other similar ML self-improvement benchmarks.
Yakun from OpenAI has directly prioritized, "We will have an automated AI researcher." I did a podcast with Yita from Gemini, who's also basically plugging his own training logs into Gemini to improve his own code, and I'm like, "At some point, you don't need to be here."
I'm—
You're skeptical. Okay. Say more. Say more.
I'm not speaking for everyone at Meta. I'm a relatively longer-timelines, quote-unquote, person at Meta. We have Nicolo, my colleague, who helped out with AI 2027, who's on the shorter-timelines end. This is not a view—
Which is officially AI 2028 now, which, you know—
That's—
One year passed; we move it back a year. Okay.
Yeah. I think my view would be—not that Nicolo's view is necessarily different, just so I'm not speaking for other people at Meta—a PaperBench, let's say, perfectly measures not only reproducing papers but, in fact, producing novel research papers.
That's just a part of this R&D production process. There's also, like, your GPUs are constantly failing. Can you get someone to go to the data center and fix them in the appropriate way? Can you call up the water company when the cooling breaks down? Et cetera, et cetera, et cetera. I'm not aware of benchmarks tracking that in particular.
My point is more that there's this very long tail of things potentially involved in R&D that would perhaps need to be fully automated in order to lead to a capabilities explosion. I expect we're measuring, in some ways, only a small proportion of those capabilities, and so I expect the capabilities needed for the full loop to close to come somewhat later.
Yeah.
That's a contrary view.
I don't think so. I think that's a reasonable take. Something that does surprise me when you're talking to capabilities researchers is that you guys don't have an enumeration of the capabilities that matter. I think you implicitly do in the choices that you make.
But I think it's almost important—I always imagine, like, a wagon wheel. That's the terminology. I don't know who came up with this term. But here's the 10 things we care about, and here's where everything is on those 10 benchmarks. I feel like capabilities tracking is just tracking: okay, what's that list, and then where are we on that list?
I almost feel like this need to reduce everything to a single number is actively working against that because it reduces any form of nuance: it's insufficient here, like the calling-a-data-center thing. So we're fine, and it's like, actually, we should just not invest anything in that area because that's the danger zone.
Yeah. I think I couldn't agree more that time horizon, for instance, like many other single numbers, is one number that's collapsing an enormous amount of really important detail.
Dimensionality, yeah.
I don't know how to come up with that list of 10, and I challenge you—
I'm working on it.
—if you're able to come up with that list of 10.
Yeah, I'm working on it for code.
I'll be very interested to see it for code. My intuition is that we'll come up with a list of 10, and it will turn out that there's a secret 11th thing that—
Yeah.
—we thought was important, but it was difficult to pre-specify ahead of time. And now it seems obvious that, even ahead of time, if we'd had that foresight, it would've been helpful to add.
I think that the security community does this by versioning year by year, right? So this year, the top 10 are blah, and we would just publicize it to everybody, so everyone knows what the top 10 is. And next year, we'll have a different top 10.
It obviously is stochastic, and we should update our assumptions, but it's broadly useful to have that list as a public service.
Alessio Fanelli
You also had this research on the slowing AI improvements based on AI compute, and you mentioned that, in a way, you could tie the AI time horizon to the growth in compute. Can you say more about that? It's, in a way, unintuitive because compute growth is not always tied to how much compute every single model needs. It's kind of a broader market thing.
Yep.
Yeah. How did you get the 2 together, and what were some of the findings that you had?
7. When Compute Growth Slows
Yeah. Maybe for a second, let's take time horizon very literally. We don't have the qualms about it that we've just been discussing. It makes sense to continue extrapolating it into the future. What are some important forces that might cause it to rise more quickly? Some of the things we've just been talking about: automated R&D versus things going more slowly.
One of the most obvious forces that might cause it to go more slowly is if inputs slow. One important input is compute. I think we all have the intuition that, to some extent, if compute growth slows, which we expect it to at some point in the not-so-distant future, then capabilities will slow. But by how much? It's a big question.
The suggestion in this paper is that, if you think algorithmic progress—coming up with the transformer, coming up with RLHF, all of this stuff, better learning-rate schedules—is itself a function of compute, because you need compute to discover it.
The gains from transformers show up much better with scale. If you don't put in those resources, you'll never find out that this is the superior algorithm. You need to run a ton of experiments, and each of the experiments can be quite compute-expensive. Not to say that no labor is involved—obviously, people are working on this.
But if you think it's ultimately bottlenecked by compute, then algorithmic progress slows down too if compute growth slows down. If you think about time horizon, or whatever your favorite measure of AI capabilities is, as being a function of algorithms in some sense and compute in another sense, and both of those components halve when compute halves—compute is halving, and algorithmic progress halves because compute is this important input—then you might expect time horizon growth to halve. Some of these major capabilities milestones that we might be interested in would then be significantly delayed.
I think there are so many caveats to that picture. There clearly are some types, at least, of algorithmic innovations that did not require a lot of compute to create, and some that took a lot more compute input. If you expect that no compute input is required, we could just survey researchers for the best ideas and then immediately put those into training the frontier models. Then there would be no slowdown of algorithmic progress from a slowdown in compute growth.
Of course, all of this is counteracted by the possibility of capabilities explosions, or AIs providing significant labor at making AIs better, even short of capabilities explosions. But just analyzing the compute force on its own might lead to significant slowdowns, depending on the degree to which it makes sense to call algorithmic progress basically determined by compute versus not needing compute to come about.
swyx
Do you think of compute on a per-lab basis? Because there's one way you can model this out: improvements slow down, not every company is able to stay in business, and then their compute gets recycled back into the other labs, which then grow compute again. There's almost a benefit to the heterogeneous distribution of researchers and compute. But I'm curious how much you care about broader compute—compute being out there for people—versus the big labs having more and more compute.
Yeah. For the paper, we use OpenAI data and OpenAI projections. I think this applies more broadly, but we used that as a kind of case study. I think the argument I just laid out goes through if you're not interested in compute at all and you just talk about dollars: What are the dollars going into models? Will algorithmic progress slow if the dollars going into them slow? The whole argument works, and that works, I think, at an industry level or at a lab level, and so on and so forth.
I agree that things like certain labs going out of business, labs consolidating, or these kinds of industrial-organization questions would be very important. I'm laying out an extremely simple picture. The real picture is not as extreme, but that's the basic picture.
Alessio Fanelli
We have examples of xAI being said to be distilling from Claude, right? So people share compute in indirect ways, let's call it. I think that's also very interesting. I'm just curious what OpenAI numbers you had. Is this the $500 billion for Stargate or something else?
This is from their previous tax returns—the amount they've spent on R&D compute.
Timeline?
And then from The Information reports earlier this year, some projections that OpenAI has for how much they'll spend on compute R&D in the future, converting that from dollars back into FLOPs.
Yeah. It's interesting because—
Back into FLOPs, sorry.
And obviously, all the labs—but particularly OpenAI in the last 3 months—have basically thrown $10 billion each to every single compute provider on the planet to develop alternatives to their current approach, which is very interesting. But I also say don't discount Meta's compute spend, don't discount xAI's compute spend, and don't discount DeepMind's compute spend, all of which you have basically zero visibility into, right? If you're looking at a single company, maybe that's authoritative, but then the total spend could be a lot higher.
Yep.
It's interesting. I also observe that people like Dylan from SemiAnalysis do tend to very strongly tie model progress with compute clusters coming online. The people on the model/API side don't see it, but this is all downstream of our 10,000-GPU cluster just coming online. It takes 6 months to do it, and therefore Grok 5 will be here. It's pretty mathematically deterministic there.
Yeah, it seems right to me.
Yeah, it's fascinating.
swyx
Yeah. From the lab side, they must see something in the early checkpoints to go ahead and keep investing 18 months from now. I wonder what the time gap is between finishing—
Alessio Fanelli
Yeah.
swyx
—a good pre-training run and going live. That's probably 9 months, 12 months, something like that.
Alessio Fanelli
I think Mistral is actually pretty open about this. The plans for Mistral 3 and 4, I think they've been pretty open about the number of GPUs and the direct timeline from coming online to when they ship the model. It's pretty set. I don't have a clear timeline in mind, but I would say 4 to 6 months.
swyx
Yeah.
Alessio Fanelli
But yeah, the competition is very tight. One of the things that's also very interesting is seeing when labs throw away models because they failed—when their run came behind someone else's run that was better. Then they say, “Oh, we can't release this anymore.”
swyx
Yeah. That's the biggest risk with prediction markets on model performance, actually. Just to tie back—
Alessio Fanelli
Failed runs? Yeah.
swyx
I think in December there was the question of who was going to have the best model by the end of 2025.
Alessio Fanelli
Yeah.
swyx
I think there was a lot of activity in the last few months when the GPT-5.1 model came out and—then I guess Gemini, because they just threw that out. It means that Gemini is coming out next week, and so trade that.
Alessio Fanelli
Yeah. Do we want to talk about Manifold—
swyx
Yeah.
Alessio Fanelli
—while we're on the topic?
swyx
You were the most profitable Manifold Markets trader. I mean, there's obviously a lot of talk about insider trading on these markets, especially in AI. I've seen it with a lot of the embargo news that we get. I'm like, “Man, people are trading a million dollars on this market.” There are thousands of people who know the actual information. If you didn't have insider-trading information, how would you think about modeling these things out, and do you think it's a worthwhile thing? For example, who's going to have the best model in 3 months? Do you think that's a prediction market where you can build some sort of strategy alpha?
8. The Prediction Market Edge
I guess the naive prior, without any extra information, is just: In 2025, for what percentage of time did which model providers have the top model as measured by time horizon? You could do it for any old benchmark. I think that's something like 5% xAI, 50% OpenAI, 45% Anthropic. Don't shoot me if I'm completely incorrect.
Alessio Fanelli
No, DeepMind.
I don't think it's the case that a DeepMind model was at the frontier of time horizon at some point in 2025.
swyx
Oh, time horizon. Yeah.
Yeah, different things for different measurements. Maybe that's the same prior that you want to apply. xAI was coming online at the beginning of the year, so maybe naively you want to raise xAI a bit. Yeah.
I'm always curious when I see people betting on these things that are obviously not—there's no real basis to—
Alessio Fanelli
I was just going to say—
swyx
A broader—
Yeah. Go on.
Alessio Fanelli
What's your secret for Manifold Markets alpha?
Yeah, I see. So the secret, if you read this article about how I became the number 1 most profitable trader on Manifold, which sounds very nice and impressive, like I must be so good at predicting things, is that it mostly comes down to this 1 market where Manifold had opened up a charity program. The market is on how much is going to be donated through this charity program by the end of its 1st month.
The market opens, or I first see it 5 days in, and it's giving a linear projection of how much has been donated so far. Let's assume that the per-day amount keeps getting donated every day until the end of the month. As a person who gives money to charity sometimes, I noticed that you could manipulate this market by giving more to charity and so moving it more up.
I think the strategy was to put a ton of mana, this fake currency that's used on Manifold, into the option that was above the linear projection. People keep betting against you because it doesn't look like that's happening. I haven't actually done any donations yet. Eventually, they cotton on to what's happening: someone's going to make this donation to move it over the edge, and they're betting on that.
And then I did it again into the next category, once people had started betting on that category above the linear projection. Again, people bet against that, and I mopped up those fake internet points. Then I think I did it once more as a bluff. The bluff failed, but the previous 2 worked out, and then I ended up donating—I can't remember exactly how much it was. Not so much. Something like $5,000.
No, it's all for a good cause.
Yeah, yeah. I gave, I think, and won lots of fake internet points on the market, and so became the number-one most profitable trader. Slightly legitimately. There's nothing about that that was outside of the rules.
Exactly.
swyx
This is called—
Also, I used to have less respect for my forecasting abilities.
Yeah. This is called prediction markets with high agency, as you actually go and—
Alessio Fanelli
Yeah.
swyx
—the future is what you make it. So I think the broader lesson is the classic difference between Manifold Markets and Polymarket. Polymarket is only real money, right? So is the whole fake-internet-points thing a worthwhile pursuit or a waste of time? Should people actually want to use real dollars? Maybe that's one question.
The other question is prediction-market ethics, which I think is always just going to indirectly come to assassination markets. Even if you ban the phrase “will someone die?”, some other proxy for “will someone die?” will happen to be an assassination market.
Yeah. I'm good friends with the Manifold Markets co-founders. I love them very much. My view on the social value of prediction markets—which was always the dream, right?—is that it would be nice to have calibrated probabilities on events that matter: whether this country is going to war with that country, things that really matter to people. It would be nice to have high-quality information.
But when I look at real examples that have come about in the past year, it doesn't seem to me like those examples are so socially valuable. I'm not sure about assassination markets in particular. I'm sure those would be overruled, hopefully. But I think gambling-like behaviors are socially costly, and the value of higher-quality information is real. Is it worth that disbenefit of people trading away their money? It's not so clear to me.
Yeah, yeah. Price discovery has a cost. Sometimes that is gambling, and that is the stock market. It funds a lot of corporate America.
Yeah, yeah. A lot of the stock market—
The guy used to be one of those—
—big firms playing against other big firms. It doesn't have the same character as sports-betting markets. To take an example on the other extreme, they have this very different character. You might imagine that at least one direction prediction markets could go is this sort of big players playing against retail, and that maybe has a more worrying dynamic.
I don't keep so closely in touch with the space, but at least something like that, you can imagine, would be concerning.
Alessio Fanelli
I think at a large enough scale it becomes profitable for some of the companies to do it, if they can get the markets on it.
Yeah. Right.
I think now it's still small enough numbers compared to the rest. Which company has the best AI model by the end of January? $28 million of trading volume.
swyx
Oh, my God.
Wow.
Alessio Fanelli
It's just crazy. Why are people trading 20—? It's just crazy. But I was having dinner with somebody this weekend and—
Wait, wait a second. I think it's totally possible to—I'm not going to do it. I think it's important that Meta employees not be making bets on prediction markets like that. But I think it's totally possible, in principle, to have a guess for the answer to these kinds of questions.
Oh, but other people know exactly. Gemini 3's score on the FrontierMath benchmark by January 31st is like some—
I see, you presented—
People at Google already know what the number is.
Right.
You know what I mean?
Yeah, yeah.
It's—
I think if you believe in the benefits of price discovery, then this is—
Exactly.
—this is legitimate—
swyx
Yeah, yeah. They actively encourage insider trading, and this is a way for insider information to pseudonymously leak out. As long as whoever is traceable to that thing bears the consequences of leaking the information, that's okay.
People, I think, have been fired for trading on insider information. The only step up from that is the government coming in and saying, “This is actually illegal. Put you in jail for that.” But I think for now it's self-policing.
Alessio Fanelli
Yeah.
swyx
Yeah.
Alessio Fanelli
Retroactively, you can't really do that. All right. If you have any embargo news, press N—no.
swyx
We do work with people under embargo, and we don't trade on it.
Alessio Fanelli
Yeah. We would have made a lot more money trading on embargo news than we've made on anything else. What else? What are other interesting model-evaluation trajectories, or anything that you're not doing at Meta that you've maybe seen other people do that you find interesting, or that you would like more people to do?
9. The Next Evaluation Frontiers
One project that I think is interesting is AI Village, which I think possibly both of you would have come across. These are very open-ended goals given to a village of agents, and they try to accomplish them. Setting up a merchandise shop is maybe one of them, organizing an event in a park, building a human-subjects experiment, this sort of thing.
I have a number of questions about exactly what I should learn. They're using old models as well as new models in this quote-unquote village. The models are relying a lot on vision capabilities, which we spoke about models not being so capable of today, this sort of thing. But the vibe of models trying to achieve open-ended things instead of benchmark-like tasks—the vibe that's a bit more like vending machine bench in some ways—seems like a very interesting direction for the science to go, or something that comes with a lot of cons but attacks some of the cons of benchmarks in a pretty interesting way.
I think seeing the ways in which these models trip up, seeing the ways in which they're derpy, is an important source of information. I'd be interested in more work like that coming about. I think that's one of them.
Another is transcripts as an extremely interesting source of information. This is the models taking actions, then seeing outputs, and then using those outputs to commit the next action, and so on and so forth, on benchmark-style tasks or, even more interesting, on in-the-wild deployments like you might find in your own Claude Code usage, Codex usage, Cursor usage, et cetera.
That has the con of being less experimental, less clean and scientific in some way. It's more selected, quote-unquote. The tasks that you get AIs to do are obviously the tasks you expect them to have some chance of succeeding in, so you're not just giving them any sort of task. If I see the models doing something extremely impressive or potentially unsafe in some sense, subverting user preferences, it's not clear how often that kind of behavior would happen given the previous history.
But it's a massive data source. There's a huge amount of information there, and I'd love people to be working more on that sort of thing.
As we mentioned, there are a lot of problems with time-horizon or developer-productivity work. I think it has been important evidence. I think it's moved the field forwards, but it's far from perfect. I think there are lots of other directions there that look very interesting to me.
Maybe one that I'll call out is this difference between whether models pass unit tests, whether they succeed by SWE-bench-like scoring, a kind of METR-like scoring, benchmark-style scoring, versus whether their solution would be merged into main. That is, whether the solution adds tests where it should or doesn't, whether it follows existing patterns in the codebase, and whether it makes sure that its changes speak to other parts of the codebase in appropriate ways.
That seems very interesting to me. I think model capabilities probably are lagging behind there somewhat versus what you might see on SWE-bench-like scoring. I can keep going, but—
swyx
These are the novel research things that you were referencing earlier, right?
These are novel research things, yeah.
Just to comment on the AI Village thing first, you mentioned a lot of stuff. I even want to double-click on the transcript stuff. The AI Village ties back to one of our highlights of last year, which was Noam Brown's conversation about how he's actively working on multi-agents that are cooperative instead of competitive.
Mm-hmm.
The basic idea is that we can do more as a team than we can do individually. The agents are the friends we made along the way.
Mm-hmm.
And I think that's great. On the DeepMind side, the way they phrase it is literally having an open-endedness team, which I think is a topic that reemerges once a year.
Alessio Fanelli
Yeah, I think it's unclear what open-endedness does for us, and this is a core divide in terms of studying these things as life forms, potentially new artificial life forms, versus tools for us that serve us. Maybe open-endedness means that there is no goal, and if you're just trying to evaluate this as, “What does it do for me?” that's completely wrong. You will never get anywhere with that because they are just living their lives as artificial life forms.
In some sense, the gold-standard evaluation that I would like to do, if I was looking to learn the most about the questions that I'm most interested in—the degree to which AIs might automate or accelerate R&D—I'd quite like to just give the AI a bunch of affordances, type into the AI, “Automate R&D, go,” and see what it does. I suspect that wouldn't work today, even with all the affordances, because it would fall over on its face when working with resources and handling resource use in ways it's not so capable of today. It would struggle at some types of long-horizon tasks, et cetera.
In some ways, I think benchmarks face difficulties in capturing this sort of thing. And AI Village—or AI Village-style things, with these more open-ended goals and seeing how models pursue them—gives some color to this sort of thing, to seeing models fall on their face. I think these more open-ended goals will become more and more important over time.
I agree to some extent. In the extreme case that I just mentioned, you're going to provide them documentation about how this part of the company works and that part of the company works, and so on and so forth. It's not purely open-ended, but it's pretty open-ended. It's more open-ended than the kinds of problems that we're giving them today.
Yeah. Standard open-endedness.
Yeah. If models are excellent when used with a detailed issue description and something very clearly specced about what they're supposed to do, that's interesting, but it's a very different thing, I think, from being able to automate R&D. I'm interested in how far we are away from that, and in some ways, this speaks more directly to that sort of thing.
Yeah.
swyx
We had the Terminal-Bench guys on the podcast. How do you think about harness benchmarking, in a way? If you look at their leaderboards, the same model with different harnesses has a difference of, like, 10 percentage points. Does that seem interesting? I don't know if you build a harness in Meta or not—do you always pick the best harness or compare them?
Yeah, let's say how we pick harnesses at Meta. This is not what I work on in particular. I'm not an expert. But roughly, we build harnesses to get models to be as performant as possible on a dev set of tasks and some held-out set of tasks, and then we use those same harnesses, trying to make sure they're not overfit to our main suite of tasks.
On the one hand, I do have the intuition that there's a lot of juice in scaffolding. It's easy to overstate how much juice there is because of this overfit problem. If we were building a scaffold to do as well as possible on our test tasks, then it would do much better than the scaffold that was built only on our dev tasks. In some sense, that would feel illegitimate or not interesting, or you wouldn't expect that to generalize to some other set of tasks, potentially.
On the other hand, a lot of work has gone into building scaffolds that make models as performant as possible, because we are interested in upper-bounding the capabilities of models when thinking about whether these models might or might not be dangerous. I do have faith that these scaffolds are a lot better than the first thing that people might try, because so much effort has gone into them.
Yeah. It's interesting because I do want to overfit as a customer of the models. You do want to overfit to your task specifically, and I think sometimes people underestimate how much value you can get out of it, but—
Yeah, I think if you have a kind of mechanical workflow or something that you're imagining automating, and there's some place where more stochastic intelligence would be nice inside of that, like deciding where to route customers to on customer calls, something like that, I feel like that makes a lot of sense. But for this sort of more general, in particular thinking about helpfulness in software engineering thing, I'm not sure I have that same—
Taking an example, I work in TypeScript.
Yeah.
Alessio Fanelli
Yeah.
swyx
If I build a better linter that's private to me, or a better test suite, like a better Playwright replacement, in theory, I'm overfitting the model to perform better, right?
Yep.
It doesn't really matter to me.
Yep.
I'm not trying to report on model performance. I'm trying to build the best thing.
Yeah, but if you had a model build the linter—
No, I agree. I agree. I think that's the question of—
Yeah.
Okay, should I just wait for the next model? You know what I mean? It's like, at what point should I be building the better scaffold? No, I'm wrong. All scaffolding is going to get washed away.
Yeah.
But on a realistic schedule, what am I supposed to do this week?
Alessio Fanelli
He would say that.
swyx
On a realistic schedule, what am I supposed to do this week?
Alessio Fanelli
Yes, those can simultaneously be true: that all scaffolding will be washed away, and that scaffolding today is valuable.
swyx
Right.
Totally.
Yeah.
Totally. Or within a model generation, it's valuable, and across model generations, it's not so valuable. I'd say at best, I'm an acceptable software engineer.
Right.
I'm intentionally not investing in engineering skills because the AIs are getting so good. Maybe that's the wrong decision. If you expect, as I think you should, capabilities to keep going up and up, it forces difficult trade-offs about how you spend time today, because maybe it won't be so helpful in 6 months' time.
Take a sabbatical. Or if you live in Europe, you can just take 6 months off or something.
Yeah, but then in 6 months, you might want to take another sabbatical for the next 6 months.
Perfect.
Alessio Fanelli
Just to wrap up, what do we expect out of METR in 2026? What does success look like in 2030? I don't know if you have a sort of broader vision. And then maybe on a personal side, we can talk about the karaoke stuff, but METR as well.
Yeah. From METR, I think you're going to see more, hopefully, high-quality capabilities evidence—the kind of thing you saw in the past with time horizon and the developer productivity work, along the lines of what we've been describing. So, some of these future research directions.
We also have some monitoring research directions that I'm not so expert in, thinking about whether we can successfully apply safeguards to models attempting dangerous tasks. There's a whole line of work there.
Is that an interpretability dimension, or what kind of safeguards?
Usually, this is black-box, not white-box, in my understanding of current work, so it's not using interpretability. But you can imagine, in principle, doing something more white-box. Then there's this risk-assessment work that takes into account how capable we think models are, what their propensities are, and whether we can track, using safeguards, the kinds of things that the models are doing. Do we think these models pose large-scale harms? You can expect to see much more of that in 2026.
And then I think productivity or something. There are a lot of people with great talents who are not going to work quite as well in a scrappy environment working on frontier science, and that's the thing we do.
swyx
I just want to prime people for what the valuable skills are in this new age, because I think the more people articulate what the positive directions are, what is hard to hire for, that's what we guide our audience towards improving themselves, and I think that's important.
Hmm.
Alessio Fanelli
Did you have a karaoke question, or?
swyx
I don't know.
Are you going to sing on the podcast?
I've never done it.
Come on.
I've been wondering—
“Can't Help Falling in Love.”
What is this karaoke thing that you organize? Are you a musician?
A musician might be exaggerating it, but I hit instruments and noises come out. I've hosted a couple of these live-band karaoke events, which is like getting a group of friends together, with people accompanied by bands, singing karaoke to an audience of 50, 100, 200 people. It's great fun. I think people should be doing more of this. I look forward to seeing you both at the next one.
I will do that at one of your events. Yeah, it's one of those things where it's weird, because I used to be in a cappella a lot.
Oh, wow.
And I just think it's a dying form. And I just watched this video that was really good about the 2010s wave of a cappella, from, like, Pitch Perfect and Glee to Pitch Perfect—
Yeah, yeah.
What's that group?
Pentatonix?
Pentatonix, exactly. And that's where it died. It's very interesting to see how it's dying as an art form in general, and how new formats have taken over. And I don't know, it's weird for humans also, because now I'm also, let's call it, more interested in synthetic song generation or DJing, anything like that. The human voice is actually more commoditized. It doesn't really matter who sings it.
Ah, I don't know. I feel like there's a kind of transcendence to singing in person that AI-generated songs are not providing me.
That's good. That's good. Yeah, yeah. I do think that we humans always want that.
Yeah, yeah, yeah.
But I'm not sure humans in the year 3000 will want that. It's one of those weird things. Thank you for coming on. It's great to have you as a human in person here.
Thank you so much for having me as a human.
Yeah. Someday we'll interview AI versions of you.