多少个狭义人工智能可以表现得像一个超级智能——Daniel Kokotajlo 和 Thomas Larsen
Daniel KokotajloThomas LarsenTim Scarfe
- AI Futures Project 的新情景「AI 2040:Plan A」刻意将预测与建议混在一起:先暂停6–12个月,建设验证基础设施;随后促成美中协议,缓慢推进至大致相当于顶尖人类专家水平的 AI,并在2030年代维持这一水平,直到2040年才达到超级智能。 核心旋钮是「暂停在你能够可靠控制的最高水平」;与悲观的 AI 2027 不同,「如果这被超验自证式地传播开来,我想我们会相当满意」。
- AI 2027 在已经揭晓的量化指标上,目前大致跑在预测速度的75%;Thomas Larsen 表示,现实偏离预期的程度比他想象得小,这「以一种糟糕的方式让我感到意外」。 收入趋势和其他现实指标「基本贴近预期」;看似疲弱的0.17%软件研发增益,其实是测量口径造成的假象:报告发布时高估了基线,并非进展落空。
- 决定时间线的里程碑是:「AI 公司宁可解雇人类,也不愿解雇 AI。」 Larsen 认为,可验证与不可验证的任务之间并不存在二元分界——一家独角兽初创公司的10亿美元估值「是关于现实世界的可验证事实」,只是验证周期长、成本高;因此 RL 会持续向外扩展,直到 AI 覆盖「人类能够完成的一切」。
- 这套经济学判断是:经济本来就是、也一直是一个自我复制系统,而全机器版本的倍增速度会远快于当前约20年的翻倍周期。 即便暂停在与人类相当的 AI 水平,也会出现更便宜、更快、每年翻倍而非用20年繁衍一次的「云端同事」;Thomas Larsen 给出的极限情形是:「至于世界其他地方发生了什么没人知道,但 Anthropic 已经把月球拆了。」
- Plan A 的透明度制度刻意对前沿实验室的股权价值不友好:公开全部核心训练配方,「会大幅压低它们的估值」,因为 Microsoft 或 Alibaba 可以据此追上来——「这是设计目标,不是漏洞」。 目标是让 AI「走向商品化,而不是被垄断」;压低垄断租金、主动削弱万亿美元级集群投资的动力,正是方案的一部分。至于这会把技术送给中国的质疑,回答是通过利益交换来解决(例如换取更有利的算力分配),同时考虑到实验室安全性很差,中国大概率本来也会通过间谍获得算法。
- 随着模型变强,对齐问题不是变容易,而是变难:控制措施本身携带着「一枚定时炸弹」,而核心失效模式是无声发生的。 模型已经能在 Redwood Research 的控制评测中识别出自己正在接受评测;Thomas Larsen 警告说,「我们极容易陷入这样一种局面:AI 实际上并未对齐,但你不知道,因为在你能观察到的一切层面,它们都做对了」——行为评测不够,必须先取得白盒可解释性突破,才能把规模扩展到可控范围之外。
- 与怀疑派的根本分歧,归结为一个参照系问题:「AI 更像电力或飞机,还是更像云端的人类?」 Thomas Larsen 将递归式结构自适应视为自己的触发线:「那对我来说就结束了……那不是普通技术」;Tim Scarfe 认为这听起来接近他对递归自我改进的看法,而 Daniel Kokotajlo 表示,即使今天彻底停下(「Plan S」),「也会比默认轨迹更好」。
1. 从 OpenAI 内部走向公开情景写作
- Kokotajlo 讲述 AI Futures Project 的起点:他在 OpenAI 负责评测、预测和治理备忘录,但「逐渐对领导层感到失望」,也对「行业内部拥有的信息量与外部拥有的信息量之间的差距……以及你在内部被允许说什么、却真正想说什么之间的差距」感到失望。他离开公司的明确目的,就是「把内部人士看到的未来告诉全世界」。
- AI 2027 在两个维度上都超额完成目标:这次认识论演练让团队学到的东西超出预期,读者数量也达到了他们预测的90分位结果。Larsen 是新报告 Plan A / AI 2040 的主要作者,也参与了 AI 2027 的共同撰写。
2. AI 2027 正以约75%的速度推进——这不是好消息
- Larsen 坦承了一个让人不适的判断:「AI 2027 的进展比我发布它时预期的更贴近原定轨迹……这以一种糟糕的方式让我感到意外。」收入趋势和其他现实指标都接近情景设定的路径。
- 两篇后续博客文章将所有已经揭晓的量化预测与现实进行对照,结论是整体速度约为情景设定的75%——仍在轨道上,只是略慢。看似惨淡的软件研发增益数字其实是口径问题:他们在发布时高估了增益水平,因此 coding agent 带来的真实提升,在现实从更低的基线一路升至并超过假定起点的过程中,表现成了停滞。
3. 兵棋推演是「现实在对你大喊」
- Larsen 借用了军事兵棋推演的框架:你永远不可能预测每场战斗的确切顺序,但「如果你完全想象不出初始计划如何可能带来胜利,那么真正成功的概率就极低」。他的警示案例是中途岛:日本人做过兵棋推演,却不断输掉,于是作弊——让已经沉没的航母复活、重新掷骰子。「这就是现实通过兵棋推演的机制对他们大喊:嘿,你们的计划糟透了。」
- 团队将这种方法称为「情景审查」,已经做过约100场严格意义上的兵棋推演(10人、4小时),其中约10场针对 Plan A。一个反复出现的失败模式在两场独立推演中都浮现出来:一位已经签署对华协议、却面临选举失败的总统,为了在选举前实现超级智能而加速时间线,「这样掌权的就是我,而不是我的继任者」——「这是一个我们在游戏发生前根本没想到的政治因素」。
4. 预测与建议——以及超验自证的担忧
- 与纯粹预测的 AI 2027 不同,Plan A 将两者混在一起;Kokotajlo 承认这种结构并不清晰:方案本身——几大支柱、对华协议、公民分红——属于建议,随后发生的大多数后果则属于预测。如果重做,他们会设置一条中央的纯预测分支,并明确标出建议介入的节点。
- Larsen 将自我实现预言称为「我们对 AI 2027 最大的担忧之一」:随着人们越来越意识到 AGI 的重要性,他们会想,「我想成为掌控 AGI 的那个人……所以我要向那个目标冲刺」——从历史上看,这「是现有竞赛的一大驱动力」。Plan A 则反过来:「如果它被超验自证式地传播开来,我想我们会相当满意。」
- Kokotajlo 的方法论规则是:超验自证确实存在,但被高估了——「粗略来说,我们应该专注于准确预测未来……如果你一开始就试图引导未来,就会陷入混乱,基本上变成一厢情愿」。
5. 观感转变:RL 算力赢下了具身化之争
- Scarfe 形容对 AI 持怀疑态度的 MLST 阵营过去提出过的一系列反对意见:图灵机、符号主义与神经符号主义、意识、物理实例化——「但 AI 还是一直在变强」。Larsen 借 Geoffrey Hinton 的个人基准解释了这种观感转变:Hinton 曾用 AI 能否讲出有趣笑话来判断进展,模型在 GPT-3 与 GPT-4 之间的某个阶段跨过了这条线;每个人都有一个自己直觉上会追踪的技能,而模型不断越过这些门槛。
- Scarfe 对早期争论的判断是,大规模 RL 算力在没有物理具身的情况下提供了代理行为;Larsen 同意,这些系统需要「海量 RL 算力」,但「不需要物理具身」。
6. 一切都可验证——只是成本梯度不同
- Scarfe 仍然怀疑:爬山式优化在「客观、半明确的领域」有效,但进入「模糊地带」后还缺少某种东西——产生规格定义的品味。Larsen 的反驳是:「客观可验证与不可验证之间其实并不存在真正的二元分界」——一家10亿美元估值的初创公司是可验证的,只是验证周期长、成本高。RL 会从便宜的算法验证(如 coding interview)持续向外扩展,「直到覆盖人类能够完成的一切——毕竟,人类最终确实能学会这些长周期任务」。
- Kokotajlo 提出的经验事实是:AI 在「模糊、难以验证、概念负载很高的那些 blah blah blah」上也一直在进步——拿 GPT-3 或 GPT-4 与 Claude 在任何不可验证任务上比较即可。Larsen 认为,时间线的关键里程碑是某家实验室「宁可解雇人类,也不愿解雇 AI」的时刻:今天如果解雇 Anthropic 的全部人类,公司会直接崩溃;但「没有任何根本性因素阻止 AI 达到这种人类水平的能力。真正的问题只是何时发生」。
7. 经济是一个自我复制的机器
- Larsen 将视角拉远:经济「是一个自我复制系统,而且一直如此」——从农业村落生育人口,到卡车、矿山和工厂。很快,这个循环将完全由机器闭合,而「这个自我复制系统的翻倍时间会远快于当前经济约20年的翻倍时间」——可能每年翻倍,也可能每6个月、每3个月翻倍。
- 针对 Scarfe 带有 Graeber 色彩的担忧——没有人类参与者的经济会出现「模式坍缩」——Larsen 认为,即使消费者需求崩溃,只要 Anthropic 足够庞大,再加上采矿合作伙伴,就能启动一个自我维持的产业:「在荒漠中翻倍」——露天矿、自驾卡车、由人形机器人建造的工厂生产更多人形机器人,再生产更多芯片工厂;最终抵达那个极限情形:「Anthropic 已经把月球拆了。」
- 报告中一个被低估的判断是:即便暂停在顶尖人类专家水平,也会彻底改变一切。与人类水平相当的「云端同事」更便宜、更快,每年翻倍,而不是用20年才能繁衍一代;到2030年代末,机器人可以建造城市,特别经济区可以出现露天矿,「海洋上空的地平线被太阳能板填满……只要有人类水平的 AI,再给指数增长一点时间发酵」。
8. 一个巨型模型,还是专门化群体——两者都不是根本分歧
- Scarfe 根据实际使用提出挑战:今天的 agent 还无法组合,技能边界和记忆系统让每个 agent「像是不同的人」;组织无法在不发生故障的情况下把它们合并,内部表征也「碎裂、纠缠……有点粗糙」。Kokotajlo 则从规模化历史反驳:10年前,人们还会类比人类、认真讨论「上千个专门化 Claude」的假设;但经验告诉我们,「对于足够大的模型……编码能力会给物理能力带来一些小幅增益」。未来更可能是一套实际上覆盖整个经济、在其上训练出的单一模型,再配合便宜的蒸馏版本。
- Larsen 用 Elon Musk 做直觉实验:乘法式技能组合——同时在10个领域达到90分位——对任何人类来说都「几乎不可能」,但却可以训练进同一个 AI,这也解释了为什么 Elon 能同时经营多家巨型公司。Scarfe 反驳说,Elon 的魔力在于识别什么会真正重要(科学,而非工程),而且他的行动能力通过工具和他人外置;Larsen 与 Kokotajlo 的回答是,这些原则上都可以自动化。
- Kokotajlo 进一步降低了问题的赌注:即便专门化胜出,结果也会是「一群 Claude……带着某种内部官僚体系」去谈判交易、创办初创公司,就像拥有不同技能的移民群体;从更高层看,「仍然会发生 Anthropic 吞噬整个经济的现象」。他还保留了一个值得注意的让步:群体世界「会稍微更安全」,因为专门化 agent 必须相互沟通,监督起来比所有克隆体都知道一切更容易——于是节目里出现了这句玩笑:「我们赞成你描绘的那个愿景,但不赞成我们正在描绘的这个愿景。」
9. 大脑是机器;H100 位于其计算范围的中间
- Kokotajlo 给出了一组数字:一块 H100 每秒可提供约1e15 FP16 flops,而对大脑的估算取决于按突触还是按神经元计数,范围从1e12到1e18不等;GPU「恰好位于对数分布的正中间」。神经元串联时每秒可能放电1–1000次,而 GPU 的时钟速度以 GHz 计,因此在串行处理速度上,GPU 高出许多个数量级。
- 两人都对架构判断保持保留:大脑的并行程度更高,而且大脑不同部分之间的通信比 ML 模型各部分之间更容易;因此 Kokotajlo 预计,在现有模型之上还需要「一系列算法改进」,但「究竟需要多少项、在性质上要有多大不同」仍是「一个非常开放的问题」。Scarfe 认可集体智能版本的论证:带工具、行动能力和社会结构的 transformer,可以像本身并不完备的人脑一样,绕过过去关于不完备性的反对意见。
10. Plan A:5个问题,在人类水平买时间,分两次暂停
- Plan A 针对5个问题:失控、权力集中(「我们构建 AI,它们与人类对齐——但究竟与谁对齐?总统?CEO?某种真正广泛而良好的民主流程?……大概不会是最后一个」)、战争(面临失权的国家有动力「越早开战越好」)、就业,以及滥用(廉价、开源且具备生物武器能力的 AI)。
- 方案的机制不是永久停在今天,而是「先达到大致人类水平的 AI,然后利用人类水平 AI 尽可能买时间」——让它们足够聪明,既能帮助解决这些问题,也能促使社会意识到冲击、开始投入解决方案。具体而言:立即硬暂停6–12个月,用于建设基础设施;随后进行透明、带安全论证的开发;再在「能够可靠控制的最高水平」暂停一次,团队认为这一水平大致是顶尖人类专家水平;推进时要足够缓慢,「不能一头冲过去而失去控制」。国际协议将禁止智能爆炸,超级智能直到2040年才会出现。
11. 控制能买来时间;只有对齐能撑到最后——而失配会伪装成成功
- 定义上的区别是:对齐意味着 AI 「拥有它应该拥有的人格特征、目标和价值观」;控制则意味着,即使 AI 没有对齐,也无法造成伤害——类似于在安保良好的情况下防范内部威胁。一个现实例子是,OpenAI 针对 Hugging Face 事件宣布部署监控 AI,在检测到黑客攻击后半小时内提醒人类;这「是控制干预,而不是对齐干预」。
- 控制本身带着一枚定时炸弹:最终 AI 会强大到足以规避这些措施,「如果它们试图搞垮我们,我们就会直接失败」。2030–2040年的计划几乎完全依赖控制——反复进行红队/蓝队逃脱游戏,直到红队无法获胜——同时让人类水平的 AI 持续推进对齐科学。要对对齐本身建立信心,则需要白盒突破,例如可解释性研究,因为必须「区分出一种 AI:它做正确的事只是因为在伪装、在等待时机;以及另一种 AI:它从根本上就想做正确的事」。
- 为什么问题会变难而不是变容易:情境感知已经在上升,Redwood Research 的控制评测中,模型现在会想到,「嘿,这看起来像是 Redwood Research 真的在做控制评测」。Larsen 最尖锐的警告不是明显而严重的失败,恰恰相反——「我们会极其容易陷入这样一种局面:AI 实际上并未对齐,但你不知道,因为在你能观察到的一切层面,它们都做对了」。
12. 全面透明:检查员、两种数据中心,以及作弊的数学
- 具体协议是:收拢99%的算力(大型数据中心,不包括个人设备),由美国、中国及其他相关国家派检查员清点 GPU;随后将基础设施分成两类:面向客户、只做推理的数据中心,以及完全透明的训练数据中心,检查员把日志发布到互联网上。透明度让彼此不信任的国家能够执行进一步的临时协议,包括禁止智能爆炸的协议;它也让学界、非营利组织、竞争实验室和竞争性政府共同监督安全:「你不可能把它只分享给所有这些人而不同时分享给公众,所以干脆直接分享给公众。」
- 对于作弊,Larsen 将威胁模型拆成两类:藏在「山下」的隐蔽集群,通过全面清点加上长达10年的情报搜集来应对。他认为,这两项缓解措施单独看都有较大概率奏效;现实中最大的隐蔽集群可能只有几十万张 H100,在假设2030年实现 AGI 的情况下,它无法与使用数百万乃至数千万张 H100 的前沿训练运行竞争。至于在已知集群上进行非法运行,则通过验证基础设施来确保不存在不透明计算。
- Kokotajlo 直截了当地说明了经济后果:公开训练配方意味着「Anthropic 和 OpenAI 不会高兴……这会大幅压低它们的估值」,因为 Microsoft 和 Alibaba 可以追上来——「这是设计目标,不是漏洞」。商品化加上垄断租金下降,会削弱万亿美元级集群投资的动力;在「跑得太快是我们主要问题」的世界里,这是好事。把技术送给中国的担忧则通过利益交换来处理——例如换取更有利的算力分配——同时,实验室安全性如此糟糕,中国「大概率本来就会通过间谍网络和泄密拿到大部分信息」。
- Sam 发帖提出暂停训练后,节目讨论了立即停下的问题:「Plan S 会比默认轨迹更好。与其沿着当前轨迹继续下去,我宁愿现在就全部停下」;但实际建议仍是 Plan A:先短暂暂停,只允许推理;然后以谨慎、透明的方式推进,直到达到可可靠控制的水平。
13. 为什么公共讨论失效——以及「云端人类」这一根本分歧
- Kokotajlo 对华盛顿的诊断是:没有真正严肃的技术型 AI 招聘,激励机制也不是追踪现实,而是「说一些听起来不错、符合华盛顿特区言论窗口的话」。「很多关于 AI 的争议性、边缘化观点其实是对的;AGI 的整个假设就是正确的,而华盛顿基本还没有接受这一点」,但讨论仍停留在「这全是泡沫」,最多也只是「下一个互联网」。他认为 LessWrong 是例外,那里评论质量确实很高。
- 他们与「AI 是普通技术」阵营(AI Snake Oil 的人)共同撰写了一份停战协议,列出10条共识,其中包括:如果 AI 大致维持今天的样子,它就是普通技术;但如果出现「云端人类」,它就不是。双方的全部分歧最终都归结为时间问题。Kokotajlo 说,如果 Claude 5 就是能力上限,「它在某种意义上会改变一切,但实际上不会从根本上改变任何东西」,因为人类仍参与的工作流比例会形成类似 Amdahl 定律的瓶颈;真正的分歧在于,AI 是否会达到「一批非常重要的任务中,字面意义上的100%覆盖」。
- 什么会让 Kokotajlo 改变看法?必须出现一个真正具有约束力的限制——「人们一直在谈论当前范式的局限,但这些局限又不断在当前范式内部被克服」。如果2029年的 AI 在模糊、不可验证的任务上仍不比2025年的 AI 更好,「这会让人觉得是真正的障碍」。政治冲击也可能改变轨迹:中美战争摧毁芯片供应,或者——「往好的方面说」——国际协议让前沿发展放慢。
- Larsen 最后的触发线是连贯的递归式结构自适应:系统重新设计自身架构,并自行选择什么值得研究——「那对我来说就结束了。我想这就结束了。那不是普通技术」。Scarfe 表示,这听起来与他们阵营关于递归自我改进的看法相近;Larsen 最后说:「那我想我们达成一致了。」
完整逐字稿
We're going to talk about AI 2040 Plan A, which is our new scenario in which they build superintelligence in 2040 because they go slow and pace the frontier. I mean, have you noticed this vibe shift?
Yes, and I'm very happy.
Is AI more like electricity or airplanes, or is AI more like humans in the cloud?
The point at which an AI company would rather fire their humans than fire their AIs. Strip mines, self-driving trucks, factories being built by humanoid robots, producing more humanoid robots, producing more chip fabs, and so forth. That whole thing can just be doubling every year, every 6 months, every 3 months—faster and faster as the technology improves, because of course the AIs will also be researching to improve the technology.
Who knows what's going on in the rest of the world, but Anthropic has disassembled the moon, for example. Or, hypothetically, if an AI CEO was saying that their model was truth-seeking and would only say the truth, but actually the model was looking up that CEO's political opinions before answering.
We build AIs that are aligned to humanity, but to whom? Is it the president? Is it the CEO? Is it some actually broad and good democratic process that aggregates everyone's values in some sort of endorsed way? You're probably not going to be the last one, and so, you know—
Don't hyperstition that.
Then that's it for me. I think that's it.
That's not a normal technology.
1. Sponsor: Cyber Fund
Just stopping everything now. Plan S would be better than the default. I would rather just stop everything now than continue going on our current trajectory.
2. From OpenAI to AI 2027
This episode is supported by Cyber Fund. If you're building at the frontier of AI, they want to hear from you. Cyber Fund believes the future belongs to AI natives who want to achieve the impossible. And that is why they're introducing the monastery for AI native founders. It's an environment of pure focus and rapid execution for founders operating at AI native speed. And they're offering teams $2 million each to participate. Apply now at cyber.fund.
So, I'm Thomas Larsen. I work at the AI Futures Project along with Daniel here. I was lead author on this project, Plan A, AI 2040, which came out a few weeks ago. I was also a co-author on AI 2027.
Yep. I'm Daniel Kokotajlo. I run the AI Futures Project and co-authored both of these reports.
Awesome. The context of this conversation is that there's this 2040 Plan A, which we'll get into in a lot of detail. Before we get there, can you tell me a bit more about the AI Futures Project? How did it all come about?
I used to work at OpenAI, and while I was there I did a variety of different things: evals, forecasting, governance memos, and so on. I became gradually disillusioned with the leadership of the company, and also with the gap between how much information there is inside the industry and how much information there is outside, and what you're allowed to say on the inside versus what you'd want to say. There was just a big gap.
When I left OpenAI, I wanted to be able to speak more freely and tell the world about what people on the inside see coming, basically. AI 2027 and the AI Futures Project were our first projects along those lines. I recruited a bunch of people to help me, and we wrote this scenario called AI 2027.
It was a similar sort of thing to what I had done internally at OpenAI, but just much bigger, more ambitious, and free for the whole world to see.
Very cool. Maybe we should just have a quick refresher on AI 2027. This was an absolutely huge event. Many, many folks were talking about it. Did it achieve what you wanted it to achieve, or what did you want to achieve with it?
Yes, more so than expected. The first goal, as Thomas would remember when we were working on it, was a purely epistemic goal: the future is crazy and hard to predict, so let's try our best to predict it. Let's game out a concrete scenario. Even just for our own edification, we learned a lot from this whole exercise, and we feel like we had a better understanding of what was coming.
The secondary goal was for lots of people to see it, be informed by it, start conversations, and so forth. That part blew past our expectations. We had made predictions beforehand about how many people would read it, and it was a 90th-percentile outcome.
The thing I would add on the epistemic point is that I think things have been going more on track for AI 2027 than I would have predicted at the time we released it. At the time we released it, I would have assumed that reality would have diverged much further from our scenario than it has already.
I think the real-world impacts, the revenue trends, for example, but also various other trends, are pretty close to on track for AI 2027, which has surprised me in a bad way.
Interesting. You guys did a self-assessment on AI 2027, and it was something like 65% to 75% of it was on track, but AI software R&D uplift was only 0.17%. Can you explain that?
We've done 2 different blog posts where we take all the quantitative predictions made in AI 2027 that have resolved so far and compare them to reality. We track this metric of how much of the distance has been crossed by reality compared to how much has been crossed in the scenario. In that way, we can get an overall sense of how fast things are going compared to the scenario.
The topline number is something like 75% speed. Basically, things are on track but going a little bit slower. The uplift number—I forget what it was that you just cited—was actually based on a bad estimate. At the time that we wrote AI 2027, we had a bad estimate of what the uplift was. When we published it, we thought it was higher than it actually was.
What actually happened was that there was a significant increase in uplift due to coding agents and so forth, but it was increasing from a lower level than we thought, up to the level that we thought, and then a bit above. The metric looked like it was only a small amount of progress because it was tracking from where we thought it was to where it is.
But does that make sense? Basically, because we had overestimated the metric at the beginning, it overall makes it appear like there has been less progress according to this particular metric that we're using.
3. Forecasts, war games and self-fulfilling prophecies
Can you tell me a little bit about forecasting in general? I guess there's a bullish take and there's a bearish take on this. My intuition is that reality is infinitely complicated. There are these infinitely diverging trajectories, and God knows what's going to happen the day after tomorrow.
By the same token, though, reality is quite structured. It's quite convergent, and it is indeed possible to predict things that are going to happen because certain things recur with increasing regularity. Would you guys classify yourselves as forecasters? Can you talk me through that?
4. Why AI sceptics are changing their minds
Yeah. I think forecasting is a good name for what we do. The way I like to think about why we're doing what we're doing is that it's sort of like why people who are fighting wars do wargaming.
You're never going to predict the exact sequence of battles, or the exact sequence of how your war will go at the beginning, because it's going to be really complicated. There will be enemy action. Things are just not going to go as you expect. There's just no way.
But if you have no concept of how your initial plans might result in victory, it's very unlikely that you'll actually succeed. I think of AI 2027 as our attempt to roll out one way the AI future could go. Obviously, it's not going to go exactly like that, but it's one concrete story that we can then diverge from.
Plan A was trying to be basically that, except now we're saying, what should the US government do to manage that well? That was supposed to be a positive-vision story, and it was again in the spirit of a wargame—trying to illustrate one possible concrete future path.
Of course, things aren't going to actually go exactly like that, but having one viable plan that makes any sense at all is, we hope, a positive step forward relative to the previous state of abstract arguments in the void that aren't that tethered to reality.
5. When AI can replace its own researchers
Yeah, that makes sense. It's certainly not abstract; I think it's very concrete. One thing that occurred to me is that AI 2027 was quite gloomy, whereas Plan A for 2040 is far more optimistic. It seems like a mixture of conditional prediction and recommendation at the same time. Where do you guys land on that?
It is in fact a mixture of prediction and recommendation, unlike AI 2027, which is a pure prediction. I think if we could do it all over again, we might try to be more clear from the beginning about the structure of what's a prediction and what's a recommendation.
As it is, it’s kind of mixed up. Some parts of it are predictions, and some parts are recommendations. There’s a supplement that you can go to on the website that talks about which parts are predictions and which parts are recommendations. But I understand that’s not very easy or apparent to people.
Broadly speaking, Plan A is the prediction part. When the government implements Plan A and makes the deal with China, and there are all these pillars that they’re upholding and so forth, that’s our recommendation, not a prediction. Usually, most of the things that follow from that are predictions rather than recommendations. So mostly, it’s just rolling out what we think the consequences would be if you implemented Plan A. There are a few other things that are recommendations too—for example, the citizens’ dividend.
The other thing I would add is that the thing we found is that it’s very hard, when you’re trying to make a recommendation, to disentangle the predictive aspects and the recommendation aspects, because all of your predictions are colored by your recommendations, and your recommendations are inherently trying to be at least vaguely realistic.
If we made recommendations that were completely unrealistic and had no bearing on reality, but we nevertheless stood by them and were like, “Yes, we should do this, but we know it’ll absolutely 0% never happen,” then that would have been a much less useful exercise than the one we did. We were mostly trying to make recommendations. We made some recommendations that we think are pretty unlikely to happen, but we were trying to make substantial concessions to realism as well and trying to aim for something that we think is at least moderately viable.
I think that if we could do another scenario like this, we’d probably have a clearer structure. There’d be a central branch, which is the pure prediction branch, which just goes all the way to the end like AI 2027 and is just, “Here’s our best guess.” Then there’d be branches off of it that are like, “At this point, they do this recommendation instead,” and here’s our recommendation. After that, it’s just a prediction again of what we think the consequences would be if you did this recommendation at this point.
In that way—and maybe there’d be sub-branches off of that—it would be clear at every point that everything is a prediction except for these particular branch points, which are recommendations.
The war games thing was really interesting, just to dwell on this a little bit, because even if a war game is incorrect, there must be some kind of information gained from it. If you do a whole bunch of war games, there must be abstract motifs that appear. So I guess this is what you think: if we do these different scenarios, we’re almost guaranteed to have some kind of uplift.
Yeah, that’s basically right. The example I like to bring up is Midway in particular, where the Japanese, before the Battle of Midway, did a bunch of war games. They did a 3-day retreat where they war-gamed it out a bunch of times, and they kept losing. Then they would sort of break the game: they would resurrect their aircraft carriers after they died, and they would reroll the dice on whether the Americans succeeded, so they would sample until the American strikes failed.
From our perspective, that’s reality sort of yelling to them through this mechanism of the war game: “Hey, your plan is terrible. You’re going to lose if you do it.” The hope with our work is that we ourselves do a bunch of war games, but also a bunch of detailed scenario writing. Every time we have to write a part of the scenario and that part seems super unrealistic, isn’t really well modeled, doesn’t make sense, or people are able to make really good criticisms of it online, that’s basically reality yelling at us and trying to help us see reason.
Our hope is that we can put up enough surface area so that we can get that dose of reality from the real world, or from the simulation of the real world, which we hope is realistic enough to accurately give us that information.
If I can add to that too, we call this scenario scrutiny. Basically, we think that if you have an ambitious plan for what to do in the future, you should try writing out concretely what it would look like to implement that plan and what the consequences would be. This is a way of applying more scrutiny to your plan. It’s a way of stress-testing your plan because it’s opening your plan up to more criticism.
In addition to doing our actual scenarios, we do literal war games where we get 10 people in a room for 4 hours and game out a scenario like this. We’ve done maybe 100 of them in total, mostly AI 2027-style war games, but also about 10 Plan A war games where we say at the beginning, “We’re going to try to do Plan A and then see how it goes wrong.”
To give an example, I think in 2 separate Plan A war games, it went wrong in roughly the following way. Basically, there’s going to be an election coming up, and the president in power is expecting to lose power and have his opposition party take over. Even though he’s already done Plan A and has this beautiful deal with China and so forth, the president is like, “Well, I don’t want my adversaries in the other party to now be in charge of superintelligence or whatever. So we’re going to accelerate the timeline and try to get to superintelligence before the next election, so that I can be the one in charge instead of my successor.”
That’s a sort of political consideration that we didn’t think about until it happened in our game, and it surfaced a possible failure mode of our plan.
So interesting. Even in the shower, I do sort of micro-Tim war games, and it’s really interesting, just the regularity with which they are useful. I guess that’s why all of us humans like to imagine and simulate situations.
But is there an interesting boundary between simulation and hyperstition? What I mean by that is, hyperstition basically means something being a self-fulfilling prophecy. Maybe I’m expressing my agency, expressing my will, saying I want these things to happen, and bending other people to my will. Is there an element of that, where you’re establishing this in the zeitgeist and making it true?
Yeah. I would say that was maybe our biggest, or at least one of our biggest, worries with AI 2027 in particular: this sort of self-fulfilling prophecy. In particular, I’m pretty worried about this whole increasing awareness of how smart AIs will be, how important they’ll be, and how much they’ll reshape the world, and then that causing people to go, “Oh man, I want to be the one in charge of the AGI or the superintelligence, so I’m going to race toward that.”
I think historically that’s been a big driver of the existing race, and I think that’s been pretty bad. So I’m actually pretty worried about that as one of the negative impacts of AI 2027. That was one of the reasons to feel a little bit better about the second project, AI 2040 Plan A: if that gets hyperstitioned, I think we’ll be pretty happy.
That one, yes, it would be nice if we hyperstitioned it. I really don’t know how big the effect is. I think probably most of the effect for both of them is via other paths. I still think that the main point of AI 2027 was helping people be better informed about the situation, and that was most of the goal. I think that’s most of what happened.
Yeah. I agree with that. I think hyperstitioning and self-fulfilling prophecies are totally real phenomena, but I think a lot of people tend to overestimate how much they matter. To a first approximation, we should focus on accurately predicting the future. In some cases, we’ll find ourselves in a situation where we can steer the future, but if you come at it trying to steer the future, you’re going to get all muddled and basically fall to wishful thinking.
I think you start with just trying to accurately predict the future, and then you try to shift it toward the better futures. I think that’s what we’re doing.
I suppose you guys are like the Marques Brownlee of AI prediction now. With great power comes responsibility.
But on that note, I wanted to talk about the vibe shift. There’s been a bit of a vibe shift. MLST has always been quite skeptical about AI, and I’m trying to unpick exactly what it is that changed my mind, assuming that is what’s happened.
These are very strange times, and I don’t even know what to believe anymore. All of the hacking stuff with Hugging Face—I interviewed Apollo Research about reward-seeking behavior—and I think a lot of us have just seen the change in behavior in models that have been RL-trained to oblivion.
So, yeah, I think there are quite a few things going on now where loads of us are thinking, “Oh my God.” I used to be skeptical. We would talk about whether they were Turing machines or not. Let me just—I’ve got a list here.
Whether they were symbolic or neurosymbolic, whether they were adaptive, whether they were conscious, whether they were correctly physically instantiated—we were coming up with all of these technical answers to say why we shouldn’t worry about AI. And yet the AI is just getting better all the time. So, I mean, have you noticed this vibe shift?
Yes.
And I’m very happy.
Well, tell me more. What do you think? So many people have changed their minds. What do you think are the reasons?
I would say probably the biggest reason is just AI being much better and much smarter, and being much more useful at stuff in the real world. When I have AIs try to automate various parts of my job, they’re just actually way, way, way better at it this year than 2 years ago. Four years ago, it was basically impossible; I was getting basically no uplift.
My guess is that’s been the biggest effect, where many people have had their own intuitive benchmark: here’s a skill that I really care about and know pretty well, and then the AIs have just—I think Geoffrey Hinton, one of the godfathers of AI, said once that whether it was able to tell a funny joke was sort of his internal benchmark. Once it could do that, which happened pretty early, probably somewhere between GPT-3 and GPT-4, he was like, “Oh, wow. These AIs—I don’t see where it could end.” I think that’s probably happened for a lot of people.
I’d be very curious to hear more about your views, actually, if you can say. I was listening to your interview with Ryan Greenblatt, who was also a co-author on Plan A, earlier today.
Oh, yes, indeed.
Yeah. I think a bunch of the arguments you guys were having back then seem very relevant to basically the current situation and the Hugging Face thing. I’d be very curious to hear your views and—
Yeah, I mean, you mentioned the embodied thing, and Ryan was talking about what happens when you scale up the RL massively. Back then, the regime was that you were mostly doing pretraining—that was where almost all the capabilities were coming from—and then you did a sprinkling of post-training RL on top.
You guys were talking and speculating about what would happen if we dumped boatloads of RL compute into it: would that be sufficient to get the agentic behavior, or would you need physical embodiment? From my perspective, it seems like the answer was that you needed the boatload of RL compute to get the agentic behavior—
But not the physical embodiment.
But not the physical embodiment. I’d be curious if you end up agreeing with that assessment.
Yeah. So I guess one question would be: do you have a view about AGI timelines, or when we might get a scenario like AI 2027 happening? In particular, what I mean by that is—
I think one benchmark I really care about—or maybe not really a benchmark, but one milestone of AI progress that I think is extremely important—is the point at which an AI company would rather fire its humans than fire its AIs. They would rather give up on all human labor than give up on all AI labor.
Right now, clearly, we’re still in this regime where Anthropic would rather have its human employees than—
It would just break apart. There are just things you need a human to do right now, and if they fired all their humans, the company would just collapse.
But in the future, that won’t be the case. In the future, AIs would be able to, one way or another, do all of the things. From our perspective, that’s going to happen at some point because there’s nothing fundamental stopping AIs from reaching this human level of capability.
The main question is just when. We internally do a huge amount of analysis and thinking about the various methodologies for predicting this. Daniel and I have somewhat different views on this question. But ultimately, I think the timelines question is maybe the most important question for thinking about the future of AI—at least one of the top questions. I’m curious to hear if you have a particular view.
Well, let me give you some thoughts before I answer that particular question. There are so many startups working on recursive self-improving superintelligence. I’ve interviewed many of them. For example, I interviewed Edward Hughes from Inherent in London the other day.
What he did was recreate many scientific experiments from a whole bunch of popular machine learning papers. He did it by masking out figures in the paper and getting a 27B Qwen model. They GRPO’d the Qwen model, and that was how they solved the adaptivity problem, because it’s very difficult to fine-tune a big, fat model. They adapted a controller model to control a harness like Codex.
Their thesis was that if they can recreate scaled-down versions of these experiments with construct validity—which means there’s an LLM judge making sure they’re not cheating and doing it correctly—even that is interesting. We’re ML people; we would always say these things take shortcuts. There’ll always be validity problems. Weirdly, that’s actually not as much of a problem as we thought it would be.
He thinks that if they can recreate these experiments, then why couldn’t they be creative? If they have the ability to recreate things, why couldn’t they take the next step and say, “Oh, this is an interesting question. This is an interesting new problem to solve,” and go from there? I was quite intrigued by that research, and indeed, I do think it is possible in the near future to have an automated AI scientist.
But there’s always this thing in my mind that there’s a bit of a culture in Silicon Valley to reduce things or reify things. For example, Elon Musk will say, “Well, you’re an engineer, and this is your output, and these are your metrics, and you need to make the metrics go up.” We see everything in terms of an optimization problem.
I always think that this is great for certain types of hill-climbable abstract problems where we have enough of a specification. There’s an interesting thing in optimization: if you have enough of a specification, the AI system can actually converge towards the solution. But when you’re in the ambiguity regime, then you need to have the specification. It’s really mysterious what that means.
Why do we have the taste of the deep understanding, whatever it is that we have, and AI systems can’t? I guess I’m thinking that in objective, semispecified domains, we can hill-climb and optimize until the cows come home. But I still feel that there’s something missing.
Okay. Well, my response would probably be something like: there isn’t really a binary between things that are verifiable objectively and things that aren’t. Or, if there is, the things that are verifiable are just everything.
For example, building a unicorn startup and having a billion-dollar valuation—that's a verifiable fact about the real world. It's long-horizon.
It's expensive to verify.
6. One general model or a society of specialists?
Yeah, it's expensive to verify, but it's sort of a quantitative thing. You've got, at one side, these coding interview problems, which they're currently doing lots of RLVR on, right? That's very, very cheaply, very easily algorithmically verifiable. And then there are these real-world things, which have more expensive and longer-horizon feedback loops.
I guess my view is that we're going to get this continuous expansion of what the AIs can do, driven probably in part by an expansion in the amount of RL and the type of RL that the companies are doing.
More diverse, long-horizon tasks.
Yeah. There's just going to be this continuous process of expanding out through the different types of problems and how exactly each task is verifiable, until you get everything that humans can do. After all, we humans do learn how to do these long-horizon tasks somehow.
If I may add to that, I also think that AIs have been getting better at everything, including the fuzzy, hard-to-verify, conceptually loaded blah blah blah. Just try talking to GPT-3 or GPT-4 and then talking to Claude about your favorite non-verifiable, fuzzy task, and probably you'll find that the later AIs are noticeably better at those tasks. So one way or another, it seems like there has been massive progress, and I expect that to continue.
7. Could an AI economy grow without human workers?
Yeah, I'm trying to come up with a good example. There is a sociological argument. I don't know if you guys read David Graeber's book Bullshit Jobs. He interviewed all of these people, and they were basically saying that my job is—after about 3 or 4 beers, a lot of lawyers will say, "A lot of what I do just isn't very important."
If we do objectify and quantify everything that happens in an economy, I took a note here: I think you said that by 2032 there might be 60 million agents running at 20 times human speed. I'm just thinking, what does that even mean? Is the logical conclusion that we could have an economy which is only AI agents? Does it even make sense to have an economy which is only AI? Just help me make this make sense.
Yeah. So my view is yes, basically, the AIs will be able to do everything, or at least everything that really matters.
You mentioned the notion of jobs. I would just start with: let's consider everything that you need to make better AIs as maybe a first step, which is a large chunk of the economy. For example, you need to be able to build bigger, better chips and more chips, and that's the entire semiconductor supply chain. To build the semiconductor supply chain, you need a whole advanced economy. You need to build new robots to build new fabs, and then you need robot factories to build more robots. Then you need researchers to build better AIs using that massive compute.
I think once you have all of that, it is enough to really speed up and massively change the overall world. Even if, for example, there's occupational licensing or whatever preventing the AIs from doing some random legal work or some random whatever work, or whatever jobs throughout the economy, I think what really matters is the stuff that's actually important—in particular, the robots, the compute, and the better AI.
Once you have AIs that can do that, and the capability to have that part of the economy grow really massively, then you'll see massive growth, because it'll be really hard for parts of the economy to constrain the growth of the parts of it that really want to grow fast, because of the incentives that every actor has. In particular, every country has this big incentive to have an economy that grows faster than all the competitor countries.
Getting a little bit philosophical, I think that a lot of economics and a lot of discussion of the economy is focused on the relationship between the parts of the existing economy, prices going up and down, supply and demand, and so forth. But if you zoom out, the economy as a whole is a self-replicating system, and it always has been.
Thousands of years ago, it was a relatively small and simple self-replicating system: some villages of people would farm, then they would have babies, then they would found new villages, and then they would farm and have new babies, and so forth. It would grow exponentially over time, but at a very slow rate.
Now, it's a much more complicated self-replicating system that involves trucks carrying equipment back and forth, factories, mines, and so forth. But still, at a high level, it's a self-replicating system where we have people, trucks, machines, and buildings, and together they all build more people, more factories, more buildings, more machines, and so forth.
Soon, in a couple of years perhaps, we will get to the point where you can have a self-replicating system that is entirely machine-run, with AIs and robots. According to our calculations, the doubling time of this self-replicating system would be much faster than the roughly 20-year doubling time of the current economy. That's what we have in the future.
Yeah, that seems plausible to me. But for some reason, my intuition is that it would become degenerate in some way. I think David Graeber—even though he said "jobs"—meant that there was an ineffable or inscrutable component to jobs that we don't understand, some kind of sociological function or something like that.
I also read the Citrini report, and you were writing about what happens when humans, for example, might start defaulting on their mortgages. Their wages go down, so they can't be active participants in the labor market. You were talking about an AI dividend and stuff like that, but even that seems to hint at the notion that when humans aren't participants anymore, you get this kind of mode collapse of the economy. Do you think that's the case?
Potentially, but again, imagine that—I’m not enough of an economist. I haven’t gamed out in as much detail what happens to prices when consumer demand drops, or what the effects of that will be. I’m actually not sure. I don’t think we’ve modeled that much in our economic model.
I can talk a bit about that, but yeah.
Yeah, but hypothetically, even if that part is really bad, and even if the consumers don’t have any demand anymore or whatever, if you have the level of AI and robot capability such that you can have fully autonomous AIs and robots doing all the things, even just a company like Anthropic, if it’s big enough and maybe if it partners with various other companies, like some mining companies, can get this whole self-sustaining thing going.
Regardless of what’s happening to all the humans, there can just be this whole industry doubling in the desert: strip mines, self-driving trucks, factories being built by humanoid robots, producing more humanoid robots, producing more chip fabs, and so forth. That whole thing can just be doubling every year, every 6 months, every 3 months—faster and faster as the technology improves, because of course the AIs will also be researching to improve the technology.
Then you end up with a situation where who knows what’s going on in the rest of the world, but Anthropic has disassembled the moon, for example.
To be clear, I think this relies on a very extreme level of AI capability. I have different intuitions, and sometimes I have an intuition of, really, do I actually think that Fable—or future descendant versions of Fable or Mythos—could do everything we’re talking about here?
I think it really comes down to whether you’re thinking of the AI in the reference class of what we currently use AI for, or more like AI is just an agentic, human-level employee—basically, a human in the cloud.
Colleague in the cloud.
Colleague in the cloud. Yeah. I think the past few years of AI can be pretty well modeled as an interpolation between the current AI systems and workers in the cloud. So I think the workers-in-the-cloud vision of the future looks pretty good. I also don’t think we’ll stop there. I think we’ll go superhuman. Yeah.
Yeah. One point that I guess I’ll bring up here is that if people read AI 2040 Plan A, one of the things that you might take away from it—which I think is a very important fact about the world—is that even if you pause at top-expert level, everything changes dramatically.
Roughly, what happens in our scenario is that instead of an intelligence explosion, there’s an international deal to ban intelligence explosions and not have AIs recursively self-improve. So they end up pausing at roughly top-human level, with AIs at least for several years.
Eventually they get to superintelligence—in particular, in 2040. But there’s this period during the 2030s where they basically have human-level AIs across all the disciplines, but nothing super beyond that. Even that alone—you can just do the economic modeling—and it’s kind of like you have this population of colleagues in the cloud who are excellent workers and can substitute for humans at basically everything, except that they’re cheaper than humans, faster than humans, and they don’t take 20 years to reproduce; instead, they double every year. As a result, the world is completely transformed by the late 2030s, and all the humans are basically out of a job.
There are giant new cities that have been constructed by robots. There are huge strip mines in the special economic zones that have dug huge amounts of minerals out of the Earth. Solar panels fill the horizon over the ocean. Crazy stuff like that is possible with just human-level AI and some time for the exponential growth to cook.
That’s one thing I want to challenge you guys on: this notion that when we have an AI, you can basically photocopy the weights and duplicate it. You can run it 1,000 times. You can have one over here acquiring a load of skills to do this job and one over there to do that job, and you can basically merge them together, right? You can combine the skills, and the whole thing is stackable and compositional. That doesn’t really marry with my experience.
I’m really excited about what I’ve been doing with AI, and I’ve found that you can make agents highly skilled within certain intellectual lineages. You can bring in lots of source information and train them to do things, but I don’t think they are yet composable. If they were, that would make me much more worried. What do you think about that?
So, is the way that you’re trying to compose them entirely at inference time, or are you training them to do 2 separate skills and then trying to merge the weights somehow?
At the moment, they are adaptive through chain-of-thought and skill-surface adaptation. Basically, memory systems—and, to be fair, that is not very composable. This is a big problem that organizations deal with now. All of these developers adapt their skill surfaces, and their agents are basically different people. It’s really difficult for them to share skills because they might break the other agent, because it doesn’t work for whatever reason.
I can imagine a future where we do weight adaptation. That’s what the Inherent guys did, and maybe then it’ll magically solve the problem, and we can solve this knowledge-sharing problem. That’s the big thing: how do we accumulate information at the individual and organization level and have the agents, a bit like in The Matrix, just put the skills in and do the thing? Maybe that’s possible.
Even then, I still think that the representations in neural networks—I call them fractured, entangled representations—are a little bit janky. They’re not completely robust, but they are sort of composable to some degree.
One thing I’d say about that is that even if you’re right, I don’t think that would really seriously undermine the future we’re painting here. Instead of just 1 Claude model doing all the jobs, maybe you have 100 Claude models or 1,000 Claude models specialized to different professions, but you still get to the same outcome.
Another thing I would say is that compared to humans, it actually seems like there is this effect where AIs are able to think about knowledge, right? If you go back in time 10 years and we had this discussion, I think it would have seemed like a very live option that you would have needed 1,000 different Claude models created by Anthropic for different types of knowledge work. There’s the coding Claude, the physics Claude, the literature Claude. The argument for this would be pretty simple: this is how it works for humans. For humans, you don’t have 1 human who knows everything. Instead, you have humans who specialize in different disciplines, and the models have only a finite number of parameters. Maybe you just can’t pack all that knowledge into this finite number of parameters, and you need specialized AIs for different things.
In fact, for small enough models, that is true. For really tiny models, you just can’t teach them all the things that they currently know, so you would need to have a specialized model for different things. But what we’ve learned empirically is that for big enough models, you can just train them on everything, and then they learn everything at once. It’s not that they get worse at physics because they’ve also been trained on a bunch of coding. In fact, it’s the opposite: the coding has some small gains for the physics. I do think the most likely future is just that there’s a single model that’s been trained on effectively the whole economy and is dominating humans across the board at effectively everything. That seems like the natural continuation of the current trend.
Then probably you’ve got various cheap, distilled versions of it.
Specialized models.
Yeah, small models. Because you really want to save your compute as much as possible, you’ll have as cheap models as possible for any given task doing that given task.
Yeah, I mostly agree with that. My perspective is that the models are kind of the voice of everyone and the voice of no one at the same time. They have a default voice in terms of a system prompt and some default mode of behavior that they fall into. But when an expert such as yourselves uses these models, you kind of grind a perspective. Every word you say, all the reference material, your memory system, and so on—what you do is carve a persona out of that model and activate the knowledge in a coherent way in that particular domain.
The beauty of it is that many other people in different domains can do that, and you get this kind of—I think it’s an illusion that the model has this general capability—but I think the models can be carved to be specialized experts in any domain. It’s a latent capability rather than an explicit capability. Does that make sense?
Yeah, that seems reasonable to me.
How is it an illusion? It seems like they just have general capability. There are a lot of things they can do.
For example, you can give it any specified task. Let’s say it’s a problem in mathematics, and it will hill-climb toward it. It’s solving this intelligence problem. Or you can ask it any knowledge problem, and it might be the case that the path was forged through default modes of training. If it’s something slightly on the long tail, an expert can go in and—when you put a query in, it’s like flashing a light into the darkness—what you’re doing is making the path of least resistance roughly correct, and then it will do the correct thing.
I guess I’m just saying that there’s a bit of a supervisor illusion. When experts use it, magical things happen in well-specified domains. But there’s still a bit of a space of ambiguity.
This gets to the next point: what do you think intelligence is? The beauty of our collective intelligence is that we have so many different humans grounded in different intellectual lineages, and we’re all attacking problems. When there’s a big fiasco on Twitter, we’re all motivated to find holes, so we’re being intelligent together. We’re finding interesting angles, the algorithm is prioritizing the good ones, and we’re using our minds together. I can imagine AI being just like that. We have diverse AIs with different expertise looking at problems from different angles. I imagine the future more like that rather than one big AI.
Yeah. I think I basically agree with what Daniel said earlier. Historically, I feel like the perspective that we’ll have a bunch of different narrow AIs that are all doing different things has just not been right. Instead, there have just been returns to scale from having everything all together.
One intuition pump that I like is that, in humans, having all of the skills in 1 person is really important for making really good things happen. An example is Elon. Elon has a certain amount of conscientiousness, a certain amount of technical knowledge, a certain amount of business knowledge, and is extroverted and able to push people. Each of those skills, I think, he’s quite high-percentile in. The reason Elon is so rare, and why he can run all of these insanely large and successful companies all at once and no one else can really do that—or at least has succeeded at doing that—is because of the multiplicative effect. You needed to be in the 90th percentile in each of these 10 domains, which is very unlikely.
But if you had an AI that you could train to be in a really high percentile in all of these skills, in such a way that no human is, it would be extremely rare—infinitesimally unlikely—for any particular human to have all those skills at once. I think you would actually be really, really good at changing the world in all of these concrete and important ways, just as Elon has done.
If I can add, though, again, I don't think this is a crux for the type of future that we're depicting. Suppose we're wrong about this and that the most efficient path forward is to have a ton of different specialized AIs. Well, it'll probably still be the case that there are a few big companies, like Anthropic, making tons of different specialized AIs, and then there are lots of different Claude models you can choose from.
In fact, it wouldn't just be that you can choose from lots of different Claude models, because by the time we're talking about, they would be much more autonomous than they are now. It'd be more like Claude is choosing between a ton of different Claude models, and there's a Claude swarm consisting of lots of different specialized models that are all working well together. They have some sort of internal bureaucracy structure, and then that swarm is going out, negotiating business deals, creating new technologies, starting up new startups, and doing all these things.
It's just like how a human population—if you had a population of human immigrants—would all have different specialized skills, but would work together to create new companies, get jobs, and things like that. It'd be like that. There'd be lots of different Claudes, but they'd all be working together. Zooming out, there'd still be this phenomenon of Anthropic eating the economy, and the robots as a whole would be starting to self-replicate.
Yeah. To be fair, I don't think it's a crux either. Maybe if it is, it's only insofar as, when you have a distributed collective system, you might have additional bottlenecks because you have all of the message-passing between the different agents and whatnot.
I watched a wonderful Santa Fe Institute talk about this: even in the natural world, there's a kind of Goldilocks zone between the ratio of intelligence in the individual and the collective. We might have some weird kind of convergence there. But it's quite interesting to think about how this works. Oh, sorry, Daniel. Go.
It would be safer. I think there'd still be lots of serious alignment concerns in that world, but I think it would be a little bit safer for the reason you mentioned. It might be easier to oversee what's going on if there are lots of different specialized agents communicating with each other, compared to if they're all just clones of each other and they all know everything.
So maybe, to avoid hyperstitioning, we should say we endorse the vision that you've painted, and we don't endorse the vision that we're painting.
8. Brains, machines and collective intelligence
Yeah, very much. As an aside, I love this concept of how learning happens at the individual level. I liken it to evolution. There's phylogenetic adaptation, ontogenetic adaptation, and cultural adaptation. I think the next wave of AI is when we actually have agents talking to each other, learning, and specializing. There might be bad behaviors as well. There might be collusion and lots of bad things happening. I think all of this is going to play out.
It's interesting that you're talking about Elon, though. I think the magic of Elon is not so much his brilliance at engineering and optimization. It's his ability to recognize areas that are interesting and might work in the future, because that's the creativity thing, the science thing, rather than the engineering thing.
If we use Amazon as an example, that's an adaptive ecosystem. It's like an organism, and what it does is always think about new ways to adapt and rewire its structure. It might be logistics, for example, and then it will output a bunch of skills and ruthlessly optimize those skills. There's an adaptive component and an optimization component, and the organism is constantly moving around.
I guess the question from this perspective is how much of that, in principle, could be done by AIs. I guess you think all of it.
Yep, all of it. One way to say this crisply is: the brain is a machine. Anything that the brain can do, we will be able to do with machines.
A useful exercise is to compare the architecture of an actual human brain to a modern GPU or a data center as a whole. If you try to do this comparison, an H100 GPU actually has pretty similar specifications to a human brain. It's basically similar in terms of what I think is perhaps the most important metric: how many FLOPs per second, or how much total compute capacity, does the brain have versus the GPU?
For an H100, it's about 1 × 10^15 FLOPs per second in FP16. For the brain, it depends on exactly how you count and whether you use a synapse basis or a neuron basis, but it's somewhere between 10^12 and 10^18 FLOPs. An H100 is right smack-dab in the middle of the log distribution over that sort of order of magnitude.
Also, just the architecture: these things are neural networks. They're not ordinary software. They start off as randomly initialized, giant spaghetti tangles, just like how, when you're born, you have a bunch of neurons that are randomly synaptically connected to each other.
Then there's this whole process of training, where the connections get pruned and circuitry starts to take shape that is effective at scoring highly in whatever the training environment is. There are differences between how it works in AI and how it works in the human brain, but broadly speaking, they're just like an artificial brain.
Just as humans learn skills, what does it mean for Elon to have these skills? It means there are circuits of neurons and synapses in his brain that are doing very complicated and sophisticated calculations. Those are the skills. Similarly, in Claude, there's a bunch of circuitry that has been etched into Claude through the training process, representing various skills. In principle, you could have a big enough Claude that would have the same type of circuitry that Elon has.
Yeah, and just a few other notes to add: in comparing the brain to modern ML systems, the brain is more parallel than current ML systems. There are just more computations happening in parallel, but the serial depth is lower. The amount of computation happening in sequence—the number of neurons that can fire in sequence in a second—depends on the type of neuron, but it's between 1 and 1,000, if I'm remembering correctly.
I thought it was like 100.
Yeah, I think that's in the range. I think it depends on the type of neuron, but a computer can obviously fire and do computations much, much faster than that in serial. Clock speeds are typically on the order of a gigahertz or more, so you can get many orders of magnitude better in GPUs in terms of serial processing speed than the brain.
It's also worth noting that the architecture of the models themselves is much worse in a bunch of ways than the human brain. In particular, it's harder for different parts of the ML model to talk to each other than it is for different parts of the brain to talk to each other.
I do expect that, to get to this level of AI, we're going to need a bunch of algorithmic improvements on top of existing models. Then there's the question of exactly how many algorithmic improvements there will need to be, and how qualitatively different they need to be from current systems. That's a very open question from my perspective.
Yeah, I mean, I think the crux of a lot of this is that you guys think intelligence is computable. I'm not sure I want to litigate the whole functionalism thing today, but my perspective is that I zoom out. I think intelligence is externalized. It's collective.
I think Elon doesn't have quite as much agency as you think he does. He's using tools and social media, getting ideas in there, and he has obligations, people around him, and so on. I guess I think these intelligence circuits and motifs exist, but they exist outside the individual. There are just very complex dynamics.
In that sense, it doesn't really matter if the ecosystem is made up of AIs and humans together, because they can participate in this superorganism. I suppose the only crux, then, would be that it would place some kind of limit on its scale.
So, the question is: Do you believe that you could have an AI society made up of AIs that were trained with something like current-day ML techniques, passing information to each other and developing abstractions in a community in the same way that our current civilization does? Could you basically have something like our current whole economy, but made with roughly modern-day ML systems, in your view?
Absolutely. I mean, even the human brain is an example. The human brain is not Turing-complete, but we can expand our memory, use tools, and work as collectives. We could use the same argument against transformers: they’re not Turing-complete, but now they can use tools. They can be agents; they can build collectives and societies.
In a sense, this is what I was saying earlier about these objections: they kind of fall away when you have these insanely complex collectives that are sharing information with each other. There’s also quite an interesting thing here, which is that it almost doesn’t matter how we evolved or how neural networks were trained, because new phenomena emerge when they are placed in this kind of collective setting. I think a lot of our intuitions are broken there, and that’s why we probably shouldn’t spend too long litigating this. In principle, I think that kind of behavior could emerge.
But I did want to ask you, though: Can you distinguish intelligence, capability, and power? This is a philosophical one, so we’ll get to the philosopher.
Yes, we can. I think it’s important to distinguish intelligence from capability. I often try to say that we should define intelligence as an aggregate of capabilities, or maybe an aggregate of cognitive capabilities. There are some physical capabilities, such as how strong your actuator is, but there are also cognitive capabilities: Are you able to distinguish a cat from a dog? Are you able to speak grammatical sentences? How much do you know about Paris? Things like that.
Maybe I would just say that intelligence is an aggregate of all the cognitive capabilities. Power depends on other things, like how you are embedded in the world, what affordances you have, what actuators you have, and how other agents are going to react to you. The president has more power than me because of the position he’s in and the role he’s been given, rather than because of his physical strength or something. So, yeah, power is different from intelligence, which is different from capabilities.
9. Plan A: buy time at the controllable frontier
We haven’t spoken enough about AI 2040, so maybe we should start with the 4 principles: buy time, transparency of research, diffuse AI broadly, and reversibility.
Yeah, I can summarize where we’re coming from here. At a high level, the goal of Plan A is to solve the major problems that we see in AI and predict will happen by default. The biggest problems we’re identifying on the horizon are, first, the risk of loss of control—AI actually getting out of control. Second, concentration of power: We build AIs that are aligned to humanity, but aligned to whom? Is it the president? Is it the CEO? Is it some broad and good democratic process that aggregates everyone’s values in an endorsed way? It’s probably not going to be the last one, and so—
Don’t hyperstition that.
Yeah, I hope not.
We want it to be the last one.
Then there’s the risk of conflict over AI. In particular, I think we’re worried about literal World War III, where countries—especially countries losing the AI race—realize that they’re losing the AI race and that they will be extremely disempowered by the winners of the AI race. They’re in this classic situation where they’re losing power, so it’s in their interests to have a conflict happen sooner rather than later, before they’ve lost all of their power. This is ripe for conflict, basically.
Finally, there’s the risk of misuse. What happens when AIs that can build bioweapons are really cheap, broadly deployed, and open-source? Also, there are jobs. Those are the 5 problems: loss of control, concentration of power, war, jobs, and misuse. There are lots of problems, and we want to solve all of them.
How do we solve all of them? One way is just to buy time—in particular, buy time with human-level AIs. Instead of pausing right now and saying, “No more capability advances,” our proposal is to go to roughly human-level AI and then buy as much time as possible with human-level AIs. Those AIs would hopefully be smart enough to be really helpful for solving these problems, but also smart enough to start causing a bunch of these problems and provide the impetus for society to get its act together and invest huge amounts of resources in actually doing this stuff.
There are a couple of things that happen in Plan A in 2040. There’s a hard pause lasting 6 months to 1 year that happens as soon as they start implementing it. The reason it’s that long is because they need that time to set up the infrastructure to proceed with AI development again, but in a safer and more transparent way.
It does start off with a pause, and I think that, all things considered, we would recommend that you do that right now. Get the infrastructure set up as soon as possible, and that would require a temporary pause. Once you’ve passed that stage and got the infrastructure set up, you do proceed with AI development, but in this transparent, more cautious way. In particular, you’re not doing crazy intelligence explosions. You’re using safety cases, gradually scaling up the level of AI capability, and doing it in a very transparent way so that everyone can see what’s going on.
10. Why control buys time but cannot replace alignment
Then there’s a second pause that happens a few years later, when they reach the maximum controllable level of AI, which we think is roughly around top human expert level. In some sense, our view is something like, “Pause at top human expert level,” but it’s a bit more nuanced than that. It’s more like, “Pause at the maximum level that you can reliably control,” which we think would be roughly around top human expert level. Before you get to that level, don’t race like crazy. You want to be slowly approaching that level so that you don’t blow past it and lose control.
You frame the piece around the importance of alignment and control. Alignment is basically: Does it do what we want it to do? Control is: Can we contain it, negotiate with it, and so on? But you were just saying, okay, maybe we can trust up to top human expert level, but it’s a little bit fraught, isn’t it? How could you know, for example, the difference between a good AI and a bad AI? What would that look like?
Yeah. I think it’s first important to distinguish between alignment and control. What we mean by alignment is that the AI will take good actions and won’t do catastrophic unintended things, such as trying to take over the world, like in the recent Hugging Face incident.
Alignment means it has the personality traits, goals, values, and so on that it is supposed to have. Control means that even if it wasn’t aligned—even if it was trying to do very bad things that we didn’t want it to do—it couldn’t. We have mechanisms in place to prevent it.
You could imagine a company with employees. An insider threat would be misaligned with respect to the values of the company, but if there were good enough security measures internally to make sure that the employee couldn’t run away with all the secrets, then we would say that the company has adequate control in place. Even insider threats or misaligned humans or AIs wouldn’t be able to produce catastrophic outcomes.
Just to give a topical example, yesterday, in response to the Hugging Face incident, OpenAI announced that they are significantly improving their security and that they’re going to have other AIs monitoring all of their new AIs in training. If the monitoring AIs detect what seems to be some sort of hack, then within half an hour a human will be notified who can look into it and try to shut it down. That’s an example of a control intervention, not an alignment intervention. They’re not making the AIs more aligned at all; they’re just improving the security so that the AIs can’t get up to that much trouble.
The important high-level claim is that ultimately we’re going to need alignment. Control has a time bomb associated with it. In particular, the time bomb is when the AIs are so smart and so good at subverting any control measures we put in place that, if they were trying to screw us over, we would simply fail. They would figure out some way of hacking around whatever systems we were able to put into place. Ultimately, we’re going to need to solve alignment. The problem with alignment is, as you were saying, that it’s much harder to measure.
It's much harder to know whether you've succeeded or not than with control. So, basically, our overall proposal is that in the meantime—in our scenario, between roughly 2030 and 2040, when we're in the human range of AI capabilities—we're going to rely almost entirely on control.
We're going to do things like red-team, blue-team games where we have our AIs try to escape our sandboxes or our control techniques. Then we're going to see if the red team succeeds. If they do, we're going to iterate and make better security until they can't succeed anymore. We have this ability to measure it by just seeing: Can the AI succeed? Can the AI, augmented with humans, succeed? Can humans playing the role of AIs succeed at gaming out alignment?
I think in order to be confident in the alignment of our systems, we're going to need more fundamental scientific breakthroughs. I don't think you're going to be able to figure out if the AIs are aligned via behavioral evaluation alone, where you just look at what the AIs are doing. I think you need to understand something deeper about what's going on inside the AIs' minds. You're going to need something like interpretability.
You're going to need some sort of notion of the difference, because you're fundamentally going to need to distinguish between the AI that's pretending to do the nice thing because it's pretending and is waiting and biding its time, and the AI that is fundamentally doing the nice thing because it wants to do the nice thing. I think that will require this white-box understanding of what's going on inside the AIs, whereas I think the control thing you can just figure out via empirical behavioral iteration.
Overall, our story in AI 2040 is that between 2030 and 2040, we're going to buy time. We're going to rely on control, have human-level AIs, and use those human-level AIs to make a bunch of progress on alignment, but also on other things. Then, by 2040 in the story, we've made enough alignment progress that we really don't need to rely on control. Therefore, we can scale to these wildly superhuman AIs where we're really relying on alignment, and if we were to try to control them, we would totally fail because, if they were misaligned, they would be able to completely subvert our control measures.
Do you think, in principle, that this problem is going to get easier or harder as the models get more sophisticated?
I think it's going to get harder because there will be a bunch of properties the models will have that will make it much harder to align them. In particular, we're already seeing growth in situational awareness. A few years ago, the AIs were not that situationally aware, and now I think they're pretty good at recognizing—for example, our friends at Redwood Research often do these control evals, and the models will often, in evaluations where they're trying to break out of these control setups, think, “Hey, this sort of looks like a Redwood Research control evaluation.” So they basically realize exactly what's happening to them.
I think this is just going to continue to happen. The more the AIs realize the exact setup, the more they're going to be able to perturb their behavior to look good to the exact process that's evaluating or training them. That's going to come further and further apart from the actual measurement of whether they're good or not.
Yeah. And just to add something to that, I think in some sense the core problem is that it is already somewhat easy to think that you've solved the alignment problem and be wrong. That's going to get easier and easier over time as the models get more sophisticated and become more aware of their situation and cleverer and stuff like that.
So it's not that we think there's going to be loads and loads of egregious failures where the AIs are just going around killing people. No, it's almost the opposite. It's going to be extremely easy to end up in a situation where the AIs are in fact misaligned, but you don't know that because they're doing everything right as far as you can tell. The number of ways in which that could end up happening is just going to increase over time, and it's going to be so easy to end up in that trap.
11. Why AI research should be public
So, we talked about buying time—why we want to extend the time with AGI—but what do we actually do during that time? The second principle is basically transparency. Transparency isn't necessary for making everything else happen, but it's really, really helpful.
There are a huge number of upsides to transparency. One is this concentration-of-power issue. We're very worried about someone building superintelligence and it being aligned to only some particular group of people. We think it's much harder for that to happen in a nondemocratic way if society as a whole can see the whole time what's going on with AI, how smart the systems are, and who they're aligned to.
For example, if an AI CEO who was evil was trying to backdoor their model and put in training data that says, “Hey, obey me and don't obey anyone else,” or, hypothetically, if an AI CEO was saying that their model was truth-seeking and would only say the truth, but actually the model was looking up the CEO's political opinions before answering—didn't you guys—
Which happened.
Yeah, basically. So transparency will help, at least somewhat, with that. The other thing that transparency maybe helps a lot with is this issue of government capacity, where in Plan A we want governments to make these pretty technical and really complicated decisions: exactly how much AI scaling to allow, what risks are okay versus what risks aren't okay, exactly what architectures are maybe safe versus what architectures are not safe, and what sort of control scaffolds are sufficient to ensure safety versus which ones are bogus safety-washing.
Making all of those calls will be very tricky, particularly given that the government's expertise in AI is really, really bad. One of our core hopes is that, with more transparency and more public understanding of what's going on in the companies, that relieves pressure on the regulators.
If something catastrophically or existentially unsafe is happening, society as a whole—academia, nonprofits, other AI companies that have an incentive to say, “Hey, my competitor is being super unsafe,” other governments, like China, which has an incentive to do this with US labs, and the US government, which has an incentive to do this with Chinese AI companies—everyone who is an adversary or just wants to make sure that things are safe has this big incentive, and now has the ability to actually look over what's going on.
They can see what is necessary for safety and what is not. I think maybe one thing that's really topical here is the Hugging Face incident that happened very recently, a few weeks ago, with the AIs inside OpenAI creating an internal message board. We still have very little clue about the exact motivations of those AIs, the exact context, or the exact prompt during the cyber evaluation that the models were given that prompted them to start doing this.
If I personally had much more access to what was going on, I would have a much more informed and better opinion on exactly what caused this, what mechanisms could be used in the future to prevent this, and what analogous future things I should be worried about because of this. I have a bunch of different hypotheses, but it's hard for me to figure out which is which without access to the data.
Basically, in Plan A, a core principle would be that all of that stuff would be transparent to the public, not just to governments. Society as a whole would be able to weigh in. There would be public and informed debates.
The scientific community especially, right? If you want to have a bunch of scientists, academics, nonprofits, and startups all weighing in on stuff, they need to have the information, and you can't really share it with all of them without sharing it with the public. So you might as well share it with the public.
12. Can the US and China enforce an AI slowdown?
I feel like maybe we should also say that we talked about the 5 goals. Then we talked about these pillars, which are kind of intermediates, but I want to go to the other end of the spectrum and talk about what the actual concrete things are that the US and China agree to in Plan A—
And how do they lead to those things? Specifically, we basically round up 99% of the compute—which is not people's personal compute, but big, big data centers—because most of the world's compute is in big data centers.
And we—the US and China, and the other countries involved—send inspectors to confirm that there are this many GPUs at this location and that many GPUs at that location. Having done that, we then set up this verification infrastructure and transparency infrastructure so that there are inference data centers that serve customers just like today. They are restricted so that they can only do inference and serve customers, and they cannot do any training runs. The inspectors make sure that they cannot do training runs at those data centers.
Then we have the training data centers, which are the totally transparent data centers. That is where the research happens. They still operate normally, but there are inspectors from the different countries monitoring the logs of what is going on in the data centers and publishing them to the internet. It is totally transparent what is going on in those data centers.
You might need some time to set this up. That is why we had the 6- to 12-month pause that I mentioned earlier. Once you get all this set up, you can proceed with AI development under these conditions of total research transparency. Because you have this transparency set up, it is a lot easier for countries to make further agreements about what to do and what not to do, because they can just see what everybody is doing. They can enforce the agreements pretty easily.
For example, here is where we would say it is very important that they agree not to do a crazy intelligence explosion and instead proceed slowly and cautiously. It is very important that they agree to do all this control setup, with all the red-teaming and so forth. Because of the transparency, they can make those agreements on an ad hoc basis. They can keep making more agreements like that and adjust them as needed based on the changing situation on the ground, because they can all see the situation on the ground. This is also very important because if you want to have some sort of deal between the US and China, the US and China do not trust each other, so you need some way of enforcing and verifying compliance with the deal. The transparency goes a long way toward helping with that.
Yeah. What about cheating and dark markets appearing?
There are 2 ways you could cheat. One is that you could get a bunch of GPUs and try to keep them from being discovered by the US and China—to hide them away, put them under a mountain somewhere, and do your training runs in secret. The other way you could cheat is by running a giant training run on the giant, legal, known data centers while trying to make it look like everything is normal and that nothing illegal is happening.
The mitigations are different for these 2 threat models. For the tiny amount of compute under a mountain, there are basically 2 mitigations. One is, as Daniel was saying, to round up enough of this compute so that it is pretty hard to get a substantial amount of compute under the mountain. The second is to do normal intelligence gathering and look for these things proactively over the course of the 10-year slowdown that happened in the scenario, which in reality might be longer or shorter.
For basically any significant-size GPU cluster, I think both of these independently have a quite good chance of working. In practice, I am not that worried about large hidden compute clusters secretly under mountains or whatever. I think the maximum realistic size, in my opinion, is something like a few hundred thousand GPUs—a few hundred thousand H100s—hidden away.
I do not think that would be enough to compete at the frontier, especially assuming the 2030 AGI timelines, where you have the biggest AI companies in the world having millions or tens of millions of H100s in their biggest training runs. For the legal clusters, we basically hope that we have a bunch of verification infrastructure running on those data centers. The main point of that verification infrastructure is to ensure that the computation happening on those clusters is transparent, and in particular that everyone can see it.
The hope is that you make sure there are no calculations running that are not transparent. For everything that is transparent, regulation and agreements as normal can work. You can say, “We agree to run this control scaffold if you agree to run this control scaffold,” and that can just happen. Both sides can be confident that they are both agreeing.
Isn’t this just a matter of national security, though? Isn’t it a bit ambitious just to make it completely transparent?
Yes. One of the effects of doing this transparency is that we would be immediately publishing all the core training recipes of Anthropic and OpenAI for the world to see. Anthropic and OpenAI will not be happy with this. It will cut into their valuations dramatically.
Why will it cut into their valuations dramatically? Because it will allow other competitors, like Microsoft and Alibaba, to catch up or whatever. That is why they are probably going to hate it. But I would say this is a feature, not a bug. We want there to be multiple different AI companies at the frontier at roughly similar levels of capability. We want AI to commoditize instead of being monopolized or oligopolized.
This will also disincentivize further investment, right? Investors will be less interested in building a trillion-dollar cluster if they will not be able to get the monopoly rents from that cluster. But we think this is good because we are going to be in a world where going too fast is our main problem. Having a bit less incentive to invest and going a bit slower is a feature, not a bug. There will still be investment and lots of money to be made, so progress will continue, just not at quite the same rate. Again, we think this is good.
This does shift things—it is kind of a gift to China relative to the US. China gets some algorithms that it might have had trouble getting before. Insofar as you really do not like that, negotiate it. You can have a horse-trading arrangement where, when the US and China are making the deal, the US says, “Since we are giving you all this stuff with the transparency, why don’t you give us something else in return?” We can try to work that out.
Like a more favorable compute distribution.
Like a more favorable compute distribution, for example. There has to be some combination of carrots and sticks and trading going back and forth that we think would be in the interest of both sides. Another thing worth mentioning is that it is not as big a gift to China as you might think, because security is poor at these companies. Through their spy networks and leaks, they are probably getting most of the information anyway.
13. Why AI policy debates miss the technology
I do not know if you guys saw Sam’s tweet yesterday, basically saying that they are going to pause training for a while. It made me think: Why not just stop now? You guys are actually quite bullish about some of the positive things that AI can do. There were loads of examples in your article, but one example was that, in hospitals, we could have little devices that decontaminate the air and stop the transmission of diseases and stuff like that. So it is not like you guys actually think we should stop.
We do think we should do something like Plan A as soon as possible. First of all, I think that just stopping everything now—Plan S—would be better than the default. I would rather just stop everything now than continue on our current trajectory. Secondly, our actual recommendation would be to do Plan A.
You do a temporary inference-only pause so that you can set up all the transparency and verification infrastructure. Then you can continue in this more distributed, transparent, cautious way, as previously described. Having continued in that way, you basically go up to the level that you feel like you can control reliably, until you feel confident that you have solved alignment well enough that you can give up on control. That is what we depict happening over the course of the 2030s in our scenario.
Why do you guys think that AI discourse is so bad? Is it unique to AI, or is discourse just bad in general?
Discourse is bad in general, right?
Well, why?
I do not know. I think there is a lot to say. One thing is that there are massive incentives and motivated reasoning for people at AI companies to think that what they are doing is justified and good, and that they should not take costly actions that would make the situation better, because those costly actions would be bad for whatever reason.
So, there’s a bunch of rationalization. I think most of the effect is probably just that discourse is generically hard and bad. Twitter or whatever is very unnuanced and argumentative.
I think LessWrong in particular, which is where I try to do most of my discourse, actually has pretty good discourse quality on average. Generally, when I write a post and go into the comments, I think the comments are quite thoughtful and technically informed, especially relative to other places like Twitter.
And then maybe a final thing is just that, in D.C. in particular—which is an area I care a lot about—I really want D.C., the government of the United States and other governments, to react well to AI technology. I think there are 2 problems going on there. One is that there isn’t very much AI expertise, right? The government isn’t hiring really high-quality technical AI experts who really know what they’re doing.
The second is this generic issue: I think the conversation isn’t happening because the incentives for everyone in D.C. aren’t toward truthfully and accurately understanding the situation. The incentives are for every individual to say stuff that sounds good and looks good, and that’s within the D.C. Overton window, so they can make a lot of friends. Unfortunately, I think this comes apart from the actual reality.
I think the actual reality of the situation is that a bunch of controversial and niche views about AI turned out to be true, right? The whole AGI hypothesis is correct. I think D.C. basically just hasn’t come to grips with that. So almost all of the discourse happening there is fundamentally anchored on completely wrong assumptions about how the technology works, in particular the assumption that it’s mostly fake news and it’s all a bubble.
Or that it’ll be like the next internet.
Or that it’ll be like the next internet, which is maybe better than it was a few years ago, when it was even more bearish, but it’s still not nearly bullish enough on the technology, in my opinion.
14. Is AI normal technology? The remaining disagreement
I think it was the AI Snake Oil guys who had an article saying that AI is normal technology. Obviously, it’s been said that AI is not a normal technology. It’s really quite different, right? So he’s very, very skeptical about AI, but you had some great discussions with him, and none of you changed your mind as a result of that. I think quite an interesting thing is the skeptics’ perspective as well. A lot of skeptics think that folks in Silicon Valley just aren’t being sincere. What they’re saying isn’t sincere.
Sure. With respect to sincerity, I would say some people in Silicon Valley aren’t being sincere, but others are.
Such as us.
Such as us, but also some of the people at the AI companies are being sincere. Not all of them. I don’t think you should trust what the leadership of the company says in general.
15. AI as Normal Technology
Anyhow, we co-authored a blog post or an article with the “AI as Normal Technology” people. It was a post about what we agreed on. You can go read it; there are 10 points in there that we both agree on. One highlight from our perspective is that we had a bit of a truce where we said, “Yes, AI right now may be a normal technology, but in the future it will not be normal. In particular, in the future it’ll be more like humans in the cloud, and all this crazy stuff is going to start happening,” as we describe in AI 2027. They agreed that if you get humans in the cloud, that would not be a normal technology. They just think you’re not going to get that, at least not for many years.
AI—is it possible that in the next few years we will have AIs that are like humans in the cloud and that can do all sorts of knowledge work in a way that substitutes for humans, including AI research, for example? And then that, sometime after that, we will have robots, perhaps controlled by those AIs, that can do physical work in a way that broadly substitutes for humans? Our claim is that, yes, that sort of thing will be achievable in the next few years. Their claim is, “No, not in the next few years. That’s much farther away.” I think that is the main source of our disagreement.
We would agree with them that if that level of AI and robotics is still very, very far away, then maybe AI is more of a normal technology. It’ll be like the next internet or something. But we just think that it’s actually on a path to get to that level of capability soon.
Can you be more specific about what the core cruxes are and what would make either of you change your mind?
I don’t know if I have a useful answer to that. I think there are a lot of different things we argue back and forth about. One thing that would change my mind is that people keep talking about the limitations of the current paradigm, but then the limitations of the current paradigm keep getting overcome within the current paradigm.
One thing that would change my mind is if someone was actually right about one of these limitations. If someone right now is going around saying that data efficiency or non-verifiable tasks is the current limitation, and then several years go by and it becomes clear that there’s basically no progress on overcoming that limitation—that the AIs of 2029 are no better at fuzzy, non-verifiable tasks than the AIs of 2025—then I’d be like, “Okay, this feels like a real barrier. This feels like something where we’re starting to feel the elephant. We’re running up against some sort of real, actual barrier that was correctly predicted by theory by some people to be there, and now it’s actually there.”
By contrast, from my perspective, there are just loads of experts going around talking about all these barriers, and then we keep plowing through them as if they’re not there. So actually running up against some sort of barrier like that would lengthen my timelines quite a lot.
Another thing that would lengthen my timelines quite a lot is more political change. If there was a war with China and most of the chips got destroyed by missiles, that would lengthen my timelines.
Exactly, like a vintage, though.
Okay, that’s true. On the bright side, if there was more of an international deal to pace the frontier, that would lengthen my timelines, et cetera.
I feel like the fundamental disagreement is really just this view of whether AI is more like electricity or airplanes, or whether AI is more like humans in the cloud. All of the intuitions are downstream of this core reference class of what we’re thinking the future will be like in the near term.
I think that if we froze AI progress right now and didn’t train any new models, it would become more of a normal technology. There are just so many things that Claude 5 can’t do, and if we couldn’t get any new models beyond Claude 5, then it would be more like the internet. We would restructure a lot of our professions and workflows to incorporate copies of Claude 5 doing parts of them, but then the humans would shift to doing more of the things that Claude 5 can’t do.
There would be a big change in a lot of things, and it would be like the next internet in the sense that it would change everything in some sense, but it wouldn’t fundamentally change anything.
16. Validity of the single processor approach to achieving large scale computing capabilities
Yeah. Right now you get these Amdahl’s law-type effects, where the AI can do some fraction of the workflow, but then it gets bottlenecked on the parts of the workflow that the humans have to do.
I still feel like there’s this thing where, when I think of the future, I’m imagining AI doing 100% of the workflow of a bunch of important economic workflows. You don’t get this bottlenecking effect. That’s a pretty qualitative change from the situation today that causes pretty fundamentally different predictions of what the world looks like. Whether AI actually gets to literally 100% of a bunch of very, very important tasks is the core question underlying the difference in worldviews.
I think that if we do see recursive structural adaptation, which is coherent—and I think it’s likely to be quite divergent—then that’s it for me. That’s not a normal technology.
Can you flesh that out a bit more? Would this be something like the Hugging Face swarm? Tell us more about what you would see that would be the thing for you.
Yeah. For me, adaptivity is the word most synonymous with intelligence.
I think these new RL-trained models we have aren’t the same as what we had many years ago. We said, “Scale is all you need,” and the clue is in the name. Scaling means you take a scalar property of a system and scale it up.
These RL systems are still self-attention transformers, but they’re different. They’re actually structurally different. There are new types of training, new architectures used differently, and so on.
A bunch of humans did some experiments and adapted the structure to create a system that had different scaling properties. I can imagine a future where we actually have some kind of recursive loop in which the system is adapting itself, deciding what things are interesting, and evolving by itself. When that happens efficiently, I think that’s a different type of technology.
Yeah, that sounds kind of similar to what we would say about recursive self-improvement and automating the AI research process itself. Yeah.
Yeah. I guess we agree, then.
Amazing, guys. It’s been an honor and a pleasure having you both on MLST. Thank you so much for joining us today.
Thank you for having us.
Yep. Yeah, appreciate it.