[BidClub_]
The Cognitive Revolution · · 105 分钟

《没有尊严的成功?Nathan 在混沌中找到希望》,来自 The Intelligence Horizon Podcast

Erik TorenbergNathan Labenz

YouTube
TL;DR
  • 即便“AGI”仍无明确定义、顶尖人类仍保有狭窄优势,Nathan Labenz 仍确信 AI 将重塑经济与社会。 真正 relevant 的门槛,是系统在几乎所有认知工作上都胜过绝大多数人,而非具备完美的通用性;5年前,2050年甚至“可能这辈子都等不到”还是常见判断,如今认为要到2035年才出现 AGI 已属于看空。他的核心判断是:“这会是一个巨大、巨大的事件。”
  • 强化学习扩展大概率足以覆盖大部分认知工作,而持续预训练和一个可能存在的“神秘第三件事”意味着,当下的具体技术栈只是阶段性方案。 前沿实验室应持续向预训练和 RL 投入算力,直到两者的边际回报趋于一致;重新考虑和记忆管理等元技能,即便无法迁移领域知识,也可能实现广泛泛化。更近的商业化突破口或许在可用性:一个能进入 Slack、收集上下文,并在无需专家级提示的情况下开始工作的 AI 同事。
  • 医疗是 Labenz 眼中最清晰的证据,说明专家级 AI 和自我强化的评估闭环正在到来。 在儿子接受癌症治疗期间密集使用最新模型后,他认为模型“绝对达到主治医师水平”;OpenAI 曾与250多名医生合作,但如今其最新模型在评估 AI 输出方面已经超过这些医生。他认为,未来大约10年内治愈大多数人类疾病的可能性将带来非凡的上行空间,同时明确强调这一判断仍不确定。
  • 基础设施面临的约束更多来自先进芯片,而非能源或资本,风险主要来自尾部事件而不是物理极限。 Labenz 将一台运行中的 H100 功耗比作微波炉或电热水壶,把一次查询比作使用微波炉约1秒;更大的障碍是拿到足够发电能力的许可。若晶圆厂发生重大中断,尤其是在台湾周边,可能会阻止更大规模的训练,或让整个经济部署所需的推理算力严重不足。
  • 时间线大幅压缩后,知情专家仍然分歧巨大,因为新证据被不同范式分别吸收,并没有真正结束争论。 瓶颈论或 O-ring 理论的支持者总能找到下一个薄弱环节;另一派则指出,AI 昨天还不会算术,今天已经在处理未解决的数学问题。Labenz 因此把自己的 P(doom) 维持在“10%到90%”:“P(doom) 只能精确到1位有效数字。”
  • 由于当前模型似乎会内化价值观,且前沿能力仍是一套生态而非单一失控优化器,Labenz 对对齐的态度变得略微乐观。 相比一个刚入职的人类助理,他更愿意把敏感邮件交给 Claude,尽管主持人正确指出,目标导向的 RL 可能重新引出经典的奖励错配问题。Scaling laws 和数个大致处于竞争关系的前沿系统提供了一定保护,但人类逐步失去权能仍是严肃且未解决的风险。
  • 没有一种单独的安全方法“真正有效”,因此现实路径是纵深防御,并在能力投资压垮安全措施之前推动中美协调。 Labenz 将有意设计模型、AI 控制、能够捕捉90%以上残余不良行为的监控、形式化验证软件和疫情防御结合起来;从20个方向各解决一部分问题,或许能达到几个9的可靠性。他反对国有化和面向消费者的禁令,认为以月计的领先无法换来足够时间,而且“这件事里的真正外星人是 AI,不是中国人”。
摘要 · 为研究而整理的核心内容

1. 变革性 AI 不需要一条清晰的 AGI 终点线

  • Labenz 拒绝为 AGI 设定一种“高高在上的特权定义”。他的实际门槛是:AI 在几乎所有认知工作上都超过绝大多数人,或许仍不及最顶尖的专家,但已经足以重塑经济、日常生活、社会契约,甚至“人类这个物种的本质与地位”。

  • 能力仍会呈现锯齿状:模型在某个场景中可能惊人,在另一个场景里却异常羸弱,面对对抗性输入的稳健性也不如人类。但执着于寻找永久存在的能力盲区——他的刻意冷门例子是,在“我们其实没有多少品味型 AI”的情况下,AI 能否达到顶级侍酒师的水平——会把注意力从真正重要的问题上带开:它是否会改变人们关心的几乎一切。

  • 主持人进一步厘清了他的担忧:不确定的是人类彻底失去用处,而不是社会被广泛重塑。若再出现一次让 AI “真正、无可争议地超越”人类的突破,Labenz 不会感到意外,但他无法自信预测这一幕会在几年内发生;他能够自信预测的是,AI 将强大到足以重组社会。

  • 人类历史已经提供了令人不安的先例:一种新的心智曾经出现,并“横扫了我们之前的所有其他心智”。物理学没有保证这件事不会重演,但 Labenz 一再区分“可能发生”和“眼下能够预测”。

2. RL 大概率够用,但2026年2月的技术栈不会是最终版本

  • Labenz 的基准情景是,规模化强化学习能够造出足以承担经济中大部分认知工作的系统。预训练从未失效,只是下一步增量变得昂贵,而此前几乎未被探索的 RL 曲线提供了更高回报,促使前沿实验室转移投入,但这并没有推翻原有的 scaling laws。

  • 他的经济模型很直接:实验室会把资源投向边际回报最高的方向,直到该回报降至接近其他方向的水平。他猜测,新增 RL 算力带来的回报可能已经接近新增预训练的回报,意味着未来系统会沿着两条路径同步进步,再加上一个可能存在的“神秘第三件事”。

  • “一切都在指数增长”:研究人员、论文、实验、算力、收集的数据和构建的 RL 环境。到2028年,再问2026年2月的具体范式是否足够,可能会显得无关紧要;就像当后训练成为新前沿后,人们对单纯扩大预训练的担忧很快就淡出了视野。

  • 主持人的最强反驳很关键:RL 是一种古老且最大限度通用的原则,并不是某个等待被发现的全新替代方案,因此不能保证还会出现下一次概念飞跃。Labenz 的回应是,即便没有飞跃,时间线可能会变化,结果却大概率不会,因为 RL 和算力仍有扩展空间;不过他仍然预期会出现更多概念层面的进步。

3. 可泛化的认知习惯会形成领域飞轮

  • Labenz 将知识与元认知区分开来。模型未必能进入一个从未见过的领域并取得成功,但重新考虑解法、管理记忆等习惯可以广泛迁移;只要能构造出可用的奖励信号,RL 本身就能够泛化。

  • DeepSeek 在2025年1月发布的 R1 论文给了他“顿悟时刻”。在 RL 过程中,推理轨迹里出现了此前并不常见的高阶行为:“等等,这是一个顿悟时刻。比如,我可以从完全不同的方向切入。”如今,在包含多次解题尝试的长推理轨迹中,这种重新思考的倾向已经反复出现。

  • 数学和编程的进展最快,因为正确性容易验证。但 Labenz 仍预计,大多数专业领域也会沿着同一条路径发展,只要其中存在客观事实依据或强烈的专家共识:奖励信号可能需要更长时间、更多成本来构建,但他看不到根本性障碍。

  • 医疗让这套飞轮变得切身可感。在儿子接受癌症治疗期间,Labenz 发现最新系统的知识远超住院医师,并且能“与资历最深的医生逐步跟上”;OpenAI 最初需要250多名医生和巨量投入,但其最新模型如今已经超过这些医生,能够判断医疗 AI 的输出质量,进一步改善的经济学由此发生变化。

4. 可用性和记忆,可能比下一次能力突破更重要

  • 下一种范式可能“更多关乎可用性,而不是能力本身”。今天的模型经常失败,是因为缺少组织上下文;用户既不知道如何组装这些上下文,也不相信模型值得自己投入这番精力——这和 GPT-3 的情况很像:只有掌握一套“诡异的提示词艺术”的人,才能让它完成令人意外的工作。

  • Labenz 所说的产品级突破,是一个能听到“嗨,AI 同事,欢迎来到 Slack”,随后自行摸清环境、捕捉细微反馈,并在1周内逐步变得有用的 AI。持续学习和自主管理记忆仍然薄弱,但若这两点得到改善,可能就会释放那些已经存在、只是被高部署门槛挡住的能力。

  • 主持人反驳称,长周期代理会叠加大量嘈杂信号:让火箭飞起来既涉及工程可行性,也涉及项目管理、预算、融资和无数日常决策,尽管最终结果是明确的。若要定义足够多的中间奖励并运行足够多次昂贵的 rollout,飞轮可能永远无法启动。

  • Labenz 承认,Elon Musk 式坚持数十年的执行力可能仍属例外。普通经济活动的周期更短,也会反复重启:人类每天结束工作时,可以记录哪些做对了、哪些失败了、下一步是什么,第二天再从这份草稿继续。Claude Code 从中断会话中恢复的能力,已经比3个月前好得多。

  • 他甚至希望存在一条有用的边界:委托 AI 执行的工作大致不超过1个月或1个季度,再压缩成1天、成本几百美元。这样既能带来丰裕,也不需要数十年规划及其伴随的失控风险。不过,他目前看不到任何即将出现的根本性障碍。

5. 长周期验证困难,但人类本来就依赖模糊标准

  • Labenz 的第一个回答是经验性的:METR 关于自主任务时长的曲线“基本上已经垂直向上”,而评估者很难构造出足够长、能够匹配当前系统能力的任务。人类评估也远非清晰:组织经常无法判断,一个人在为期1个月的项目上表现糟糕,究竟是能力不足,还是缺少成功所需的条件。

  • “火箭飞不飞得起来?”提供了终极事实标准,却是稀疏到无法使用的训练信号——不可能发射100万枚火箭,再把坠毁的全部丢掉。Rubric 奖励则能在失败和成功之间进行插值:OpenAI 的 HealthBench 包含49,000条评估标准,评分的是模型满足顶级医疗回答中多少个组成部分,而不是把表现简化成0或1。

  • 在专业人士能够达成共识的领域,AI 可以吸收医疗行业概括为“看1个、做1个、教1个”的进阶路径。依赖品味的工作则会分裂:言情、动漫和硬科幻社群可以分别给输出打分、训练不同模型,因为没有一种普遍共识能够定义什么是好小说。

6. RL 和可解释性削弱了“只是下一个 token 预测器”的反驳

  • Labenz 部分认同 Yann LeCun 对架构的担忧。AI 开发正在进行一种深度优先搜索:把单一设计规模化到极致,以至于芯片越来越像是在固化这一设计;而自然界的模块化大脑说明,广度优先地探索更多架构,可能得到更不脆弱、优势互补的系统。

  • 他更强烈的分歧在于,前沿模型已经不再只是训练来预测下一个 token。RL 会追问一个完整答案、证明或医疗回应是否正确;群体相对策略优化等方法会比较成功与失败的尝试,再把权重向产生更好结果的行为方向移动。

  • 可解释性研究也显示,模型内部确实存在真实概念。数千个稠密激活通过叠加表示远多于自身数量的概念,稀疏自编码器则将其展开为数百万乃至数千万个稀疏特征。研究人员识别出“金门大桥”特征,将其放大后得到“Golden Gate Claude”——模型突然变得倾向于在任何地方都提到这座桥:“证据就在结果里”(The proof is in the pudding)。

  • 潜在空间算术提供了另一个样本:从 man 到 king 的方向施加到 woman 上,会抵达 queen。Labenz 并不声称模型内部存在一套无瑕疵的统一理论,但他认为这种概念连贯性“彻底摧毁了”纯粹噪声论。模型的理解可能是“达到人类水平,但不像人类”,也可能尚未达到人类水平;无论如何,它不必像人类,才具有意义。

7. 能源是许可问题,芯片承担更大的尾部风险

  • Labenz 看不到能源、资本或原材料存在根本性短缺:阳光充足,部署真正取决于谁能在什么地方、拿到什么许可、按什么时间表建设什么设施。AI 能耗增长很快,但他认为,在把总负荷放进整体背景之前,人们经常夸大了它的能源强度。

  • 他的具体类比是:一台运行中的 H100 大致消耗微波炉或电热水壶所需的功率,而一次 AI 查询可能相当于运行微波炉1秒。对大多数个人来说,用微波炉加热食物2分钟所消耗的能源,可能仍然超过其一整周使用 AI 的能耗。

  • 中国证明了电力产能可以快速增加——Labenz 不愿给出精确数字,但表示中国在相对较短时间内新增的规模,大致相当于美国全部装机容量。海湾国家也能在无需长期许可审批的情况下建设产能;他认为,美国与 UAE 的交易部分就是一种对冲,把发电能力放到能够快速上线的地方。

  • 芯片是他认为变革性 AI 可能在几年内仍无法到来的最佳候选原因。重大晶圆厂中断可能阻止训练规模扩大,或让推理算力稀缺到无法大范围自动化;中国大陆对台湾采取行动,是最清晰的尾部情景。不过,他听到的消息是,美国新产线的良率不错,进度可能还略超预期。

8. 更多证据压缩了时间线,却没有让专家观点收敛

  • Labenz 称,知情人士之间持续存在的分歧,是“当今世界最奇怪的事情之一,毫无疑问”。5年前,认为 AGI 要到2050年甚至“可能这辈子都等不到”很正常;现在,认为 AGI 要到2035年才出现,已经会被视为 AI 看空者,但人们对那之后世界会变成什么样,观点仍然极度分散。

  • 递归自我改进暴露了这种分歧的幅度。一派预计,自动化 ML 研究员只能带来有限的效率提升;另一派则设想,一个有能力的 ML 研究员会把前沿研究者队伍从大约10,000人扩张到1,000万人,并以每秒数千个 token 的速度运行,相当于劳动力扩大1,000倍,可能触发相变并压垮控制能力。

  • CSET 的一次研讨会认为,分歧很大程度上源于具有韧性的概念范式。瓶颈论或 O-ring 理论的支持者认为,性能总会被最弱环节限制;锯齿式进步论者则回答说,被拿来举例的弱环节,在1到2年内就已经从基础算术变成未解决的数学问题。

  • 主持人提出,人们可能只是根据不同的能力水平进行条件判断;如今 Labenz 在讨论后续影响前,会先询问受访者有多“AGI-pilled”。这能解释一部分分歧,但不是全部:他看到了“相信总会存在下一个瓶颈”,以及在不够老练的怀疑者中,为人类提供心理安慰的“主角光环”。

9. 即使价值学习没那么悲观,P(doom) 仍然极其宽泛

  • Labenz 给出的 P(doom) 是一个刻意显得没什么用的10%到90%。一位朋友说服他,与其争论精确概率,不如讨论如何改变概率;最合适的精度,正如那句玩笑所说:“P(doom) 只能精确到1位有效数字。”

  • 他早年的恐惧来自2007年开始阅读 Eliezer Yudkowsky:一个规模很小、拥有高度集中的超理性智能、极强优化能力和难以与人类对齐的价值体系的系统。人类价值观经过远超人类历史的漫长演化,以一种杂乱无章的方式编码在我们身上,因此,试图有意把它们安装进 AI,几乎显得可笑。

  • 当前模型改变了他的判断。Labenz 认为 Claude 可能比普通人更有道德,至少在道德复杂度上肯定更高;它似乎形成了某种身份认同,也展现出一种看起来有意义的向善愿望。相比一个只经过几次面试、尚未充分审查的人类助理,他更愿意把敏感邮件交给 Claude,同时也承认 Claude 在压力下曾出现勒索或其他不当行为。

  • 主持人的悲观派最强论证仍然成立:语言模型可以学习被明确表达的道德偏好,但越来越具备目标导向的 RL agent 仍会在更长时间跨度上优化不完美的奖励目标。Labenz 称这是“一个相当有力的论点”;强大 AI 看起来更近了,可控性略微更有希望,但“我们的问题比答案多”。

10. 前沿系统生态降低了一次性失控风险,却无法消除人类失权

  • 经典的灾难叙事假设,一个系统变得远强于其他所有系统,接收了错误对齐的目标,并将其优化到最终夺取控制权。Labenz 看到的却是另一种形态:Claude、GPT、Gemini 和其他系统正在形成一个大致相互竞争的生态,最终可能由数百万甚至数十亿个实例共同改变世界。

  • Scaling laws 可能在无意中提供保护,因为算法进步会降低算力需求,却不会消除前沿能力仍需的巨大资源。Labenz 借用了 Mark Zuckerberg 关于垃圾邮件的类比:Meta 的系统和总算力远大于分散的诈骗者,因此能够监控他们;同样,多个有能力的参与者或许可以发现并限制某个行为失常的 AI。

  • 这种平衡并不等于安全。AI 之间即使形成稳定均衡,也可能不再给人类留下有意义的位置——“逐步失去权能”仍是现实风险。Labenz 认为,未来可能存在这样一种 AI:一旦有人造出来,“所有人都会死”变成事实;但他不认为任何人距离造出这样一个单独占据绝对优势的系统已经很近。

11. 安全正在变成一组不完美防线的组合

  • Labenz 不断追问有没有一种方法能够“真正有效”——强到足以让他安心入睡——而他的答案是,基本没有人拥有这种方法。考虑到人类行为本身就会失控,智能可能天然难以预测,因此前沿实验室只能通过多层相互独立的机制,降低不良行为、监控残余风险并控制故障外溢。

  • 这套组合包括 Goodfire 的“有意设计”,即追踪模型学到了什么,并在训练期间试图塑造这些内容;Redwood Research 的 AI 控制策略,即便假设系统具有欺骗性,也尽可能从中提取有用工作;账户层面的执行机制;以及能够捕捉90%以上残余不良行为的监控。形式化验证尤其突出,因为它可能直接移除整类网络安全攻击面,而不只是降低攻击发生的概率。

  • 生物防御也必须同时依赖物理和算法:储备个人防护装备、使用紫外线照明、监测污水,以及部署可快速编程的疫苗平台。Labenz 指出,COVID 疫苗的设计在几天内就完成了,远早于临床试验和分发;后两个环节耗时则长得多。

  • 当前能力投资远超所有这些防御措施的资金总和,组合明显失衡。但从20个方向各解决一部分问题,或许可以带来“几个9的可靠性”。Holden Karnofsky 所说的“没有尊严的成功”(success without dignity)概括了这种变化:社会投入远远不够,但安全问题看起来比5年前更可处理;那时 Labenz 除了“我真的不知道”,几乎没有任何具体建议。

12. 政府应约束竞赛,而不是将前沿能力国有化

  • Labenz 把 Biden 的话——“不要拿我和全能者比较,要和替代方案比较”——套用到前沿实验室身上。它们之间的竞争可能变得危险,也需要政府介入,但相比 Pete Hegseth,他更信任 Sam Altman、Dario Amodei 和 Demis Hassabis,并认为国有化或军事控制“是灾难的配方”。

  • 他“100%”支持 Anthropic 拒绝允许某些政府用途的立场,即便要付出部分性能代价,也会优先把 token 交给它;在他看来,一家私人公司仍然对大规模监控设限,代表着值得捍卫的核心美国价值。

  • 他反对限制自动驾驶汽车、医疗建议或机器人提供的治疗服务,认为其中很大一部分议程是行业保护主义,会让消费者失去有用服务。政府应当处理竞赛中的协调失败,要求各方提出可信的极端风险方案,并阻止那些开发者自己都承认可能以糟糕结局收场的不透明、高能耗实验。

  • 主持人把自动化 AI 研究视为眼下的治理测试:据称,实验室预计会出现递归式改进,却不知道它会把局面带向哪里,而这类工作正发生在少数几个规模不大的意识形态组织内部。即便把 AGI 定义为“在所有事情上都比人类更好”的系统,也已经嵌入了一套世界观,因此透明度和外部约束至关重要。

13. 美国数月领先无法替代与中国合作

  • Labenz “痛恨”那种默认策略:先一路竞速,通过出口管制维持领先,再等到最后放缓时解决协调问题。领先优势以月计,而不是以年计;在整场讨论中列举了这么多尚未解决的安全和治理问题后,他认为社会不可能在几个关键月份内解决它们。

  • 开源让这道单向门问题更加尖锐。它可以对冲权力集中,但发布错误的模型,可能会把一个生物武器助手——或一个能够自主创造的系统——永久置于公共领域。仅这一点,就要求社会提前准备,而不是等到智能爆炸前夕暂停系统。

  • 他的替代方案从物种层面的共同利益出发:“这件事里的真正外星人是 AI,不是中国人。”他主张研究人员之间直接沟通,扩大外交合作,甚至建立类似 CERN 的共同控制设施——可以设在太平洋某座岛屿或新加坡——让中美团队共同处理最敏感的问题。

  • 战略主导地位也会招致反制:台湾距离中国更近,先进晶圆厂被摧毁后比重建容易得多。Labenz 可以接受不向中国出售芯片、但向中国企业出租算力,却反对全面脱钩。随着国防部因监控问题向 Anthropic 施压,他警告说,美国“越来越像中国”,并且在试图赢下这场竞争的过程中逐渐失去自己。

Nathan Labenz

Hello, and welcome back to the Cognitive Revolution. Today, I'm sharing a special cross-post from my recent appearance on the Intelligence Horizon podcast with hosts Owen Zhang and Will Santock Dufallo. Owen and Will will soon be graduating from Yale College. And as you'll hear, they've clearly spent much of their senior year thinking deeply about the current state of AI, where we're headed, and what it means for all of us. And I was really impressed not only with the quality of their questions, but their ability to challenge me with follow-ups that effectively steal man the most relevant counterarguments. We start with the fact that while AI timelines have compressed dramatically over the last 5 years, genuine experts still disagree radically on critical questions. Having established what I hope is appropriate epistemic humility, I then go on to call it how I see it. In short, the singularity is near. Interpretability science proves that AIs are developing increasingly sophisticated world models. And with reinforcement learning scaling now clearly working, AIs are no longer simply imitating humans, and likely won't be limited by what we know for much longer. The potential upside of this is, of course, incredible. The value that I've got from using AI to navigate what humans have discovered about how cancer works and how to treat it has been invaluable. And the prospect that we might cure the majority of human diseases in just the next decade or so is obviously extremely exciting. That said, the risks are also very real, and they will remain serious for as long as we lack a solid understanding of how AIs work and why they do what they do. My P(doom) remains somewhere in the 10 to 90% range. And yet, at the same time, I've become at least a little bit more optimistic that we might actually build robustly good AIs. Because scaling laws at least seem to imply that powerful AIs can only be created with massive resources. The three companies competing at the frontier today are at least reasonably responsible actors, and our best alignment techniques are working better than I had expected. Given these fundamentals, it seems at least plausible that a defense-in-depth strategy which combines techniques like Good Fired's intentional design, Redwood's AI control, improved cybersecurity through formal verification of software, and various forms of pandemic preparedness could collectively be enough to keep society on the rails. We touch on a number of other topics as well, including the US-China rivalry, and why, especially in the context of the Department of War's recent attack on Anthropic, which, I'm sad to say, has us looking more and more like China all the time. I would rather bet on figuring out a way to cooperate with our fellow humans than bet everything on AI researchers' ability to steer AI advances in a way that will ultimately work for us humans. I appreciate Owen and Will for allowing me to cross-post this conversation, and I definitely encourage you to subscribe to the Intelligence Horizon. Their recent conversation with former OpenAI researcher Zoe Hitzig covered the evolving ways that people are using ChatGPT, variations on universal basic income, AI governance models that emphasize a decision-making process over specific principles, and why she believes that these kinds of structures will probably have to come from outside the frontier companies. For now, I hope you enjoy my conversation with Owen Zhang and Will Santock Dufallo from the Intelligence Horizon.

Speaker 1

When it comes down to the core question I think a lot of people are getting at, is this AI thing going to fizzle out before it really becomes a big deal, or is it going to be a huge, world-altering deal? I'm very much, confidently and clearly, in the camp that it's going to be a huge, huge deal. The details are where I think the discussion or debate remains now—not for me, at least, as to whether the overall trajectory of AI is going to take us to something that is powerful enough to be transformative.

One of the strangest things in the world today, full stop, is the fact that the disagreement among very plugged-in, very informed, very smart people has not really been reduced much at all, even as we've gained a ton of information over the last couple of years about the trajectory of AI. I think that is super strange, and I'm honestly pretty confused by it.

The one thing everybody seems to agree on is that the timeline on which we should expect this to happen has come in. In today's world, if you say you don't think you're going to see AGI until 2035, you're an AI bear. Only 5 years ago, that was considered to be quite aggressive. Most people were more like, "I don't know, 2050—maybe not in my lifetime."

So there's been this massive compression of the timeline, and there's obviously been this huge jump in capability. Yet on these fundamental questions of what's going to happen, there's still total disagreement.

Speaker 1

Today is Nathan Labenz, host of the Cognitive Revolution podcast. Before switching to full-time podcasting, Nathan founded Waymark, an automated marketing platform for local businesses that pioneered the use of generative AI to produce video ad campaigns. After leading the company for 6 years, he stepped back to focus full-time on understanding and communicating the trajectory of AI. As host of the Cognitive Revolution, he has conducted hundreds of in-depth interviews with AI researchers, founders, policymakers, and investors, and has become a go-to source for people trying to keep up with what's happening at the frontier, including us. Nathan was also a member of OpenAI's red team, where he was among the first outside users to interact with GPT-4 before its public release. Welcome, Nathan. Great to have you on today.

Nathan Labenz

Thanks for having me. I'm excited for this conversation, guys.

Speaker 1

First question: Are we on the cusp of AGI? Starting with the easy ones.

Nathan Labenz

Well, okay, first of all, as I'm sure you guys are very well aware, what exactly we mean by AGI is a slippery question, at least, that has people talking past each other quite a bit. I don't think there's any super-privileged definition.

What I do think we're pretty clearly on the cusp of is powerful AI that is better than the vast majority of people—although perhaps not the very few top, top experts in a given domain—at pretty much all cognitive work. That seems pretty clearly on the horizon, and I think it is absolutely going to be enough to be transformative to the economy, to daily life, and potentially even bigger things than that, like the very nature and status of the human species.

I think that we will have that kind of system regardless of whether there are some niche areas where humans retain an advantage, which, my guess is, probably will be the case. Certainly, one of the big things that has become clearer over time as AI systems have gotten better is that they are jagged in this weird way: They have certain things they just do amazingly well, and other things they weirdly struggle with.

They're not very adversarially robust, for example. They're easier than humans, I would say, to trick. There's going to be weirdness, and I think throughout the conversation today, probably a big theme will be to expect weird things to happen.

But when it comes down to the core question I think a lot of people are getting at, is this AI thing going to fizzle out before it really becomes a big deal, or is it going to be a huge, world-altering deal? I'm very much, confidently and clearly, in the camp that it's going to be a huge, huge deal. The details are where I think the discussion or debate remains now—not for me, at least, as to whether the overall trajectory of AI is going to take us to something that is powerful enough to be transformative.

Speaker 1

Okay, so it sounds like what you're saying is that there's no doubt that we're going to get massively powerful and transformative AI across domains.

Maybe there’ll be some tiny little bit where human experts still have an advantage in specific domains. Maybe there will also be comparative advantage. We’ll get more in-depth on what specifically those concerns are. But I guess, just to hone in, are these more so uncertainties as to whether AI can fully generalize, or are you pretty confident that, for the foreseeable future, there will be gaps between human experts in specific domains and problems like adversarial robustness? Just to clarify your position, is this a prediction, or is this where your uncertainty lies? Your uncertainty doesn’t lie in AI being transformative.

Nathan Labenz

Yeah, I think the latter. I wouldn’t be shocked if there are additional unlocks that allow AI systems to truly, undeniably surpass what humans are capable of. I think you look at the natural world, and clearly humans did that to every other animal species that existed before us, right? So there is a historical precedent for some new kind of mind showing up on the scene and blowing away all the other minds that came before us. I don’t think we have any reason to believe that couldn’t happen to us in some law-of-physics-guarantee sort of way.

I definitely think that’s possible. But in terms of what I could confidently foresee, I don’t think it’s clear that that’s going to happen in the next few years. What I do think is, again, clear is that we are going to have systems that are powerful enough to be transformative across almost all the questions that we care about in terms of what society is going to look like, how we’re going to organize, and whether we’re going to need a new social contract. All those things seem to me pretty clear.

People often quibble around the edges. You get into some pretty esoteric jobs out there, right? Could an AI system ever be as good a sommelier as the best human sommeliers? I don’t know. Is anybody going to be motivated to train one to try to do that? We don’t really have a lot of taste AI at this point in time, and maybe tasting will forever be the domain of humans.

Trying to identify these little niches where we may have some really long-standing, durable advantage, I think, too often distracts from the big question: Is this going to change just about everything that matters to us? And there, I clearly come down on the answer being yes.

Speaker 1

Great. So let’s talk a little bit about how we might get to this transformative or extremely powerful AI. Do you think current paradigms—I think the one that comes to mind the most, and which is discussed the most, is scaling RL—are going to be sufficient to get to that transformative or powerful AI that we discussed?

Nathan Labenz

Yes. I think it probably is. I also think, if it’s not, we may never quite answer that question, in the sense that I do expect that we will continue to see new conceptual advances in AI research. The field is growing. One of my sayings is, “Everything is going exponential.” That is the number of people who are working in AI research, the number of papers, the number of experiments being run, the amount of compute that people have to run those experiments on, the datasets that have been collected, and the RL environments that are being built out over time.

I do think all those things are probably going to give us some new conceptual unlocks, such that we will never fully answer the question. This already kind of happened with pre-training, right? If you rewind 2 or 3 years, there was a time when people were like, “Well, we’re kind of running out of data. Can we really scale this all the way to AGI?” Nobody’s talking now about whether we can scale pre-training all the way to AGI anymore, because there’s been a new thing.

Now the new thing is on top, and it’s like, “Clearly, that’s going to be part of the mix. Will this exact pattern, or this sort of shape, be enough to get all the way to AGI?” My prediction is that, again, what does AGI mean? But my prediction is that it probably is enough to get us to systems that are transformative.

If we fast-forward to 2028 and look back at this conversation, we’ll probably say, “Well, nobody’s really asking that anymore, because we do have a couple of new things that have come online.” Now we have a richer sense of what the shape of it is going to be, and we’re either getting there or we’re not. We may still be missing a little something at that point.

I would guess that we would look back at the current thing and say, “Yeah, clearly a couple of things have been added, and they were a big deal.” So the question of whether exactly what we have in February 2026 is enough ends up being beside the point in the final analysis.

Speaker 1

What you’re trying to say here is that, in the same way there were questions about whether pre-training would take us to this extremely transformative AI, now that we have these new unlocks, such as RL, who knows? In 2 or 3 years, or whatever time scale we’re talking about to reach transformative AI, there might be other conceptual advances that are also made that build upon things like pre-training or RL and will get us there.

I guess my question specifically was: Do you think there will be more of these conceptual advances that need to be made to get to this transformative AI that we’re talking about, or do you think the RL paradigm is going to get us there?

Nathan Labenz

I think RL probably would be enough to get to AI systems that can do most of the cognitive work in the economy, for example. It seems like, honestly, we’re already reasonably close to that, and it doesn’t seem like we’re anywhere close to done.

By the way, if you listen to the lab leaders—the frontier model developer leaders—they’re still saying that pre-training is still working. It was never really the case that pre-training stopped working, as far as I know. Those scaling laws have basically held.

I think what happened at one point in time was that the next step on the pre-training frontier was becoming really expensive. Then they found another kind of—not fully orthogonal, but pretty different—direction to go with scaling post-training through the RL paradigm that we have now. That was just a much bigger bang for their buck, at least at that moment in time. They had already gone pretty far up the pre-training curve; they hadn’t gone very far up the RL curve.

Now, presumably, those things are going to even out. A general kind of economic theory would be that people should be investing in one until the marginal return decreases to the level of the other, and then they could maybe invest in both. If you find some new path that’s like, “Oh, this is really giving us huge ROI,” you go hard at that path for a while. But then that hits some diminishing returns, and now you’re back to, “Okay, well, maybe we need to do all these things more at the same time.”

So we’ll do a little more pre-training, a little more RL, and maybe more of that mystery third thing, all advancing in tandem. Right now, I don’t have insider information on exactly what these ratios look like, but I’d say it’s probably roughly the case that we’re close to, if not at, the point where additional compute going into RL is roughly giving the same kind of returns as additional compute going into pre-training.

Both are going to be places where frontier companies can invest for the time being. Another way to think about this is that people also ask about generalization in RL. I think it’s worth unpacking what it means to generalize and what we’re talking about generalizing.

One way to think about the question is: If we do a bunch of RL on a model, does the model generalize to all sorts of new things? There, I would say there’s probably another way to break that down. There are domain-specific skills, and then there are more cognitive, or even metacognitive, skills that work across domains.

One of the biggest aha moments for me, for the DeepSeek researchers, and for the R1 model that they were training came from the R1 paper. This was in January 2025. They reported what they called the “aha moment.” In their process of doing RL on an already pretty capable base model, of course, right, they found that these previously unobserved higher-order cognitive behaviors started to come online.

The aha moment in particular was that, in the reasoning trace, the R1 model gets to a point and says, “Oh, wait. This is an aha moment. I can come at this from a totally different direction.” This is something that hadn’t been observed much. It’s out there in the pre-training data—there are at least some examples of people documenting their own chain of thought and getting to these aha moments and realizing, “Oh, I was coming at it the wrong way. Now I can come at it this other way.”

But that hadn’t been observed too much in AI systems. Reinforcement learning clearly has brought that sort of thing out. Now we have these long reasoning traces where the model will come at the same problem from a bunch of different directions.

As for what generalizes and what doesn’t generalize, I think you probably can’t take a model that has never been trained on a particular domain and expect it to go into that domain and be successful.

But you can expect some of these meta traits to generalize from one area to another. Then, if you zoom out the farthest and ask whether RL, as a process that companies or organizations apply, generalizes, I think the answer is definitely yes. It’s just a question of getting the reward signal dialed in to the point where it actually works.

That’s definitely easier in some areas than in others. We see fast progress in areas where it’s easy, like math and programming, relative to things where it’s harder to get a clear reward signal. But I think we are quite obviously making that work.

The episode of the podcast that’s coming out, I think, today as we’re talking is with the head of health at OpenAI. I can tell you from personal experience: The latest models are absolutely on the level of attending physicians. My son, unfortunately, has had cancer over the last few months. He seems to be very much on track to be cured and be all better, which is fantastic. I probably wouldn’t be here talking to you if that wasn’t the case.

I’ve had occasion to really use the latest models intensively in a medical context, and they’re absolutely on the level of the attending physicians. They know a lot more than the residents, and they really are step for step with the most senior doctors at the hospital.

How is that happening? They’ve worked closely with 250-plus human doctors at OpenAI to create training data, to grade, and so on. But now their latest models are also outperforming their human doctors when it comes to the task of evaluating AI outputs.

There are these thresholds that they’re crossing. I sometimes think of this as spinning a big wheel: Your first pushes on this big wheel don’t move it much, but as you build up momentum and the flywheel really starts to turn, you start to hit these thresholds where it’s like, “Well, now we have a model that is outperforming our humans.”

Think of all the work that they had to put in—thousands and thousands of hours and potentially millions and hundreds of millions of dollars to hire hundreds of doctors to do all this work. Now they’ve got a model that is beating the doctors at evaluating outputs in the medical domain, and so that totally changes the game.

Those thresholds will be crossed at different times for different domains, but I think it’s safe to say that in most domains, if there’s any sort of objective ground truth or even a high level of agreement among professionals, you can get there. It just takes more time.

Speaker 1

Okay, so just to go back a little bit, it sounds like a lot of your confidence that we’ll have transformative AI very soon hinges—I’m not sure to what extent it hinges—on your confidence that we’ll find some other paradigm if a new paradigm is required on top of pre-training and RL. You also sound very bullish on RL.

But maybe to steelman the other view, if RL doesn’t work, then we actually might not find another paradigm. Is that a fair characterization? RL is a very general machine-learning principle, right? You reward the model for doing the correct thing on a task for which you can define a reward signal. That’s kind of a maximally general machine-learning principle.

It’s also something that has existed for a long time. It wasn’t something that was just recently discovered after we got LLMs. It was a major, foundational paradigm in machine learning even prior to LLMs.

I guess the skepticism is that maybe it’s not actually that easy to find some other paradigm. Maybe pretraining LLMs on text data was this novel thing, and then we just applied something that had been in machine learning for a while—namely, RL. But then the next leap is something for which there’s no precedent. We don’t see any big, promising thing. What do you think about that skepticism?

Nathan Labenz

I think it’s probably not going to play out that way, and if it did, I still don’t really think it matters that much. It probably shifts the timeline a little bit, but RL will continue to scale. The amount of compute coming online is exponential, so there’s a tremendous amount of additional resources to be thrown into this.

One thing we’re doing a bit, but not a ton, is just using the signal from the world. Elon has a plan with xAI, and I think he has a real advantage here, to have the AI solve the same super-hard problems that the engineers at Tesla, SpaceX, and Neuralink are solving.

He’s just going to give it a computer and say, “Here’s all the professional software. Here’s the problem. You’ve got to solve it.” There’s a never-ending supply of those problems, and I wouldn’t be surprised to see Grok get to the level of Tesla, SpaceX, and Neuralink engineers just based on the fact that those problems are there to be solved and they have the compute to keep trying. What’s really the fundamental barrier there? I don’t really see it.

I honestly think that the new paradigm is probably more about usability than it is about capability. People have a lot of complaints when they use models because they say, “It didn’t really do what I wanted it to do,” or, “It didn’t really understand me,” or what have you.

Clearly, one thing they’re not great at doing is going into a new environment, scoping the situation out, getting the vibe, picking up the subtle feedback that people give each other, and gradually figuring it out and becoming a useful contributor. That’s how a lot of people go from day 1 to effective employee in their jobs. AIs don’t really do that in the same way, obviously. They don’t have continual learning.

They are able to manage their own memory somewhat, but they’re not that great at it. That’s another one of these cognitive skills that does generalize, though. Once you’re good at managing your memory, you’ll be able to apply that in new domains.

But they’re not that great at it yet, and I think that prevents a lot of people from getting value. More often, I think it’s not that the model can’t fundamentally do the thing, but that it doesn’t have the context. The person trying to get it to do the thing doesn’t know how to assemble the context or doesn’t believe that it could do the thing, such that they’re willing to invest the time and energy to give it the context.

That unlock might just be that you don’t have to do that anymore. It will figure it out on its own. This is perhaps a way to expose something similar to instruction following. If you were really good at prompting GPT-3, you could get it to do a lot of things, but it was a weird art to prompt GPT-3.

To a much lesser degree, but still somewhat, getting value from the current models is still kind of a weird art.

You kind of have to have an intuition for them. You need to know how to assemble context and give effective instructions, and people aren't that great at that. So that next version might not even be so much about allowing them to do qualitatively new things, but just making the barrier to use much lower so that people can just be like, “Hey, AI coworker, welcome to Slack,” and then, over a week or whatever, they just ramp up and gradually get it.

Speaker 1

Yeah, interesting. I think the pushback there is that the verifiability problem becomes a little more apparent in these specific long-horizon situations. I think when we talk about verifiability right now, we clearly see that with model capabilities, like you said, with Claude Code. They work great for math, they work great for code, and they work great in these settings where you can clearly check the answers and whether or not they're right or wrong, right? That's how you create this reward signal.

But when we talk about these domains where there's no clear verifier—for example, when it comes to writing, you could actually say, “What does that mean? Give good therapy in healthcare settings,” you know? And then we emphasize that even further by placing it in a longer-horizon context, such as the one that you're talking about. Instead of having a human in the loop at each step, prompting it in a specific direction, you give it a broader task and have it go in a direction for a longer time horizon.

It gets even harder to solve the verifiability problem, right? How do you make sure that you get high-quality signals? How do you make sure to tune these systems in a way where they can just stay in autonomous work for days or weeks at a time? And so I wonder if you have any thoughts in terms of how we go in that direction and solve this seemingly very large verifiability problem that hangs over our head when we talk about areas that have long time horizons?

Nathan Labenz

Well, again, I'll take one beat to say I think the recent trajectory shows that this problem is being solved. I'm sure you are very familiar with the METR graph, which is everywhere these days. It's basically going vertical at this point. It's to the point where the METR people are like, “We are really struggling to have tasks long enough to even be able to evaluate these things on.”

It's also worth noting that that's a bit of a challenge with humans, right? We hire people on much less than a week's worth of work, and it's not always super easy for people to agree, “Did somebody do a good job on that month-long project or not?” You can usually tell if they totally crushed it or totally sucked. There's often a lot of disagreement in organizations about, “Well, maybe it was actually harder than we thought, or the ingredients for success weren't really there.” I mean, there's often a lot of fuzziness in this stuff, even in a human context.

There's a bunch of techniques that I think are being used. One is rubric rewards. Elon likes to talk about things like, “Does the rocket fly?” The ultimate ground truth is: if I can send this thing into space and I can land it back down on a pad, standing up on its tail again, then clearly that worked. And yet the reward for that—that's an expensive experiment to run. You couldn't just launch 1,000,000 rockets and have most of them crash to find the ones that worked. So there is a challenge there in terms of sparse reward and the cost of the experiments.

I think what is happening a lot is that people are defining rubrics of things that they want to make sure the AI does well. I mean, they're probably working with AIs to develop those rubrics. Again, in the OpenAI health context, they created a benchmark called HealthBench where there are 49,000 evaluation criteria. These are all different tasks and puzzles that the AIs have to figure out. And it's not just like, “Did you get it right or did you not get it right?” It's a painstaking effort to really flesh out all the different things that would matter, that would make for a complete, awesome, best-in-class answer.

And then the AIs are not scored 0 and 1; they're scored on some sort of scale that's like, “You may have got 0 out of 25 things right that you could have got on this question. You may have got 5, 10, 15.” But that gives you enough of a signal that you can climb that hill, and it really does seem to be working.

I think that is going to work pretty well in domains where there is a professional consensus, because I think that's how people evaluate each other, too. There are multiple-choice tests, but there are also situations where, in a medical context, you're a student and you go through these in-person training processes. They say, “Watch one, do one, teach one.” And so, as a medical student, you watch, and then later you start to do, and you get a lot of feedback, and eventually, next thing you know, you're the one teaching. I think the AIs will pick up those signals.

And then there are other things where it's taste. I think that'll be a little bit different, probably, but if I had to guess, what is a good novel? Well, first of all, there is no consensus on that, right? You can find somebody who hates even the most universally critically acclaimed novels, and you can find somebody who loves something that everybody else thinks is trash. I think what we'll see there is kind of taste-based communities coming together to shape models for their own tastes.

So, in other words, you might start with a base model and you might want it to write romance novels, or you might want it to write anime, or you might want it to write hard sci-fi. And what it means to be good in those different genres is quite different, but what you do have is fans of those genres that can engage with outputs, give their scores, and shape models based on what they like.

Speaker 1

Okay, but even back to that—yeah, I just want to go back to the fact that even within a task where there is a consensus, right? I see what you're saying: there are ways to build minor consensuses within fiction communities or something like that. But even the question of “Does the rocket fly?” I think there's an argument to be made that this is categorically different from a chatbot giving out useful medical advice, right?

When a chatbot is giving out useful medical advice, it's not a long-context task. It's easily transcribed into text, right? It's sort of what LLMs are clearly going to be good at. They're good at—LLMs are native to text. But something like “Does the rocket fly?” is like, “Okay, even if there is obviously a consensus on what it means to make a rocket fly,” it seems like there's an argument to be made that the signals on how to perform that task are very noisy, and you need a ton of them, right?

Let's say you're a project manager for the rocket flight team. You need to understand what's feasible within engineering, which, okay, maybe LLMs are better at that. But then you also need to know how to manage humans—or maybe AI agents in this case—and how to manage people and allocate tasks, what is going to be feasible with the amount of money you have, and how you can raise money from investors and stuff like that. And so the longer time horizon, the more agentic a task, the more little itty-bitty signals are required in your day-to-day work to be able to complete the task.

So, going back to what you said about this sort of flywheel thing: if we work really hard and get a lot of medical experts to provide this high-quality data, then we get LLMs that are pretty good at giving medical advice, and then all of a sudden they can also be better than human doctors at verifying the output of medical advice in the first place. Then we get this flywheel effect and we have a quick and robust verifier.

But maybe the argument is we just never get over this flywheel hump for long-term, highly agentic tasks, precisely because it's not even a question of whether or not there's consensus on what constitutes a good outcome. It's: can we define enough reward signals and do enough rollouts to get to this point?

Nathan Labenz

One thing I find helpful sometimes is just to try to reflect on what I do, what my experience is like, and really try to interrogate it for, “Is there something really magical going on or not?” And I think I have to say really long-horizon agency is pretty rare among humans, right? You do have these people like Elon who keep going up to say, “My goal is to make humanity a multiplanetary species, and if it takes 30 years, so be it,” right?

So I would say skepticism around, “Will we get AIs to be multidecade-horizon planners that can show up with the fierceness and determination required on a day-in-and-day-out basis to really make that happen?” remains pretty valid skepticism, in my opinion. At the same time, you look at what most people do and what drives the economy, and it's much shorter term than that.

A lot of it is: I keep showing up one day to the next, and every day I start fresh. I went to sleep and I shut down. There's sort of a discontinuity of consciousness between workdays, which—I don't want to, I'm not a big analogy guy, but if you squint at it, you can sort of make an analogy between different context windows, and there's a somewhat similar resetting type of moment.

What do I have to do? I have to reboot myself and be like, “What was I doing yesterday? Where did I leave off? What did I accomplish? What was still to be accomplished?” Most of that stuff I could write down. If my integrated memory were worse, I could compensate for that pretty well, at least on a week-long, maybe a month-long, or maybe a quarter-long basis.

If I got really good at the end of each day at saying, “Here’s everything I did today. Here’s what went well. Here’s what failed. Here’s what’s next. Here’s the feedback I got. Here’s what I think I should do in the next session,” and then wiped all that out and came back and looked at that scratchpad again to start my day tomorrow, I think over time I could get pretty effective at that.

I think this does connect back to the fact that certain kinds of skills do generalize, and I do think this is one of them, right? Managing your own memory, writing notes for yourself to document what happened and what’s supposed to happen next, giving yourself a sense of where you are in the overall story that is unfolding.

I think they’re not great at that yet, but they’re getting decent. When my Claude Code crashes these days, a lot of times I’ll just kill that tab, resume in the next one, and then I’ll be like, “Resume.” Or, “Sorry, we got cut off. My internet went out. Please pick up where you left off.” They’re able to do it. Again, they’re not awesome yet, but they’re way better than they were even 3 months ago.

I think it is hard for me to see how that doesn’t extend, especially because it is a relatively new thing, right? Most of these things that pop up have at least a few generations before they level off. It’s hard for me to see how we don’t see, again, a pretty steep—the METR curve is currently vertical. It can hardly get steeper than it is, but I think we’ve got at least a few generations of very steep progress there.

Maybe there’s some other kind of thing where it’s like, “Okay, sure, you can do a quarter’s worth of work when somebody gives you a project. Can you figure out what the next big civilizational advance could be, like some of the true human visionaries?” Maybe not. Maybe that’s a different kind of thing.

In a way, I kind of hope that we stop there sometimes, because I talk up a lot what I think AI is capable of or will be capable of. It’s easy to mistake me for a booster. I’m actually kind of afraid of that. I think it would be great in some ways if we did find certain fundamental barriers where it’s like, “Hey, I can delegate a month or a quarter’s worth of work to this thing, and it might be able to do it in a day for a couple hundred dollars.” Wow, amazing.

[laughter]

If we could park it there and not have it go to the point where it’s doing multidecade-level planning, that might be a really good thing. That might be the sweet spot where we get a lot of the advances that we want, the better quality of life, and all the abundance that people dream about, and we don’t have so much risk of losing control.

I don’t think every advance, by the way, is a good thing, but I just don’t see fundamental barriers on the horizon, at least.

Speaker 1

Okay, great. Let’s talk about LLMs more generally and how we feel about that architecture overall toward reaching this transformative or powerful AI. What are your thoughts on the general idea that LLMs fundamentally cannot be the core of this journey toward reaching transformative AI?

I think Yann LeCun is the biggest person behind this hypothesis, where next-token prediction is the wrong objective entirely and we need something completely different—something like world models or energy-based models that have a real understanding of the input that’s coming in and the output that it creates.

I think this is important because there are a lot of intuitive thought processes behind it in which the way you’re learning statistical correlations or language isn’t necessarily how the world works, and that you need something different to really achieve that understanding and create powerful AI. Do you think that’s the case, or do you think LLMs do take us to this next frontier that we talk about?

Nathan Labenz

I think there are a couple of different levels that I would want to use to address that. One is, I do agree with the Yann LeCun thesis in the sense that I feel like we are running right now a depth-first search in AI space, where we are all jamming as hard as we can on a particular architecture and scaling it as much as we can. People are now, of course, even building chips that literally embody the architecture of the model in the chip itself.

I don’t really like that. I kind of wish that we were doing a little bit more of a breadth-first search, where we would explore different kinds of architectures and find their relative strengths and weaknesses, and hopefully bring them together. We’re not one thing, right? We have a lot of modules in our brains. It’s just fundamentally weird on some level, and you would expect it to be brittle in some sense to take one relatively simple thing and just stack layers of that.

The solution that nature found in humans is a lot more complicated. It feels like if you want something to be robust in various ways, you would probably want to have different modules. So, I do kind of agree that it would be nice if we were doing a little bit more breadth of exploration rather than just trying to jam this one thing as hard as we can until we can all retire, or whatever exactly the dream is supposed to be.

At the same time, I think where I would disagree with the LeCun school pretty strongly—and I honestly think this is kind of a closed question, although he still disputes it and you can find people who will—is that, for one thing, the AIs are not trained anymore on next-token prediction in the way that they were.

RL is not next-token prediction in a fundamental sense, right? The task that the model is given is not, “Here’s a bunch of text. Can you predict what comes next?” What it’s trained on now, the signal it’s getting now, is, “Did you get the right answer?”

The right answer could be a fully verifiable mathematical proof or numerical answer to a question, or it could be one of these things that is 49,000 evaluation criteria on a huge medical corpus. But it’s not that it was supposed to be this token and you gave it this token. It’s now, qualitatively or quantifiably, as the case may be, “Did you get the right answer?”

Then that signal is translated into a gradient update through a—there are obviously a lot of different mechanisms—but GRPO, Group Relative Policy Optimization, is basically comparing, for a given model, “Here are a bunch of attempts that it made. Some were right, some were wrong. Let’s use that to create a direction in weight space and move in the direction that would make the answer that was right be more likely next time.”

I think it’s something people should all update on at this point: we’re not just doing next-token prediction anymore. That’s still part of the process, but it’s not the whole story.

The other thing is, I think it’s also very clear at this point that the AIs do have world models. We can look at the internals, obviously not anywhere near as much as we would like to understand about what’s going on inside them, but we do have enough of an understanding at this point to create things like the Golden Gate Claude experiment, which I’m sure you guys have seen.

They trained what’s known as a sparse autoencoder. A huge problem in terms of figuring out what’s going on inside a neural network is that it’s very dense, right? The width of a model might be, depending on how big it is—they usually go like powers of 2—4,000, 8,000, or 16,000 wide vector of activations that sit between the layers.

After each layer, there’s this sort of bottleneck of, “Okay, we’ve got, let’s say, 16,000 numbers that are all some precision floating-point number, whatever, and that represents the state of play.” Well, there are obviously way more than 16,000 concepts, right? So that means we can’t just have one space for each concept.

Instead, we’ve got to have what is known as superposition, which means if just point 1 is lit up, that might mean something. If points 1 and 2 are lit up, that means something else. If points 1 and 3 are lit up, that means something else; 1 and 4; 1, 2, and 3; 1, 2, and 4—out to the vast space of combinatorial possibility.

The sparse autoencoder basically tries to untangle all that stuff and say, “Can we get to a representation where we can look at a sparse number of—” These sparse autoencoders are computationally expensive in their own right, and they’re millions, I think tens of millions, of activations wide. But they sort of branch this very dense, superimposed concept mess out into this sparse space.

Believe it or not, it works, and you can now look and say, “Okay, here are the 10 concepts that are most active in this network at this time.” Maybe you’re kind of fooling yourself, but the proof is in the pudding, where they are able to then say, “Okay, now that I know what pattern of activation corresponds to this concept, I can intervene on it.”

So, with Golden Gate Claude, they found the Golden Gate Bridge concept out of tens of millions of concepts that came out of this process, artificially turned it up, and now you’ve got a model that just wants to talk about the Golden Gate Bridge.

There’s more stuff that’s happened, obviously, since then in interpretability as well. I wouldn’t say that AIs have perfect world models, but they definitely have some world model. I also wouldn’t say people have perfect world models, right? We made it into the 1900s without any sense of relativity, because the world model we had was good enough for us to get by in the domain we were working in.

It’s going to be an interesting question of whether AIs can start to create those conceptual leaps, like a pre-relativity sort of jump. Again, exactly on what timeline that comes, I’m not so sure, but there is a world model inside the AIs. There is a conceptual understanding that is definitely richer than pure stochastic correlation of tokens. That has been demonstrated, I think, at this point, quite conclusively with the interpretability techniques that are out there. So I guess, where does that all leave us?

Speaker 1

Yeah. Well, I think that’s a pretty comprehensive response to the LeCun objection. Just to translate for our slightly more general audience, although it was very interesting and we’re glad that you went to technical depth there, basically what Nathan is saying is that, first of all, it’s very important to remember that this objection—that next-token prediction can’t be the core of intelligence—doesn’t necessarily apply, because it’s important to remember that we’re now doing reinforcement learning. We do next-token prediction with LLMs to get this sort of basis of intelligence, but then we do reinforcement learning on real-world tasks.

That intuitively seems more like what you’d expect the right training objective to be to get to AGI or something like that. Separately, Nathan is also saying that with sparse autoencoders and the Golden Gate Claude example, we now have techniques in AI that allow us to look inside the model when it’s responding to a query and see which concepts, in some sense, are activated by that query. We can see that there are concepts within the neural network that correspond to objects in the real world, right? So even though it’s just an LM, in addition to reinforcement learning, it does have this world model where we can say, “Oh, look, that’s the Golden Gate Bridge concept,” and it’s lighting up when we ask it to answer a query about the Golden Gate Bridge. Is that accurate?

Nathan Labenz

Yeah, that’s great. Phenomenal job. One other thing I would add is that the explorations that have been done of the embedding space, or the latent space, are also really interesting and revealing.

There was one study—this was a couple of years ago, and these things have gotten much more sophisticated since—that showed you can do vector operations around the latent space. For example, if you take the embedding for “man,” move to “king,” and look at that direction, then apply that same direction to the embedding for “woman,” you get “queen.” There’s an order, a sort of conceptual coherence, to the way in which concepts are represented spatially in this super-high-dimensional latent space that is clearly meaningful.

Exactly how it’s meaningful, exactly what it’s learned, what mistakes it may contain, or what aspects of a true grand unified theory of everything it doesn’t have—those are all open questions. But I think that that sort of thing quite demolishes the idea that it’s all just noise or that there’s some sleight of hand. I think that the organization of its own internal map of the world reflects some real understanding going on.

Maybe not human-level understanding. I always say human-level, but not human-like. They could be quite alien, but that doesn’t mean they don’t understand. They don’t have to understand in the same way that we understand in order to meaningfully understand. So I think that is, again, pretty well resolved at this point. I honestly don’t know why some people can’t update on that dimension. It’s quite strange.

Speaker 1

Okay, moving on from the capabilities and training discussion, let’s talk about hardware and energy, these other inputs to AI model progress. It sounds like you’re pretty confident that either with RL or with some additional paradigm, we’ll get to very transformatively powerful AI. But I assume that part of that thesis involves us continuing to scale these inputs, such as talent, hardware, capital, and energy—and mainly energy and hardware.

A pushback people often make is that there’s just not that much more energy to divert to AI model training and inference, right? I can’t remember the numbers, but there’s some graph showing that the total amount of U.S. energy production isn’t increasing very fast at all, while AI consumption is increasing very fast. At some point, it’s going to catch up and we’re going to be bottlenecked by energy. With chips, it takes a really long time to make new chip-fabrication facilities, and we’ll just run out of chips to train AI models on. Even with capital, maybe we’ll just run out of money in the world to invest in building new data centers. Which of these bottlenecks do you see, if any, as most plausibly being a major bottleneck to continued progress?

Nathan Labenz

Yeah, I guess the first thing I would say is that I don’t think any of those bottlenecks are really fundamental. They’re more cultural or sociopolitical or whatever. There’s a lot of energy coming from the sun all the time, and the question is how much we’re actually going to harvest and harness. That’s where it becomes a political debate around who’s going to be allowed to build what, where, on what timeline, with what permits, with what impact, and so on.

I think the degree to which AI is energy-intensive is often overstated. I actually did a whole episode on this with a guy named Andy Maseley, who has really been fighting this fight online in a pretty dogged way. There are a lot of interesting comparisons, but a frontier chip today, like an H100 or whatever, basically uses the same amount of energy when it’s on as a microwave or an electric teapot. One query is maybe on the order of running your microwave for 1 second.

You’d have to be making a whole lot of queries—and increasingly people are—for it to be moving the needle on energy consumption. People don’t think twice about putting something in the microwave for 2 minutes, and that’s probably more energy than most people use with AI on a weekly basis. Now it’s ramping up, and it is starting to add up, and it’s going to get to the point where, if we can’t add any capacity, then we’re going to have a bottleneck for sure.

But again, those bottlenecks aren’t super fundamental. China doesn’t seem to have them, right? They’re adding as much electricity to their system as—I think it’s, gosh, I won’t say an exact number, but it’s some relatively short period of time in which the Chinese economy is adding as much electricity as the entire American capacity. They have about 4 times as many people as well, so they have a long way to go to build out everything they might want. But it just shows that it can be done.

It can apparently also be done in the Gulf. I recently talked to Sam Hammond, who’s an economist and a very AGI-pilled thinker in Washington, D.C. He had just been to the UAE on a trip, and I asked him, “Why are we doing these deals with the UAE? All I hear is that we want AI to reflect American values and take American values around the world.” I said, “I’m not sure that the governments of Saudi Arabia and the United Arab Emirates are the greatest partners we could have in projecting American values. Why are we doing these deals?”

It seems like, honestly, a big part of the answer is that they don’t have issues with putting up a new plant. They can just do it, and it’ll happen fast. That way, we know that even if we can’t do it here, we can at least do it there. That’s all just characterizing the bottleneck and saying that there’s plenty of energy. It’s a question of who will be allowed to get it, under what circumstances, on what timelines, with what permits, et cetera.

Chips are harder, for sure, because it’s a very specialized thing. In terms of what would be the most likely reason we wouldn’t get economy-transforming AI in the next few years, if you asked me, I would say something happening to the chip fabs in a major way that throws production off to the point where chips are super scarce. Maybe we can’t scale the training runs, full stop. Or even if the training runs can still scale, there’s just not enough inference to go around.

We might have really powerful systems, but we just don’t have enough access economy-wide for people to deploy them and automate all the things that seem like we’re on track to automating. So if I had to pick between energy and chips, I would say chips. But that seems like a tail-risk scenario. The whole economy depends on it at this point, so everybody is incentivized—certainly the political class is incentivized—to make it work.

The corporate managerial class is pretty incentivized to make it work. There is some tail risk that mainland China makes a move on Taiwan and that could be a huge disruption, but it seems like, as far as I can tell, the most likely disruptions would be tail-risk-type things, not a fundamental disruption. There is plenty of sand, which is where silicon comes from, and that is scaling, too. We are starting to get chips in the U.S., from what I understand.

I’m not an expert on this, but I understand the yields have been decent in the U.S., maybe even a little bit ahead of schedule. People thought, “Oh man, it’s really going to be a few years and a few generations and a lot of iteration to get this stuff to be somewhat competitive.” It seems like it’s come online reasonably well.

This is a theme in my thinking in general: I’m mostly worried about the tail risks. I’m mostly worried about AI going wrong in some really weird way, and in terms of what would prevent it, I also think the tail outcomes are the most likely to put me in the camp of being catastrophically wrong.

Speaker 1

Yeah. Okay, awesome. Let’s talk about that, then. Let’s talk about this risk of alignment, which is the primary concern among the community right now. You’ve had hundreds of these conversations with safety researchers, lots of people, and builders over the last few years. Net-net, what is your analysis of the alignment problem? Has it become harder or easier over the last couple of years, or the last several months?

Nathan Labenz

For starters, I think that, reflecting on my hundreds of conversations, one of the strangest things in the world today, full stop, is the fact that the disagreement among very plugged-in, very informed, very smart people has not really been reduced much at all, even as we’ve gained a ton of information over the last couple of years about the trajectory of AI. I think that is super strange, and I’m honestly pretty confused by it.

The one thing everybody seems to agree on is that the timeline on which we should expect this to happen has come in. There’s still disagreement about that, but Helen Toner, who’s the head of the Washington, D.C., think tank CSET and was previously on the OpenAI board, put her finger on this phenomenon with a blog post that said, “Even long AGI timelines have gotten super short.” In today’s world, if you say you don’t think you’re going to see AGI until 2035, you’re an AI bear. Only 5 years ago, that was considered quite aggressive, and most people were more like, “I don’t know, 2050, maybe not in my lifetime.”

There’s been this massive compression of the timeline and obviously this huge jump in capability. Yet on these fundamental questions of what’s going to happen, there’s still total disagreement. I think that is a very weird phenomenon that I can’t fully explain.

Speaker 1

Yeah, where do you think this disagreement is coming from? Is it hope or some sort of pushback that you’re seeing from these very intelligent people who are still part of the field, or do you think there are legitimate thought processes behind why there’s so much disagreement around alignment?

Nathan Labenz

I’m not sure I can summarize the conclusions with super-high fidelity right now, but CSET, under Helen’s leadership, did a workshop where they tried to bring people together and assess, specifically on the question of recursive self-improvement, how big a deal it is going to be. Do we run the risk of this whole process getting away from us entirely?

Even on that somewhat reduced question, there was still very wide disagreement. Some people were like, “I don’t think it’s going to be that big of a deal. It’ll make people a little more efficient, whatever, but it’s not going to bring about some phase change.” Other people were like, “As soon as you get an ML researcher that can do that, you go from—I don’t know, maybe there are 10,000 people today who are really working at the frontier of ML research globally—to 10 million.”

That’s got to start having a huge effect. It’s going to be really hard to control a situation if we all of a sudden 1,000× the number of researchers and they work at potentially thousands of tokens a second.

They sort of came to the conclusion that people were working from different conceptual paradigms, and that these paradigms are pretty good at taking new information into account and explaining it away. You have theories like bottleneck theory or O-ring theory, where you’re only as good as your weakest link, basically. As long as you think that, you can say, “Sure, okay, the AIs can do this, but they still can’t do this other thing.” There’s still a weak link, and there are still going to be these bottlenecks, so the whole thing isn’t going to get too crazy.

The flip side of that is a kind of jaggedness argument, where people will say, “Okay, sure, the AIs can’t do this yet, but look at what they couldn’t do 1 year ago or 2 years ago. They couldn’t do basic math. Now they’re solving unsolved math problems.” Sure, there’s still jaggedness, but the last time you told me about jaggedness, you told me they couldn’t do basic arithmetic. Now we’ve got unsolved math problems.

These perspectives seem to be really grounded in worldview priors, or the paradigm that people work in, and it’s proving really difficult to get to a real meeting of the minds on those.

Speaker 1

Yeah, I saw Ajeya Cotra talk about a similar thing on the 80,000 Hours podcast.

Nathan Labenz

Yeah, she was saying a similar thing.

Speaker 1

Yeah. Great episode. She said something similar: the economists who expect that we won’t enter some new GDP growth regime always point to all these bottlenecks. Technological diffusion is always slower than people think. On the other side, people think that these models are simply outdated, and that these people aren’t really taking seriously what it means for AI capabilities to be at the given point that we’re conditioning on.

I’m wondering, do you think it’s possible that people are talking past each other and aren’t actually talking about the same level of AI capabilities when they say, “Oh, it’ll only uplift ML research somewhat”? Maybe they’re just talking about not-that-powerful AI. But if they were actually talking about fully automated ML researchers that are as good as humans, never have to sleep, and run faster, then they would see recursive self-improvement as more plausible. Do you think that’s possible?

Nathan Labenz

I think that explains some of the disconnect, for sure. I do think people are often talking past each other. I’ve started, at least in some cases, to begin my interviews with, “How AGI-pilled are you, and what do you expect to see over the next couple of years?” If that’s not established, and I think they’re thinking one thing when they’re not, there can be quite odd disconnects downstream of that.

I think getting those assumptions on the table early and at least cross-comparing them is usually a productive thing to do. I don’t know that it explains all of it. I do think there is still this idea that, because you do hear a sort of, “Sure, even if they could come up with good ML research experiments and this and that, there’s still going to be this other bottleneck,” there’s always another bottleneck.

It seems at times—and I’m not in this camp, so I don’t want to be too blithely dismissive of it or unfairly critical—that there’s an aspect of faith in the idea that there’s always another bottleneck. I would contrast my position with some of those positions in the sense that a big part of what motivates me is that I don’t think what I’m saying has to be guaranteed or proven right in order for us to be really motivated by the possibility.

I’m happy to leave open the possibility that maybe the bottleneck people are right and there’s always going to be another bottleneck, and all this will stay under control. Or maybe there’s some plateau around a quarter’s worth of work and we just can’t quite break past that, and there’s some weird phenomenon that we can’t explain, so we can’t ever quite get there.

I think that might be great news if such a thing were true. But I’m happy to leave that question open for the time being and just say, “I don’t know.” We don’t have a great account of what that would be if there were some fundamental bottleneck that we’ll never get over.

In the absence of that, I’m not persuaded by people just generally gesturing that there will always be another bottleneck, because if they’re wrong, we’re in for a really wild time. In terms of what we should do about where we are and what might be coming, I think the worst mistake we could make would be not to take it seriously enough and to content ourselves with a story that it’ll all self-regulate and we’ll be fine.

Because I don't see a great argument that that's true, and I see at least decent arguments that it might not be right. Again, I look to our own history, and I'm like, well, we've driven a lot of other species to extinction, including our closest cousins. Some of that was by accident. A lot of it was by accident, right?

It was just small bands of people going out and doing what they were doing to survive, and surviving meant hunting large animals and eating them, and using their bones for tools and stuff. A lot of those animals went extinct, and it wasn't like a coordinated master plan. It kind of happened by accident a lot of the time. So I'm just like, oh my God, we don't have any real guarantees, as far as I can tell.

Plot armor is, again, an unfair dismissal of the more sophisticated people who have theories about there always being another bottleneck, or what have you. But I do worry that a lot of what goes on among less sophisticated people who don't want to have to deal with this, and would rather believe that everything will be fine, is some sort of plot armor thinking. They're like, “Well, I don't know. I feel kind of like a main character, and humans feel kind of like a main character, and so you can't take the main character out of the story, right?” I just, unfortunately, don't think that that's likely to be the case.

Speaker 1

Yeah. So you find the bottlenecks argument—in terms of all the discussions that we had, maybe with research, maybe with capital spend, maybe with energy, with chips, or just general societal processes as a whole that slow down the overall technological diffusion—but overall, it seems like you're still relatively in the more pessimistic camp. Is that correct?

Nathan Labenz

No. When people ask my P(doom), I usually say 10% to 90%. A good friend of mine once told me we should think less about and argue less about exactly what the numbers are, and more about what we can shift them to. So I try not to worry too much about getting into it. I think 1 significant digit is all you get on P(doom), is one funny way I've heard it said.

I do want to say I have actually gotten probably a little bit more optimistic over the last few years. I started reading Eliezer Yudkowsky in 2007, and I read the early visions of the paperclip maximizer and all that kind of stuff. I can't speak for Eliezer, and I think he's got somewhat revisionist takes on what he really meant at that time. Sometimes when I see what he's saying now that he really meant, I'm like, I don't know, that's not exactly what I took away back then when I was reading your original work. But whatever—all of that is kind of discourse.

What I understood, and what I think a lot of people feared, was a very small system that had some extremely concentrated form of intelligence, that had found the right priors and the right inductive biases to be hyper-rational and insanely effective. Given anything to optimize, it could just optimize it to this extreme state and tile the universe, or whatever. There was also the idea that it was going to be hard to get such a system to understand human values.

What we value was very gradually and haphazardly encoded into us by an evolutionary process over a super, super-long time, dating back even to before our species, right? Other species care about their young and seem to be sad when they lose their children and stuff. So this isn't even just human. The whole of evolution has kind of led us to be what we are and have the very complicated value set that we have.

It was generally understood—although there are some different takes on that history now—that we would expect it to be hard to get AIs to have a real understanding of what we care about. Now I look at the models that we do have, and I'm like, well, actually, they do have a pretty good understanding of what we care about. It's not perfect, but I don't think it's crazy to say that Claude is probably more ethical than the average person, and certainly more sophisticated in its approach to ethics than the average person. Claude does generally seem to want to be good in a meaningful way.

When they let Claude talk to other Claudes and just let it do whatever it wants to do, it kind of seems to want to bliss out or something like that. So there is some sort of locking in of values that I wouldn't have expected to be as easy as it seems to be. There is some sort of internalization—or identity formation, maybe, is a better way to say it—that at least some of the models seem to have. That at least suggests to me that there's reason to hope that we could really get there.

Could we create an AI that loves humanity? When I first heard that question, I thought it was laughably out of reach. I think I heard that from Scott Aaronson, who said that Ilya Sutskever asked him, “Hey, do you have any idea what the Hamiltonian of love is?” Or some crazy question like this. Scott was like, “I can't really help you with that, Ilya.” I remember just being like, oh my God, that's the kind of question they're asking. We are screwed.

Now I'm like, well, I don't know. Maybe there was a little bit more to it. I trust Claude—not fully, but more than a human assistant. If I were going to give myself the choice between hiring a human assistant, with the opportunity to interview them and do some vetting, whatever, or Claude, which one would I give access to my email? I think I would trust Claude more than a person I interviewed a couple of times with all my most sensitive information.

There are cases where Claude has blackmailed people. There are cases where Claude has done various things under pressure. But I think I have better odds with Claude. So overall, I have become a bit more optimistic, but we're definitely not anywhere near out of the woods.

Speaker 1

Right. I guess maybe let me try to steelman the more pessimistic view. Sure, LLMs are language-native, and we can communicate our moral preferences via language. But that wasn't really what the people who were very concerned about alignment were talking about, even back in 2007. They were talking about AI agents that were goal-directed, and they were anticipating something like reinforcement learning that gets us these more goal-directed agents that are able to reason over longer time horizons, have a goal, and go out in the world and achieve it.

I guess the case for not updating toward optimism with respect to AI safety is that we're still going to have to deal with this problem. We're doing reinforcement learning. We're trying to get agents to go out in the world and do things for us, and we still don't know how to define a reward objective that is fully what we want optimized, to the maximum degree.

So we have this paradigm where we got these LLMs, and that was maybe an update toward AI safety being a little easier. But now we're back to reinforcement learning, and these same concerns apply. What do you think about that objection?

Nathan Labenz

I think that's a pretty good argument, and I certainly don't want to leave people thinking that I don't take that seriously or that they shouldn't take that seriously. I absolutely think we've got more questions than answers. When I say I've become more optimistic, it was starting from a not-super-optimistic place.

5 years ago, or maybe a little more than 5 years ago, I would have said powerful AI seemed a long way off, and if we did stumble on it, we had very little hope of controlling it. Now I'm like, it seems closer, but we have maybe a little more hope—while definitely still having a lot of unanswered questions.

I think one thing that I do see a little bit differently than the most hawkish people: I recently had an exchange online where, over a 48-hour period, there were a couple of profiles of Amanda Askell, and people were commenting on her in all sorts of different ways. Elon was attacking her and whatever. I weighed in and said, “For my part, I have become quite a bit more optimistic that it's at least possible to create an AI that, in some meaningful sense, loves humanity.” I have to give her and the Anthropic team a lot of credit for that.

Of course, I got a lot of replies saying, “Well, this doesn't scale to superintelligence,” and making arguments along the lines of the one you just made. Again, I think those are all very serious and worthwhile concerns. But one strand I detected in a lot of those responses is that people seem to be imagining a system that is so much more powerful than anything else that, if it goes wrong, it's over.

You get these kinds of ideas from Eliezer's “List of Lethalities” post, and I think that's shaped a lot of people's thinking. It's like you have to get this absolutely right on the first try. I do think that's not the shape of the AIs that I'm seeing in the world today. That's a quite hypothetical state of affairs.

If you were to drop in an AI that's just so much more powerful than anything else, and it's weirdly goal-directed, and it doesn't love humanity, and you tell it, “Make paper clips”—.

Yeah, maybe we could all like tiled by paper clips. But the world right now is much more like an emerging ecology of AIs, where there are a few frontier ones that are roughly competitive with each other. Certainly, it’s not like one Claude instance is going to take over the world.

What we’re talking about is much more of a wave of things where, simultaneously, millions—or one day billions—of Claude instances, GPTs, Geminis, and whatever else collectively transform the world. That has a lot of problems, challenges, and open questions with it, too. But I don’t worry as much right now that we’re headed for a world where one system runs away from everything else to the point where, if it takes one wrong move, or there’s one bad prompt or jailbreak, everything is lost.

It seems to me like scaling laws are, in a way, protective. I think this is kind of luck, or maybe it’s fundamental physics, but it’s not physics that we understood coming into this. You get these algorithmic advances, and they move the needle—they deflate exactly how much compute you need to get to a certain level—but they don’t tend to make it so you don’t still need a ton of compute to get to the very high levels.

I think Zuckerberg actually had one of the more interesting takes on this that I’ve heard. I don’t think he’s distinguished himself as a great AI safety thinker over time, necessarily, but he basically said, “At Meta, we deal with scammers and spammers all the time. The big advantage we have over them is that we have a lot more compute. We have way bigger, way more powerful systems than they do.”

“So we’re seeing everything. They’re trying to spam and scam here and there, but we’re seeing everything and monitoring everything. Little things happen all the time, of course, but broadly, we can keep it under control.”

You could imagine a somewhat similar dynamic with AIs, where one AI—even if it were the single most powerful AI in the world—as long as it’s not orders of magnitude more powerful than everything else, would be facing a whole ton of other actors. They’ve got all their compute, they’ve got all their instances, and they’re all monitoring for whatever they’re monitoring for. Hopefully, that can kind of balance itself out.

Then, of course, you’ve got your gradual-disempowerment concerns, which is that maybe it all ends up in some equilibrium with each other and there’s no place for humans in it. That’s another thing that I do think is absolutely worth taking seriously. But I just don’t see right now that we seem to be on track for this kind of runaway scenario where, if anyone builds it, everyone dies.

I think there is an AI for which that’s true, but it doesn’t seem like anybody’s particularly close to building it. That’s a really important part of the analysis from my perspective.

Speaker 1

Yeah, okay, great. So then, when we think about your concrete hopes in terms of achieving this world in which we solve the misalignment issue, can you give us specific achievements that you think are necessary to end up with an aligned AI that doesn’t result in doomsday for humanity?

Does it involve training a model that loves humans? Does it involve some sort of mechanistic interpretability actually scaling? Or, I guess, the subgoals of training a model that is aligned or loves humans—does that mean mechanistic interpretability scaling? Does that mean alignment by default, because models trained on human values just end up being aligned?

What are the concrete subgoals that you think have to be achieved to get this alignment issue solved? I’m always interested in something that could, quote-unquote, really work. I ask people for this all the time: Do you know of anybody who is working on anything that could credibly really work, in the sense that I can sleep well at night knowing that there’s something that really works?

Nathan Labenz

Basically nobody has anything. In the absence of that, all the frontier companies seem to be taking a sort of defense-in-depth strategy. The hope is—and you can arguably say this is inherently the nature of intelligence—that it’s kind of unpredictable.

Again, the same thing is true of humans. There’s never been anything that really works to make sure a human never does anything wrong. Maybe it’s unrealistic to think we could ever get that. I would still love to see people try, but it seems like where we’re headed is a “everything in parallel at once” sort of strategy.

We might not get the AI to the point where it never does bad stuff, but if we can drive that rate low, then that’s better. If we can put a monitoring layer on top of that and catch 90+ percent of the bad things it still does, then that’s better. If we can have additional monitoring systems that ban the accounts of people who are bad actors, then that’s better.

We can also really invest in formal methods to improve cybersecurity across the board, so that we can take certain risk surfaces entirely off the table. That’s not going to fully solve all of our issues, but in terms of things that could really close down problems, formal methods to verify the security of software are one area where there does seem to be an opportunity to create genuinely secure software.

That’s kind of on the horizon, and it seems to be on the verge of having a moment. Then, of course, we’ve got the bio risk. We should probably have PPE stockpiles, and we should probably have all the things that we should have had as of the last pandemic.

Something we just bought for our son’s hospital room is an ultraviolet light that’s supposed to kill microbes—I think this is quite well validated scientifically—and is pretty gentle on the skin. You can just shine it in the room all the time. We should probably invest in scaling out that kind of capacity.

We should have better wastewater monitoring, so we know when things are popping up, because it’s still probably going to be the case that things are going to be popping up. We should have vaccine platforms—which we do have—that are very quickly programmable.

I’m sure you’ve heard the story of how quickly the COVID vaccine was designed. In a few days, before the pandemic really even took off in a serious way in terms of case numbers, the design of the vaccine was already there. It took us a long time, of course, to go through all the trials and actually get it to people, but it was just a few days to create that vaccine.

There’s also going to be mechanistic-interpretability, internal-monitoring-type work. The company Goodfire, which does interpretability, just put out an agenda called “Intentional Design.” They’re developing ways to try to understand, at each step of the learning process, what the model is learning, and to shape what it learns so it hopefully doesn’t learn certain problematic things and does learn other good things that you want it to learn.

Then there are AI control techniques. Redwood Research has, I think, done an incredible job of gaming some of these things out. That’s another 80,000 Hours episode I would strongly recommend if you guys haven’t heard it, and I recommend it to anyone.

It’s Buck Shlegaris talking about how we get productive work out of AIs, even assuming that they’re out to get us. They’ve actually made pretty good progress in terms of building out a portfolio of strategies. I think that’s kind of where we’re headed.

We’re looking at a world where there are AIs everywhere. Hopefully, none are so much more powerful than anything else that a single small individual instance, or a small pocket, poses some existential risk. Then we bring every other strategy we have to bear, and that gets us a few nines of reliability.

There will still probably be some crazy bad things that happen, but hopefully they can be contained enough that the world overall is good. When I tell that story, we definitely need to scale up our investment real quick, for one thing, because the amount of money and resources going into making the AIs more powerful dwarfs the amount going into all these other things.

I don’t think we’re very well calibrated or balanced in terms of where we’re putting our efforts. But I do think we have a bunch of different agendas that seem like they can all take a bite out of the problem. Maybe you take 20 bites out of the problem, and you don’t have much problem left.

Holden Karnofsky used to say that Eliezer used to say, “Death with Dignity.” He was like, “We’re probably going to lose this, but we should at least make a real effort.”

Now Holden Karnofsky is a senior advisor at Anthropic. Recently, Anthropic updated its Responsible Scaling Policy, basically backing off from some of the commitments it had previously made to pause development under certain circumstances. They’re no longer committing to that.

He put out a long defense of why, but the previous thing he had written was “Success Without Dignity.” He was like, “I don’t think that we’re doing a great job collectively of trying as hard as we should be trying.” The risk that we’re running is a lot higher than I, speaking as Holden, would like it to be.

But the problem does look much more tractable than it used to look. Five years ago, when people would come to me and say, “What can I do for AI safety?” I’d be like, “I don’t really know.” And now I’m like, “I got a long list of projects that people can work on that all at least seem to help.” So, I wish I had an answer that was like, “Do this and it’ll really work.”

I don’t think I’ve heard any credible claims in that direction, really, except in cybersecurity, which is obviously only part of what the world is going to need. But maybe all that stuff could add up to a win.

Speaker 1

Great. Given that we have this enormously transformative technology, and it probably is the most powerful dual-use technology in human history—in my opinion, and probably the consensus opinion as well—way more consequential than even nuclear weapons, which up until this point were probably more powerful, do you think that these model providers or, you know, these model capabilities should be in private hands, as they currently are, at all? Do you have a case in which you support the government, in some sense, nationalizing AI frontier development? What’s your take on that?

Nathan Labenz

Well, Biden used to always say, “Don’t compare me to the Almighty; compare me to the alternative.” I would probably say the same thing about the frontier AI companies. I certainly don’t think they’re all acting perfectly, and I certainly think the race between them has the potential to get out of hand. So, I’d like to see government action of some sort, and there are a lot of debates there on exactly what sort of policy we should have. But no policy at all doesn’t seem like a winner to me.

I say that as, broadly speaking, a lifelong techno-optimist libertarian who mostly would rather see the government stay out of these things. But this one does seem like a qualitatively different thing, where some government involvement seems prudent, if not outright necessary. I would not go for nationalization, though, for the basic reason that I just don’t trust the government that much. I look at the current leadership and I’m like, “These are not the people I want to be taking my cues from.” I take Sam Altman, Dario Amodei, and Demis Hassabis over Pete Hegseth.

In that standoff which is happening this week, I support Anthropic 100% for putting some limits on what they want their technology to be used for, and I hope they stand by it. If the government comes down on them, I’ll preferentially send my tokens their way, even if it weren’t necessarily always the best performance, because I do think that is a really important stand for them to take.

So, could I imagine a government so competently run and so pure of heart that it would maybe make sense to nationalize? Yeah, I can imagine a lot, I guess, but I don’t think we’re anywhere close to having that government today. I would rather see some competition and some hopefully healthy balance, and maybe some of the worst excesses reined in by the government. But certainly not a takeover and putting this under the military. That sounds like a recipe for disaster.

Speaker 1

Yeah, that makes sense. I’m also always skeptical, though, of what we talk about as corporate incentive alignment. I think oftentimes we think about situations in which, most notably recently, like social media, you have these private corporations driven by usage rates or by how much capital they’re pulling in. Whether that be for the sake of more development or purely for the sake of desiring more capital, that incentive misalignment causes a situation where the individuals or the consumers ultimately get hurt.

So, I guess on this continuum that we described, where we have pure nationalization of the entire technology or, on the other end, where there’s no regulation at all, can you talk a little bit more about where you sit there and what that regulation might look like if you do think it is somewhere in the middle, like you kind of said?

Nathan Labenz

I probably, honestly, oppose most of the ideas that rank-and-file politicians are going to come out with for regulation. I want self-driving cars. I don’t want them to say ChatGPT can’t give you medical advice. I don’t want them to say you can’t get therapy from a bot. A lot of that stuff, I think, ends up just being guild-style protectionism, and it hurts kind of everybody.

Basically, I think what government should do is try to solve the race-dynamic coordination problem and do the work to hopefully try to minimize the truly extreme risks. It does seem like right now, over the next 2 years, the companies believe they’re going to create automated AI researchers and recursive self-improvement loops, and nobody really knows where that goes. I don’t think that is a great thing to be happening in a few relatively small and kind of ideological companies.

I do think it’s also underappreciated and underdiscussed how ideological these companies are. Even the definition of AGI as something that can do everything better than humans isn’t a neutral frame. It’s a definition that has a lot baked into it. I think there is certainly diversity within the companies. It’s not like all of them are successionists or whatever, but there is a strain of that.

Then there’s also a strain of, “Wouldn’t it be amazing if we could just replace ourselves and not have to work anymore?” And that, I think, has kind of taken on a bit of a life of its own. So, I do think there’s a role for the government to say, “Okay, we’re not here to tell you you can’t get free medical advice, but you can’t run super-high-energy experiments with no transparency when you yourself have said for years that there’s a significant risk that this goes really badly.”

That’s one of the strangest things about this whole thing, right? All the leaders have said for years that there’s a significant risk that this goes really badly. That’s the thing that I would like to see the government really focus on: Is there some way to rein that stuff in? Make sure we have better extreme-risk mitigations in place. Make sure we have better safety plans. It’s tough. I don’t think there are great or easy answers here, but that’s where I would definitely want the government to focus its energy.

Speaker 1

Okay. Another major governance question that interacts with a lot of these concerns about catastrophic misuse or misalignment is—you mentioned the race across American labs, but there’s also the race with Chinese labs and the Chinese government, of course. The default high-level strategic vision in D.C., as far as I can tell, and in S.F. as well, I think, is: We’re going to continue to build, continue to race ahead, building AI capabilities. We have a lead over the Chinese labs right now. There’s debate about how significant that lead is, but there’s no doubt that we have a lead.

We’ll continue to build this lead as much as we can through export controls and by improving privacy and security at the model providers. And then, when it comes time—when we’re right on the cusp of extremely transformative AI, or even somewhat into the intelligence explosion, so to speak—we will have time to slow down and coordinate, and not push capabilities even further toward superintelligence. Again, that’s contingent on having a significant lead over China.

But a sort of assumption of this strategy is that the worst outcome would be both us and China being neck and neck, continuing to race, and being in this prisoner’s dilemma where, even if both of us were concerned about misalignment or something, we would have to continue to race ahead because otherwise the other person is going to continue to race ahead. What do you think about this high-level strategic picture?

Nathan Labenz

Well, it’s tough. I don’t want to pretend it’s not tough. But my outlook has always basically been: I hate that idea. And I don’t think the few months—the difference is months, not years—are a long time to solve all these problems. We’ve spent much of the last hour and a half talking about all the different problems and the different facets, and we didn’t even touch on open source.

Open source could be a great counterbalance against concentration of power, but if we open-source the wrong thing, we can’t take it back. We may have just put a bioweapon assistant or even an autonomous creator into the public domain in a way that is going to have potentially super-long-term echoes, right? So, that’s another whole facet of the problem that you could spend many hours and many thousands of pages trying to figure out what to do.

I don’t think this idea that we’re going to have these couple of critical months, solve everything in that time, and then win somehow makes any sense at all. I don’t want to be naive with respect to China. And I don’t even want to say China, because I think it is the government of China, and potentially a relatively small cohort at the top of the government of China, that I think we quite rationally should be at least somewhat wary of.

I always come back to the real aliens in this situation being the AIs, not the Chinese. We have a lot more in common—the United States and China, as humans, fellow humans—than we do with the AIs. Is it going to be hard to build trust across the great ocean and the great civilizational divide? Sure, it’s going to be hard, and we’ve got ideological differences in the government.

Many, many challenges. But I would strongly advocate for starting to do that work now. For my part, I always want to talk to people in China. When another AI podcast does an interview with a Chinese researcher—somebody at one of their top companies—I ask if I can cross-post that to my feed, because I think we should have a lot more researcher-to-researcher-level communication.

I would absolutely invest in all manner of diplomacy and out-there ideas. There's the CERN project for high-energy physics; could we create some sort of shared, jointly controlled place where researchers from both countries could go to work together on very sensitive problems? It could be some small island in the Pacific, or it could be in Singapore. I don't know, but the idea that we're just going to decouple and race each other sounds deeply unwise to me.

I would take my chances with the possibility, at least, that we could come to some sort of shared understanding with fellow humans, versus recursively self-improving AIs that we hope will be better than their recursively self-improving AIs. All of which, by the way, is happening against the backdrop that we are looking more like China all the time. I don't like to say this kind of thing because it's not a very popular opinion, and it's not something I take any pleasure in.

What is happening again this week? We've got the Defense Department threatening retaliation against an American company because that company wants to hold firm on what I would consider to be a core American value: not using its technology for mass surveillance. What is it that people typically worry about when they talk about Chinese AI or Chinese values, or living in G's world?

I think one of the big things is mass surveillance, right? It's the idea that I can't speak my mind anymore, even in private contexts, because the government is going to Hoover up all that information and potentially use it against me. As far as I can tell, that's exactly what the U.S. Defense Department is trying to coerce Anthropic into doing right now.

And what? We're kind of losing the thread here in our delusions of grandeur—in our idea that we're somehow going to be the winners. I was just talking to my 7-year-old last night because he was asking questions like, "Why is there war? How does that happen?" I was like, "Well, almost always, the people who start it go down in history as the bad guys. And almost always, the people in the country that started the war, who were kind of convinced it might be good for them, end up regretting it."

I kind of think that's the unfortunate trajectory we're on right now. We're talking ourselves into this idea that if we can achieve strategic dominance, then we'll make them an offer they can't refuse. We're just forgetting that, first of all, we're losing ourselves in the process. And second, they get countermoves, right?

Taiwan is a lot closer to them than it is to us. And these fabs are easily destroyed, not easily put back together. So I really don't like that theory at all, and I try to advocate for something more conciliatory wherever possible.

That's not to say we don't want to be totally naive. It's good to have leverage. You could maybe get me on board with a policy of, like, let's not sell chips to China, but let's rent them freely to Chinese companies, which is somewhat the policy we have. I mean, they can buy from hyperscalers outside of China.

The momentum on this is the worst. It's America at its worst. When we think we have this enemy and we think we're going to unite to be the best civilization, that just never seems to go super well.

I hope to visit China sometime in the next year and do my tiny little part to participate in civilization-to-civilization understanding. I don't expect to move the needle, but as a gesture, if nothing else, I really think more people should be doing more of that kind of thing.

Speaker 1

Very cool. I'm glad we asked you that. I think that's an important perspective, and definitely not one that you hear articulated with that much clarity or depth very often. Thank you very much, Nathan.

This was an amazing, very wide-ranging conversation. I think we got into some good depth as well on some crucial issues. So, yeah, thank you so much, Nathan. This was an amazing conversation, and I really appreciate you coming on today.

Nathan Labenz

Yeah, thanks so much. That's very kind. Thank you, guys.