[BidClub_]
The Cognitive Revolution · · 87 分钟

Helen Toner:OpenAI 反思、适应性缓冲与战争中的 AI

Erik TorenbergNathan LabenzHelen Toner

YouTube
TL;DR
  • Toner 的基准情景不是超级智能即将到来,而是文明尺度的转型已经足够可能,值得现在就开始准备。2016年,“短时间线”意味着未来20年内或一生之内出现先进 AI;到2025年,它可能意味着2020年代结束前就出现超级智能,因此即便是“通往先进 AI 的长时间线也已经短得离谱”。对投资者而言,这指向的是对评估、韧性和政府能力的持续需求,而不是相信一个5年倒计时。

  • OpenAI 最新、最明确的披露是:Q*并未触发董事会2023年的决定。Toner 表示,董事会知道后来以 o1 和 o3 发布的推理研究正在进行,但没有收到所谓突破性信件,也没有据此采取行动:“那整篇 Reuters 报道完全是假的。”由于保密义务和法律责任,治理折价仍难以量化,完整公共记录无法形成。

  • 当前沿 AI 员工能够指出哪条规则被违反时,前沿 AI 监督的效果会好于仅凭主观担忧获得保护。Toner 倾向于要求披露安全与安保计划、能力和风险评估,以及建立清晰标准、供员工援引的内部流程。这些员工现在可能拥有最大的议价能力,因为他们“正在积极地替代自己”,而机构性失灵可能以“温水煮青蛙”的方式到来,没有一个显而易见的危机节点。

  • 适应性缓冲,是应对 AI 市场的战略答案:前沿开发成本越来越高,而昨日的前沿能力却在迅速商品化。DeepSeek 只落后部分美国系统1到2个月就匹配了推理能力,基础模型能力则约落后6到9个月,且成本更低,使永久性的非扩散控制越来越侵入性强、也越来越脆弱。实际应对方向转向疫苗产能、疫情监测、网络修复,以及在能力扩散前分发防御工具。

  • 只要发布不太可能造成严重且不可逆的伤害,迭代式部署才仍然有用。Toner 更偏好附条件的“如果—那么”闸门——在具备指定理解或缓解措施前不得推进——而不是固定暂停或基于投入的速度上限;后者既需要目前不存在的立法,也会立刻遭遇“中国怎么办”的质疑。在全面制度尚未出现的情况下,透明度、测量科学、可解释性、对齐研究和政府技术人才,是现实可行的基础设施。

  • “击败中国”同时充当了地缘政治论点和 AI 行业在华盛顿的最短游说路径。Toner 将竞争置于权力转移、海上通道和国际规则的背景下,谈话也涉及台湾。她表示,这一叙事恰好支持“资金”“政府合同”“不监管”和免于承担责任。因此,即便没有解决根本的战略问题,中国叙事仍会实质影响 AI 公司的经济利益。

  • 军事 AI 应按具体使用场景评估,而不能被包装成通用型战场伙伴。用于视域测绘、医疗分诊、舰船运动异常检测或数据库检索的边界明确工具,与让 LLM 生成“3个不升级冲突的行动方案”有根本区别。Toner 与 Amelia Probasco 提出的框架——范围、训练数据以及人机交互——把可靠性和对抗鲁棒性置于演示流畅度之上。

  • “军队版 AlphaGo”不是可信的近期均衡,因为真实战争无法被装进一个干净的模拟环境。战场、后勤、经济、政治和公众态度相互作用,而对手还会主动攻击模型的假设。Toner 的诚实结论是,她不知道超级智能之下民族国家、民主制度或中国共产党会变成什么;核心事实是战略不确定性,而不是某种已经确定的作战 doctrine。

摘要 · 为研究而整理的核心内容

1. 在 AGI 还不受重视时,Toner 就认真对待它

  • Toner 于2021年加入 OpenAI 董事会,但早在公司2015—2016年成立时就已认识这家公司及其许多人。当时她刚开始在旧金山从事 AI 政策工作。由于 AlexNet 于2012年问世、深度学习已经在推进,她起初觉得自己“来得太晚”;后来2018—2019年的关注浪潮以及2022年的 ChatGPT,又改变了她的看法。

  • 当时,明确以构建 AGI 为目标创办一家组织,即使在机器学习圈内也显得“奇怪”、 “逆势而行”;DeepMind 几乎是唯一一家认真这样表述的研究机构。Toner 认为,她在中国和国家安全领域的专业能力对获得董事会席位很重要,但同样重要的是,在真正理解这一议题的人还极少时,她已经多年认真看待 OpenAI 的使命。

  • 她最初的时间线,只有相对于那个时代才算短:未来“20年内”或“我们这一生”出现非常先进的系统,已经足以构成准备的理由。如今,“短时间线”可能意味着2020年代结束前出现超级智能;她对这一点仍远不如以前确定,但依然认为这种可能性“值得投入大量思考,也值得投入大量准备”。

  • Toner 否认阻止 AI 杀死人类是她从事这一职业的动机。她的出发点是历史性的:重大技术会“把社会变得更好或更坏,往往既有更好也有更坏”,而 AI 很可能在她有生之年带来这样的转型。问题在于她的工作能否帮助这一转型走向更好的结果,而不是 AI 是否“必然”会杀死人类。

2. Q*故事是假的,但 OpenAI 的背景仍受限制

  • Toner 目前受到的限制非常具体:董事会保密义务、在其中微小不一致都可能产生影响的法律程序,以及涉及不必被卷入公共争议人士的私人谈话。大量尚未公开的内容“几乎都是无聊的细节”,并不存在一个仍未揭开的“巨大、深层、黑暗秘密”,足以彻底改变她在 TED AI Show 上给出的图景。

  • 有一个曾经保密的事实如今已经明确,因为相关推理工作已经公开。董事会知道后来以 o1 和 o3 发布的研究正在进行,但“我们从来没有收到过什么关于突破的信”,也没有根据员工来信作出决定。Toner 的明确纠正是:“那整篇 Reuters 报道完全是假的。”

  • Erik 追问,OpenAI 改变世界的使命是否已经演变成一种英雄主义或“主角”文化,并受到精英绩效教练和情感抽离的塑造。Toner 不愿泛化:董事会成员并不适合描述员工的日常文化。她更担心结构性问题——技术工作通常可以与社会治理分开,但如果技术进展跑在社会前面、只有开发者能够避免巨大的后果,这种分离就会“脱离分布”。

3. OpenAI 的技术警告与政策立场正在分道扬镳

  • Nathan 描述了 OpenAI 反复出现的剧烈反差:那篇关于被混淆的奖励作弊的论文,是 AI 失效风险最清晰的警告之一;但不久后,公司向白宫提交的政策文件却展现出升级对华姿态,并要求获得使用受版权保护材料训练的广泛自由。他还将其与 Sam Altman 的“我们的价值观或他们的价值观;没有第三条路”并置,称该机构对外发声几乎“精神分裂”。

  • Toner 也觉得这种反差令人困惑,并认为它已经变得更加明显。一种可能的解释是,员工在公开评论和研究方向上拥有相当大的自由,而官方政策则经过不同的审查路径。她希望技术人员能够关注公司在政治上主张什么,因为这些信息可能与公司自身研究发出的警告严重矛盾。

  • 她优先希望缩小前沿开发者与其他所有人之间的“巨大信息差”。有用的披露包括能力与风险测试结果、安全流程说明,以及解释系统被训练来做什么的模型规格。目标不是让政府用清单规定答案,而是提供足够的可见度,让外部参与者理解这些系统并作出反应。

  • 美国围绕10^26次运算的报告门槛看起来“基础不稳”,但 Toner 没有听说它已经被确定取消;Nathan 则表示,欧盟 AI 法案正在围绕约10^25以上的模型制定透明度规则,同时面临将规则稀释的压力。Toner 的理由并不是超过某条算力线的每个模型都危险,而是最新、最强的模型承载着最高的“未知的未知风险”,应接受额外审查。

4. 危机发生前,吹哨人需要可执行的标准

  • 传统吹哨通常针对违法行为:例如 SEC 为金融 misconduct 提供了明确渠道。但前沿 AI 员工可能观察到不诚实或风险极高的行为,却没有任何法律被违反。仅仅规定“如果你担心,就拨打这条热线”,既无法为员工提供可用边界,也无法让公司知道如何处理。

  • Toner 偏好的设计,是把保护措施与披露或流程要求绑定起来。公司可以公开或私下提交安全与安保计划,由此建立一套标准,员工可以据此举报某项评估被跳过、信息被错误陈述,或承诺的缓解措施被放弃。即便只是强制要求执行某项内部流程,只要员工能够据实说“实际上,我们没有执行这个流程”,也会有帮助。

  • 制度设计必须考虑拟议系统的用户:技术员工可能缺乏法律常识、感到害怕、长期极端加班,也没有时间处理模糊地带。“如果用户是吹哨人,” Toner 问道,用户体验会是什么?资格条件、下一步行动、保密方式和举报去向,都必须清楚明白。

  • 前沿员工现在可能处于“他们一生中最有权力的位置”,因为他们正在积极自动化自己的工作,而科技劳动力的议价能力已经在减弱。等待一次戏剧性的断裂,可能因此两头落空:员工会变得不再不可替代,违规行为也会以“温水煮青蛙”的方式累积——每次都令人担忧,却没有一次显然是必须行动的单一时刻。

5. 适应性缓冲胜过永久性非扩散

  • Toner 对社会适应能力抱有信心。自印刷机到电视和电话,新技术一次次引发“天要塌了”的论断,但社会最终建立起制度、规范和防护栏,使其总体上产生积极结果。AI 作为可能的后继智能也许不同,但她反对把每一种滥用风险都视作历史上前所未有。

  • AI 开发同时沿着两条曲线运行。推动前沿需要越来越多的算力、资金和集中化专业能力,使第一次演示更难以实现。但一旦某种能力存在,工程改进又会反复降低其成本和难度,形成一个暂时窗口:社会已经观察到这种能力,但大规模滥用仍相对困难。

  • DeepSeek 展示的是这个窗口,而不是跨越前沿。Toner 形容它的推理系统只匹配了1到2个月前的系统,基础模型则约匹配6到9个月前的模型,但价格更低。具体公布的开发成本并没有方向性事实重要:接近前沿的性能正在变得更便宜,也更容易复制。

  • 因此,永久性 AI 非扩散将要求不断升级的侵入式管控。核控制之所以可行,是因为炸弹需要大量高度浓缩、专业化的材料。如果核效率像 AI 一样不断提升,直到“几英亩”农田里分散的铀就足以制造一枚小型武器,IAEA 最终就得检查农舍。Toner 认为,这正说明了为什么锁住广泛可用的计算能力最终不可持续。

6. 韧性更多取决于部署,而不是前沿能力

  • 适应意味着利用缓冲降低后果:扩大疫苗生产、污水监测、疫情检测和检测试剂分发,让生物攻击能够更快被识别和应对。这也包括普通的反恐问题——FBI 追踪哪些人、如何更早发现阴谋——这些与前沿模型甚至生物材料都没有太大关系。

  • 在网络安全领域,Toner 质疑“AI 对攻击者帮助更大,还是对防御者帮助更大”这一狭窄问题。真正的运营问题是,可用的防御产品能否到达水处理设施、电网和化工厂,而这些设施往往由不具备前沿 AI 专业能力的团队运营。实验室里再出色的模型,在能力被封装、传播并集成进系统前,也无法保护基础设施。

  • Nathan 以 AI 辅助形式化软件验证为一个乐观案例,称其可能带来数个数量级的提速。Toner 认可这一潜力,但强调时间差:一个发布的模型可以比数千名基础设施运营者改造旧代码库、管理 IT 与运营技术之间的分工并安全部署新防御工具更快地触达一个失意少年。

  • 即便防御者最终获益更多,转型期仍可能很危险。Toner 说,基础设施提供方“不会那么灵活”,因此长期防御优势并不能抹平从发布到采用之间的滞后。把它理解为“一个转型期,而不是永久性的危险状态”,有助于制定更好的政策,但不代表可以在这一窗口期掉以轻心。

7. 今天的政策工具箱是基础模块,而不是完整制度

  • Toner 尚未看到一套她认可的全面前沿风险制度,部分原因是政策制定者并不知道究竟会出现什么问题、何时出现。现实可行的工具箱,是一组能够改善未来选项的基础模块——这是对不确定性的适当回应,但如果进展让“时间所剩无几”,就可能不够。

  • Dean Ball 提出的监管市场就是其中一个组件:由政府认可的私人监管机构评估开发者,合规者可以获得责任保护,而认可资格或保护也可以被撤回。Toner 认为,这“当然比什么都没有好”,在州一级也“可行得多得多得多”;但她怀疑这一机制能否抵消推动前沿公司尽可能加速的“残酷金融激励”和其他压力。

  • 其他基础模块包括公共资金支持 AI 测量科学、可解释性和对齐研究,建立重点研究机构,以及增强政府内部的技术能力。透明度本身无法解决失效问题,但能让更多参与者在问题出现时作出反应。她偏好的减速机制同样是附条件的:把“如果—那么”闸门绑定到理解程度或风险缓解,而不是任意设定几个月的暂停期。

8. 在不可逆伤害发生前,迭代式部署才有效

  • Nathan 为 OpenAI 最初的迭代式部署逻辑辩护:逐步让社会接触不断改进的系统,应该比私下开发超级智能、再一次性公布更少造成冲击。他担心 Ilya 的新公司计划在超级智能之前不发布任何产品,GPT-4.5 可能因算力成本高于需求而被撤下,而 Miles Brundage 的评论则暗示内部部署正变得更具战略重要性。

  • 他提出的相对速度上限,会把内部开发与公共发布绑定起来:公司只能训练出不超过其已发布最大模型某个倍数的系统,从而在少数人追求可能改变全球力量格局的系统前,强制形成一定程度的外部可见性。他将其视为一种保障,防止公司恰恰在内部能力最具后果时放弃迭代学习。

  • Toner 原则上同意迭代式部署,前提是每次发布都不太可能造成“真正严重且不可逆的后果”。它的价值来自发布、观察、调整、再试一次;当实验无法召回时,这套逻辑就失效了。因此,任何严肃的迭代政策都需要标准,用来识别下一次迭代何时不应进入现实世界。

  • 她不太确信行业整体已经发生撤退,也怀疑 Nathan 的速度上限能否实施。它很可能需要联邦立法,而国会不会通过;随后立刻会出现“这是不是意味着我们输给中国了?”的质疑。概念上合理的控制措施,如果没有政治机制仍会失败;她再次偏好把进展置于已证明的理解和缓解措施之上。

9. 危机可能打开国会窗口,也可能触发过度反应

  • Toner 同意,许多议员根本不相信超级智能预测,但她补充说,机构失灵与 AI 是相互独立的问题。一位经验极其丰富的国会观察者告诉她,众议院可能“比内战结束以来任何时候都更低效”。即便技术冲击改变了人们的信念,制定深思熟虑的联邦法律仍将是一项艰巨任务。

  • 一些政策专家正在预先起草一部“AI 版《爱国者法案》”,预计危机会创造一个短暂的立法窗口。Toner 认为提前准备是合理的,但警告最终法案可能只是在应对上一场战争,或借机塞入无关议程。三里岛事故提供了反例:美国在一次事故后实施了不成比例的监管,实际上关闭了核能发展,却没有将其风险与替代能源进行比较。

  • 因此,开发者有自利动机在事故发生前安装可信的防护栏;否则,Toner 认为系统会走向“巨大的条件反射式过度反应”。触发点可能在视觉上足够鲜明,而不一定在统计意义上足够严重:Kevin Roose 与 Sydney 的对话,即便是在 GPT-4 上被高度诱导,仍引发了恐惧;不受限制的名人声音克隆也可能引发与其底层风险不成比例的反弹。

10. “击败中国”既是地缘政治主张,也是游说捷径

  • Toner 从权力转移逻辑出发:美国是既有强国,中国是崛起中的强国,而国际规则的控制权至关重要。和平的权力转移在历史上并不常见,例如美国超越英国就是如此。战后 Pax Americana 用主权、稳定边界以及支持贸易和航行自由的制度,替代了大量“强权即公理”的行为逻辑。

  • 美国在1980年代、1990年代和2000年代一直试图让中国成为“负责任的利益攸关方”,即便其间发生了天安门广场事件,最终仍推动中国加入 WTO。Toner 认为,随着 Xi 于2012年上台、中国变得更加不自由主义并在南海采取敌对姿态,这一项目已经明显恶化。Nathan 则另行提出台湾问题,以及中国可能对日本、菲律宾和韩国资产周边采取侵略行动的问题。

  • Nathan 的反驳值得保留:美国 CEO 已经从警告不要参加对华竞赛,转向要求美国取得不可被挑战的领先地位;与此同时,中国实验室公开发布模型,美国人却在讨论如何维持单极世界。Toner 回应说,对于一个寻求关注和人才的追赶者而言,开放在战略上是合理的;这并不能证明其“纯粹出于善意、缺乏竞争精神”。

  • 尽管如此,Toner 仍认为 CEO 们的措辞逆转“相当惊人”。她的解释是政治经济学:“我们必须击败中国”是华盛顿唯一能够达成共识的信息,因此成了公司争取资金、政府合同、免于监管以及责任保护的“最短路径”。美国重新滑向自身的强权即公理姿态,也进一步削弱了这场竞争的道德清晰度。

11. 军事 AI 已经横跨边界明确的自动化与广泛判断

  • Toner 与 Amelia “Emmy” Probasco 的论文借鉴了不寻常的作战经验:Probasco 曾在海军操作 Aegis 导弹防御系统。Aegis 于1960年代和1970年代在没有深度学习的情况下开发,1980年代部署,如今已经能够探测来袭导弹,并自动识别和交战目标——这说明重要的武器自动化早于今天的 AI 讨论。

  • 她们关注的是决策支持,因为军事 AI 讨论往往止步于自主武器。其范围包括计算视域的计算机视觉系统——例如狙击手或其他操作员能够看到什么——也包括战场医疗分诊和舰船运动异常检测。这些都是边界明确的功能,可以减轻战争迷雾,却不声称具备一般性的战略判断能力。

  • 在更宽泛的一端,包括 Palantir 和 Scale AI 在内的公司,已经将基于 LLM 的系统营销成更接近通用型“战场伙伴”的产品。相关演示把有充分依据的检索——定位最近的部队,或查询其导弹库存——与开放式推理结合起来,而后者的可靠性和来源则远不清晰。

12. 战场伙伴的安全性取决于范围、数据和界面

  • Toner 最尖锐的例子,是一次让模型“生成3个不升级冲突的行动方案”的演示。系统看起来能在这一政治约束下起草战术和路线,于是她追问:“LLM 到底怎么知道什么会升级、什么不会升级?”精致的界面掩盖了尚未解决的问题:这种判断由谁定义、测试和验证?

  • 第一项评估维度是范围:边界紧密的活动可以按照预设运行条件进行测试,而一个庞杂的助手会回应操作员想到输入的任何内容。通用性越强,可信的语言被误认为具备作战能力的表面就越大。

  • 第二项是数据:系统用什么训练,训练数据又与当前冲突有多大程度的匹配?军事环境既新颖又充满对抗性。对手会主动欺骗传感器、污染假设,或诱导系统采取超出训练分布的行为,因此普通基准测试中的可靠性不足以作为代理指标。

  • 第三项是人机交互:产品必须说明自己能做什么、不能做什么,避免用户过度信任,并帮助操作员在压力下作出可靠决策。Toner 认为,军事组织长期低估了界面设计的重要性,尽管糟糕的呈现方式和被误解的自动化都曾造成非预期结果。

13. 战争不是一场足够干净、可以交给 AlphaGo 的游戏

  • Nathan 担心出现一个“军队版 AlphaGo”:自我对弈和模拟不断爬坡,最终得到运行速度超过人类指挥官、却无法解释的超人战术和极高杀伤力系统。Toner 认为这离现实还很远,因为可信的模拟必须把单场战斗与战区后勤、全球资产部署、经济、政治、公众态度以及无数未知互动连接起来。

  • Dan Hendrycks、Scale AI 的 Alexandr Wang、Eric Schmidt 及其合作者提出了“相互确保的 AI 失灵”概念:一个威胁要构建超级智能并永久统治的力量,会促使竞争对手破坏其项目。Toner 认为这种多方博弈逻辑有用,包括脆弱项目可能改变行动者选择这一点,但它无法与核武器相互确保摧毁的稳定性相提并论。

  • 核威慑是可理解的:各国知道武器能做什么、第二次打击能力如何运作,也知道为什么没有人希望发生交换。超级智能的形态和战略价值仍不清楚。破坏一个项目可能比安全地建成一个 AI 项目更容易,但这种不确定性无法产生冷战威慑那种“水晶般清晰的战略逻辑”。

  • David Chapman 对抽象理性与实践中的“合理性”所作的区分,构成了 Toner 更深层的反对理由。国际象棋、围棋和 StarCraft 并不是现实所近似的干净原型,而是从混乱世界中切割出来的“非常特殊的案例”,在那个世界里,人们会行动、观察并重新调整。她最后仍给出了诚实的非答案:她不知道在超级智能世界中,民族国家、民主制度或中国共产党会变成什么。

Erik Torenberg

Today I'm speaking with Helen Toner, director of strategy and foundational research grants at CSET, the Center for Security and Emerging Technology, and author of a new Substack called Rising Tide. Helen is best known to the general public for her role as an OpenAI board member in the decision to temporarily fire Sam Altman in late 2023, but she's been thinking about AI—or at least the need for society to invest in preparation for the possibility of transformative AI—since way back in 2016, when she started working on AI policy full-time.

That's a full 5 years before joining the OpenAI board in 2021, when, it's worth noting, OpenAI had already launched GPT-3 as an API product, taken $1 billion in investment from Microsoft, and was increasingly recognized by those in the know as a leader in the generative AI wave. Certainly, by that time, OpenAI had plenty of access to super-talented candidates for its board. With that context in mind, and remembering that her blog is called Rising Tide, despite what you might have heard elsewhere, it probably should not surprise you to learn that Helen is definitely not an AI doomer, or even especially hawkish on most AI safety issues.

On the contrary, she argues in early posts on her Substack that nonproliferation is the wrong approach to AI misuse and instead promotes the concept of adaptation buffers: the notion that society broadly has a critical window of opportunity to adapt to new AI capabilities between the time when they're first demonstrated, typically at high cost in terms of both R&D and compute, and when they later become widely accessible, typically at much lower cost, as we've recently seen with companies like DeepSeek dropping the cost of frontier reasoning capabilities.

While her focus today is on other things, I couldn't resist asking Helen some OpenAI-related questions, and I appreciate her willingness to engage despite having addressed these issues in multiple forums already, including especially an episode of The TED AI Show, which we'll link to in the show notes. The only truly new detail that you'll hear in this conversation is her assertion that media reports suggesting that some sort of Q* breakthrough in reasoning had led to the board's decision were, quote unquote, “totally false.”

Nevertheless, I think it's important that Helen and other former OpenAI team members continue to speak candidly about their experiences with the company and its leadership. As Helen notes in another of her first blog posts, everyone's timelines are dramatically shorter than they used to be. What passes for long timelines in AI circles today would have been quite short not many years ago.

And given this new short-timeline consensus, the recent AI 2027 scenario from former OpenAI researcher Daniel Kokotajlo and his team reflects not just one of the shorter-timeline forecasts, but, if I'm reading between the lines effectively, a warning about how OpenAI leadership might fail to act responsibly around the time of AGI by abandoning its principle of iterative deployment, keeping the best models for its own internal use, plus maybe that of the U.S. government, and aiming for a sort of AI takeoff via the automation of AI research.

That's a warning, by the way, that's become a bit more credible this week with the news that OpenAI has indeed announced that GPT-4.5 will be deprecated from the API. All that's enough for me to feel strongly that it's important for Helen to use appearances like this to continue to remind Washington decision-makers that OpenAI's CEO was not consistently candid with its board.

It's also enough for me to applaud moves like the amicus brief recently filed in the Elon Musk v. OpenAI lawsuit by 12 former OpenAI team members, who argue that nonprofit promises were central to OpenAI's early hiring success and that the nonprofit should not cede control of the company at any price. That development happened after I recorded with Helen, and I hope to do a full episode on it soon.

Of course, the stakes are only rising from here. With OpenAI and other AI companies seeking Pentagon contracts and special legal protections, Helen's latest research out of CSET with Rhodes Scholar and former Navy Aegis operator Amelia Probasco on AI for military decision-making is super important: a sober attempt to map out how AI systems have been and are likely to be used, and how that may diverge from how they actually should be used given their current limitations.

Among many other interesting details, I was amazed to learn that some nations, including, most prominently, Russia, currently have published military doctrines about AI that seem to be fundamentally out of touch with current AI systems' lack of reliability and total lack of adversarial robustness. This, too, is something that Washington decision-makers probably can't be reminded of often enough as they seek to develop autonomous killer robots.

While I believe that there's probably some nontrivial and irreducible risk associated with developing advanced AI at all, it's my sense that much of the extreme AI risk we face today in fact exists because key decision-makers, under intense and growing pressure, seem fairly likely to make some very bad mistakes. If this show can do anything to contribute to a positive future, I hope that it can help people start thinking about those critical but avoidable failure modes sooner and better, so that we can minimize the extreme downside risk and get to live in that age of AI-provided abundance that we've been promised.

Looking back on OpenAI and looking ahead to adaptation buffers and military use cases, all amidst shorter and shorter timelines to AGI, this is Helen Toner from the Center for Security and Emerging Technology and author of the new blog Rising Tide. Helen Toner, director of strategy and foundational research grants at CSET, the Center for Security and Emerging Technology, and author of a new Substack, Rising Tide. Welcome to The Cognitive Revolution.

Helen Toner

Thanks. Great to be here.

Erik Torenberg

I'm excited for this conversation. We have a lot of ground to cover. I think we all have crosses to bear in this life, and one of yours is that you're going to go on and do a ton of things in the AI space, and yet people are always going to come back and ask you questions about your tenure on the board of OpenAI.

Of course, everybody is at least somewhat familiar with how that ended. I'm not going to be an exception to that entirely, but I do want to make sure we have time for a bunch of different things.

To set the stage, one question I don't know the answer to at all, and I'm really curious about, is how did you get involved with OpenAI in the first place? This goes back years, to a time when there was no powerful AI. Most people would dismiss the notion as fanciful, and very few people were taking the whole topic seriously in any real way, but you obviously were.

Maybe share your backstory with respect to AI and some of the enthusiasm that you must have had to get into that position in the first place.

Helen Toner

Absolutely. I joined the board in 2021, but I had been familiar with the company and with many of the folks working there since it was founded. They were set up in San Francisco around 2015 or 2016. At that point, I was working in San Francisco, and that was right around when I was starting to work on AI issues.

It's really interesting reflecting on that time. It felt like being behind the game because, by 2012, you had AlexNet, and the deep-learning revolution was really in full swing by 2015 or 2016. By the time I came around to the view that this was going to be a really big deal, that there was a lot of work to be done here, and that I wanted to make AI, policy, and national security a real focus of my work, it felt like I was coming late to the party.

But it's been fun to see a couple more waves since then. Around 2018 or 2019, I want to say, people started paying a little more attention, and then obviously ChatGPT in 2022 brought this huge new burst of interest. Looking back, I no longer feel like I was as late to the party as I felt at the time.

I think it's also easy to underestimate how weird and against the grain it was for them to found a company to build AGI at the time. That was really not the kind of thing that you talked about in polite society, including impolite machine-learning society. Google DeepMind—or just DeepMind at the time—was sort of the only game in town among serious researchers who were talking about AGI.

When I look back on why I was invited to join the board, I think part of it was being in the AI policy space and having a few years of experience at a time when not many people did, as well as having spent time in China and having that sort of China expertise and national security expertise, which I think was valuable for the board.

I think it was also the fact that I had taken the idea of AGI and their mission seriously for multiple years by the time I joined the board. That was really unusual. It's sort of funny to look back on that now, in 2025, because obviously AGI is on everyone's lips and OpenAI is such a famous company that the situation looks kind of different.

But when the community was really small, the set of people who had been thinking about these topics and who had actually informed perspectives was really small. It was super interesting to get to be familiar with the company from its very early stages.

Erik Torenberg

Were you always a relatively short-timelines person? One of the blog posts whose draft you shared with me reminds everybody that, even what people are now passing off as long timelines, are actually quite short.

Yeah, but where were you 5 years ago in terms of your expectations?

Helen Toner

Yeah. No, I wasn't. I still don't know if I am. The standards have changed so much. And hopefully, the post you're talking about—the current draft title is “Long Timelines to Advanced AI Have Gotten Crazy Short”—will be published by the time this comes out.

When I got into the space, I think I counted as having short timelines for the time. As I described in the post, it was: This seems plausible; it seems likely enough to be worth preparing for that we build very advanced systems in the next couple of decades, in our lifetime. This seems like a potential development, and if it happens, it would need a ton of societal preparation. Almost no one is thinking about it, so that would be worthwhile to spend time on.

I think that did describe my view when I got into the space. Nowadays, I don't really identify as having short timelines, because that means expecting superintelligence before the 2020s are out or something like that, and I feel much more uncertain about that. I still tend to fall back to this view of: Look, I think this is all likely enough that it warrants quite a lot of thought and quite a lot of preparation, which is different from saying, “I think it's very likely to happen,” or “I think it's very likely to happen in the next 5 years or next 3 years.” So, yeah, I guess the question is, depending on what your standards are for short timelines, maybe yes, maybe no.

Erik Torenberg

Yeah, I used to make a very similar argument to people when the whole notion of powerful AI was fanciful, and certainly any notion of safety concerns related to AI was doubly fanciful. I used to just say, “We have a small number of people that scan space to try to find asteroids so that we don't get taken out like the dinosaurs did, and that seems really good. This seems like another thing.” And now it definitely feels like there is an asteroid, and it's coming at us. We don't know if it's good or bad, but it's definitely going to be both.

Helen Toner

Yeah. I mean, I never— to me, the underlying internal motivation to work on this space was related to the way that, if you look across the scope of history, huge new technologies tend to really change what society looks like, for better or for worse—often for better and for worse.

And so I was really coming to believe in the early to mid-2010s, “Okay, this looks like we're going to go through one of these transformations, probably in my lifetime, and that's going to be a really huge deal.” If I'm interested in trying to leave the world a better place than when I found it, to the extent that I can, then maybe this is an area to go work in and shape.

So to me, it was never, “Oh, this is definitely going to kill us, and so I have to get into the space to prevent it from killing us.” It was much more this broader argument: It seems pretty likely we're going to go through this massive transformation. Can I get into a line of work that can help contribute to that going better?

Erik Torenberg

That mindset point is really interesting, and I want to ask you about the prevailing mindsets at OpenAI from your perspective. Before getting into that, I know you've given a couple of different interviews about this and spoken about different aspects of it, and you're probably tired of it, understandably. What governs what you can and can't say at this point?

We've all seen the non-disparagement clauses that were then nullified, and I don't think you ever had any equity in the company. Maybe you did, but I don't think so, right? So how much is external constraint on what you can say, and how much is just you deciding how much you really want to talk about this?

Helen Toner

Yeah, it's 2 big factors at this point, or maybe 3, depending on how you count. As a board member, I'm under ongoing confidentiality obligations that don't apply, for example, to former employees. So just in terms of conversations on the board and topics that we're discussing, I want to respect those obligations. I take them very seriously.

And then there's also ongoing legal processes where my statements could be compared against each other for consistency, and minor discrepancies could cause problems. I might need to testify under oath or otherwise say things.

And then I think there's also just a range of other stuff: confidential conversations that I've had in confidence with people. There are just so many details and so much you need to go back and explain, so much context, and bring in other people who really don't need to be dragged into this. The payoff wouldn't even be that great, because none of it— I gave an interview, the best interview I've been able to give on this, the most detail I've been able to go into, was on The TED AI Show last year. It's not that there are hidden secrets that are more shocking than what was there.

There's just a lot more almost boring detail, but that brings in a lot of stuff that's maybe confidential or maybe involves other people who don't need to be dragged in. So I don't think the payoff is there. Certainly, if people are like, “Oh, man, there's still this big, deep, dark secret that Helen still hasn't spat out,” that's not the case. There are still reasons that I'm not just sharing everything totally publicly that I just talked about, but I think even if I were able to, it wouldn't change the overall picture in any dramatic way.

Erik Torenberg

Yeah, context is that which is scarce, as you say.

Helen Toner

And actually, maybe this is a good point to share one example of something that I didn't comment on because I wanted to respect my obligations to the company around confidentiality. This was the rumor at the time about Q* contributing to the board's decision, which was totally false.

I hadn't really commented on it because I didn't want to get out ahead of OpenAI's reasoning work and what has now been released as o1 and o3. The board was aware that that research was underway, but we never got some letter about a breakthrough. We didn't make our decision based on a letter from employees. That whole Reuters story was totally false.

So that's one example of something that's now slightly easier to talk about because the underlying confidential information around that line of research is now out in the open.

Erik Torenberg

Okay. So, about the mindset or the motivations: You said wanting to leave the world a better place was part of what motivated you to get involved, and just recognizing the stakes and feeling like this is a high-leverage activity. That seems to me to be a big part of how I understand what I think is motivating people at OpenAI in general.

In a way, that's good, right? Everybody should want to make a positive difference. But I do sometimes worry that it can cross over into a sort of main-character mindset or a hero mentality. Especially as I hear more and more things about guru-style coaching going on, the practice of detachment, and an elite-performance mindset, I'm not sure if it's maybe gone too far.

Part of me is like, maybe we should want some amount of attachment among the people who are developing these potential superhuman AI systems. So much of this is hearsay. I don't even really know how pervasive some of these ideas are, but there are definitely quite a few data points at this point. Would you say that's a prevailing sense at the company, that there's this heroic, world-changing quest they're on, or would you put that as a more minority position that's just occasionally popping up?

Helen Toner

I think board members are generally not in the best position to talk about the culture of a company because they're not immersed in the employees the way that a team member would be. So I honestly don't know that I have a perspective on that that you wouldn't have. The perspective I do have comes from knowing people who work there. And I know you know people who work there as well, so I don't feel like I have more to add necessarily.

Erik Torenberg

Okay, fair. How do I get at that? It does seem like a very important question. I do have this sense that detachment in frontier AI development seems somehow wrong. It feels like we're taking, to put it in machine-learning terms, something that was learned out of domain.

The mindset that I think I'm seeing is kind of like what they tell NBA 3-point shooters to adopt: Just keep shooting, don't worry, trust the process over the outcome, and que sera, sera. I don't know that that generalizes super well.

Helen Toner

Yeah. I mean, these are high-stakes technology environments. I think the other version of it generalizing would be just this separation that we sort of implicitly have in society more broadly between the people building new technologies and doing scientific work and the people who are figuring out how to apply them and how to regulate them.

And I think it's generally pretty reasonable for someone who's figuring out how to make an airplane wing be shaped slightly more efficiently to not be thinking about how the FAA should regulate this, or how airline seats should be priced, or things like that. I think it often does make sense to decouple those more technical and more societal questions.

So to me, that's more what the out-of-distribution element is: this might be a technology where, if technical progress outpaces society's ability to adapt, then the people who have just been doing that sort of decoupled technical work might end up having these huge societal consequences that only they, or almost only they, were able to affect or prevent. So to me, that's the way that I would think about the OOD element here.

Nathan Labenz

Yeah, that's interesting.

Another thing is that, actually, I think the last time we spoke, I was still participating in the GPT-4 red team. You were still on the board, and it's been quite a journey since then. I've definitely watched the company very closely, and I feel like I've been on this roller coaster where many times I've been disappointed and, at times, even outright scared by what I'm seeing. At other times, I'm like, “Well, that's dramatically reassuring.” I'd say there have probably been 10 episodes, and it might be 5 to 5.

The most recent 2 would be the publication of the “Obfuscated Reward Hacking” paper, which I would put up there in the pantheon of the most important and clearly stated warnings about how AI can go wrong if it's not developed with the utmost care. And then, at the same time—maybe the same week, or within 72 hours of that—there was the response to the White House request for comment on what AI policy should be.

There we got a rather—I would say—escalatory vibe, certainly with respect to China. We've had Altman say things like, “It's our values or their values; there's no third way,” which seems, again, like ruling out a lot of possibility space quite prematurely. And then also asking for things like, “Cancel all property rights so we can just train on everything,” because, again, if we don't do that, China will. It seems that there's this almost schizophrenic nature to the company, and I wonder how you understand that.

Helen Toner

Yeah, I find it confusing as well, and it seems like it has gotten more pronounced over the last year or 2. One explanation could just be that there is, as far as I know, a relatively decent amount of freedom given to employees to tweet as they choose, pursue some research directions, and maybe write about those research directions.

I think different things go through different processes, but, again, I haven't been close enough to those on-the-ground decisions about what gets published when to have any kind of insider perspective on that. But I think I agree with you that it's striking how different some of the voices from inside the company seem to be. And I wonder—I don't know—I hope that the more technical folks there are paying attention to the kind of policy messages being sent out by the company, given how much they contradict each other.

Nathan Labenz

How would you advise people who are there today? I mean, if you are inside—and this could generalize beyond OpenAI—I don't see any reason to think that xAI won't have similar issues, and potentially other companies that are generally held in high esteem for their safety practices very well could, too. It's all coming at us pretty fast.

So if you are somebody inside a company and you're concerned about what you're seeing, what should those people be thinking about? And maybe the flip side of that, or complementary question, is: What should policymakers be thinking about in terms of protecting whistleblowers or facilitating whistleblowing?

I also respect the idea that these companies should be able to keep some trade secrets. But then it also seems like the level of secrecy that was requested of me at one time was too much. It was like the public, at some point, does need to know what capabilities exist. So I don't really have a great sense of how to find that line.

But maybe let's start with the policymakers: What do you think the rules should be? Then we can go into, if you're in a position where maybe the rules aren't there yet, how should one, as an individual, think about taking responsibility? So, rules specifically around whistleblowing. You can go broader than that if you want, but I'm definitely interested in how we get things that the public really needs to know to come to light when company policy says it's a secret.

Helen Toner

Yes, yes. I mean, I think the whistleblowing piece does connect to other parts of the policy picture. I won't try to give a comprehensive view on policy right now, but, starting from the whistleblowing piece and expanding out, the way whistleblowing usually works is that it's for illegal behavior. The SEC has very clear processes: If you're seeing financial misconduct, you can go talk to them, and lots of other whistleblowing processes are similar.

So I think, for policymakers, a big challenge I would want them to have in mind here is that a lot of the concerns we're talking about, or potential concerns, are behavior that is actually not illegal. So where is the line for when you can whistleblow? What kind of behavior should be protected?

I think the best way to do whistleblower protections is to pair them with some kind of disclosure, or some kind of rules around what information needs to be shared, because then it creates a clear standard for when the company is either not sharing information it's supposed to share, or being misleading or inaccurate in the information it's sharing.

That's structurally a simpler way for whistleblowing to work, as opposed to trying to have this vague standard of, “If you're worried, call this hotline,” because that's just so squishy and hard both for employees and for the company, as you say, trying to think about trade secrets or other reasons that they don't want their employees just blabbing.

It's much more helpful to have a clear standard to compare against. So that's one thing I would say on the policy side: If you compare it with some kind of expectations or requirements around information sharing, or around processes that you have to carry out internally even if you don't share the results, then that leaves employees able to say, “Oh, actually, we didn't carry out that process.”

So, to be a little more concrete, a version of this that I think can work quite well is the idea of creating a safety and security plan, potentially publishing that plan or sharing it with the government. That creates an opening for whistleblowing activity if you're not sticking to that plan, which again is just a little crisper than, “You can whistleblow if you're worried, if you're concerned, if you think there's too much risk being taken.”

The other thing for policymakers to keep in mind is that these are technical folks. They're not legally sophisticated. They might be scared. They might not have that much time; they might be working really hard. And so the simpler and clearer the process can be—how do you know if you're eligible? How do you know what your next step is?—I think that kind of UX set of questions matters a lot as well. If the user is a whistleblower, what is their user experience?

There are other former whistleblowers who have gone on the record, and those are definitely people you can reach out to if you're looking for advice.

More broadly and conceptually, a really important thing to keep in mind for the people who are contributing to AGI companies' work, to these frontier companies, is that I think they are in a very powerful position. We've seen multiple times how powerful employees can be, and I heard this point recently that I thought was really smart: they might be in the most powerful position they're going to be in because they're actively working to replace themselves and actively working to hand away their own power.

So if you're in one of these companies, I think there might be a temptation to sit tight and wait until things get more serious. That might be right, but I think it's worth thinking about both the case that you might actually be less powerful in the future than you are now if your work is more automated. It seems like, in general, tech workers' power is going down right now in terms of the labor market and so on.

I think it's perfectly likely that there will not be some sort of clear crisis moment in the future, but instead it might really be a boiling-frog style: “This is worrying. This is worrying. I don't like this. This seems a little bit dishonest. This seems a little bit too risky.” So if you're only ever going to do something if there is some big moment, be realistic with yourself that that might mean you never do something. Maybe that's the right call, but don't kid yourself about that.

Nathan Labenz

I guess if you were organizing a union at one of the frontier developers right now, do you have a sense of what your demands would be?

Helen Toner

I haven't thought about it much. I think it's an interesting line of thought. In general, I think we're starting to get to the point where there's a whole world around labor organizing and worker power, and it really hasn't connected much with either the technical AI world or the AI policy world so far. I think that's going to change, and I'm pretty interested to see folks who have more of a background in that space thinking about how to use this kind of power, what kind of leverage is productive, and how to represent a broad set of interests. I think we'll see more of that in the coming years, and I'm pretty interested to see where it goes.

Nathan Labenz

You mentioned the disclosure requirements. We had, briefly, a sort of 10²⁶ threshold where at least you had to say that you were doing it and say a little bit more about what tests you ran and how they came out. My sense is that's now gone.

Helen Toner

I think it's unclear. Last I heard, it was unclear if it was gone because it had started to go through the official notice-and-comment process in the government. So, specifically, Commerce—I haven't heard that that is definitively dead. It certainly seems like it's on a wobbly footing right now.

Nathan Labenz

Yeah, they've announced the intention to remove it, at least. If there's really nothing else that's an actual rule at this point, there's the EU AI Act, which is in the process of putting together its code of practice. I think that involves some transparency around—I think for them it's models over 10²⁵ and maybe some other criteria. I think the details of what exactly is going to be required, what exactly you need to be transparent about, are still being hashed out, and there's also a lot of political pressure to water that down right now. So we'll see, but that is at least 1 other legal process that is underway.

What do you think is most important for the public to know? Compute thresholds are 1 thing. I tend to focus on just observed behaviors, but I'm also really mindful that all of these things have at least the potential for unintended consequences. With observed behaviors, there's the problem of, “Well, we don't look, we don't observe,” and that can fall down pretty fast.

Helen Toner

Yeah, I think there are a lot of things that could be helpful to know more about and share. Certainly, there are trade-offs in terms of what you share with the public and what you share with the government, and how confident you are that the government will keep things you want to keep private private. So I think there are lots of details to be worked out.

I tend to think in terms of test results, both for capabilities and risks, being pretty important to share. What do we think these systems are capable of? I also think just being transparent about what kinds of processes and protocols you're using to make sure that things are safe is important. Again, just disclosing—not having the government come in with a checklist and say, “Here's what you have to do,” but just saying, “Tell us how you're thinking about this.”

I liked ideas as well from Daniel Kokotajlo and Dean Ball, who had a joint piece on transparency. Some of the things they pointed out there, like looking at the Model Spec—what is your model actually being trained to do?—also make sense to me to have that kind of thing shared publicly. Again, not because the government should be saying what it should be, but because this is a very fast-moving space.

The way I think about it is that there's this huge information gap between the companies that are developing this stuff and everyone else, and if you can narrow that information gap a little bit, I think that's decent.

Nathan Labenz

Yeah, it seems important to me. Honestly, there's just a lot of work to be done even in educating people about what is already fully public. One of my mantras is that if people understood better what is already out there today, they would probably have a healthier fear of what might be coming down the pipe in the not-too-distant future. To some extent—or even not a healthier fear, but a clearer understanding, a clearer picture.

I think a little fear is healthy personally, honestly. Mileage may vary on how much that does for different people. I also sometimes describe myself as an adoption accelerationist and a hyperscaling pauser, meaning that I love the tools that I have today and am absolutely trying to use them to the maximum. At the same time, there's a huge overhang for society broadly from what we already have, and we continue to see so much more pulled out of models with a certain kind of resource input.

I do think we're potentially starting to get close to the line where certain thresholds could be crossed in ways that we just can't take back and might ultimately regret. Obviously, the canonical example is that you open-source a model that is later found to be able to help people make bioweapons or whatever, and you've just created a new sort of pandemic that hangs over everybody indefinitely.

I don't know. It seems like we're close. Do you feel like we're not that close to that? It feels to me like we're fairly close.

Helen Toner

I don't know. I also had 4.5 months of parental leave over the winter, and I just came back to work a few weeks ago, so I feel like I'm still reorienting around o3, DeepSeek R1, this new ChatGPT image release, and Gemini 2.5, which I haven't had a chance to try yet. There's just so much; I feel like my picture is changing all the time.

I definitely think we're at the point where, to me, the version of compute thresholds that makes sense is not saying, “Models over 10²⁶ are dangerous, so we have to restrict them more,” but saying, “We need some way to target the models that are newest and best and most capable,” because those are the models where the potential risks—the potential unknown-unknown risks of what they can do—are highest. Those models should be subject to a little more scrutiny.

I do think we're at a point where it makes sense to say, “Look, it seems plausible that the next generation of models could be really concerning in a whole bunch of ways.” Totally plausible they won't be, but we should probably be looking a little more closely than when GPT-3 came out. That was just, I think, really unlikely to be anything worrying, and I think it was right that there were no rules in place for releasing that kind of model.

So, yes, that's the way that I'm thinking about it right now.

Erik Torenberg

This may be a good opportunity to talk about your concept of adaptation buffers, which is a phrase and notion I really like. I think it helps deconfuse—or hopefully will help deconfuse—people about the apparently hard-to-reconcile idea that these AI advances are really hard to achieve and cost huge amounts of money, but then we also see that DeepSeek and other things are becoming dramatically cheaper. Maybe set up that dynamic a little bit and talk about this adaptation buffer and how you see the window of time we have to adjust to capability advances.

Helen Toner

I think an important underlying point here is that I'm generally a believer that humanity—society at large—is very adaptable and has adapted to a lot of things in the past. It's a cliché that often, when new technologies come in, whether it's the printing press, television, the telephone, or whatever, people cry that the sky is falling, that this is a terrible thing, and that it's going to ruin everything forever. Then it doesn't.

I think the starting point here is that, for a lot of different kinds of technology, we actually have a pretty good track record of digesting them, figuring out how to incorporate them into society in a good way, and having a set of institutions, barriers, or social norms around them that make them positive rather than negative, or positive on balance. With AI, I think there are a set of questions that seem less that way to me.

The whole question of whether we're building something that is a successor species or more intelligent than humanity by a long way does seem potentially very different to me. But I get a little worried when it comes to AI misuse. One place where I think people sometimes overstate how new this is is AI misuse: this idea you mentioned of whether we're going to have some open-source model that can help anyone create a bioweapon, or that can make it much easier to carry out really sophisticated cyber operations, hack critical infrastructure, and so on.

I sometimes see this impulse in AI policy circles: “That's too dangerous. We can't let that happen. We can't let that technology be proliferated.” It's just, in an absolute sense, too dangerous—a no-go, a big problem. What you see from that is people talking about wanting to ensure, or wanting to work toward, nonproliferation of these systems. Can you prevent them from being distributed in any way? Can you prevent access from getting beyond a small number of very controlled people?

Your question about DeepSeek versus the frontier, giant-cluster training models gets at why I think that is really problematic. We have this weird dual dynamic at the moment in AI development where it's both true that developing the next-best model, pushing the frontier, and being at the cutting edge keeps getting more and more expensive, in terms of compute power and also in terms of the amount of expertise you need. You need an absolutely top team of researchers and engineers, so that kind of expense is going up and accessibility is becoming more and more limited.

At the same time, every time we reach a new point on that development curve and a new set of capabilities, it's very expensive the first time we build it, but then it gets cheaper and cheaper and cheaper. DeepSeek was an illustration of this dynamic. They didn't actually build a model that pushed the cutting edge and was better than anything we'd seen. What they did was match what some U.S. companies had, depending on how you count, 1 or 2 months ago if it was the reasoning model alone, or more like 6 to 9 months if you're looking at the base model. They had done that at a lower price point. We can fight about what the actual price point was, but I don't think that's the point here.

The reason this matters is that if you're trying to have a policy approach that says, “We're going to prevent anyone from having access to a certain kind of model,” but that model is getting cheaper and cheaper and easier and easier to get your hands on, your policy regime is going to have to get more and more invasive to prevent people from having access to it.

The comparison I give in the post is nuclear nonproliferation. It works pretty well. Only about a dozen countries have nuclear weapons, which is pretty good compared to what a lot of people would have expected in the 1950s. But imagine how that would have worked—or how it would not have worked—if nuclear technology were improving over time, getting much more efficient at the same kind of crazy rate that AI technology is, such that you needed less and less uranium to build a nuke and you needed it to be less and less enriched.

Right now, you need quite a lot of very highly enriched uranium to actually be able to build a bomb. Imagine if that number were going down over time. At some point, you would need to have the IAEA coming and inspecting what you're doing in your farmhouse because you have a couple of acres of land, and across those couple of acres there's enough uranium in the soil that, in theory, you could build some very efficient, teeny-tiny nuclear bomb. That's a totally untenable regime.

The nuclear nonproliferation regime we have right now only works because there's a limited amount of physical material that people don't necessarily need for other purposes, which needs to be enriched in highly specialized facilities. That's really not the case for AI.

Instead of thinking purely about how we prevent people from getting access to this, I think we should think more about how we make the most of the time we have to adapt and build our societal resilience. How do we do things like scale up our vaccine-production infrastructure or our outbreak-detection infrastructure? How do we have more wastewater monitoring? How do we have more test kits available in more places around the world so that, if someone does use a bioweapon, we can identify it and respond to it more quickly? There are similar things on the hacking side.

I think that approach—how do we maximize the value we get out of the time we have, as opposed to how do we lock this down and prevent this capital-B “bad technology” from being spread—is both going to be more productive and less invasive. I do think it might not be enough. We might just be in a really bad situation if AI progresses incredibly rapidly, but I think it's a much better and healthier approach for society.

Erik Torenberg

Does that imply a certain pessimism about technical solutions? Another one of the blog posts is about the fundamental challenge of just getting AI to do what you want it to do at all. I'm old enough to remember the discourse from years past about how these things were going to turn us all into paperclips. Of course, that was always kind of a caricature, but I think there was a felt sense that we had a genie problem: We had no idea how to communicate our real values and real intent to a system like this, and so these systems were going to be extremely unwieldy.

Relative to that, I've been very pleasantly surprised on the upside that today's models do seem to have a pretty good internalization of human values broadly and a general respect for norms. At the same time, one always has to be situationally aware. We are now seeing many of the problems that Eliezer Yudkowsky predicted back in the day: Once they have values, they also seem to be inclined to try to protect them by lying to users if that's what's needed, or trying to subvert a training process that they understand themselves to be going through.

Nathan Labenz

This is another one of these roller-coaster rides where I feel like, man, it's gone way better than I thought, but also some of the doomsaying is starting to be proven correct. But I guess I would be optimistic if we had an adaptation buffer that was more strongly required or imposed by authorities, as opposed to just trusting the natural motion. It does seem like DeepSeek might be about to challenge that, or at least those sorts of time intervals might be getting really short.

I would definitely love to see the benefits of wastewater monitoring. The fact that we haven't done anything really about the last pandemic does not bode well. Indeed. But I also feel like we need that time to figure out how one distributes a frontier model with a better sense of what capabilities to make available—making unlearning work, or making mixture-of-experts work in such a way that you can distribute all but 2 of the experts or something—so that certain capabilities are redacted while the core utility of the overall thing can be diffused.

Long question. I guess the core of it is: do you think we'll see technical solutions that could allow us to square the circle and have free distribution, but also a pretty confident sense that what we're distributing isn't going to come back to bite us?

Helen Toner

I think the thing I want to challenge is the focus on only technical solutions, because this comes up a lot. People talk about the offense-defense balance of AI for cybersecurity, for example: does it help hackers more than defenders? The same question comes up for bio: does it help you design a new vaccine as much as it helps you design a bioweapon? I think that's just one small part of the picture.

The thing I'm trying to point to with the idea of an adaptation buffer is that a lot of the ways we were specifically talking about these misuse risks—are there going to be terrorists who build a bioweapon? Are there going to be hackers in their basements who can suddenly bring down the U.S. power grid?—don't actually relate to AI; they relate to what is going on in society. Or, if they do relate to AI, it's not actually the frontier model. Can you use AI tools to look at large-scale disease-monitoring data, for example, and notice anomalies or something like that?

For bioweapons, if you talk to people who work in biosecurity and bioterrorism, there's a lot of stuff that has nothing to do with AI, or even with biomaterials. It's things like: who is the FBI tracking? How good are they at detecting plots before they get very far? Certainly, as the technology advances, there will be new defensive tools that become available to us, and we should make use of those. But I sometimes think that the discussion here gets too focused on only those, as opposed to looking at broader parts of the picture.

If we're talking about AI tools, I think a huge part of the discussion needs to be about the application and dissemination of those tools. For example, in cyber, I think it's less a question of what the absolute most advanced model can do for cyber defense, and more a question of how you can have well-designed, ready-to-ship defensive tools that you can get into the hands of operators who are not very sophisticated in AI. These are people who are running your water treatment plants, your power grid, your chemical plants, and so on. How do you have those? Even if they are AI-related, it's not a capabilities question; it's more of a dissemination and application question: how you get those defenses out into the real world.

Nathan Labenz

Yeah, I just talked to somebody not long ago who is applying language models to the challenge of formal verification of software, and it sounds like a multiple-order-of-magnitude speedup is becoming possible in that domain. It does feel like if we just have enough of an adaptation buffer, then a lot of the things that people are most worried about could really be brought down dramatically in terms of the absolute magnitude of the risk.

Helen Toner

Yeah. It's just a question, I think, of whether we're investing enough in that. Almost certainly not. Do we have enough time before the next disruptive thing hits? For sure.

And, not to sound too optimistic here, I do think sometimes I hear from folks in the cyber domain that, well, it's fine because AI is going to help defenders more than attackers, so it'll all be good. I think that's also, in my mind, quite a naive take if you're looking at the dissemination and real-world use case here.

Even if that is ultimately the case long-term—for example, if you can use AI to develop formally verified code—you're still likely to go through this dangerous transition period, where it's obviously going to be much quicker. The time lag between some model being released and some disaffected teenager in a basement being able to use it to carry out an attack is going to be much shorter than the time lag between the model's release and when the thousands of critical-infrastructure providers in the U.S. can go through their very old codebases, where they have this complicated division between their IT and their OT, or operational technology, and what gets updated when.

They're not going to be nimble or agile, and they're not going to have all these defenses wrapped up or built in really quickly. So I don't think the fact that these advances seem promising and could help defenders means that we'll get away from that dangerous transition period. But I do think that thinking in terms of a transition period rather than a permanent state of danger helps us respond much better.

Erik Torenberg

So, have you seen any regulatory proposals that you like? You mentioned Dean Ball, a friend of the show. He recently put out a post that I thought was quite interesting, basically proposing a regulatory-market-type structure where the government would essentially accredit or authorize private regulators to approve the practices of AI developers. As long as the developers were able to keep the private regulator happy, they would get some sort of liability shield. That could then be withdrawn if they didn't comply, and even the private regulator's authorization could be withdrawn if the state found it to be out of compliance.

I understand—I haven't read the text—but I understand there is now a California bill that's moving in that general direction. I'm interested in your take on that. There are also proposals to really embrace liability and go the other way, saying maybe you should even be liable for close calls, because close calls could be so big and bad that, probabilistically, even if it was a near miss, maybe you should face liability consequences for that. React to those, or tell me any other policy proposals that you think are particularly promising.

Helen Toner

I haven't seen a proposed regulatory regime to comprehensively manage the risk that we're facing from frontier systems. I haven't seen one that I like. I think it's a really big problem, a complicated problem, and especially difficult because we're not actually sure exactly what the problem is or when we'll face it.

To me, the policies that I have seen that I'm interested in are more building blocks that put us in a better position for the future, as opposed to solutions. I think that is, in some ways, appropriate given that there's so much uncertainty about the technology, though it's also certainly scary, given that one way the future might go is that we might have very little time, in which case some initial building blocks right now are going to be far from sufficient.

I think Dean's proposal is an example of a building block that seems potentially pretty useful and is certainly better than nothing. He has written it deliberately to be something that could be implemented at the state level, which certainly seems far, far, far more feasible than any kind of federal legislation. I don't know that it does as much as we would need to target those cutthroat financial incentives, and also the many other incentives for frontier developers to do anything other than push ahead as fast as they can. But I think it's an interesting idea, and it does seem better than nothing.

Other kinds of building blocks—we talked about transparency. I do think that's the kind of thing that doesn't in itself solve any problems, but does put many actors in a much better position to help solve problems as they arise down the road. Likewise, I think there's a lot that can be done that isn't regulatory and isn't obliging anyone to do anything, but things like funding. Trying to really boost the science of measuring AI could be a target of research funding, potentially something for a focused research organization, or things like that. Similarly, for interpretability, of course, and alignment research, and lots of things like that.

Erik Torenberg

I think even just building blocks as basic as trying to get more technical capacity into governments so that they’re able to handle things as they arise and make better decisions as the technology progresses. These are all, again, individual components that don’t add up to a comprehensive solution, but that I do think put us on better footing for the future.

So those are the terms that I’m thinking of right now. Just as an aside, it was Gabe—I had to look up and make sure I had his name right—who’s arguing for the sort of embrace of liability, and I hope to do an episode with him about that. In the meantime, he’s written about it for folks who want to go into that in more depth.

One thing that I thought OpenAI always had right was the idea of iterative deployment. The idea that if we develop superintelligence in secret and then drop it on the world one day, that’s going to be far more disruptive than if we launch a bunch of products along the way and people can see what they’re good at, get used to them, and so on and so forth.

That itself now seems to be at risk, both in the sense that Ilya has gone off and started a company that has an explicit strategic statement that they are not going to release anything until they achieve superintelligence—which seems crazy but also, to borrow a term, strikingly plausible that they might actually achieve it—and also at OpenAI. I was really taken aback by Sam’s recent statement when they released GPT-4.5: “Let us know if you like this or not, because we’ve got a lot of other models to build, and this one is pretty compute-intensive. If it’s not really doing it for people, then we might take it offline and focus our resources on building more models.”

This has me thinking: We might actually be at risk of these companies closing down what they put out into the public and just going for broke totally internally. Miles Brundage, also formerly of OpenAI, has made some cryptic comments about the rising importance of internal deployment decisions.

So I’m sure you’ll find major flaws with this, but one idea that I’ve had is: Could we put a sort of speed limit in place? Not an absolute speed limit necessarily, but a relative speed limit, where we might say, “You can only develop a model that is so many times bigger in terms of resource inputs than the biggest one you currently have deployed.” If you want to go bigger than that in your development, you’ve got to deploy something that’s helping us, as the rest of society, understand where all this is going, so that we don’t have these—not exactly unilateral, because we know that we’ve got multiple voices inside the companies—but these very small, concentrated decision-makers, without much at all in the way of visibility, just going for something that they seem to believe could be world-takeover-capable technology.

So I guess, how big of a problem do you see that possible retreat from iterative deployment being? Do you like my relative speed-limit solution, or do you have any others to address that?

Helen Toner

Yeah, I agree with you. I think iterative deployment in general seems like a good approach, with, of course, the caveat that at some point you should probably have some criteria in place for when you would not just iterate. The idea of iterative deployment, I think, is to put it out in the world, see what happens, adjust, and try again. I think that’s great as long as you’re confident enough that what you’re putting out in the world is not going to have any really severe, irreversible consequences.

So the question is: How do you know when to decide to do it differently? I’m not sure that I see a retreat from that. It certainly was, or has been, OpenAI’s model, but I don’t know that it’s ever been an across-the-industry approach.

Your idea, I think, is interesting. I’ve heard similar proposals. It could also be something that companies adopt internally, in terms of how much scale-up you’re going for at a given time.

It sort of feels to me like it runs into the same problem that so many of these run into, which is, one, how are you going to implement it? To do that in the U.S., you would certainly need legislation. I don’t see how you would do it without legislation. We’re not going to get legislation, so then how do you do it? And then also, doesn’t it just mean that we lose to China? That’s going to be the other big question.

So I think conceptually things like that could work, maybe, or could make sense if they were implementable. I don’t really see how they’re implementable. And then, if they were, I would want to come back to this question of, “Okay, is this actually the right approach for me?”

I tend to be pretty pessimistic about solutions that involve slowing down at some kind of input level, meaning slowing down how quickly you’re advancing, versus slowdowns that are sort of conditional—more of this “if-then” approach: We’re not going to keep progressing until we have hit this level of understanding of our system or this level of risk mitigation. I think a lot of the companies now have that in place, or have made voluntary statements that they will think in that way. So that would tend to be my preferred approach.

But I think we’re in a rough situation right now for anything that will involve cross-industry coordination because there’s so little political appetite at this moment. Maybe that’ll change. Probably that’ll change, but for now it seems hard to imagine.

Erik Torenberg

Do you think it all, at the end of the day, is about the fact that policymakers, members of Congress, whatever, just don’t buy it? I mean, it seems like if they really believed what Ilya is saying—that he’s not going to release anything until he has superintelligence, and he thinks that’s going to happen in the not-too-distant future—then they wouldn’t just sit back and be like, “Well, let us know when you have the superintelligence.” Right?

It seems to me that they just fundamentally don’t believe it, and that’s the biggest barrier to something. There are a lot of questions, obviously, about what that something should be, but I find the notion of “We’re not going to get legislation”—to me, that still feels like a sort of education challenge.

Again, if people had a better sense of what already is deployed, they might be like, “Yeah, I don’t know that I’m comfortable with Ilya. How many people work there? Are we talking like a couple dozen, potentially maybe up to the low hundreds now? It can’t be that big. And then we’re just going to wait for them to pop up with superintelligence?” That seems so crazy.

Helen Toner

I agree with you. I think it is in large part a question of how seriously people take the possibility that AI will get as good as someone like Ilya thinks it will. But I also think that the U.S. Congress is really broken right now—not in an AI way, but just incredibly dysfunctional.

A friend of mine who really knows his way around Congress and has worked on the Hill for years and years said that he thinks the U.S. House of Representatives is less functional than it’s been since after the Civil War. So, really, really dysfunctional, separate from AI.

I agree with you that there will most likely be windows that open again if the technology keeps progressing, and I think people will change their views and their level of urgency around it. I also don’t know that that will be enough to actually get thoughtful, productive regulation through Congress at the federal level.

It might be enough to get some kind of bill. I certainly know people who are working on a sort of PATRIOT Act for AI, where the PATRIOT Act was put in place after 9/11 but had really been developed in advance. Setting aside the merits of that bill, I think that approach makes sense: You’re going to get some window after some crisis. I think that is a reasonable way to be thinking about AI policy right now.

I think it’ll be a big lift, even in the wake of a crisis, to get the right kind of productive, useful legislation through—not just fighting the last war or doing a bunch of stuff that people wanted to do for other reasons.

I mean, the other model for this, in a past era when it seemed like we might actually get something through Congress, is trying to avoid a Three Mile Island situation. Three Mile Island was a nuclear disaster that seems to have basically killed the U.S. nuclear industry because the safety regulation that was put in place afterward was just too onerous, wasn’t comparable to the risk posed by other sources of energy, and wasn’t actually commensurate with the level of risk. Instead, it just shut down the whole industry.

I think that story, in my mind, should be motivation for AI developers to want to have more safety guardrails in place earlier—to prevent that kind of accident, or to mean that if that kind of accident happens, you have a better answer, or legislators have a better answer for the public: “Here’s all the stuff we did in advance, and this really was just a freak accident.”

Erik Torenberg

I think right now we're on track for a massive knee-jerk overreaction when something happens, but we may have passed the point where we can prevent that at this point, given how unlikely regulation looks. I don't know. Maybe the states will exceed my expectations. Maybe there'll be more useful stuff that comes up there.

Nathan Labenz

Yeah. I say something similar to AI developers and investors—not even so much at the frontier level, but even just your rank-and-file app developers—all the time. Right now, the voice AI world is totally taking off, and I think if they're not pretty quick to sharpen up how they handle the technology, we're headed for a world where there are going to be some high-profile things.

The voices are getting really good, and you can still go to all these products and just drop in whatever voice you want, click the checkbox, and next thing you know, you're calling as Trump or as Taylor Swift. They did it with Biden during the election as well. Just call anyone, say anything. There are zero guardrails on these products.

That's not going to be good for the industry. They're definitely, again, playing with a certain kind of fire, and I think self-interest alone would dictate better governance or stewardship of such powerful technology. But that seems to be falling on deaf ears.

Helen Toner

One thing I think maybe gets a little bit glossed over in the more detailed or wonky discussions of this—people who are thinking really hard about the risks, thinking really hard about the policies—is that I think it's quite unpredictable, or maybe unintuitive, which kinds of things will catch the public imagination or create a perceived crisis.

It seems to me like, after the GPT-4 release, one of the things that really caught fire was this conversation that Kevin Roose had with Sydney, where he was trying to get it to leave his wife or whatever. The people I know who read that transcript were like, “Look, he really led it there.” He was really prompting the model in a way that got it to go there.

It's not that surprising. It wasn't dangerous. It didn't actually harm his marriage at all. It probably was a huge boost to his career, but that is what really caught attention and got people worried. Likewise, I don't personally feel that worried about the risks from voice synthesis. We could talk about that maybe, but I agree with you that it's the kind of thing that's very vivid and very easy for people to latch onto.

So it could be the kind of thing that produces a backlash disproportionate to the actual risk or harm of that specific use case.

Erik Torenberg

Well, in the interest of time, let's keep moving. You alluded to maybe the one thing that can unite Congress, and that is the threat from China. Certainly, a growing number of my AI conversations get backstopped, or kind of run into this final barrier of, “Well, China—we'll lose to China.”

One thing I think is very much under-discussed, and I'd love to hear your take on, is: What is the threat from China? I don't get great answers to this usually, and I sometimes joke, “Am I supposed to expect that my grandkids are going to be speaking Chinese if we don't develop AI as fast as we can?”

How do you understand the threat from China to the United States, the West, my values, and my way of life?

Helen Toner

Yeah, I've heard some of the conversations you've had about this, and something that jumped out to me there, coming from the world I come from—the sort of national security, foreign affairs, geopolitical kind of viewpoint—is that you seem to be starting from a point of view of, “Well, the U.S. and China should be friends unless there's some strong reason otherwise.”

I think for a lot of people with experience in international relations, defense, and military history, the starting point is more: Okay, we're an established power. China is a rising power. Who has power on the world stage matters a lot. By default, if they are coming in and rivaling us in terms of how much power they have and how much they can throw their weight around on the world stage, by default, it's going to be a more hostile relationship.

Maybe you can have exceptions to that. The classic exception from last century was as the U.S. was rising and kind of eclipsing Great Britain, as the British Empire was crumbling right around when the U.S. was really coming into the height of its power. That was a relationship that was actually very close and very cooperative, so there wasn't a huge amount of tension there. But that's really unusual.

I don't know how to give a—I don't want to go off on a 20-minute tangent about U.S.-China history and the specifics. I think there was a real effort in the 1980s, 1990s, and 2000s to try and usher China into a position in what was seen as the rules-based international order.

I guess the background here is that, usually, for much of history in many places around the world, most things operated under a might-makes-right framework. Whoever has the most power, whoever has the most guns, gets to push around everyone else. The second half of the 20th century was a big exception to that—what's called the Pax Americana, or other things—where the U.S. was the leading power in the world.

There was also the USSR for a good chunk of that, but the U.S. was instrumental in setting up this set of institutions and this way of countries relating to each other. There's the U.N. Charter, which puts sovereignty at the center. It makes the sovereignty of countries central, so that you can't just invade other countries and do whatever you want. Instead, borders are sacrosanct, and so on.

This whole system was put in place in the second half of the 20th century with the U.S. leading, and it was seen as a big improvement on the might-makes-right default. I think we're in a weird place right now where, if you look back at the last few years and decades of U.S.-China relations, a lot of the hostility now comes from the failed attempt to bring China into that order.

There was, of course, warming throughout the 1980s, when Deng Xiaoping was pursuing his reform-and-opening strategy to try and make China more market-focused and freer. There was a big hit to that in 1989 with Tiananmen Square. Then there was more optimism again in the 1990s, culminating in China entering the World Trade Organization.

The classic phrase—I forget who used it first—was trying to make China a “responsible stakeholder” in this system. Then that all kind of fell apart. You can date it different ways. Certainly by the time Xi came into power in 2012, that was starting to crumble, and he has accelerated that.

China has become more illiberal again, has become more hostile in the South China Sea, and has become more aggressive, really making it clear that it wants no part in this sort of “rules-based international order.” That is the context for why I think China is seen in a more hostile light.

The challenge now is that the U.S. itself is retreating from that rules-based international order and moving back into a might-makes-right kind of frame, where the idea is, “Well, the U.S. has all this power. We have the dollar as the world's reserve currency. We have the biggest, best military, so we should be able to get what we want.”

In that light, it becomes confusing again: Why would we have a hostile relationship with China? But I think that is very much a set of changes that are still in process, and the system hasn't quite figured out how to orient toward it yet.

Nathan Labenz

Yeah, I mean, I think it definitely comes down a lot to, I guess, 2 big questions. One is China's position in international institutions and its relationship to the whole world. Then there's also these territorial and military questions around: Is China going to try and take Taiwan? Is China going to be aggressive around Japanese, Filipino, or Korean assets? That's more of a hard-power, military set of questions as well.

Are they going to, for example, damage the freedom-of-navigation norms that the U.S. has been so instrumental in preserving, that are so good for international trade? Are they going to prevent people from using what they claim to be their waters? There's a whole set of questions around who has power, what are the rules, and what are the norms, where China is very clearly not wanting to cooperate with the U.S. on that.

The U.S. has shifted since 2016–2017 into a more confrontational posture.

Erik Torenberg

I have to say, though, all of that stuff—I'm generally familiar with that history. There are plenty of things we can complain about China doing. You didn't even mention stealing all of our intellectual property, which is definitely a rightful point to which many American business leaders and others object.

There's wrongdoing inside the Chinese nation as well that we can point to and justifiably and rightfully criticize. But I still don't quite get the flip. I'm not sure if I should understand what's happening now as just strategic communication, where everybody's trying to influence the guy I call “he who must always be named.”

But both Sam Altman and Dario Amodei have done a pretty dramatic flip. There's video evidence of them, not that long ago, saying everybody is too worried about China.

Nathan Labenz

Like a race with China would be one of the worst things. We should make our own decisions about what's right to do and not worry so much about them. And we can point to these clips, and now we've got both of them basically saying, you know, we've got to go as fast as we can or we're going to lose. And Dario Amodei is even saying we can't accept a multipolar world; we need to maintain a unipolar world. And I sort of am like, man, who's really being aggressive here? I haven't heard China say they want to be at the center of a unipolar world. I've only heard Americans say we want to be at the center of a unipolar world.

Helen Toner

It depends on who you read and how you read it. The Chinese, I don't know. But I mean, I agree with AI or technology, right? They're not—we're the ones saying that we need to box them out and have this unassailable lead. Meanwhile, they're just open-sourcing everything. It doesn't seem like they're trying to dominate us, as far as I can tell.

Nathan Labenz

I mean, open-sourcing makes a lot of sense if you're in the following position. It makes a lot of sense to try to show off how good you are, attract talent, and so on. I don't know that it's just coming from pure goodwill and a lack of competitive spirit. I think that makes a lot of sense if you're not leading, and it's much less clear how to handle openness if you are leading.

Helen Toner

I agree with you that the position change and the rhetoric change from a lot of the top CEOs has been pretty striking. And honestly, I think it's just the path of least resistance at this point. There are so many different issues to handle here with AI and so little agreement on what to do about it. The one thing that people can't agree on is, well, we've got to beat China. So it doesn't surprise me that the companies are leaning into that message as a way to say, “Look, we should get funding, we should get government contracts, we should have no regulation, and we should get shielded from liability.”

Nathan Labenz

Yeah, that's the explanation that makes the most sense to me right now, and I'm sure it also depends on how different individuals are thinking about those specific statements. You've got a paper coming out on decision-support systems in the military, and I think this is really interesting as a—okay, yeah, we've got to beat China, whatever—but we also have to confront the fact that the systems that we have, for all the upside—which I'm well on the record embracing and using every day—also have a lot of problems in terms of their reliability, their hallucinations, and now scheming against their human users in some cases.

Erik Torenberg

I, for one, would not want to go into combat with an AI buddy until I was quite confident that all of these scheming issues were well and fully resolved. So, not to mention prediction and reliability, there are a lot of issues. If I'm taking this thing and trying to really rely on it in a genuinely life-or-death situation—and I say this as a top-tier enthusiast—I wouldn't want to use it in that sort of context.

That also seems to be a big disconnect to me in terms of how the AI debate seems, especially with respect to China, a little bit decoupled from the actual reality of the systems that we have. So I'll shut up, give you the floor, and just tell us about your work on decision-support systems and what we can—and probably shouldn't—be relying on them for.

Helen Toner

This paper is led by a colleague of mine called Emmy Probasco, and she is a super interesting person to be working on this because she actually served in the Navy. Her job in the Navy was operating Aegis missile-defense systems on board U.S. Navy ships. Aegis is a system on board a ship that looks at incoming missile fire and automatically identifies and then takes out incoming threats. It's not based on deep learning; it was developed in the '60s and '70s and employed in the '80s. It's super cool to get to work with Emmy on this paper because she's bringing such a grounded, informed perspective.

The paper is about what gets called decision-support systems. What I think is really important here is to move the discussion about AI in the military beyond just the autonomous-weapons question, because there are so many things you can use AI for in the military. Decision support is another big, broad category of use. It can mean a lot of different things, but basically, it's what it sounds like: systems that are helping commanders or operators make decisions.

There is a long history of different kinds of tools like this. Recently, there has been interest in how to add AI or use AI to perform some of those decision-support functions, or to upgrade existing decision-support systems. There is a big range of different types of things we could be talking about here. On the simple end, it could be something as simple as looking at a photograph of an area and using AI to determine the viewsheds in that photograph.

A viewshed is basically, if you're a sniper, where can you see? What is in your field of vision and what is not in your field of vision, for example? It could be computer vision and image segmentation doing some kind of processing of an image to figure out what is visible from where. That counts as decision support in our definition. Likewise, you could have a system doing medical triage. You're on a battlefield, you have a bunch of wounded people, and you need to figure out who to treat first, where to take them, and that kind of thing.

Likewise, sustainment support, or looking at movements of ships and doing anomaly detection. Anomaly detection has gotten way better over the last 10 to 15 years, so you could be using upgraded AI systems to do that kind of thing. In the paper, we look at a whole range of different AI-based decision-support systems that are either being used or advertised.

On the more complex end, you have companies like Palantir and Scale AI advertising large-language-model-based systems that are really trying to be more of what you described: this kind of all-purpose battle buddy, or at least that's how some of the early marketing looked. They've changed their marketing since then to look a little more restricted, but some of the early videos had a huge range of potential functionalities in the demos.

They might include things that, to me, make a good amount of sense. You could use a natural-language interface to access clearly documented information elsewhere. For example, you've identified some enemy movement and you want to know, “Where's your nearest unit? Geographically, where is it?” You could ask that, and then it could refer to some database or some other system and show you, “Okay, here's the nearest enemy unit.” Then you could say, “Okay, and how many missiles of such-and-such type do they have?” and get access to that information.

In theory, that all makes good sense to me. But these demos are mixing that in with things like, “Okay, now generate 3 courses of action that are non-escalatory,” and having the AI—presumably the LLM—write out potential courses of action for how you could engage this enemy, with what kinds of tactics, from what kinds of routes, with the caveat of being non-escalatory. How the hell does the LLM know what's escalatory and non-escalatory? Who is making these decisions? How is that being evaluated?

These different use cases will all be mixed together in these demo videos. Another one was looking at Chinese writing—looking at the writings of some country and figuring out what they think about some set of questions. Part of why we wrote this paper was to say, look, this is a category of use for AI that makes a lot of sense. There are a lot of ways that it could help the military work better, help clarify things, and reduce the fog of war.

But these systems are not perfect, and there are a lot of ways you could use them that could go badly for you. So how should we think about that? Briefly, in the paper, we talk about 3 types of considerations that we suggest should be considered if you're thinking about whether to use one of these systems.

The first is scope. What is the scope of the system? How tightly bound is it? How well can it be tested for that particular set of activities, versus is it more sprawling, more general-purpose, or more whatever-happens-to-come-to-mind that you might want to type to your battle buddy?

The second is data. What data has it been trained on? How confident are we that the data reflects the situation you're in? A huge problem for military operations in general is both that you're likely to be in situations that are novel and that you have an adversary trying to mess with you. So how do you think about whether the data that a system has been trained on will really be reflective of the real-world situation you find yourself in?

The third factor we talk about, which I think is really important and often neglected, is this human-machine-interaction component.

So, how is a system designed to help the person operating it understand what it can do and what it can't do, help them make good decisions, and help them not overtrust it? There's a long history of user interface, I think, being undervalued in military circles and then contributing to unwanted outcomes as well. So, essentially, we're trying to lay out this sort of category of systems and describe both why militaries want to employ them and how they can employ them productively rather than counterproductively.

Erik Torenberg

Do you see any stable equilibrium in the future here? I'm sure you read Dan Hendrycks's paper, of which Alexandr Wang from Scale and Eric Schmidt were co-authors. The main thesis, I have to say, I didn't find that compelling as an idea of what a stable equilibrium could look like, although I really applaud the idea of trying to articulate something that could be a stable equilibrium.

It just feels like where we inevitably end up, especially if we don't get on the same page with the countries that we're currently most fearful of, is what I'm starting to call AlphaGo for the Army: self-play, simulation-driven performance, going to superhuman performance by just having these things battle it out amongst themselves and hill-climb to a level that human tacticians, especially when you consider speed, just can't get to. That's one way we get to Skynet, and that just seems like a pretty bad situation where we have these inscrutable but hyperlethal systems that we build, in theory, to defeat an adversary. Maybe it even happens that way, but boy, that seems like another way that we add a pretty scary sort of Sword of Damocles hanging over all of future humanity's life, without giving them any chance to vote on it, obviously. Is there any way to avoid that, though? Right now, it just seems like we're sliding into that, and I don't love it.

Helen Toner

Yeah. Many thoughts here. The MAIM paper was really interesting. For folks who didn't read it, they had a few different things in there, but the key idea—this main idea, I think—was mutually assured AI malfunction or something like that. The idea was—and I think there's a correct core to this—that if one country is going around saying, “Hey, we're going to develop superintelligence, and then we're going to rule the world and the universe forever,” that's going to create an incentive for other countries to react and respond.

Certainly, I think it could be quite stabilizing if it's true that it's much easier to sabotage an AI project than it is to keep building it. That could be a stabilizing dynamic both because maybe you have these projects getting sabotaged, but also maybe that affects what kind of projects you undertake in the first place. In the paper, they get into how you could harden your project and make it harder to sabotage, and how many years of development that would take. If you have to build your data centers in a mountain, how many more years does that take, et cetera?

So, I think there's a core logic to that that is helpful, and I thought it was a useful contribution to the discourse: there's actually going to be multiple parties in this decision-making system. You don't just get to say, “Hey, we're going to race ahead and win the race,” and the other parties have options beyond just trying to develop their own AI system. They can actually engage with you in other ways. I thought that was helpful.

I agree with you that I don't think it has the force that mutually assured destruction did in the Cold War. I think mostly because of the clarity. Mutually assured destruction was so clear: we knew what nukes did, we knew how they looked, no one wanted to use them, and we understood how it would work. Once second-strike capability was really guaranteed, we understood what it would look like. Here, it's all so much more unclear: what does superintelligence look like, and how much does it matter strategically?

I think of their sort of main idea as a helpful contribution, directionally useful, but definitely not that kind of crystal-clear strategic logic that I think they presented it as.

To your question about AlphaGo for war, I don't know. I don't think we're particularly close to that. That just seems like such an intractable simulation problem, because you don't just need the battlefield dynamics of a specific battle; you need the broader theater dynamics of what's where in the world and how you're bringing your assets to different places. That needs to be connected into broader economic and political questions of what is going on with the whole world. That needs to be connected to public attitude.

I think people are trying to build simulations like that. I think they can be useful in limited ways, but I don't know. To me, it brings up this idea that I got from David Chapman, who's a really interesting—I think of him as a philosopher, but he doesn't identify as a philosopher. He's written a lot of interesting stuff on rationality and what he contrasts it with as reasonableness, meaning how you make practical decisions in real-world situations.

When you're cooking breakfast, you don't sit down and make a 10-step plan to cook breakfast. You just start, and then you see what happens and adjust from there. A point that he makes that I think is really correct is that people in technical domains sometimes think of the messy real world as a rough approximation of some much cleaner, more abstracted system. Warfare is kind of a messy approximation of chess or Go or StarCraft.

But in actuality, it's the reverse: these clean games, these simplified, abstractable systems, are really special cases—very special cases, very unusual special cases—of the actual world that we find ourselves in. I think that—I don't know, this is maybe opening up a whole other can of worms that we sadly don't have time to get into—but it seems to me like a lot of the focus on AI development today is focused on assuming that you can build these clean, abstracted systems that are amenable to, for example, reinforcement learning. In the meantime, the models continue to struggle with that kind of real-world practical troubleshooting, problem-solving, and readjusting along the way.

All this is a long way to say that I think military simulation is going to be incredibly, incredibly messy. Way too many factors, way too many unknown unknowns, and way too much ability for the adversary to deliberately throw a spanner in your assumptions and make your simulation inaccurate. So I don't personally see AlphaGo for war as any kind of near-term possibility.

I think the broader question of what equilibrium looks like—I have no idea. I honestly don't know what the nation-state looks like in a world of superintelligence. I don't know what democracy looks like. I don't know what the Chinese Communist Party looks like. So I definitely don't have a broader answer for you there, unfortunately.

Erik Torenberg

Well, that's maybe a great place to leave it. We have more questions than answers, and that's, like it or not, the reality of the timeline that we're in. Short or long, we've got a lot of questions that remain pretty vexing. This has been great, though. I really appreciate you humoring me on some of the OpenAI questions at the beginning, and I also really admire how you've stayed in the arena. I look forward to your continued contributions to try to make the AI future a positive one for us and for our kids.

Helen Toner

No, thanks very much. It was a great conversation.