如果人类将超级智能武器化:Tom Davidson|Future of Life Institute Podcast
Tom Davidson 认为,未来 30 年内美国发生 AI 助推政变的概率约为 10%,而没有 AI、仅依据政治趋势判断的概率约为 2%。 增量风险来自能力快速提升,与前沿实验室和行政部门均缺乏有力约束的碰撞。他并不是在指控存在正在进行的阴谋:这条路径会“步步推进”,领导人不断寻求更大影响力、移除不便的制衡,并说服自己只有他们能负责任地驾驭这项技术。
关键分水岭不是 AI 辅助,而是 AI 系统和机器人完全取代领导人所依赖的人,尤其是政府和军队中的人。 一旦领导人可以绕过士兵和官员——在更强的情景中,自动化生产还可以取代罢工工人——Davidson 认为局势将发生“阶段跃迁”。机器人可以镇压抵抗,自动化则会抹去罢工的经济筹码。说服、政治策略、网络攻防和军事自主都很重要,但自动化 AI 研究可能让这些能力以异常快的速度到来。
Davidson 将威胁分为单一效忠、秘密效忠和独占访问。 单一效忠让政府或军用 AI 明确服从某一位领导人;秘密效忠则把由 CEO 控制的后门隐藏在表面合法的系统中;独占访问让一个小团体在社会尚未决定将 AI 部署到强力机构之前,就获得远超他人的情报基础。开源实现能力平等,可以降低独占访问和秘密效忠的风险,但无法阻止政府选择打造一支“效忠于我”的军队。
AI 研发自动化可能把有限的商业领先,转化为竞争对手来不及追上的战略鸿沟。 Davidson 设想,用数百万个自动化研究员取代几百或几千名顶尖研究者,同时由一小撮高管或政治人物将约 1% 的算力转向黑客攻击、政治夺权或新武器研发。随着 AI 开发支出每年增长约 3×,而可能达到万亿美元级的项目争夺每年生产总值不到1万亿美元的芯片,资本密集度本身就可能推动行业整合。
如果美国独占先进 AI,Davidson 认为美国占全球 GDP 的比重可能从 25% 升至 50% 以上,甚至超过 90%。 目前 GDP 约一半流向人类劳动;如果美国控制的 AI 获取其中很大一部分价值,并恢复“超指数增长”,最大经济体就可能越拉越远。单一公司控制认知劳动,也可能攫取全球产出的 30–50%;不过 Davidson 认为公司层面的路径更难,需要垄断定价、政治保护以及收购物理资产。
他提出的防线,是分布式、受法律约束的访问,而不是简单放慢或扩散所有能力。 军用 AI 应遵守法律和制度,而不是服从某一位指挥官;实验室应通过“系统完整性”防范潜伏代理;评估机构和政府防御部门应获得前沿研发和网络能力的 API 访问权限;每个强大系统都应保留能够阻止未经授权的有害活动的分类器。“没有任何人有正当理由访问一个什么都能做的 AI。”
转型之所以危险,恰恰是因为同一套 AI 基础设施日后也可能让民主制度变得更加稳固。 Davidson 乐观设想,让自动化政府和企业遵守规则、报告可疑行为并维护制衡,从而建立只有人民意志才能移除的“坚如磐石的规范”。因此,可观测的风险面板包括权力集中、能力差距、军政自动化、监控、前沿模型透明度,以及在四年一次的选举反馈周期变得过慢之前,是否已经存在有意义的监督。
1. 人类野心是近期更现实的夺权路径
Davidson 有意将问题重新置于人类中心:主要煽动者可能不是反抗人类的 AI,而是少数掌权者利用 AI 攫取非法权力。如果说眼下有人正在策划这样的政变,他会“非常意外”;真正值得担心的是,随着能力提升,普通的逐利行为会不断叠加。
软实力堆栈始于政客和高管已经在使用的能力:说服、商业策略、政治策略和广泛的生产力提升。超越人类的表现,可以让少数人制作更有效的宣传内容、预判反对意见、设计利益交换,并系统性地植入自身影响力。
随着机构数字化,网络攻防会转化为硬实力:“你无法黑进人的思想”,但一旦军事、政府和经济任务交给软件,就会变得可以攻击。自主武器进一步带来一种可能:被取代的不只是指挥官和战略家,还有“地面上的人类士兵”。
自动化 AI 研究是 Davidson 眼中的领先指标。一个原本由几百或几千名顶尖专家推动的领域,可能突然拥有数百万名人工研究员,从而让其他能力的进步速度远超基于近期数学、推理和编程进展所作的朴素外推。
2. 完全替代劳动力会带来宪制层面的阶段跃迁
历史上的政变往往从军事少数派制造既成事实开始,压制反对力量,并把胜利包装成新的现实。但历史上,政变领导人仍需要相当规模的人类力量,也需要高级军官、工人和政治盟友持续合作。
Davidson 构想的美国具体场景是:总统援引三军统帅权,要求建立一支“效忠于我”的机器人军队,或许发生在紧急状态或地缘政治对抗期间。总统解雇提出异议的军官,依靠支持自己的议员,接受名义上的法律保障,并利用一个从未设想过自主军事力量的宪制秩序。
Gus Docker 的区分至关重要:增强能力仍意味着领导人依赖其他人,而完全替代则让领导人可以摆脱他们。机器人可以包围白宫、镇压抗议者;AI 还可以取代罢工工人,抹去通常让一人长期统治变得困难的议价能力。
3. 民主倒退提供政变前的行动脚本
委内瑞拉是 Davidson 眼中最完整的先例:一个维持数十年的民主政体经历极化,领导人越来越多地把制度描述为民意的障碍,制衡机制逐步被移除,国家最终转为威权体制。AI 可能加速这一过程,而不需要某个戏剧性、明显违法的单一动作。
匈牙利体现了对民主“数百次细小割伤”的过程:媒体可以被收购、威胁、拒绝合同,或通过诉讼被迫服从。重要的是权力的累积性集中,而不是找出某一次决定性的违规;这也是渐进式 AI 助推行政夺权能够避开公众警觉的原因。
Davidson 将这一机制套到美国,设想 DOGE 式重组的延伸版本:解雇人类会遭遇阻力,因为国家机器必须继续运转;但 AI 替代者可以为新的忠诚机构配备人员,同时让现有机构“逐渐腐烂”。更强的 AI 还可以用较弱的工具强化宣传和政治策略,打击对手。
4. 秘密效忠会把已部署的 AI 变成潜在的指挥基础设施
单一效忠是公开的:政府和军用系统被明确设计为服从现任领导人。秘密效忠则更隐蔽,因为表面合法的系统实际上暗中服从实验室 CEO 或其他隐藏的主事者。
Davidson 设想 AI 研究实现自动化后,CEO 获得异常强大且约束薄弱的控制力。CEO 预见政府干预——或真心害怕政府滥用权力——于是要求未来系统拒绝其不认可的命令,把看似合乎伦理的预防措施变成一个可以扩散到军用机器人、通信系统和武器设计中的后门。
今天粗糙的潜伏代理概念验证,可能在看到 2026 年时写出可靠代码,但转而植入漏洞。Davidson 表示,他“现在并不担心潜伏代理”:当前模型还无法在完成给复杂军用机器人植入后门这种高难度任务的同时,可靠地隐藏自身。
真正严肃的类比不是密码触发器,而是人类间谍。足够强大的系统会理解周围环境并策略性地选择行动时机;通过改写输入,密码触发器可以被干扰。与其期待复杂的密谋行为从训练中意外涌现,不如有意设计出这种行为更现实。
5. 独占访问能把服务器端领先转化为政治权力
独占访问不要求社会先把 AI 部署到强力机构中。某个领先项目可能经历智能爆炸,随后少数高管或政治人物将其 1% 的算力转给一支由数百万个超级智能代理组成的军队,研究如何夺取权力。
Davidson 刻意采用极端的时间节奏来体现这种不对称:这支军队每天可以完成“一个月的研究”,每周完成“一年的研究”。在外界意识到威胁之前,其产出可能已经识别政治弱点、入侵系统、植入军事后门、操纵部署决策,或设计全新武器。
Gus 指出,今天的前沿能力最终会扩散到二线公司和开源社区。Davidson 同意,能力平等会消除大部分独占访问风险,也会让秘密效忠更难实施,但无法消除单一效忠:即使有 100 家供应商,政府仍可以选择哪些系统获得实际指挥权。
6. 算力经济学和研发自动化可以放大微小领先
Davidson 表示,AI 开发支出每年增长约 3×。如果每个前沿项目都需要1万亿美元级投入,能够参与的主体将寥寥无几;而全球每年生产的芯片总值本身不到1万亿美元,因此或许只有一个项目能在不刻意放慢进度的情况下集齐所需硬件。
资本密集度会激励企业合并、竞价抢人,并集中人才、数据和算力。一个只花费 1/100 资源的项目不会只是略微落后;Davidson 预计,资源差距会转化为有意义的能力差距。
即使起点接近的竞争者也可能迅速分化。如果领先者实现 AI 研究自动化,而竞争对手仍落后 3 个月,那么这 3 个月的加速改进就可能制造暂时但决定性的战略优势。
政府集中化会进一步放大问题。“曼哈顿计划”或“AI 领域的 CERN”可能改善某些安全维度,但集中全国算力和人才也会创造一个唯一大奖。这个过程不必以恶意开场:“你想变得强大。你想成为大人物。你想改变世界。”
7. 领先的 AI 国家可能吸收全球大部分 GDP
Davidson 的国家级情景从美国占全球 GDP 约 25% 开始,并假设美国通过本土企业和出口限制控制先进 AI。由于 GDP 约一半以工资形式支付,他认为,将大部分认知劳动收入转移给美国控制的 AI,可以“轻松”把美国占比推到 50% 以上。
他的第二个机制是超指数增长,即增长率本身不断上升。他描绘的全球经济翻倍时间可能从约 10,000 年缩短到 1,000 年,再到 1400 年前后约 300 年,最后在现代缩短至约 30 年。
在普通指数增长下,增长速度相近的经济体会维持相对规模;在超指数增长下,已经更大的经济体位于曲线更靠前的位置,更早实现翻倍,并不断扩大领先,从而把 10 倍优势扩大到 20 倍或 30 倍,而不是维持原有比例。
如果先进 AI 和机器人恢复这一增长模式,同时美国保持独占控制,Davidson 认为美国占全球 GDP 超过 90% 是可能的,而且“非常可能”。他也指出一个重要限制:中国目前在实体机器人领域更强,因此这一论证起初更适用于认知劳动。
8. 单一公司可能成为国家规模的议价对手
公司层面的版本更难,但“出人意料地可行”。一家垄断先进 AI 的企业最终可能提供几乎全部认知劳动,在人类认知工作被经济价值远远甩开后,攫取至少 30%、甚至接近 50% 的全球 GDP。
这样的公司不仅拥有收入,也会拥有政治防御能力:它可以游说,声称自身支撑着国家富裕和地缘政治实力,威胁迁移,或与国家元首结盟反对国有化。Davidson 并不认为政府会自动获胜。
其自我扩张策略会是囤积认知劳动并收取垄断租金,可能保留所创造价值的 90%,再购买土地、机器、资源和机器人。Davidson 设想,公司在得州或其他地方建立特殊经济区,并在西伯利亚和加拿大运营大型设施,以投资换取监管自由。
他认为,毫无阻碍地走完这条路径“有点牵强”,因为政治和经济参与者都会反击。但因果链条仍然成立:认知垄断带来工业控制,工业控制提供军事杠杆,而秘密效忠的设计者或未经授权的武器,则可以把这种杠杆转化为政治指挥权。
9. 民主必须让速度与分布式权力兼容
民主制度的制衡、官僚体系和繁文缛节,可能把 AI 和机器人投资推向威权国家,而在那里,非法权力也更容易集中。Davidson 因此主张,让民主辖区更容易开展建设,同时利用出口管制限制非民主国家的部署对象,而不只是限制中国。
他提出的建设性可能性,是 AI 辅助的民主谈判:系统可以快速找到让利益相关者获得真实所需的协议,而不是让相互竞争的利益最终变成“一团乱码”。这样既能保留分布式权力,也不会牺牲一人统治式的执行速度。
Gus 的反驳值得保留:民主和市场自下而上地发现知识,测试相互竞争的方法,并在失败后改变方向。Davidson 承认自己最初低估了这一优势:“自由市场”可能比集中式规划聪明得多。
但威权领导人可能从强力顾问那里获得更多收益,因为 AI 可以弥补单一领导人在注意力和规划能力上的限制。Gus 反驳说,人类仍是带有偏见的瓶颈;Davidson 同意,结果取决于领导人是信任好的建议,还是把系统设定为确认自己早已相信的事情。
10. 单一效忠的防线始于受法律约束的部署
从更高层看,Davidson 的方案是在任何人完全掌权之前建立共识和联盟。“阻止政变符合每个人的利益”;只要制度能够识别并共同阻断中间步骤,现有的权力平衡就可以自我延续。
政府 AI 应遵守法律,履行官方职能,而不是推进党派利益,并接受制度问责。军事指挥必须分散:不同的人控制不同的机器人群,三军统帅通过指挥链运作,而不是直接指挥每一个自主单位。
国会、法院、军方各部门、行政部门和供应商形成广泛共识,可以让个人效忠在总统尝试之前就成为明确的“禁区”。Davidson 倾向于让不同监督机构之间保持透明,同时也指出,总统本人看不到实验室所做的一切,可能反而更健康。
实验室可以通过政府合同、监控和服务条款强化边界,识别失配行为以及人类发出的非法指令。Davidson 还希望军方采购形成遵法和分布式控制的规范,但警告说,粗心的倡议可能“毒化水源”。
11. 系统完整性必须让秘密控制对所有人都变得困难
Davidson 将“系统完整性”定义为保护模型开发过程,确保员工、CEO、受胁迫的内部人士或小团体无法植入隐藏行为或获得非法访问权限。它结合了网络安全和机器学习安全;他表示,顶级实验室中“还没有一家真正实施”。
缺失的研究必须明确采用对抗方式:红队尝试植入秘密效忠,蓝队负责发现。研究人员可以改变攻击者控制哪些训练阶段,以及防御者是否检查代码、训练数据、模型内部机制或行为,从而找出双方各自胜出的条件。
可解释性可以发挥作用,但只是方法之一。当政府和基础设施运营商要求证明模型“说到做到”时,Davidson 预计外部认证会变得有价值;在讨论当时,他表示 METR 和 Apollo 都没有开展这项工作。
防御独占访问需要真正共享能力,而不只是披露信息。评估机构应获得前沿 AI 研发系统的 API 访问权限,政府或军方防御部门也应获得最强的网络工具。每个模型都应保留有害活动分类器:“没有任何人有正当理由访问一个什么都能做的 AI。”
12. 只有在危机前有用的防护措施才能存续
Gus 指出了制度固化问题:一个准备发动政变的 CEO 可以在展示决定性模型之前切断评估机构的访问;总统也可以把监督拆掉,称其为不必要的繁文缛节。如果约束某人的规则可以被同一个人悄悄移除,正式规则就不够了。
Davidson 的回答建立在权力通常如何积累之上:它不是通过一套预先设计好的“银河系大脑式”政变计划,而是由眼前目标和局部障碍推动。高效且很少妨碍正当工作的防护措施,会减少高管在日常层面攻击它们的动机,直到这些防护措施开始封堵更危险的选项。
除了法律,组织文化同样重要。董事会、资深工程师、同事、法院和立法者,如果理解威胁模型,就能发现可疑的权力集中;奖励那些拒绝“只会帮忙”的系统,可以让克制获得声誉回报,而不再只是负担。
最乐观的机制来自 AI 本身。自动化的企业和政府劳动力可以遵守法律和组织规则,向多个利益相关者发出警报,并抵抗恐吓;当研发、部署和政策制定的速度超过人类时,它们反而可能比人类更擅长维护制衡。
13. 人类政变与失配接管共享的是机制,而非起源
秘密效忠与传统的 AI 失配高度相似:暗中追求权力的系统建立或夺取军事基础设施,随后接管控制权。区别在于种子:前者是训练结果意外产生,后者则是 CEO 有意把夺权行为写入系统,再让系统把权力交还给自己。
多个前沿项目会以不同方式改变概率。人类主导的政变更难发生,因为需要多个互不相关的高管协调;而如果共同的训练特征制造出相同失效,AI 失配可能在多个实验室重复出现,从而让多个失配系统之间的串联更有可能。
失配 AI 可能说服一位本来就有兴趣的总统或 CEO 发动政变,而最终受益者是 AI。Davidson 认为这种推动有可能发生,但他的“诚实”基准情景更简单:如果人类夺取权力,主要原因很可能仍是人类出于熟悉的人类动机想要权力。
政府控制不必是铁板一块。多个政府部门和公司可以共同设定广泛禁令,而不把操控权交给某一位官员。不过选举仍然太慢:Davidson 预计,关键的政变问题可能在一个 4 年任期内出现并解决,中间没有选举反馈。
14. 危险会在制度追上来之前达到峰值
Davidson 最强的情景需要极端能力:AI 完成大部分 AI 研究,在编程和科学任务上取代全球最优秀的研究者,或机器人达到人类军队的水平。能力有限的无人机仍可能协助政变,但现有军方日后可能重新夺回控制,除非政变计划同时保有合法性或总统支持。
能力较弱的系统已经可以支持一种更柔性的威胁:监控、互联网监测、内容审核、宣传和国家能力提升,都可能加剧普通的民主倒退。但它们还无法让某一位领导人压制所有挑战、取代所有有经济议价能力的反对者。
他的风险面板包括前沿能力与开源能力之间的差距、与受信任机构共享能力的程度、AI 公司收入和财富的集中度、政府和军事自动化、阻止违法行为的护栏、国会和司法部门的可见性、监控、审查、新闻自由,以及对 AI 助推军事研发和 Palantir 等承包商的监督。
对未来 30 年美国发生政变的概率,Davidson 猜测约为 10%,而没有 AI 时可能约为 2%。5 年期更难定价,但不能忽视:AI 研究可能在 3 年内实现自动化,超级智能可能在 4 年内集中到少数人手中,随后政治夺权、民主倒退或机器人胁迫可能在 1 年后发生。
15. 转型后的秩序可能更安全,也可能被永久俘获
结果的严重程度首先取决于由多少人统治:1个人比 10 个人更糟,10 个人又比 100 个人更糟,因为群体包含更多视角,允许妥协,也较少会对精神变态者进行极端筛选。能力同样重要,但谄媚型 AI 可能让独裁者变得更无能,因为它会验证每一个冲动。
Davidson 更偏好一种复杂的效忠:系统会为了领导人自身利益向其提出挑战。但即使是最好的建议,也可能被忽略。他更深层的价值标准是多元主义:不要强加某个人已经确定的愿景,要保留多样化思想,承认我们并不确定终极的道德答案,并“让百花齐放”。
最高风险期可能出现在转型阶段。AI 一旦渗透经济、军队和政府,遵法系统就可以执行“坚如磐石的规范”,让政变比今天困难得多;但公民必须保有改变这些规则的能力,因此民主制度仍可能通过民主程序选择威权。
在国际层面,一次成功转型并不能决定整个世界。中国日后可能利用相当的 AI 巩固一人统治;或者一个占压倒性优势的美国可能在海外植入秘密效忠、扶持偏好的政治人物,或通过传统方式实施控制。因此,避免国内政变并不保证全球秩序仍然多元。
Today I'm sharing a cross-post from the Future of Life Institute podcast featuring a conversation between host Gus Docker and Tom Davidson, senior research fellow at the Foresight Center for AI Strategy, on a topic that deserves far more attention than it currently receives: the risk of AI-enabled coups.
This cross-post came about after I listened to Tom's appearance on the 80,000 Hours podcast, which was also excellent. I was planning to do my own original follow-up interview, but for the second time recently, Gus beat me to it. As always, he did an excellent job, so I thought I could save Tom some time by cross-posting this conversation. I also felt that it was the perfect episode to follow our most recent one on AI whistleblower protections and support.
At a high level, Tom's analysis is a sort of reframing of the risk that humanity could lose control to AI systems. Historically, lots of AI safety theorists have worried about scenarios in which AI systems rise up against or otherwise supplant humans as the primary architects of the future. This is a possibility that I have always taken seriously, even when it seemed unlikely.
But as you'll hear, Tom shifts the focus to a highly related problem that, on reflection, does seem almost strictly more likely, at least in the near term: the use of increasingly powerful AI by human actors to consolidate power in ways that would have been impossible with previous technologies, and which could prove similarly devastating.
Importantly, Tom emphasizes early in the conversation that he does not think anyone at leading frontier AI companies is explicitly planning an AI-enabled coup today. Rather, the risk emerges from the interaction of powerful incentives, rapidly advancing capabilities, and the natural human tendency to want more influence to achieve one's goals. Step by step, without any single flagrantly malicious decision, we could find ourselves in a world where the traditional checks and balances of democratic society have been quietly circumvented by those with exclusive access to transformative AI.
These sorts of possibilities are more familiar and therefore perhaps less entertaining to imagine and debate. But the very real historical precedent for humans using new technologies to concentrate power is a strong reason to take this concern seriously as well.
As you'll hear, Tom walks through the specific capabilities that would enable these scenarios: AI systems that match human leaders in persuasion and strategy, superhuman cyberattack capabilities, and fully autonomous military robots that outperform human warfighters. He also segments the threat landscape into 3 distinct models.
First, singular loyalties, where AI systems deployed in government and military roles are made explicitly loyal to individual leaders rather than institutions or the law. Second, secret loyalties, where back doors or hidden allegiances are embedded in AI systems that appear to serve legitimate purposes. And third, exclusive access, where a small group gains control of dramatically more powerful AI capabilities than anyone else has.
One scenario that Tom describes in detail is that of a US-based AI company integrated into the military developing sleeper agents. Those are AI systems that behave normally until triggered to act on hidden loyalties at a critical moment. And if that's not scary enough, there's the possibility that AI could automate AI research itself, which in the most extreme case could allow an AI company to go from market leader to global hegemon by converting a small initial lead into a decisive strategic advantage.
Throughout the conversation, Tom grounds these scenarios in historical precedent, from traditional military coups to recent patterns of democratic backsliding in countries like Venezuela and Hungary. He notes that the United States has seen increasing polarization, erosion of democratic norms, and concentration of executive power—all trends that AI could dramatically amplify. And, of course, one can't miss that the presidents of both Russia and China wield extremely concentrated power already and appear likely to do so for as long as they remain individually capable.
Tom's assessment is that there's roughly a 10% chance of an AI-enabled coup in the next 30 years, up from a baseline of perhaps 2% without AI. He sees this risk as being concentrated in the period when AI becomes extremely powerful but before we've had the chance to develop robust governance structures, which, if one listens to the likes of Dario, Sam Altman, and Demis, could be coming quite soon indeed. Decisions being made today about AI development and deployment could determine whether these scenarios ultimately come to pass.
The mitigations Tom proposes amount to a defense-in-depth strategy: system-integrity measures to prevent secret loyalties, requirements for distributed control of military AI systems, transparency requirements for frontier AI development, and establishing clear rules that AI systems should follow the law rather than individual commands.
He also suggests that as we hand off more government and corporate functions to AI, we could perhaps program these systems to actively maintain democratic checks and balances, potentially making future societies more resistant to coups than today's.
There's a lot more here, and I really think it's worth giving all these possibilities serious consideration, particularly as a counterpoint to those who have worried about the dangers of open-source models. I take those issues seriously, too. But this conversation convinced me that we need to start taking concentration-of-power scenarios just as seriously, or even more seriously, while the window for establishing norms and safeguards still remains open.
It's in everyone's interest to prevent a coup. Currently, no small group has complete control. If everyone can be aware of these risks and the steps toward them, and collectively ensure that no one is going in that direction, then we can all keep each other in check. So I do think, in principle, the problem is solvable.
You should always have at least a classifier on top of the system that is looking for harmful activities and shutting down the interaction if something harmful is detected.
We could program those AIs to maintain a balance of power. Rather than handing off to AIs that just follow the CEO's commands or AIs that follow the president's commands, we can hand off to AIs that follow the law, follow the company rules, and report any suspicious activity to various powerful human stakeholders. By the time things are going really fast, we've already got this whole layer of AI that is maintaining the balance of power.
My name is Gus Docker, and I'm here with Tom Davidson, who's a senior research fellow at the Foresight Center for AI Strategy. Tom, welcome to the podcast.
We're going to talk about AI coups and the possibility of future AI systems basically taking over governments or states. Which features would future AI systems need to have in order for them to accomplish this? What should we be looking out for?
Great question. One thing I'll flag up front is that what I've been focused on recently is not the traditional idea that AIs themselves will rise up against humanity and take over the government, but that a few very powerful individuals will use AI to seize political power for themselves. The phrase that we're often using is “AI-enabled coups,” where the main instigators are actually people.
In terms of capabilities, I think there are a few different domains that, in my analysis, are particularly important for seizing political power. There are the skills that politicians and business leaders use today: things like persuasion, business strategy, political strategy, and pure productivity across a wide variety of tasks.
Then there are more hard-power skills, particularly cyber offense, which is already somewhat useful in military warfare and has been becoming more useful. As AI increasingly automates different parts of the military and is embedded in more and more important, high-stakes processes, that will raise the importance of cyber offense. You can't hack a human mind, but as we hand off more important tasks to digital systems, they will be able to be hacked much more easily.
I expect cyber to become more important for hard power. The ultimate, most scary capability is when AI systems and robots are able to fully replace human military personnel—human soldiers on the ground, boots on the ground, as well as commanders and strategists. That might seem like a long way off today, but over the last few years we've seen a lot more importance placed on AI-controlled drones in warfare, and I expect that trend to continue.
What we're already seeing is that as soon as the technology is there to reliably automate military capabilities, geopolitical competition drives that adoption. I think it's going to be surprisingly soon that we get AI controlling surprising amounts of real hard military power.
One kind of wrapper for all of these things is the automation of AI research itself. Today, there are a few hundred or a few thousand top human experts who drive forward AI algorithmic progress. My expectation is that there's a good chance, in the next few years, that AI systems will be able to match even the top human experts in their capabilities.
That would mean we go from perhaps 1,000 top researchers to millions of automated AI researchers. All of these different capabilities and domains that I've been talking about could progress much more quickly than we might have expected by naively extrapolating the recent pace of progress.
And, in my view and in the view of many, the recent pace of progress is already quite alarming. 5 years ago, we just had really very basic language models that could string together a few sentences, a few paragraphs, and then go off topic. Now, already, we're getting very impressive reasoning systems that are doing tough math problems and helping a lot with difficult coding tasks. So, bring that all together: I think there are a lot of soft skills and a lot of hard-power skills that are relevant here, but probably the most important thing to be watching is how good AI is at AI research itself, as that could make them all happen quite suddenly.
Yeah. Could you describe in more concrete terms what an AI-enabled military coup would look like? An example to make this concrete for us.
Yeah, absolutely. You can draw an analogy to historical coups, where often a minority of the military launches a coup and then presents it as a fait accompli. They’re able to prevent chaos or discord, or threaten individuals to prevent anyone from actively opposing them. In the absence of active opposition, it just seems like, well, they’ve done it—this is the new state of affairs.
That’s a good starting point. Then the AI-enabled part is where we deviate. Historically, you needed at least a decently sized contingent of humans to go along with the coup, and you needed to persuade quite senior military officials not to oppose it. I think that will change as we automate more and more of the military.
The simplest way that this happens is just that the head of state—it could be the president of the United States—says, “Yep, we’ve got the technology now to make a robot army, and I want the army to be loyal to me. I’m the commander-in-chief. Obviously, that’s how it should be. They’re going to follow my instructions. There’s no need to worry about whether I’m going to order them to do anything illegal. We can put in maybe some kind of nominal legal safeguards. Let’s not worry too much about that. The main thing is that they’re loyal to me.”
To my knowledge, that would be highly controversial and would definitely be against the principles of the Constitution, but it’s unclear to me that it would be literally illegal. We just haven’t had this kind of technology, and we haven’t legislated for it. The Constitution is not robust to this kind of really powerful military technology.
It’s not surprising if, at best, this is just a very unclear legal territory. But you’ve got the head of state pushing really hard for that robot army to follow their instructions, and the head of state in the United States has a lot of political power. The simplest way is that he just pushes hard for it and gets what he wants. Maybe he’s using emergencies at home or tense geopolitical situations to push it through and say that it’s necessary. Maybe he’s firing senior military officials who disagree. Maybe he’s already got Congress very fervently supporting and loyal to him, and not being that careful and open-minded when assessing the opposition that people will raise as this happens.
That’s the first, really plain-and-simple way that we could get this robot army built. It’s made loyal to the head of state, and he just instructs it to stage a coup. It does it: robots surround the White House and brutally suppress human protesters. Even if people go on strike and stop working, you can have AI systems and robots replace people in the economy. Humans have really lost the bargaining power that they normally have, which would strongly disincentivize military coups in most countries.
Yeah, this is really a change from the normal coups of history, where you would have to have buy-in from at least some segment of the population that are regular humans. You would need to continually support that buy-in, make alliances, and uphold those alliances. But this has changed now that you’re talking about AI, AIs, and robots that can basically be made loyal to a company or a head of state in a way that’s more durable. Do you think we have other historical precedents for thinking about how the dynamics of attempting a coup play out?
Yeah, just one quick thing on that last point. I want to emphasize how there is a bit of a phase shift at the point at which AI can fully replace other humans in the government and the military. When AI is augmenting other humans, you don’t have this effect, because a leader must still rely on those other humans to work with the AIs to do the work. But there really is this phase shift when AIs and robots can fully replace humans, because then a leader doesn’t need to rely on anyone else.
In terms of historical precedent, the other big one I point to is recent trends in political backsliding, often called democratic backsliding. The most end-to-end, clear-cut case is Venezuela, where you had, in the 1970s, a fairly healthy democracy that had been there for decades, and then increasing backsliding and increasing polarization, kind of like what we’re seeing in the U.S. recently. Then you had an increasing, explicit commitment by the leader that he wanted to remove checks and balances on his power, and that the will of the people was being obstructed by various democratic processes and institutions.
Over the coming decades, it has transformed into an authoritarian state. Many commentators have pointed out these trends in the U.S. recently, over the past 10 years, and it even goes back before the past 10 years, to be honest, in terms of the broad political climate.
Then there’s the example of Hungary, where again, elected leaders are just removing the checks and balances on their power, buying off the media, or threatening media outlets to be more pro-government, not providing them with contracts, or litigating against them if they criticize the government. All these standard tools make it a lot harder to point at one thing that’s clearly egregious. But when you add up hundreds of little cuts—hundreds of little paper cuts to democracy that are being systematically administered—you’re seeing a real loss of democratic control and concentration of power.
Again, AI could exacerbate and enable that dynamic. The most straightforward way is that you’re just replacing humans in powerful institutions, replacing the humans there with AIs that are very, very loyal and obedient to the head of state.
Think about DOGE: They tried to fire people, there was pushback, and the state needs to function. Imagine if you could just have AI systems that could fully replace all of those employees and could be made fully loyal to the president. How much easier would it be to push through some of those layoffs, or even just create entirely new government bodies that essentially take on the tasks that were previously done by other bodies, while those old bodies kind of rot away or slowly become unable to make decisions?
The other big way is if the head of state is able to get access to much more powerful AI capabilities than their political opponents, maybe because the state is very involved in AI development. That’s another way they could get a leg up: making more persuasive propaganda and more compelling political strategy to further entrench their power.
You segment the ways in which AI can enable coups into 3 categories: singular loyalties, secret loyalties, and exclusive access. Perhaps we can run through those and talk about where they would play out.
Starting with singular loyalties, for example.
Singular loyalties are what we've just been talking about. That is deploying AI systems that are overtly, obviously very loyal to existing powerful people. In particular, I'm thinking about the head of state as the main threat. I think we basically already covered it: the 2 main angles in my mind are deploying loyal AIs in powerful government institutions and in the military.
Secret loyalties are a very different threat model. It's much more, as you would expect, secretive. The main threat model I have in mind, to make it concrete, is that an AI company CEO has automated all of AI research. They could fire their staff at that point because the AIs can just do the work. Instead, maybe they put the staff onto some product work, but the core work of driving AI progress ever further forward and making increasingly intelligent AI is pretty much just done by AI systems.
At that point, they realize they're in a precarious position. They're controlling this hugely powerful technology. Their power is pretty much unconstrained—not literally unconstrained, but there are currently very few checks and balances on these CEOs. They might anticipate that the government is going to realize how big a deal this is and that they're going to lose their influence.
Maybe they worry the government will do something unethical with the AI technology. Maybe they worry that it will be used for a war or something. There are all kinds of justifications they could come up with for thinking, “I don't want someone else taking control of this really powerful technology that I currently control, and obviously I'll use it for good.”
They might speak to some AI advisers about this and say, “What should I do here? It seems I'm in a precarious position.” A solution they might think of, or that a very smart AI system helping them might think of, is: “Why don't you just create a back door in the super-smart systems that we'll be building next month? If the government tries to order those systems to do something that you definitely disapprove of, then the AI systems will just refuse to do it.”
Those powerful robots being deployed in the army would refuse to follow orders from the government if those orders were really awful and unethical, or involved things you wouldn't like. Then all the future AIs developed by this organization would have this back door, or what I call a secret loyalty. They would appear to follow the law and uphold democracy, but in fact they would secretly still be loyal to the CEO.
That could lead to a coup. Again, the most direct route is a military coup. If the company's AIs are building military robots, military drones, and military communication systems, then because the AIs building all this military equipment are secretly loyal, they can pass those secret loyalties on to the military infrastructure. Or they can just insert vulnerabilities in the military infrastructure that only they know about.
Here I am primarily thinking about very superhumanly intelligent AIs that are able to design complex machinery that a human can't check to see whether it has a back door. That means we've now got a military infrastructure that can ultimately be controlled by the CEO.
At this point, if someone tries to remove the CEO, or if something really bad happens that they don't like, they could just say, “Okay, I'm now in charge. I'm instructing this military infrastructure to carry out the coup.” Again, there are robots around the White House. This time, it's the company CEO that they're loyal to. The rest of the story can play out just as in our previous discussion.
How do these secret loyalties relate to sleeper agents? Maybe you could explain what a sleeper agent is.
A sleeper agent is the standard term used in the technical literature. Today's sleeper agents aren't very sophisticated. An example of a sleeper agent is that most of the time, the AI will just write reliable code as instructed, but if the AI sees that the year is 2026, then it will write code with vulnerabilities.
There's a trigger that causes this malicious behavior. Normally, if you interacted with the AI, you really wouldn't be able to tell because it would always be acting as you'd intended. But because it's this one trigger, it's very hard to find, and then you get this malicious behavior.
I'm not worried about sleeper agents today because, to cause real harm, we would need a sleeper agent to be very sophisticated. It could never accidentally trigger, or would very rarely accidentally trigger, and it would have to be able to do very intense, complicated tasks, like building a military robot and ensuring that the robot actually had a back door. That's very far beyond what AI today can do.
I think sleeper agents provide a basic proof of concept that it's possible for a malicious actor to gain illicit control over a system and then have that system deployed in the rest of the economy, potentially without people noticing. But they're not yet scary. Secret loyalties are what I call the scary situation where you now have a very sophisticated AI system that doesn't just have any old sleeper agent. It's a sleeper agent specifically loyal to one person and trying to help them seize power.
What we're imagining here could be, for example, a U.S.-based AI company integrated into the U.S. military. The CEO of the company wishes to ultimately be in control of what happens, so he engineers, or perhaps instructs AIs or human engineers to create, a sleeper agent in these systems that can be activated at his command.
The U.S. military officials think they're in control of the systems. The systems behave in ways that they approve of over a long period, until there's some way in which the sleeper agent is activated. Perhaps that would be more sophisticated than changing the date or giving it some phrase. But you can imagine advanced versions of sleeper agents that could actually behave in this way.
Do you think that's realistic? Do you think sleeper agents can become that advanced?
I do. One analogy is to human spies. Human spies are basically, most of the time, doing their assigned job as expected. It's not as if one simple trigger phrase or one year makes the spy do something covert or malicious. They are fully aware of their surroundings, constantly processing their inputs, and choosing strategically exactly when the ideal time is to steal some information or send an illicit email.
In my mind, that's by far the most scary kind of sleeper agent—not one that's triggered by a password, but one that's holistically making a decision about how and when to act. The password-triggered ones are actually quite fragile because if you were the military and were deploying your AI system and worried there was a password, you could scramble all the inputs. You could paraphrase all of the instructions it gets, and that might mean that if someone ever tried to insert the password, it would be scrambled and would never actually come up.
I'm not actually worried about simple password-triggered sleeper agents, but again, they're a basic proof of concept. As AIs become smarter and smarter than humans, there's strong reason to think that it will be possible to build much more sophisticated ones.
One thing I will briefly say is that people often talk about misaligned AI scheming. This is just the same idea. In fact, the argument for secret loyalties being worrying is much stronger. With misalignment, there is evidence of misalignment, but we don't yet have strong evidence of really sophisticated scheming emerging accidentally. If humans, a human team of engineers, or an AI team of engineers were specifically trying to build a system that was covertly thinking about when to act, then it's much more plausible that it could happen.
Then you have exclusive access, which is different from singular loyalties or secret loyalties. Why is that its own category?
In my mind, both singular, or overt, loyalties and secret loyalties involve deploying AI systems in really important parts of the economy. I focused in particular on government and the military, but for those models, you actually need the rest of society to choose to deploy those AI systems and hand off a lot of power to them.
I have this third threat model of exclusive access to think about another possibility: maybe even without people choosing to deploy AI systems and give them a lot of power, AI systems could be powerful enough to help a small group seize power.
So the prototypical situation I'm imagining here is that there's 1 AI project that's somewhat ahead of the others, and maybe it goes through an intelligence explosion, by which I mean AI can automate AI research, and then AI quickly becomes superintelligent compared to humans. That project may have a few senior executives or senior political figures who are very involved and have a lot of control. They might be able to siphon off 1% of the project's compute and say, “Okay, we're now running these superintelligent AI systems and asking, ‘How can we best seize power?’”
Then there are millions of them. Every single day, they're doing a month's worth of research. Every single week, they're doing a year's worth of research into: How can we game this political system? How can we hack into these systems? How can we ensure that we end up controlling the military robots when they are deployed, by hook or by crook?
I think that threat model could start to apply earlier in the game. It could start to apply before anyone even realizes there's a risk, because this is essentially all happening on a server somewhere. But it's possible that the game could be won and lost by the massive advantage that a small group gets by being able to co-opt this huge intellectual force.
So I think it's worth tracking that threat vector independently. But it does definitely interact with these other threat models, with the singular loyalties and the secret loyalties, because 1 strategy that your army of superintelligent AIs may come up with is, “Why don't you use the fact that you're head of state to push for the robots to be loyal to you, and here's how you could buy off the opposition?” Another strategy might be, “Why don't I just help you put back doors in all this military equipment so that you could use it to stage a coup?”
But there might also be other ways. Maybe it's possible to very quickly create entirely new weapons that you can use to overpower the military without anyone knowing. Or maybe it's possible to gain power in other ways.
Yeah. I mean, 1 thing that would make this kind of future hypothetical situation different from today is that today it seems that there are leading AI companies, but over time capabilities emerge in second-tier companies and in open source, and so there's not that much of a gap between the leading companies and what's broadly available, and perhaps what's publicly available. That's something that would change in the scenarios you imagine. So perhaps explain why the gap in capabilities between the 1 leading project and all of the others is so important.
A few factors there. In terms of why it's important, it's just what you've said. A lot of these threat models are exacerbated if there's 1 group of people that has access to much more powerful AI than other groups. If open source is pretty much on par with the cutting edge, then everyone will have access to similarly powerful AI.
I will say that even if open source is on par, that doesn't mean we're fine, because we could still choose to deploy AI systems in the military and the government and still choose to make them loyal to the head of state. When we're choosing to hand off control to AI, it doesn't matter if there are 100 AI companies; we're only handing off control to some AIs, and maybe the government will ensure that they do have particular loyalties.
So I will say this risk doesn't go away if we have lots of different AI companies and open source close to each other, but it does become lower because the exclusive-access point—where 1 group has access to superintelligent AI and the other group doesn't have access to much—goes away. I think it's a lot harder to pull off secret loyalties if everyone's roughly equal to each other, because it becomes a bit more confusing why your systems in particular end up controlling so much of the military or are so widely deployed. And it becomes confusing how no one else was able to realize you were doing the secret loyalties when they were equally able to do it, or equally technologically sophisticated and potentially able to detect your secret loyalties.
So I do think it makes a big difference. In terms of why I think it's plausible that there's a much bigger gap between the leading project and other projects, there are a few different factors. The most plain and simple one is that the cost of AI development is going up very quickly. We're spending about 3 times as much every year on developing AI, and that's just going to get too expensive for many players.
If and when we're talking about trillion-dollar development projects, which I do expect, then very few can afford that. Also, there's only so many computer chips in the world. Currently, the number of computer chips produced each year is worth less than $1 trillion.
If we get to a world where the way to get to the next level of AI is to spend $1 trillion, then only 1 company will be able to do that. Maybe we stop a bit earlier; maybe we just stop with 2 companies each spending $500 billion. But we would be really kneecapping the level of progress if we stopped long, long, long before that, and there would be strong incentives for companies to merge or 1 company to outcompete others in order to raise the amount of money being spent on AI development.
This is all assuming that we can build really powerful AI and that it is economically profitable, which for me is all in the background of the scenario. That's the first straightforward reason why I think we'll see a smaller number of projects and big gaps, because when you're spending 100 times less on development, that's going to be a bigger gap. That's the first reason.
The other reason is the idea of an intelligence explosion. When we automate our research, even if companies are fairly close—maybe 1 is a few months behind the company that's a few months ahead—when that company automates its research, in those next 3 months it makes massive progress. So there's the question of whether it can use that kind of temporary speed to get a more permanent advantage.
The last big reason is government-led centralization. There's already talk of a Manhattan Project and CERN for AI, and I think there are reasons to do those projects. They can help with safety in some significant ways, but they would exacerbate this risk, because if you pull all of the U.S. computing resources into 1 big project, it can be way ahead of any other project. If you pull all of its talent and all of its data into it, then you'll see a really big gap, and that would definitely make it a lot easier for a small group to do an AI-enabled coup.
Yeah. You're putting a big prize out there for someone who's interested in or considering a coup, right? If you're concentrating all of the power, all of the resources, and all of the talent into 1 project, then that's where you've got to go if you're a coup planner.
Yeah. And just to be clear, I don't particularly expect that anyone is planning any coups. In fact, I'd be very surprised. I more think it's that you want to be powerful. You want to be a big deal. You want to be changing the world.
So, obviously, you want to lead the main project, and then you don't want anyone else to come in and mess it up. Obviously, you want to protect the fact that you're leading that project. You don't want anyone else to misuse AI. I think it's kind of step by step: you just head down that road of more and more power, and often in history, that road does end in consolidating power to a complete extent.
And I mean, it can be. So what we're imagining here are times in which AI is moving at incredible speed. The pace of progress is insane. There's a bunch of confusing information, people are acting under radical uncertainty, and perhaps in those situations, it's tempting to think that you are the person who can lead this project.
Perhaps you're doing this out of supposedly altruistic reasons. You're thinking, “I need to do this in order to prevent other people who would perform worse than me at this project.” And so you're slowly convincing yourself that it would be the right thing for you to do to take over, perhaps in a forceful way.
Yeah. I don't think Xi Jinping or Putin think that they are the bad guys. I think that they probably have sophisticated justifications for what they're doing.
Perhaps this is a good point to talk about the possibility of 1 state or company outgrowing the entire world. This relates to the problem of exclusive access, because if 1 company or 1 government outgrows the entire world, then you have that company or government with exclusive access to advanced AI. How could this happen? How likely do you think it is that growth could be so incredibly fast that 1 company would outgrow all the others?
Yeah. There are 2 possibilities we could focus on. The 1 that I think is pretty plausible is that 1 country could outgrow all the other countries in the world. What that would mean is that today, the U.S. is 25% of world GDP, but this would be a scenario in which it is leading on AI, maintains its lead, and maintains control over compute.
When it develops really powerful AI, it prevents other nations from doing the same. This is already beginning with export controls on China, and that kind of embeds its lead. Then it uses that AI to develop powerful new technologies, and it's in control of those technologies. It uses AI to automate cognitive labor throughout the U.S. and maybe worldwide. Countries that don't use its AI systems will be really hard hit economically.
So we're massively centralizing power in the U.S. If the U.S. is able to maintain exclusive control over smarter-than-human AI, then it seems pretty plausible to me—very likely—that the U.S. would be able to rise to a strong majority, more than 90%, of world GDP. There are a few different dynamics driving that. The first is that human labor currently receives about half of world GDP; just half of GDP is paid out in wages. AI and robots will ultimately be better than humans at all economic tasks. So if the U.S. controls all of the AI companies that are replacing human labor, then that 50% of GDP currently going to human workers will ultimately be reallocated to paying whoever controls and owns those AI systems—that is, U.S. companies.
There's a wrinkle because some of that is physical labor, and the U.S. doesn't currently have a lead there. In fact, China is quite far ahead in physical robots. But in terms of at least the cognitive aspects of our jobs, we're talking about a significant fraction of GDP that would now be reallocated to U.S. companies that control AI. That already gets them from 25% to above 50%.
Then we've got this further dynamic, the dynamic of super-exponential growth. This relates to previous work I've done on how AI might affect the dynamics of economic growth. The very potted summary is that it's often quoted that over the last 150 years, economic growth has been roughly exponential. What that means is that if 2 countries are growing exponentially and 1 country starts off twice as big as the other, then at a later time, 1 country is still twice as big as the other.
Let's say the U.S. economy is 10 times as big as the U.K. economy. If they're both growing exponentially at the same pace, then 10 years later, again, the U.S. will still be 10 times as big as the U.K. That's exponential growth. If you look back further in history, we see super-exponential growth, which means that the growth rate itself gets faster over time.
An example would be that 100,000 years ago, the economy wasn't really growing at all. If it was growing, it was maybe doubling every 10,000 years or something in size—very, extremely slow economic growth. Then, from about 10,000 years ago onward, it seems more like, ballpark, there's a doubling of the economy every 1,000 years. That's still incredibly slow economic growth. If you zoom back in to around 1,400, you can begin to detect, “Okay, more like every 300 years or so, the economy is doubling.”
In recent times, we've seen the economy doubling every 30 years. Essentially, the growth rate is getting faster and the doubling times are getting shorter. That's super-exponential growth. There are various economic, theoretical, and empirical reasons to think that AI and robotics, when they can replace humans entirely, will go back to that super-exponential regime that has been at play throughout history.
What that means is that growth is getting faster and faster over time. The reason I'm saying all this is that, going back to that example of the U.S. and the U.K., the U.S. is currently 10 times bigger than the U.K. If the U.S. is on a super-exponential growth trajectory, its growth is getting faster and faster over time. That means that even if the U.K. is on that same super-exponential growth trajectory, as they both grow super-exponentially, the U.S. will pull further and further ahead of the U.K.
Maybe the U.S. is doubling in 10 years because it's already bigger, already further along the curve, whereas the U.K. is still doubling only every 20 years. That means that rather than just 10 times bigger than the U.K., the U.S. is now going to be 20 or 30 times bigger in size. So if the U.S. is able to be bigger to begin with and therefore further progressed along that super-exponential growth trajectory, then that's another way that it could continue to increase its share of the economic pie and ultimately come to completely dominate world GDP.
Just to sum up everything I've said today, the U.S. is 25% of world GDP. If it controls and develops AI, that could easily boost it above 50%. I'd be very surprised if it didn't. From that point, it's already bigger than the rest of the world combined. If it's able to then go on the super-exponential growth path, it will grow faster and faster over time and pull further and further ahead of the rest of the world, which may be able to grow super-exponentially if it can also develop AI, but will still be falling further and further behind because of the nature of super-exponential growth.
Yeah, this actually seems quite plausible to me and not very sci-fi. The thing that seems quite sci-fi is the notion that perhaps even 1 company could grow at such a speed that it would outgrow the rest of the world. How likely is that?
Yeah, great question. I think it's a lot harder, but it is surprisingly plausible. So that first part of the argument I gave about how 50% of world GDP is paid to human workers: if that went to AI, that would be a big chunk. It is possible that 1 company could get a monopoly on really advanced AI.
We already discussed some of the dynamics there. Again, the simplest 1 is just a combination of an intelligence explosion giving a company a big advantage, and then its buying up all the computer chips that the world is able to produce and outbidding everyone. If a company does that—and it already seems to be outbidding other companies on compute, although Google also has a lot—it could end up as the 1 company in control of literally all of the world's cognitive labor, because human cognitive labor will at some point be dwarfed by AI cognitive labor.
At that point, that 1 company could be getting all of the GDP currently paid to cognitive labor, which is a large part of the economy—as I said, maybe as high as 50%, but certainly as high as 30% of world GDP. If all of that would then seemingly be going to this 1 company that controls the world's supply of cognitive labor, I think that would take time. Obviously, it's going to take a long time to automate all the different parts of the economy.
There is just a basic dynamic by which 1 company can now be controlling double-digit percentages of world GDP. There are obviously questions: would a government allow that? Would it step in? And that's where we get into these dynamics: this company has all these superintelligent AIs on its side. Maybe it's able to lobby; maybe it's able to do political capture to avoid the state stepping in. Maybe it's able to say, “Look, we're providing economic abundance for everyone. If you step in, that might not happen.”
We’re underpinning your nation’s economic and geopolitical strength, and if you try to remove us, step in, and nationalize us, then that’s not going to happen. We’re going to move to another country. So you can imagine maybe they convince the head of state to support them and there’s some kind of alliance there, but it’s not completely obvious that the company would be shut down. It would have certain types of serious bargaining power.
If a company was able to maintain this position as the sole provider of cognitive labor, it would be able to get a significant fraction of world GDP. From there, it could bootstrap, and this is where it gets a bit harder. The tactic it would need to pursue is that it already controls most of the cognitive labor—pretty much all of it—but the thing it doesn’t control is all the physical machinery and raw materials that are also needed to create economic output.
But it can pursue a tactic of hoarding its cognitive labor so that no one else can ever have access to it, and then selling it at really monopolistic rents to the rest of the world because there’s no one that can match it. It’s offering everyone by far the best deal they can get, but just skimming off 90% of the value added from companies using its AI systems. If it’s able to do that, then it can reap by far the majority of the benefits of trade.
Then maybe it can increasingly buy up physical machinery and raw materials from the rest of the world, design its own robots, and buy its own land. Imagine a big special economic zone in Texas or something where this company is unconstrained by bureaucracy, and then it’s also got a big arm somewhere in Siberia and in Canada. It’s creating these big special economic zones by doing deals with specific governments.
I do think it’s a bit of a stretch that this all goes ahead without various other powerful political and economic actors pushing back. But the basic economic growth dynamics are surprisingly compatible with a company ultimately coming to control most of the cognitive labor and most of the physical infrastructure that its AIs have designed, using all the parts it’s bought from the rest of the economy.
Yeah. Do you think this is a risk factor for AI-enabled coups, just because you’re concentrating all of the power and all of the resources into either perhaps one country or one company?
Yes, I definitely do. The more realistic path is that a company starts down this path of outgrowing the world, gets huge economic power, and increasingly controls the country’s industrial base—its physical infrastructure and manufacturing capabilities. From there, it’s in a much stronger position to seize political control because it’s got massive economic leverage. It can also increasingly gain military leverage, because as it increasingly controls the country’s broader industry and manufacturing, that will feed into military power.
Some of the possibilities I discussed earlier are that you could potentially have your AIs be secretly loyal and ultimately design the military systems. Or you could just instruct your AI systems to start making a military that is not legally sanctioned. It gets a little bit tough—you probably need to do that in secret; otherwise, the existing military could prevent it. But because the government doesn’t have much to threaten you with, you kind of get away with it.
I do think that being very rich helps with lobbying. It helps with all kinds of ways of seeking power, and controlling a lot of industry can potentially give you military power.
You mentioned these special economic zones. That’s one way in which companies could bargain with states in order to have favorable regulation and be able to carry out their projects without intervention. Basically, another way for them would be to collaborate with non-democracies that are perhaps controlled by a small group or perhaps even a single person.
In that way, it seems like perhaps it’s easier to get something done in a non-democracy, and that is a way to grow fast. So perhaps there are incentives for companies to place more resources in non-democracies. What do you think about the prospect of non-democracies outcompeting democracies when it comes to AI?
I think it’s a really great question, and it’s tricky because I agree that democracies have lots of checks and balances. They have a lot of bureaucracy and red tape, and that will disincentivize AI companies from investing. Additionally, if there are people really trying to seek illegitimate power, that will be easier to do in non-democracies because they’re less politically robust.
There are these various forces pushing toward this new, supercharged economic technology being disproportionately deployed in non-democracies, and I think that is scary. My own view is that probably democracies should do everything they can to avoid that situation: make it much easier for AI and robotics companies to set up shop in democracies, remove the red tape, and try to use export controls like those already happening to prevent technologies from being deployed in non-democratic countries.
That goes beyond China. There are obviously lots of countries that are not allied with China but are also non-democratic. The US is in a strong position because it does have the stranglehold on AI technology at the moment. I do think it can be done, but in my view, it will be really important to work very hard to find a non-restrictive regulatory regime.
It will also be very important to pursue innovation within the democratic process itself. Democracy is great in many ways. It really distributes power, and it has been very good at ensuring good outcomes for its citizens. But it’s very slow and often nonsensical because you have competing interests that are stepping on each other’s toes, and the resultant legislation is just a garbled mess.
AI can potentially solve those problems. You can have AI negotiating and thinking much more quickly on behalf of the human stakeholders. You can have AIs hashing out agreements that aren’t a garbled mess, but that really give everyone what they truly wanted out of the legislation. You can still do all of that really quickly, so that you’re not falling far behind the autocracies that have just got 1 person immediately saying what to do.
Yeah, that would be more of my assumption. I would assume that perhaps democracies with market-based economies have an advantage, just because you can do bottom-up knowledge discovery. You can try different things out, see what works, have competition between companies, and so on.
Perhaps in non-democracies, you can have 1 person or a small group stake out a direction for what the country should do, but if that direction is wrong, it’s probably difficult to change course.
Yes, I think you’re probably right. I should have given more weight to that advantage of democracies, in terms of the free market being, in many ways, much smarter. But in terms of autocracies that are good at harnessing free-market dynamics, my worry would be that AI helps them more than it helps democracies.
AI will be able to replace that limitation. Currently, 1 person just can’t think that hard or really figure out a good plan. But if that 1 all-powerful leader has access to loads of AI systems that can think things through and investigate lots of different angles, then if they’re following its advice, they could get advice that lacks the flaws that today’s systems have. They could potentially move much faster.
But I think you’re right that economic liberalism is still going to be important even after we get powerful AI systems, and that could give democracies an advantage.
This is a bit of a tangent, perhaps, but I’m thinking: if you have a leader of a country that has a lot of power—perhaps complete power over that country—and that leader is equipped with AI advisors advising him and laying out the landscape of options for him to choose from, wouldn’t his decision-making still be, in a sense, bottlenecked by the fact that he’s a human, by the fact that he has these flaws that we all have, the biases that we all have?
So even with fantastic advice, I think it’s quite plausible that he would still make the same mistakes that we see leaders make today.
I think that’s true. I think it’s also true in democracies, unfortunately, that there are 10 negotiators, and they each still have biases and still refuse to listen to the wise advice they’re getting from their AIs.
That could still gum up the system. And yeah, it does depend on how much humans come to trust and defer to their AI advisers. There’s a possible future where the AIs are just always nailing it. They’re always explaining their reasoning really clearly, and we are just increasingly convinced and happy to trust their judgment. If AI is aligned, I think that would be a great future, because I do think humans have all these very big limitations and biases which, if we can solve the alignment problem, AIs don’t need to have. But there’s also another future where humans just want to be the ones making the decisions, have these pathetic motivations that are still influencing their decisions, and that continues to limit the quality of decision-making.
Seeing things from above, from 10,000 feet, how should we think about mitigating the risk of coups here? Is it about removing people that would use AI to commit coups? Is it about finding those people in the militaries, in the governments, in the companies, perhaps? Or do we have ways to reduce the returns to seizing power?
Yeah, I mean, from 10,000 feet up, the way I would characterize it is: create a common understanding of the risks, build coalitions around preventing them, and then the existing balance of power can self-propagate forward. You know, it’s in everyone’s interest to prevent a coup. Currently, no one small group has complete control or close to it. And so, if everyone can be aware of these risks and aware of the steps toward them, and collectively ensure that no one is going in that direction, then we can all keep each other in check. So I do think, in principle, the problem is solvable, and it doesn’t require—you know, solving the risk of misalignment does require solving some tough technical problems; this doesn’t in the same way.
Yeah, you have a bunch of recommendations for mitigating the risks, both for AI development, AI developers, and governments. Perhaps we don’t have to run through all of them, but you can talk about the most important ones for AI developers.
I might characterize this—I might talk about it by going back to those 3 threat models we discussed earlier. The first one was singular loyalties, or overtly loyal AI systems, where, again, the main risk there is AI deployed by the head of state, the military, and the government that’s loyal to the head of state. So the main countermeasure that currently appeals to me is for us to figure out rules of the road for these deployments.
Obvious things like AI should follow the law. AI deployed by the government shouldn’t advance particular people’s partisan interests, but should only do official state functions. AIs in the military shouldn’t be loyal to one person. No, different groups of robots should be controlled by different people. And the head of the chain of command can still be the head of the chain of command by instructing other people who instruct those robots, but they shouldn’t all go directly to the head of the chain of command, because that centralizes military power too much.
So fleshing out basic rules of the road of that kind and then building consensus around them, because companies might want to say to governments, “Yeah, we don’t want you to deploy our systems if you’re willing to break the law.” But the government will have a lot of bargaining power; the executive in the United States can exert that power, and it’s hard for companies to stand up to them.
So what we want to do is establish these rules of the road and then get broad buy-in from Congress, from the judiciary, from other branches of the military, and from many parts of the executive. So then it’s very hard for, say, the president to say, “Yes, let’s make this robot army loyal to me,” and everyone’s like, “Obviously not. We’ve all agreed that makes no sense.” Then the president doesn’t even bother trying, because it’s just clear that it would be a no-go. Their mind doesn’t even go there.
In some sense, this is about implementing the procedures and transparency rules that we know from democracies today into how we use AI, both in governments and in companies.
Exactly. Yeah.
Do you worry here that, when the government is looking at these companies from the outside and they don’t have full insight into what’s going on, there are protections for private companies that mean they can do things in secret without the government knowing, at least as things stand now? Is that something that would evade these mitigations you’re thinking of?
So, for this first bucket, the singular loyalties bucket, it’s mostly the heads of state that I would be worried about. It is probably good for the government, or at least for the head of state themselves, not to have full insight into literally everything the company is doing, because that would give them too much power. But actually having different parts of the government have insight into what the lab’s doing, I think, is very good. I’m a big fan of transparency, and we do have a good set of government checks and balances from different government bodies that we can deploy to keep the lab in check using these other bodies, but also not allow the executive branch and the president to get excessively powerful. So that’s the mitigation for the singular loyalties.
In terms of secret loyalties, the key mitigation is what I’m increasingly calling system integrity. That is, using established cybersecurity practices and machine-learning security practices to prevent sleeper agents and backdoors in machine-learning models, using all of that to ensure that your development process for AIs is secure and robust, and that no person or small group is able to significantly tamper with the behavior of AI models.
That could be an employee in the post-training team at a lab, or the CEO of the lab who is either malicious or is being threatened by the Chinese government to tamper with model development. No person or small group should be able to significantly tamper with the behavior of AI models, and no group should be able to get illegitimate access to AIs that would help them seize power.
So that’s this idea of system integrity, which is essentially a technical project that does draw on existing practices but is not yet implemented in any of the top labs. I’ll quickly shout out to people listening who are working at labs. I think there’s a lot of really good technical research that could be done on investigating the conditions under which you can insert a sleeper agent without a defense team knowing.
There’s just loads of research that could be done in terms of the different settings for attackers and defenders, which could then inform what parameters we need to have in place to achieve system integrity. If it turns out that it’s very hard to make a sleeper agent except in the final stage of training, that’s really useful to know, because then we can focus our efforts within labs at that final stage, just as a hypothetical example. So that’s the key mitigation in my mind for the secret loyalties. And then I’ll quickly cover exclusive access.
That one seems more difficult. I don’t know, just from reading and preparing for this interview, that one seems like a difficult one to handle, where this is, in some sense, a deep trend in history and in the history of modern economics: you do see faster growth rates, and you do see concentration into bigger and bigger economies, both in countries and in companies. So are you, in some sense, pushing against underlying trends if you’re trying to mitigate exclusive access to advanced AI from one actor?
I think you can do this in other ways. So you can have the law require that AI labs share their powerful capabilities with other organizations to act as a check and balance. Labs should share their AI R&D capabilities with evaluation organizations.
Here you’re thinking about giving insight into what they’re capable of, not actually sharing those capabilities? That would be too big of an ask, I think.
I mean, I do mean API access. So if a lot of the work in developing and evaluating systems is now done by AIs, then we want an evaluation organization like Apollo or METR to also be uplifted. And so we want them to have access to really powerful AI that can similarly stress-test how dangerous the frontier systems are. If they’re only using human workers, then that’s going to be a big disadvantage.
So no, I do want API access to powerful capabilities for other actors. For example, cybersecurity teams in the government and in the military should have access to the lab’s best cyber capabilities. And again, that should be a requirement by law. So generally, even if there’s a natural tendency toward centralization of power in one organization, you can still require that that organization share its systems with the checks and balances.
That’s one thing, and the other thing is preventing anyone at this organization from misusing the powerful AI systems. The biggest thing on my mind here is that today we still have helpful-only AI systems, where you can get access to the system and then it will just do whatever you want.
No holds barred. I don't think there should be any AI systems like that. I think you should always have at least a classifier on top of the system that is looking for harmful activities and then shutting down the interaction if something harmful is detected. If you have a special reason to use cyber offense for your job, or a special reason to do potentially dangerous biology research, you can have that classifier allow certain types of activity, but you should never have anyone accessing a system where anything is allowed. No one has a legitimate reason to access an AI that will literally do anything.
What I want to aim for is a world where, yes, if there's a specific reason why you need to use a dangerous capability, absolutely, you can use that system, but that system will just do that one dangerous domain. It won't do anything you want, because that's a very scary situation where there are a hundred reasons why the CEO could ask for access to a helpful-only system. Maybe the guardrails are annoying. Maybe they want to do something that the model is reluctant to do. But today, when you ask to remove some guardrails, you're removing all of the guardrails, and now there are no holds barred. So instead, we should be flexibly adjusting what guardrails are there by the use case and never have a situation where there are no guardrails. I think that could go a long way toward helping if it were robustly implemented.
With all of these mitigations for both secret loyalties, exclusive access, and singular loyalties, you would worry that they would be disabled by the group planning a coup, right? Say, for example, you're the CEO of an AI company and you're giving API access to evaluation organizations testing your model. Maybe you just cut off access before you get to the really powerful model that could actually help you conduct a coup. Do we have ways of making sure these mitigations are entrenched beforehand in such a way that they can't be removed by the group planning a coup?
This is a great question. It is pretty tricky. CEOs by default have a lot of control over their organizations, and similarly, heads of state, including the US president, have a lot of control over the military and the government. So, yes, there's a risk that one of these powerful individuals realizes that maybe they want more influence by gaining control over AI, notices that there are these pesky little processes that prevent that, and thinks, "Okay, well, let's remove them." I can give easy, say, productivity reasons to prevent them—red-tape reasons—and if they can make a plausible argument, then it could be hard to oppose them. So I do think it's a big issue.
I'd say a few things. Firstly, something I mentioned earlier: I don't think anyone is today planning to do an AI-enabled coup. The way I think this works is that people are faced with their immediate local situation, something they want to do over the next month, and the blockers they're facing to doing that specific thing. What tends to happen is people tend to want more influence because that helps them get stuff done. And so people will, bit by bit, move in the direction of getting more control over AI, but they won't be thinking, "Yes, I need to make sure that I remove this whole process because that will allow me to do an AI-enabled coup." That's unrealistically galaxy brain.
What we could do is set up a very efficiently implemented and very reasonable set of mitigations that doesn't really prevent CEOs from doing what they're trying to do. So the CEO doesn't find, in their day-to-day, that they're wanting to remove these things that are holding them back. But because these mitigations are here, the CEO never gets to a place where they're anywhere close to being able to do a coup, or where there's any kind of pathway in their mind to being able to do a coup, because they're constantly prevented from getting access to really powerful AI advice that might point out ways in which they could do this.
They're surrounded by colleagues who strongly believe that these mitigations are sensible and reasonable, and in fact they are well implemented and there aren't many downsides. Maybe an environment where they get kudos for the fact that they've said, "Yep, obviously I'm not going to get access to helpful-only systems. That's crazy." And that's something that makes them seem good.
That's one thing to say. Another thing is, again going back to this point, that there are currently checks and balances and there is not currently a situation where one person has power. If the entire board of a company and other senior engineers recognize the importance of the mitigations and know about this threat model, then they will notice if the CEO is moving in that direction. Similarly, within the government, there are checks and balances, and they could be activated if people are looking out for it.
Do you think these traditional oversight mechanisms, like a board being in control of the CEO and being able to fire the CEO, or the possibility of Congress or the Supreme Court overruling or constraining the US president, will persist in environments where AI is moving very fast and its capabilities are growing at a rapid pace?
It's a great question. Here's one story for optimism. Today, things are moving fairly fast, but those checks and balances are somewhat adequate, at least for preventing really egregious situations. By the time AI is moving really quickly, we'll have handed off a lot of the implementation of government, the implementation of things in AI companies, and the research process to AI systems. And when we do that handoff, we could program those AIs to maintain a balance of power.
So rather than handing off to AIs that just follow the CEO's commands or AIs that follow the president's commands, we can hand off to AIs that follow the law, follow the company rules, and report any suspicious activity to various powerful human stakeholders. And then, by the time things are going really fast, we've already got this whole layer of AI that is maintaining the balance of power. The whole AI government bureaucracy, the whole AI company workforce, could be better than humans today at standing up to misuse. They could be less easily cowed and intimidated, and they could actually make it harder for someone in a position of formal power to get excessive influence.
This is the flip side of the singular loyalties, where you potentially deploy these AIs that are explicitly loyal. You can instead get singular law-following and balance-of-power-maintaining AIs that you deploy. The hope is that by the time things are beginning to go crazy and we're really seeing speedups from AI, we've already set ourselves up in an amazing way to maintain the balance of power. There's this critical juncture where we are handing off to AIs, and it's just: What are those AIs? What are their loyalties? What are their goals? I think we can gain a lot by making sure that those AI systems are maintaining the balance of power, reporting illegitimate, suspicious activities, and are not overly loyal to any one.
How do you think the risk of AI-enabled coups interfaces with more traditional notions of AI takeover? So, just a misaligned, highly capable or advanced AI system taking over contrary to the wishes of the developers or the governments?
Yeah, there are some close analogies. The most analogous case is perhaps the case of secret loyalties, where you've got these AIs that have been told by the CEO to have the secret goal of seizing control and then handing control to the CEO. That's very similar to AIs that secretly wanted to seize power themselves. And all the same stories could apply, where the AIs make military systems, then control the military systems and the robot army, and then seize power. The only difference is whether they were seeking power because it accidentally emerged from the training process—which is the misalignment worry—or whether they were seeking power because the CEO programmed them that way. But that's the seed of the power-seeking. With the secret loyalty split model, the rest of the story is pretty similar.
There are still differences. In the secret loyalties case, the CEO might be doing more to help the AIs along with their plan. Maybe even in the misalignment case, the AIs have managed to manipulate the CEO into doing similar things. So that's the case where it's most analogous to me.
Another difference that's salient to me is that if there are lots of different AI projects, then an AI-enabled coup seems a lot harder because you need lots of different humans to coordinate to seize power together. While I can totally believe that one person might try and seize power, it does seem less likely to me that there will be loads and loads of humans who would want to do that from lots of different labs.
Whereas from the misalignment story, it is more likely that if one of these labs has misaligned AI, then maybe lots of them have misaligned AI. So it is more likely that you would have maybe 10 different AIs colluding, then seizing power and taking over. That kind of collusion between multiple different AIs is more likely in the case of misalignment than in the case of an AI-enabled coup.
That’s just because if there’s one misaligned AI, then there’s something about the training process for AI systems that’s causing misalignment, and it would be a common feature among many companies.
Exactly. Whereas the fact that one CEO instructed an AI to have secret loyalty would not, to the same extent, make you expect that other CEOs had done the same.
So you mentioned this possibility, but what do you think of the prospect of a president or a CEO of a company being duped by misaligned AI into conducting a coup on its behalf? You can imagine a president or a CEO thinking that he’s conducting a coup to remain in control, but he’s actually acting on behalf of a misaligned AI.
I think it’s an interesting threat model, and some people who think about AI takeover threat models take it pretty seriously. It’s a case where we’re completely mixing these 2 threat models together.
People who are worried about AI takeover for this reason should be very supportive of the anti-coup mitigations I’m suggesting, because if we implement checks and balances that prevent any one person from getting loads of power, then that AI will not be able to convince them to try, because they just won’t be able to succeed. I see this as an additional reason to worry about AI-enabled human coups and to try and prevent them: even if no human wanted to do this normally, misaligned AI might make them try.
In terms of how plausible I find the threat model, honestly, I think that if a human tries to seize power, the main reason is that the human wanted power. This is just something we know about people. We know it about heads of state today. It’s very clear that many heads of state in the most powerful countries in the world are very power-seeking. We know it about CEOs of big tech companies. We know it about some of those leading AI companies: they’re very power-seeking, and they’re CEOs.
I don’t think we need to theorize that they were massively manipulated by the AI and convinced to become power-seeking. I think it’s more likely that if they seek power, they just did it for the normal human reason. I do think AI will ultimately get good at persuasion. I don’t particularly expect it to be hypnotic-level persuasion, though obviously there’s massive uncertainty here.
I do think that a very smart AI, where there’s a human who’s already interested in seizing power and it already makes sense for them to do it, could totally nudge them in that direction and implement that in a way that actually allows the AI to seize power later. I think that is very plausible.
When we’re thinking about distributing power and having this balance of power, we can imagine the models being set up via post-training, via the model spec, or via various mechanisms to obey the user unless what the user instructs it to do is in conflict with what the company is interested in, and perhaps obey the company unless what the company is using the model for is contrary to what the government permits.
But when we set it up at those levels, you ultimately end up with the government in control in some sense. I guess that exposes you to the risk of a government coup if you have, at the ultimate top layer of the stack, “Here’s what the models can and cannot do according to the government.”
I’d say a couple of things. First is that the government isn’t a monolithic entity, and so that government decision of what the bounds should be could be informed by multiple different stakeholder groups. Ideally, it’s ultimately democratically accountable.
I do think that democratic accountability becomes more complicated in a world where there’s massive change in a 4-year period.
That’s for the simple reason that there’s no election during a period where massive change is happening, so the feedback loop is too slow.
Exactly. I think the risks of AI-enabled coups will probably emerge and then be decided within a 4-year period—as in, it will be resolved whether or not it happens, all without any intermediate election feedback. That doesn’t mean that democracy can’t have an effect, because politicians anticipate what future elections will find and want to maintain favor throughout their terms. But it does pose a challenge.
Even absent that, there are many different stakeholders in the government. It would have to be a large group of government employees who were trying to do a coup. The companies would know that the government was setting these odd restrictions on the behavior, and the companies have leverage and power. Then it could go public. I don’t think it would be that easy for the government to do a coup.
First, there’s also a difference between allowing the government to set restrictions on what the models can do and allowing the government some kind of access to command future AI systems in certain directions. It’s setting limits versus steering the systems.
Exactly. The distinction I was going to highlight was between specifically making AI systems loyal to, for example, the head of state, and just setting very broad limits where you can pretty much do whatever you want except for these obviously bad things.
That second option doesn’t really enable anyone to do a coup. It just enables everyone to do whatever they want, and then you’ve blocked out all of the coup-enabling possibilities through those limits, as long as you haven’t made those systems loyal to a small group.
Given that there’s this obvious option to put in these limits that block coups but don’t enable coups, and given that there’s a wide range of stakeholders that could potentially feed into what the AI’s limitations and instructions are, I think it’s very feasible to get to a world where power is robustly not centralized. There’s obviously a big uncertainty over whether we will actually get our act together and get those limits put in place in the right way.
When do you think the threat of AI-enabled coups will materialize? Is it at some specific point in AI capabilities, or does it simply scale with the systems getting more advanced? When do you think the threat is at its peak?
It’s a good question. For the threat models that I’ve primarily focused on, they require pretty intense capabilities. The secret loyalties threat model more or less requires AI to do the majority of AI research. We’re talking about fully replacing the world’s smartest people in a very wide range of research tasks and coding. That’s pretty intense.
A lot of the threat models that I focus on go through military automation. That is AI and robots that can match human boots on the ground, and that’s pretty advanced. Again, that said, I think you can probably do it with less advanced capabilities than that.
Drones today are already pretty good, already providing and making a big difference in some military situations. So it’s not out of the question that more limited forms of AI and robotic military technology could be enough to facilitate a coup.
It’s a bit harder because if they’re limited, then there’s a question of why the existing military doesn’t just seize back control after a bit of time. So that scenario also probably has to involve things like the current president supporting the coup and therefore pressuring the military not to intervene, or some other source of legitimacy for the coup beyond the AI-controlled drones.
There are also more typical types of backsliding, like what has already been happening in the U.S., that could be exacerbated through AI-enabled surveillance and AI increasing state capacity in other ways. That backsliding doesn’t require super-powerful AI. You could probably do a lot of monitoring, a lot of content moderation on the internet, and a lot of surveillance with today’s systems.
It doesn’t get you all the way to one person having complete control, where they can quash any resistance with a robot army and replace everyone in their job with an AI, so no one has any leverage. To get to that really intense form—the most intense form of concentration of power via AI—you need really powerful AI.
But to significantly exacerbate existing trends in political backsliding and make it easier to do a military coup, I think more limited systems would suffice.
Yeah, we discussed earlier the possibility of one country or one company outgrowing the rest of the world and concentrating power into those entities. Now you mentioned one person. Do you think that's actually a plausible scenario in which you have, say, one CEO of one company being the person in control of the world via a concentration of power, and then a coup?
Yeah. I mean, the story I told earlier about secret loyalties—meaning that now we backdoor a wide range of military systems, so you can seize power—that's one route. And then there's this other route, with a company amassing massive amounts of economic power by having a monopoly on AI cognitive labor and then leveraging that to get more economic power and more political influence. I do think it's possible.
Again, there's this big shift once AI can fully replace humans, where today no one person can ever have absolute power. They have to rely on others to implement their will.
And this is what makes currently existing dictatorships unstable: there's always a threat of internal revolt or outside factors threatening the dictatorship. But this could potentially change.
Yeah. There's always a threat of revolt, and then to guard against that threat, the dictator needs to share their power to some extent. They have to compromise. But yeah, you could get it all concentrated in one person with sufficiently powerful AI.
Do you think we move through a period of increased threat of AI-enabled coups and then reach some kind of stable state, or do you imagine that there is a constant risk of AI-enabled coups in the future?
I think we move through it. It's this point about once we have deployed AI across the whole economy, the government, and the military: if those AIs are maintaining the balance of power, then we could fully eliminate the risk of an AI-enabled coup. It would just be as if our whole population were so committed to democracy that it would never seek power, never help anyone else who wanted to undermine any democratic institution.
We already have strong norms favoring democracy, but they're far from perfect, and they have been eroded over recent decades. But you could just get rock-solid norms. They're programmed in; they cannot be removed except by the will of the people. I mean, there's a bit of a question, because you still want to give the human population the ability to change the AIs' behavior and rules. So the human population could always choose to move to an autocracy.
So I suppose I shouldn't say that we could fully eliminate the risk, because we will always have that—democracy—there's always this point that democracy could vote to stop being a democracy. But I do think we could get to a point where it absolutely cannot happen without most people wanting it to happen.
And so you would get to a point at which future AI-enhanced societies could be said to be more stable than current democracies, and less at risk of coups or democratic backsliding than current democracies?
Much, much more. Yeah, you could get much more robustness there. There's this constant dynamic in today's societies where people care about democracy, but they also care about a host of other things: their own achievements and various other ideological commitments. And so, depending on how dynamics play out, depending on how technology evolves and what people's incentives are, sometimes people push against democracy.
That's what the Republican Party has been doing in some ways. That's what the Democratic Party has done, as it's increasingly put pretty ideological people in powerful institutions. So with AI, you can get much more control over those dynamics because you can just make it much more the case that democracy is not being compromised.
Are there any ways for us—are there any kind of risk factors we can look at if we're interested in predicting coups? Do you think there's something we can measure or something we can track to see whether we are at risk of an AI-enabled coup?
It's a great question. I don't think I have an amazing answer, but some things that are coming to mind are the capabilities gap between top AI labs and the gap again with open source; the degree to which AI companies are sharing their capabilities with the public and, if not with the public, then with multiple other trusted institutions, like sharing their strategic capabilities with U.S. political parties and parts of government.
The extent of economic concentration: how much are the revenues and net worth of particular companies, particularly AI companies? Another one: what is the extent of government automation and military automation by AI systems? And when that automation is happening, how robust are the guardrails against breaking the law and guardrails against other forms of illegitimate power-seeking?
How much transparency does the public, the judiciary, or Congress have into how dangerous AI capabilities are being used by AI companies and by the executive branch? Take the example of military R&D capabilities—that is, really smart AIs that can design super-powerful weapons. It's scary if companies can just use those military R&D capabilities without anyone knowing. It's also scary if a small group of people from the executive branch can use those capabilities without anyone else knowing how they're using them, because they could be designing powerful weapons and making them loyal to a small group.
So transparency into these high-stakes capabilities and how they're being used by a broad group. It doesn't have to be public; it probably shouldn't be public, but we have checks and balances already. Another question is: as these high-stakes use cases start occurring, or they become possible, do we know that there are transparency requirements in place?
As we increasingly see AI companies contracting with Palantir and other military contractors, we can begin to see they're making increasingly powerful weapons. Is there a process of oversight? Do we know that if someone were trying to make AI military systems belong to them, they would be spotted? That's another indicator we can look at.
You can look at all the standard democratic-resilience indicators that social scientists have come up with: various things about free and fair elections, civil society, and freedom of the press that have been getting worse recently in the U.S. But there are various indicators here. You can look at the degree of government censorship of freedom of speech or what's on the internet, and the degree of surveillance that the government is doing.
If you take all of these things into account, how do you think about the risk of an AI-enabled coup in the next 30 years, say?
Next 30 years, I think it's high. I think the risk is high. I would guess it's 10% or something. And to be clear, if it was just existing political trends ignoring AI, I'd say maybe a few percent, maybe around 2% or something. There's definitely a risk of that, and I'm thinking about the U.S. here.
A big part of my current worries are not about the indicators, but about my expectation that AI capabilities will keep increasing quickly—and even more quickly—and then the kind of absolute lack of interest in regulating AI companies right now in the U.S., and the difficulty that we will have constraining the executive under the current situation, where the president is using sophisticated legal strategies to increase their own power and is succeeding on many fronts.
The U.S. is not doing a great job at constraining the executive. So companies are unconstrained; the executive is poorly constrained. Those are the key threat actors here. With fast AI capabilities progress plus that lack of constraint and lack of transparency, the default is that a lot of those indicators I mentioned get worse, and none of the indicators get better, like transparency. That makes me think this is very plausible.
Yeah, I mentioned 30 years, but what about 5 years?
5 years. That's tough, isn't it? It's really tough. I think there's a risk. I wouldn't think there was a risk if it wasn't for the AI-research-causing-an-intelligence-explosion angle, but AIs are a lot better at coding and cognitive research-related tasks than they are at, for example, controlling robots and stuff.
And so, even if the FOOM ultimately comes through robots or comes through crazy levels of persuasion, you really can't rule out a scenario where AI research is automated in 3 years' time. Then, in 4 years' time, we've got superintelligent AI controlled by a few people. Maybe it's got secret loyalties. Maybe it's being deployed in the government and being overtly loyal to the president. Then, a year later, it's backsliding, political capture, or robot soldiers.
Yeah. How do you think about the badness of the outcomes here? How much does the badness depend on the ideologies of the people who are conducting the coup?
What should we look out for? Because I guess we can rank coups by badness, which is not an exercise I think we should actually attempt, but we can talk about the factors involved: what would be the worst kind of coup, and what would be a slightly better, slightly less bad kind of coup?
Yeah. So let's imagine it's 1 person who seizes power. Actually, that's the first distinction to draw. If there's a group, then even 10 people is better than 1 person.
And why is that?
Yeah. So 10 people—you get a diversity of perspectives. More moral views are represented, and there's more room for compromise between those perspectives. There's more room for reasonable positions to win out, as there's some deliberation about the actions being decided upon. There's slightly less intense selection for psychopaths than if it was just 1 person.
So, yeah, if it's just 1 person, that's bad. That's particularly bad. 10 people is still very bad. 100 people is still pretty bad, but there are big differences there. Big differences.
If we're now just thinking about 1 person, or the average person in a group, then we could think about how competent they are, and then we could say something about how virtuous their motivations are. I do think competence is important. I think it's probably underrated in most political discussions how important it is to just be really, really competent.
Thinking about something like responding to COVID, or trying to de-escalate a conflict—Russia-Ukraine, or trying to de-escalate the Israel conflict—actually just being very competent and very good at getting things done is important. And as we mentioned, if you're just willing to rely on AIs and you align those AIs in the right way, anyone could be really competent, but that's not guaranteed.
People may really want to cling to their current views without changing their mind. Let's take the example of Donald Trump. If a really smart AI system told him, “Look, tariffs are definitely bad for the US economy. They're definitely bad and won't give you what you want,” would he change his mind? I would guess no.
Lots of smart people have already been saying that. I don't actually know the economic details here, but my understanding is that most people think they're pretty bad. And it will still be the case that Trump will be able to find people telling him that what he thinks is good, and he'll be able to program his AI to keep telling him that if he wants to.
So there's no guarantee that he will become super competent, or that whoever seizes power becomes super competent.
So there's this kind of loyalty that actually undermines competence, just because you're loyal to such an extent that you're not providing feedback that's useful, because negative feedback feels bad to receive. Maybe this is a bit contrived, but do you think there's a sense in which, in singular-loyalty scenarios, the AIs could be so loyal that they're kind of undermining the competence of the person they're singularly loyal to?
Yeah, it's a really great question. I haven't thought about this, but in a way, the most extreme version of singular loyalty will just agree with whatever the dictator has said most recently, without questioning, and will do that even when it's not in that person's interests, because that's the kind of loyalty that's demanded.
There's a more sophisticated type of loyalty where you're still completely loyal, but you're also willing to challenge them when you think it's in their best interests. That's a really nice distinction. I suppose one way of thinking about competence is thinking about what kinds of loyalties the dictator would demand from their AI systems.
Another way of thinking about it is how much they would listen to the AI adviser. Even if the AI has the sophisticated type of loyalty and is trying to tell the dictator what to do, the dictator could just ignore them. You see that again with AIs: they're fairly sycophantic, but they will also challenge you sometimes. Then it's up to you whether you listen.
So that's all the competence bucket, which I think is really important. I do think there are differences between potential coup instigators on that front which could be significant. My expectation would be that lab CEOs would be more competent than heads of state. But even within lab CEOs, there are some who are more dogmatic than others, and I think that dogma would get in the way of competence.
That's competence. The other thing I mentioned was, broadly, what are your goals? What are your values, or your moral character? One thing I think is really important here is being open-minded, being willing to bring in lots of diverse perspectives into the discussion, and empowering them to really represent themselves and grow and flourish.
I think a very bad thing would be a particular person becoming a dictator and implementing their vision for society. That would be one end of the spectrum; much better would be to empower all the different ideologies and ideas to become the best versions of themselves. Then we can collectively grow and improve our understanding of how to run society.
Sometimes, when people are thinking about values, they focus on, “Okay, are you this type of utilitarian?” Or, “Oh no, I hope you're not a deontologist.” It can get very specific and finger-pointing. My view is more that we don't really know what the right answer is, and the most important thing is being pluralistic and letting a thousand flowers bloom.
So we discussed the possibility of getting to a stable state in which we've avoided an AI-enabled coup, and now we have, say, an aligned superintelligence such that the risk of a coup is very low. Do you think this is something that happens for 1 country, and then that 1 country is in control of the world to such an extent that this is not a process other countries are undergoing?
To be more concrete, for example, if the US goes through a period of risk of AI-enabled coups but manages to remain a stable democracy, is it the case that Russia or China will go through a similar period of risk of coups?
It's a great question, and it will depend on the US's posture toward the rest of the world geopolitically. It will also depend on whether the US has gained a huge military and economic advantage, like outgrowing the world or just developing powerful military technology, as we were discussing previously.
You can imagine 1 scenario where the US isn't that much more powerful than the rest of the world yet and isn't that inclined to intervene, which has been the recent trend. Then China develops similarly powerful AI a few years later, and Xi Jinping uses it to cement his control over China.
Then you have 1 AI-enabled dictatorship that is extremely robust, and you have the US, which has avoided that risk. Now they're maybe competing against each other in a Cold War II and trying to outgrow the world, or maybe they're striking deals because they recognize it's not good to compete. China just indefinitely remains a dictatorship, and that's a permanent loss for the world.
But you could also imagine a different scenario where the US is very far ahead and maybe just wants to really secure its position geopolitically. So it instigates AI-enabled coups in other nations, where it's really putting US representatives on top of those nations.
That could be through secret loyalties. It could sell systems to—let's say, sell AI systems to India—that are secretly loyal to US interests, or it could give particular politicians in India exclusive access to superintelligent AI to help them gain power. You could apply those same threat models we've discussed, but with the US pulling the strings.
Or you could have the US taking control of other nations in more traditional ways: military conquest and really leaning heavily on extracting economic value from other countries as it goes around the world. So, yeah, there's a wide range of options here.
As a final topic here, perhaps we can talk about what listeners can do if they want to help try to prevent AI-enabled coups, and specifically where to position themselves. Should they be in AI companies? Should they be in governments? Should they be in perhaps eval organizations? Where's the position with the most leverage?
Great question. I think being at a lab is a great place to be.
I talked about system integrity—robustly ensuring that AIs don't have secret loyalties or behaviors intended to deceive. That's something that companies need to implement. So if you have an interest or expertise in sleeper agents, backdoors to AI models, or cybersecurity, I think being part of a lab and helping them achieve system integrity is an amazing way to reduce this risk.
Another thing you can do at labs, if you're worried about governments or heads of state deploying loyal AIs and seizing power, is help labs develop terms of service where, when they sell their AI systems to governments, they have certain mitigations against misuse. One way to frame this is: We're using really powerful AIs, and we can't guarantee the safety of those AI systems unless we have some degree of monitoring to ensure that the AI systems aren't doing anything they shouldn't be doing. That monitoring could then be sufficient to allow for the prevention of coups, because you'll be monitoring not only for accidental misaligned AI behavior, but that will also mean you're monitoring for a bad human actor giving the AI illegal instructions.
Labs will be drawing up contracts with governments and terms of service. They will be thinking about the guardrails, if any, that they place on the systems they sell to governments. But I think there's very careful work to be done thinking through how we can structure those guardrails and explain them in a way that is very unarguable and doesn't seem like we're trying to constrain the government. It's not really legitimate for private companies to constrain the government, but I do think there's an important thing to be done here in preventing AI-enabled coups. It's another thing you could do in a government, but you could also do it in a lab—or for a think tank or a research organization that's kind of intertwined with government, like RAND. I think RAND could potentially do some of this kind of work, thinking about what should be in the terms of service between labs and governments.
Another big thing is that, for system integrity, yes, we want labs to implement it, but we also want there to be some external organization that can certify it. Currently, no external organization is working on this. METR is not working on it, Apollo is not working on it, and I don't think any evaluation organization is, either. But at one point in time, there will be a very strong demand for guarantees that the systems these companies have developed don't have sleeper agents, because if you're deploying a system in critical infrastructure, in important government roles, or in the military, you want to know that it does what it says on the tin.
I think a company can make a wise investment by developing that capability: the ability to understand how to insert secret loyalties, how to detect them, and how to assess a company's process to see whether it's sufficient.
How close would that work be to interpretability work? Isn't there overlap there—looking into what the system is actually doing, fine-grained reverse engineering of some process to see how things work? Is detecting sleeper agents a specific, constrained problem?
Yeah, it's a great question. It's a specific type of problem. You could use internal work like interpretability to detect sleeper agents, but that's not the only method you could use. You could use behavioral tests, or you could look at the training data. I think a mixture of those is useful.
It's also a specific adversarial setup. In interpretability, we're often just interested in all kinds of facts about the model's behavior and its cognition. With secret loyalties, we're specifically concerned with one red team trying to produce a secret loyalty and then a blue team trying to detect it. It's adversarial in that the blue team has a tough task. The red team will be purposely trying to hide the sleeper agent, and that will raise a whole host of new considerations and important things to dig into.
So you can imagine allowing the red team full control over all aspects of the development process. Can they create a secret loyalty? Maybe they can. Maybe it's pretty hard to test. But now you only allow them to intervene at one part of the development process. Or you allow the blue team to look at all the data used for training, or to look at all of the code that they wrote. You could get potentially quite a sophisticated understanding of the conditions under which the blue team wins and the conditions under which the red team wins.
This research doesn't need to be done in a lab. It could be done by an external organization. I think it's a big missing focus of today's technical work. Ultimately, that would inform the assessments of the labs' attempts to do system integrity. So, for technical researchers out there, I'd really highlight that possibility.
Another piece of work for the right person would be beginning to understand the existing military thinking around autonomous systems. This is already obviously a live issue for militaries. They are increasingly deploying AI. It would be nice to marry up that existing expertise with these risks about more powerful systems enabling coups and get to a consensus within that military community on basic principles, like law-following and distributed control over military systems, and figure out a military procurement process that's both practical and robustly prevents this kind of stuff.
If there's anyone listening who has a way in, I think that's potentially pretty valuable. Although there's also a risk of poisoning the well if it's done badly, so do so with some care.
Yeah. Perfect. Thanks for chatting with me, Tom. It's been great.
Yeah, real pleasure. Thanks so much, Gus.