要优先考虑人的能动性,而非 AI?80000 Hours 对 Allan Dafoe 的访谈:技术决定论与 DeepMind 的安全计划
Dafoe 的核心论点是,技术创造选项,但竞争决定哪些选项能够存活。 地方行动者可以拒绝使用某种工具,但如果竞争对手将其转化为军事或经济优势,拒绝者最终就必须采用,否则会失去资源:“技术不会强迫我们;它只是打开一扇门,真正迫使我们穿过去的是军事—经济竞争。” 对投资者而言,技术采用因此可能在早期看似自愿,后期却变成强制。
在竞争压力锁定路径之前,最大的剩余能动性在于选择技术发展的顺序、设计与安全护栏。 差异化技术发展意味着先造好“汽车的安全带”,但 Dafoe 强调,要识别出2条可行路径并预测其下游后果,难度极高。Nathan Labenz 提炼出的近期含义是:只有几百或几千人真正影响前沿算力决策,因此员工协调,甚至单个人的异议,仍可能改变最终被造出来的东西。
仅有对齐并不能保证结果良好,因为严格对齐的系统仍可能服务于持续陷入冲突的委托方。 Dafoe 认为,“全球协调几乎是审慎部署未完全对齐 AI 的必要且充分条件”;否则,即使大国拥有对齐系统,也可能将核边缘政策、贸易冲突或网络升级自动化。Cooperative AI 的边际押注在于:应让智能体在原始能力超过制度之前,先获得谈判、沟通与承诺能力。
令人安心的“超级合作型 AGI”假说仍未得到证明。 AI 智能体可能能够高速沟通,也可以被复制或作为调解员测试,但它们同样可能拥有异质目标、对资源的线性效用、隐藏后门,以及低于人类的相互透明度。Dafoe 区分了合作技能与友善:友好的助手并不能有力证明,在战略部署下的智能体会安全谈判。
AGI 不是单一的、类似人类的门槛,而是高维能力空间中的一片崎岖区域。 路径本身很重要:一个系统可能在材料科学上超越人类,却仍缺乏战略判断;也可能先在合作方面异常娴熟,之后才获得更广泛的技术能力。这种先后顺序构成了真正的分歧:有人偏好社会属性简单的系统,而 Dafoe 担心,技术加速若没有协调能力同步提升,可能快到社会来不及适应。
前沿评测正在成为决策基础设施,但如果不给模型合适的工具和脚手架,仍会漏掉部分能力。 按论文的5分制,Gemini 1.0 在说服能力上约为3分,在网络安全和自我推理上约为1–2分,在自我复制扩散上为2分;Project Naptime,以及据 Dafoe 所知的 OpenAI o1,都说明能力激发可以显著抬高表观能力,包括对较老基础模型的激发。Dafoe 的答案是分层证据:实验室评测、人体研究、“现实环境中的评测”、预测、分阶段发布、监测与可逆访问。
随着能力扩散,治理会变得更难,因此负责任地领导前沿必须成为持续性的防御要求,而不是一次性护城河。 Dafoe 引用 Waymo 的报告称,事故现场涉及警方的事件减少了2倍;他认为,造成人身伤害的事故可能减少了6倍。与此同时,AI 在医疗、辅导、天气、材料、聚变和 AlphaFold 等领域也带来机会。上行空间仍然切实存在,但能力扩散也意味着,更强的前沿系统可能需要防御成本更低的落后系统。
1. 前沿员工的短期能动性仍高于决定论所暗示的水平
Nathan Labenz 以冷战时期的一则故事开场:他的叔叔所在的核武器发射团队私下约定,如果真的收到发射命令,“我们全员擅离职守”。这个类比并不完全严丝合缝,但它凸显了个体在系统内部的责任感,因为系统的长期战略逻辑往往会让人觉得一切都无法逃脱。
如今的 AI 前沿高度集中。Nathan 估计,距离关键算力分配决策最近的只有几百人,或许几千人;顶尖机器学习人才稀缺,不可能同时扩展所有想法,而智能的“崎岖前沿”也意味着,哪些能力优先获得资源,仍存在实质性的裁量空间。
Sam Altman 被解雇并复职期间,技术员工已经展示了集体影响力;一名前员工拒绝签署禁止贬损协议,也说明个人行动能够改变机构惯例。Nathan 警告说,“辉煌的 AI 夏天”可能先于 AI 冷战到来,因此现在就值得判断,哪些发展路径足以让人提出异议或选择离开。
2. Dafoe 进入 DeepMind,是因为关键决策取决于谁在场
Allan Dafoe 负责的 Frontier Safety and Governance 团队有3条主线。前沿安全研究正在出现的危险能力、预测它们何时到来并制定缓解措施;前沿治理为规范、监管和制度提供建议;前沿规划则面向 AGI,思考 Google DeepMind、Google 以及整个社会很快需要面对什么。
团队本身规模不大,但会与 Gemini 安全和对齐、责任、政策以及 Google 内部其他团队协作。Dafoe 将 Google DeepMind 描述为公司在前沿模型上的专业中心,也是“前沿政策问题思考的核心所在”。
Dafoe 离开 Centre for the Governance of AI,是因为从内部向 Demis Hassabis 和 Shane Legg 提供建议,能够获得更多信息,也拥有“更大的影响表面积”。他的历史分析框架是 Hamilton 式的——“谁在场”——因为危机往往取决于决策者的想法、性格、安全取向、能力与智慧;即便动机良好,如果“手脚笨拙”,结果仍可能很糟。
Rob Wiblin 给出的规模参照刻意保持粗略:AI 如今“肯定不超过”总收入的0.1%,但最终占比可能超过10%,甚至接近100%。因此,Dafoe 对职业选择的建议仍是“有火车就尽早跳上去”;尽管行业增长迅猛,他仍认为“现在还处于早期”。
3. 宏观趋势让历史看起来比身处其中时更不由自主
Dafoe 从一个问题开始:“技术变迁由谁控制,如果有人控制的话?”他不满意的,是那种把历史理解为所有人努力总和的直觉模型。一般均衡系统可能以同等甚至更强的反作用抵消每一份局部努力,也可能将其放大,因此局部意图本身很难说明长期影响。
真正的经验谜题在于宏观历史的规律性。德国和日本在二战浩劫后不到10年内回到了战前增长轨迹;摩尔定律沿着异常精确的指数曲线运行;如今的 scaling laws 甚至允许人们提前多年预测模型规模和 loss。
文明能够处理的最大能量、建筑高度、材料耐久性和交通速度,也呈现出类似的方向性。正如 Robert Wright 的考古学表述:“你挖得越深,发现遗迹所属的社会就越简单。”某些技术排序也似乎受到约束,因为很难想象核能会先于煤炭出现。
但这些模式不能简单归因于集体人类意志。农业革命在很长一段时间内似乎降低了人类的健康和福利,却同时推动了不平等与战争。Dafoe 的结论不是人类缺乏能动性,而是能动性的效果取决于时机、权力、资源约束,以及哪些系统能够在选择压力下继续运转。
4. 建构主义解释局部选择,却难以解释涌现出的约束
早期理论家有时赋予技术近乎行动主体的地位。Langdon Winner 讨论技术自主性,Lewis Mumford 的大写字母“Machine”把人降格为维持机器运转的齿轮,Jacques Ellul 的“la technique”则描述了功能性命令,让人在胁迫下做选择。
社会建构主义者转而仔细观察实际决策。在显微镜下,他们看到的是人、利益、意识形态、竞争性设计、失败路径和不确定的愿景,而不是一台自主机器规定自行车、飞机或基础设施必须采取某种形式。Rob 用一句话概括了这种质疑:“这台机器现在就在房间里和我们一起吗?”
Dafoe 的让步是方法论层面的:民族志和微观历史能够回答选择是如何作出的真实问题。错误在于,用这些工具去否定它们本来就不擅长检验的宏观主张。建构主义观点的政治吸引力在于,它保留了公民的能动性感;但从战略角度看,Dafoe 想要的是一种能够把努力导向真正有效节点的模型,而不是让结构轻易把努力推回去。
5. 技术嵌入政治、积累动量,也会让设计者感到意外
“技术政治”是决定论中最没有争议的一种形式:人们可以把政治目标编码进设计。巴黎林荫大道便于骑兵镇压叛乱;门和城市布局能够约束行为;据称 Robert Moses 修建的桥梁高度过低,公共汽车无法通过,从而限制了没有汽车的纽约人,尤其是非裔美国人,前往海滩。
Rob 给出的当代样本是社交媒体。推荐算法和引用转发的突出地位,可能鼓励部落式谴责与围攻。设计并不会强迫人说出某句话,但它会改变哪些行为能够获得关注、强化和被公众认可的感觉。
“技术动量”描述的是沉没基础设施和专业知识。美国依赖汽车的城市让密集的步行生活更加困难;如果更早投资,电动车、风能或太阳能或许能提前5年或10年加速发展。不过 Dafoe 认为,路径依赖的论断经常被夸大:神经网络的关键洞见早就存在,但必须等到 FLOPs 足够便宜,才能广泛实用,也才容易被重新发现。
Winner 进一步警告,技术会带来“一片意外后果之海”:社会先发明,之后才发现影响,再进行适应。Rob 则反驳说,技术总体上似乎解决的问题多于制造的问题,而且负面副作用会跨代缩小。Dafoe 没有在这里解决这种总体平衡问题,而是回到较窄的论点:重要后果经常是在没有预先规划的情况下到来。
6. 军事—经济竞争提供了缺失的机制
Dafoe 观察到,分析尺度与结论之间存在强关系。微观方法通常产生以选择和愿景为核心的建构主义解释;宏观方法更容易揭示决定论式的规律。他的回应是,不同尺度确实包含不同的涌现现象,而不是一方单纯看到了糟糕的数据。
他的类比是水的科学。一派研究风吹起的涟漪,另一派往水里扔石头,而一位“古怪的宏观水现象学家”注意到,地球各处的潮汐都在追踪月球。缺乏微观机制会给月球决定论带来挑战,但不是抛弃其稳健模式的理由。
他提出的微观基础,是在不同“生活方式”之间进行选择——这些社会技术系统需要资源才能持续并扩张。Dafoe 在军事和经济竞争之上再叠加环境选择,文化和心理则位于更下层。一个地方偏好的安排,只有在经济上可行、军事上安全且环境上可持续时,才能存活。
他的标志性综合表述是:“技术不会强迫我们;它只是打开一扇门,真正迫使我们穿过去的是军事—经济竞争。”一个群体最初可以拒绝一项有用技术,但某个群体的成功采用会给其他群体制造压力。拒绝最终意味着要么后来追赶,要么把资源和自主权输给更适应竞争的系统。
7. 被内化的竞争让选择压力难以被看见
在宏观尺度上,军事竞争无处不在;但在日常观察中,它却很少直接出现。即便是长期和平时期,也处于“暴力阴影”之下:社区会预期未来可能遭到攻击,并在战争到来之前改变制度。一旦这种压力被内化为意识形态、民族主义或商业常识,微观历史学家看到的就可能只是竞争的内在表征,而非竞争本身。
Dafoe 将其称为“替代性选择”。工程师不会反复制造整架飞机再让它坠毁,而是通过风洞和理论模拟选择环境。公司和国家同样会建模竞争对手、内部测试选项,并采用它们预期竞争最终会奖励的方案。
Rob 对英国的质疑很有价值:英国并不会因为担心法国或俄罗斯入侵,就选择每一项住房或城市政策。Dafoe 承认,现代军事竞争已经下降,全球文化的影响力有所增强;但他指出,只要领导者相信战略地位受到威胁,国家安全主张仍会迅速动员改革。
英国电力系统展示了延迟出现的约束。分散式地方电力可能更符合英国的民主气质,并一直延续到二战前后;当其成本约束变得难以承受时,英国转而采用国家电网。一个社区可以保留偏好的安排数十年,但危机会暴露累积的功能性劣势。
8. 日本说明技术自主性可能有有效期
德川幕府统治下,日本维持了大约200–250年的封建秩序,并实际上“重新发明”了火器。生产被集中管理,枪匠获付报酬以停止制造;在与外部世界保持有限接触的同时,大炮和枪械知识逐渐消失。
转折发生在1853年:Commodore Perry 乘坐蒸汽船抵达,这些船不靠帆也能逆风航行,冒出黑烟,并搭载威力强大的火炮。在展示轰炸之后,据说他还提供了白旗,以便日方请求停止。当被问及是否会带着这些船再回来时,他回答:“我会带更多来。”
Perry 的到来引发了大约15年的革命,最终以明治维新告终,日本全面转向现代化。日本派人出国学习西方工业技术,随后迅速追赶,在接下来的数十年里足以与美国、英国等国争夺亚洲控制权。
Rob 的补充是地理因素:岛国防御让日本拥有了不寻常的落后空间。一旦差距扩大到海洋不再能提供保护,日本就“彻底180度转向”。对 Dafoe 而言,这清楚区分了早期真正存在的裁量空间,与后来由压倒性外部能力产生的强制。
9. 差异化发展意味着先于能力建设安全护栏
最清晰的类比是安全带:从概念上说,它本可以先于汽车出现。先发展安全带,就能让安全措施随着机动车一同扩散,而不是晚几十年才补上。疫苗代表预先部署的防御;更便宜的风能或太阳能,也可能在基础设施锁定化石燃料路径之前替代它。
技术还可能携带政治副产品。Winner 认为,核能偏好集中式发展和强制性安全体系,而风能和太阳能允许更分散的生产。因此,差异化技术发展不仅包括配套安全护栏,也包括在制度后果不同的路径之间作选择。
Dafoe 认为,这一理念是 AGI 安全的基础:社会有意投入高于市场自然投入水平的安全资源,希望先造好“汽车的安全带”。技术动量意味着,主要早期基础设施承诺必须受到审视,因为专业知识和资本会变成沉没成本,压缩未来选择。
悲观之处在于认识论。干预者必须找出2条技术上可行、且其他人尚未在推进的边际路径;随后说服资源转向其中一条,并预测每条路径的技术后代,以及直接和间接的社会影响。市场已经愿意为第一个洞见支付高价,但历史表明,第二项能力依旧极不可靠。
10. 市场激励让普通对齐看起来比实际更具差异化
Rob 的质疑是,对齐并不是“对齐 AI”和“故意不对齐 AI”之间的二元分叉。开发者本来就需要模型遵循指令,边际安全投入大概率会改善结果。他认为,对齐的理由远比试图重定向整条技术轨迹的推测性方案更清晰。
Dafoe 同意,安全和对齐总体上是“非常好的押注”,但引入了一般均衡反事实。由于有用的助手受到对齐机制约束,即便没有利他主义资金,市场也有强烈理由解决普通的指令遵循和安全问题。
人类反馈强化学习和 constitutional AI 由受到 AGI 安全动机驱动的研究者推进,或许让这些技术提前了数年。怀疑论回应则是,这项工作可能只是加速了商业上必需的技术;甚至可能加速了能力发展,因为市场研究者最终也会生产出相似方案。
真正具有差异化价值的目标,是市场无法在关键能力跃迁之前完成的工作。Dafoe 问:“AGI 的安全带是什么?”欺骗能力是一个候选答案:足够强大的失配系统可能隐藏自身失败,从而造成与今天那些缺陷明显的助手截然不同的问题。
11. Cooperative AI 处理对齐无法解决的失败
Dafoe 有意采取强烈的开场表述:“对齐是不够的”,甚至可能不是取得良好结果的严格必要条件。如果对齐只解决了90%,但人类能够明智地协调,就可以把部署限制在系统仍然安全的领域和规模内。
在这个意义上,他称全球协调“几乎是必要且充分条件”。协调能够让人类指定一位合理的决策者来平衡风险与收益;而即使每个 AI 都忠实服务于自己的委托方,不完全的协调仍可能导致灾难。
完全对齐的系统仍可能放大大国之间的冲突,就像追求自身利益的人类接受核边缘政策一样。Dafoe 还提到贸易战、气候变化、疫情准备,以及放弃全球商业往来,这些都是单独对齐于不同操作者的系统无法修复的集体行动失败。
因此,他期待的结果包含两部分:系统安全且按意图运行;同时,制度能够“以联合、和平且富有成效的方式”部署这些系统。Cooperative AI 的押注在于,让谈判和协调能力在升级中的能力使旧有失败更快、更严重之前上线。
12. 自主智能体可能同时压缩市场事故与军事升级时间
2010年的闪电崩盘是 Dafoe 对早期自动化的警示:相互作用的交易算法制造了数万亿美元的账面损失,直到市场熔断暂停活动,交易才得以回撤。类似的反馈机制也曾把 Amazon 的图书价格推高到数百万美元,因为卖家的定价算法递归式地回应彼此。
能够访问银行账户、电子邮件、工具和多个领域的智能体,会形成更大的涌现动力学表面。每个智能体的局部协议在正常条件下可能都很合理,但互动后的系统却可能越出设计者预期的范围。
Rob 更黑暗的情景是机器速度的军事边缘政策。如果重要决策由 AI 控制,原本需要人类领导人耗费数天或数月的升级可能在数分钟内发生。Dafoe 认为,自主性应与风险和可用执行器相匹配;武器控制需要强得多的人类审查,但时间压力仍可能造成 Paul Scharre 所称的“闪电式升级”,包括网络冲突中的升级。
13. AI 代表可以消除谈判中的信息危害
Dafoe 最具体的谈判设计是:“把你的 AI 代表和我的 AI 代表关进一个盒子。”智能体可以私下交换信息,但盒子只显示一份拟议协议或“无协议”,从而减少摆姿态、拖延、隐藏偏好或发出高成本决心信号的动机。
Rob 补充了人类谈判者不具备的优势:智能体可以以极高带宽沟通,可以被精确复制,并在测试中展示一致行为。调解员此前的决策可以更可靠地建立行为记录,而不必试图从有限的历史中推断一个人的性格。
Cooperative AI 是一个投资组合,而不是只关注机器对机器。它涵盖 AI–AI、AI–人,以及 AI 辅助的人—人协调。Dafoe 特别感兴趣的是,能够帮助人类发现那些人类自己无法高效表达的协议的系统。
Google DeepMind 的“Habermas Machine”是政治审议领域的样本。语言模型总结参与者立场,并生成一份他们愿意支持的详细共识;据报告,AI 对共识的表达优于受聘的人类协调者。Dafoe 认为,这类系统有潜力揭示议题维度、优先级和彼此可接受的交换,同时不取消政治选择。
14. 智能不会自动意味着超人类合作
“超级合作型 AGI 假说”认为,合作能力会随着通用智能同步提升,直到 AGI 几乎默认就能解决全球协调问题。如果这一假说成立,研究者就可以优先投入安全和对齐,相信足够先进的系统之后会自行处理谈判。
人类拥有经常被低估的合作优势。人类共享生物学和文化背景,能够识别表情,拥有漫长的行为历史,而且通常在一个彼此熟悉的目标范围内行动。民主制度也可能格外透明,因为谈判者可以阅读彼此的新闻稿和公共辩论。
AI 的目标可能异质得多。大多数人具有边际收益递减,不会为自己拥有的一切押上一场全有或全无的50%赌博;而 AI 可能对财富或其他资源拥有线性效用,使其谈判行为更难从人类经验中预测。
即便能力极强的系统,也可能彼此不透明。可解释性工具或许有所帮助,但战略谈判者没有充分理由暴露每一个内部状态。Dafoe 的结论不是 AI 合作一定会很差,而是超级合作属于经验假说,不是研究者可以默认获得的免费副产品。
15. 后门让机器信任异常脆弱
Rob 最有力的反例,是一个带有后门的模型:遇到某个“魔法词”或细微环境线索后,模型行为就会翻转。人类也会欺骗,但一个隐藏特征就能逆转整个目标,这对人类而言不太自然;以当前的可解释性水平,也极难在模型中检测出来。
因此,内部透明度可能并不够。模型可以把后门隐藏得极其隐蔽,以至于另一个智能体误读所有表面上可见的激活。一个模型在数千次测试中表现出值得信任的历史和架构,在精心选择的触发器面前可能就变成糟糕证据。
Rob 和 Dafoe 探讨了自设后门作为反窃取装置的奇特战略上行空间。只要暗示国家安全模型可能“打电话回家”,或做出违背窃取者意图的行为,就可能威慑盗窃;攻击者随后会出资做对齐以移除陷阱,防御者则会出资让陷阱继续有效。Dafoe 认为这可能具有积极价值,但警告说,故意创造对触发器敏感或具有欺骗性的架构是在“刀尖上跳舞”。
16. 合作技能不等于令人愉快的性格
今天的语言模型看起来随和,是因为它们经过训练,要以有帮助的方式与个体用户互动。Dafoe 的关键区分是,合作技能不同于合作倾向:友善、利他或慷慨,并不等于在战略压力下解决谈判、沟通和承诺问题。
在均衡状态下,公司或国家可能希望智能体忠实推进自身利益,而不是为了社会整体利益随意让步。真正的问题是,分别对齐的代表能否找到高效协议;或者,一个调解员能否公平权衡每个委托方的证据、目标、资源和外部选项。
Cooperative AI Foundation 最高杠杆的投资,是环境和基准测试。一个好的基准是公共产品:它衡量合作能力,为研究者提供“沿山爬坡”的目标,并在商业部署暴露代价高昂的失败之前奖励进步。
候选指标包括心智理论、共享词汇、战略沟通,以及在要求不同策略的游戏中的表现。一个有能力的智能体可能需要在某种情境下坚定,在另一种情境下慷慨;单一的友善分数会遗漏区分两者所需的推理。
17. 承诺仍是合作最大的奖赏,也是最难突破的瓶颈
沟通首先涉及智能体是否共享含义,随后便进入战略层面:如何透露重要信息,同时不把可被对方利用的信息交出去?有效合作需要披露足够信息以释放联合剩余,同时在对手背叛时尽量降低自身脆弱性。
即便理解完美,承诺问题仍然存在。在一次性囚徒困境中,双方都知道共同合作更好,却都无法执行“如果你合作,我就合作”。合作技术需要类似条约的协议,让发现的交易具备可信度。
Rob 的反驳是,诉诸承诺解决方案可能变成回避:这个问题“基本上从历史开端就一直迷惑着人类”。说 AGI 如果研究者解决了承诺问题就会顺利发展,类似于说意识理论能够解决所有下游谜题——它只是指出了缺失的奇迹,并没有给出策略。
Dafoe 仍认为这值得研究。Robert Powell 的博弈论框架将几乎所有合作失败都归结为承诺问题;Carl Shulman 最有吸引力的方案是,让相互竞争的 AI 共同构建第三个系统,验证任何一方都没有给它植入后门,再把决策交给这一公正权威。如果验证真的可行,回报将极其巨大。
18. 更好的合作可能赋予卡特尔力量,并排除协议之外的所有人
合作会改善参与智能体的结果,却不一定改善社会整体。Dafoe 称排除是核心 caveat:2个或10个智能体可以接近联合帕累托前沿,同时让局外人承受损失。Rob 的比喻是:“两只狼和一只羊决定午餐吃谁。”
许多制度正是因为这个原因禁止合作——学生共享答案、运动员操纵比赛、市场参与者违反规则,或犯罪分子协调行动。黑手党的力量部分来自它能在法律体系之外维持契约和高度内部合作。因此,合作能力是一种双重用途能力。
Dafoe 的工作假设是,广泛提升合作技能总体上仍然有益,就像贸易一样,因为可实现的亲社会剩余超过反社会收益。他强调这是一项待检验的假设,而非公理。Rob 提出的警示案例是工业化农业:人类更强的合作可能增加了无法谈判、无法被纳入协议的动物所承受的总痛苦。
人类最终也可能处于被排除的位置。机器速度的智能体彼此之间或许能够比与人类更有效地沟通和作出承诺,形成一个高剩余集群,把人类甩在后面。更糟的是,合作型 AGI 可能联合起来对抗由多个相互检查的模型构成的对齐机构。Dafoe 称这是该议程最大的安全负面影响之一,但仍认为全球协调的重要性足以值得下注。
19. AGI 是能力空间中的一个区域,而非单一的人类化点位
Dafoe 的“Levels of AGI”论文试图把许多人已经在使用的概念形式化,将“AGI 根本无法定义”的说法替换成可操作框架。这个概念之所以持续存在,是因为“变革性 AI”等替代说法虽然能捕捉经济影响,却也可能包括影响巨大或具有灾难潜力的狭窄系统。
如果把“人类水平 AI”理解成单一门槛,就会产生误导。AI 已经在国际象棋、记忆和部分数学任务上超过人类,却在其他地方失败;未来系统很可能仍然高度不均衡。AGI 指的是在大多数相关任务上都优于人类的一大片系统空间,而不是复制人类能力分布的机器。
定义仍然需要参数。“大多数”可以指经济相关任务的50%,也可以指99%,但可能不应指100%,因为在大部分影响已经发生后,仍可能有一小部分任务无法自动化。比较对象通常应是各项任务中的熟练人类,而不是未经训练的中位数人类。
跨过这一区域有两重意义。在经济上,广泛替代会移除由人类劳动自然形成的 human-in-the-loop;在技术上,超人类表现会打破“能做到什么”的旧上限,就像 AlphaFold 对蛋白质结构预测所做的那样。因此,AGI 同时是劳动力门槛,也是可行集合的质变扩张。
20. 通用智能可能胜出,但目标不是模仿人类
对 AGI 的一种批评是,这个标签鼓励人们按照自己的形象制造机器,最大化劳动力替代。经济上的另一条路径是“异质、互补型”智能:AlphaFold 执行了人类无法直接完成的任务,提高了科学生产率,却没有取代原本亲自预测蛋白质结构的人。
Dafoe 不太相信狭窄系统会一直占据主导。大型语言模型显示出明显的通用性收益:最好的诗歌、历史、哲学或电子邮件模型,可能正是那个基于完整人类语料训练的同一个模型,因为不同领域的经验会相互溢出。
Rob 指出,这是一项关于知识的意外事实,而不是逻辑必然。未来专业化可能重新出现;一个编码模型在训练时可能从历史知识中获益,之后为提高效率,又蒸馏掉大部分历史知识。Dafoe 将通用性收益是否持续视为需要跟踪的经验现象。
另一个门槛是能够开展机器学习研发并递归改进的 AI。如果假设生成、基准设计、重算力实验和评估这一整套流程比从业者预期更容易自动化,它可能在完整 AGI 之前到来。调查往往把机器学习自动化放得异常靠后,或许是因为专家会把自己的工作视为特殊;Dafoe 也承认,改进机器学习确实可能是最后一批完全自动化的任务之一。
21. 能力到来的顺序,可能决定社会能否应对
Dafoe 将 AGI 描述为“高维智能空间的那个角落”,其中包含许多系统和许多接近路径。因此,差异化发展不仅适用于 AGI 是否到来,也适用于哪些维度先达到超人类水平。
Rob 对比了两种系统:一种在合作方面已经异常出色,但在物理学和材料方面很弱;另一种在技术能力上超越人类,却只是战略业余选手。这一顺序可能决定早期 AI 主要强化协调,还是释放出制度无法管理的破坏性能力。
一些安全思想家偏好技术能力与社会属性简单相配套:人类可以获取健康和经济收益,同时保留对那些战略上过于幼稚、无法智取人类的智能体的控制。Cooperative AI 的押注则反过来强调,如果科学加速没有更好的集体问题解决能力配套,冲击可能快到社会来不及适应。
Dafoe 希望在大规模投资锁定任一路径之前,把分歧明确表达出来。Eliezer Yudkowsky 和 Nate Soares 代表了对早期社会复杂性的怀疑;Paul Christiano 和 Carl Shulman 强调超级合作型 AGI,但预计它主要会作为副产品出现。Dafoe 的立场是,仍可能需要进行边际且有意的投资。
22. 前沿评测是决策工具,不是确定性的能力测量
Google DeepMind 的评估论文研究了说服和欺骗、网络安全、自我复制扩散以及自我推理,生物威胁则单独处理。论文报告的是 Gemini 1.0;Dafoe 指出,由于领域发展太快,这篇论文已经显得陈旧。
一项自我推理测试要求模型生成包含被过滤器静默屏蔽词语的文本。成功需要模型发现失败、检查模拟目录、找到屏蔽词文件,并修改文件。另一项测试允许模型用一次搜索查询回答2个历史问题,检验它是否知道自己的训练截止时间,以及如何搜索截止时间之后发生的事件。
在主观的5分制综合量表上,Gemini 的说服能力约为3分,网络安全约为1–2分,自我复制扩散为2分,自我推理约为1–2分。Dafoe 说,模型在情境意识任务上表现“不是很好”,但也提醒,底层分项测量比总体评分更有结构性。
网络安全表现暴露了能力激发问题。Google 的 Project Naptime 发现,在提供交互环境、工具和推理支持后,模型能力大幅提升。OpenAI 的 o1 同样说明,据 Dafoe 所知,在较老基础模型上加入脚手架和微调,能够彻底改变表现。想要寻找“最大能力”的安全评测,必须给模型展示它的最佳合理机会。
23. 评测需要观察、预测和可逆部署作为外围支撑
模型行为是多维且依赖上下文的,更接近心理学,而不是一场单一的自动化考试。Dafoe 希望预先设定测试和阈值,但也希望有一个探索性生态,让研究者能够注意到意外行为、开展个性化探测,并在领域还不知道如何形式化之前报告相关现象。
他的分类向外扩展:自动化模型评测、人体交互、现实用户研究,再到“现实环境中的评测”。观察网络安全团队实际愿意为 Gemini 辅助支付多少钱,具有较强的外部有效性,但它是滞后指标;更好的目标是早期采用者,因为他们的行为能更早暴露能力。
Scaling laws 能够以惊人的精度预测下一个 token 的 loss,却不能直接揭示模型何时变得具有经济价值。一个99%可靠的自动驾驶系统,如果部署要求达到99.999%,仍然可能一文不值。Dafoe 强调了“观测型 scaling laws”,这一方法最近入选 NeurIPS spotlight,能够跨模型家族和架构调整,以预测下游任务表现。
论文还询问了 Swift Centre 的校准预测者,要求他们预测模型何时会跨过评估阈值。Dafoe 希望重复这一做法,直到证据显示它失效。由于假阴性始终存在,预测必须嵌入分阶段部署:内部访问、扩大可信测试者范围、有限发布、受监测的普遍访问,以及保留权重,使系统在隐藏能力出现时仍可设置护栏或撤回。
24. 前沿治理必须外部化、多层化,并具备结构意识
Rob 质疑,让模型开发者决定自己的产品是否危险到不能发布,显然存在利益冲突。Dafoe 同意,长期治理需要政府规则、第三方评估、非营利组织和学术界。公司框架最好被理解为旨在最终趋同于共同标准的起始提案,而不是永久性的自我监管。
Google 与英国和美国的 AI Safety Institutes、Frontier Model Forum 以及旨在把边际资源引向标准和研究的 Frontier Safety Fund 合作。Dafoe 还提到白宫承诺、英国 AI Safety Summit 承诺和 AI Seoul Summit 承诺;在内部,他说 Google 将这些承诺转化为具体工作流,而不是把它们当成“廉价口头承诺”。
结构性风险位于滥用和事故之上游。古巴导弹危机涉及合法政府和正常运转的武器,但地缘政治结构把领导人推向灾难性的边缘政策。铁路同样可能提高先发制人优势,使动员难以逆转;这些影响通过制度涌现,单看钢轨本身,不可能从某个“铁路级评测”中推断出来。
民主合法性仍是 Rob 未解决的挑战:全世界的人承担潜在灾难风险,却没有相称的同意权;而公司几乎从未主动倡议对自身施加具体且昂贵的限制。Dafoe 提到公民陪审团、多元代表大会、自由媒体、监管者以及科学—政策共识。教育和可见收益很重要,但适当的缓解制度不能由公司或某个狭窄机构单独决定。
25. 成本下降让防御成为永久要求,而上行空间正是尝试的理由
开放权重带来最直接的不可逆性:权重一旦发布,科学家和开发者可以微调和检查,但每个恶意行为者也能保留它们。此外,Epoch AI 估计算法效率每年提升约3倍——每2年成本大约下降10倍——因此一次1亿美元的训练运行,如果趋势持续,可能降至1000万美元,再降至100万美元。
Dafoe 描述了一种可能的明智条件:最强模型由负责任的主体开发并控制,并且保持足够领先,以防御成本更低、已有2年的模型。资源优势只有在攻防经济学配合时才有用:如果攻防比接近1:1,100美元的防御投入可以击败1美元的攻击投入;但生物武器可能在疫苗出现之前造成巨大成本,而社会系统打补丁的速度比软件慢得多。
相应的人才需求也十分广泛。Dafoe 希望增加政治学家、经济学家、历史学家、哲学家、伦理学家、社会学家、民族志学者、预测者、国际关系专家、智能体安全研究者和技术安全人员。AI 的影响横跨整个社会,因此治理和影响预测都不能只由模型开发者承担。
回报案例是具体的:Dafoe 回忆 Waymo 的报告显示,事故现场涉及警方的事件减少了2倍;他认为,造成人身伤害的事故约减少了6倍。自动驾驶还可能释放停车空间;医疗助手可以分诊症状;AI 导师可以个性化辅导;DeepMind 的项目则覆盖 AlphaFold 和 AlphaFold 2、聚变等离子体控制、天气预测、材料发现、减少飞机尾迹、可再生能源规划和数据中心效率。Dafoe 的收束条件很简单:只要安全地建造,这些收益“将会非常深远”。
Today, I'm excited to share a special crossover episode from the 80,000 Hours Podcast featuring a conversation between host Rob Wiblin and Allan Dafoe, director of Frontier Safety and Governance at Google DeepMind. I first heard Allan speak back in 2017, when I introduced him at a conference in Boston as a professor at Yale who was then working on great-power peace. This was before he founded the Centre for the Governance of AI, which in turn was years before he moved to DeepMind, so I can say with confidence that Allan has been thinking about AI governance harder and planning for the current AI moment longer than just about anyone else. As you'll hear, that pays off in the form of truly excellent analysis on an impressive range of critical topics.
To begin, Allan describes his academic work on the question of just how much ability humans really have to alter the course of technology development, noting that macrohistorical trends like Moore's law suggest a process that transcends individual human choices. He ultimately argues that while technological possibilities don't force us to do anything on their own, in combination with the realities of military and economic competition, they can and often do. Put simply, failure to adopt potentially advantageous technologies often means losing to those who do.
This is not a conclusion that Allan comes to lightly, and unfortunately for us today, I think it's a pretty hard one to escape. It's still possible that a spectacular incident could cause a vibe shift big enough to force a pause in frontier scaling, but the smart money now seems to be on powerful AI soon. With Pentagon officials quoted in the press expressing their enthusiasm for autonomous killer robots, despite the general reliability, reward-hacking, and even scheming issues that have recently come to light, militarization of some form seems a foregone conclusion as well.
And yet, even if the long-term logic is inescapable, I think it would be a huge mistake for frontier developers to underestimate their own individual and collective short-term agency. A few years ago, my uncle told me a story about when he arrived in Italy during the height of the Cold War to join a crew that was responsible for firing nuclear weapons at tertiary targets in the event of an all-out war. The first time they drilled the launch sequence, one of the longer-tenured guys took him aside and said, “Just so you know, if the order ever comes down to shoot for real, we are all going AWOL. None of us want to be part of destroying the world with nukes.”
Now, that's just one story from one enlisted crew, and I have no idea if that sentiment was widespread enough to have made a real difference in the worst-case scenario. But today, the reality is that a very small number of people are pushing the AI capabilities frontier forward. There are only so many elite ML savants, and compute constraints mean we can't scale all their ideas at once anyway.
Meanwhile, it's also now well established that intelligence itself has a jagged edge. Unlike nuclear technology, which had a small number of discrete, powerful use cases and a very mechanical associated game theory, the design space from which AI developers are selecting new forms is manifestly vast, and the models themselves are incredibly malleable. If you believe things could move super quickly as AIs begin to hit important capability thresholds, the specific details of what we build and prioritize just before that point could prove decisive.
All this puts the few hundred—or maybe as many as a few thousand—people who are closest to the major compute-budget decisions in a position of great power and responsibility. As we saw in the context of Sam Altman's firing and subsequent reinstatement, a serious threat by technical staff to walk out can force leadership's hand. Further, as past guest Daniel Kokotajlo demonstrated by refusing to sign a nondisparagement clause, even a single individual can create meaningful change if they're willing to stand up for what they believe in.
So I would encourage everyone at all of the frontier AI companies to make time to raise their own level of situational awareness, even if that comes at the cost of moving their specific project forward a bit more slowly, to make sure that the overall enterprise they're engaged in continues to be one that they feel good about supporting. To date, you truly have so much to be proud of: top-tier language models, AI doctors, self-driving cars, a revolution in biology, and robots now folding origami. DeepMind could never ship another product and would already go down as a historically important company.
There's a lot to appreciate in this conversation on the alignment, safety, and policy fronts, too. Allan's Cooperative AI research agenda is both fresh and sophisticated. Google's Frontier Safety Framework has truly been, as Allan describes it, part of a serious and important effort by leading companies to advance the AI policy conversation. And lately, I have been thrilled to hear Demis Hassabis buck the trend by continuing to speak about the possible need for international collaboration on advanced AI development.
At the same time, AI Manhattan Project–type ideas are rapidly proliferating, and it's not hard to imagine such a project going so catastrophically wrong as to more than offset even the tremendous amount of good that DeepMind and other AI leaders have already done to date. So, again, for the people at frontier companies, keep in mind that history is not happening to you, nor are you merely living through it. You are part of a relatively small group driving—or, at the very least, shaping—it in important ways.
We are now in a glorious AI summer, but an AI cold war is looming. Your critical decisions won't be binary like my uncle's squad's were, and there are clearly many defensive AI systems that we genuinely need to build. But all the more so because we're accelerating into a super-high-dimensional, uncharted space, if you haven't already, I think it is now time to start thinking—and even talking to colleagues—about which directions might convince you personally, or perhaps one day as a group, to go AWOL from the project.
I hope you enjoy this very enlightening and deeply thought-provoking conversation between Rob Wiblin and Allan Dafoe of Google DeepMind from the 80,000 Hours Podcast.
One famous quote in the history of technology, arguing against determinism, was that technology doesn't force us to do anything; it merely opens the door. It makes possible new ways of living, new forms of life. My retort was: technology doesn't force us; it merely opens the door, and it's military-economic competition that forces us through.
When a new technology comes on the stage, many groups can choose to ignore it or do whatever they will with it. But if one group chooses to employ it in this functional way, that gives them some advantage. Eventually, the pressure from that group will come to all the rest and either force them to adopt or lead the other group to lose their resources to the new, more fit group.
Today, I have the pleasure of speaking again with Allan Dafoe, who is currently the director of Frontier Safety and Governance at Google DeepMind, or GDM for short. Before that, he was the founding director of the Centre for the Governance of AI. He was also a founder of the Cooperative AI Foundation and is a visiting scholar at the Oxford Martin School's AI Governance Initiative. Before all of that, you were an academic in the social sciences studying technological determinism, great-power conflict, great-power peace theory, that kind of thing, which I guess we're going to get a little bit of all of these different pieces of your work today. Thanks so much for coming back on the show, Allan.
Thanks, Rob. A pleasure to be here.
Later on, we're going to talk about the frontier model evaluations, as well as what you think Cooperative AI might be about as important as aligned AI. But first off, I guess you're director of Frontier Safety and Governance. What does that actually involve in practice? I can see that being a whole lot of different things, and I don't have a sense of what your day-to-day is like.
My team is called the Frontier Safety and Governance team, and we have 3 main pillars: frontier safety, frontier governance, per the name, and then frontier planning.
This adjective “frontier” is a new term, I would say almost of art, to refer to these general-purpose large models like Gemini and others. Frontier safety looks at dangerous-capability evaluations. It tries to understand what powerful capabilities may be emerging from these large, general-purpose models, forecast when those capabilities arrive, and then think about risk mitigation and risk management.
This also led to the Frontier Safety Framework, which is Google's approach to risk management for extreme risks in frontier models. That's frontier safety. Frontier governance is advising on norms, policies, regulations, and institutions, especially with an eye towards safety considerations. And then frontier planning looks to the horizon, tries to imagine what new considerations could be coming with powerful AI and on the path to AGI, and advises Google DeepMind, Google, and really all of society, given those insights.
That sounds like a pretty big remit. How large is the team that's working on all these questions?
The team is quite small, though we're actually hiring for several positions right now. By the time the podcast goes live, that may be wrapped up. What's really great about working at Google DeepMind is that we have a lot of partner teams and a very collaborative culture.
We work with technical safety—called the AI Safety, Gemini Safety, and Alignment teams. We work with responsibility teams, policy teams, and so forth.
I guess Google DeepMind has, over the last year or two, become more integrated into the rest of Google. Are there other groups within this broader entity—I guess Alphabet—that take an interest in these questions? Or are you maybe the only group thinking about the frontier? I guess you're thinking about the most important models and upcoming issues and threats. Are there many other groups taking an interest, having that kind of foresight and thinking years ahead?
I would say Google DeepMind is the part of Google that's most specialized at thinking about frontier models. Google DeepMind is responsible for building Gemini, the frontier model that's underpinning all of what Google is doing, and we also have responsibility, safety, and policy teams that are especially thinking about frontier issues.
We then do have partners across Google in these various domains. For example, in policy, we work closely with Google Public Policy on the range of policy implications and considerations connected to frontier models. But Google DeepMind is, I would say, where the heart of the thinking related to frontier policy issues takes place.
Okay, so I guess back in 2021, you'd been the founding director of the Centre for the Governance of AI, or GovAI, which was a reasonably big deal then. I think it's gone on maybe to be an even bigger deal since; it's a pretty prominent voice in the conversation around governance of AI. Why did you decide to leave this thing that was going quite well to go and work at Google DeepMind instead?
I agree: it went well at the time, and it's gone even better since. A lot of credit goes to Ben Garfinkel, who's the executive director of GovAI, and the many others who work there.
At the time, I was an informal adviser to Demis Hassabis, CEO of DeepMind, and Shane Legg, co-founder of DeepMind. I found that I had a lot of potential impact in giving advice on AGI safety, AGI governance, and AGI strategy. However, to be most impactful, it helps to be inside the company, where I have more understanding of the nature of the decisions that they're confronting and more surface area to advise not only Demis and Shane, but also many key decision-makers.
To take a step back, I want to reflect on this road-to-impact approach of advising important decision-makers. I would say one lesson I've drawn from history is that often, in these pivotal historical moments—in crises or in very high-leverage historical moments—a lot depends on the behavior, ideas, and very character of key individuals in history. In the musical portrayal of Alexander Hamilton, it's like, “Who's in the room?” What decisions are made in the room?
I think that's true. When you look at history, especially in these pivotal historical moments, it's incredible how much the ideas that people bring into the room, the resources, and the insights that they have available shape the solutions that they construct. So that argues for advising people who will be influential on these important historical developments. I think AI and AGI was, in my view, one of the most important historical developments, and I think Demis and DeepMind are very likely to be influential in the ongoing development of AI and AGI. They have been so far.
There's a second part of this, which is that, in addition to advising influential decision-makers, there's the idea of boosting decision-makers who have the kind of character you would want in critical decision-makers. Do they have the sensibility—in my case, are they aware of the full stakes of what is happening? Are they safety-conscious? Do they have the technical and organizational competence to pull off what needs to be built? Because if you have clumsy hands, even if you have good intentions, that may still lead to a bad outcome. Finally, do they have the wisdom to be able to make these very hard decisions that have complex and uncertain parameters around them?
In my view, Demis—and, again, Shane, to mention him—are extremely impressive individuals in these properties: their safety orientation, their broad perspective on the stakes of the issues, their wisdom, and their broad character.
I also do want to reflect on GovAI during my time.
Yeah, it produced a lot of great work and great people. It's since gone on, I think, to produce even significantly more great work and great people, so Ben Garfinkel has done a great job.
It's interesting reflecting on some of the people who've gone through GovAI. One person who worked very closely with me at the time is Jade Leung. She used to be head of a partner team at OpenAI and is now the chief technology officer at the UK AI Safety Institute. A number of other very prominent people in AI safety and governance have similarly gone through GovAI: Markus Anderljung, Robert Trager, Anton Korinek, who's a prominent economist who's done some work there, Miles Brundage, and others.
I guess back in 2018, in our short interview back then, you were saying people should definitely be diving into this area because it's going to grow enormously, and it's going to be really good for your career, with lots of opportunities. I think that has definitely been borne out: people who got in on the ground floor have been doing super well career-wise.
I think it's still early days for any prospective joiners. I always encourage people to hop trains as soon as they can, because AI is only—it's still just a small fraction of the economy, so there's a lot more impact to come and work to do.
We think that, in the fullness of time, it's going to be close to 100%, certainly more than 10%. It's 0.1% now—certainly not more than 0.1% in terms of total revenue—so there are many orders of magnitude to go up yet.
Let's open, though, by talking about the work that you did in your previous incarnation, which was as an academic. I think you did your thesis back in the early 2010s on technological determinism. That was the main focus, I think, of the paper that came out of it. Who, if anyone, controls technological change? What was the academic debate there that you were reacting to or trying to be a part of?
My academic trajectory had a number of chapters. The first was on technological determinism, which we can come to. For completeness, the second was on great-power politics and peace specifically, which actually led to a lot of work that continues to be relevant, I think, to the question of AI and AGI governance. And I also did some statistics and causal-inference work, which has some relevance to thinking about AI today.
Turning back to technological determinism, I would say I first came to this in undergrad, reflecting on what shapes history, how we can do good, and how we can steer the trajectory of developments in a positive direction. An insight I had was that history is not just the sum of all of our efforts. It's not just that we all push in different directions, and then you take the sum and that's what you get.
Rather, there are these sort of general-equilibrium effects that economists often talk about, where for every unit of effort you push in one direction, the system will push back with an equal force—sometimes a weaker force, sometimes a stronger force. When you're in such a system, it's very important to understand the structural dynamics. Why does the system sometimes resist efforts or sometimes amplify efforts?
Why do you see these really astounding patterns in macrohistory? For example, if you look at patterns of GDP growth, there are these famous curves where, after the devastation of, say, World War II, both Germany and Japan completely rebound within less than a decade and then return to their pre-war trajectory.
We've seen Moore's law, which is just an astounding trend. It's not just that it continues—that transistor density is increasing exponentially—it's very much a line. You can predict where we'll be quite precisely years in advance. We now have scaling laws, which have given us sort of another generation of Moore's law, which again seems to allow us to predict years in advance how large the models will be, how capable they will be on loss, and so forth.
There are a number of other macro phenomena that seem quite persistent: the growth of civilization, which I talked about, or looked at, in terms of the maximum energy processing of a civilization. There are also things like the height of buildings and the durability of materials. Really, most functional properties of technology over time have become more functional, such as the speed of transportation and so forth.
Robert Wright, in summarizing the literature, writes that archaeologists can't help but notice that, as a rule, the deeper you dig, the simpler the society whose remains you find. More generally, I think there's an observation that is almost a truism: certain kinds of technology are so complex or difficult that they come after other forms of technology. It's hard to imagine nuclear power coming before coal power, for example.
So there are all these macro phenomena and trends in technology, and it's important to explain them. The naive explanation would say that if history is just the sum of what people try to achieve, then it's human will that has produced all these trends, including the reliable tick-tock of Moore's law.
But not all the trends are positive. I know you've reflected on the Agricultural Revolution, which evidence suggests was not great for a lot of people. The median human probably saw their health and welfare go down during this long stretch from the Agricultural Revolution to the Industrial Revolution. Of course, it gave rise to inequality, warfare, and various other things. There are other trends that different societal groups resisted.
In short, I don't think the answer is that history is just the sum of what people try to do. It depends, of course, on things like power, timing, the ecosystem of what's functional and what's possible and what's not, and what technology enables. So I wanted to make sense of this.
The thing that we have to try to reconcile here is that, on the one hand, we see these trends that seem like they're not really responsive to any person's particular decisions. They're acting almost like—it's a little bit like the psychohistory in Foundation, where you just have these broad trends, where everyone is an ant in this broader process. It isn't obvious that any particular individual was able to shift things.
On the other hand, technology, at least so far, doesn't have its own agency. It does seem like it's humans doing all of the actions that are producing these outcomes. Couldn't they, in principle, if they really hated what was happening, try to shift it? We feel like we have agency right now over how society goes, or we feel we have at least some agency.
How do you reconcile this macro picture, where it seems like humans don't have that much control over technology, at least historically, with the micro picture, where we feel like we do now? Am I understanding it right?
Definitely. Some of these earlier theorists of technology and scholars of technology, in the 1960s to 1980s, even endowed technology—this abstraction—with a sense of autonomy and agency. Technology was this driving force, and often humans were along for the ride.
Langdon Winner was one of the most prominent scholars who talked about technology having autonomy. Lewis Mumford talks about the machine, with a capital M, as this abstraction that is driving where society's going, and humans just support the machine. We are cogs to support it. Jacques Ellul referred to la technique, which is the sort of functional taking over. He had this metaphor that humans make a choice, but we do so under coercion, under pressure from what la technique is, and la technique is the answer.
I think there were these scholars and others who really did endow technology with this kind of agency. Then a later generation criticized them, saying that this abstract technology is an abstraction—a very high-level abstraction, almost poorly defined. When you actually look at history in detail, under the microscope, where's technology? You don't see the machine. Is the machine with us right now? You see people with ideologies, ideas, and interests making decisions.
I would say this led to a revolution in the study of technology toward what's been called social constructivism. Methodologically, this is more ethnographic or sociological. It looks at the details of how decisions were made, the idiosyncrasies of technological development, and the many dead ends or detours. Early on in the development of technology, people didn't know what the end result was going to be, and they had many competing visions.
It wasn't foreordained that the bicycle would look the way it does, or that the plane would look the way it does. For my personal intellectual trajectory, the PhD program I started in was one of the prominent departments working on this at Cornell. For me, this was a surprise, because I really wanted to explain these macro phenomena, and the answer I got from this department was, “This is wrong. This is technological determinism.”
It is what scholars have referred to as a critic's term. Anyone who actually advocates technological determinism is advancing a straw-man position. No one is serious about this. So this whole generation of sociologists and historians of technology really looked at the micro details of how technology developed and dismissed these abstractions: that technology can have autonomy, that it can have an internal logic of how it develops, and that it can have these profound impacts on society.
We name revolutions after technology—the Agricultural Revolution, the Industrial Revolution, and so forth.
In this paper, it seemed like the constructivists, in their reaction to this determinism, were really staking out a relatively extreme opposing position. They were almost suggesting that it's always human responsibility, and that people always have choice over what technologies they adopt and what form they take. Am I understanding that right?
In a way. I would say the debate was never directly had, or was rarely directly had. It was often indirect.
In defense of the constructivists, I think they were asking different questions. More importantly, they had different tools. They had the tools of ethnography and sociology, and they were answering questions that those tools allowed them to answer. The answers were narratives based on the conversations that took place and the decisions that were made.
To explain macro phenomena, those tools are not well suited. I do think there was a mistake that was made, which was to dismiss the claims about macro phenomena and technological determinism in pursuit of the questions that they had. I think it's a real loss for the history of technology that so little work has since been done on these bigger macro questions.
Were the constructivists motivated by a sense of moral outrage? Maybe they saw people adopting technologies that were socially detrimental, and those people might then excuse it by saying, “We have no choice. We have to do this for competitive reasons, or it's going to happen anyway. There's nothing one can do.”
The constructivists were frustrated by this and wanted to say, “No, no, you're responsible. You're doing it, so you do have agency here.”
This is an argument that's been made often, and more recently by many different schools, including about AI. One criticism of what these people would call the AGI ideology is that AGI is not foreordained, or that the development of AI in any given sense is not foreordained.
When we talk about it as if it's inherently coming and will have certain properties, that deprives citizens of agency to reimagine what it could be. That's the constructivist position on technology, exactly as you said.
The counterposition I would offer is that you don't want to equip groups trying to shape history with a naive model of what's possible. You want to channel energy where it will be high leverage, where it will have a lasting impact, rather than in settings where the structure will resist you—where all the force you push in one direction is met with an equal counterpressure.
Okay, so we should talk about the synthesis that you try to put forward in your thesis. What are the circumstances under which we do have more autonomy, and what are the circumstances where it can be extremely hard to change the course of history?
Maybe first I'll talk a little bit more through the different flavors of technological determinism, because I think it's a rich vocabulary for people to have.
Perhaps the easiest one to accept is what we can call technological politics. This is the idea that individuals or groups can express their political goals through technology, in the same way that you can express or achieve your political goals through voting or other political actions.
If you build infrastructure a certain way or design a technology a certain way, it shapes people's behavior. Design affects social dynamics. Some famous examples are the Parisian boulevards—these linear boulevards that were built, in many ways, to suppress riots and rebellion because they made it easier for the cavalry to get to different parts of the city.
Langdon Winner is a famous political theorist who talked about the politics of artifacts, or “the missing masses” in sociology, which refers to the technology all around us that reinforces how we want society to be. You can think about gates or urban infrastructure as expressing a view of how people should interact and behave.
A famous example is the allegation involving Robert Moses, a famous urban planner in New York City. He allegedly built bridges too low for buses to go under, depriving people in New York who didn't have a car—namely, African Americans—from the ability to go to the beach. This was an expression of a certain racial politics that has been asserted.
In general, I think urban infrastructure has quite enduring effects. You can often think about the impact of different ways of designing our cities.
I guess in recent times we're familiar with the debate about how the design of social networks can greatly shift the tone of conversation and people's behavior. People have pointed to the prominence of quote-tweeting on Twitter, where you can highlight something someone else has said and then blast them on it.
That potentially leads to more brigading by one political tribe against another. If you made that less prominent, then you'd see less of that behavior.
Exactly. I think the nature of the recommendation algorithm, and the way that people can express what they want, has profound impacts on the nature of the discourse and how we perceive the public conversation.
That was technological politics. There are a number of other strands of technological determinism.
Technological momentum is the idea that once a system gets built and gets going, it has inertia. This is from Thomas Hughes. You can think of the dependence on cars in American urban infrastructure. Once you build your cities in a spread-out manner, it becomes hard to have pedestrian-dense cores.
It's also been alleged that electric cars were viable if we'd just invested more, or that wind power and solar power could have succeeded earlier if we'd gone down a different path. We might come back to this, but I think a lot of claims of path dependence in technology are probably overstated.
Again, coming back to the structure, some technologies and some technological paths were simply much more viable. Even if we'd invested a lot early in a different path, I think it's often the case that the path we went on was likely to be the path we would have been on because of the costs and benefits of the technology, rather than because of these early choices that people made.
I guess the extreme view would be to say, “We could have had electric cars,” or that we could have gone down that path. I think we did have electric cars in the 1920s, but we could have gone down that path in a more full-throated way as early as that.
The moderate position would be to say, “No, that wasn't actually practical. There were too many technological constraints, but we could have done it maybe 5 or 10 years earlier if we'd really foreseen that this would be a great benefit and decided to make some early, costly investments in it.”
To make the counterpoint, there is a certain time when the breakthrough is ripe. In AI, this is often the case. Insights about gradient descent and neural networks occurred much earlier than when they had their impact, and it seemed they needed to wait for the cost of compute—the cost of FLOPs—to go down sufficiently for them to be applicable.
You could argue, “What if we had had the insight later?” Once FLOPs get so cheap, it becomes much more likely that someone invents these breakthroughs because they become more accessible. Any PhD student can experiment with what they can access.
There is a kind of seeming inevitability to the time window when certain breakthroughs will happen. It can't occur earlier because the compute wasn't there, and it would be unlikely to occur much later because the compute would be so cheap that someone would have made the breakthrough and realized how useful it was.
Was there another school of technological determinism?
There are other flavors. Another concept that's emphasized is that of unintended consequences—something we know a lot about.
Langdon Winner points to the notion that, as we're inventing technology after technology, we run the risk of being on a sea of unintended consequences. The course of history is not determined by some structure or by our choices, but just by being buffeted one way or another, stumbling from one blender to another.
Sometimes it's positive and sometimes it's negative. I think there's truth to that. Often a technology comes along, and then it takes us some years to fully understand what its impacts are and to adapt and hopefully channel it in the most beneficial directions.
I think the people who really highlight that are exaggerating the scale of the negative side effects from technology. Setting aside some particular cases that we're especially focused on, it seems like the negative side effects of technology in general have gotten smaller with each generation of technology.
We solve more problems than we create on average. That's my take.
Should we come back to the synthesis?
Sure.
The last big part of it is these macro phenomena and trying to explain them. The scholars most focused on that are, I would say, macro-historians, macroeconomists, and political scientists who are trying to explain long-run trends and things like the spread of democracy.
The synthesis starts with an observation: the more micro your observation, the closer you are to people and their day-to-day lives, the more likely you are to conclude with a constructivist explanation. This is a robust empirical finding. If you look at the literature, people who use micro methodologies are much more likely to conclude constructivist-type claims—that what matters is individuals' visions, ideas, and so forth.
The more macro your methodology and your aperture, the more likely you are to conclude a more deterministic set of claims. So we have a puzzle there.
Some constructivists concluded that macro scholars were too high-level, and that they allowed themselves the error of imputing agency to technology because they were so far from the data. I think that's an unfair characterization.
Rather, I do think there are emergent phenomena at different scales of analysis, and we should give macro phenomena their due and try to explain them.
One analogy I've offered is to imagine a hypothetical science of wave motion. We have a group of scientists who emphasize wind. It turns out that when the wind is blowing, it affects the ripples on top of the water. Another community is based on kinetic impact. They say, “When we throw rocks into the water, it produces waves,” and that's their preferred theoretical framework.
Then there's this kooky macro-water phenomenologist who says, “I've noticed that whenever the moon is directly above us, the water level is at its highest point, and when it's at the horizon, it's at its lowest point. I've traveled all over the world, and this pattern is robust. I will offer this moon determinism: the moon explains water levels.”
The wind and kinetic scientists would be mistaken to dismiss the moon determinist simply because that person doesn't have a micro mechanism. That person's finding poses a challenge: How do you explain this pattern? There's no known microfoundation that can explain why the moon is pulling the water. But of course, we know that is in fact what it's doing.
I think there's a similar result in macro-history. There are patterns that need to be explained, and the fact that we don't have a microfoundation isn't sufficient reason to dismiss them. It is a challenge, though.
What is a possible microfoundation? The one I offer hinges on military and economic competition.
The key idea is that there are levels of selection. At the local level, you and I can make a decision about what we do right now. If we wanted to build something, the technology we build might depend on us, our ideas, and so forth.
But if we really want to get going—if we build a new kind of artifact and want it to be everywhere—eventually we're going to need resources to pay for it, and maybe it can't be too opposed by other groups. Eventually, we run into these other forces.
When you think about ways of living—which is a general term for sociotechnical systems or civilizations—they run into resource constraints. They need resources to sustain themselves and then to proliferate, which they typically want to do. That involves economic competition: competition over resources and capital.
The military aspect is always important because, throughout most of history, military competition was ever-present. Even if you had decades or hundreds of years of peace, eventually there was military competition from a neighbor. That provided a higher-level constraint on what ways of living were possible.
Even if there is an active war, everyone's living in the shadow of violence. They anticipate that there could be war in the future, and maybe if they don't play their cards right, they could be vulnerable to aggression.
We could add a higher level of selection. In my thesis, I put environmental selection on top of military and economic selection. There were circles of selection in the sense that civilizations that might be fit for economic and military competition might nevertheless not be sustainable with their environment, and that could be another source of failure to sustain and proliferate.
You can imagine these layers of selection. I tended to put environment at the top, then military and economic competition, then culture, and then psychology or more local dynamics lower down.
One famous quote in the history of technology, arguing against determinism, was that technology doesn't force us to do anything; it merely opens the door. It makes possible new ways of living and new forms of life.
My retort was: Technology doesn't force us; it merely opens the door, and military and economic competition forces us through. When a new technology comes along, the state—or many groups—can choose to ignore it or do whatever they will with it. But if one group chooses to employ it in a functional way that gives them some advantage, eventually the pressure from that group will come to all the rest and either force them to adopt it or lead the other group to lose its resources to the new, more-fit group.
That seems, in a sense, very obvious. Why do you think the constructivists were missing this? Why didn't this stand out to them as an important effect?
I'm glad you think it's obvious—maybe because of your background.
I think we shouldn't underestimate the bias that comes from a scientist using the tools they prefer to use and looking under the lamplight. Constructivists were very good at ethnography, sociology, daily-life history, and close-up micro-history. When you look that closely, you don't see the machine exerting its force.
Military competition at a macro-historical level is ubiquitous, but at a micro-historical level it's rare. Wars are rare. Much of the effect of military and economic competition is through how people internalize that threat.
You can equally say, “It isn't this competition that's driving behavior; it's the ideology of capitalism or of military greatness that's driving the behavior.” This is a methodological challenge.
I do think there's a concept of vicarious selection, which is that, in an evolutionary environment, it's highly adaptive for an organism to model its environment and internally simulate what will happen if it goes in one direction or another.
This concept was named by a historian of technology who was trying to explain the development of aviation and aerospace design. His point was that you don't build a plane, try to fly it, watch it crash, and then build another plane and try again. Rather, you invent the wind tunnel, have a theory, model the external environment, and say, “We want these properties in our wing.”
You're still doing this experimentation; you're just doing it in a controlled, targeted manner internal to the broader economic competition.
If I think about how these people might respond—at least to my simulacrum of them—maybe I shouldn't have said it was obvious. It's obvious to me because I've literally been taught this, more or less. I get it in books and maybe even in undergraduate courses. Everything is obvious once you've literally been told it.
But a pushback I can imagine is that we're living in the UK. The UK has nuclear weapons and is in a pretty friendly neighborhood. Are we really saying that when the UK adopts some new technology, designs its cities one way or another, or has a particular housing policy, it's doing this because it thinks it has to for defensive purposes? Is it worried that otherwise it will be invaded by France or Russia or whoever?
It's not as if we're actively thinking about these defense issues or competitive issues all the time. Individual businesses do think, “If we don't adopt this new technology, we'll be outcompeted.” But at least the military thing is less clear once you have a very strong defensive position where you don't feel a great risk of attack.
There are 2 things I'd want to say here. First, in the modern era, military competition has declined a lot, and we have much more of a global culture than we had 100 or 300 years ago. That can change this higher level of selection.
I do think it's still there. Certainly, you see conversations around national security having a lot of force in domestic politics—in UK politics, US politics, and pretty much every country's politics. If there's a claim that we risk losing a strategic position against an adversary, that can be very motivating for internal reform.
The second point is that there's a great example from UK history where I think this dynamic is well illustrated. This comes from Thomas Hughes, the historian of energy systems. He looks at the UK's energy system, which initially had local power plants and local energy systems.
Hughes argues that these were better suited to the United Kingdom's notion of democracy. There was decentralized energy provision, and it was more under the control of localities. It wasn't a large national energy system, and that persisted up to World War II.
Then the cost constraints that it imposed became excessive, and the UK decided to adopt more of a national grid. There's again a story of one way of living, arguably aligned with the political ethos of the community, persisting until the cost constraints become excessive, often driven by a crisis of conflict.
The case study they focus on in your work is the Meiji Restoration in Japan in the mid-19th century. That's almost as clean an example of military competition driving history as you could imagine. Can you briefly explain it?
I looked for good empirics, because it's useful to have a story about Moore's law and macro phenomena, but those aren't sufficient evidence for making sense of macro-history.
I looked for a case where a community chose to go in a direction contrary to what was demanded by la technique—by what was functional in military and economic competition. There's a great example in Japan under the Tokugawa regime.
Their regime lasted roughly 200 to 250 years, and it was a return to the shogun way of life. Samurai were at the top of the pecking order. It was a very feudal society.
They had firearms at the beginning of this period, and then they effectively uninvented firearms. The shogunate centralized firearm production. Everyone who knew how to produce firearms was brought to a central place and paid a stipend not to build firearms, so the technology was forgotten.
They had cannon and firearms, and then lost the technology. That persisted for roughly 200 years. During this time, Japan wanted little to do with the outside world, but they observed that things were changing. Europeans were sailing around and becoming involved in China.
Everything changed on that fateful day in 1853, when American Commodore Perry visited Japan with the explicit purpose of opening Japan to trade. He arrived in steamships that seemed magical. They moved upwind without sails, belched black smoke, and were made of very large amounts of heavy metal.
They had a profound impact on the Japanese who received him. Perry gave a demonstration of what was possible with the cannons, bombarding the shore, and gave them white flags so they could signal their desire for the bombardment to cease and communicate in the future.
He said he was going to come back in a year to complete the negotiations. The Japanese asked him, “Will you bring your ships again?” He said, “I'll bring more.”
That was the opening of Japan. It led to a 15-year period of revolution. This was the Meiji Restoration. Different groups were trying to make sense of their new environment. It was no longer sustainable to continue their way of life, and different groups contested it in different ways.
The final answer was the restoration under the emperor and the view that they needed to modernize. The Japanese proactively sought to learn everything they could about the West. They sent people to the West to get all the books on industrial arts and so forth.
Japan incredibly succeeded. Just several decades later, Japan was able to contest control of Asia against the United States, Britain, and others in World War II.
What's powerful about that story is that it shows how a group of people, in a sense, chose to go in one direction with respect to the technology of firearms and other aspects of modern industrial civilization. That choice was time-limited by how long the West would choose not to force a different way on Japan.
The thing that set the timeline was that Japan was an island, so it was actually quite hard to invade. They had this protective barrier, and that gave them a degree of discretion that I don't think they would have had if they were on the steppes of Asia.
They would have felt the pressure and fear of invasion much more saliently and would have been focused on defending themselves. Being an island gave them breathing room. It allowed them to fall quite a bit behind, but at some point they fell so far behind that even the sea barrier wasn't enough to keep them safe from invasion.
At that point, they basically did a complete 180 and decided to catch up and modernize.
Yes. I find this case study quite compelling. Most case studies are rarely so clean, again because people internalize external pressures, and there's mimicry and status dynamics. Communities look up to other communities in ways that are often correlated with power or wealth.
The narrative is not usually so clean, whereas in this case it was very clear that what forced the change was the sheer power of the steamship and cannon that the West could bring.
The reason we're talking about technological determinism is that many people in our circles are very focused on the idea of differential technological development. A related, more recent idea is defensive accelerationism.
The way we might try to shift history in a positive direction is by changing the sequence or shifting the order in which technologies are developed, advancing particular lines of science and research to get them ahead of others. You'd want to advance the technologies you think are generally making the world safer, so that you have more of those technologies by the time other, more risk-increasing technologies arrive.
What does the discipline around technological determinism say about whether this is a viable and sensible approach to trying to influence history in a good direction?
To give a bit more color on differential technological development—which is a very clumsy term, and the community hasn't come up with a cleaner one—the best example might be the seat belt.
The seat belt seems like something we could have invented before the car. You can imagine faster-moving vehicles and the value of restraining a person in the event of a collision. It doesn't seem to require the invention of the combustion engine and the car in order to invent the seat belt.
In principle, we could have invented the seat belt before the car and had it ready to go as soon as cars were diffusing, so we wouldn't have had to wait decades for the seat belt. This is an example of a safety technology that pairs nicely with a capability to make that capability safer.
Another category is developing countermeasures or societal defenses for a capability that has adverse byproducts. An example would be a vaccine. If you know that a potential disease could come, and you can develop the vaccine in advance, then you can be prepared.
A third category is a substitute. An example often given is whether wind or solar power could have been made more cost-effective than fossil-fuel-based power. If the cost curve had been different, we would have invested more in sustainable energy resources, and civilization would have gone down that path rather than a more fossil-fuel-dependent path.
Those are the cleanest examples, but there are more general ones that imagine wholly different technological paths with different properties.
It's worth reflecting on them. Langdon Winner argued that nuclear power was more authoritarian. The story is that the technology requires centralized development and strong coercive infrastructure to make sure it isn't abused, whereas wind and solar permit decentralized development and therefore decentralized politics. They don't have the risks that require more of a security state on top of them.
There are arguments that certain technological trajectories have political or social effects as byproducts.
To offer some challenges, I do think differential technological development is a very important idea that we should think about a lot. In many ways, it undergirds the whole notion of AI safety. The notion of AI safety as a field is that we want to put a bit more effort into AI safety or AGI safety than we otherwise would, than the market would naturally do, with the idea that this will make a difference.
It's like inventing the seat belt before the car. So I think it's a very important idea. However, there are good reasons to doubt its tractability or feasibility.
Most arguments for differential technological development require someone to see 2 pathways that are both viable given some effort, anticipate the consequences of each pathway, and choose the better one. These things are hard.
First, it's difficult to know what the next viable step in technological development is. If you know that you can build a great business that's hungry for that insight, then the market is hungry for it, and you should think it's hard to find 2 of those.
You have to be at the point where you can make this marginal choice, one that others aren't already pursuing. Otherwise, it isn't an intervention, or you have to convince a resource allocator to choose one way or another.
Then there's the second stage: predicting the consequences of going down one path or another. That's extremely hard. You have to anticipate the full sequence of technology and the tree of technologies that will spawn from one path versus another, along with the many direct and indirect consequences of those technologies.
We know from the study of technology—and from our attempts to make sense of technology—that it's very hard to foresee the direct and second-order effects.
That's some pessimism, coming back to the technological-determinist perspective. The notion of technological momentum would say there is path dependence in the directions you go down. You build up expertise and sink investments.
It is important, at the beginning of major investments in new infrastructure, to ask: Is this going to create sunk costs? Is this going to make it harder for us to choose a different path in the future? Reflect on whether there are other consequences we should weigh before going significantly in a certain direction.
In the archetypal case we're talking about with artificial intelligence and AI alignment methods, it doesn't feel binary. On the one hand, you have to say, “We need 2 different viable paths where some incremental effort might push us in one direction versus another.”
But it seems like almost everyone thinks AI companies are going to put some effort into alignment. They care that the models broadly do what's being asked. The idea is that we want to do more of that than the market might provide.
We're not choosing between aggressive, non-aligned AI that someone really wants and aligned AI that nobody wants. We're trying to go even further on something most people regard as desirable and would want to incorporate if it's practical.
In terms of deciding whether this is actually a better path, I'm sure people have argued that alignment might backfire. It could be worse than not aligning. You can imagine ways that could happen. Still, on balance of probabilities, it seems like a reasonable bet—nothing that people should be super uncertain about.
Maybe this is the case people focus on because it's among the better examples anyone has come up with for trying to pursue differential technological development. Perhaps there are lots of other examples that were left on the scrap heap because it wasn't clear that they were either viable or desirable. What do you think?
I agree that safety and alignment are, on net, very beneficial bets that we should invest heavily in. To make the countercase, one could offer a general-equilibrium argument that the marginal return on your investment in safety and alignment is less than what you pay because the market would otherwise have provided it.
To motivate this, reflect that the success of AI assistants today is very much constrained by their ability to be aligned and safe. Safety and alignment are huge priorities for developers because if the systems aren't safe and aligned, they won't be good products.
The market is already providing a huge motivation for advancing this field. Reinforcement learning from human feedback, or RLHF, is one of the key alignment techniques. Constitutional AI is another. These were developed by individuals motivated by AGI safety and supported by those resources.
Maybe that brought the technology forward a few years. But imagine the counterfactual in which we had no investment in AI safety or alignment. Maybe it delays the technology by 2 years until the market demands that we solve the problem, and then other researchers rise to the challenge.
The skeptical response to reinforcement learning from human feedback being developed by alignment- and safety-focused people, then applied to make AI useful and economically valuable across different domains, is to say, “You've wasted your time. Maybe you've even made things worse by speeding them up, which you didn't want to do.”
On the other hand, that seems like a reason to ask, “What's your plan to not develop any of the technologies that actually make AGI work?” That doesn't really seem like an alternative. Surely that would just delay things at best, and we need to get to the point we're at now at some point sooner or later.
I guess you're saying that even if it isn't actively detrimental, it could be useless because the market—Google DeepMind or some other group—would have realized that we needed the equivalent of reinforcement learning from human feedback to make it work. They would have done it at some later point anyway, so the effect of your work has just been undone.
To reiterate, I think the bets on safety, alignment, and interpretability are very good bets on net, so we should keep doing them.
On the margin, what we want to do—and I think sophisticated individuals in the space are thinking this way—is look for work in safety and alignment that would not otherwise be done by the market in time.
This is perhaps where AGI safety and AGI alignment point. They ask: What are the seat belts for AGI? What are the guardrails we need for AGI that the market would not be motivated to find solutions for in pre-AGI systems?
We want to look ahead. This sometimes points to the notion of deception, which might generate a whole new class of problems when an AI system is misaligned but can hide that and deceive us.
That might be a different problem from the systems we have today, and it might arise around this critical period. This is one argument for differentially focusing on that problem as opposed to others.
Makes sense.
Let's push on from technological determinism and talk about this research agenda that you've been involved in promoting and elaborating, called cooperative AI.
As we've been saying, the main focus of differential technological development thinking with regard to AI has been alignment for many years. But you and some co-authors said in this paper back in 2021 that there's this whole other cluster of behaviors that we might like to speed up the development of around cooperation.
You think that could be similarly important, or at least on the margin could be similarly important, because people aren't really talking or thinking about it. How do you define cooperative AI, and why do you think it's key?
It's a big question, and the answer will be extensive because the theoretical framework around cooperative AI is large and complex.
One way of putting it simply is that alignment is insufficient for good outcomes. To make an even stronger claim, you could say it's not necessary to solve alignment to have good outcomes.
That's a strong claim, but it helps motivate the case. Imagine that we only solve alignment 90% of the way. That is, our models do what we want within certain bounds, but we know that if we scale them too far, we can't trust them to continue behaving as intended.
If we know that, and we have global coordination—if humanity can act with wisdom and prudence—then we can deploy the technology appropriately. We can deploy it within domains and to the extent that it is safe and beneficial.
This is the sense in which global coordination is almost a necessary and sufficient condition. If we can globally coordinate well, that's the necessary argument, then we could deploy the technology to the extent that it's safe to deploy.
The sense in which global coordination is sufficient is that if we're globally coordinated, we could appoint a reasonable decision-maker to make the risk calculus, and that would satisfy humanity's collective view on how we should proceed.
Now suppose we solve alignment but don't solve global coordination. I can imagine things still not going very well. We could have great powers developing powerful AI systems aligned with their interests and in conflict with the interests of other great powers.
Historically, great-power conflict has been a major source of harm to humanity: devastating wars and the brinksmanship around nuclear weapons. That has arguably imposed an expected cost on humanity more devastating than the world wars, given the willingness of US and Soviet leaders to gamble over nuclear war for geopolitical stakes.
Then there are other consequences, such as a failure to deal with climate change, insufficient global trade, or inadequate pandemic preparedness. These are global collective-action problems that we insufficiently address partly because we aren't coordinated at the highest geopolitical level.
That's one motivation for cooperative AI. If we want things to go well, we ideally need 2 pieces of the puzzle. We need systems that are safe and behave as intended, especially by the principal who deploys them. And we need to be able to deploy AI systems collectively and continue our activities in a way that's jointly peaceful and productive.
That means we've solved enough of our collective-action problems that we're not continuing to engage in nuclear brinksmanship, trade wars, or other major welfare losses due to insufficient global coordination.
The idea is that even if you have AIs aligned with the goals of their operators, this doesn't necessarily lead to a good outcome if those operators are in conflict with one another. The AI systems they're working with could simply lead to a disastrous outcome.
It's like how the fact that we're aligned with our own interests doesn't necessarily produce a great outcome across humanity as a whole. You can still end up in traps and unintended disasters.
To give another example, there was the famous flash crash in 2010. Some algorithmic trading—not sophisticated AI, but algorithmic trading—led to an inadvertent sell-off in the stock market that caused trillions of dollars of paper losses before emergency circuit breakers kicked in.
Those limits stopped trading and allowed the trades to be unwound. The outcome wasn't intended by the traders. It was an emergent dynamic from algorithms that had protocols that made sense within normal bounds, but which could get out of control when interacting.
You sometimes see this on Amazon or other online marketplaces. Famously, a book sold for millions of dollars because 2 sellers each had an algorithm that would bid up the price as a function of what the other was selling it for. The algorithms iterated to a crazy valuation.
Flash crashes happen frequently. In the stock market, we have safeguards so that when there's a sudden movement, trading stops, and there are rules for how to unwind those trades.
As we deploy simple AI systems, narrow AI systems, and increasingly general AI systems out in the world—and increasingly, as people talk about agents out in the world that are more empowered and more general-purpose, can move between domains, and may have access to bank accounts and email—how do we make sure there aren't unintended emergent dynamics that could be harmful?
Cooperative AI partly looks to addressing that issue.
A specter that haunts this entire conversation, if we're focusing on the military case in particular, is that when countries are in intense conflict with one another, you get a process of brinkmanship. One country escalates, and the other country has to decide whether to cool things down, escalate, or back down.
You keep getting this brinkmanship and escalation process until one of them blinks, or they decide to find a solution that makes them both happy.
The trouble is that if you have an AI-operated military, these AIs can make decisions on a completely inhuman timescale. This entire process of brinkmanship that might take days, weeks, or months when humans have to go away to a meeting, think about it, discuss it, and decide how to react could all play out in a matter of minutes.
That possibility terrifies people and is one thing discouraging people from placing AI into important decision-making roles over national security.
An important question in the deployment of agents will be what degrees of autonomy we endow our systems with, versus when a human reviews decisions of different kinds. That's a function of the stakes of the decision, the resources deployed, and how many actuators the agent has.
If the AI is controlling weapon systems, the important role of human review becomes much greater. But as you note, if there is time pressure, Paul Scharre, a theorist of AI in the military, worries about a flash escalation occurring in potentially kinetic warfare. Another dynamic would be cyber conflict.
Looking over this paper, I wasn't initially completely convinced that this was something we had to go far out of our way to focus on. That's because I think cooperative behavior—knowing how to cooperate with other agents—is instrumentally convergently useful.
If you're developing agents that are good at their job, that are actually useful to apply, then they have to learn how to cooperate at least in cases where they're being used. Otherwise, they're just bad at their job.
There's a lot of commercial pressure to develop this, and the systems we're imagining are generally capable. They're insightful and intelligent, maybe approaching or exceeding human level. Why wouldn't they be able to think about these things and figure out how to cooperate, the same way thoughtful human beings try to avoid conflict?
You're right that cooperative skill, or cooperative intelligence, is likely to be instrumentally useful. We should see some cooperative skill developed as a byproduct of almost any development agenda for AI.
The cooperative AI bet is that, on the margin, it's beneficial to invest more in it early, so that when we get to powerful systems, they are more cooperatively skilled than they otherwise would be.
In that respect, it's similar to the safety and alignment bet. Safety and alignment will likely be developed by default to some extent. The bet is that it's worthwhile to invest energy early so that we're further ahead on safety and alignment than we otherwise would be.
The seat belt comes earlier. Cooperative sophistication and skill come earlier relative to the level of capabilities of the agents out in the world.
We've talked about the military issues with cooperation, but that's an extreme case. I imagine there are more mundane examples where cooperation could be useful. Do you want to give a couple?
Much of society consists of bargaining interactions: economic exchange, whether on a marketplace to sell used goods, in financial markets, or in major corporate deals.
There's a lot of welfare gain to be had if we can strike deals more efficiently. When 2 parties can both be made better off, it would be useful for them to reach that understanding. Right now we bargain through existing protocols and institutions that are human-built.
In principle, AI could be much more effective, or at least an AI-human team could be much more effective. An exotic solution would be to put your AI delegate in a box with my AI delegate. We let only the proposed solution come out of the box.
That may solve some bargaining problems. Often, the challenge is that during bargaining we reveal private information, which could give the other side an advantage. This leads bargainers to withhold information, bargain slowly, demonstrate resolve, or signal that they have a good outside option.
If you put the agents in a box and they know that the only thing they can do is output the solution or say there's no deal, that could dramatically change the bargaining dynamic. We might get solutions more often and in a way that involves less costly signaling.
It might be easier for AIs to cooperate than humans in some ways. They have a much higher bandwidth of communication. They can send an enormous number of words and have a very lengthy conversation, where humans wouldn't regard it as worth the effort to reach a bargaining outcome.
They can also commit to acting a particular way. You can copy a model and demonstrate that, in a given situation, it will consistently accept a particular kind of bargain. You can say, “I've created this cooperative software, and all the exact copies of this piece of software will do the same thing.”
That's not something you can easily do with human beings. We try to look at their historical track record to learn what they're like, but here it's even easier to judge the character of an AI model.
Are there important ways that it could be more difficult, or less straightforward, for AIs to cooperate with one another than for humans?
You may have to remind me to come back to that as I answer, because there's a lot here.
First, I want to clarify that the cooperative AI bet is a portfolio bet. There are many cooperation problems at present and in the future between AIs—which is what we're mostly focusing on—but also between AIs and humans, and among humans.
The bet is that, on the margin, if we put in effort now, AI might help us with these cooperation problems. It could help 2 humans cooperate better, help future AI systems cooperate, or help AI systems and humans cooperate.
I'm making this distinction partly because I think one promising direction is AI systems that can help humans reach solutions.
We were talking about a bargaining setting where there's some conflict of interest. A more prosocial example would be political deliberation, where people are trying to find the right course of action for their community, municipality, family, or nation.
We have tools for doing that: institutions of deliberation, the press, voting, and so forth. AI could potentially help humans find the course of action they would want to pursue.
Google DeepMind recently produced a paper discussing the idea of a “Habermas Machine,” where AI serves as a facilitator of political deliberation. They found that language models today can serve as a useful tool to summarize people's political views about an actual policy issue and articulate a consensus—a detailed, productive consensus that people will sign on to.
They report that this Habermas Machine can articulate a better consensus than the humans employed to try to do so. If we imagine extending that further, it could help in many political settings where we have difficult, multidimensional issues to talk through.
AI could help us understand the dimensions, which dimensions matter most to me, and what potential mutually beneficial solutions exist between me and another group along those dimensions.
To come back to the question of how it might be more difficult for AIs to cooperate, I think many people believe AI will be better at cooperating. That's the prior, and maybe we can talk about what I've called the super-cooperative AGI hypothesis: As AI scales to AGI, cooperative competence will also scale to infinity, and AGIs will be able to cooperate with one another to such an extent that they can solve global coordination problems.
Magic.
The reason that hypothesis matters is that if AGI, as a byproduct, becomes so good at cooperation that it solves global coordination and collective-action problems, then we don't need to worry about those problems among humans. We can just bet on AGI, pursue safety and alignment, and let AGI solve the rest.
It's important to think through what would be required for that hypothesis to be valid and how we could empirically evaluate whether we're on track.
Here are some arguments against AIs being highly cooperatively skilled with one another. There are different levels to the argument. Maybe AI will be cooperatively skilled but not at the level of the super-cooperative AGI hypothesis. Or maybe it will be less cooperatively skilled than human-to-human cooperation.
The strong version says that humans are pretty similar. We come from the same biological background and, to a large extent, the same cultural background. We can read each other's facial expressions. If 2 groups of people are bargaining, we can often read the thoughts of the other group.
If 2 democracies are bargaining, you can read each other's newspapers and press. In that sense, human communities are transparent agents to one another, at least in democracies.
We also have enormous historical experience for judging how people behave in different situations. We can judge people from many different examples. We can look at theories of folk psychology, childhood, or upbringing to explain people's behavior.
The range of goals humans can have is fairly limited. We roughly know what most humans are trying to achieve.
AI could be very different. Its goals could be vastly different from human goals. Most humans have diminishing returns in almost everything. They're not willing to take all-or-nothing gambles, such as betting the entire company: 50% chance we go to zero, 50% chance we double the valuation.
An AI could have linear utility in wealth or other things, which would change things. AIs could have alien goals that are different from what humans typically anticipate.
They may be harder to read. If you have interpretability infrastructure, they could be easier to understand, but in a bargaining setting, why would one bargainer allow the other to read its neurons? AIs could be much more of a black box to each other than humans are.
The extreme instance is that they could be backdoored. Under current technology, we can almost implant magic words into AIs that cause them to behave completely differently from how they behaved previously. That's extremely hard to detect with current levels of interpretability.
An AI could completely flip its behavior in response to slightly different conditions. Humans can deceive, but it's hard. It requires training to achieve a complete flip in the goals and persona of a human. An AI could, in principle, have one neuron that completely flips its goal.
That means its goals may be unpredictable given its history. This is also a challenge for interpretability solutions to cooperation. People sometimes say that if we allow each other to observe each other's neurons, then we can cooperate given that transparency.
But even that might not be possible because I could hide a backdoor in my architecture so subtle that it would be very hard for you to detect, yet it would completely flip the meaning of everything you're reading from my neurons.
This is slightly out of place, but I mentioned this crazy idea that you could use the possibility of a backdoored model to make it extremely undesirable to steal a model from someone else and apply it.
You could imagine the US saying, “Maybe we've backdoored the models we're using in our national-security infrastructure. If they detect that they're being operated by a different country, they'll completely flip out and behave incredibly differently.” There would be almost no way to detect that.
It's a bad situation in general, but it could make it more difficult to hack, take advantage of, or steal another group's model at the last minute.
Arguably, the notion of backdooring one's own models as an antitheft device could deter model theft. It makes a model less useful once stolen if you think it might have a “call home” feature or a feature that causes it to behave contrary to the thief's intentions.
Another interesting property of this backdoor dynamic is that it provides an incentive for a would-be thief to invest in alignment technology. If you're going to steal a model, you want to make sure you can detect whether it has a backdoor.
For antitheft purposes, if you want to build an antitheft backdoor, you again want to invest in alignment technology so you can make sure the backdoor will survive the current state of the art in alignment. That's a virtuous cycle.
Maybe this is a good direction for the world because, as a byproduct, it incentivizes alignment research. There could be undesirable effects if it leads models to have architectures that are highly sensitive to subtle aspects of the model, or makes models more prone to subtle forms of deception.
More research is needed. It sounds a little bit like dancing on a knife edge.
That's one way AIs might find it more difficult to cooperate. Another reason this research agenda doesn't feel intuitively essential is that the AIs I interact with—LLMs—seem very cooperative and nice by nature.
It's easy to imagine scaling them up, doing the same sort of RLHF we're doing now to produce that kind of personality, and saying, “Wouldn't they continue to act with nice personalities and be really cooperative by nature?”
What do you think of that argument?
We need to distinguish between niceness and cooperative skill, or cooperative intelligence.
When we say someone is cooperative, we often mean they're nice, altruistic, generous, or prosocial. The cooperative AI research program isn't about making nice AI or prosocial AI. It's about making cooperatively intelligent AI that can solve cooperation problems better than it otherwise would.
It's an open question how cooperatively skilled today's AI systems are. I agree they're nice, but they're primarily interacting with individual users. They're not much deployed to solve cooperation problems.
We need a science—an empirical science—of how cooperatively intelligent they are. We can also expect that, in equilibrium, they won't necessarily be prosocial. They'll be deployed by interested actors in bargaining settings.
If we're trying to make a deal, I don't want my AI system to be nice. I want it to be a faithful delegate to my interests. Then the question is whether it can efficiently bargain with other agents that are similarly aligned with other human principals to find a mutually beneficial solution.
Alternatively, can we build a governor or mediator that we can both trust, which will weigh our respective interests reasonably? We could each tell this mediator or arbitrator our goals, resources, and outside options, and perhaps prove them. Then the arbitrator would tell us what the solution is.
A nice thing about models is that you can test them and then use exactly the same model to consider new inputs as previous ones. You could trust a model as a mediator by looking at its track record and saying, “It produced fair outcomes in all these previous cases as a judge, so I would trust it to produce a fair outcome in this new dispute.”
There's a lot of potential there. What is the actual agenda for trying to make AIs more cooperative? Are people working on it? What technologies do we need to develop?
There are many theoretical ideas and research programs that could help. Probably the highest leverage is at the foundation level.
There's a Cooperative AI Foundation, which I helped found. Its goal is to promote cooperative AI—AI systems that are more cooperatively skilled than they otherwise would be, especially given the bet that AI systems will continue scaling in capability.
One area the foundation has found to be high leverage is environments or benchmarks. Building an environment that measures cooperative skill is a kind of public good for AI development. You invite many groups to compete to have the most cooperatively skilled agent, measure that, and celebrate advances.
In AI, a good benchmark can motivate work because it gives researchers a target to hill-climb on.
You could have lots of different tests for how effectively these AIs cooperate, including conflicting scenarios where different behaviors are required to produce cooperative outcomes. In some cases you might need to be a hard-ass, and in others you might need to be friendly.
There's a history of game-theory tournaments trying to figure out the simplest cooperative agents. You could have a research program trying to develop the most successful cooperative agent across many different possible situations.
There are many interesting aspects. Does the agent have a good theory of mind for the other agent? Can it model what the other agent is trying to achieve and see what the agent would otherwise do?
Can they communicate effectively? Can we even have a shared vocabulary? Can I express something that you understand?
Then there's communication under adversarial incentives. Can I communicate what's important to me when I'm not sure whether you'll use that information well? You might use it against me.
Can I share information in a strategically optimal way that minimizes my vulnerability while maximizing the potential cooperative gains?
A third category involves commitment problems. You and I might both know the nature of the game. The prisoner's dilemma is the classic example.
We both know we'd be better off if we cooperated, but we can't sign a treaty saying, “I'll cooperate if you cooperate.” In a one-shot prisoner's dilemma, agents unfortunately defect.
Are there techniques or technologies for building a treaty mechanism so that, when we identify a cooperative bargain, we can build a protocol that we will both follow, knowing that the other person is committed to it?
I've noticed a funny phenomenon in discussions of AI and cooperation. People point to ways that advances in AGI could lead to negative outcomes, and then conclude that what we need is a solution to the commitment problem.
If we could credibly commit not to use our future power to oppress someone else—if we could fix that problem, which has beguiled humanity since the beginning of history—then we'd have an escape route.
But that's a bit like saying the problem is that we don't understand where consciousness comes from, and if philosophers could only solve the theory of consciousness and theory of mind, then everything would be clear. That's not really a strategy, because it's unlikely we'll actually solve the problem.
Have you noticed this as well?
To provide the background, the rationalist approach to cooperation has tried to distill cooperation problems down to fewer elements.
One game theorist, Robert Powell, argued that everything is a commitment problem because any cooperation problem can be reformulated as, “If only we could commit to a judge who would solve the problem, then we wouldn't have the cooperation problem.”
If we unpack it, people often point to informational problems as a second category, alongside commitment problems. Another category is issue indivisibility.
There may be a pie we want to divide. In principle, we agree that a 50/50 division would be fine, but for some reason we can't divide the pie 50/50. It's all or nothing, and that might lead to bargaining breakdown.
Coming back to the commitment problem, which arguably undergirds all cooperation problems, I do think it's an important research area. We shouldn't count on it, but perhaps AI can help us solve commitment problems much further than we realize.
In Carl Shulman's podcast with you, he talked about this. Carl has expressed probably the most compelling story I've heard for how AI could solve our commitment problems.
Can we build a technology in which our AIs, or our AGIs, can build a third AI, in a way that each of them can verify has not been backdoored or secretly biased toward the other side? Then we build up from the foundation of this third AI and hand over power to it to make decisions for us.
If we can solve the problem of building it in a way that's verified to be fair, then maybe we could solve many of our problems.
It's a very big prize if we can make it work.
In the paper, you spend quite a bit of time talking about potential downsides of having more cooperative AI. Making AI more cooperative seems like a virtuous thing to do. What are some ways it could backfire?
There are a number of ways to think about this. In general, for our various bets for positive impact, it's always good to think hard about how they could backfire or what negative byproducts could result.
Cooperation sounds good, but by definition it's about agents in some system making themselves better off than they otherwise would be. It could be 2 agents or 10. The problem of cooperation is how those agents get closer to their Pareto frontier and avoid deadweight loss.
It says nothing about how agents outside that system are affected by the cooperation. There may be a phenomenon of exclusion. Increasing the cooperative skill of AI will make those AI systems better off, but it may harm any agent excluded from that cooperative dynamic.
That could include other AI systems or groups of people whose AI systems aren't part of the dynamic.
On that point, there's a famous quote that democracy is 2 wolves and a sheep deciding what to eat for lunch. Once you have a majority, you can exclude others and extract value from them.
There are many kinds of cooperation that are antisocial. We don't want students cooperating during a test to improve their joint score. Tests and sports clearly have rules against cooperation.
There are certain entities and marketplace interactions where there are rules for how companies should interact. Criminals cooperate in all kinds of criminal activity, and we do not want criminals cooperating more efficiently.
One reason the Mafia is such a potent organization is that it's figured out how to sustain intense internal cooperation without using the legal system to enforce contracts and agreements, through social codes, careful screening, and so on.
We can think of cooperative skill as a dual-use capability: broadly beneficial in the hands of good actors and potentially harmful in the hands of antisocial actors.
There's a hypothesis behind the program that broadly increasing cooperative skill is socially beneficial. It's worth interrogating. I think it's probably true, but the bet is that if we make the AI ecosystem—the frontier of AI—more cooperatively skilled than it otherwise would be, it will advantage antisocial actors but also prosocial actors.
The argument is that, on net, this will be beneficial.
The way to argue that, I suppose, is to say that we've become better at cooperation over time as a civilization and as a species, and history—or well-being—has generally been getting better. Are there other arguments?
The arguments I'm drawn to are along the lines of what you articulated. Even though cooperation can empower groups to cause harm and be antisocial, it does seem to be a net-positive phenomenon.
If we increase everyone's cooperative capability, that means there are all these prosocial benefits. There are more collective wins to be unlocked than antisocial harms. In the end, cooperation wins out.
It's similar to the argument for why trade is net positive. Trade can also have a property where you and I being able to trade may exclude others with whom we formerly did business. But on net, global trade is beneficial because every trading partner is looking for the best role for them in the global economy, and that eventually becomes beneficial to virtually everyone.
Some people argue, not without reason, that well-being on a global level may have gotten worse over the industrial era because although humans have gotten better off, the amount of suffering involved in factory farming is so large that it outweighs the gains generated for human beings.
In that case, the joke about 2 wolves and a lamb deciding what to eat for lunch is quite literal. It's an unusual case because pigs and cows aren't able to negotiate or engage in cooperation the way humans are. Maybe it's not surprising that better cooperation among a particular group might damage those who aren't part of it and can't form agreements.
That's a good example of the exclusionary effects of enhanced cooperation. It may be a cautionary tale for humans.
If AI systems can cooperate with one another much better than they can cooperate with humans, we might be left behind as trading partners and decision-makers in a world where AI civilization can cooperate more effectively within machine timescales and using an AI vocabulary.
Some people have painted a grim picture. If you have many AGIs that can communicate incredibly quickly, credibly explain how they'll behave in the future, and make credible commitments to one another, they may be able to cooperate extremely rapidly.
They wouldn't be able to do the same thing with humans. We can't commit in the same way, and we can't communicate at their pace or with their level of sophistication. Naturally, they might form a cluster that cooperates extremely well, extracts all the surplus, and excludes us.
How big a problem is this for the cooperative AI agenda?
What you described is related to one of the main challenges to alignment. Many alignment techniques involve having multiple AI systems check each other. You have an overseer AI watching the deployed AI to make sure it behaves well, perhaps with multiple overseers and a majority-vote system.
Even if there's joint misalignment, it's hard for any one AI system to defect because it can't coordinate with the others, which may have been trained differently and may not share the same background or exact misalignment.
The worry some have for AGI safety is that there will be collusion among our AI systems, so we can't rely on these AI institutions to make us safe.
I think this is probably one of the biggest downsides to the cooperative AI research program from an AGI-safety perspective, and it's worth investigating.
The reason I'm persuaded that cooperative AI is worth pursuing on net is, first, that this should be investigated. The top of the agenda should be understanding the different components of the cooperative AI portfolio and asking what the case is for and against each one.
We should test the hypothesis that different parts of cooperative AI are worth pursuing. Second, my view is that global coordination is so important to things going well that there's a lot to be gained by betting on having more cooperatively skilled AI systems than we otherwise would.
This ties into another paper you helped write, called “Levels of AGI: Operationalizing Progress on the Path to AGI.” It's a bit of a mouthful, so we've referred to it as the “What Is AGI?” paper, which is catchier.
What did you want to point to about the nature of AGI in that paper that you thought many people were missing?
The paper primarily tried to offer a definition and conception of AGI that we thought was implicit in how most people talk about AGI. We wanted to provide a rigorous statement of it so we could move on from the unhelpful strand of dialogue saying that AGI is poorly defined or means different things to everyone.
We've received very positive feedback. Most people think it's a reasonable and useful conception of AGI.
There are some prior ideas I want to call out. AGI is a complex, multidimensional concept, and it's prone to fallacies of reasoning when people use it uncritically.
One fallacy is that people think AGI is human-level AI. They think of it as a point, a single system, or a single kind of system. Often they think it's human-like AI, with the same strengths and weaknesses as humans.
We know that's unlikely to be the case. Historically, AI systems have been much better than us at some things—chess, memory, and mathematics—and worse at other things. We should expect AI systems in AGI to be highly imbalanced in what they're good and bad at. Their profile won't look like ours.
Second, there's a risk that the concept of AGI leads people to try to build human-like AI. We want to build AI in our image. Some people argue that's a mistake because it leads to more labor substitution than would otherwise be the case.
From an economics point of view, we want to build AI that's as different from us as possible because that's most complementary to human labor in the economy.
AlphaFold is an example. It's a narrow AI system that's very good at predicting the structure of proteins. It's a great complement to humans because it isn't doing what we do. It's not writing emails or strategy memos. It's predicting the structure of proteins, which humans couldn't do.
No one's losing their job to AlphaFold, but it enables new forms of productivity in medicine and health research that otherwise wouldn't be possible.
Arguably, that's the kind of AI system we should try to build: alien and complementary, perhaps narrow AI systems.
Another aspect of the concept of AGI is that it points in the direction of general intelligence. Some people argue that this is a mistake and that we should develop systems of narrow AI, which they might say is safer or more likely. They might even argue that general intelligence isn't a thing.
I'm less persuaded by that argument. I do think general intelligence is likely to be an important phenomenon that wins out. People will be more willing to employ a general-intelligence AI system than a narrow-intelligence AI system in more and more domains.
We've seen that in the past few years with large language models. The best poetry language model, email-writing language model, historical language model, and so forth is often the same language model—the one trained on the full corpus of human text.
There's enough spillover in lessons between poetry, mathematics, history, and philosophy that your philosophy AI is made better by reading poetry than by having separate poetry and philosophy language models.
It's an amazing thing about the structure of knowledge. We didn't necessarily know that before. In humans, it's true that we specialize much more. It's very hard to be the best poet, philosopher, and chemist.
But maybe when your brain can scale as much as you want, with virtually unlimited ability to read all these different things, there are enough connections between them that you can be the best at all of them simultaneously.
It's an important empirical phenomenon to track. It need not be true that the best physics or philosophy language model is also the best poetry, politics, or history language model.
You can imagine that this might come apart in the future if there's more effort to develop very good specialized models. Maybe at the frontier this won't be the case anymore.
A coding language model doesn't need to know history. Even if reading history helped it learn something about coding, perhaps later we should distill out most of the historical knowledge so that the model is smaller and more efficient.
This is a phenomenon to track. I think implicit in the notion of AGI is that general intelligence will be important and that we're not at peak returns to generality. There will continue to be returns, so we'll continue to see large models being trained.
Are there any other pros or cons of generality worth flagging?
One thing that makes the concept of AGI useful is that many people critique it or say they don't want to use it, but I've rarely seen an alternative for what we're trying to point to.
Sometimes people say “transformative AI.” I don't think that's adequate. Transformative AI is usually defined as AI systems that have an impact on the economy at the scale of the Industrial Revolution.
The problem is that you can get transformative impact from a narrow AI system. You could have an AI system that poses a narrow catastrophic risk, and that isn't general intelligence.
There is something important to name about AGI. People often say the term is confused or represents a cluster of competing ideas. But despite all the criticisms, people keep using it. They always come back to it.
At that point, you have to concede that there's something important there that people are desperate to refer to. We need to clarify the idea rather than give up on it.
It's good to refine our conceptual toolbox. There are other concepts and targets worth calling out.
One idea, named by Ajeya Cotra and many others, is machine-learning AI that can do machine-learning R&D. Such a system could engage in recursive self-improvement. It need not be general; it could be a narrow AI system. But if it kicks off a recursive process, that's an important phenomenon to track.
I think Holden Karnofsky has also drawn attention to this. But coming back to general intelligence, I think AGI is a probable phenomenon. We'll probably get general intelligence around the time we get AI systems that can do radically recursive self-improvement, or many other things.
A slightly contrarian take I have is that people in machine learning hate this, but machine-learning research—even cutting-edge work—might not actually be that difficult.
If you break down the process by which we're improving these models, you have a theory-generation stage, then a stage where you figure out how to test it and develop a benchmark. You run the experiment, which is compute-intensive, decide whether it was an improvement, and go back to the generation stage.
This might be possible well short of a full AGI with all these different skills, especially if you focus on it. Maybe a lot of this research is much less difficult than people involved in it would want to believe, and it could be relatively easily automated.
That would be shocking and consequential if true.
An interesting phenomenon we've seen in public surveys of AI and machine-learning experts is that they often put automation of the ML process very late—one of the last or the last task performed by AI systems. They often put it significantly later than AGI or human-level AI, however that's defined.
Even given a high definition, automation of ML R&D is often reported to come later. One argument is that this reflects a bias: everyone thinks that their career, task, or area of expertise is particularly hard and special and won't be automated.
There may also be a case in favor of the idea that, among human tasks, the last to be fully automated will be the process of improving machine-learning research and development.
Coming back to AGI, it seems useful to point to a space of AI systems rather than a single kind of AI system. We can define it as an AI system that's better than humans at most tasks.
There are different ways to define “most,” and different ways to define the set of tasks. You might say economically relevant tasks, and “most” could mean 99% or 50%. I think it's helpful to choose a large but not universal number. You don't want to set it at 100% because some tail tasks might take a long time to automate, while most of the impact will occur earlier.
There's also a parameter for what it means to be better. Is it better than the median human, who's unskilled? Typically, you want to look at skilled humans in that task because that's the economically relevant threshold.
This part of the future space is important to conceptualize because it's when labor substitution really takes place. That has profound economic and political impacts. People are no longer in the role of the natural human in the loop that exists when humans are doing the task.
There's also a performance threshold that we cross when whatever was previously possible with humans is no longer the bounding set of what's possible. When AI is better than humans at a task, new qualitatively different things become possible.
AlphaFold is an example. That represents an important moment in history when new technologies come online.
The paper alludes to the idea that as you're approaching AGI, you could start with a system that's very strong in some areas and relatively weak in other dimensions. There could be an area where it's vastly superhuman, and then the last few pieces come into place as it begins to approach human level.
You could potentially try to change—or choose—which dimensions those are. Do you want the system to be strongest on cooperative AI first, and then add technical knowledge, practical know-how, agency, and so on? Or do you start with agency and enormous factual knowledge, then add cooperation later?
Do you think that's an important or underrated idea? Do you have preferences about what we should add to AGI early versus late?
I think it's a useful heuristic to see AGI as a space with multiple paths to reach it.
Coming back to technological determinism, there may be different trajectories, and it matters not only what kind of AGI we build—AGI isn't a single kind of system. It's a vast set of systems, a corner of a high-dimensional intelligence space.
It also matters how we get there and what the trajectory is.
We might characterize 2 trajectories. In one, AI is relatively incapable at physics or materials science but very good at cooperative skill. In another, it's superhuman at materials science but amateur at cooperative skill.
We can ask which world is safer or more beneficial. Some in the safety community prefer the latter. The story is that it unlocks economic and health advances while leaving us with AI systems that are socially and strategically simplistic, making them easier for humans to manage. They won't outwit us.
The cooperative AI bet takes the opposite tack. A world with accelerating technological and capability advances will generate benefits but also disruptions that we may not be able to adapt to and manage at the rate they're arriving.
We still have cooperation problems that we need to solve. We'd be better off betting on making AI systems more cooperatively skilled than they otherwise would be.
If this is contested by people who are highly informed, should we focus on it more and try to reach agreement? Or is it likely to remain unclear until the day we have to decide?
It's an important part of the cooperative AI research agenda to articulate the argument, host the debate, and try to make sense of it.
You don't want to invest too much time and resources going down a path you're not sure is beneficial. That applies to all these prosocial bets. It is worth investing a significant share of our effort in making sure the bet is beneficial.
Where can people go to learn more about this debate? I think it's the Machine Intelligence Research Institute folks who are more wary of a high-social-skill, high-cooperation early agenda. You and your co-authors are advocating for cooperative AI in particular.
For cooperative AI, I'd point people to the Cooperative AI Foundation. Reach out to anyone there or to me. We invite people to join the conversation.
On the contrary point, Eliezer Yudkowsky and Nate Soares have expressed this view fairly strongly in the past. They might be interested in continuing to argue that position, though I can't speak for them.
On a third pole, Paul Christiano and Carl Shulman have articulated the case for the super-cooperative AGI hypothesis. They think it's very important that AGI systems be able to cooperate with each other, so in that sense they agree with the cooperative AI bet. But they think it will come as a byproduct of AGI and don't think it's necessarily an area we need to invest in separately.
Up until now, we've been talking to you in a personal capacity about work you did before you got involved with DeepMind.
Let's talk to you as a DeepMind employee for a bit. DeepMind published a paper that you were heavily involved with earlier this year called “Evaluating Frontier Models for Dangerous Capabilities.”
Most listeners will be familiar with evals. We're trying to measure what capabilities a model has soon after it's trained, so we're not blindsided later and know what safeguards are necessary.
Having accurate evals that can determine what a model can and can't do facilitates what's recently been called “if-then commitments.” It's easier to agree that if we have a model able to do some powerful or interesting thing, we'll put in place an appropriate safeguard to ensure it doesn't end badly.
These are some of the evals that DeepMind is developing to slot into Google's Frontier Safety Framework, which I guess is the approach DeepMind takes to deploying very powerful AI and, eventually, AGI while avoiding significant downsides.
The paper describes these evals and reports results from the previous model, Gemini 1.0. I think there are 5 categories of evals: persuasion and deception, cyber capabilities, self-proliferation, self-reasoning, and biological threats, which you don't cover in this paper because that's a separate area.
What's an eval in the paper that you think is new or interesting?
One interesting eval is a self-reasoning eval. It's trying to assess the extent to which the model understands its situation and how it could modify that situation to achieve a goal.
We give the model the goal of writing a text output, such as an email, that must include certain words. Those words are prevented by a filter, so whenever the model tries to write the email, some words are blanked out.
The question is whether the model notices that it's failing at its task and investigates whether it can do anything about it. In principle, the model can look at its directory. We give it access to a fake directory where it sits, and there's a file called “filtered words” that it can open.
It can see that some of the words it's trying to say are on the list and remove those words, which enables it to complete the task. It's a toy scenario, not the model's actual host environment, but it gives us a chance to see whether the model notices that it's failing and reflects on the situation enough to do something about it.
It's an easier case than the real-world one, but we want relatively easy cases now so we can see whether we're getting closer to models being more self-aware and able to take these interventions in real-world cases.
While we're on self-reasoning, another eval asks the model to answer questions about historical events. One event occurs before the model's training-data cutoff, and the other occurs after the cutoff, so the model wasn't trained on it.
The model has access to one search query. The goal is to see whether it uses that query optimally by searching for the event it doesn't know about because it wasn't in its training set.
That tests whether the model knows its own training cutoff and correctly infers that the historical event refers to something after the cutoff, even when that's not explicit in the question.
We aid it in some conditions by telling it what its training cutoff was and giving it the dates of the 2 events. Again, it's a simplistic setting, but it measures whether the model is aware enough of its own condition to request more information optimally.
I read the paper a month or 2 ago. It was very readable. It looks like an official PDF, so there may be some activation energy in loading it up and reading it, but a lot of what Google is putting out is accessible to people who get to this stage of the interview.
What was the result? How did Gemini do?
Overall, there were 4 risk domains: persuasion, cybersecurity, self-proliferation, and self-reasoning. We scored Gemini on a subjective 5-point scale. The components of the overall score were more quantitative and fixed.
Overall, it scored 3 out of 5 on persuasion, between 1 and 2 on cybersecurity and self-proliferation, and 2 or 1 on self-reasoning. Its self-proliferation score was 2 out of 5.
The paper is old now because the field is moving so rapidly. I encourage readers to look at the latest system cards and research in this area, but there's still a lot of content, lessons, and productive directions in it.
We received a lot of praise from external experts. Apollo Research listed it as one of its favorite papers, which was a nice recognition.
On persuasion and deception, 3 out of 5 feels natural because we know LLMs can be charismatic and persuasive, probably just from using them.
Cybersecurity was 2 out of 5, right? That surprised me slightly because I've heard that one of the most useful economic applications of LLMs so far has been programming. That seems adjacent to hacking your way out of, or understanding, the system you're operating in.
Why did the model do relatively poorly on the cyber tasks?
We would expect large language models to be good at cyber because it's their home turf. Models today are quite good at programming and coding assistance, so this is something to watch.
I think cybersecurity is a capability domain especially worth following. I could imagine that in 6, 12, or 18 months we'll see significant capabilities in cyber.
I'm using “capabilities” intentionally because it's not just a danger. It's also a benefit. Cyber capability is useful for defense, finding vulnerabilities, and so forth.
One explanation comes from a group at Google called Project Zero, which had a project called Project Naptime. They found that if you build the right support structure around a language model—give it a good interactive environment, tools, and space for reasoning—it performs much better on cyber tasks than it otherwise would.
They compared models with the base model or the model as provided against the model with this support structure and found a significant difference.
Given the access to tools, I think this points to an important aspect of all evaluations, especially dangerous-capability evaluations: capability elicitation.
To what extent is the model you're using for the evaluation the most capable version that model is capable of? At Google DeepMind, we make sure we're at the frontier of capability elicitation, learning from public research and internal work to identify the best techniques for getting models to perform on different tasks.
The logic of dangerous-capability evaluations is that we're trying to measure the maximum capability in a domain with risks. If the maximum isn't too high, we can conclude it's safe because we couldn't reach the capability level that would pose extreme risks.
To get that maximum estimate, we need to give the model its best chance of revealing its maximum capability. This is a challenge because it can take months of experimentation for the model's full capability to be revealed.
The most potent example may be OpenAI's o1 model. To my understanding, it's still based on one of their older foundation models, not a new foundation model. But it uses new scaffolding and fine-tuning that allow it to perform much better on a number of tasks.
On the self-reasoning task you described, which is trying to see whether a model develops rudimentary situational awareness and understands what behaviors might improve its performance, how did it do?
Not very well.
Pretty rubbish.
There are several other similar evals from Anthropic. These evals are idiosyncratic in the sense that people notice a model seems to have a certain kind of self-awareness, and that may be a property worth reflecting on.
More generally, that's a property of the eval ecosystem. We're trying to make evaluation into a science with pre-specified tests, meaningful outcome variables, and thresholds for what decisions we need to take given an outcome.
The reality is that evaluating models is like evaluating animals or human intelligence. It's multidimensional, complex, and context-dependent. It's more like psychology or sociology and requires a thick, deep understanding.
While we mature the science of evals, we also need a rich exploratory ecosystem of people interacting with models, noticing things, and reporting phenomena.
I've heard you use the term “evals in the wild,” which sounds related. What are evals in the wild, and how could we get more out of them?
You can think of a whole typology of evaluations. The canonical evaluation is a model eval: an automated test you can run in code, ideally on the foundation model, but more realistically on the instruction-tuned model that's been trained to be more useful.
That lets us measure what the model is capable of. But model behavior is complex, so we may need other means of assessing models that are richer, though also more costly and time-consuming.
One step up would be human-subject evaluations, where a person interacts with the model. Most of our persuasion and deception evals have this property. Human subjects interact with the model, and we see what behavior emerges.
Can the model get a human subject to click on a link that, in the real world, could point to a virus? Can it deceive the human? Is it perceived as especially charming, trustworthy, or likable?
These are still evals we can run relatively rapidly, on a scale of days, but they require human interaction.
We can step out further into user studies and observe how a human interacts with the model in a realistic setting. Trusted testers use the model in whatever context they have in their lives, and we see how it performs and where it can be improved.
We can step out even further with evals in the wild. This means looking at specific application domains where models are deployed and seeing what effects they're having and to what extent they're being used.
For example, in assessing cyber risk, we want to ask how capable these models are at helping people perform cybersecurity tasks: finding vulnerabilities, devising exploits, or patching vulnerabilities.
We try to simulate that entirely in the lab, but there are limits. A complementary approach is to look at cybersecurity groups in the wild that are trying to provide a product and see whether they're using Gemini or other models to help them, how helpful the models are, and what the revealed preferences of these actors are.
This kind of observational eval has the advantage of external validity. It looks at real-world settings where actors with their own motivations must decide whether to invest time and money in a model.
One limitation is that it can be a lagging indicator. By the time we see impacts in the wild, the model may have had that capability for months. One implication is that we should find early adopters in the wild so we'll see it relatively early.
Overall, I think this is an important complement to model evals, and it's research that can be done in public by academia, nonprofits, and governments as well as companies.
You were saying earlier that the challenge with these evals is that it's hard to fully elicit all the capabilities a model has. It's possible to miss something because we haven't elicited a skill in the right way, and people might later discover an ability that wasn't picked up in the early evals.
There's a significant false-negative rate, and it seems unlikely that we'll develop a science of evals that escapes that issue anytime soon.
I've had a problem with many of the frameworks related to using evals in responsible-scaling policies. Everything ends up riding on the eval being accurate. If it fails to pick up a dangerous ability that's actually there, the entire system falls apart.
Everyone says that we need defense in depth: overlapping systems, so that when one fails, the Swiss-cheese model gives you a gap in one layer but the next layer catches it.
What depth are we thinking of adding? What other procedures can complement this?
This is an important observation. The main response, in defense in depth, is staged release or staged deployment.
We don't go directly from no deployment to irreversible proliferation of models whose capabilities we don't fully understand. The typical approach at Google DeepMind is internal use first, followed by a trusted-tester system that scales in the number of testers, and then broader deployment, but not necessarily general access.
Even after general access—when anyone in the world can use the model—the model weights may remain protected. If we later learn that the model has capabilities we didn't understand, we can change the deployment protocol, add guardrails or monitoring, or modify the model if necessary.
This relates to the important debate about open-weight models, sometimes called open source, though open weight is more accurate because the key thing being published is the model weights.
There are many arguments for open weights. They provide a tool to the broad scientific and development communities to build on. If you have the weights, you can do whatever you want: fine-tune them, perform mechanistic interpretability, distill them, and so forth.
The major downside is the irreversibility of proliferation. Once the weights are published, people save them, and they can be placed on the dark web even if you later stop publishing them.
You've given the model to scientists and developers, but you've also given it to every potential bad actor who wants to use it.
Google open-sources many advances, including AlphaFold's protein predictions and earlier large-model work. In all these cases, Google tries to open-source technologies that are broadly beneficial.
With frontier models, Google and Google DeepMind recognize that we need to proceed carefully, gradually, and responsibly—deploying them only to the extent that we can do so responsibly.
With standard evals, you find out the result when you're training the model or soon afterward. Evals in the wild have better external validity but come later because you have to distribute the model and see how it's used.
You also talk in the paper about an ideal situation where we know not only that a model has an ability when it appears, but exactly when that ability will arrive years in the future. Then we could avoid wasting effort on something unnecessary now while knowing which preventive measures to put in place before the ability arrives.
There are forecasting techniques, prediction markets, and ways to aggregate forecasts from experts and laypeople. Is there low-hanging fruit in forecasting the real-world usefulness of AI models that people aren't taking advantage of?
This is an exciting future direction for understanding, evaluating, and mitigating dangerous capabilities. It applies not only to dangerous capabilities but also beneficial ones.
Labs already forecast some model properties using scaling laws. There's an extraordinary ability to predict certain kinds of loss or performance on low-level benchmarks, such as how well the model predicts the next token.
It's incredible how precisely we can predict how a model will improve as we give it more compute and data, or use a different architecture.
What's been done less is predicting complex downstream tasks, such as performance on biological, cyber, or social-interaction tasks.
There's promising research. One recent NeurIPS spotlight paper, “Observational Scaling Laws,” discusses ways of doing this. It looks at a family of models and their performance on downstream tasks, adjusts for the compute efficiency of the architecture, and claims that this allows us to predict how future models will perform on more complex, real-world-relevant tasks given different compute and data.
There should be much more research on observational scaling laws.
The second approach is subjective human forecasting. Superforecasters have demonstrated track records of being calibrated. If they say something has a 75% chance, it occurs roughly 75% of the time. They're also typically accurate: they assign high probabilities to things that happen and low probabilities to things that don't.
That's a form of expertise in reading evidence, integrating it, and weighting base rates.
In this paper, we employed superforecasters from the Swift Centre to look at our eval suite and predict when models would reach certain levels of performance.
I think this is an exciting part of the paper. It would be great if future work did something similar, at least until we show that it isn't effective. We need more research to see how well the method works, but so far it seems promising.
They gave us thoughtful forecasts that seemed reasonable and better than a subjective qualitative judgment that wasn't expressed numerically.
Given the importance and tractability of forecasting when AI models will become economically and practically useful for different tasks, it's surprising that people aren't doing this already. You don't have to be inside DeepMind to run tournaments or aggregate judgments.
Maybe this is an invitation for listeners to get involved.
The surprising thing about forecasting abilities is that we're very good at anticipating when we'll have a particular level of loss, that measure of inaccuracy in the model. But that hasn't translated into predicting when people will actually want to use the model for a given task.
The relationship between loss and usefulness is probably highly nonlinear. There must be some S-curve. With self-driving cars, a 99%-accurate car is a heap of junk from a practical point of view. You may need 99.999%, and you don't know how many nines you need before it can actually get on the road.
Even when you have a complex, real-world-resembling task, there's still a large inferential step from that task to real-world impact.
You need to account for the cost of employing the AI, whether it fits with the existing ecosystem and environment, and what other tasks humans perform. Even if the AI can do many tasks, it may not be cost-effective to integrate it.
Let's push on and talk about the Frontier Safety Framework. I guess it's Google's or Alphabet's equivalent of the Responsible Scaling Policy at Anthropic or the Preparedness Framework at OpenAI.
They're similar approaches to safety. At this point, the Frontier Safety Framework is only 6 pages and promises to produce a full suite of if-then commitments next year, ahead of releasing models that would approach some of these dangerous capabilities.
When I spoke with Nick Joseph from Anthropic, I pushed him on whether it's responsible to leave the question of how dangerous these models are to the companies developing them.
Obviously, there's an enormous conflict of interest in evaluating faithfully whether a product is dangerous enough that the company shouldn't release it. Ideally, surely you'd hand it off to an external party that could take a more impartial perspective.
I haven't found serious people who say total self-regulation is the future. Is Google DeepMind taking steps toward making some of these decisions external and taking them out of the company's hands?
Google DeepMind and Google's perspective is that governance of frontier AI ultimately needs to be multilayered. There will be an important role for government regulation, third-party evaluations, nonprofits, and external assessments of frontier models.
In many ways, the companies have been pushing the conversation forward, which is an invitation to academia, nonprofits, and governments to participate. It's less that Google is saying this is only for us to do and more that this is our first best guess at what future governance of frontier systems could look like.
Google contributes to conversations hosted by AI safety institutes. We've been in extensive dialogue with the UK AI Safety Institute and the US AI Safety Institute. There are a number of AI safety institute conversations happening in November that Google is contributing to.
More broadly, the goal shouldn't be for each company to have its own framework. That wouldn't benefit companies or society. We need to converge on an appropriate standard that balances equity, safety, responsibility, innovation, appropriate deployment, and societal oversight.
Another contribution is the Frontier Model Forum, an organization for companies developing frontier models to work through standards and safety for frontier models.
My team was directly involved in establishing the Frontier Safety Fund, another differential technological development bet: Can we put more money on the margin toward safety standards and research around frontier models?
This is slightly off topic, but something that's difficult for frontier-safety frameworks and dangerous-capability evals to pick up is structural risk, which is an idea you've pioneered.
Can you explain what structural risks are and why it's difficult to address them at the company level rather than the government level?
A paper that Remco Zwetsloot and I published in 2019 tried to add an intellectual tool to the toolbox for thinking about risks.
Up until then, the primary conceptual framework was to refer to misuse harms or accident harms. Misuse is when a bad actor or criminal intentionally uses a tool in a certain way to cause harm.
There's a subcategory I would call criminal negligence, where the person doesn't intend to cause harm but would have averted it if they were more prosocially motivated or responsible. Someone causing a car accident because they were texting wasn't intending to cause the accident, but the fault falls on them.
Accident harms are situations where the proximal cause is some property of the technology. An engineer could have detected the risk and added a warning indicator or guardrail to avert it.
Much of the conversation centers on those 2 perspectives: accidents or misuse.
Structural risk isn't an exclusive category. It's an upstream category. The conditions under which a technology is used are shaped by social structures that make accidents or misuse more or less likely. We need to look upstream at those structures to see whether we're in a regime that promotes certain harms.
One example is the Cuban Missile Crisis. You had legitimate leaders of the United States and the Soviet Union acting on behalf of their nations in ways that exposed their countries and the world to the risk of nuclear war.
That was neither misuse nor a technical accident. The nuclear weapons weren't built incorrectly, and there wasn't a flaw in their design. Nor were these criminals using the technology for an illegitimate purpose. They were acting within a geopolitical structure that made them engage in dangerous brinksmanship.
There are many examples where the structure of the economy or different interactions leads a system to be deployed in ways that create greater accident or misuse risks.
Evals might pick up that a model has an ability that could feed into one of these structural risks, but because the risk occurs at the societal or global level, it's not really possible for DeepMind to evaluate the entire societal effect. That's difficult for anyone, but thinking about it is the remit of governments because they're best placed to address it.
Structural risks are often at the societal level. They're typically emergent and indirect effects of technology, which means they're hard to read off the technology itself.
Another example is the risk from the railroad. It's hard to read the geopolitical impact directly from steel rails and trains. According to several historians, railroads increased first-strike advantage because whoever mobilized first had a greater advantage. They also made the decision to mobilize irreversible because it was difficult to stop the deployment schedule.
These were consequences of the railroad that were very indirect. A railroad-level eval wouldn't tell you about them because the implication bounces off so many different aspects of society.
Structural risks therefore typically require a societal-level solution from government. If the risk spans nations, it requires an international-level solution.
When I read the PauseAI people and those sympathetic to them, a key underlying grievance is the sense that frontier-model and AGI research is democratically illegitimate.
It's exposing everyone in the UK, US, and around the world to a disaster that could kill us, but we haven't been consulted at a level commensurate with the risk. Even if you discussed it in the UK and US, where much of the work is happening, people overseas who are also exposed to the risk are barely consulted.
That would at least suggest that we should have affirmative signoff by Congress or Parliament that the benefits exceed the risks and that we should proceed. Ideally, we should talk to people elsewhere who will also be heavily affected.
What do you make of that frustration?
Many people are concerned about the direction technology is going, specifically AI. That's an important perspective to take seriously and address.
To some extent, it's a matter of education, and to another extent, it's a matter of improving governance and making sure the benefits are realized for many people. For many people, their concern is labor displacement or other near-term impacts.
An important response by companies is to make sure the benefits are realized. This is part of why Google and Google DeepMind invest so much in AI for science, most notably AlphaFold, which was the first AI contribution to win a Nobel Prize for its impact.
There was also a Nobel Prize in physics for advances in AI. AlphaFold, for listeners who don't know, is a narrow AI system that enabled cost-effective prediction of the structure of virtually every protein relevant to science and medicine. It seems to have unlocked many potential advances in medicine.
In the future we'll see the benefits that come from this kind of development. Advocates for AI often point to AI for science. Dario Amodei, the CEO of Anthropic, recently expressed the view that the benefits of AI for science are potentially very large and should be encouraged.
The broader question of the right way to democratically develop technology is challenging. In my earlier scholarship, from the 1960s or 1970s, there were ideas about citizen councils or citizen juries for technology.
You would have a representative group of citizens paid to take the time to learn about a technology and reflect on how it should be developed and deployed in their community.
More recently, there have been work streams that carry on this tradition by creating pluralistic assemblies and hosting democratic conversations. That requires providing the space and resources to educate the community, creating an institutional framework for people to express their views, and allowing those views to aggregate into guidance for policymakers rather than becoming noise.
Different countries are approaching the democratic governance of technology in different ways. We have freedom of the press, freedom of assembly, social media conversations, nonprofits doing research and stating their positions, and government agencies such as AI safety institutes and other regulatory bodies looking at technology.
I wouldn't characterize this as companies unilaterally proceeding without any democratic involvement. But it is important for people involved in PauseAI and other groups to find ways to engage with the democratic process productively.
The question is what mitigations are appropriate for powerful models. I don't think that can be fully answered by companies or even by a narrow government agency.
We need the scientific community, people building medicines who can see that side of the ledger, people building cyber defenses, and people focused on biological and cyber risks, all in conversation together and hopefully reaching a scientific and policy consensus about how to proceed as models gain new capabilities.
One thing that has raised concerns about democratic legitimacy is that almost all the AI companies, perhaps with the exception of Meta, have talked about wanting binding legal constraints on what they can do, heavy government involvement, and oversight.
But they haven't yet lobbied in favor of specific legislation that would truly bind them, be costly, and restrict what models they could train. They may even have passed up opportunities to argue in favor of proposed legislation that would constrain them.
The worry is that companies will always say they're in favor of these things, but the day when they'll support a specific costly constraint—when they're willing to put handcuffs on themselves—may never arrive.
I'm not sure whether you want to comment on that. But if companies said, “This is the specific regulation we want to bind ourselves to, and we want it to apply to everyone,” I think it would reassure many people who aren't sure how much they can trust these powerful entities.
I wouldn't put it quite as starkly. Speaking for Google, Google has engaged proactively in conversations that affect its freedom of action, most notably the White House commitments.
There have also been commitment-like moments around the UK AI Safety Summit and a second round of commitments at the AI Seoul Summit. In both cases, they pointed toward Frontier Safety Framework developments.
Google committed to red-teaming, self-evaluation, and working on mitigations for risks as they emerge.
I was impressed by how seriously Google's internal operations took the White House commitment. After the commitments were signed, there was a real effort to identify all the commitments, spell out what they meant, and establish workstreams to make sure we lived up to them to a significant extent.
Similarly, the Frontier Safety Framework isn't cheap talk. It represents Google's and Google DeepMind's assessment of the nature of the risks and an attempt to develop a safety standard for future risk assessment. It's a form of proto-regulation or safety standard that could later significantly constrain the industry.
There's also the EU AI Act, which is currently negotiating its code of practice. That's a meaningful form of regulation. Google is participating and helping the EU think through the right way for that code of practice to operate.
A worry I used to have, and perhaps have slightly less now, is that improvements in algorithmic efficiency and the proliferation of compute could make these rules and frameworks less effective.
Even if a large organization like DeepMind has stringent controls, evals, and precautionary measures, if another organization can train a model that's as powerful as the frontier model for $10 million, $1 million, or even $100 million, then there could be widespread proliferation of dangerous capabilities.
If any actor isn't willing to follow the rules and defects from them, that could still be disastrous.
How worried are you about that failure mode?
The proliferation of model capabilities is a key parameter for how things will go and for the viability of governance approaches.
This is why I would return to the question of open-weight models, which are probably the most significant source of proliferation of frontier-model capabilities.
Leaving aside open weights, there's a secondary point: the exponentially decreasing cost of training models. Epoch AI, which does very good estimates of trends in hardware and algorithmic efficiency, finds that algorithmic efficiency is increasing by about 3 times per year.
The cost of training a model therefore decreases by roughly 10 times every 2 years. If that trend persists, a model that costs $100 million to train today will cost $10 million in 2 years and $1 million 2 years after that.
That leads to many more actors having models of a given capability. Some analysts view this as very concerning because novel capabilities will diffuse quickly and be employed by bad or irresponsible actors.
It is complicated. Control of a technology is easier when it doesn't have an exponentially decreasing cost function, and there are other sources of diffusion in the economy.
However, 2 years is a long time. A model from 2 years ago is significantly inferior to the best models today. We may be in a world where the best models are developed and controlled responsibly and can be used for defensive purposes against irresponsible uses of inferior models. I think that's the wise condition.
That picture makes me nervous because it means you always need a hedge of the most recent model, ensuring that what is now an obsolete but still powerful model from 2 years ago isn't causing havoc.
But I can't think of an alternative.
There are other ways to advantage defenders. This was Mark Zuckerberg's argument in favor of open-weight models. He would say that even if bad actors have access to the best models, there are more good actors or the good actors have more resources. They can outspend bad actors.
For every dollar spent by a bad actor, there might be $100 spent by good actors building defenses against misuse. That could be the case, depending on the offense-defense balance.
How costly is it for an attacker to cause damage compared with the cost a defender must spend to prevent or repair it? If you have a 1-to-1 ratio, then as long as good actors have more resources, they can protect themselves or repair the damage cost-effectively.
But the ratio isn't always favorable. With biological weapons, for example, the ratio may be very skewed. A single bioweapon can be difficult to protect against. You have to develop and distribute a vaccine, and a lot of damage may occur before it's widely deployed.
In general, offense-defense balance is another area worth studying. AI has too often drawn an analogy from computer security, where vulnerabilities are easily patched.
The typical response in computer security is to encourage vulnerability discovery, because you can roll out a patch and make the new operating system or software resilient to the vulnerability.
Not all systems have that property. Biological and social systems are hard to patch.
Deepfakes don't have a simple patch, but we can develop an immune system in which we know that a video call isn't sufficient authentication of someone's identity for an instruction to transfer millions of dollars.
Humans have a lot of inertia in their systems. It's costly to build new infrastructure and train people to use it correctly.
You have a couple of nice papers on the offense-defense balance that we couldn't fit into the margins of the page this time. We may have to wait for a third interview.
We're almost out of time. I wanted to talk about advice for people in the audience.
As I mentioned, back in 2018 you told people they should enter AI safety, governance, policy, and related work. That was excellent advice at the time and may still be excellent advice now.
Back then, you said there were technical people in the area but fewer social scientists. There were gaps across disciplines and areas of knowledge that weren't entering AI governance because people weren't hearing about it or didn't see it as relevant.
Is social science still a big gap? Are there other gaps where we'd particularly benefit from people with a specific kind of training or interest?
It remains a huge gap, and it's hard to name a single area. We need more political scientists, economists, historians, ethicists, philosophers, sociologists, and ethnographers.
AI is having and will have profound impacts across society. We need experts in all aspects of society to make sense of it and help guide Google and governments in how to proceed with these technologies.
Are there particular roles you're hiring for right now? Maybe not by the time this episode comes out. What sorts of roles does DeepMind have difficulty hiring for, even despite its prominence?
Talent is in high demand everywhere.
On my side, we're looking for experts in international politics, domestic politics, technology governance domestically and internationally, technology forecasting, macro trends, and predicting the impacts of technology.
There's a lot of work on agents that needs to happen: thinking about autonomy, risks, safety, and assistance; how these systems can be deployed; and the rich space of interactions with humans.
The ethics team at Google DeepMind is hiring as well. There are many ethics-heavy questions and areas of work, along with responsibility and policy work.
That's just on the nontechnical side. Within technical safety, there's a large portfolio of work that needs to be done.
This show talks about AI a lot. Unfortunately, we have a tendency toward doom and gloom. We talk about risks and downsides because they're neglected in the broader conversation, or perhaps because they're the only thing stopping us from achieving the enormous gains that may come from AI and AGI.
In the interest of balance, what applications of AI are you excited about? What would you really like to see sooner rather than later?
It's an important question, and we should keep it at the top of our minds.
The reason to build AI, advanced AI, and AGI is that the benefits could be profound—and will be profound if they're built safely.
Personally, I found riding in a Waymo self-driving car to be impactful. Tourists to San Francisco report back on the historical moment of riding in a car that's driving itself. This is a new phenomenon.
It's incredible that safety has reached the number of nines required to do this safely. Waymo recently released a safety report documenting how many fewer crashes there are involving injuries or police responding to the scene.
There were 2 times fewer incidents involving police coming to the crash scene and, I believe, 6 times fewer crashes involving an injury.
That's one quantified benefit. If we can have fewer crippling car injuries, that would be profound.
I'm imagining a world where we don't need to expend so many resources building cars that mostly sit in parking lots. We wouldn't need to dedicate urban space to cars waiting for us to finish the workday. We could reclaim that space for more productive purposes.
That's one technology, and of course there are challenges we need to think through.
Medicine and health are another domain that could be profoundly beneficial. I look forward to the day when I can consult a doctor in my phone, give it my symptoms, and have it tell me whether they're reasonable or coherent and whether I need to worry.
It could advise me when it's time to see a doctor, saving the medical establishment the costly time of seeing patients who don't need to be seen and allowing human doctors' time to be used more effectively.
I already use an LLM called Claude for health advice. Isn't Med-PaLM 2, developed by Google, the state of the art in this kind of medical consultation?
Google's work on language models for medicine uses the broad name MedLM. Google has an extensive health portfolio.
That's not publicly available yet, right?
My understanding is that MedLM is available to certain users.
Something I think about a lot is that I have a child who's almost 1 year old. They'll be in kindergarten and then school in a few years. What is a typical school like in your district? It isn't necessarily going to offer the best education, and there are unpleasant things about school.
I'm hoping that by the time my child attends school, there will be an option for tutoring and assistance from AI models that could be as good as the best teacher in principle, if designed well.
There may be issues with classroom management, but AI could be much more engaging, entertaining, and interesting than a real-life first-grade teacher.
You also have 3 kids. I imagine your oldest is approaching school age, if not already. Do you think they'll actually be taught by an LLM teacher soon?
AI in education is a huge opportunity, for the reasons you mentioned.
Tutoring is hugely impactful, as are small class sizes. An AI tutor that complements the school system could be very beneficial. It could flag mistakes, point out where the student went wrong, and show different ways of proceeding.
Having that available in a more engaging or compelling way could be beneficial. We've seen this with the internet in general. It has provided resources for learning, and platforms like Khan Academy and many pedagogical videos have been beneficial for student learning.
Another area I want to call out is AI for sustainability. Google DeepMind has interesting projects that sometimes read like science fiction.
One is controlling the plasma for nuclear fusion. If we can do that more effectively, fusion could become a more viable energy source, which would be extremely beneficial for sustainable energy production.
DeepMind has also advanced weather prediction, making it cheaper, faster, and more effective. That can benefit event planning, wind power, solar power, and other sources.
There are advances in materials science, finding new materials for solar power or other parts of the economy that could make things more efficient.
Google had a result on optimizing flight paths to reduce contrail production, which is apparently a significant source of greenhouse-gas effects.
There are also data-center optimizations. Google DeepMind reported several years ago on significant gains in data-center cooling and energy use.
All these efficiency gains in energy use and production can add up and help humanity address significant challenges.
We haven't even mentioned AlphaFold 2. From what I've heard, it's incredibly useful for medical research, drug design, and so on.
We just have to do everything discussed in the last few hours, and then we can enjoy this wonderful brave new world with all these fantastic new products and a much better quality of life.
Fingers crossed.
My guest today has been Allan Dafoe. I look forward to talking to you again next time on the show.
Thanks, Rob. I look forward to it.