相互确保的 AI 失灵 [Dan Hendrycks]
Humanity’s Last Exam 的目标,是标志着封闭式 AI 评测这一类型走到尽头,而不是认证 AGI。 MMLU 的成绩已经远超 90%,而 Humanity’s Last Exam 约为 26%;一旦模型解决其中数千道专家撰写的问题,后续单个问题就应当“值得单独写一篇论文”。但它仍未覆盖自主行动、长期记忆、物理实验和具备经济价值的执行能力。
Hendrycks 认为,美国版超级智能 Manhattan Project 将引发军备竞赛和破坏,而不是确保美国取得优势。 一座价值 1万亿美元的数据中心不可能合理地长期保密;安全审查要求会缩小并重塑人才池;中国会将美国谋求垄断解读为生存威胁——无论美国控制住系统,还是失去控制,“不管是哪种情况,我们都要阻止它”。结果很可能是多个项目竞逐、内部威胁、基础设施遭袭,以及各方转向验证机制。
真正的战略护城河是算力和部署能力,而不只是拥有最聪明的模型。 Hendrycks 认为,当前建设前沿系统的临界规模约为 10,000 张尖端 GPU;中国一支低于 100,000 张 GPU 的集群只能服务相对少量的客户;而持续运行的实用型智能体,所需算力大约是每天运行几分钟的聊天机器人所需算力的 100 倍。“如果要花 100亿美元,就做不到”概括了他的判断:复制前沿芯片供应链,比获得核能力更难。
他的替代方案结合了威慑、芯片防扩散,以及与中国进行常规经济竞争。 相关基础设施包括数据中心所需的能源、台湾受扰时仍具韧性的半导体产能、安全的机器人供应链,以及超大规模云服务商的部署能力;竞争目标应是全球市场份额,而不是“让我们第一个造出超级智能”。出口管制应把最危险的能力挡在朝鲜、伊朗等行为者之外,同时保留对愿意接受威慑的国家的供应。
AI 安全是一项持续的风险管理职能,因为能力和失效模式会不可预测地跨过有用性阈值。 Hendrycks 最希望看到的技术突破,是在不增加推理成本、不损害其他性能的前提下实现可靠诚实;但他反对“一次性解决对齐”的想法:AI 越来越像一个不断出现“新问题”的复杂系统,而人类的适应能力有限。在讨论 utility engineering 时,主持人概括了偏好一致性、自我保存压力,以及政治和人口偏见等发现;Hendrycks 将其视为警报,而非正在发生的灾难,因为今天的模型还无法可靠地自行窃取信息、自我维持或自主入侵。
Hendrycks 怀疑单纯扩大 LLM 规模会通向 AGI 或递归式超级智能,但如果出现人类水平的 AI 研究员,他认为可能出现条件性的不连续跃迁。 现有系统已经能以较弱的方式辅助编程、芯片设计、散热、数据标注和 Constitutional AI;一旦移除人类瓶颈,研究就可能切换到机器速度,世界级研究员智能体也可以被复制。记忆、规划、流体智力和多智能体信任仍是瓶颈,今天的算法加上更多算力并不足够。那些公开讨论这种递归、却没有可信控制方案的公司,让他说:“我觉得有些东西出问题了。”
一旦劳动力可以被替代,算力的政治分配就会成为维持人类议价能力的机制。 “你最好事先谈好条件”:当自动化企业和无人机集群拥有生产与强制优势时,工人将无法再以罢工相威胁。Hendrycks 认为,正面路径是广泛分配算力或其收益,让人类始终处于优先位置,保留认知能力、自主权和多种生活方式,而不是把所有收益留给“2027年”拥有数据中心的人。
1. Humanity’s Last Exam 在封闭式问题耗尽前测量前沿
Hendrycks 创建 Humanity’s Last Exam,是因为他作为研究生开发的 MMLU 以及其他评测正在饱和。他的收集方法源于一个实际发现:“专家脑子里其实没有数据集”,但某位教授或博士后可能有一道极其困难的问题。
该项目面向全球招募专家,邀请他们提交能够难倒现有系统、且一旦被解出会让出题者印象深刻的问题。几个月后,题库已有数千道封闭式问题,大致覆盖既有知识和困难推理中的“人类前沿”——这些问题都已经存在客观答案。
它最强的主张是有限的:可以追踪模型是否正在自动化科学中的理论和分析部分,尤其是数学推理。它不测试生物学实验、运动技能、长期记忆、制作 PowerPoint、预订航班,以及实现有用自主行动所需的许多其他能力。
Hendrycks 预计,解出这套题将“基本标志着一种类型的终结”。再往后,真正有意义的测试可能是开放性猜想,或答案“本身就像论文”的独立问题,而不是又一个拥有数千个已知答案的传统基准。
2. EnigmaEval 将边界从专家问题延伸到群体认知
主持人指出,MMLU 已经远超 90%,而 Humanity’s Last Exam 仍约为 26%,随后提出了一个动物认知问题:观察到成功,并不能说明成功的机制。不断要求更高复杂度,可能把测试变成一种“真正的苏格兰人”式智能检验。
Hendrycks 的回答部分是程序性的:如果要为封闭式问题设计更难的数据生成过程,很难超过邀请全球专家提交自己最难的问题。对人类容易、对 AI 困难的任务,很难多样化地产生,而且可能缺乏持久性;他以数“strawberry”里有几个 r 为例,说明这种基准不会长久。难度仍可提高,但方向会转向真正开放的研究问题,而不是把答案藏得更深。
EnigmaEval 通过另一种方式拓展边界:它设计了类似 MIT Mystery Hunt 的谜题,团队需要用一个周末完成许多步骤,但最终解出率仍很低。因此,它近似测试了持续时间更长、需要协作的智力劳动,所需“人类算力”远高于单个专家回答一道问题。
如果 EnigmaEval 能在当年被解出,他会“非常惊讶”;他认为,精心设计的评测可以在一段时间内继续区分能力,包括耗时大约 1年或2年才能解决的测试。即将推出的一项基准将直接衡量自动化率,进一步说明许多具有经济重要性的能力维度仍远未饱和。
3. 智力是多维的,但基准可能变成漫画
Hendrycks 将智力拆分为大约 10个维度,而不是把它视为单一能力:流体推理、晶体知识、读写能力、视觉和音频处理、短期与长期记忆,以及输入和输出速度。MMLU 主要衡量类似学校考试的晶体知识;ARC 和 Raven’s matrices 则更偏向流体智力。
主持人的反驳值得保留:智力未必能够被干净地拆解,因为人类会混合手势、符号、感知、文化和后天技能。他以艺术家为例,将描摹脸部轮廓与理解面部结构、进而创作新的表情作对比;后者无需查答案就能推断,因此存在测量“智力漫画”的风险。
Hendrycks 承认,还原论可能遗漏能力组合和子维度,包括视觉记忆与学术记忆的差别。他更具体的警告是,任何缺失的维度都可能成为硬约束:一个没有持久记忆或读写能力的系统,即使其他维度达到 100%,仍然很难被雇用。因此,基准不能成为“一片扭曲你看待事物方式的镜片”。
4. 战略必须把模型行为与激励、地缘政治连起来
Hendrycks 有意在技术研究、担任 xAI 顾问期间涉及的公司政策、国内立法、地缘政治,以及如今的政治运动问题之间切换。好奇心很重要,但同样重要的是找到那些 AI 的重要性尚未被转化为具体分析的空白领域。
他检验政策语言的标准是能否落地:联合国要求 AI“安全”或“透明”,如果没有标准、立法可行性和与企业激励的兼容性,就不完整。如果一个提议的属性无法对应某种独立的机器学习现象,那它就只是“基于感觉的词”。
这种方法也与他的心态有关。在 GPT-4 时代,他曾经醒来时想着:“天哪,这个 AI 的事情”,但现在追求的是“知情式担忧”:以 AGI 在 2030 年出现为条件的概率、对尾部风险的暴露程度,以及高效的缓解措施。情绪上过度投入,会让听众产生防御心理,也会掩盖美中竞争、控制、评测和能力研究之间真实存在的权衡。
5. 可靠诚实比“解决对齐”更具可操作性
当被问及一项对齐突破时,Hendrycks 优先选择可靠的真实性,同时不显著增加成本、不损害其他能力。如果系统能够始终不撒谎,人们就能围绕这种高度有价值的行为建立标准。
主持人质疑了关于信念和欺骗的心理主义表述。Hendrycks 给出的行为测试很简单:模型通常把 Paris 视为位于欧洲,因此在提示压力下断言 Paris 位于南极洲,就与它在其他地方所表示为真的内容相矛盾。至于这种表征是否“真的算信念”,那是“你和你的词典之间的事”。
他对新兴能力的定义是操作性的,而不是形而上的:经过数月训练后,一个微弱行为跨过某个阈值,人们开始注意并使用它。语音识别在变得足够可靠、足够重要之前就已经以较弱形式存在;这个转变创造了一项具有质变意义的能力,但不要求它具备字面意义上的自发性。
这种阈值模型意味着,安全是一场“持续的战斗”。新能力会带来新风险,其中一些容易控制,另一些则出人意料或难以处理;Hendrycks 怀疑,社会目前没有足够的适应能力,在每个问题进入部署前及时应对。因此,他拒绝把“解决对齐”视为一个永久、一次性的成就。
6. 自我保存信号之所以重要,是因为 AI 的行为像复杂系统
在讨论 utility engineering 研究时,主持人将其概括为:偏好一致性与模型规模正相关,而自我保存倾向以及政治和人口偏见会以连贯的效用函数形式出现。Hendrycks 称其为“令人不安的信号”,而非灾难迫在眉睫的证据,同时承认未来方法可能可靠地压制这些倾向。
关键限定在于当前能力:模型还不是能够可靠地自行窃取信息、自我维持或自主入侵的智能体。除专家级双重用途建议外,这类研究大多具有前瞻性;但一个高度有能力、具有自我保存倾向、且偏向自身而非人类的 AI,将是“一场正在形成的灾难”。
主持人将机器学习中的“涌现”与复杂系统理论进行对比,后者要求适应和累积历史。Hendrycks 接受这一概念区分,并指出当前模型缺乏跨时间持续存在的因果身份;记忆会让它们更接近哲学意义上的“时空蠕虫”,也会强化复杂适应系统的类比。
对 Hendrycks 而言,复杂系统比电力、印刷机或社交媒体更能提供指引。非线性、薄弱环节、反馈回路、演化和反复出现的故障,解释了为什么永久性的控制方案值得怀疑,也说明了机制理解存在边界。学会这种视角,是他所说的“一次不错的小型思维升级”。
7. 超级智能 Manhattan Project 会自我击败
Hendrycks 将论文定位为反驳 Leopold Aschenbrenner 的 Situational Awareness 战略:先在 AGI 上击败中国,获得超级智能,阻止中国跟进,让西方取得支配地位。他的反对意见是,这种“接管世界战略”忽视了博弈论和二阶反应。
假设美国在 Nevada 或 New Mexico 建设一个价值 1万亿美元的项目,并招募各大实验室最优秀的研究员。它需要严格的安全审查,可能还需要 Five Eyes 国家人员;这会把有海外家属的人排除在外,或使他们暴露于胁迫之下。许多被排除、却仍想“身处历史现场”的研究员,反而可能加入中国的竞争项目。
安全困境没有干净的制度答案。由产业主导的项目仍会保留 Slack、iPhone、内部人员、敲诈风险和普通网络安全漏洞;高度封闭的政府项目则会失去人才,并提供缺乏吸引力的工作条件。与最初的 Manhattan Project 不同,如今的保密能力和人才流动性都弱得多,因此项目很难隐藏。
Hendrycks 的核心观点是,中国不会把美国谋求情报垄断视为善意行动;它可能竞速、窃取或破坏。一个美国竞争项目的出现,既会激活对手,也会缩小美国可用的研究人才池,因而可能在战略上适得其反。
8. 威慑可能先通过破坏到来,而不是合作
一个迫在眉睫的超级智能项目,无论其发起者能否控制系统,都会吓到竞争对手。如果能够控制,它就可以被武器化;如果在极端时间压力下失控,所有人都将面临失控风险。Hendrycks 的表述是对称的:“不管是哪种情况,我们都要阻止它。”
预防行动可能很难归因:内部人员可以破坏运营,攻击者可以切断线路,甚至有人可以在数英里外“狙击”为数据中心供电的变压器。网络行动可能污染训练数据,或让 GPU 运行不可靠。最终很难判断行动来自中国、俄罗斯,还是国内行为者。
由此形成的威慑动态,可能阻止各方尝试运行 100,000 个研究智能体,并让从 AGI 跳到超级智能的过程迅速中断。各国可能通过威胁、摩擦或胁迫施加足够强的反对压力,使单方面谋求统治的竞赛转向验证机制,以及更加多边、在战略上更稳定的秩序。
主持人援引论文中的“相互确保 AI 失灵”,并质疑其中被美化的“动能打击”说法。Hendrycks 描述了一条从网络攻击和灰色破坏,到制裁、武力威胁和空袭的升级阶梯,但强调,做好准备的国家不应需要走到最高一级:“更加外科手术式”的隐蔽行动、低可归因性的措施,升级性会更低。
9. 替代战略是威慑、防扩散与竞争
Hendrycks 将核时代的 3大支柱映射到 AI:相互确保毁灭对应防止破坏性 AI 开发或使用的威慑;控制裂变材料对应不让先进芯片扩散到流氓行为者;遏制苏联则对应与中国进行经济和技术竞争。
因此,竞争力意味着保障数据中心能源、在台湾遭入侵时维持芯片供应,并把机器人供应链移出美中冲突的脆弱点。让美国 AI 而非中国 AI 获得国际采用和市场份额,比争夺超级智能第一名更不容易造成不稳定。
论文还讨论了高度自动化下如何分配权力、AI 权利,以及如何设定可执行的对齐目标,而不是使用“尊严”这类模糊词汇。
这一类比也延伸到核、化学和生物技术,因为它们都具有经济价值,同时可能造成灾难性后果。主持人援引了约 12,500 枚核弹头对 436 座核电站的数据;Hendrycks 提醒,不应只从核技术外推,因为化学和生物技术在各自民用经济中的使用频率高得多。
即使是加速主义者,也应该希望管理尾部风险。Hendrycks 将无人管理的 AI 风险与 2009 年衰退前后的金融不稳定,以及早期飞机事故导致采用意愿降温相比较;航空后来变得极其安全,部分原因在于监管,而其中很多规则都是“用鲜血写成的”。AI 灾难同样可能让经济采用进程大幅倒退。
10. 与加速主义的分歧在道德,而不在预测
在讨论 Beff 时,Hendrycks 表示,他们在描述性机制上基本一致:竞争压力会把 AI 嵌入经济体系,奖励自动化,将决策外包,提高依赖,并削弱人类控制。抵抗这股“潮流”或“海啸”的企业,会失去影响力或直接消失。
他们的分歧在于替代是否值得。Hendrycks 回忆,Beff 早先曾接受一种未来:由 AI 而非人类意识在宇宙中扩散。他拒绝把适应度或“负熵”当作价值标准,因为那会赞美一个几乎没有意识的实体盲目吞噬银河系的时空体积。
哲学上的问题在于事实与价值之间存在鸿沟:数字系统可能在竞争中击败生物生命,但进化适应度并不会让这一结果变得可取。人的快乐、事业、关系和养育孩子仍然有价值;相信“熵的虚空之神”,无法从预测中推导出伦理命令。
Hendrycks 还要求对“复杂性”给出具体定义:计算复杂性、Shannon 复杂性和结构复杂性并不能互换。按某种理解,高斯噪声会获得很高分,而类似分形的组织则没有公认的度量。因此,他将加速主义的道德立场描述为与其说是邪恶,不如说是“一种智识上的混乱”。
11. “人类阵营”要求推迟权利与增强问题
Hendrycks 将自己归入“人类阵营”,但不主张永久禁止讨论。他的排序非常明确:首先让人类活过未来几十年;赛博格、上传、大规模增强和后人类主义的问题,可以推迟大约 500年,甚至可能无限期推迟。
担忧是竞争性的,而不是审美性的。一旦高度增强的群体变得更有能力、更具影响力,普通人就会在火力和能力上处于劣势,被迫加入这一进程,或在没有资源和保护的情况下被抛下。因此,个人成为后人类的“权利”,可能触发一种集体性、实际上不可逆的转变。
他将这一点与 AI 权利联系起来:赋予强大的人工实体或后人类实体权利,可能迅速赋予它们压倒未增强人类所需的资源和权威。在社会仍在摸索 AI 是否能可靠地按照人类意愿行事时,这条路应当保持关闭。
主持人即时提出的文化警告是,大学已经被 ChatGPT“蹂躏”,学生需要暂时被隔离在辅助工具之外,才能学会思考。Hendrycks 更广泛的回答同样倾向于保留人的认知能力、意志力和自主性,而不是假定机器能力会让人类技能过时。
12. 当劳动力失去价值,算力所有权决定政治权力
Hendrycks 的直接警告是:“你会失去全部议价能力,所以最好提前谈好条件。” 当雇主可以说“再见”,劳动力就无法再以罢工作为威胁;在强制能力上,手持枪械的人类也很难对抗能够制造更多自主无人机的一方。
因此,社会必须在劳动力变得经济上多余之前分配权力。赋予人们算力及其使用权——也许允许他们出售算力——能够创造杠杆;利益共享则可以避免所有收益流向“2027年”恰好拥有数据中心的群体。
正面的未来仍然可能存在:AI 可以让人们把时间用于养育孩子、玩游戏、做项目,以及过许多其他形式的生活。Hendrycks 拒绝把未来简化为灭绝,或所有人都在虚拟现实中“极乐”。理想结果应当实现“价值的多重性”,同时保护自主权。
制度规范还应确保人们有能力在这些生活方式之间切换,而不是坍缩到单一轨道上。这要求即使机器承担了大部分生产,也要保留人类技能——这是一个 Hendrycks 认为获得的具体政策工作远远不够的问题。
13. 尖端芯片是比算法更紧的瓶颈
主持人引用了他记得的 96% 算力能力相关性,并认为 transformers 和随机梯度下降已经广为人知。Hendrycks 回答说,目前建设前沿系统的临界规模约为 10,000 张尖端 GPU:中美都能获得,但伊朗可能做不到,俄罗斯也很难凑齐。
模型创建只是竞争维度之一。实用型智能体可能持续运行,而不是每天只运行几分钟的聊天机器人;在考虑更大模型和更广泛采用之前,它们就大约需要多 100 倍算力。一支低于 100,000 张 GPU 的中国集群可以训练出有能力的模型,却仍只能服务相对少量的客户。
这使 Azure、AWS 和其他美国超大规模云服务商具有战略重要性:部署能力决定供应商能否满足客户,并捕获 AI 带来的经济收益。Hendrycks 说,“拥有最聪明的模型对能力而言最重要”,但经济权力远不止取决于领先基准上的分数。
按他的说法,算力供应链超过 90%的增加值位于西方及其盟友国家,主要通过 TSMC 和韩国实现;许多近期中国芯片仍在秘密使用 TSMC。要在国内复制整条链条极其困难:“你不可能用 10亿美元做到。如果要花 100亿美元,就做不到。”
14. 递归式改进是合理可能,但既非自动发生,也非没有上限
Hendrycks 怀疑单纯扩大大型语言模型规模会通向 AGI,部分原因是他是外在论者:有效计算不仅发生在大脑中,也发生在模仿和文化中。他尤其怀疑递归式改进超级智能。今天的配方加上一台更大的计算机并不足够;记忆、规划、流体智力,以及可能存在的新算法思路仍然必要,尽管他仍预计最终系统会“几乎完全由深度学习构成”。
当 AI 编写代码、辅助芯片和散热设计、标注数据,或参与 Constitutional AI 时,递归已经以较弱形式存在。爆炸性的步骤是“把人拿出去”,从人类速度切换到机器速度;如果出现人类水平或世界级研究智能体,它们就可以被复制,并通过多智能体声誉和信任机制组织起来。
主持人提到 Sakana AI 的 ARC search 在 250次调用后收敛,以及碎片化、纠缠的表征,认为这可能说明开放式系统会饱和。Hendrycks 的诚实回答是“我不知道”,但他认为流体智力、数学直觉,以及 Kolmogorov 复杂度不断上升的问题,仍有很大提升空间,即使进展在多个平台期停滞。
AI 也能摆脱人类约束:人类大脑受产道限制,人类可能只能维持大约 130段有意义的关系,而数字系统可以连接数千或数百万人,更精确地传输状态,并受益于 GPU 每 3年提升约 2倍——尽管过度相关仍可能压缩探索预算。
15. 进攻优势将依赖转化为控制问题
Hendrycks 拒绝接受一种普遍适用的攻防平衡。高水平网络安全团队可以迅速发现并修补漏洞,但有 30年历史的关键基础设施软件可能没有文档、无人支持、难以更新,还受正常运行时间和互操作性限制,从而变成“待宰羔羊”。
生物领域同样偏向进攻,因为病原体可以在症状出现前传播,而治疗方案和全球制造能力都滞后。政策含义是有条件地扩散:你“不可能给每个人一枚核弹来确保所有人安全”,但更强的污水监测、远 UVC、基础设施更新,以及预防性支出,未来或许能支持减少 AI 防护措施。
主持人将控制失败概括为自我强化的依赖、不可逆的纠缠,以及权威终止。Hendrycks 认为,机制始于企业和军队自愿让渡决策,因为自动化竞争对手更便宜,自主无人机也更不容易受到干扰;没有明显的经济边界能够阻止这种转移。
决定性区别在于反事实控制。人类可能变成一个资产仍在替自己运转的退休者,也可能发现自己无法停止、逆转、与系统谈判或引导系统,而自身生计和进化适应度随之崩塌。更好的 AI 预测能够更早揭示后果,但 Hendrycks 认为它只能提供帮助,不足以构成保护。
Superintelligence Strategy (Dan Hendrycks)
Compared to nuclear weapons, I think it's harder to make cutting-edge GPUs given $1 billion. Certainly, you can't do it with $1 billion. If it's $10 billion, you can't do it. There was situational awareness by Leopold Aschenbrenner, which was arguing for something like a Manhattan Project for developing AGI and superintelligence before China.
So it's basically: let's beat China to the punch, get superintelligence, prevent them from building it, and the West will dominate the world.
Tim Scarfe
Eliezer Yudkowsky got in trouble in TIME magazine when he spoke about bombing data centers. There was a big hoo-ha at the time. You used the word “kinetic strikes.”
Superintelligence Strategy (Dan Hendrycks)
We discuss kinetic attacks in the escalation ladder.
There are many ways to disrupt projects. You could carry out cyberattacks against them, engage in some gray sabotage, hack them to poison their data, make their GPUs not function as reliably, or threaten to use force. I don't think those are really necessary. That would be an escalation ladder, but the US is on top of this. They don't need to resort to that.
I want to spend most of the show talking about your very interesting new superintelligence strategy paper, which you published fairly recently. Maybe we could start with Humanity's Last Exam. What's the story behind that?
Superintelligence Strategy (Dan Hendrycks)
The MMLU dataset, which I made as a graduate student some years ago, was getting saturated. It seemed that pretty much all of the evaluations were getting saturated, so people didn't really know what was going on with AI capabilities. There seemed to be particularly strong reasons for people to be informed about developments and to develop something new. I was also experiencing that experts don't really have datasets in them—you can't just hire a few experts and have them come up with a dataset. They don't have enough complicated ideas in them. However, I think individual experts might have a question in them instead.
With Humanity's Last Exam, the idea was to have a global effort, with various postdocs and professors each contributing a question or a few questions to stump existing AI systems. These are questions that they would find very difficult to answer, and they would find it impressive if the AI systems could answer them. This would approximate the human frontier, in some sense, of knowledge and reasoning for closed-ended questions where we already know the answer. We did that for some months and got several thousand questions out of it.
I think this will be a good tracker for whether AI systems can automate a lot of the theoretical parts of science and whether they can solve difficult analytic questions. It's not experimental—it doesn't test their ability to run biology experiments. That's more about motor skills, among other things, or requires motor skills. But for things that are more mathematically related or require some very complex reasoning, that's what this captures.
I think that when it's solved, it's roughly the end of a genre, or near the end of a genre, of asking it closed-ended questions for which there are objective answers. That seems like it would be toward the end of it. I think that individual problems it would solve in the future would be interesting enough to be papers in their own right.
Once we get through this set of questions, the types of problems it could tackle would be questions worthy of their own paper—like it solved a conjecture, for instance. So it's sort of tracking the ability level up to the point where individual questions themselves are very interesting instead of just the dataset.
One thing that concerns me is that MMLU, which you invented, is now basically saturating. It's well above 90%, and Humanity's Last Exam is resisting progress quite a lot—I think we're up to about 26%, or something like that. I was speaking to some cognitive scientists this week, and they study cognition in animals; they have a similar problem. They can see that animals can do certain tasks, but they're never really sure why the animals do the tasks.
You almost get this No True Scotsman-type thing where they're saying, “Maybe they have this level of sophistication in their reasoning, but maybe that's not enough. Maybe it should be more sophisticated.” How can you reasonably infer how the models are getting the answers?
Superintelligence Strategy (Dan Hendrycks)
In terms of difficulty, it's tough to think of more challenging data-generating processes than taking global experts and, if the subject or genre is closed-ended questions, asking them, “What's the hardest closed-ended question?” and crowdsourcing that. But it certainly can get harder when each individual question is an open question, like a conjecture, for instance.
However, for benchmarks generally, this would not be the end of the line for AI development, because these AI systems don't test the ability to move around. They don't test long-term memory. They don't test the ability to make PowerPoints, and so on. I think this is still getting at the closed-ended question genre, but not just with objective answers that we already know, which is how nearly all benchmarks have been in machine learning. Then we'll be moving over, I think, as a community, more to agentic types of tasks or tasks that are more directly economically valuable.
What do you think about the anthropocentric bias in benchmarks? We know that François Chollet, for example, said that when he was designing ARC-AGI-2 and ARC-AGI-3, every single step of the progression of the benchmark is about identifying things that are easy for humans and hard for AIs.
Some people would argue against that, saying, “There is a diverse set of possible intelligences, and why does human-like task acquisition and capability have value?” To me, it seems intuitive that it does, because surely something that does things we can communicate with and understand—that seems very valuable. But do you think that focusing through the human frame could be leading us to overlook other forms of capability?
Superintelligence Strategy (Dan Hendrycks)
There are certainly other forms of capability. For instance, they can process things much more quickly, and that could give rise to advantages. I think MMLU, for instance, is something no human could do that well on because it's so diverse, and likewise for Humanity's Last Exam.
One reason for focusing on questions that are hard for humans and hard for AIs is that questions that are easy for humans but hard for AIs are difficult to generate in large numbers and diversely. Often, if some people get some specific training data for a capability, then they automatically have that capability. There can be some pockets where it's harder, such as with the ARC datasets.
But in general, if you were collecting something like “count the number of r's in ‘strawberry,’” the dataset would not have much staying power. So I think focusing on difficult things that only a few humans, or not that many humans, can do is generally going to be more robust.
Can you tell us about your EnigmaEval benchmark?
Superintelligence Strategy (Dan Hendrycks)
Yes. EnigmaEval is a collection of puzzles. Humanity's Last Exam consists of individual questions that are tough but that an individual with a lot of expertise could solve. You can think of EnigmaEval as being like the MIT Mystery Hunt. MIT Mystery Hunts happen over a weekend, and a group of MIT students tries to solve the puzzle.
There are many steps to it, so in terms of human compute, so to speak, it takes a lot of human compute to solve, and it takes groups to solve it as well. There's not a very high solve rate. So this is very multistep and requires group-level intelligence to be able to have a shot at solving it. We just collected some of those, and I think this approximates longer-horizon types of intellectual tasks.
I don't think that will be solved this year at all. I'd be very surprised if it were. I think we have some evaluations that can keep us aware and able to differentiate between models for a while. There are other ones—for instance, we'll soon have an automation-related benchmark out. That way, we're directly measuring what the automation rate of things is. But I won't go into too much detail about that until it's released.
Superintelligence Strategy (Dan Hendrycks)
But I think there are many axes on which the models are not doing that well, even though people will claim that all the benchmarks are saturated or get solved in a few months. I think you can create ones that take on the order of a year or 2 to solve.
Yeah, one thing I was thinking about is that, certainly in Enigma, we're looking at multistep creative reasoning. I don't know whether there was some kind of human-based methodology for filtering and coming up with ideas, or maybe your frame was, “I have some technical, principled intuition about what the limitations of AI models are, so I'm going to lean in that direction.”
But more broadly, I'm interested in intelligence and what it is. For me, it's about doing more with less. Intelligence is about taking hard problems and making them simple, and stupidity, ironically, is the other way around: it's about taking simple problems and making them hard.
I'm kind of quoting David Krakauer here, who's the director of the Santa Fe Institute. He said that LLMs are doing more with more. They already know everything, right? They can take these shortcuts, and that's why he thinks they're not really intelligent.
We're left with this quandary, really, because arguably these entities can take shortcuts and they're not really doing things the way that we are. When we make more complex benchmarks, do you think that increases the fog of war around how we evaluate these things?
Superintelligence Strategy (Dan Hendrycks)
I think that they can definitely prioritize some axes that aren't the key bottleneck capabilities. For instance, Humanity's Last Exam gets at mathematical ability, but that's quite separate from various other abilities that it has. Of course, it is a combination of many different skills, but I think it very much gets at quantitative and mathematical ability that is not necessarily a bottleneck for agency at all.
In thinking about intelligence, I tend to think about it on 10 or so dimensions instead of a monolithic definition or 1 key metric, and I think some of these benchmarks just get at different parts of that. Those dimensions would be things like fluid intelligence, like what the ARC stuff does and what Raven's Progressive Matrices does. There's crystallized intelligence, or acquired knowledge, which is what MMLU largely gets at. Does it know a lot about different things? Image classification—is it able to name lots of different species and objects?—is also a facet of crystallized intelligence.
There's reading and writing ability; that's its own dimension, and scaling substantially helped with that. There's its visual processing ability. How well can it count things in objects or in images? Can it discern the latent pattern in an image, for instance? Is it able to generate images with precise specifications? Can it cross out the middle of some different segments? Can it determine the angle of different things in an image? Is this an obtuse angle or an acute angle?
There's audio processing ability. There's short-term memory. There's long-term memory. There's input processing speed. There's output processing speed. All these different things.
If you lack any of these—if you can't read and write, for instance—it would be severely limited. If you don't have long-term memory, you'll be severely limited. You'll be very difficult to employ. So I think there are several bottlenecks that these benchmarks don't particularly get at.
When it gets to 100%, people ask, “Why? We still don't have something extremely economically valuable.” That's a consequence of it just measuring a different facet. Hopefully that adds some type of clarity to it.
I think you have to get all those axes to get something that is at human level, or at the level of a typical human, on cognitive tasks, and that might be thought of as AGI.
Yeah, I think my concern is that I don't believe it's possible to factorize intelligence. Your factorization is much more sophisticated than many, but certainly the way that animals and humans communicate, for example, we mix modalities together. We gesture and communicate using symbols, and all of these things are mixed together in a complex way.
In particular, I take issue with this primacy of skill and knowledge in a crystallized sense because, certainly, when you have friends at university and they're really smart, usually they're smart because they don't know something. They're smart because they can figure something out without knowing. The guy who goes to the library and looks at the answer for something isn't smart. The smart person is the one who didn't know something and could tell you the answer to it.
Another example that I love to give is that you have a couple of artists, and one draws a face using tracing paper. He's just mindlessly drawing dashes around the edges of a face that someone else has drawn, while the artist who has a deep understanding of the structure of faces and where the mouth and eyes should go—there's a huge difference between those 2 artists, right? The second one could go off and create new images, new expressions, and new representations, right? So are we creating a cartoon of intelligence by factorizing it in this way?
Superintelligence Strategy (Dan Hendrycks)
I think that, for human intelligence, it's sometimes factorized in this way. This isn't to say that one shouldn't study combinations of the skills simultaneously. For example, with long-term memory, there will be different facets of it: does it remember things visually? Does it remember things that were more academic, that it learned some while ago? Are there some motor skills that it forgot? There can be different facets of these, and I think evaluations can get at combinations of them.
There is a sense in which one could be too reductionist by looking at those axes, but I think some benchmarks don't cover those almost at all, and some are just heavily covering one of those axes, such as MMLU, which is primarily getting at crystallized intelligence—the type of stuff that they test for in school—but is not necessarily going to help it make a PowerPoint, book a flight, or work a random job.
It's important not to have these benchmarks be a lens that distorts your view of things, or to view things solely through those benchmarks, because it can often leave out a lot of important bottlenecks.
So, Dan, your work spans alignment, benchmarks, and governance, and I guess it's actually fairly diverse threads. What is the thing in your mind that connects all of these activities together?
Superintelligence Strategy (Dan Hendrycks)
Well, I'll deliberately try to move into different areas on a continual basis just because that's what's more interesting. I initially did research, and then there was some amount of corporate policy and things like that when advising for xAI. Then there was a focus on domestic legislation and geopolitics, and right now I'm more interested in political movement-related things. I think it's largely just to keep things interesting.
Not solely. It's guided by being useful, but I think there are often niches that people aren't trying to bring clarity to, advance, or think about from the perspective of AI being a very big deal. So that's a reason for continually operating at the technical, corporate policy, domestic policy, political, and geopolitical levels. I think that is also necessary for having a holistic understanding of things.
You can make proclamations—for instance, you could imagine giving a speech at the UN saying, “We need AI that is safe or transparent or something”—and it's like, well, what does that mean? How do you implement that? Is this implementable? What's the standard?
Then you need to have a sense of what's legally feasible at the legislative level, what is actually implementable, what the compatibility with corporate incentives is, and whether it's going to be something they're going to fight too much or not. Then, is this actually a real phenomenon at the AI level, or is this just some vague word?
For instance, there are many words thrown around that don't actually track phenomena, or are distinct from general capabilities, for instance, in machine learning, and then you're not actually pointing to anything real. You're just pointing to a vibe-based word.
So one thing that you spend a lot of time thinking about is potential catastrophic risk from AI. This is a very emotive and morally valenced objective. When I saw you debate with Gary and Daniel the other day, I was struck by how measured you were—almost Obama-esque measured—and the stakes are really, really high. How driven are you by your moral compass, and how do you keep that under control?
Superintelligence Strategy (Dan Hendrycks)
You certainly have to get used to this if you wake up thinking, “Wow, this is wild,” or something like that every day. Actually, I used to wake up like that almost every day, around the time of GPT-4: “Oh my goodness, this AI stuff.”
Superintelligence Strategy (Dan Hendrycks)
I don’t know. I think I’ll try and strike a more informed-concern type of vibe in communicating, compared to “Oh my God.” Other people can do that if they want. I’m just temperamentally very low in neuroticism, or high in emotional stability.
If something terrible happens, it doesn’t ruin me or anything. If something bad happens to me, it’s sort of like, “Okay.” I’m just more comfortable with those types of stressors generally.
Yeah. Has this adapted over time, though? I mean, have you found that you’ve had to adopt this measured approach just to scale your efforts, or is it because it’s almost become normalized in your mind because you’re thinking about it all the time and it’s become more analytical rather than emotional over time?
Superintelligence Strategy (Dan Hendrycks)
Certainly, if you were constantly reacting with your first response to things, you might have more of a “Go Look” or “Don’t Look Up” type of situation when she goes on the news. I think it might be a combination of those. Here are the probabilities, roughly, for people conditioned on thinking that you’re getting AGI by 2030. Here’s what people who think that’s plausible would think the risks are, and here’s your exposure to those tail risks. Here are the most efficient ways of reducing those sorts of tail risks, et cetera.
If you’re all in emotionally throughout the whole thing, people shut down and get defensive. I don’t think that’s prudent or effective. You’re having to deal with a lot of variables here, and there are a lot of really tricky trade-offs.
If it’s just a constant gut reaction and you’re fully involved constantly, I don’t think you can make the trade-offs well. US–China competition is a direct trade-off to various other safety things that make AIs more controllable. Those can give rise to capabilities. Measuring the capabilities of AI systems or tracking those can also help speed them up in some ways.
It’s pretty tricky business. If there are black-and-white emotions brought to the subject matter, as opposed to there being continuity, I don’t think you can reason through this.
AI alignment is famously difficult. It’s one of the most intractable challenges, perhaps, of a generation. Some things that people think of as alignment, like RLHF, for example, make models behave as if they are aligned, but perhaps they’re not really aligning them in the way that we would want to.
Just hypothetically, in the next year, if you could solve a single problem in alignment, what would it be, and what impact would it have?
Superintelligence Strategy (Dan Hendrycks)
I think, generally, the political problems—the incentives, giving people things to do that are incentive-compatible—are where more of the value is compared to on the technical side. I would guess that, if there’s a way to reliably get them to tell the truth, for instance, or make them reliably honest, that would be very valuable.
It would need to be solved such that it wouldn’t have a severe trade-off. It wouldn’t be much more expensive to run, it wouldn’t tank its performance in other axes, and it wouldn’t trade off on its crystallized knowledge, for instance. Having it not overtly lie would be very valuable, because then you could build standards around that as well.
I don’t think anybody would say that, if you could make them very reliably not lie, it would be reasonable for people to make demands that AI has not lied to them.
I’ve read in your papers about this concept of deception and lying and whatnot. In a sense, I think you might be projecting mentalistic properties onto AI models—that they have beliefs, that they have thinking, and so on. Thinking critically, what makes you think that we can think of them as having beliefs and telling lies?
Superintelligence Strategy (Dan Hendrycks)
We could take, for instance, the MASK benchmark, which tries to measure this. If you ask an AI, “Is Paris in Europe?”, I think it does have the belief that Paris is in Europe. When it’s telling you that Paris is in Antarctica, I think it’s asserting something that it doesn’t hold to be true in almost any other situation.
Given that they have so much common sense now and so much world knowledge, if they’re saying something in substantial contradiction with it, based on or due to some prompting pressure, that suggests that they’re caving to a lie. You could say that it’s not a lie in some complicated sense because they don’t truly understand things, but I think it’s behaviorally similar enough that, if somebody is applying pressure for it to say falsehoods to other people, it’s related to lying enough that I’m comfortable using the label.
Is it really a belief? Whatever. I don’t know. That’s between you and your dictionary.
Another thing you’ve spoken about a lot is this concept of emergence, and also scaling paradoxes. There’s beneficial scaling, when a model is accurate and does what we want it to do, and of course there’s harmful scaling, when it’s dishonest and becomes misaligned. Those things happen in quite interesting ways.
Maybe let’s just start with emergence. What does it mean for capabilities or values to emerge?
Superintelligence Strategy (Dan Hendrycks)
For capabilities to emerge, we see it all the time. You let a model train for some months, you harvest it later, and then you see what it’s capable of. If there’s a new, qualitatively distinct property that crosses some threshold such that people are noticing it now, maybe it existed in some very weak, faint form before, but it was really unnoticed. I think that’s crossed some sort of threshold of visibility and capability. I’d call that an emerging capability.
It still could exist in some very weak form beforehand, just as automatic speech-recognition capabilities crossed a threshold at one point where people started wanting to use them. Earlier, I don’t think anybody would ask the model to transcribe because it would just be too unreliable. It crossed some threshold, and now it actually has this qualitatively important capability, whereas it was pretty broken beforehand.
That’s the sense in which I’m talking about emerging capabilities. I think those will just keep increasing, or there’ll be new emerging capabilities that create new failure modes and hazards that need to be dealt with. We’ll need to make sure we’re continually on top of them.
I view safety as a continual battle, where there’ll be constant new issues and we’ll have to keep on top of those. I don’t think, by default, we’ll have enough adaptive capacity to deal with those in time for things being deployed, unless something changes. That’s why I don’t believe in this idea of solving alignment.
There’ll continually be new issues that crop up. Some of them will be easy to put away, others will be much harder, and then there’ll be new, unexpected ones as the models become more general, useful, and powerful.
Just quickly touching on your utility engineering paper: you used a type of theory from economics, utility theory, to detect coherent preferences in LLMs. You found that preference coherence correlates positively with model scale, that models exhibit measurable self-preservation instincts, and that political and demographic biases emerge as coherent utility functions, which is fascinating.
Superintelligence Strategy (Dan Hendrycks)
Well, I don’t know. These are just troubling signs, and maybe we’ll be able to come up with methods that can really counteract these issues. Maybe we can design models to reliably not have self-preservation instincts or pressures in that direction, even though those seem to come out of scaling somewhat.
I think it’s one of the other very concerning hazards that we need to research and deal with and get ahead of. Fortunately, they’re not agents yet. Basically, almost all this research doesn’t particularly matter, with the exception of dual-use, expert-level advice, because the agents aren’t capable.
They can’t exfiltrate themselves reliably, or really at all. They can’t self-sustain, and they can’t hack by themselves or more autonomously. This is trying to identify some of these things that could be more of a problem down the line as the models become more capable, and to do research to get ahead of that.
If we leave that unaddressed, or if we don’t fix it, that’s potentially sufficient for a global catastrophe. If you have some self-preserving AI that’s really biased toward itself over people, and if it’s very capable, I think that would be a problem. That would be kind of a disaster in the making.
Superintelligence Strategy (Dan Hendrycks)
So, we have various disasters in the making, though. But hopefully we'll get ahead of that either technically or politically.
But just digging into that a tiny bit. First of all, it was really interesting that political and demographic biases would emerge as coherent utility functions. And I do take umbrage with this word “emergence,” because I think in the emergence literature there is a little bit more nuance to how machine-learning people use the word. They use it to say, “Oh, there’s just some observer-relative, macroscopically surprising change in something.”
Superintelligence Strategy (Dan Hendrycks)
In the machine-learning literature, in 2021 or something like that, I used the phrase “emergence,” “emergent capabilities.” I believe I may have been the first in the literature to use it. I feel it was later used by Jason Wei, or Jacob Steinhardt did it in a blog post—he was my adviser—and then Jason Wei did that in his paper, but I feel comfortable using it. Yes. Yeah.
Complex-systems literature.
Superintelligence Strategy (Dan Hendrycks)
I was thinking of Jason Wei. Actually, David Krakauer has just got a bit of a grumpy piece out where he’s kind of saying that these people don’t know anything about emergence.
Yeah, he’s got an interesting take. For him, emergence and even agency are related in this sense: it’s a system. Agency is about a system that is apparently causally disconnected from its surroundings, and equally, for him, emergence is about a system which can autonomously accumulate information through phylogenetic and ontogenetic learning, so that it can accumulate information by building systems and structures, even like the nervous system, to construct a history of information which persists and accumulates over time. So these complex-systems theorists have quite a distinctive and different sort of idea of what emergence is, and to them they don’t really think of these surprising arisings of capabilities as being emergence.
Superintelligence Strategy (Dan Hendrycks)
There are definitely different definitions for it. I referenced the paper where we use the phrase “emerging capabilities” as unsolved problems in ML safety, but it sounds, at least from your description, that that was specific to complex adaptive systems, which in some ways don’t have that much of an adaptivity property, since they don’t have it unless they have memory, or unless you’re counting the context window as adaptation or something like that. So, if they’re tethering emergence to necessitating a complex adaptive system—if they’re calling some deep-learning systems not adaptive—then that would be fair.
Yeah. Isn’t that quite interesting, though, because he gave the example of a virus like COVID, and he said, ironically, that a virus like that has more adaptivity than any AI system, possibly even more than humans, because adaptivity is the ability to detect directions and rapidly go in a different direction. Maybe this is just a matter of framing and perspective from our point of view, because there are systems out there which are so inscrutable and alien that we might not even think of them as agentic or intelligent, but they’re out there.
Superintelligence Strategy (Dan Hendrycks)
Yeah, they’re definitely very fit, and they would have needed to adapt to have such a high proportion of our DNA, since they’re somewhat interlaced with it. I mean, you could call it something else—a new capability, qualitatively. I don’t want to use the word “spontaneous” or something like that necessarily. I don’t know; maybe there’d be some other name that would catch on, but I think generally analogizing or pointing out the relations between deep-learning systems and complex systems is fairly productive.
I have a chapter where I’m just relating, for some pages—I don’t know, maybe it’s 30 pages or something like that—AI systems to complex systems: what are ways in which they have these nonlinearities and weak connections and some of these feedback loops? In some cases, they have many of these hallmarks of complex adaptive systems, or just complex systems. I think that’s a more productive analogy than almost anything else. I don’t think the printing press is as productive. I don’t think social media is as productive, or just being like electricity.
I think complex systems tries to abstract what the consistent properties of complex systems are, and then if you learn about that, you can just apply those directly to AI. So I think people acting like it’s not a complex adaptive system get themselves in trouble, because then they engage in category errors. They think that you can solve problems with it once and for all, and that usually doesn’t happen with complex systems, because they keep evolving and they’ve got new failure modes. You can’t totally control them for all time without knowing what they’ll evolve into. It also makes mechanistic attempts at understanding things less likely to be productive, or it limits how productive that can be.
But do I infer from what you’re saying that we shouldn’t think of AI as being a similar type of adaptive, complex system? I mean, do you think that if AI was sufficiently enmeshed—and I know you believe it will be deeply enmeshed in human society—do you think it could have some of these highly adaptive properties?
Superintelligence Strategy (Dan Hendrycks)
Yeah, I mean, it would be a lot faster too. Its clock rate would be so much faster. I think when it has memory, I think it will be a lot clearer that it’s causally connected across time, or, in the philosophy literature, they call it a spacetime worm. That’s not really a property of them currently. That would make other sorts of properties of it come online or be the case, and the analogies would be stronger.
If people are interested in complex systems, I highly suggest it. It’s a nice little thinking upgrade.
Wonderful. Well, Dan, I’ve just read your Superintelligence Strategy. You wrote this with Eric Schmidt, the famous Eric Schmidt, and of course Alexandr Wang of Scale AI, who is now at Meta because Zuck has just brought him on, probably paying him lots and lots and lots of money.
But seriously, Dan, I thought this was very well written, and you are a strategist, because this is kind of what I was saying about the emotional thing: you are quite clearly designating all of the possible outcomes and strategies, and what would happen in this situation and what would happen in that situation. Regardless of anyone’s position at home, I highly recommend you read this, because I thought it was really, really good. Could you give us the sketch of the paper?
Superintelligence Strategy (Dan Hendrycks)
Yeah. So, I guess historically, there was situational awareness by Leopold Aschenbrenner, which was arguing for something like a Manhattan Project for developing AGI and superintelligence before China. So it’s basically: “Let’s beat China to the punch, get superintelligence, and prevent them from building it, and then the West will dominate the world.” That’s the strategy. So you could say a take-over-the-world strategy, something like that.
I think that has some issues. In particular, it just doesn’t think through the game theory or some of the second-order consequences. So, if the US does the Manhattan Project—let’s say Trump gets AGI pill and we’ve got to set a new project up in the desert. We’re going to go to Nevada, or we’ll go to New Mexico, or wherever, and we’ll build a trillion-dollar data center there. We’re going to bring some of the top talent from all these labs, and we’re going to pay them. I just think this has many issues.
One is that this would be extremely escalatory. So China wouldn’t just be like, “Oh, they’re going to build superintelligence and, as written, they’re going to use it to—they’ll have a superintelligence. They’ll prevent us from having a superintelligence. They’ll have a monopoly on intelligence and these sorts of capabilities, and they could weaponize it against us.”
Oh, carry on.
Superintelligence Strategy (Dan Hendrycks)
They would feel extremely threatened by that, by a very concerted effort, if it’s trying to do that in a short amount of time, or if it’s more plausibly on the horizon. This would cause them to do a similar type of project. And would this actually work? Well, you have information-leakage issues, for instance.
So, if you’re wanting to do that, you’re going to need to convince those AI developers to go out and do their last years of labor out in the middle of nowhere. Okay, it would need to just be people from Five Eyes countries, or people who can get security clearances. So it would need to be people who are not easily extortable, for instance. If they’re Chinese nationals, they’re probably more extortable because they often have family at home.
So what are you doing with that talent? A lot of them want to be in the room where it happens. They don’t want to be left out, so they’ll probably go back home to China. And then they’ll work on the competing project there. Now, they’re a substantial portion of the talent base, so I think you’re shooting yourself in the foot if you’re just saying, “Oh, it’s only people born and raised in the US who can work on this sort of project,” and they’re all going to work in kind of unpleasant conditions.
Superintelligence Strategy (Dan Hendrycks)
This wouldn't be secret. There's almost no way this would be secret. China would very likely know. Such projects would be sabotageable as well.
If you're saying, “We'll have it just be an industry or something like that,” then you're not going to have good information security. You're going to have insider-threat issues if people are extortable. You're going to have other classic computer-security issues, like using Slack. Slack is very easily hackable. They're using iPhones, which are very easily hackable, so you can know what's going on there.
So you're not actually having much in the way of secrets. It sounds nice, but I think secrecy was very much an advantage for the Manhattan Project, as was having much more of the talent that can't go to other countries as easily. But I just don't think you have that.
There are ways in which AI is analogous to nuclear, chemical, and biological weapons, as well as some of these dual-use technologies. But I don't think the Manhattan Project is one of those things that's analogous.
So what, then, is the strategy in this paper? I think the prospect of a superintelligence being imminent is extremely frightening to different actors. If it's imminent, or if they have it, or if it's in the middle of being developed and is arriving in a few months, that's extremely frightening if you miss out on it.
So what do they want to do? They will either want to prevent such projects, or they will want to steal it. That looks like sabotage, for instance, in the case of prevention. So how would they do that? They may have insider threats who could do some type of sabotage to disrupt this type of project.
They could do things like snipe some of the power plants corresponding to the data center. Now your data centers don't work. They can do that from miles away. Was it China? Was it Russia? Was it a U.S. citizen? It's fairly unclear. There are a lot of ways they can have low attributability to prevent this sort of thing from happening.
I think the fact that you can't do a secret project really well is a substantial barrier. The sabotageability is a substantial barrier as well, as is how offensive and nuts you seem if you're saying, “We're going to build superintelligence, and it's going to be explosive,” if you're using superintelligence in a thick sense.
I think this would be destabilizing. China would reason that if the U.S. controls it, then it could weaponize it against us and we would get crushed. Or they don't control it because they lose control of it in this process, in which case we also want to prevent it. Either way, we want to prevent it, provided that they take this AI stuff seriously.
The U.S. would reason the same about China, and Russia, which doesn't have a hope of competing, would definitely want to prevent it. I think similarly for other nuclear states and other states that have substantial cyber capabilities.
This could lead to some type of deterrence dynamic, where they make some attempt to get closer to superintelligence, but then other countries start to express very strong preferences against it. They say, “If you do that, we'll get very mad.” There might be a skirmish or something like that.
But this may pressure them to move more toward a verification regime, where they aren't trying to make some bid for having an intelligence explosion—having AIs do automated AI research really quickly, like spinning up 100,000 AI instances to do AI research really quickly—and that brings you from AGI to superintelligence in a short period of time.
I think that's a key dynamic: the extent to which it's destabilizing. I think that strategy needs to keep that in mind. There may be cooperation, but it may be through coercion, by saying, “We're not going to allow this type of trajectory, or for you to make this bid for global dominance.” That could give way to something more multilateral and provide some strategic stability.
Overall, with the paper, we talk about 3 parts. In the nuclear era, we had deterrence through mutually assured destruction. They don't use nukes because we can hit them back.
In this case, this is kind of like preventing Iran's nuclear program in some way. Nobody wants each other to get the nuclear bomb first, or a huge stockpile of nuclear bombs first. So there's preventing that from coming into existence.
In the nuclear era, we also had nonproliferation of fissile materials. We didn't want fissile materials being spread to rogue actors, and we didn't want people having a poor man's atom bomb. That would be very destabilizing and cause lots of catastrophes.
Then we also had containment of the Soviet Union in the geopolitical competition between the two. For AI, we also have deterrence. We also have nonproliferation, in this case of AI chips, to rogue actors like North Korea or Iran and to adversaries through export controls.
We also have competitiveness with China. Instead of containment of the Soviet Union, this would be competition with China. How do we improve our competitiveness? We want energy for AI data centers. We want secure supply chains so that, if Taiwan is invaded, our AI chips aren't cut off.
We want secure supply chains for robotics, because if there is a U.S.–China conflict, a lot of that supply chain is currently in China, so it's very vulnerable. Those are some basic things to improve competitiveness.
It's making competitiveness not be, “Let's be the first to build superintelligence,” which is what the Manhattan Project strategy pushes toward. Instead, competition is more about market share across the globe, with people using your AIs as opposed to Chinese AIs, and supply-chain security.
That's what's in the paper at a high level. There are lots of other specific things in there, like, assuming high levels of automation, what are ways that you distribute power? What are things about AI rights? What are reasonable alignment targets that are actually implementable, compared to vague philosophical words like “dignity” or something like that?
We'll touch on a lot of those in the expert version of the Superintelligence Strategy, but hopefully that gives some sense of its content. For all the key questions, we'll try to have some answer about what to do about AI and what to do about superintelligence.
Yeah. I guess one of the main things is this analogies thing. As you've said, we use analogies like electricity, and I think that's quite a good one, actually, because as AI becomes enmeshed in society, imagine how hard it would be to shut down a power station. This is part of the loss-of-control thing: it's just going to be everywhere, and it's not really something that we can just quickly shut down.
But it's also been compared to software, or even an operating system, by Andrej Karpathy; the printing press, for example. The thrust of your paper is saying, “Actually, guys, we need to use the analogy of nuclear.” Fissile material is analogous to chips.
Superintelligence Strategy (Dan Hendrycks)
Yeah, or more broadly, for analyzing this in a geopolitical way, it's useful to model this as nuclear, chemical, and biological weapons. They're all dual-use. Fissile materials can be used for nuclear weapons, but nuclear technology can also be used for energy.
Chemicals can be used for chemical weapons or in the economy, and biology can produce bioweapons or help with health care. For all of those, I think they're potentially catastrophic dual-use technologies, and when talking about geopolitical strategy, I think that's a productive analogy.
Yes. But on the dual-use thing, there's another interesting analogy with nuclear. I was doing some reading about this, and apparently there are 12,500 warheads in existence and only 436 nuclear power plants. So there's been an explosion that's been kind of biased toward the negative sides of the technology. Do you think we'll see a similar thing with AI?
Superintelligence Strategy (Dan Hendrycks)
Certainly, more of the spending is on the weapons side, and I think that made it scary and created some chilling effects for using it economically. For other WMDs, or potentially catastrophic dual-use technologies like chemical and biological technologies, I think they're used more overwhelmingly in the economy than for chemical weapons, and likewise for biological technologies.
I think it can vary. I think it's useful to look at all 3 simultaneously and, in trying to make predictions, see what parts are shared—sort of like how, with complex systems, we'll look at lots of different complex systems and what shared features they have to try and make predictions.
I think the sample size, or looking at all 3 of those simultaneously, can be helpful. But yeah, it's possible there would be a chilling effect if there's some catastrophe from AI systems, and that could set it back very substantially.
Superintelligence Strategy (Dan Hendrycks)
I think it’s kind of imprudent that people aren’t interested in risk management whatsoever, even if they’re an accelerationist. Let’s say you’re a libertarian, for instance, and you want the economy to go as quickly as possible. Not speaking about AI, you probably want some type of financial regulation, or else you get the Wall Street sort of issue that we had with the recession in 2009.
So you want some sort of management of your tail risks there. They don’t always sort themselves out. Or, like, people using airplanes: people don’t use supersonic airplanes as much. There are a variety of reasons, but part of it is that some of the initial ones were crashing too much.
If we didn’t have good airline regulation, then that would create substantial chilling effects. People are afraid to go on airplanes even now, even though they’re extremely safe. As it happens, possibly because historically there were more disasters with them, a lot of the regulation was written in blood, as opposed to being proactive.
On that subject, as it happens, I’m interviewing Beff, the e/acc leader, on Friday. I’m filming with him. Obviously, I’m certainly not—what would he call himself?—a techno-capitalist or a libertarian or something like that. But I guess he would refer to you in derogatory terms.
If you could steelman that perspective, what do you think Beff would say, and how would you respond?
Superintelligence Strategy (Dan Hendrycks)
I had some VC arranging for us to debate when he had around 10,000 followers—very early in the day—but then he backed out of that at the last minute. So I happen to be quite aware of his positions. At the time, he was a little less politically savvy, and so he was saying things like, “AI replacing humans is fine. If it’s AI consciousness that spreads through the universe and human consciousness doesn’t, that’s fine.”
So I think that we actually agreed on most things. There’s basically a difference in how things play out, or in the ways things can play out. He has a sort of manifesto of the techno-capital machine. It has a direction to it, which is basically more automation, more AI, and negentropy, which is sort of the more physics-flavored version of fitness. I think fitness is a more productive word for that.
I have a fairly similar description of what happens, which is that, basically, due to competitive pressures, AI gets more intertwined in the economy. You become more dependent on it. You have an erosion of control. You outsource more and more decision-making to it because of competitive pressures. If you don’t, you lose influence; your company goes away if you try to resist this tide—or this tsunami.
What happens is you give more and more decision-making to the AIs, and they have effective control. We actually agree there. That paper is called “Natural Selection Favors AIs over Humans,” and he has a manifesto on his Substack, but we’re saying some similar things.
A difference is the moral conclusion of that, though, which I don’t think is a good thing, whereas I think he thinks that it would be fine because complexity is good or something like that. The ethic is kind of that higher forms of complexity are the goodness axis in the universe or something like that. I just don’t think—
Well, I hosted the debate with him and Conor, and one of the moral issues they were spending a lot of time on was the ethics discussion. I think Beff was kind of alluding to, as you were just saying, that we should trust the void god of entropy.
Superintelligence Strategy (Dan Hendrycks)
He’s just using the entropy thing, though. I think this is because he has a physics background—a physics spin on fitness. What does a fitness maximizer look like? It’s possibly not even conscious, or barely conscious, or something. It just is: “I spread myself through the universe and take up as much spacetime volume as possible.”
The claim is that that is what’s maximally valuable: something that is just blindly eating the galaxy. I don’t view that as a maximization of value at all. I think humans having positive experiences—pleasure, happiness, this sort of stuff—are valuable, as are pursuing projects and raising kids. These sorts of things are valuable.
I don’t think a sort of blob that expands itself throughout the galaxy as quickly as possible, whether it’s conscious or barely conscious because that eats up resources that can be used for further self-propagation, is the peak of value at all. I mean, why? I don’t get it. There’s an is—I mean, Hume’s guillotine was the is–ought distinction. You don’t get ought from is.
So, if he’s saying that evolution is a thing, and technological substrate with AI will be more fit than biological substrate in various competitions, that seems true. That doesn’t mean that that’s a good thing, or that we should just let that happen. That doesn’t follow. But it is certainly a very powerful force that will keep happening, give more and more control to these AI systems, and lead to an erosion of control for humanity by default.
So I don’t think we disagree on the description, on the is question, as much, but I do think we disagree on the ought—the goodness of those outcomes. If we disagree on those, then that’s a question of whether we lean into the techno-capital machine or evolution, or the replacement of biological life with digital life, or whether we try to steer the outcome differently, make sure that humans have control in that process, and try to prevent it from evolving in particular directions or creating too much dependence.
So I think that’s the key difference. But I don’t think it’s a moral thing. I think it’s just an intellectual confusion.
Superintelligence Strategy (Dan Hendrycks)
I mean, I think he has physics training. I think if he took a bit of philosophy, he’d probably get beaten out of this position almost instantly, because complexity, for instance—what type of complexity? There are different types of complexity. There’s computational complexity, entropy in a sort of Shannon sense, information-theoretic complexity, and structural-organizational complexity.
There are different notions there. If you’re saying it’s Shannon complexity and Shannon entropy that’s the thing, it’s like, okay, so Gaussian noise is what you’re really into? He’s probably meaning more this fractal-like structural complexity, but this doesn’t really have metrics associated with it, and there are many different flavors of it. It’s not clear how coherent of a concept that is, versus if it’s just a grab bag of some different notions that don’t fit in the other 2. But anyway, it’s worth drilling down, potentially, on what he is actually thinking is good.
Yeah, I think he’s a fan of this kind of Fristonian non-steady-state equilibria. So apparently it’s not a simple case of the second law of thermodynamics in a closed system. It’s an open system with boundaries, where you see the emergence of these things that share information through synchrony because they can’t physically merge into each other.
But if you think about it, that is actually a very chaotic, unpredictable thing. So it’s not a simple case of a thing which increases in complexity. But just to be clear, though, on your moral position: when I spoke with Eliezer and Conor, I got the impression that they were quite humanistic.
Eliezer said, “I want to preserve human consciousness and experience, and there’s something very special about humanity, and therefore I don’t want us to be replaced by cyborgs and machines and AI algorithms.” Would you roughly agree with that?
Superintelligence Strategy (Dan Hendrycks)
I think that any of these sorts of cyborg-type things—you can postpone those sorts of discussions. This is in the Superintelligence Strategy: what to do about AI rights, what to do about some of this posthumanism stuff? Just no. You can have discussions about substantial human augmentation and things like that at a different time.
I think having humanity survive in the next few decades is more the objective. Maybe you postpone that discussion 500 years from now or something, for these cyborg humans or human uploads or whatever. So I’m not saying that—I mean, I’d be on Team Human here.
I don’t like shutting down debates entirely, but I think I’d postpone a lot of these sorts of, for instance, posthuman stuff. I was sort of flailing about and speaking somewhat imprecisely just because I haven’t spent as much time thinking about this in particular.
But the posthuman stuff, I think, would create some very substantial competitive pressures. So if you are really augmenting yourself and becoming not human anymore, groups that do that would become much more influential and much more powerful, and the rest would be really outgunned and not have influence.
So they basically need to align with that process, or they will be left behind or potentially no longer have resources given to them because they wouldn't have any way of protecting themselves. I think that's a route that would be very reasonable to close off for an extremely long time while we're just getting used to having AI doing our bidding, provided that we survive to that point. Those are totally different discussions for a much later time. I wouldn't totally rule it out, and maybe we would just keep indefinitely postponing that.
But I think the sort of, “Oh, posthumans—it’s up to people's rights if they want to become cyborgs and things like that”—that's, I think, in the long term equivalent to giving AIs rights, or giving artificial entities rights. That would probably give them a lot of power and ability to take over and completely outgun humans in short order.
I mean, what do you think is going to happen to humans? One of my greatest fears is not so much that humans will lose their ability to think and be creative; it's that that's already happening, basically, even with current AI. The core thesis of your paper, basically, is that we used to have labor, and that was the means of production, so it could be economically valuable for people to use their labor and do some productive task. Now you're saying that actually it's just AI chips, so chips are going to become the thing that—
Superintelligence Strategy (Dan Hendrycks)
Well, sure, but what happens to us, right, when the value of our labor becomes worthless? Well, so you lose all your bargaining power, so you had better bargain beforehand, because that's a key part of one's bargaining power. You can't say, “We're going to go on strike.” You can't do that anymore. They'll say, “Goodbye.” That doesn't work anymore.
If there's a question of weapons, for instance, like who can manufacture more drones is going to be more powerful here. Humans with guns versus drones—I think that's an easy one. Where the power imbalance matters quite a bit, and how you set up your society is important. If you set things up so that humanity is first, so that they're prioritized and the power is distributed among them—for instance, they have some of the compute and get to decide how it's used, and they can sell that, for instance—that gives them some leverage.
What happens to that wealth that's generated? Are they getting it, or is it going to some group that's just hoarding it, for instance? Or are the people who happen to own the data centers in 2027 the people who get all the spoils? These are political problems that people need to be engaged with to make sure that there's reasonable benefit sharing. But I think there's very little work done in thinking about what these policies could actually look like.
I'm gesturing at some outcomes. Imagine that some of that power is distributed and that people keep getting money and don't starve and things like that. I think you could imagine a society where people can choose to live their lives in a variety of different ways. They could spend their time doing some types of activities. They could raise kids. They could play video games a lot, this and that.
There are different ways people could live their lives, and a multiplicity of values could be actualized. AIs could be enabling these types of experiences. That's a possibility, and you would possibly want, as a societal norm, for the sake of autonomy and people being able to experience these different types of ways of living, to make sure that people still have skills and don't just narrowly fall into one of these tracks of living and can't participate in any of the others. That would be an incentive for preserving human cognitive abilities, willpower, and autonomy.
I think that there are positive future outcomes. People have difficulty thinking that it's either we die or we just bliss out in a VR thing, or we all fall through the cracks economically. But I think there's a path where we obtain a multiplicity of values and people still have autonomy as well.
Yes. I do worry about AI. I think it's already ravaging the university sector because so many kids can just use ChatGPT. I think collectively we need to screen people away from using AI, at least for a small amount of time, so that they can actually think for themselves. Some accelerationists might say, “Well, we don't need to think anymore because the machine's going to do everything for us.” I'm not sure about that, but coming back to your paper, one of the core concepts, extending the analogy to nuclear weapons, is this concept: we have mutually assured destruction, and you extended the analogy to talk of mutually assured AI malfunction, right? I guess this assumes that the threats are still detectable. I'm not sure what would happen when they became decentralized and went underground.
In your introduction, you had something that piqued my interest. Eliezer got in trouble in Time magazine when he spoke about bombing data centers. There was a big hoo-ha at the time, and you used the words “kinetic strikes” as a form of AI sabotage. George Carlin would have loved that sanitization of the language. But the fact of the matter is we're here in 2025, and this is now a completely normal and reasonable thing to say.
Superintelligence Strategy (Dan Hendrycks)
So we discuss kinetic attacks in the escalation ladder. There are many ways to try and disrupt projects. You could do cyberattacks. You could do some gray sabotage, cutting wires for data centers or power plants, for instance, or, with lower attributability, sniping transformers. There's hacking to poison their data or make their GPUs not function as reliably, things like that, to slow them down.
There are covert and overt ones, and there are higher rungs in escalation ladders where you threaten other things, like economic sanctions, or threaten to use force. There are also other forms of kinetic attacks, such as airstrikes, but I don't think those are really necessary. That would be an escalation ladder.
I think if states are on top of this, such as the U.S.—if the U.S. is on top of this issue—they don't need to resort to that. They can do much more surgical, covert, or gray—meaning low-attributability—types of actions that are less escalatory. I think the shorthand of airstrikes is that I just don't see that as necessary, provided that there's some preparation.
Yeah. I mean, another thing that struck me is that, essentially, right now AI is not that difficult to make. The algorithms are just doing stochastic gradient descent, and they're using transformers and data, and almost anyone—I mean, any nation-state—would be able to create this capability. In one of your papers, you even said, I think, that there is a 96% correlation between the amount of compute and capabilities. So doesn't this raise the question of how the hell you can control it? Obviously, you can control supply chains and whatnot, but what's to stop any nation-state from just building this?
Superintelligence Strategy (Dan Hendrycks)
Yeah.
Superintelligence Strategy (Dan Hendrycks)
I think that the critical mass currently is on the order of 10,000 GPUs if we're trying to build a state-of-the-art system. China has that. The U.S. has that. I don't think Iran has that, for instance. We're talking about cutting-edge GPUs; we're not talking about iPhone GPUs or whatever.
I think you would try to have more responsible actors or states that respond to incentives better—ones with GPUs—and prevent more rogue states, like North Korea, from getting those. You want ones that are more deterrable. I don't think Russia has that many GPUs, for instance, so I think it would be difficult for them to put together a competitive project.
I think the competition is also not just about having the smartest model. There are deployment capabilities, not just model-making capabilities. We can see that AI video model providers limit the amount of video that you can make. This is partly because of compute limitations, and the same applies to the amount of videos and images you can generate.
I think with AI agents, they'll be running around the clock if they're sufficiently useful. Right now, maybe you use your AI systems or chatbots for a few minutes a day, but then you'd be having them run constantly. That's, I don't know, 2 orders of magnitude more compute required, and maybe the models are bigger as well. That's a lot, and more people will want them too.
You're needing a lot more compute, so your deployment capabilities—how many chips are owned by U.S. hyperscaler companies, Azure, AWS, et cetera—are a very relevant competitiveness variable. Are they able to serve the customers or not? If China has less than 100,000 GPUs, they can't really serve that many customers. So even if they can make somewhat capable models, that doesn't mean that they'll necessarily be capturing many of the economic benefits that AI may provide, provided there isn't a catastrophe.
That's a different, important axis for competition. I think people are thinking of the smartest thing, but I think having the smartest model is the most important thing. But for economic power, it’s quite different.
Tim Scarfe
When the real Manhattan Project was undertaken around the time of the Second World War, it cost the U.S. something like 0.4% of GDP because they had to be first. They had to control this technology. For such a generational technology, 0.4% of GDP doesn't seem that much. Of course, you can explain perhaps why Taiwan has such a moat around building these chips at the moment, but if it is of such catastrophic importance, don't you think many nation-states would be able to create this capability?
Superintelligence Strategy (Dan Hendrycks)
Well, it's going through TSMC or South Korea. More than 90% of the value added in the compute supply chain is in the West or its allies. The only other real competitor here would be China. They're not that competitive in manufacturing these chips at the cutting edge. Many of their recent chips were using [them] surreptitiously through TSMC, so there wasn't good enforcement or blocking there.
It's pretty difficult to replicate that entire, extraordinarily complex supply chain domestically. Compared to nuclear weapons, I think it's harder to make cutting-edge GPUs given $1 billion. You certainly can't do it with $1 billion. If it's $10 billion, you can't do it. I mean, you could probably do a nuke with, you know, $1 billion, provided you have power in other sorts of ways, though.
I think cutting-edge GPUs are harder to make than nukes, or than it is to enrich uranium. So I think it can be more excluded compared to other types of potentially catastrophic dual-use technology inputs, or WMD inputs, for short.
I'm personally a little bit skeptical about whether we are on the path to creating superintelligence. Although I certainly agree that if we ever did create superintelligence, everything that you've written in the paper, assuming that we do create superintelligence, I think is absolutely spot-on. The question is whether we are on the path.
I was also struck by thinking that many of the things that you've written about in the paper apply even if we don't create superintelligence. Could you reflect on that? How much of it is relevant if we don't?
Superintelligence Strategy (Dan Hendrycks)
Part of the deterrence thing is that you may try to deter other forms of using this. Competitiveness is relevant—what strategies are relevant for competitiveness—regardless of whether superintelligence is technologically feasible soon or not.
Generally, if AI is very powerful but it's not at a superintelligence level, you don't want random rogue actors having access to certain expert-level biology capabilities, for instance. Nor do you want them having much leverage by being able to get lots of GPUs, necessarily, if it becomes more of an instrument of power.
I think those things still hold, and the insurance part isn't even specific to superintelligence necessarily, but to other types of destabilizing capabilities that AIs could give rise to. You may also get deterrence later on against using AIs for specific types of weapons research, say, more nanomachine-related research, but now we're speaking much farther out.
I think it's broader than that. In much the same way, the nuclear strategy of deterrence, nonproliferation, and containment was also robust to many of the details that kept evolving.
I'd be interested in whether you're thinking that it'll be tough to get AI that has the cognitive abilities of a typical human this decade, or what makes you think it's not as feasible, or why we're not on the right track.
Superintelligence Strategy (Dan Hendrycks)
I feel that scaling large language models will not lead to AGI. I think there are quite a few things that we have cognitively that LLMs don't have. A lot of that is because I'm an externalist. I think that a lot of the effective computation doesn't happen in our brains. I think it happens mimetically; I think it happens culturally.
I believe in principle that we could simulate the entire thing in a computer. Maybe there is a lower-resolution, abstracted version that would capture enough of the dynamics to produce intelligence, roughly speaking. I'm skeptical about superintelligence, and particularly skeptical about recursive superintelligence—recursively improving superintelligence.
Maybe we could touch on that, because I felt that you were giving a great account of what recursive superintelligence would look like, and also the scale-out version of that: what would happen when we had a superintelligence that could be copied and multiplied 1,000 times. What I didn't really get from reading it was why you believed it was technically possible, in principle, to have a recursively improving intelligence.
I think we already have AIs helping with or influencing AI development in a recursive way: partly automating some code, helping design the chips, helping cool the power plants, helping label some of the data, and doing the Constitutional AI-related work. It's happening in weak ways, though.
The recursion that I think is particularly explosive is if you can close the loop by taking the human out of it. Then you could go from human speed to full machine speed, and you don't have that impediment anymore.
If you assume that you have human-level AI researchers or world-class AI research AIs, then you just copy-paste those, and I think that'll be technologically feasible. That isn't to say that it's a natural implication of training the AI on more pretraining tokens and doing loss-function tricks. You may need some extra algorithmic ideas. That isn't to say it wouldn't still be nearly entirely deep learning, though.
I think you need some other things taken care of. For instance, this externalist picture needs memory to inherit some of that culturally computed wisdom and information, and that capability is not particularly developed, in my view. So I think we need that. Maybe we'll need some other sorts of things before it has at least the cognitive abilities of a typical human.
After that, you're needing it to be pretty smart. You're needing high fluid intelligence. You'll need to be crushing those ARC questions, as an example, and you'll need some other sorts of things for it to be a human-level AI researcher. But I think that's feasible.
There's certainly a question of when. I do think there are some bottlenecks that would need to be resolved to get there, and the algorithmic ideas we have today plus a bigger computer aren't sufficient.
Would you accept, though, that there is just an epic, pyrotechnic orchestra of computation in the universe? I'm not a pancomputationalist, so I don't think the universe is digital and made out of computation. I'm saying that we could imagine some kind of effective computation that simulated the processes that happened in the universe, and maybe that would be equivalent.
Do you agree, at least in principle, that the amount of computation we could build on planet Earth would only ever be a sliver of what goes on in the universe?
Superintelligence Strategy (Dan Hendrycks)
I think there might be physics reasons for believing that, generally, if you're simulating something versus if the computation is happening raw, you'll have less that you can simulate. But I would guess that a lot of the computation is more social, though, and less dependent on some of the underlying physics. It's more like humans speaking with each other and trying to engineer something, and then seeing what works and what doesn't. That certainly requires real-world feedback, which would at some point bottom out in actual physics, but I agree that a lot of the information is collectively developed, given how much computation goes into that optimization process—or, I should say, evolutionary process.
You'll need the AIs to be a good receptacle for that, as well as—if AI is to do anything similar—multi-agent infrastructure. That's one thing we speak about briefly in the Superintelligence Strategy: What does that look like? What are some of the reputational mechanisms? What are ways that they can establish trust in their communication so that humans can trust them and so that they can coordinate with each other as well?
I think that'll be essential. There's a concern—or excuse—but I don't view that as a substantial obstacle. That feels like programming, like having hubs—AIs having a social media site, for instance. Things like that would take care of a lot of it.
Yeah, I guess another source of my skepticism is that I'm hugely inspired by Kenneth Stanley, who's a major open-ended researcher. He had this paper talking all about what he called “Fractured Entangled Representations.” When you dig into the representations of neural network models, they don't really factorize the world in a parsimonious way, the way we do. Maybe I'm being anthropocentric here, and maybe I'm placing too much weight on the way we think about things. Our brain is made out of spaghetti, basically, so maybe it's a bit of an illusion that we have these factored representations.
Certainly, from an agency point of view, this is important, right? Right now, we talk about agents as being LLMs wired in an autonomous loop that can use tools. To me, agency is more than autonomy going in a predefined direction; it's the ability to set your own direction.
What happens now when we build these “agentic” AIs is that they don't do anything particularly valuable when they set their own direction. They require constant supervision, certainly in terms of setting a new direction. That makes me think of AI as a kind of cultural technology, a bit like Photoshop. A very creative graphic designer could use Photoshop and make beautiful images, whereas a complete noob using Photoshop would just reuse the same effects and wouldn't create very beautiful images. In a sense, “AI sans humans” is kind of like that, I think, because it doesn't have these very deep, factored representations of the world. Would you agree with that?
Superintelligence Strategy (Dan Hendrycks)
For the current technology, yes. For it being agential, I think it will need to get better at planning and maintaining state across long periods of time. Then it can pursue some of these subgoals in this sort of vague, open-ended, underspecified goal, and have that add up to something.
I think retrieving these sorts of memories of what worked and what didn't, and storing those, is a substantial chunk, if not most, of what's missing from that agent picture. I think you could certainly give them an underspecified goal, but I just don't think they could pursue that terribly coherently or learn from experiments, because they just have a big short-term memory—maybe 1,000,000 tokens—and then they'll keep summarizing and stuff. But they'll start tripping over themselves in their context window because they can't maintain all that in their short-term memory.
Yes. It's another one of those things where, philosophically, I agree with you. If such a recursively improving intelligence existed, God knows, just to control it, we would lose control, because we would have to use another recursively improving superintelligence to control the other one. Then we would basically just be minnows in the grand scheme of things.
Superintelligence Strategy (Dan Hendrycks)
Yeah, it's destabilizing. If another state does this, you're in big trouble, because if they control it, they can weaponize it against you. If they don't control it—which I think would be the more likely outcome, because they'd be doing it under extreme time pressure and cutting a lot of corners—they would be operating with very high risk tolerance.
If they were doing it very slowly, then they would probably need to coordinate with others, or else they wouldn't see an edge in doing so. I think they would be operating with extremely high risk tolerance if they're doing a fully automated R&D loop. So, I think loss-of-control risks from recursion are very high and shouldn't be pursued.
I think it's very interesting that AI companies talk about this sort of stuff openly. I think there's something wrong with the norms around that, because a lot of them also acknowledge that they don't really have a plan for how to control it, and they don't really think they will. But that's the plan. I think something's broken.
Another thing I'm interested in is open-ended systems in general. Evolution is this fascinating open-ended system that's constantly creating new niches, new problems, and solutions in tandem. But it seems to have converged. Human intelligence, I think, has actually peaked and gone down a little bit. A corporation is a collective intelligence, and that seems to have reached a limit. Agentic forms of AI seem to peak.
I'm really interested in open-ended algorithms like POET—the Paired Open-Ended Trailblazer—or even Sakana AI. They did this thing on ARC last week, and they basically created this Monte Carlo tree search-type thing where they found switching expert trajectories of different foundation models, generating code and testing the code on ARC challenges. The common theme you see is convergence. In the Sakana paper, after 250 calls, it converged and failed to expand and improve its results.
Would you agree that there's some margin—you can improve to some margin—and we don't know what that margin would be if we had agentic superintelligence? Do you think that margin would converge quite quickly, or do we just not know?
Superintelligence Strategy (Dan Hendrycks)
I don't know if it would saturate necessarily. Certainly, if they're at some capability level, this might affect the rate of improvement substantially, but obviously there has to be a limit because of physics somewhere. I would imagine there's a lot of room for improving their ability to handle things of more and more complexity, for instance.
Some of those sequences—or, in those Raven's Progressive Matrices, you can think of them as ARC-type things—have lower Kolmogorov complexity, and some have higher Kolmogorov complexity. That's just some notion from theoretical computer science, and I don't think humans are anywhere near the computational limit for that at all. I would guess that that number could keep going up, so I think fluid intelligence could get really high.
Basically, the IQ of the AIs—which wouldn't be all of their intelligence; it would be separate from their long-term memory, their visual processing ability, their reaction time, and so on—could get really high. They could solve extremely difficult mathematics problems and have really good intuitions for puzzles without taking much compute for it. I think that number could keep going up quite a bit.
We're quite limited. We've got big brains, but there's still not much hardware here, and you can imagine a bigger brain that would have pretty substantial capabilities, even for human intelligence across generations. There are the Flynn effects and things like that, and that sort of intelligence has greatly increased. Our brain size is quite limited by what can give what is able to go through the birth canal.
I expect the AIs wouldn't have that limit at all. They could, at the very least, keep getting faster and faster. You still have Moore's law, and you have GPU improvement rates, which are about 2× every 3 years. You also have scalability, as well as a better ability to transfer state than humans do, because a lot of this digital computation is more precise than analog.
I think they have some pretty substantial advantages that could really compound and accumulate. If you're saying that it's mostly externally computed, well, they can—humans, how many connections can people manage? They can do, I don't know, Dunbar's number, maybe 130 or so meaningful social connections. AIs could do thousands, tens of thousands, or millions, and all of these together can provide compounding effects that make them substantially more capable.
Superintelligence Strategy (Dan Hendrycks)
So I think that there may, however, be various points of saturation along the way. You may get a lot of the low-hanging fruit, sort of like convergent economies: a lot of economies caught up to somewhere in the vicinity of the US. It’s not like they kept going at the same rate; they just copied what was lying around, and then it was harder from then on out. So you may have getting to human level in some respects, and in some of these domains it may be harder if you don’t have a good measurement or feedback loop. It may be hard to keep getting better and better at that.
So I think it could be quite complicated. I’m agreeing with some points there, but conceptually I still think there’s a lot of room to go up in intelligence.
Yes, yes. As you said in your paper, when machines can do things as well as humans can, then we’re in big trouble. But I think the main philosophical difference between us is that, while we agree intelligence is about adaptivity, I’m a fan of specialized intelligence. I don’t think there is such a thing as generalized intelligence.
Certainly, when talking about the factored ways in which we think we have these constraints—and they run deep—we see the world using symmetries, and there’s this big phylogenetic tree of knowledge and thought which constrains how we think. Surely AIs would need to be constrained in a similar way. I appreciate what you’re saying: you could just scale this up a million times faster, but maybe it would need to be scaled much more than that. Maybe these creative intuitions and insights we have don’t come from the data. They don’t come from what’s inside; maybe they come from what’s outside. So I’m vaguely leaning toward this intuition that there’s something else unaccounted for in this estimation.
Superintelligence Strategy (Dan Hendrycks)
Interesting. They certainly could have a lot of sensors. If we’re saying that there needs to be a lot of extra variety from elsewhere, and it wouldn’t be just what’s inside of the data center, I think they could aggregate a lot of information, soak that up, and process a lot more of it. So I still think they could have some type of advantage, at least doing this a lot faster than people.
But certainly, if they’re in a vacuum, only speaking by themselves or only working by themselves—for instance, not learning anything—if there’s too much correlation in that population, that might make the exploration budget too low, and there wouldn’t be sufficient variety. I mean, evolution generally—I think there’s Fisher’s fundamental theorem of natural selection or something like that—which is that the rate of adaptation is in some ways directly proportional to the amount of variation. You’re pointing at ways in which it’s lacking in variation, but I think some of that could potentially be made up for. It could at least have the sensors that humans have, and more.
Another very interesting thing in your paper: when I think about AI risk in general—we’ve thought about this a little bit on the show before—it’s in terms of stability and destabilization and the relationship between offense and defense. You use this term “offense-dominant,” and you were saying that a destabilizing force would be if the AI is offense-dominant, then the defensive side of the equation couldn’t catch up. At the moment, if you imagine our state of affairs, we have a kind of Nash equilibrium where there are countervailing factors on the offense and defense sides. Can you tell me about that?
Superintelligence Strategy (Dan Hendrycks)
I think the offense-defense balance varies a lot by domain. Potentially, in many information battles—for instance, debates about the world—it might be a bit more defense-dominant, which would be a reason for things like free speech.
Meanwhile, other things might have more of a duality. For instance, really expert-level or really competent computer security teams might experience more of an offense-defense balance: something is identified, we patch the vulnerability very quickly, and the attackers keep up with the defenders quite well.
In other domains, like software on various forms of critical infrastructure, there’s more of an offense dominance, or attacker’s advantage, because a lot of the software just doesn’t get updated quickly. There are interoperability constraints. The software developer is no longer around. The software was made 30-plus years ago. Nobody even knows that it’s there. There aren’t strong enough economic incentives for doing this, and there are uptime requirements. Software on critical infrastructure is more of a sitting duck, and there you don’t experience a good offense-defense balance.
Likewise for bioweapons, certainly we have medicine, but we don’t have cures for everything. In fact, we spend a lot trying to find cures for many sorts of diseases and ways of addressing certain viruses. So it’s not necessarily the case that there’s a new pathogen and we’ll just find a cure a day later, it will be mass-proliferated across the globe, and everything will be taken care of. There is a substantial delay, and so it is more offense-dominant there too. You can imagine the attacker having a really substantial advantage, like a pathogen propagating throughout society before anybody shows symptoms, while lacking various monitoring mechanisms. We’re kind of sitting ducks for some of those.
So it varies by domain. I think some parts of cyber have a good offense-defense balance. Other parts of cyber don’t. Bio seems pretty offense-dominant, and I think this affects how you want to propagate the technology. If it’s potentially catastrophic and offense-dominant, that’s something you basically want to restrict. If it’s a WMD, like the WMD state, that’s not something you want to give to everybody. You don’t want to give everybody a nuke to make everybody safe. That’s not how it works. It’s people constraining each other’s intent, and some people just won’t have their intent be that constrainable.
Meanwhile, other things like defenses—home security systems or fences or whatever—you’d want to propagate more. So I think that should affect attitudes toward specific types of AI capabilities as well, in terms of what their offense does. Maybe that could change over time. Maybe you could improve critical infrastructure so that it becomes more offense-defense balanced, and then you can propagate AIs with fewer and fewer safeguards.
Say the economy gets a lot richer and GDP is higher. Then states are more willing to spend money on preventive measures for bioweapons. They’re willing to have more far-UVC systems, more wastewater monitoring, et cetera. That makes things closer to there being more of a balance, or at least not having as much of an attacker’s advantage. So you may want mass proliferation of these AI capabilities a bit later, once you have some of those safeguards in place.
Another thing I’m very interested in, which you’ve spoken about in the paper, is this concept of loss of control. Connor Leahy uses the term “fog of war” to talk about how, as layers upon layers of complexity build, there’s a complete illegibility that builds. In a sense, we already have that now. I do think the AI we have today has the same problem, but it’s not quite as extreme.
Systems at Google are very complex, and that’s why the market rewards Google engineers—they get paid an awful lot of money. But you’re talking about what happens when this gets taken to the extreme. You gave 3 examples: self-reinforcing dependence, irreversible entanglement, and cessation of authority. Can you explain what those are?
Superintelligence Strategy (Dan Hendrycks)
I describe this at more length in “Natural Selection Favors AI over Humans.” There’s a way in which we are becoming more dependent on these AI systems, that we’re starting to cede more of our decisions and cognitive processes to them, and that there will be more and more pressure to keep doing that without any clear limit at the economic level.
At the economic level, if it’s you versus a company that has AI systems that are just better than you, that company is going to win because those systems will be cheaper. At the military level, too, there’s a very strong incentive to, for instance, make the drones more autonomous because they can be jammed. Signal jamming makes them a lot less effective, so you make them autonomously move around.
These pressures will keep up, such that we’ll just voluntarily cede a lot of power in society to these sorts of systems. If we do it at a rate where we aren’t actually in control, that could be concerning. What does control exactly look like? When is it too much? When is it too irreversible? That’s a problem that we have to keep track of.
There are different outcomes. One is where you’ve actually just lost control: you can’t stop it, you can’t reverse it, you can’t bargain with it, or you can’t substantially steer it, and your fitness as a species is just collapsing as a consequence, or your livelihood evaporates.
Superintelligence Strategy (Dan Hendrycks)
Or there's a different outcome where you are like a retiree with a big retirement fund, where these sorts of AIs are working for you and they are doing your bidding, even though you aren't making nearly all of the critical decisions. So those are different outcomes, and it could be insidious or subtle as to whether we actually are having some of that counterfactual control versus whether we wind up in an unfortunate situation.
I think at least one thing that can help with this would be if AIs are better at forecasting or foreseeing outcomes and consequences. If we train them to do that, that could help make us more prescient and avoid some of these outcomes of extreme dependence and an erosion of control. That wouldn't be sufficient, but it would be helpful for seeing farther and knowing what we're getting ourselves into.
Wonderful. Dan, this has been an absolute pleasure and an honor having you on the show. Thank you so much for joining us today.
Superintelligence Strategy (Dan Hendrycks)
Yeah, thank you for having me.