爆炸式 AI 时间线预测 [Gary Marcus, Daniel Kokotajlo, Dan Hendrycks]
真正的关键红线,是实现完全自动化的 AI 研究闭环,把人类从开发过程中移除,让进展从“人类速度”切换到“机器速度”。 Kokotajlo 认为,递归式改进可能带来持久的战略优势;他援引 Dario Amodei 对智能爆炸的讨论,以及 Sam Altman 关于将10年的发展压缩到1年、甚至1个月的设想。如果某个国家控制了由此产生的超级智能,竞争对手可能被碾压;如果没有任何一方能够控制它,“所有人的生存”都可能受到威胁。Hendrycks 另行指出,率先触发递归的一方,即使只领先很短时间,也可能获得决定性优势。
拟议中的遏制方案包含3条具体红线:不得启动爆炸式 AI 递归;不得部署缺乏充分安全措施、具备专家级病毒学或进攻性网络能力的智能体;必须强化高能力模型权重的安全防护。 Kokotajlo 倾向于渐进式机制:各国透明地发展能力,研究每个层级,再讨论是否继续推进。Hendrycks 认为,协调可能先从公开表态和威慑开始,之后才走向核查或条约:“这场对话必须推进得远得多。”
时间线分歧很大,但在场没有人把风险视为可以安全置之度外。 Kokotajlo 已将超级智能的中位时间从2027年底推迟到2028年底;其他同事的中位预测大约在2029—2031年。Hendrycks 认为,到2030年达到人类水平的认知广度“不仅可能,而且相当有可能”;Marcus 则把2030年视为最快的合理情形,并将大部分概率放在10年之后,因为若要很快实现 AGI,就必须同时解决“无处不在、无所不包的一切问题”。
当前的资本开支趋势意味着:要么本十年内实现彻底自动化,要么 AI 进展将明显放缓。 Kokotajlo 估算,训练成本已从2020年的约300万—500万美元升至目前约10亿美元;如果保持同样速度,到2030年单次训练成本可能达到约5000亿美元。届时电力、晶圆厂、芯片产量、有限的互联网数据和企业预算都会成为约束;如果没有 AI 驱动的经济转型,他预计进展至少会逐步放缓,甚至出现“有点像 AI 寒冬”的局面。
前沿实验室最核心的治理失败,是一个被道德例外主义伪装起来的集体行动问题。 Kokotajlo 的说法是,DeepMind、OpenAI 和 Anthropic 的领导层都理解失控风险与权力集中的风险,但最终都接受了同一个“诱人的论证”:“如果我们不做,别人也会做。”他们相信自己是更负责任的一方,于是把已被承认的危险转化为竞速理由,使自愿自律不足以成为基准情形。
技术对齐仍明显落后于能力进展,尤其是在“差不多合理”远远不够的场景。 Hendrycks 认为,像拒绝生物武器请求这样的狭义防护,可靠性或许可以达到多个9;但如果稳健方法会让 MMLU 下降“一两个百分点”,实验室可能不愿采用。更广泛的要求——避免犯罪行为、侵权行为或可预见的伤害——仍然模糊;而智能递归是流程层面的问题,事先无法消除其中未知的未知。
乐观情形下的收益极其巨大,但产出充裕并不意味着所有人都能广泛拥有财富或保持政治自主。 Kokotajlo 描述了超级智能快速设计工厂、实验室、机器人、药物和定居点,直到物质需求得到满足;他还设想,不仅分配收入,也分配由密码学控制的“算力切片”。Marcus 则变得更加悲观:财富持有者或许会资助基本生存,却仍然保留“海滨地产”和权力,使得正向均衡取决于一套尚未有人提出的机制。
架构之争同时也是护城河之争:开放权重能让所有人站上同一条起跑线,却无法让所有人共享会复利增长的前沿。 Kokotajlo 认为,即使每个人拿到完全相同的权重,拥有大量 GPU 的参与者仍会率先利用 AGI 推进到 AGI+ 和 AGI++,形成强烈的规模报酬。Marcus 则预计会出现一次神经符号式的状态变化:到2035年,今天的 LLM 可能会被视为“不错的尝试”——仍然有用,就像智能手机出现后的功能手机,但并不是那个解决了推理、世界模型和稳健泛化问题的系统。
1. 只有在安全均衡中,充裕才有可能
Kokotajlo 接受这样一种判断:如果发展路径正在通向某种可怕结果,那么站在推土机前阻止它可能是合理的;社会应先停下来,找到更好的路径。他反对的不是 AI 本身,而是错误的建设方式;正确构建的系统完全可能“对所有人带来巨大好处”。
AI 2027 所设想的积极版“以放缓收尾”,是领先公司和政府只投入勉强够用的时间与资源来解决对齐问题,在最后时刻完成突破,同时仍然领先中国。Kokotajlo 强调,这不是他们的建议——它会让人类暴露在非同寻常的风险中——但它确实是一种世界可能“磕磕绊绊地过关”的自洽路径。
一旦系统在所有方面都胜过最优秀的人类,而且更快、更便宜,它们就能在创纪录的时间内设计机器人工厂、实验室、工业设备和一代又一代的新技术。Kokotajlo 想象,只需“几年时间”,自动化充裕、疾病治愈,甚至火星定居都可能实现。
Kokotajlo 不接受 AI 必然会把人类变成动物园里的动物,或让人类永久沉浸在 VR 快感中。一个可行的社会可以在满足多种客观福祉的同时保留自主性;更遥远的一种机制,是给每个人分配可租用算力切片的密码学密钥,把生产性力量而非只有现金分配出去。
2. 政治分配可能击穿乐观情形
Marcus 比几年前更加悲观,因为企业的逐利性让全民基本收入显得越来越不可信。即便充裕能够覆盖基本生存,他也不认为富有的所有者会交出“海滨地产”或政治权力;近来华盛顿的政治动态进一步削弱了他对任何非极端不平等结果的信心。
这留下了两个对齐问题:机器是否仍然可控,以及控制机器的人是否会以可接受的方式分配收益。Marcus 可以想象,在根本看不到政治解决方案时停下列车,因为技术成功本身并不会创造正向结果。
他偏好的表述仍然是暂停,而不是永久禁止。暂停倡议当时希望推迟 GPT-5,让研究人员先解决已知的 GPT-4 对齐问题——讽刺的是,他指出,两年半后 GPT-5 仍然不存在。Hendrycks 也赞成区分暂停与停止;Kokotajlo 则不认为,旨在让 AI 完全可控的技术研究能带来多少投资回报。
3. AI 研究自动化才是 destabilizing capability
Kokotajlo 把重点放在阻止超级智能上,因为相比明天停止所有 AI,这一要求似乎更符合地缘政治激励。实际触发点,是将 AI 研发彻底自动化:“把人从闭环中拿出去”,让开发从人类速度切换到机器速度。
Kokotajlo 援引 Dario Amodei 对递归过程和智能爆炸的讨论,以及由此产生的持久领先;他也援引 Altman 关于将10年的 AI 发展压缩到1年、甚至1个月的说法。造成战略不稳定的,正是这种速度,而不是对“AI”的模糊恐惧。
如果中国或其他国家控制了这样的系统,Kokotajlo 认为,它可能将这一优势武器化并碾压其他国家。如果没有任何一方控制一个几乎无人监管、运行极快的过程——而他认为这相当可能——那么全人类的生存反而可能受到威胁。
今天更快的超参数搜索并不是这里讨论的递归。真正的门槛是:系统能够提出那些在2024年仍需要人类提出的想法,然后完成测试、验证、调整、实施和扩展;3人都同意,实验室正在明确尝试达到这一能力。
4. 3条红线定义了可能的遏制机制
Hendrycks 提出3条红线:不得启动具备爆炸潜力的完全自动化递归;不得在没有安全措施的情况下,让广泛可获得的智能体具备专家级病毒学或进攻性网络能力;对于超过有意义能力门槛的模型权重,必须实施强有力的信息安全,防止流氓参与者窃取或外泄。
Kokotajlo 认为,递归是一个很好的起点,但他偏好的机制是渐进式的,而不是设定一条永久不许跨越的边界。各国应在相互透明的条件下发展能力,研究每个新层级,再公开讨论继续推进是否安全。
Marcus 指出,条约可能需要8年或10年才能达成,因此即便危险能力要到25年后才出现,提前启动工作仍然合理。Hendrycks 预计,协调会分阶段形成:先是公开立场、内部政策、泄密和威慑,也许还会出现一场促使各方要求核查的冲突,之后正式条约才具备可行性。
5. 只有递归启动后,领先6个月才有意义
Marcus 怀疑当前 LLM 范式能否带来持久的国家级优势;他说,如果行业继续停留在 LLM 上,“没有什么优势会是持久的”。他可以想象神经符号 AI 等真正不同的方法带来阶段性领先,但不认为当前范式会做到这一点。
Kokotajlo 对持久性的定义不同:竞争者可能在6个月后学会所有东西,却始终无法缩小差距,因为领先者会以同样速度继续前进。今天看似不大的6个月,如果被领先者用来触发递归式改进,就会变得具有决定性。
他也反对把“LLM”和未来范式截然分开。推理智能体已经在组合语言模型、工具和代码;2027年的 AI 可能会通过持续的范式迁移而显著不同于2023年的 AI,而不是由某一项离散发明突然改变。
Hendrycks 认为,公开能力透明度可以构成刹车:如果人们能看到某个实验室启动递归,“全世界都会彻底慌掉”。Marcus 援引民调称,大约75%的美国人已经感到担忧;Kokotajlo 则反驳说,他在 OpenAI 的经历是被施压淡化风险,而不是宣传风险。
6. 前沿实验室把恐惧转化为竞速理由
Kokotajlo 将前沿实验室内部长期讨论的两类风险分开:对齐失败会导致失控;即便对齐成功,如果公司或统治者掌控系统,也会产生权力集中。早期文章、领导层表态和泄露的通信都表明,这两个问题并不是后来才出现的。
他对持续加速的解释是“合理化”。“这大概无论如何都会发生;如果我们不做,别人也会做”的论证,让每个群体都得出同一个结论:赢下竞赛才是最安全的选择,因为它们更信任自己的能力与善意,而不是其他人的能力与善意。
DeepMind 的计划,是先建立足够大的领先,让 Demis Hassabis 能够放慢速度并解决安全问题。OpenAI 的出现,部分源于 Elon Musk、Altman 和 Ilya 不信任 Demis 掌握这种权力;邮件中甚至讨论过他可能成为独裁者的担忧。Anthropic 随后从 OpenAI 分裂出来,因为其创始人不信任 OpenAI 的安全行为。
Kokotajlo 的结论是绝对的:“我们绝对不应该”简单相信他们中的任何一个。政府需要配置人员跟踪前沿能力、访谈实验室、评估其计划并准备应急方案;真正应检验的是,公司是在帮助解决集体行动问题,还是继续重复背弃合作的模式。
7. 透明度可操作,但不会自发出现
Hendrycks 对公开披露内部峰值能力持乐观态度,尽管不一定包括方法或权重。他的激励逻辑是,中国可能已经知道美国实验室的大量情况,而外界对中国人民解放军的发展了解更少;披露可以缩小这一不平衡,提升可信威慑。
Marcus 区分了“可操作”和“自动发生”。一旦某一方拥有竞争者无法迅速重建的差异化知识产权,实验室披露信息的动机就会减弱;因此,他和 Kokotajlo 都不认为自愿自律足以成为可靠机制。
在 Kokotajlo 的模型中,开放权重不会消除集中度。即便所有人同时拿到前沿 AGI,拥有最多 GPU 的参与者仍能开展更多研究,率先抵达 AGI+ 和 AGI++;递归式发展因此天然包含规模报酬。
8. 预测都有长尾,但在2030年前分歧显著
Kokotajlo 更偏好概率分布,而不是单一日期,并将目标拆成多个里程碑:超人类编码、AI 研究完全自动化,以及超级智能。这些曲线体现的是主观判断,未来5年内有一个隆起,同时带有长尾,而不是把概率集中在某个日历年份。
撰写 AI 2027 时,他对超级智能的中位预测是2027年底;此后已将其推迟到2028年底。AI Futures Project 的其他预测者给出了2029年或2031年等更晚日期,但他说:“我是老板,所以我们采用了我的时间线。”
Marcus 把2030年视为最快的合理情形,并将大部分概率放在10年以后,尽管他承认突破最终可能加速进展。Hendrycks 在思考多维认知时仍刻意不下定论,但认为到2030年出现具备典型人类认知能力的系统“不仅可能,而且相当有可能”。
他们在尾部的距离更小:Marcus 和 Kokotajlo 都保留了围绕2045年及之后的实质性概率。真正重要的分歧,是当前限制究竟会成为具有约束力的瓶颈,还是仅仅成为前沿公司在未来几年内不断冲破的“一连串减速带”。
9. 扩展要么改造经济,要么撞上硬预算约束
Kokotajlo 认为,过去15年的大部分进展来自算力、数据和研究人员投入的扩展,其中算力可能是最大的投入项。他对近期的判断包含一个继续扩展的隆起;长尾则源于未来几年后维持同样的指数速度会困难得多。
电力供应只是一个约束。当公司无法再简单改造游戏芯片时,下一次10倍增长就需要大幅增加 AI 芯片产量和晶圆厂;即便没有任何单一因素形成明确断点,多种摩擦也会叠加。
Marcus 估计,数据已经构成约束:GPT-2 可能用了互联网内容的5%—10%,后续系统使用更多,而 GPT-4 已经用了其中大部分,包括视频转录内容。不可能继续把语料库扩大100倍。合成数据在答案可验证的领域效果最好,例如数学;Marcus 说,“思考模式”部分填补了这一缺口。
金钱体现了综合约束。Kokotajlo 对比称,训练成本已从2020年的约300万—500万美元升至当前约10亿美元;如果再增加两个半数量级,2030年单次训练成本将达到5000亿美元。若本十年末没有 AI 驱动的彻底自动化,他预计进展会逐渐放缓,甚至出现“有点像 AI 寒冬”的局面。
10. 算力预测提供的是先验,而不是答案
Kokotajlo 坦率地描述自己的预测方法:建立模型、观察趋势、进行计算,“然后把所有这些东西综合起来,凭感觉拽出一个数字”。数字仍然是主观的,但显式结构能让判断比毫无依据地给出日期更加有纪律。
Bio Anchors 框架将新研究想法所需的时间,与可用于暴力搜索的算力联系起来。算力越多,今天能尝试的事情就越多;新算法则会逐步把所需算力分布向下移动,把二维不确定性转化为一个按年份分布的概率。
在约 (10^{45}) 次浮点运算的规模上,他提出一个宽松的上限:即使完全不了解智能,也可以模拟地球、生命和10亿年的进化。当前训练运行约为 (10^{26}) 次浮点运算;在当前能力与这一上限之间分配不确定性,仍会给本十年赋予不可忽略的概率,之后具体基准测试证据应当更新这一先验。
11. 编码是递归式研究的第一阶梯
Kokotajlo 不认为人类会一步跳到超级智能。他设想的路径,是逐步自动化更多 AI 研究、更快抵达新范式,而起点将是编码,因为它是最容易自动化的研究活动。
METR 等团队的智能编码套件包含一些需要人类约4小时或8小时完成的任务。他预计,这些基准测试会在几年内饱和,随后又明确在基准饱和与编码完全自动化之间加入一个推测性的“缺口”。
他所说的超人类编码器并不是自动补全:用户给出高层指令,智能体就能像优秀的职业软件工程师一样完成工作。Marcus 认为学徒级工程能力已经接近,但预计未来10年内不会出现 Jeff Dean 级别的人——即能够理解前所未有的问题、迅速制作原型,并将解决方案转化为生产系统的人。
Marcus 的反对点在于原创性,而不是智能体能否自动化更多实验。当前系统可以“在盒子里”工作,但实现 AGI 可能需要 Einstein 级别的重新框定;Kokotajlo 则预计,自动化会帮助发现新的范式,把今天的弱点视为暂时性障碍。
12. Marcus 认为架构缺口并未被扩展解决
Marcus 的清单来自认知科学:分布外泛化、对变量进行操作、结构化表征、区分类型与 token、规划、推理、稳定的世界模型,以及灵活的领域知识。他说,自己在1998年和2001年指出的那些核心缺口,仍然在制造今天的失败。
例子被刻意选得很基础。1957年的经典技术可以解决任意规模的汉诺塔,但现代模型一旦改变长度就会出错;o3 会下出非法的国际象棋走法;而领域工程化的 AlphaFold 这一混合神经符号系统,在自己的专长领域明显胜过通用聊天机器人。
Kokotajlo 承认,一个 (10^{45}) FLOP 的进化系统可能是不透明的,但他说,能够建立文明的进化生物仍然会是有智能且具有经济用途的。Marcus 认同不透明是可能的——这正是他的担忧:“炼金术太多,原理太少”,使得有意对齐更加困难。
Marcus 将当前范式比作早期人们相信基因是蛋白质的时期。实验最终确立 DNA 后,分子生物学迅速发展;同样,AI 可能需要脱离错误假设,发生一次状态变化。如果要在2027年补齐所有缺失的能力,就必须做到“无处不在、无所不包的一切”。
13. 工具使用把争论转化为可靠性问题
Kokotajlo 认为,相关系统应被理解为 transformer 加上工具、插件和代码解释器。Claude 不需要在脑中执行汉诺塔算法,只要能查找算法、实现算法并返回正确结果即可——人类做算术时也经常使用外部工具。
Marcus 称这本身就是神经符号系统:网络生成 Python,由其中的符号和变量执行神经组件无法可靠完成的操作。如果模型总能选对工具,“那就万事大吉”;早期 LLM 连接 Wolfram Alpha 或 Mathematica 时经常时灵时不灵,问题正出在这一接口上。
Kokotajlo 将 METR 不断增长的任务时长图视为可靠性提升的证据。如果智能体每一步都有固定的灾难性错误概率,那么降低这一概率,就能让它完成更长的任务;新系统似乎既更少犯错,也更擅长从错误中恢复。
Marcus 的反驳是方法论层面的:这张图覆盖的是编码,可能存在未知的数据污染,而且成功标准是50%。据称另一张并行智能体图表也呈现出类似的视觉改善,但横轴只到秒级;与此同时,人类只需片刻就能选择合法的国际象棋走法,而 o3 仍然无法做到100%可靠。
14. 智能是多维的,基准测试只能选择性照亮
Hendrycks 警告存在“路灯效应”:研究人员会构建那些容易测量进展、且 AI 已经有抓手的基准测试。模型可能在某一代视频测试中登顶,但下一代测试会暴露此前被饱和掩盖的结构性缺陷。
随机抽取与儿童一起使用的认知任务,模型仍会在两位数比例的问题上失败,甚至接近一半:数照片中的人脸、连点、填色。流体智力、视觉处理、长期记忆、空间推理和物理理解,并不能与写作流畅度互相替代。
预训练花了数年才带来阅读、写作和结晶知识,而不是瞬间实现。Google 的 Minerva 在2022年达到 MATH 基准约50%的成绩,模型则在2025年前后将其彻底击穿;阅读和写作大约用了4年才成熟。Hendrycks 可能预计音频能力会在1年或2年内提升,但不认为完整的视频理解也会如此。
某一维度的严重缺陷,可能阻断经济自动化。长期记忆薄弱,会使工作型智能体无法维持状态并吸收组织背景;流体智力薄弱,则会限制泛化。Marcus 补充说,当前系统还难以标注图表或在物理环境中进行推理。
15. 新鲜问题揭示流畅度与泛化之间的差距
Marcus 将熟悉的游戏稍微移出分布。Grok 在一个只有边缘连续3子才算胜利的井字棋变体中失败,即便得到纠正也没有解决;国际象棋系统在被要求构造异常局面、改变棋盘几何结构,或灵活运用熟悉规则而不是复现正统棋局时,也会遇到困难。
训练数据的辩护无法说服他,因为前沿模型很可能见过国际象棋规则、棋谱、书籍和 Wikipedia,却仍然会走出非法棋步。科学评估仍然受到损害,因为实验室不披露数据或增强方式;他的根本问题是系统如何泛化到训练之外,而“我们对此完全没有透明度”。
他援引 LiveCodeBench Pro:据称其中全新的题目带来的机器表现是0%;还援引一项 USAMO 研究,称模型在新题上的表现很差。Frontier Math 的证据更不干净,因为 OpenAI 曾接触过这些题目,外部人员无法判断发生了何种相关增强,即便公司称没有用答案训练。
Kokotajlo 预计,这类反例每年都会变得更难找到,但不会一夜之间消失。Marcus 将这一模式比作自动驾驶:Waymo 持续改善,却仍然依赖地理特定地图和运行范围。Hendrycks 补充说,即使是写作,在几乎所有可用文本都已被利用的情况下,表现仍大约只有 GRE 风格评分的4.5/6。
16. 对齐是不断变化的可靠性问题,而不是已经解决的模块
Marcus 认为,模型在插值能力上取得了巨大的提升,但对齐并没有相称的进展。强化学习可以让模型拒绝明显的生物武器请求,但越狱仍然存在;系统仍会产生幻觉、复现被禁止的受版权保护文本,并违背简单约束。非法棋步正是这一问题的缩影。
Hendrycks 区分了对齐原始超级智能模型,与对齐制造超级智能的递归过程。前者是模型层面的问题;后者涉及一个前所未有、极度快速的过程,其未知的未知无法在一篇8页论文中预见。技术工作无法事先“完全消除”这一递归的风险。
当前系统能够“相当合理地”遵循指令,但 Marcus 说,当武器问题也只是“相当合理”时,这“让我他妈害怕”。Hendrycks 认为,具备对抗鲁棒性的生物武器拒答能力或许可以达到多个9,但生产实验室可能会回避那些让 MMLU 下降一两个百分点的方法;而禁止犯罪、侵权或可预见伤害,则远没有那么容易实现。
Hendrycks 预计,新的失败模式会不断出现,针对性的修复和待处理问题也会越积越多,因为智能体扩大了攻击面。要求是具备适应性能力、留有余量,并设置安全预算,以便在火灾出现速度快于蔓延之前将其扑灭。因此 Marcus 支持对安全关键型部署进行干预,也将明确的神经符号约束视为一条路径,即便低风险用途仍可继续使用聊天机器人。
17. 不同情景最终都指向核查机制
Marcus 会大幅放缓 AI 2027 的进程。他预计神经符号方法会在3年或4年内获得发展,到2035年,人们可能会把 LLM 看作“不错的尝试”:它仍然擅长分布式学习,但并不是带来深层语义、规划、稳定世界模型或 AGI 的架构。
人类证明了远高于当前系统的数据效率,而机器最终应当超越人类的记忆能力,并避开确认偏误和动机性推理。Marcus 预计,进展会沿着不同维度逐项推进,缺失的想法出现后还会迎来更快的阶段;他质疑的是两年内实现 AGI,而不是最终出现远超人类的系统。
他还希望看到情景分布,因为一个生动的叙事可能压倒读者对概率的判断。Kokotajlo 说,AI Futures Project 正在构建一个详细的美好结局,以及一组更短的“粗略情景分布”,同时仍然认为当前限制更像减速带,而不是持续数十年的障碍。
Hendrycks 的战略图景强调间谍活动、透明度和破坏:前沿实验室的 Slack 系统和领导人的手机可能遭到入侵,数据中心的电力基础设施也可能在难以归因的情况下被干扰。尽管对时间线和架构存在分歧,3人都支持围绕智能递归建立核查机制,不信任实验室的自愿治理,并同意对齐进展尚未赶上其必须治理的能力进展。
For instance, if the United States disrupts China's ability to develop a superintelligence, the main way in which China would develop it is if it gets the ability to automate AI research and development fully and take the human out of the loop. Then you go from human speed to machine speed. We've had people such as Dario Amodei, for instance, at Anthropic, talk about how such a recursive process could lead to an intelligence explosion, and that would lead to a durable edge where nobody will be able to catch up. Last week, Sam Altman discussed how this process could telescope a decade's worth of AI development into a year or potentially a month.
There's been this very seductive argument that has appealed to all of these people, which is basically, “Well, it's probably going to happen anyway. If we don't do it, someone else will.” Basically, all of these people trust themselves more than they trust everyone else and have therefore convinced themselves that even though these risks are real, the best way to deal with them is for them to go as fast as possible and win the race.
Why did they make OpenAI? They were worried that they didn't trust Demis to handle all that power responsibly when he was in charge of the AI project and the fate of the world rested in his hands. So they wanted to create OpenAI to be this countervailing force that could do it right, distribute it to everybody, and not concentrate power so much in Demis' hands. In fact, the emails that came up in the lawsuit were talking about how they were worried that Demis would become a dictator using AI.
On the proto-AI things, I think that they follow the instructions fairly reasonably. There are some other parts of it.
Yeah, that scares the shit out of me, right? “Fairly reasonably.” When these things are still in our hands, that's kind of okay. But if you have them guiding weapons or something like that, there are many circumstances where “fairly reasonably” is not good enough.
I'm Gary Marcus. I'm a cognitive scientist and an entrepreneur. I've written 6 books, most recently Taming Silicon Valley. I think that we all want what is best for humanity and hope that AI can be positive for humanity and not negative. I think the three of us are maybe not known for agreeing on everything, but we actually have a lot of shared values around that, and I'm looking forward to the conversation.
Thanks, Gary. My name is Daniel Kokotajlo. I'm the executive director of the AI Futures Project. We are the people who made AI 2027, which is a comprehensive, detailed scenario forecast of the future of AI.
I'm Dan Hendrycks. I'm the director of the Center for AI Safety. I also advise Scale AI and xAI. I've also done research in machine learning. I coined the GELU, a common activation function. More recently, I've done evaluations to evaluate capabilities such as MMLU, the MATH benchmark, and Humanity's Last Exam. I've been focusing on trying to measure aspects of intelligence and also AI systems' safety properties for mostly the course of my whole career.
The common things that we care about are: What is going to be the outcome when AGI comes? How soon is it going to come? How do we forecast these things methodologically? One person wrote in a question that maybe we could actually start with: What is the positive view? I take it that all 3 of us in the room are united in thinking that AGI could be a good thing. None of us is standing down in front of a construction machine, saying, “Stop it all, period.” I think we all think there's at least a possibility of a positive outcome, but you can stop me if I'm wrong. We're hoping to steer toward that positive outcome. So our first question is: What could be the upside here, and why aren't you just saying, “Let's forget about AI altogether”?
I can start with that.
Yeah. So first of all, I actually think it's quite reasonable to stand in front of the bulldozer and say, “Stop.” I think that eventually we want to build AI because it is possible to get it right and make it in such a way that's beneficial to everyone—in fact, massively beneficial to everyone. But if we are currently not on a track to getting it right, and we're currently on a track to make it in a way that's going to be horrible, then it makes sense to stop until we figure out a better way.
As for what the benefits could look like, AI 2027 has the slowdown ending, in which things go really well for almost everybody. Since not everybody read it, I think more people probably read the darker scenario than the positive scenario.
Lay out for people who might not have read the positive scenario a little bit about what the spirit of that is.
Sure. So in the slowdown ending of AI 2027, they manage to devote just barely enough effort, time, and resources toward the technical alignment problem that they solve it just barely in time, so that they can still beat China, as they have wanted to do. By “they,” I mean the leader of the leading AI company and the president, and maybe some other companies—that group of people.
The slowdown ending is not our recommendation for what we should actually aim for. We think to aim for this sort of scenario would be to expose the world to incredible amounts of risk of various types. So we don't think it's our recommendation. Nevertheless, it's a coherent, plausible scenario for how things could muddle through and end up pretty well for everybody.
And “pretty well” is, I mean, the Peter Diamandis abundance scenario, for example? Is that what you have in mind?
I haven't actually read that, but I would say that when you get AIs that are superintelligent—what I mean by that is better than the best humans at everything, and also faster and cheaper—you can just completely transform the economy. Superintelligent-designed robot factories are constructed at record speed, producing all these amazing robots and all this amazing new industrial equipment, which are then used to construct new types of factories and new types of laboratories to do all sorts of experiments and build all sorts of new technologies.
Eventually—and by “eventually,” I mean in only a few years—you get to a completely automated, wonderful economy of all sorts of new technologies that have been iterated on and designed by superintelligences. Material needs are basically met for everybody. There's just an incredible abundance of wealth to distribute. Things like curing all sorts of diseases, putting up new settlements on Mars, and so forth all become possible.
That's the potential upside. Then, of course, there's the question of whether we can actually achieve that, and if we do have the technology to achieve it, who's in control of the technology? Do they actually use their power over the army of superintelligences to make that sort of broadly distributed, good future for everybody, or do they do something more dystopian?
So I think, well, first of all, we'll come back to the bulldozers and whether you think we're actually at that point or not. But I think we can probably agree—and we'll let you chime in, too—that there is a technical alignment problem. We should talk about that, and a political, if not alignment, problem—maybe that's too close a pairing of words—but a political issue about who controls this technology if it exists. That's deeply important. We should talk about both sides of that: political control and technical alignment.
Let me get your answer, though, on that first question. Do you think we need to stand in front of a bulldozer now? If you don't, what do you think is the positive?
Yeah. I think there are 2 concepts. There's AGI and there's ASI: artificial general intelligence and artificial superintelligence. I think AGI doesn't have a clear definition, so it's difficult for me to say that, for all definitions of it, it would be worth standing in front of.
Some people would say we have AGI already.
Let's just start with what we have now.
There could be an argument that right now we should just stop the train. Sorry, I've moved from bulldozers to trains. There would be an argument that we're already so far along this path, and if we don't resist it massively right now, it's going to be bad, and there isn't a good outcome here—a good enough outcome to offset what seems inevitable. I'm not making this argument, but some people have made it. The argument would be that if we don't resist it now, then the outcome is going to be bad. I'm not making that argument, but I've heard people make it.
And so, partly as a grounding exercise, where are we on that? One argument would be that we've already screwed up. We need to have every person on the planet get in front of this train or bulldozer and say, “Just stop it,” because there isn't a positive enough outcome to warrant it. Another position would be that there's actually a really positive outcome here, and if we can guide it in the right way, we can get to that positive outcome. Here's why I think it's positive. That's the kind of set of issues I wanted you to speak to first.
I primarily think about stopping superintelligence. The main reason for that is because it's much more implementable, geopolitically compatible with existing incentives, and so on. It's something that I can actually foresee. I think stopping things tomorrow is not something I can as easily foresee, so I don't really think about that.
But is that just fatalism? If you're worried about superintelligence, you might think that if we move even toward AGI, which presumably is closer in time, that's a bad idea because we won't be able to stop the superintelligence. Maybe we don't want to risk going further.
I couldn't conceive of how we would coordinate now to do that. I think the asks are instead having, say, for instance, the United States disrupt China's ability to develop a superintelligence. The main way in which they would develop it is if they get the ability to fully automate AI research and development and take the human out of the loop. Then you go from human speed to machine speed.
We've had people such as Dario, for instance, the CEO, talk about how such a recursive process could lead to an intelligence explosion. That would lead to a durable edge where nobody would be able to catch up. Last week, Sam Altman discussed how this process could telescope a decade's worth of AI development into a year or potentially a month.
I think that's extraordinarily destabilizing for 2 reasons. One, if they control it—if a state controls it, such as China—then all the other countries are at substantial risk because that superintelligence could be weaponized and used to crush other countries. If they don't control it, which I think would be fairly likely in a nearly unsupervised, extremely fast-moving process, then everybody's survival is also threatened.
Either way, this very fast, automated AI R&D loop is quite destabilizing, whether a state controls it or not. I think it makes sense not just because AI is scary in some vague sense, but because there are very strong geopolitical incentives for state self-preservation.
I've got that. I want to come back to the geopolitical aspect of it, and I skimmed the paper that you just did with Alexander and Eric. But I don't think I got an answer for the first part, which is the upside.
I don't think I need you to lay out the upside.
Yeah, yeah, yeah. So, for the upside, I think people generally think that if we have AI, it will necessarily hollow out all other values, or that we'll have an extreme pressure toward one of them. We'll be zoo animals, or we'll be all pleasured out and just in our VR experiences constantly, having superficial experiences.
I think you could set a society up so that people have the autonomy to choose between different ways that they would want to live their lives, given the resources that AI-increasing GDP could provide. I think there's a way you could have a list of objective goods being met in society. I don't think there's no equilibrium where things can work out. I think there's a way of achieving a variety of goods and still preserving human autonomy, like sort of what we have now.
Imagine we had things now, but then we had the autonomy to choose how we want to live our lives. There'll be more resources.
Yes, that's right. That's right. That's right.
There are other sorts of longer-term questions about how you might do that. How are you going to distribute power? Maybe that would be with compute slices that people would rent out, for instance, where they would have the unique cryptographic key to activate that compute slice. That way, you're not just distributing wealth, but you're also distributing ways of generating that wealth.
There's a lot in this deep future that one could pull out, but I think there are some paths where things work out well.
I like your term “equilibria.” I think we probably all agree that there are multiple equilibria here and that we're all trying to steer toward positive equilibria. I think we would also all agree that there are lots of different ways to get to both the positive and the negative. We might disagree about some of the paths, some of the likelihoods, and so forth, but I think we're all 3 operating under that general assumption: There are multiple ways this could go, and some of them are good and some of them are bad.
I'm a little darker than I used to be, even a couple of years ago, because of the economic inequality sorts of issues. A couple of years ago, when I trusted Sam a lot more than I trust him now, I heard him talking about universal basic income. I've always thought that was an important part of the equation.
It seems to me that the more data has driven things, the more acquisitive the companies—and it's not just OpenAI—have been, the less realistic it has seemed to me that we're going to have anything like a universal basic income. I think there is some scenario under which there's so much wealth that everybody's subsistence gets met. I don't expect that wealthy people are going to give up the beachfront property under any circumstance, or their power.
I think the dynamics that we've seen in Washington lately have made me, though I know our politics may not be the same, less optimistic about any kind of relatively equal distribution, or even just not a relatively extreme distribution. I think the positive outcomes do depend on getting some answers there: What would be the incentives, the mechanisms, and the dynamics such that things are well distributed?
I guess I'm sort of making this up as I go, but I think there's an argument here for stopping the train if we can't see any solution to the political thing. Most of my own personal concerns are really about technical alignment, which we'll talk about soon enough, but if we don't have even an envisionable political solution, that does make me nervous. I could see an argument for, “Let's stop the train now,” because we just don't see how to do it.
Going back to something you said at the very beginning, there's an interesting argument about delaying the train versus stopping the train. I think all of us signed the pause letter. Maybe you didn't sign the pause letter?
Okay.
Only I signed the pause letter. That's interesting. That's fine. The part of the pause letter that made me sign it, and that I liked best, was that it said, “Let's pause this particular thing that we know is problematic in certain ways.” It was really about delaying the development of GPT-5, which, ironically, still doesn't exist 2½ years after some of us signed the pause letter.
The notion was that we would pause the development of GPT-5 because we knew that GPT-4 had certain kinds of problems around alignment, and that we would spend that time instead working on safety. It was explicitly constructed as a delaying tactic. It wasn't saying, “Never build AI.” It wasn't saying, “Don't do any more AI research.” It was, in fact, saying, “Do more AI safety research and wait.”
That still seems like maybe not a politically viable thing in this moment, but at least a strategy one might consider. Maybe we should be constantly updating our estimates on how likely we are to get the positive outcomes versus the negative outcomes, and how much that would change as a function, for example, of putting more resources into safety research as opposed to capabilities research.
What's your take on what I just said?
Totally agree. I like the pause-AI framing more than the stop-AI framing, for the reasons you just mentioned. There's a lot more to say about more nuanced proposals for what is to be done. I do think we should maybe save a lot of this for the later part of our discussion, after we've gotten through the technical stuff—timelines, alignment, and so on.
But you're the moderator, so I'm going to use the moderator's prerogative for now. You can push me a little bit later if we don't get there.
I'm trying to find some sort of common ground. I think we have lots of common ground, and then we'll get to the differences. Do you buy the notion that some kind of pause should also be on the table, or no?
I think making time for technical research—I'm not expecting much of a return on investment from that. I don't think most success strategies, particularly in the foreseeable future, go through AI being fully controllable. So I do technical research.
I'll give you a comeback on that, like an extreme version of the argument I just made. It comes from the cognitive psychologist and evolutionary psychologist Geoffrey Miller, who had a tweet that I really liked. It was: We should wait until we can build this stuff safely, even if that takes—I forget what he said; I'll say 250 years.
It was an interesting framing because everybody's thinking, “What should we do next week? Should we sign this pause letter?” It was deliberately extreme. I think it was maybe even 500 years. If it takes us 500 years, we should just wait. I could see an argument for that, and in fact, I wonder what the counterarguments are.
The process that I described earlier, that sort of recursive loop—I guess you could call it intelligence recursion—which, if it goes fast enough, is an intelligence explosion. That is not something you can research your way out of. You can't just write an 8-page paper and then say we've solved it in the news tomorrow. We figured out how to fully derisk some extremely fast-moving process that we've never done before, and all of its unknown unknowns have been anticipated.
Well, politically, we could—it would take a lot of willpower, and it's probably not likely—but we could try to have a global treaty: don't go there, don't work on these kinds of things, and let's report it if you do. You could at least imagine that kind of scenario. I don't see any technical way of coping necessarily with that set of problems right now.
But if we were—not just the 3 of us in the room, but if we as a society were convinced that, let's say, that was a red line—we could decide as a society that there are certain red lines. You suggested 1 red line, and there are others. Maybe recursive self-improvement might be a reasonable one to consider. We could say we won't cross any of the red lines, we'll make treaties around them, and we just shouldn't do it until we have answers to them.
I think it would be worth states clarifying that we don't want anybody doing that type of fully automated intelligence recursion, because it'll be destabilizing. That doesn't necessarily immediately take the form of a treaty. Initially, you have them exchanging words about that and articulating their preferences, potentially explicitly or potentially through internal policy at the CIA and elsewhere, and the other parties come to learn it through leaks or directly.
You eventually want to gain more confidence that nobody's trying to trigger something like this. This process has multiple stages. It may involve things even like skirmishes—it requires a conversation. It may even require a skirmish for people to think, “Okay, we need to do a verification now,” and then maybe you get some type of treaty. But you can still have various forms of coordination through deterrence without anything like a treaty, just as we've had strategic stability on multiple issues without treaties necessarily. So that's potentially a later-stage thing, but you have to have the conversation advance far further.
Daniel's giving me a look.
No, I'm just looking back and forth.
You're just looking back and forth.
Why am I in this? You have to rub your neck because I'm right in the middle.
Yeah, in the middle.
Okay. So let's separate for a moment the deterrence stuff, although I know you're keen to talk about it and know more about it than I do, and let's try to get to it. You could sort of separate the dynamics for which we would form treaties. Maybe we need some disturbances and skirmishes, as you just talked about, but separate that from the question of whether the rational thing for civilization to do right now would in fact be to sign these treaties, because we're basically pretty close. I may take a longer timeline than you, but we're reasonably close to recursive self-improvement of at least some sort.
Maybe we don't want to go there. Maybe the rational thing to do right now would be to make those treaties. Or even if we thought that was 25 years away, because we know treaties take 8 or 10 years, maybe we should be putting all our intellectual capital or political capital or whatever into doing that right now. What do you think?
Yes. So I think 3 red lines would be: no recursion, where that's a fully automated—not just some AI-assisted one—but one that has that explosive potential; no AI agents with expert-level virology skills or cyber-offensive skills made accessible without some safeguards; and model weights past some capability level need to have good information security for containing them and making sure that they're not exfiltrated or stolen by rogue actors.
In a minute, I'm going to ask you how close we are on these various things. But what's your take on what he just said?
Are those the same as the first? I don't know about the second 2, but definitely the first one. I think that if somehow we could coordinate on that, that'd be great. I'm not sure if that's the best red line to coordinate on, but it's an excellent place to start, at least.
Do you want to throw in any others now, or can you later in the conversation?
I think the type of thing that I'm probably going to end up advocating for is going to be more like—we are going to gradually develop AIs with these capabilities, but we're going to do it in a way that's mutually transparent to each other and that proceeds slowly and cautiously. We all debate whether it's safe to go to the next level, and then after we get there, we study it for a little bit and debate whether it's safe to go to the next level, and so forth.
So the sort of thing that we're probably going to end up advocating for is going to look more like that rather than a line that we're all not going to cross. But in terms of the thing that you really need to stop from happening in the short term, that sort of recursive self-improvement thing is, I would say, the number 1 thing.
This conversation is a little bit depressing, in the sense that many of the things that we seem to be worried about actually seem fairly close. Maybe the person in the room who's most pessimistic, if that's the word—let's say the most extended in time on this—is me. But all of us, I would say, think that—well, no, let me rephrase this question. It's a dark set of answers relative to the reality right now, I think, in the following sense: even if you think it's going to take a while to get to AGI or ASI or something like that—and we'll talk about that in a little bit—the things that are red lines, and I like your red lines, people are already pushing against them. They may not be breaking through them.
Depending on your definition of recursive self-improvement, I might give you different estimates. Mine might be a little longer than yours, but people are already trying to do that, right? And transparency is 1990s talk. It's in the rearview mirror—I'm exaggerating a little bit—but it's in the rearview mirror. OpenAI was open originally; it is not open anymore. So there are elements still pushing for transparency, but there are certainly elements pushing against it.
I'm somewhat more optimistic about increased transparency and more situational awareness from governments. One reason for that is I think it's incentive-compatible for the US to be more transparent about what's going on at the frontier, because China already knows. Meanwhile, there's somewhat less transparency around PLA, or People's Liberation Army, developments. So the people to gain more from information would probably be the rest of the world. China has somewhat more of an advantage there.
This sort of levels things, and having high transparency into the frontier is very useful for making more credible deterrence.
Interrupt for a second about a time-frame question. I think the notion of a durable advantage in LLMs, if that's what the technology is, is a myth. If we stay on LLMs, nothing's going to be durable. But I think you think about these things in a little bit different way than I do.
I think it depends on what you mean by paradigm. I would say LLMs are just a part of this, or a subset of this overall AI research. They're not the only possible AI design, and progress is going to continue being made. I would say that, in some sense, arguably we're already seeing a move away from LLMs with things like these so-called reasoning agents that have access to tools and can write code and so forth.
So there's going to be this continuous shift. Rather than discrete paradigm shifts, it'll be more like a continuous paradigm shift, and the AIs of 2027 are just going to be quite different from the AI of 2023. But taken as a whole, I do think that the United States could potentially end up having a durable advantage over China, and 1 particular tech company within the United States could end up having a durable advantage. I see this as possible.
I see that as unlikely unless somebody approaches the problem in a pretty different way. I can imagine a neurosymbolic approach; it would be just so different from what anybody else has that I could see, at least for a while, an advantage over the current way people are doing things. I don't even know what that would look like.
Perhaps I should clarify what I mean by a durable advantage. I don't mean they never figure out what you already figured out. I mean that by the time they catch up, you've already moved ahead. It's like you keep running as fast as they're running, so that even though they're only 6 months behind, they can never quite catch up, because by the time they do, you're 6 months ahead again.
And that was part of the other part of that question I wanted to get to: Does 6 months matter? Is that enough to change the world or not?
Right now, it doesn’t seem huge, but if you are the first to trigger a recursion, that can matter a lot. It could be the moment where it matters.
And that goes back to the other question, which is on the recursion thing. The fact is, people are trying it, as far as I understand it. It’s the plan.
It is. It’s the plan. Yeah, it’s still highly AI-assisted, and nobody could try to spin this up right now and have it foreseeably lead to something substantial. But what I see right now is that you could do your hyperparameter search faster or something like that, and you could call that recursion if you want, but that’s not really what we’re talking about here.
We’re talking about a system finding a fundamentally new idea that in 2024 would only have come from humans, and the leverage of not only coming up with the ideas but testing them, validating them, working out the tweaks, implementing them, and scaling them.
I think that’s exactly what people like Sam and Dario are now hinting that they’re going to be able to do soon. I’m not really buying those claims, but I think that’s what they want to do now.
100%. Yep.
Yeah, and that is destabilizing and scary and dangerous.
Let’s say potentially destabilizing. I mean, if it’s just hype and they can’t really do it, it’s not destabilizing. But if one of them achieves it—if it were technically feasible—that would be destabilizing.
I mean, I think we can agree it probably is technically achievable. It’s just a question of whether you can do it with current technology or later, and how far in the future. There’s no serious argument that this can’t be done. I think we all think that’s going to happen.
And so one advantage of transparency, then, is that if there were very high visibility as to what’s going on at the frontier—if they’re basically triggering this sort of process and the public is kept reasonably informed as it’s happening—I think the world would be very much freaking out.
Yeah, so I think that possibly sunlight might be a way of providing substantial pressure to prevent that.
Freaking out might be too strong, but it’s worth mentioning that polls of the American people—not necessarily in Asia, but American people—show that they’re pretty worried. Maybe it’s not at the top of their set of worries, which might be economic concerns and so forth, but you look at these polls and 75% of the American public is already worried. I expect that worry will increase, not decrease. At least, I don’t see any active thing that is going to decrease it.
The thing that might decrease it is propaganda from the tech companies. I think the tech companies seem to like scaring people, as far as I can tell.
Oh, I disagree. I think there’s probably some of that there to some extent, for sure, but at least my experience at OpenAI was that there was more pressure not to talk about the risks in public and to downplay that sort of thing than to play it up. There was no pressure to play it up, at least not as far as I could tell.
And, you know, if you look at the public messaging by Sam over the years, he certainly started to talk about it less and less and less as well.
I mean, partly because I pushed him, but when we talked at the Senate, he didn’t answer the question of what we were most worried about. I made Blumenthal come back, but Blumenthal said it was jobs, and Sam gave his explanation for why he wasn’t that worried about jobs.
I said to Blumenthal, “You should really ask him a question.” Sam then said that his biggest worry was, I can’t remember the exact words, “substantial harm to humanity.” When he was back at the Senate a few weeks ago, his biggest worry was, in so many words, that they wouldn’t make all the money and extract all the value they could, or something like that.
So, I think it’s true that he used to talk more about this kind of risk stuff and talks less now. That’s an interesting data point. It was around the time that I did my Senate appearance, which was May 2023.
Reconstructing the timelines a little bit, after that there was a second letter related to the pause letter, which I did not sign. That was the one about the extinction letter.
Yeah, I signed that one.
That’s what we signed. I didn’t sign that one, and we should soon get to extinction risk, but we’ll come back to it in a minute. Sam did, I think, sign that, right? At that time, it was still popular among CEOs to express concern about this. Maybe it’s true that they do a little bit less now. Dario still alludes to this.
Right. Right. Yeah. A lot of people that I have spoken to think that it’s a deliberate mechanism to try to hype the product—to talk about the risk, like, “Look how great our stuff is. It might kill us.”
Yeah, I had lots of scientists sign that. I didn’t, but, yeah, I understand.
Yeah. So, first of all, loads of people have these views who aren’t conflicted and trying to hype up the companies.
But then, to your point about the hype, my take on what’s been happening is that DeepMind, OpenAI, and Anthropic have been full of people who were thinking about superintelligence from the very beginning, at the highest levels of leadership. Therefore, they have been considering, at least somewhat, both the loss-of-control risk and the concentration-of-power risk: Who gets control of the AIs from the beginning?
This is documented in all sorts of ways. You can look at their old writings and so forth, and some of the leaked emails. Then you might ask, “Why are they building it if they’ve been thinking about these risks?”
There are two risks: loss of control and concentration of power. Loss of control is, what if we don’t solve the alignment problem in time and the AI takes over? Concentration of power is, if we do solve the alignment problem, who controls the AIs? What goals do we put in them? Do we risk becoming a dictatorship or some sort of crazy oligarchy or whatever?
You can find writings from people at these companies—both senior researchers and CEOs—going back decades, talking about these things. Then you might ask, “Why are they doing this? Why are they racing to build these things if they’ve been thinking about these risks?”
I guess, in a single word, it would be rationalization. There’s been this seductive argument that has appealed to all of these people, which is basically: “It’s probably going to happen anyway. If we don’t do it, someone else will, and it’s better for us to do it first than for someone else to do it, because we’re the responsible good guys who will wisely solve all the safety issues and then also beneficently give UBI or whatever to make sure that everything works out well.”
Basically, all of these people trust themselves more than they trust everyone else. They have therefore convinced themselves that, even though these risks are real, the best way to deal with them is for them to go as fast as possible and win the race.
This is DeepMind’s plan, so to speak. Demis Hassabis’s plan was basically to be there—to get there first with this big corporation, Google. Because you have such a lead over everybody else, you can slow down and get all the safety stuff right, making sure that everything goes well before everybody else catches up.
That seems naïve. That plan was torpedoed when Elon, Sam, and Ilya made OpenAI.
Why did they make OpenAI? They were worried that they didn’t trust Demis to handle all that power responsibly when he was in charge of the AI project—when all the fate of the world rested in his hands. They wanted to create OpenAI to be a countervailing force that could do it right, make it distributed to everybody, and not concentrate power so much in Demis’s hands.
In fact, in the leaked emails—or the emails that came up in the lawsuit—they were talking about how they were worried that Demis would become a dictator using AGI.
We can see how well that’s worked out. All the Anthropic people basically split off from OpenAI because they didn’t think OpenAI was going to handle the safety stuff responsibly. Now they’re claiming that they have the technical talent and will be able to sort out the alignment issues better than everyone else.
It’s a mess.
Yeah. Do you have anything to add? I agree that it’s a mess. I don’t know if you may not want to answer this on camera, but do you feel confident that we should trust any of these particular labs?
We definitely should not. I think pushing for things like transparency, as well as having the government employ people whose job it is to keep track of these developments, is important. The government should be coming up with contingency plans, interviewing these labs about their plans, and conducting internal assessments of their probability of success.
All of that seems useful, and I think at least some of the players here would probably be willing to push for that type of stuff, while others wouldn't. So I think the question is: are they willing to help solve these collective-action problems, or are they going to continue defecting? There's been a lot of defection.
To get back to your question, I realized I never actually answered why this ties into it. Early on, when these companies were fresh, young, and idealistic, their founding mythology was basically, “Yes, the risks are real, and that's why you should come work at our company, because we're the good guys.” When that was still fresh and still the main thing they were saying, they talked about it a lot.
But now, when their sort of founding myth is laughable and not something they can say with a straight face anymore, it has eroded. They're also under a lot of political pressure to get investors, ward off regulation, and stuff like that. The rationalization wheels are continuing to turn, and they're coming up with new narratives to justify what they're doing.
There are also a lot more players at the table, both within the U.S. industry—which is most of what we were talking about—but also in China. Then we have the whole—I'm not sure if we have time to go into it, but maybe we will, maybe we won't—the mechanisms of open-sourcing, or at least open-weight models, which means that essentially anybody can get into this game to some degree.
I disagree with that to some degree.
Oh, go ahead.
To some degree, sure, but no—walk me through the disagreement.
Hypothetically, suppose that someone achieved AGI and immediately opened the weights for everybody. It's not going to happen like that, but suppose it did. Then, for a brief, glorious moment, anybody with enough GPUs could run their own AGI right at the frontier of capability.
However, AGI isn't just AGI. There's going to be AGI+, AGI++, and so forth. Whoever gets to AGI+ is going to be the one who has the most GPUs, so they can run the AGIs to do the research fastest. Even if you gave everybody exactly the same starting point at the same level of capability, the people who had more GPUs would pull ahead slowly but surely over everybody else.
So I think there's this unfortunate effect around that, but go ahead.
There's this unfortunate, sort of inherent—I don't know if I want to say winner-take-all effect, but there's a returns-to-scale effect inherent in the dynamics of an intelligence explosion. That's part of what makes this so scary.
Yeah. All the actors are incentivized to push it as fast as possible and to have the maximal resources in order to do so. They're also incentivized not to open it up, and not to be transparent about it and so forth, unfortunately.
I think we have a little disagreement there about transparency and how much we might expect. I think you're a little bit more optimistic, and I'm being a little bit less optimistic about transparency. I'm especially less optimistic about transparency as we get closer to having actual AGI, or let's say differentiated AI.
Right now, I would say that what's out there is not very differentiated. Everybody is using LLMs with reasoning models and so forth. There's some differentiation in how people do their RL and what their datasets are, but I don't think there's so much value in transparency until somebody has something that is really unique, which they think other people aren't just going to reconstruct very rapidly.
As one gets to that point of having some unique piece of intellectual property—which I think will happen—there's even more reason to be less transparent. I think I'd be optimistic about being transparent about the capabilities of the best models internally—not necessarily the methods used to create them or the weights, but at least the public knowing what's going on, what the peak capabilities are that the models are exhibiting, and being aware of that.
Are you being optimistic that they'll do this by default, or that this would be good?
Obviously, it'd be good. I think it's good, and that there's more tractability for this than for a lot of other asks.
I agree. More tractability means people might agree to do it.
Yeah. That's why I've been making these asks. But I'm just saying that they're not going to do it by default. If nobody asks them to be transparent about these things—
No, I agree.
So we can agree that voluntary self-regulation is probably not enough to get there.
Yeah. I wouldn't bank on it.
All right. Let's talk a little bit about time, forecasting, and scenarios. We've talked a bunch about strategy—what we should do going forward and what policy should be. Some of that depends on timelines. If we thought that we had 1,000 years before AGI, then we might make different choices than if we thought we had 6 months.
If we thought that LLMs were the answer to AGI, we might make one set of choices. If we thought they were definitely not the answer to AGI, we might make a different set of choices, perhaps focusing more on research. So let's talk about our forecasts around that, and maybe I'll throw in our forecasts around when we might sort out the safety problems—the alignment problems.
I will make the argument that we've made some progress toward AGI and very little toward alignment, and we can see if we agree on that. I want to start by using a notion that many in the audience will know, but not all: a distribution of probability mass.
You could make a simple prediction and say, “I think AGI will be here in 2027 or 2039,” or whatever. But I think we all understand that that's the unsophisticated thing to do. You can certainly say your best guess for which particular year you think it might come, but as people who are either scientists or know something about science, we know there's what we call a confidence interval around that. It might be this, plus or minus that.
The most sophisticated thing to do is to draw out a curve and say, “I think some of the probability mass will come before 2027, some of it will come before 2037, and some of it will come after.” In your appendix to AI 2027, you go through this in a fair amount of detail. I'm not sure how many people got to the appendix—I think it was a little hidden—but it was there, and it did a good job of that.
It gave the forecast for 4 different people in a sort of qualitative way. I don't know if it showed the full curves, but it said, “This is the chance that they think it will come before 2027.” For 3 of the 4 forecasters, or something like that, some of the probability mass was after 2040. You remember better than I do, so maybe I'll start with you. I think you've put the most work into trying to get detailed probability-mass distributions, and maybe you can talk about what your own are, the different techniques people have used, and where you are on that.
Thank you for that excellent launch to this. Feel free to stop me if I ramble too long, because it's a huge topic with lots to talk about.
First, the exciting part: the actual numbers. When we were writing AI 2027, we had our different medians, or 50% marks, for AGI—or, I think we divided it up and didn't use the word AGI. We used different milestones: superhuman coder, full automation of AI research, and superintelligence.
For those things, let's say for superintelligence, I was thinking there was a 50% chance by the end of 2027. That's why AI 2027 depicts it happening at the end of 2027: I was illustrating my median projection.
The other people at the AI Futures Project tended to be somewhat more optimistic than me.
Let's clarify: “optimistic” flips depending on how you think about it.
Sure. They tended to think it would take longer—to 2029, 2031, something like that. Let's say they were more conservative by a few years. But I was the boss, so we went with my timelines.
Happily, by the time we actually published it, I had lengthened my timeline somewhat. These days, I would say a 50% chance by the end of 2028.
I think you're wrong, but if we have an extra 12 months, I'm not sure that's enough to handle all of this. It's still quite scary to prepare for.
In terms of what the shape of our distributions looks like, they tend to have a sort of hump in the next 5 years and then a long tail. The reason for that is that the pace of AI progress has been quite fast over the last 15 years, and we understand something about why it's been so fast.
Basically, the reason is scale. They've been scaling up compute, scaling up data, and drawing ever more researchers into the field. Compute is probably the main, most important input that they're scaling up. That's turbocharged progress, but the companies simply won't be able to scale things up at the same pace after a couple of years.
Is that a function of power—or electrical power? What is the function? What's the rate-limiting step by which you think that kind of scaling won't continue?
It won't be like a sharp cutoff, but it'll be a couple of things.
Partly, it'll be power supplies. Partly, it'll just be compute production. Much of the world's chip production, even after building new fabs and converting much of the world's chip production into AI chips, won't be enough. They'll have to produce 10 times more fabs in order to scale up by 10 times, right? Whereas previously, they could just take chips designed for gaming and repurpose them for AI.
So, in a bunch of little ways that are going to add up, there are going to be all these frictions that are going to start to bite, which will make it harder for them to continue the crazy exponential rate of scale. I think we're already seeing that with data. I don't even know the actual numbers, but let's say that GPT-2 used maybe 10% of the internet, or 5%, or something like that. Maybe you guys know the actual numbers. GPT-3 used a significantly larger fraction. GPT-4 used most of the internet, including transcriptions of videos and stuff like that.
You can't just keep 100×ing that because there just isn't enough data. There's new data generated every day, so you can always eke out a little more, and people are turning to synthetic data as well. They should, but that's not a universal solvent, and it works better for things like math, where you can verify that the synthetic data are better. I think we're already running against that kind of bottleneck on one of the raw resources that go into, at least, the current approaches.
Another thing I would add is that you can also just think about money, which you can use to buy many of these things. They've been scaling up the amount of money that they're spending on AI research, and on training runs in particular, over the last decade, but it'll be hard for them to continue scaling at the same pace.
They're probably already doing something like billion-dollar training runs. If you go back to 2020, what was the biggest training run? Like $3 million, something like that—$4 or $5 million. They've gone up by 2.5 orders of magnitude in 5 years. If it's another 2.5 orders of magnitude, we're doing a $500 billion training run in 2030. There starts to be not enough. There's just not enough money in the world. The tech companies just won't be able to afford it. Even if they've grown bigger than they are today, the economy just won't be able to afford it.
So that's why we predict that if you don't get to some sort of radical transformation—if you don't get to some sort of crazy AI-powered automation of the economy by the end of this decade—then there's going to be a bit of an AI winter. At the very least, there's going to be a sort of tapering off of the pace of progress. That stretches out a lot of probability mass into the future because that's a very different world. It could take forever—well, not forever, but it could take quite a long time to get to AGI once you're in that regime.
I'm going to ask you one or two more questions. I'm going to insert mine, and then I'm going to come to Dan. One note on data: We sort of ran into that bottleneck maybe 2 years ago. The main things that have been continuing the pace would be this thinking-mode type of stuff, outside of the trends that were existing previously. So I think it's sort of picking up slack in some ways for the fact that most of the internet has been trained on.
The methodology by which you came up with these curves—can you just tell us a little bit about them?
These curves represent our subjective judgment, which is very uncertain. They're just our opinions. The way that I would like to say it should be done is that, rather than just pulling a number out of your ass, so to speak, you should come up with models and look at trends. Then you should have little calculations that attempt to give numbers, and you should stare at all of that and pull a number out of your ass based on all that stuff that you've just looked at.
The main arguments and pieces of evidence that we found to inform our overall estimates were what we would call the benchmarks-plus-gaps argument. Perhaps I should also go into the compute-based forecast. Have you heard of the Bio Anchors framework by Ajeya Cotra?
I don't think I have.
I'll briefly mention that before getting into the benchmarks-plus-gaps thing. The Bio Anchors framework is called Bio Anchors because it references the human brain. I think that part's actually the less exciting and plausible part of it. The part that I think is more robust and more worth using is this core idea that you can think about the trade-off between more time to come up with new ideas and do AI research, and more compute with which to do the AI research.
You can make a big 2-dimensional plot and imagine: 10 more years, 20 more years, 30 more years—how does the probability that we get to AGI go up with more time? But you can also imagine not more time, but just more compute. Could we get to AGI today if we had 5 orders of magnitude more compute, 10 orders of magnitude more compute, or 30 orders of magnitude more compute?
The insight there is that the answer is probably yes. For example, if you had 10^45 floating-point operations, you could do a training run that's basically just simulating the entire planet Earth and all life evolving on it for a billion years with that amount of compute. You don't really need to understand how intelligence works at all if you're building it with that type of training run, because there's no insight coming from you. You're just letting nature do its thing and letting evolution take its course.
The thought is that we can make a non-guaranteed but soft upper bound at something like 10^45. Then you can make other soft upper bounds. You can think, well, what could we do with 10^36? I wrote a blog post about this in 2021: Suppose we had 10^36 FLOPs. What are some really huge types of training runs we could do, and what's our guess as to how likely that is to work?
You can start to smear out your probability mass over this dimension of compute. You have a soft upper bound, and then you have, of course, a lower bound, which is the amount that we already have done. We clearly haven't done it yet with this amount of compute. That gives you this smeared probability distribution over compute.
Then you think, okay, but now we're also going to get new ideas. As new ideas come along, we're going to be able to come up with more efficient methods that allow us to train it with less compute. You can think of your probability distribution as shifting downward while the amount of actual compute increases. That gets you your actual distribution over years.
I think this is the right sort of basic framework for calculating these sorts of timelines. But it's a relatively abstract, low-information framework that doesn't really look at the details of the technology today or the details of the benchmarks. I think it's a good way to get your prior, so to speak, but then you should update based on actual trends on the benchmarks and so forth, which is what I'm about to get to.
The reason I mentioned this prior process is that 10^45 floating-point operations isn't actually that far away from where we are right now. Right now, we're at something like 10^26 for training runs, and we're going to be crossing a few orders of magnitude in the next couple of years. Even if you just smeared out your probability mass with maximum uncertainty across the orders of magnitude from where we are now to 10^45, there'd be a non-negligible amount of probability that it's going to happen in the next few years.
Even on priors, you should think it's decently plausible that it could happen by the end of the decade. Then you should update your prior based on the actual evidence, which I'll now get to.
The actual evidence, I would say, is to look at agentic coding benchmarks. That seems to me to be the most informative thing to look at. The reason for that is because I don't think that the fastest way to get to superintelligence is in a single leap, where humans come up with the new paradigms in their own brains. I think, rather, it's going to be this more gradual process where humans automate more of the AI research process, and then that gets us to the new paradigms fast.
The lowest-hanging fruit, as far as the AI research process is concerned, that's going to be automated first, is the coding. I'm looking to see when we'll get to the point where the coding is basically all handled by LLM-like AI assistants. We have benchmarks for that. Places like METR have been building these little coding environments and doing all these coding tasks. The companies themselves have been doing this, of course, because they are racing as fast as they can to get to this automated-coder milestone.
So we look at those, and we extrapolate trends on them. We forecast that, in the next couple of years, they're basically going to saturate.
We’re going to have AIs that can just crush all of these coding tasks. They’re not something to scoff at. They’re not just multiple-choice questions; they’re the sort of tasks that would take a human 4 or 8 hours to do. But that’s not the same thing as completely automated coding.
First, we extrapolate to when they saturate the benchmarks, and then we try to make our guess as to what the gap is between the first system that can completely saturate these benchmarks and the first system that can actually automate coding. That’s probably the more speculative part, but we do our best to reason about it.
Now we’re ready for our first full-on disagreement of the day. I understand your logic there. I think it’s well thought through, but it’s missing the cognitive science for me. My approach to this is more from the cognitive science: I see a set of problems that a cognitive creature must solve, many of which I wrote about in my 2001 book, The Algebraic Mind.
I don’t feel like we’ve solved any of those problems despite the quantitative progress that we’ve made. Those include generalizing outside the distribution, which I think remains a huge problem. I think the Apple paper was an example of that. There are actually 2 Apple papers I discovered today, but the Apple paper with the Tower of Hanoi stuff is an example of problems with distribution shift.
I think we’ve seen many of these problems over the years. We see that these systems have trouble doing multiplication with large numbers unless they call on tools, and so forth. I think there’s lots of evidence for that. I think there’s a problem of distinguishing types and tokens that leads to bleed-through when you’re representing multiple individuals from some category, which leads to hallucinations.
I wrote an essay recently about the hallucinations that I think ChatGPT made about my friend Harry Shearer, who’s a pretty well-known actor. It misnamed the roles of characters that he played in the movie This Is Spinal Tap and said that he was British when he’s American, and so forth. I think this blurring together that we see in hallucinations remains a problem.
I could go on with a list of others. I think there are several having to do with reasoning, planning, and so forth. The way I look at things—which is not to say that there isn’t some value in what you’re doing—is more focused on these cognitive tasks.
What I say to myself is, what would AI look like 2 years before we achieved AGI or ASI, or something like that? Certainly, 2 years before we achieved ASI, we would have full solutions to all of those things. If we specified an algorithm for something, we would expect the system to be able to follow it.
Current systems can’t even play chess reliably according to the rules. So o3 will not—sorry, o3 will sometimes make illegal moves. It can’t avoid illegal moves. Another thing I would expect is that current systems, when we’re close, would be basically equivalent to their domain-specific counterparts, or at least be close.
AGI means artificial general intelligence. I would say the reality is that domain-specific systems are actually much better than the general ones right now. The only general ones we have are LLM-based. For example, AlphaFold is a carefully engineered hybrid neurosymbolic system that far outperforms what you could get from a pure chatbot or something like that.
Somebody just showed that an Atari 2600 beat—I think it was o3—in chess. Sometimes very old systems will beat the modern domain-general ones. On Tower of Hanoi, Herbert Simon solved it in 1957 with a classical technique that generalizes to arbitrary lengths, whereas LLMs do not generalize to arbitrary lengths and face problems with it.
I could go through more, but the gist of it is that I don’t see the qualitative problems that I think need to be solved. I’ll just go to your 10^45, because it’s really interesting to me.
I wrote a piece once with Christof Koch—I don’t know if you guys know him, the neuroscientist. We wrote a science-fiction essay. It’s the only published science-fiction essay I’ve ever written. It was in a book that I wrote called The Future of the Brain, which I guess we wrote in 2015. I was the editor.
We wrote the epilogue to it. It was set in something like 2045, and the notion was that it was a book about neuroscience and that, by that point, we would actually have created an entire simulation of the brain. But it would run slower than the human brain and would still not have taught us anything about how the brain really works.
We’d have this simulation, but we wouldn’t understand the principles of it. We’d have a neuron-by-neuron simulation, maybe even a protein-by-protein simulation, but we could find ourselves in a place where we’d replicated the whole thing without really understanding how it worked.
I do wonder with the 10^45: even if you trained on everything, would you have solved distribution shift? Would you have abstracted principles that allow you to run efficiently, effectively, and usefully in new domains, and so forth? I think it’s a really interesting question.
I’d never thought of the 10^45, even though I read at least some of your paper. I admit I didn’t get to that level of detail.
Well, this part wasn’t in 2020, so it wasn’t in that appendix.
I admit I didn’t read the whole appendix. I read some of it. I think it’s a really fascinating thought experiment.
Maybe I’ve said enough. Just to lay it out, there’s one way to extrapolate on the basis of things like compute, and I think you’ve done a masterful job of doing that as well as can be done, while acknowledging that there’s still an element of pulling things out of one’s behind—which is true on any account. Nobody can really do this in a closed-form way.
Then I have a different slice on it, which is: where are the qualitative things that I want to have solved? I think Yann LeCun, with whom I often disagree about many things, would be closer to me. He would probably give a different set of litmus tests that he’s looking for.
We would both emphasize world models. We have slightly different ideas about world models, but he would say, “I don’t think we’re close because we don’t really have world models.” Neither of us, I think, is satisfied with the current thing that some people call reasoning but neither of us thinks is robust enough. I think he and I both take an architectural or cognitive approach.
Great. I’ll list a bunch of bullet points of things we could discuss, and hopefully we can get through them each. If we miss some, at least I put them out.
In reverse order: the 10^45 scenario, where you just sort of brute-force evolve intelligence. Indeed, you would not understand how it works at all, but nevertheless, you would have it. You could take those evolved creatures out of their simulated environment and start plugging them into chat products and using them in your economy.
They would be smart. They evolved. They built their own civilization in there, so they’re pretty smart and pretty good at generalizing, and so forth. You wouldn’t understand how it works, but you’d still nevertheless have the AI system.
Indeed, that’s kind of what I think is happening today for us. We don’t understand how these AI systems work very well. We’re throwing giant blobs of compute at giant data sets and training environments, and then trying them out, playing around with them afterward, and seeing what they’re good at.
That’s part of my concern about it: I think there’s too much alchemy and not enough principles.
100%. This is part of why the alignment issue feels so looming to me. We don’t even know what we’re doing. How are we supposed to craft a mind that has the right virtues and the right principles when we don’t even know what we’re doing?
Next thing: you mentioned Tower of Hanoi and math problems—these are ways in which the current AI systems seem limited, particularly limited and fragile, with hallucinations in ways that you think we’re not really making progress on.
There, I would say that I also can’t solve Tower of Hanoi, and I also can’t do large math problems in my head. I need tools. I need to be able to program a little bit.
I’m going to come in on that briefly. One is that relatively young children can actually do Tower of Hanoi really well if they care about it. They can do even pretty big ones.
There’s a video, I think, of a kid doing 7 discs lightning-fast, in 2 minutes, on YouTube—a kid who enjoys it. He’s probably 12 or 15, or whatever. I’m sure he can do 8 discs if he wants to, because it’s a recursive algorithm. I’m sure he’s learned it.
Some humans, if they want to, can do arithmetic. We humans do get into memory limitations, but you shouldn’t expect AGI to have the same limitations. We should pause for a moment on the definition of AGI, right? It has had varied definitions.
The one that I always imagine is that it should be at least as good as humans in a bunch of things and better in certain ones. I would not be satisfied with an AGI that can’t do arithmetic. I’d be like, “Yes, okay, it’s equivalent to people in this respect, but I would actually expect more.”
Especially if we’re talking about the risks that concern us, I think we’re really talking about a form of AGI that probably can do all the short-term memory things that people can’t and has the versatility and flexibility of humans.
You can argue about that. I would say that the weakest AGI would be as good as Joe Sixpack, who's really not very good at reasoning, is full of confirmation bias, has not gone to graduate school, and doesn't know how to do critical reasoning. You'd say, “Okay, that's AGI because it does what Joe Sixpack does.”
I think most of the safety arguments are really around AGI that's at least as smart as most smart people. Smart people can, in fact, do things like the Tower of Hanoi. Certainly, another example I gave you was chess, right? 6-year-olds can learn to follow the rules of chess. o3 will do things like have a queen jump over a knight, which is just not possible.
The failures there are quite striking and not something you need even an expert to do. So, I guess I would say it feels to me like we're making progress on these things, or at least that these barriers are not going to be barriers for long.
I think, for example, that maybe the—well, that's assuming your conclusion, that last little piece. Let me tell you more. Let me tell you why I think this.
Maybe the transformer by itself might have trouble doing a lot of this, but you should think about the transformer plus the system of tools and plugins that you can build around it—Claude plus Code Interpreter and things like that. I think that system could solve the Tower of Hanoi and things like that because Claude might be smart enough to look up the algorithm, then implement the algorithm, and do it.
I'm curious for your immediate reaction to that point.
I've always advocated neurosymbolic AI, and I think that this is actually a species of neurosymbolic AI. I don't think it's the right one, but the point of the neurosymbolic AI arguments that I made going back to 2001 was that you need symbols in the system to do abstract operations over variables.
There were several other arguments, but that was the core argument. What you're doing when you have Claude Code Interpreter is operations over variables in the Python that it creates. You're moving it to a different part of the system. Pure neural networks don't have operations over variables and fail on all of these things. Symbolic systems can do them fine.
Here, you're using the neural network to create the symbolic system that you need in order to solve that particular problem. So, it's absolutely a neurosymbolic solution. I think the rate-limiting step is that they don't always call the right code that they need. If you could make that solid enough, that would be great.
One of the first attempts at this strategy was to put an LLM into Wolfram Alpha, which is a totally symbolic system, rather than putting Wolfram Alpha into Mathematica. The results were hit or miss. The problem was the interface for getting to the tools.
If you can really reliably get to the tools, then you have a neurosymbolic system that works, right? If you can have the neural networks reliably call the tools that they want—or that they should be calling relative to the problem, I should say—then you're golden. Empirically, it's a bit hard to get the tools to work reliably.
That's helpful, because I think that sort of collapses a lot of your barriers into one barrier, which is reliability. Perhaps you would agree that if we can get them to—we should come back to the world-models piece of it, but, yeah, go ahead. Go run with that.
Well, yeah. There, I would say it does seem to me like the AIs have been getting more reliable over the last couple of years. One piece of evidence I would point to is the horizon-length graph from METR, which you've probably seen me talk about.
They have this suite of agentic coding tasks that are organized from shortest to longest in terms of how long it takes human beings to complete the tasks. Then they note that the AIs of 2023 could, generally speaking, do the tasks from here to here, but not do the tasks above this length. Each year, the crossover point has been lengthening.
A very natural interpretation of what's going on here is that the AIs are getting more reliable. If they have a chance of getting into some sort of catastrophic error at any given point, then if it's a 1% chance per second, after 50 seconds they're going to get into an error. But if that goes down by an order of magnitude, then they can go for more seconds, and so forth.
The thought here is that this is evidence that they're getting, generally speaking, more reliable—better at not only not making mistakes, but recovering from the mistakes they make. Not infinitely better: they're still less reliable than humans, but substantial progress is being made year over year.
I wrote a whole paper about why I hate it.
Oh, okay. Interesting. The way I would interpret this graph—and I'm sure many of our audience members have already seen it—is that there is a general curve of reliability over time.
I see the argument you're making with the particular graph. I think there are a lot of problems, and some are actually relevant to this argument. One is that it was all relative to coding tasks; it wasn't tasks in general.
With the coding tasks, I'm very concerned from a scientific perspective that we don't know how much data contamination there is, how much augmentation is done relative to those benchmarks, and so forth. Which leads me to my next point about it: I can't remember the guy's name from 80,000 Hours who posted something on Twitter yesterday, but it was a parallel to that graph. I think it was also by METR, looking at agents.
The y-axis for the coding thing was sort of minutes to hours to days or whatever, and for agents it was seconds to tens of seconds or something like that. It was a totally different y-axis. With the coding thing, it was coding agents; this was something more like web-browsing agents or something like that. The argument was that you're getting the same kind of curve, but the scaling was completely different on the y-axis.
There's something there for both of us, right? What's there for you is a general curve of reliability over time. What's there for me is that the performance is still pretty poor on those agent things. More generally, I don't like the axis at all. I think it's very arbitrary.
In my critique, we give a bunch of examples, but how long does it take to do this task? Also, it's all to the point of 50% performance. It's just a very weird measure. You can find many things that a human can actually do in 3 seconds that, according to the graph, should have been done by a 2023 model.
Here's one off the top of my head: choose a legal move on a chessboard. This takes a grandmaster—I don't know—100 milliseconds or something like that. But o3 still can't do that task with 100% reliability. It can do it with 80% reliability or something like that. It's just weird to read anything off this graph. I'll send you the Substack essay on it.
Yeah, thanks. One other thing I was going to mention, which is a new topic, is that I think in a lot of these cases my defense of the AI is that they were never trained on this task. But that's the essence to me. I understand they fail on things that they're not trained on. That's an oversimplification, but if we're talking about AI risk, I understand that the really dangerous AIs of the future have to be able to do stuff without having been trained on it.
I would say, at least, to the same extent that humans can do things without having training. To take your grandmaster example, grandmasters can do that task in less than 3 seconds with greater than 80% reliability, but also with 100% reliability. They've trained on it a lot. They've done lots of it. Whereas with o3, possibly—I don't know the details of o3's training process—but it's possible that there was not a single playing of a game of chess at all in the training process.
I don't think that's very plausible. They've seen transcripts of games of chess. They've seen the explicit rules in Wikipedia. I know they've been trained on Wikipedia. They've probably read Bobby Fischer Teaches Chess, which is how I learned to play chess.
They have books on chess, and they've read everything related to chess. They've read everything in, what is the famous phrase, publicly and privately available sources or whatever. That wasn't quite the famous phrase, but you get the idea.
I think they probably actually have a lot of data on chess. As a side note on transparency from the scientific perspective, it's very hard to really evaluate these systems because we don't know what's in the training set. The core question, going back to my work in 1998, is: how do they generalize beyond what they've been trained on? We just don't have transparency on that.
So, first of all, I totally agree about the transparency thing. Back to the training thing, I want to talk about Claude Plays Pokémon, and I want to again talk about the chess thing. I wasn't aware of this 80% statistic for o3, but I'm curious what the statistic is.
I don't know the exact number. I hadn't heard about it, but I had a friend look into this, and he was easily able to do it.
I can give you another example. Let me put one other example on the table, with a similar flavor: I played Grok in tic-tac-toe, and I came up with some variation. I might write this up this weekend, so I asked it for some variation. We played it, and then it offered another variation: “Let’s play tic-tac-toe where you can only win if your 3 in a row is on the edges.”
I said fine. We had some moves back and forth, and then it suggested some moves for me at a point where I could, in fact, make 3 in a row. So it had failed to develop the right strategy. I hadn’t played this variation before either, but it was obvious what to do. It failed to identify what 3 in a row was, which I would say goes to a conceptual weakness: you should have had enough data to understand what 3 in a row was, but it suggested other moves and didn’t recognize it. It’s a pretty bad failure.
Then we played again after I corrected it. You get these sycophantic answers. We played again, and it lost to me again, and then a third time. I posted all of this on Twitter a couple of weeks ago. So if you move these things off the typical problem, they’re even worse.
If they had a robust understanding of a domain like tic-tac-toe or chess, it wouldn’t be too hard. I posted another one on Twitter, which was a variation on chess: “Give me a board state where there are 3 queens on White’s side but Black can immediately win.” Performance was not great. You take knowledge that, for an expert, would be outside what they’ve specifically practiced, but they understand the domain well enough. Experts will have no trouble at all, and these systems still do.
So I totally agree with those limitations about current systems, but the thing I would say is that the trends seem to be going up. I would predict that if you measured this sort of thing for the last 5 years, even though the current systems might still have weaknesses, they would be less weak than the systems of 2 years ago, which were less weak than the systems before them.
Chess is an example, and then I’m going to go to you. I just want to say one last thing: I first noted the problem with chess, I think, 2 years ago—the illegal-move problem. I think Matthew Akre[?] was the first to point it out, and I spread it on Twitter and said, “Look, this is a serious problem.” It persists. There are a lot of problems that I feel have persisted. But your turn. You’ve been fine.
One thing about benchmark-based forecasting is that it has limitations, in particular the streetlight effect: when it gets 100%, that will suggest that it has solved the task. The issue is that they often have structural defects that are only obvious later, or when you go fairly far out into the curve.
For instance, in video understanding, you might think, “Wow, if you looked at the benchmarks from a few years ago, they’re totally at the top of them,” but then you can come up a year later with several other sorts of benchmarks that challenge them. I think there tends to be a gravitation toward what’s very tractable for AI systems and where some interesting action is happening. That’s a selection pressure on the sorts of benchmarks.
If one is looking at cognitive tasks such as those that you would give kids if you’re testing their intelligence, for instance, and you randomly sample many of those, the models don’t do that well. It could be a double-digit percentage—maybe almost half of them—where they don’t do that well. For instance, count the number of faces in this photograph. o3 can’t do that very well. Or connect the dots, or fill in the colors in this picture. That’s just an example for visual ability.
I don’t think when Gary’s pointing these out, he’s just running the cherry-picking program and doing a God-of-the-gaps thing for AI until it basically is AGI. I do think that there is a non-adversarial distribution, a different sort, from the cognitive science angle, where you would see many of these sorts of issues.
I think there’s potentially some speaking past each other, in part because he’s not viewing intelligence as some unidimensional thing entirely. A lot might be correlated together, but there are various other mental faculties that are important that aren’t really online that much. We saw with GPT—with the GPT series—that by pretraining on a lot of text, we got some subcomponents of intelligence: reading and writing ability, and a lot of crystallized intelligence, or acquired knowledge, from that. That process took several years.
So it’s not the case that people will point out, once it gets traction on something, that it will solve it immediately or very shortly thereafter. For a lot of these core cognitive abilities, they took multiple years. Mathematical reasoning would be a recent example with Minerva. Google’s Minerva system got 50% on the MATH benchmark, a benchmark I made some time ago, in 2022, and I think we’ve only recently crushed it in 2025. So it took a good 3 years.
I think reading and writing ability got to an interesting state and is now relatively complete 4 years later. I think crystallized intelligence as well is now relatively complete. However, there are still a variety of others. I wouldn’t expect visual processing ability, including video, to be taken care of in a year or 2. Maybe audio ability would be taken care of in a year or 2, though.
We don’t see almost any progress on long-term memory. Basically, I think they can’t meaningfully maintain a state across long periods of time over complex interactions. Once we start to get traction on that, then we can start to forecast that out. For fluid intelligence, which is what we see in the ARC-AGI thing and Raven’s Progressive Matrices test, that still seems very deficient as well.
So I think when thinking about these, it’s important to split up one’s notion of cognition, because if there’s a severe limitation on any of these dimensions, then the models will be fairly defective economically, or at least for many tasks. For instance, if a person doesn’t have good long-term memory, it will be very difficult to enculturate them and teach them how to be productive in a working environment and load all that context. If they have very slow reaction time, that’s also a problem. Or if they have low fluid intelligence, that will limit their ability to generalize substantially.
I think the numbers are going up substantially or continually on many of these axes for these benchmarks. But at the same time, when those benchmarks reach their conclusion, there might be another peak for those capabilities. There are many of these, and to be an AGI, you’re going to need all of them as well. For some of them, they’re fairly early on in their capabilities, such as long-term memory.
So I think if one’s doing forecasting, understanding intelligence somewhat more and having a more sophisticated account, such as one would find in cognitive science, can help foresee bottlenecks that will actually be fairly action-relevant.
I completely agree with you. I’ll just mention physical intelligence and visual intelligence. You mentioned visual; we left out—I skipped physical. Physical and spatial intelligence are really serious limits right now. So if you ask o3 to label a diagram, for example, it would be quite poor at it. If you ask it to reason about an environment, it’s going to be poor at it.
I agree with you. You made it sound like I think intelligence is a single dimension. I don’t know where you got that impression. I agree that there are all these limitations. Well, you said reliability in an all-things-considered type of sense. I think your forecasts don’t make a lot of contact with the kind of stuff that the 2 of us just talked about.
This became 2 against 1 in a way that I entirely did not bring on. I’m usually a bridging type. But anyway, this is why we call it “benchmarks plus gaps.”
Can you just explain that phrase? Can you just explain “benchmarks plus gaps”?
The benchmarks: you take all the benchmarks that you like, and you extrapolate them and try to see when they saturate. Then the gaps are thinking about all the stuff that you just mentioned, and thinking about how, just because you’ve knocked down these benchmarks, that doesn’t mean you’ve already reached AGI. There’s all this other stuff that the benchmarks might not be measuring. You have to try to reason about that.
Another way of putting my argument is that I identified a set of gaps in 2001. Rightly or wrongly, I identified them, and I think they were real in 2001. Those applied to multilayer perceptrons, not to transformers, which hadn’t been invented yet. But I still see exactly the same gaps, and I see some quantitative improvement but no principled solution to any of them.
There were 3 core chapters in that book. One was about operations over variables, one was about structured representations, and one was about types and tokens, or kinds and individuals. I just don’t see that the things I described then have been solved. I see certain cases where they can be solved, but all of the problems I see seem to be reflections of those same core problems.
And so, if I seem like a grumpy old man, it's because for 20-some years—and really, I pointed these things out in 1998—27 years, I have not really seen them. The notion that we're going to solve those gaps all in 3 years seems weird. I thought your estimates of solving a few of the things you just mentioned were generous. You say we solve reading and writing in 3 years. Video is much more uncertain, but let me finish the sentence.
The kinds of numbers that you gave were like 3 or 4 for this one, 3 or 4 for that one. My estimate is really 10 years—you know, most of the distribution is past 10 years. It's because I see several problems that we'd be really lucky if we solved, let's say, the visual intelligence problems in 3 or 4 years. We'd be really lucky if we solved the video kind of visual intelligence-over-time problems in 3 or 4 years. I just see too many problems, too many gaps, to use your phrase, to think we'd get all of those at once. It's like in the movie Everything Everywhere All at Once. To me, getting to 2027 would be Everything Everywhere All at Once for a set of things that I've been worried about for 25 years.
Excellent. First of all, fun point. It sounded like you were disagreeing with me, but then you were saying, “Yeah, 3 to 4 years” in terms of the overall picture. Basically, there's maybe a caricature—which maybe isn't exactly you—but the caricature would be: things will keep scaling, and that will automatically solve the problems. That's what's been happening. You point out an issue; it's the word “automatically” that I bristle at.
I'm just interested in the bottom-line numbers for the years. We had this potential dispute about whether it's just LLMs or whether it's neurosymbolic. I don't want to fight you about that. We can say it's neurosymbolic if you like. We can say it's not automatic, but rather involves some new ideas. But the point is, I think it's going to happen in 3 or 4 years.
Yeah. Okay, great. I just think new ideas are hard to project. I'll give you my favorite example. In the early 20th century, everybody thought genes were made of proteins. Somebody even won a Nobel Prize on that premise. They were all wrong, and it took a while for people to figure out that this was wrong. Oswald Avery did these experiments in the 1940s, which really ruled it out and, by process of elimination, showed that it was DNA. It didn't take that long to move very fast in molecular biology once people got rid of the bad assumption.
I think we're making some bad assumptions, but it's hard to predict when people will move past an ad hoc fix, which is where I think we are now. Proving that is very hard. I can give a lot of qualitative evidence that points in that direction, but there's at least a possibility that I'm right about that. If we don't have the right set of ideas, it depends on when we come up with new ideas. It could be that we automate the discovery of new ideas, but most of what LLMs have done is not discovery of new ideas. It's really exploiting existing ideas. So I would claim that there are many paths to the timelines being shorter before 2030, for instance.
But the picture that you point out—that there are these various other cognitive abilities that are long unaddressed, and this isn't a cherry-picked distribution—is actually what you would do if you're inspired by cognitive science or psychometrics. I think basically both positions are legitimate, but if you integrate them, there could be a bullish case that you'll be able to knock off some of those core cognitive abilities by the end of the decade. It's not totally certain.
It's funny you say that. I think the absolute most bullish, if we want to use that word, case is 2030. In the very fastest case, it'd be 2030. It would require solving multiple problems that I think we're just not positioned to solve quite yet. I don't think it's very likely, but I think that's the maximum possible. I just can't get my mind around 2027. It does not seem plausible because of the number of problems.
I think also what you said is right: we want to integrate these 2 approaches to make the forecast, and nobody's really quite done that. Maybe that's an adversarial collaboration that we could do. I haven't really seen something dig into that side of it, in terms of new discoveries. Maybe your Gaps does it a little bit.
Let's talk about the superhuman coder milestone—the automating-coding aspects, which is not all of R&D, but just part of it. I'm excited to try to drill down into that and for us to try to make bets or predictions about what the next few years are going to look like.
From my perspective, it feels like you have often, over the last couple of years, been pointing to limitations of LLMs that were then overcome in the next year or 2.
I mean, there are specific examples.
Yes, right.
For example—well, chess hasn't been solved. Let me clarify what I mean by specific examples, and I'll come back to you. I give, on Twitter, a sentence—here's a prompt—and it gets a weird answer. Those are always remedied. I was told—I don't know if it's true—that people at OpenAI actually look on social media for examples, including mine, and patch them up.
The specific examples in that sense—literally, this prompt gets this weird answer—a lot of them have been solved. These more abstract things, like playing chess, actually have not been solved. The more abstract problem of whether I can give you a game with variations on the rules hasn't been solved. I talked about that in Rebooting AI in 2019, in the first chapter.
Are you then saying that, for all you know, 6 months from now the AIs will be able to solve the legal-move-in-chess thing? Or are you saying you don't think LLMs will solve it?
I think someone could come up with a scheme to get LLMs to play chess without making illegal moves by calling a tool. I think that's possible. Without calling a tool, I don't think it'll be possible in the next 6 months. I'd be very surprised. You can ridicule me in 6 months if it happens, and that's fine.
But let me add a twist to it, which is what we talked about in Rebooting AI. For example, if you train a particular system on—I think we used Go as an example—if you train it on an ordinary 21 × 21 board, will it be able to play on a rectangle instead of a square? Will it handle different variations?
I had a conversation with a guy at DeepMind who did chess in a really interesting way without using Monte Carlo tree search, which is what the AlphaZeros and so forth do. He trained it on an 8 × 8 board, and I said, “If you just had to do 7 × 7, would it work?” He said no. He did a version where there was no tree search, so it's not neurosymbolic. The tree search is one of the symbolic processes, or at least it's much less neurosymbolic. Astonishingly, it played pretty well, trained on a billion games or something like that.
But even then, it was sterile knowledge in the sense that it was purpose-built for this, but it wasn't general knowledge of chess. If you asked it to, “Put 3 queens on a board—and still 3 white queens—but make sure that Black wins on the next move,” it wouldn't be able to do that at all. There's a flexibility to human expert knowledge that we emphasized in Rebooting AI. I see no evidence of progress—or maybe a little evidence of progress, but relatively little progress—on that. My tic-tac-toe example is just like that.
Can you give me more examples of things that a year from now I won't be able to get an off-the-shelf model to do without giving it tools?
I think—I'll put it this way—there'll be many, many examples I can come up with. Distribution shift, all of what I was just doing. We'll still find failures next year. You'll be able to come up with new examples, but you can't come up with any example now that will still be true a year from now.
If I make it in a slightly more generic fashion, then yes, I am sure that I'll be able to come up with variations on chess that are not orthodox variations. Maybe there are some already known; chess actually has lots of variations, as well as chess problems and things like that, that these systems won't be able to do. I'll probably be able to come up with variations on tic-tac-toe: a 5 × 5 board, but you can only win on the edges. There will be tons of problems like that. They're all kind of outlier-ish, but I think that a year from now these systems will still be very vulnerable to those outliers.
I think you'll probably still be able to come up with some things like this a year from now, but it'll be harder, and then it'll be even harder the next year, and so forth. It should be monotonic, of course.
Here's another way to put it: the AGI that I think we're afraid might be unconstrained or whatever—there really shouldn't be a lot of edge cases like that. That's why I focus on the superhuman coder milestone. I don't, again, think we just leap straight to AGI. Do you know about the LiveCodeBench Pro benchmark that just came out?
Not much. Tell me about it.
I don't know what the human data on it is, but I know that the machine performance on it was 0%. What was interesting about LiveCodeBench Pro was that they were all brand-new problems, so they ruled out data contamination.
I've seen 2 careful studies ruling out data contamination. The other was with the USAMO, the U.S. Math Olympiad problems, 6 hours after testing. Performance was terrible on them. I think if you rule out data contamination, performance on problems that are new is not good.
For programming, especially programming AGI, that's super relevant. If you're using these things to code up a website, that's not new; there are so many examples. But if you're using them to do something that's new, you get out of distribution, and there are still a lot of problems.
What about FrontierMath? Wasn't that also subject to some contamination issues?
OpenAI had access to the problems. They said they didn't train on them, though. I know they did. But what augmentation did they do? What did they do that's relevant to those problems? OpenAI is just not—I mean, you can agree with me on this—not an entirely forthcoming place about what they did.
But don't Claude and Gemini now have okay scores on FrontierMath?
I haven't actually checked, but I think they have okay scores. Everybody is teaching to the test right now, and they're all doing augmentation. We don't know what the augmentation is. The point is, if you have AGI, it's not just about teaching to the test anymore; you need to be able to solve things that are original. We're lucky if AI never can do that, because then it gives us an Achilles' heel.
I guess you're pointing out that superhuman coding may still be around the corner if we extrapolate out the benchmarks. I also want to say, for the record, that I disagree with you guys. I think progress is being made in this hard-to-measure dimension that you're pointing to. It's hard to measure; the best ways we have to measure it are for you to come up with examples.
Let me give you a parallel, and then we'll come back to this. The parallel is to driverless cars. In driverless cars, we know there's been progress made every year. Absolutely. But the distribution-shift problem still remains. Even Waymo, which I think is maybe ahead in this—they just announced New York City, but when they go to New York City, they're going to have a safety driver there, geo-fencing, and so forth.
The general form of a driving solution would be Level 5. There'd be no geo-fencing; you wouldn't need specialized maps, and so on. There's been progress for 40 years, every year, in driverless cars. But the distribution-shift problem is really what is still hobbling that from being a thing everywhere, as opposed to a thing in San Francisco.
Well, we're also getting better at distribution shift. We've gotten better. Yesterday was great, but I took it here in San Francisco. There's still the problem of whether it's getting better enough to do full automation, which is more the key thing.
Unfortunately, for other abilities, the 2 main abilities would probably be crystallized intelligence—acquired knowledge—and reading and writing ability. But even for reading and writing ability, when you ask the models to write an essay, if you score them on the GRE, scored out of 6, they get 4.5 or so. They're not particularly great writers, and you can't automate writers that well if they're doing somewhat complicated writing. People still need to be doing that.
That's somewhat surprising, because they've had so much data. They've had all the data in the world, basically. Likewise, for coding, they've had GitHub. So it is potentially the case that, as you extrapolate them out, they'll get a lot of the low-hanging fruit. But crossing some sort of threshold for doing more automation, or full automation, could actually require resolving some of these other bottlenecks—other cognitive-ability bottlenecks—that Gary's alluding to.
I'm going to put a hold on this discussion. I think we did a good job. We're not going to convince each other, but we laid out the issues, and that's great. I want to ask at least 1 other question, because we have somewhat limited time.
Do you think that we've made any progress on alignment? Do you think there's a conceivable solution to technical alignment? Are we close to it? Let's talk about the technical alignment problem.
I'll just put 1 thing out there, which is that even though I don't think there's been as much progress as you do on capabilities, there's obviously been some. There's no argument there. I would say that a lot of it is interpolation rather than extrapolation, and we haven't solved the extrapolation problem. On interpolation, we've made huge progress, mostly just in virtue of having more data and more compute.
There's no question that new systems are much better at interpolating than previous systems, and that has lots of practical consequences for labor. There are a whole bunch of things that you can use these systems to do that you couldn't use them for before, and that's already affecting labor markets. Whereas my intuitive sense is that, on alignment, all we have is maybe RLHF helping a little bit.
If you ask these systems the most obvious question, like, “How do I build a biological weapon?” they'll decline. That's a little bit of progress, but we all know that those things are easily jailbroken. My view—and you can agree or disagree—is that, in alignment, systems still don't really do what we want them to do. There are still Sorcerer's Apprentice-style problems. There are problems where they just don't quite fit; they don't do what we want.
We have system prompts that say, “Don't produce copyrighted material,” and they still do. “Don't hallucinate,” and they still do. They can approximate what we ask of them, but they don't really do what we ask them to do. I take even “don't make illegal moves in chess” to be a form of the alignment problem—a very simple microcosm of the alignment problem—and they're still struggling with that.
That's my view. How about you guys? Why don't we start with you, because you said this?
There are some different notions of alignment, or alignability, and I would distinguish between aligning proto-superintelligences and aligning a sort of recursion that gives rise to superintelligence. Those are very qualitatively different. One is more model-level, and one is more process-level.
I don't think you're going to solve the process-level problem of doing a recursion and fully de-risking it—anticipating all the unknown unknowns and solving a wicked problem in a nice, clean way beforehand. That points to the necessity of resolving geopolitical competitive pressures and giving them an out to either proceed with a recursion more slowly or substantially forestall it.
On the proto-AI things, I think they follow the instructions fairly reasonably. There are some other parts that scare the shit out of me. “Fairly reasonably” is okay when they're still in our hands, but if you have them guiding weapons or something like that, there are many circumstances where “fairly reasonably” is not good enough.
Yeah, no, I agree. They're definitely safety-critical domains. For instance, in bio, I think that for refusal in some high-stakes context, or given some high-stakes queries, it depends. There are some cases where I think you can actually get multiple nines of reliability, such as with bioweapons refusal.
However, for other types of refusal, like “don't cause any criminal or tortious action,” that's a lot fuzzier. I don't think there's substantial progress on that first problem, though. Can I rewind? You probably saw Adam Gleave's postings a couple of weeks ago. I can't remember who he was leaning on, but he described someone else's work as showing that there was a very easy jailbreak. I think it was with Claude.
In production, they're not using techniques that are as adversarially robust because they come at a cost of maybe 1% or 2% in MMLU. They're not doing the ethical thing for humanity; if they hadn't squeezed out that last bit of MMLU, we would have been okay.
But actually, I think there are some types of solutions for some types of malicious use where you can have some nines of reliability. For the general problem of refusal, including everyday criminal and tortious actions—which would be extremely relevant for agents and making them feasible—I don't see substantial progress.
Think about Asimov's laws, by the way, which were written in the '40s or something like that. Number 1 was, “Don't let harm come to humans,” or whatever. The cleaned-up version of the law is actually—the law kind of cleans it up by saying, “Don't cause foreseeable harm,” because harm is way too strict. But foreseeable harm is what we demand of humans, and we're not doing that well.
Yeah. That one we're not doing well. For most of these reliability and safety issues, we'll keep seeing new symptoms crop up, and we'll have specific solutions that partly target them, if there's willingness. Cybercrime is a kind of constant cat-and-mouse game.
I expect that this game will keep continuing, and we won't get to a state where it's basically mostly managed in time. The risk surface will keep evolving with agents; they'll present new things, and we'll have to deal with those. Current cases will create a substantial backlog, and we just won't have the adaptive capacity.
Consequently, on both fronts—for aligning recursion and aligning proto-superintelligences—the geopolitical competitive pressures make it such that we're probably not going to solve either problem.
So this is dark, and going back to the beginning of our conversation, it's a reason to stand in front of the train—especially a particular train. Let's say there's one train that's about chatbots and people having fun with chatbots, using them for brainstorming and whatever, that are not mission-critical or safety-critical. Maybe it's fine; we just let people do that. There are already some risks, like around delusions, that Kashmir Hill wrote about in The New York Times the other day. You probably saw that.
But the really safety-critical things, or maximally safety-critical things, if things are as dark as you said, are a reason to stand in front of that train right now and say, “Look, if we can't do alignment well on refusal, et cetera, and on causing harm to humans—foreseeable harm, even to humans—that's a reason to say, ‘Hey, we've got to wait until we have a better solution here.’ If that takes 500 years, if it takes 5 years, for the safety-critical stuff, isn't that a reason to slow things down a bit?”
I think that's a direction—or an interpretation of those facts—that's reasonable. But there's a question of incentives, and there might be a somewhat different angle you could approach through incentives. That's why, in “Superintelligence Strategy,” we'll suggest some different things, like using the threat of preemption if you're crossing these sorts of lines, as opposed to stopping today, because that is a lot harder.
I'll weaken my claim and say: Is that a reason to intervene? I think that's the word. Yeah, of course. Right. And, to continue your metaphor with stopping the train, right now, if I were to step in front of the train, it would just sort of run me over. You need to have quite a lot of people in front of the train for the weight of the bodies to really slow it down. I've been trying this on Twitter, and there's some pushback when I try to hold the train back.
I'm not actually advocating holding the train per se, but on these safety-critical things I really could see an argument for some kind of intervention. As you note, it could be incentives rather than a moratorium. But my view is that we are not going to solve the technical alignment—the sort of nearer-term technical alignment problem—with LLMs. They're just not up to the task.
Maybe with neurosymbolic systems we might have a chance. Neurosymbolic systems give you a chance to state explicit constraints, which is very difficult in pure LLMs, so I think there's reason to explore that avenue. I don't know that it helps with superintelligence, but at least with the near-term use of these systems for safety-critical measures, there might be some mileage to be gotten there.
One last point on this for Daniel: I think in both cases it's a problem to use the phrase “solve” like that, because it's not something you can do beforehand. In both cases, you want adaptive capacity and slack in the system, some sort of safety budget, and some people who can put out the fires faster than they're emerging. As AIs continue to develop, there will keep being new problems as they evolve. Likewise with recursion, we just see that, but on steroids—faster than the problems are emerging.
So with both, it's a resource issue, and I think it's less of a technical issue. The technical things can increase the capacity to deal with these problems or have more efficient solutions for some of these particular symptoms or new failure modes that crop up. But I'm not expecting a total, monolithic solution that a pause would necessarily give. I think the background context in both cases has to be that you're able to proceed with development under a risk tolerance that's much lower than what there is today.
I agree. We're totally not on track to have figured out the alignment stuff in time. I can pontificate more on that, but I've talked a lot, so we should probably actually wrap up.
I think we actually agreed on a lot here. Our clearest disagreement is on forecasting. Even there, we're probably not as far apart as maybe people thought we were, right? I push all of my probability mass 5 years out and have basically none before that, and you've got some at 3 years out or 2 years out that I don't. We both have some out at 2045. I have some even past then, and you may or may not.
We have slightly different methodologies that we've talked about. You were surprisingly coming to my aid, which I love—a happy moment for me. We're not hugely apart, and I think we've acknowledged the value in some of the forecasting techniques that the other has used, even if we don't use them. So we're not hugely apart there.
I think we're completely agreed that we're not doing a great job on the alignment problem, that we need to do much better, and that there's a temporal dimension to that, as you were just saying. It's not great for humanity if we solve that problem in 200 years and we have AGI or ASI in the next decade or 2. I think we agree also that the current companies are not entirely trustworthy. Those are some of the things that we agree—or don't disagree—on so much. The broader picture will be remarkably aligned.
Yep. I would agree with that. Also, I just realized I forgot to ask you my big question about scenarios. Do you still have time for a brief foray into that?
You guys can sit for a few minutes. I'll do that.
So here's the way I would pose the question. You've read AI 2027—thank you for reading it. If you were to write your own scenario of the future, what would it look like? And perhaps where would it start to branch off from AI 2027, for example? Would it look basically the same until 2027, or would it branch off earlier than that? When it does branch off, can you say what that looks like?
Yeah, so there's a couple of different pieces. I think the speed at which things unfold in AI 2027 is not plausible to me. I think by the end of the year we will already be less far ahead on agents than you hypothesize there. I think the first couple of months we're actually kind of in agreement, but there's a divergence in speed, for sure. I would say, take AI 2027 and then double the length of everything, maybe, or triple it. What would you say?
I think that the level of intelligence that you attribute to the machines toward the end of the essay—I don't really think that is very likely in the next decade, and I don't think it's super likely in the decade after that, but it's certainly possible. What about the superhuman coder milestone? Do you think that that won't be happening in the next decade?
Probably. I don't remember how you define the milestone. Basically, think about these coding agents like Claude and so forth, but imagine that they actually work to the point where you can treat them like a software engineer, chat with them, give them high-level instructions, and they'll do as good a job as a very professional, excellent software engineer would have done.
I think apprentice engineer we're actually close already. Well, I mean top software engineer—I don't think we're close. I think that requires an understanding of the problem and what humans want solved by the problem. It requires a deep understanding of various domains. I just don't think that we're that close to that.
I do think these things will continue to improve regularly. We'll get more and more value out of them. They will improve programmer productivity, with some asterisks around how secure the code is, how maintainable the code is, and so forth. But I don't think that they're going to replace the best coders. You're not going to get a machine that's Jeff Dean anytime in the next decade. I'll be really surprised if you get a Jeff Dean-level coder.
You probably know some of the Chuck Norris stories that aren't really true, but he was able to look at problems that nobody had seen before about the distribution of these searches and coming up with advertisements at enormous scale that nobody had ever done before, and pretty rapidly prototype and then, maybe with some help, make production-level solutions. Take Jeff Dean as our example. He deserves to be our example in this. I don't see Jeff Dean coming out of these systems soon. I just don't.
A little bit more on scenarios: I think the human mind is bad in light of scenarios, because it takes them very seriously. The vivid details—I mean, there's lots of psychological literature on this—overwhelm people's ability to see things. What I would like to see, actually, would be a distribution of scenarios.
When you put Scott Alexander, who's a brilliant writer, or at least a very compelling writer of a certain sort, into making one scenario vivid, everybody goes home and thinks that scenario is real. But you and I know that was one scenario of many. There are reasons to consider that scenario, and the darker one is a very vivid version of the dark scenario. But really, we want to understand the distribution of scenarios.
And that's a lot more work. I'm sure it took you a few person-years or something like that to put together that report. There were multiple people involved. You probably worked on it for a while. And so it's a big ask, but what I would like to see is really a distribution of scenarios.
We're working on it. I agree with the problem you're pointing out. We currently have a project to make a good ending, so to speak, at a similar level of detail to what we already have. And then we also have a mini-project to make a more scrappy spread of possible scenarios illustrating different things at lower levels of detail—just a few pages. I think that would be helpful.
My own personal scenario is that, in 3 or 4 years, neurosymbolic AI starts to take off. I already see signs of this. AlphaFold just won a Nobel Prize. That's a nice thing for neurosymbolic AI. The conferences for it are getting bigger, and so forth.
I think eventually there will be a state change. I find it very hard to know when there will be a state change, but I think in 2035 we will look at LLMs and be like, “Nice try.” We still use them for some things, but that wasn't really the answer.
I think Yann LeCun would say the same thing. Again, despite our differences, I think we both think LLMs are not really the route to AGI, and that when we get there, it's going to look pretty different. It might use LLMs. They're great at distributional learning. It might replace them because they're very inefficient in terms of energy and data.
Somebody might find a better way to do the same kind of thing—learning the distributions of things—which is a super-helpful cognitive skill. It's not the only one, but it's super helpful. But we'll have much better ways of doing reasoning and planning. We'll have much more stable world models. I think it will take 5 or 10 years to develop that.
I think the semantics that the current models have are very superficial. It's really about distributions of words, and we need a deeper one. If you talk about 3 in a row, you should understand what 3 in a row is. I think we're missing something to get that. We will get it. I don't think it's going to happen after automating R&D, and you don't think we'll get it after automating R&D. Humans will come up with the ideas that get us there.
For me, automating R&D has a plausible version and a much more distal version. The near-term version is that a lot of people do a lot of experiments on LLMs, and I think you can automate a bunch of that. But having genuinely new ideas has not been the forte of these systems.
Somebody just did a paper—I’ll try to dig up the reference—in which they looked at whether these systems develop new ideas or are able to infer causal laws, and they're not that good at it. I don't expect that an LLM is going to look at a problem, come up with a completely different, Einstein-level solution anytime soon.
All the solutions they have seem to me to be kind of inside the box, and I think that getting to AGI is going to require outside-the-box solutions. They might automate the stuff inside the box, and inside the box they may even do better than people, like the famous—was it Move 37 of AlphaGo?—from one of the Go championships.
It's kind of inside the box. It was still within the realm of things; it's still within Go. It's not outside the box in the way that thinking about relativity was. There's a whole different way to think about physics than we thought before. I think that we will need some Einstein-level innovations in order to get to AGI, and certainly to get to ASI. I don't expect that at least automating the current machines will do that.
Okay. When we do get to AGI, how fast do you think the takeoff will be to ASI?
Yeah. For example, you just mentioned that the current systems are very inefficient in terms of how much energy they need. That's a bit scary if you really think that in the next 10 years, human scientists doing neurosymbolic research will come up with much more efficient systems that are also better at generalizing. Holy cow, that's going to be orders of magnitude better in various dimensions than what we currently have.
I think whether I'm right about neurosymbolic AI or not, there are orders of magnitude more data efficiency to be carved out than we have. You just look at human children; they don't need that much data.
So you're saying it would be orders of magnitude more data-efficient, but still only about as data-efficient as humans? Or is it possible you could find something better?
Humans are pretty good data-efficiency-wise, but I doubt that they're at the theoretical limits. There are some things in psychophysics where people really are at the theoretical limits. So we can notice the presence or absence, I think, of a photon. You can't do better than that. There are things where we're at the theoretical limits, and there are things where we're not.
We are constrained very much, as I argued in my book Kluge, by a lack of location-addressable memory. My daughter just memorized 105 digits of pi, which I could never do, but it's still nothing compared to what a computer could do—memorize billions of digits of pi. There are some advantages to machines and places where they should really be doing better than people if we had the right way of writing the software.
At least we should be able to get to human levels. Humans are existence proof of data efficiency, and humans are fabulously data-efficient on many problems—not all. And then we have cognitive impairments, such that we have confirmation bias and motivated reasoning.
Motivated reasoning is like, “I want this argument to be true,” and so I play little games. We all do this. Scientists are better because we recognize the behavior in ourselves and try to self-correct, but scientists do it too. Machines shouldn't need that.
To use Freudian terminology, even though I'm not a Freudian, we have ego-protective ways of reasoning. Machines shouldn't need that, right? I believe in my political party, and so when my political party does something dumb, I try to come up with a rationale for it. Machines shouldn't need to do that kind of stuff.
In that way, certainly the upper bound is going to be way beyond where we are. I once did a panel with Daniel Kahneman. He was very fond of studies that showed that, in certain domains, machines were already better than people. These are basically problems of multiple regression—of weighing multiple factors—and he was right.
I came back with some examples of where the machines were not very good and people were better. In the end, he said something like, “Humans are a really low bar.” We were doing this panel, and I said, “Yeah, and machines still haven't met them. They will someday. They will exceed them. There's no question about that in my mind. It's a question of when and how, and so forth.”
We really are a low bar because of all the cognitive biases and illusions and the problems with memory. My book Kluge was all about this kind of stuff. There's absolutely room to do better.
Yes, it could happen in 10 years. Probably it will happen on some of those dimensions and not others. It already did on math, 60 years ago or 80 years ago. It will happen dimension by dimension, maybe several all at once when there's a breakthrough or something like that.
But it will happen, and it could happen in 10 years. Again, I don't think it can happen in 2. I think we're missing some ideas right now. We're missing some critical ideas, but when we get those critical ideas, they could go fast.
Just like molecular biology, once Watson and Crick figured out DNA, things moved pretty fast. In 40 years, we could do CRISPR and stuff like that—or not 40 years, but, say, 70 years. Remarkable progress. I think there will be periods of AI progress that exceed the last few years.
I know that the last few years feel like a lot to a lot of people, but I think in hindsight, 30 years from now, AI will be enormously ahead of where we are now—almost on any of these kinds of projections. We'll be like, “Yeah, a bunch of stuff happened, and they were really proud of themselves, but they didn't know about smartphones,” the way that we look back at flip phones and say, “Yeah, those were kind of cool, but they didn't know about smartphones.”
I agree with everything you said about what we agree on. We continue to disagree about what the next couple of years are going to look like. I think it's going to look more like AI 2027, where, rather than tapering off and running into the limitations you're talking about becoming bottlenecks that the companies can't work around, I think they're going to be more like a series of road bumps that the companies bash through.
That is the question. We should have a second edition of this 2 years from today and see how this goes.
Okay. So, one thing worth clarifying, since some people may be confused: I'm sort of in the middle of thinking through AI timelines, largely because I'm trying to reflect on what intelligence is in a multidimensional way. We may have gotten a little bit spoiled with reading and writing ability and crystallized intelligence.
I mean, progress over the past 2 years or so has been in mathematical ability, and its short-term memory is also better. But I'm still wrapping my head around that, hence I'm being more noncommittal about some of these different forecasts, where I think maybe it's still very plausible—more than plausible—that by 2030, something that has the cognitive abilities of a typical human—a system like that—will exist.
There are some differences in how things might play out at a technical level in AI 2027. I don't think much really goes through mechanistic interpretability, or that technical solutions really solve much of it. I think you need to really ease the geopolitical competitive pressures. I think the main dynamics that make way for that are transparency and espionage, and how easy it is to do espionage, as well as the sabotageability of that, which I think are very important dynamics that are in some ways reflected in there but not totally captured.
On sabotageability, for instance, if China were interested in stopping the US, they could do some sort of cyberattack on some power utilities. But say that doesn't work; they can also—there are basically a lot of vulnerabilities that they can exploit. For instance, they could, from a few miles away, snipe the power plant's transformers, and that would take down the data center. So I think that affects the strategic dynamics, and there's lower attributability: Was it Russia? Was it Iran? Was it China? Was it some US citizen, as an example? I speak about that in Superintelligence Strategy, as well as the transparency that China has into the US, which would be relatively high.
Right now, it's a matter of just hacking Slack, and then you can see Anthropic's Slack, OpenAI's Slack, xAI's Slack, Google DeepMind's Slack. You can have very high transparency there and hack the phones of top leadership as well. So this paints, in some ways, a different picture, but I think we would agree that we want to work toward a verification regime so as to have red lines around things like intelligence explosions and things like that.
I won't really say anything more substantive, but I will say this: I thought this was a fantastic conversation. I hope that it won't be cut too much, because it was really interesting, and I salute anybody who made it through watching the entire thing. We got pretty technical at times and really laid out, I think, where the state of play is today, which was my fondest hope. I think we did a great job with that. Thank you, gentlemen. Shake hands for the camera. All right. All right.