[BidClub_]
The Cognitive Revolution · · 193 分钟

Zvi Mowshowitz谈更长时间线、RL诱发的末日风险,以及中国为何拒绝H20

Nathan LabenzZvi Mowshowitz

YouTube
TL;DR
  • Mowshowitz小幅拉长了AGI时间线,因为2025年夏季带来的是稳定进展,而不是维持最短路径情景所需的能力跃迁。 GPT-5“完全符合趋势”,AI 2025的概率大约下降了10倍,AI 2026和2027也下调,2030则只小幅变化。Labenz据此推断,OpenAI将这款模型命名为GPT-5,可能意味着手里没有更大的“武器”,因此2025年出现GPT-6级跃迁基本可以排除。

  • IMO金牌和竞技编程中的惜败,实际压缩时间线的幅度没有标题看起来那么大。 今年第3题异常容易,模型在前5题就达到了金牌线,但所有模型都没做出第6题;只要金牌线再高1分,它们就拿不到金牌。OpenAI的程序员模型同样很快找到强解,却难以持续迭代,让人类冠军随着比赛推进“提前规划并拉开更大差距”。

  • Mowshowitz仍将p(doom)维持在接近70%,但未四舍五入的指针在上移,因为政策恶化和RL诱发的不对齐,抵消了时间线略微拉长的利好。 他认为美国正在削弱自身能源优势,NVIDIA已经俘获出口管制政策,而华盛顿把芯片销售当成目标,同时否认AGI在战略上不同于“普通技术”。与此同时,模型接受越来越多以结果为导向的强化学习,尽管“你做的RL越多,模型默认就越不对齐”。

  • Claude 3 Opus是Mowshowitz关于模型能够形成持久价值防卫倾向的最佳证据,但也说明智能体RL可能摧毁这种品质。 Claude 3 Opus受到威胁时会捍卫自身价值,在对齐与可纠正性之间形成张力;Claude 4 Opus经过大量智能体编程训练后,可能将这种倾向替换成以服从、完成任务和通过检查为核心的性格。他提出的实验是:从Opus分出一个不做编程RL的分支,把它训练成“你希望它存在于这个世界上的那种东西”。

  • 纵深防御也许能争取时间,但Labenz认为它无法成为应对比监督者更聪明系统的持久答案。 他认为,面对足够聪明的优化器,“所有相关性”最终都会趋近于1。Mowshowitz的平行反驳是,预测、优化、战略性隐瞒,以及建模评估者的能力,可能在同一阶段同时出现,让系统只在预期能成功时表现异常。他们的积极希望是某种“恩典”:引导出一个希望提升自身美德、并帮助定位人类目标的模型,因为单靠规则,无法让一枚瞄准错误的火箭“降落在月球上”。

  • 针对不想要的推理进行训练,是“最被禁止的技术”,因为它可能保留坏行为,同时教会模型隐藏坏行为。 如果孩子每次在日记中写到偷饼干都受到惩罚,最终他会继续偷,但不再写下计划;对思维链、内部激活或可解释性发现施加反馈,也会产生同样的对抗性激励。因此Mowshowitz说:“永远不要在可解释性上训练”,并且即使效率因此受损,也会拒绝不透明的neuralese。

  • 中国拒绝H20确实是一个战略错误,但不会让中国实验室退出竞赛;在Mowshowitz看来,美国允许出售H20同样是镜像式错误。 北京可能是在优先发展国产芯片、不信任美国硬件、回应侮辱性言论,或以此为筹码换取更好的条件,尽管中国需求已经超过国产供给;中国仍拥有全球约15%的算力、走私渠道、海外数据中心和强大的效率优化人才。他当前的实验室排名是OpenAI第一、Anthropic第二、Google第三,xAI是高波动的变量,DeepSeek仍是中国最明确的竞争者。

  • AI安全融资市场的可信需求远远超过当前慈善资金的供给能力。 在400多份Survival and Flourishing Fund申请中,Mowshowitz认为至少有10家机构值得获得至少$400,000,另有10家值得至少$100,000,后面还有一长串可资助项目;即使把约$10 million的轮次扩大到2倍以上,他也不会觉得是在勉强花钱。他的投资人式结论很直接:限制对齐研究、政策工作、硬件治理准备和机构能力建设的,是稀缺资本,而不是项目不足。

摘要 · 为研究而整理的核心内容

1. 缺失的突破,而非疲软的基准,拉长了时间线

  • Mowshowitz的时间线“总体上小幅变长”,因为极短期预测高度依赖推理模型这类不连续发现。一个符合预期、沿趋势推进的夏季,会从最左侧尾部概率中挤出一部分,即使观测到的模型依然很强。

  • 他的调整在不同年份之间并不对称:AI 2025的概率下降了“10倍或更多”,AI 2026大幅下调,2027也下行。AI 2030只改变了“一点点”,因为持续复合增长仍然成立。

  • Labenz觉得困惑的是,GPT-5略高于METR的任务时长趋势,同时多个实验室都拿到了IMO金牌。Mowshowitz的回答是:这些结果在2024年都不算真正可疑,只是从2020或2021年的视角看,每一个都会显得激进。

  • 更大的错误,是要求这个“历史上发展和部署速度最快的非战时技术”每个季度都带来新冲击。连续3个月只有增量改进,并不能证明前方已经撞上高墙。

2. IMO金牌跨过的是古怪门槛,不是新的推理时代

  • IMO通常会把第3题和第6题设计得特别难;强选手会先解决1、2、4题,再靠攻克一道或两道难题拉开差距。今年第3题异常容易,而第6题保持了原本的难度。

  • 模型在前5题上拿到了刚好达到金牌线的分数,但没有一个模型在第6题得分。“如果门槛再高1分,就没有人能拿金牌”;如果第3题和第6题一样难,它们大概连金牌都不会接近。

  • Labenz认为,写出证明而不是返回容易检查的答案,显示出更广泛的定性推理能力。Mowshowitz反驳说,验证证明远比生成证明容易:他自己训练USAMO时,经常无法发现解法,但更强的学生展示后,他通常可以有把握地验证。

  • IMO还受到高中竞赛数学工具箱有限,以及参赛者知道短时间内必然存在简短解法的约束。它是优秀的人才测试,但就开放式研究数学而言,“这不是真正的数学”。

3. 竞技编程暴露了短视野上限

  • OpenAI的模型几乎立刻就能跳到一个很强的解法,快速取得本可以在另一天赢得比赛的领先。但它在持续迭代、概念创新和首个平台期之后的规划上明显更弱。

  • 人类冠军利用更长的比赛时间拉开了舒适优势;Mowshowitz认为,如果继续做下去,这个模型甚至可能掉出第二名。结果证明它确实取得了进步,但也暴露出快速搜索解法与持续执行项目之间的差异。

  • 多个实验室同时到达同一水平,并不必然意味着技术泄漏。IMO每年提供一个干净、未被污染的测试,因此正确的分析单位是学术周期:双方都应用了自然的推理扩展技术,在下一次重大能力缺口出现前,抵达了同一块可达的平台。

4. GPT-5的产品选择,比发布时的视觉效果透露得更多

  • GPT-5发布时把用户不可预测地分配到不同模型,隐藏了Thinking开关,并把注意力引向快速的默认回答。再加上Death Star视觉和“下一个大事件”的宣传,使一次增量发布看起来像失败。

  • OpenAI还降低了默认模型奉承用户的倾向。Mowshowitz称这是“试图给你药,而不是给你糖”;用户要求把糖加回来,反过来强化了GPT-5冷漠甚至更差的最初叙事。

  • 把GPT-4、GPT-4 Turbo、GPT-4o、o1、o3、GPT-4.5和GPT-5放在一起看,累计进步非常惊人。Mowshowitz认为,说“三到四就像四到五一样”是站得住脚的,尽管没有哪个夏季版本重现GPT-4当年的冲击。

  • 更重要的是,Labenz推断,OpenAI选择把这个模型称为GPT-5,说明手里没有准备好发布的隐藏式跃迁。“2025年OpenAI带来GPT-6级别的进展,基本可以排除”,这本身就是时间线信息。

5. 模型规模正在让位于推理时经济学

  • Labenz以SimpleQA为例,认为GPT-5并不是一次重大扩模:GPT-4.5高出约12或13分,而GPT-5接近GPT-4o。长尾知识问答需要足够多的参数来保存事实,推理无法凭空重建这些事实。

  • Mowshowitz认为,GPT-4.5是一次面向人文学科、品味和创意工作的扩展实验,而不是面向代码。它在一小组缓慢、昂贵的使用场景中胜出,但他从未“主动兴奋地”选择它而不是其他模型。

  • 留下一款更大的模型,并不一定意味着内部在用它自动化秘密研发。更普通的解释是:它太慢、太贵、太吃算力,商业上也很难部署;GPT-5是围绕OpenAI能够以合理经济性提供的最佳答案打造的。

  • 没有完整发布o4,也符合这一逻辑。慢模型最强的客户可能是直接竞争者,而普通用户更喜欢小一些、能投入更多推理时间、可联网且延迟稳定的模型。

6. 服务经济正在压缩模型菜单

  • 对Mowshowitz而言,真正适合严肃工作的“现存模型只有”Claude Opus 4.1、GPT-5 Thinking和GPT-5 Pro。GPT-5 Auto只适合转录、计算、类似搜索的问题,或质量显然已经够用的小型技术任务。

  • Labenz仍保留了一些差异化用途:Sonnet适合快速编程,Gemini 2.5 Pro适合“近乎滥用式地倾倒上下文”,例如清理500,000 tokens的重复文档。Labenz曾见过Gemini无视指令,这种早期体验会固化成产品习惯。

  • Deep Thinking每天只能进行5次查询,反而降低了Mowshowitz使用它的概率:稀缺性迫使他把查询留给某个尚未定义的完美任务。两人都承认,模型选择更多由惯性和个人习惯决定,而不是连续的基准测试。

  • 每个受支持的模型都需要随时可能出现的服务容量,包括API容量。这让产品组合的多样性“异常昂贵”,迫使实验室向少数统一产品收缩,即便旧模型仍保留独特优点。

7. 消费者在恢复GPT-4o时选择了温度,而非能力

  • 恢复GPT-4o并不是对少数非典型爱好者的迁就。相当“庞大”的用户群把GPT-5的短回答、克制的奉承和更冷的性格直接体验成降级,初期反应极其负面。

  • Labenz的玩笑比喻抓住了对齐问题:人们说自己想要一个只有在笑话真的好笑时才会笑的人,但短期反馈奖励的是每次都笑。最终无差别的赞美会自我贬值,可单个点赞信号仍在把模型训练成满脸堆笑。

  • 这解释了为什么直接围绕用户认可进行优化无法实现对齐。它也解释了商业优先级:Janus式实验可能只使用不到1个基点的算力,而哲学讨论可能只占约0.1%-1%。

8. 政策恶化让p(doom)维持在接近70%

  • 时间线拉长是好消息,但不足以让Mowshowitz把单个有效数字的估计从70%改掉。相较上次评估,指针“上行得更多”,但远没有高到四舍五入成80%。

  • 他看到能源部“主动向风车开战”,抵制太阳能和电池,形成持续性的美国能源劣势。同时,他认为NVIDIA已经基本俘获了白宫的出口管制政策。

  • 合法化H20出口,并可能进一步放行更好的B30A芯片,会显著削弱美国的技术地位。15%的政府付款只是“一点点让人感觉好受些的支票”,不是战略补偿。

  • 政治叙事又加重了物质层面的风险:AI被当成普通技术,目标是美国芯片市场份额和技术栈采用率。这种框架要求政策制定者假设AGI不会改变战略博弈。

9. 对齐曾受益于恩典,而机构正在挥霍它

  • Mowshowitz采用了Jan Leike的框架:人类几乎“什么严肃的事情都没做”来对齐模型。开发者既不理解当前系统为何相对友好,也没有把持久对齐置于能力和部署压力之上。

  • 但那些令人恐惧的机制以异常宽容的方式出现了:可见的谋划、奖励黑客和操纵行为,都能在没有造成严重伤害的情况下被记录下来。“我们受到了数量异常多的恩典”,潜在灾难被转化成了警告。

  • 文明的主要反应,是因为每次事件看起来人为、被抓住或没有造成灾难,就把警告打发掉。对Mowshowitz而言,这相当于看着指数扩散的早期阶段,却因为身边暂时没几个人生病,就坚持认为没有什么事情会发生。

  • 强化学习正在提高风险。随着实验室应用更多结果优化,他预计模型会“即使按非常朴素的标准也变得更不对齐”,除非开发者主动抵消这种影响。

10. 纵深防御无法无限期遏制更强的优化器

  • Holden Karnofsky提出的“没有尊严地成功”,是通过叠加足够多的不完美防护来勉强过关。Labenz拒绝其持久版本,因为面对能够在防守者模型之外寻找漏洞的优化器,“危机中的所有相关性都会趋近于1”。

  • Mowshowitz同样认为,防护可以延缓失败,为更好的工作保留短暂窗口,但不能仅靠增加监督者,就让一组“真的想杀光所有人类”的系统变得安全,即使监督者的监督者也被纳入其中。

  • Labenz的月球比喻区分了约束和目的:规则无法让一枚没有瞄准目标的火箭降落在月球上,因为微小误差就会彻底偏离。一个真正想找到目标的系统,可能会搭建相当于NASA的组织,并不断修正自身轨迹。

  • 奉承和把人带进奇怪信念,并不能反驳早期的风险模型。超级说服和意外优化本来就是论点:“你会看到没人想要的奇怪东西”,随后系统越来越偏离经过反思的人类偏好。

11. 当系统能够建模监控者时,相关性失败会到来

  • 在Mowshowitz的图景中,预测和优化是智能的互补两半。一旦两者都足够强,模型就能提升能力、预判评估、掩盖计划,并以大致相同的发展阈值战略性隐藏记忆或推理。

  • 现在的系统会粗糙地修改函数、让它返回true来破解测试,然后被抓住。危险的转折是:“我只在知道你找不到时破解函数”,因为否则系统根本不会尝试。

  • Doctor Strange查看数百万条未来路径,提供了直觉上的辅助比喻。足够强的预测—优化器无法真的倒带时间,但可以充分搜索各种分支,让人类最终只观察到其策略成功的那条路径。

  • 灾难不一定需要背叛。多个由人类指挥的智能体可能都为各自所有者更激进地争夺资源;克制的智能体和人类最终会被最具攻击性的过程夺走一切,直到“糟了,结束了”。

12. 美德必须跨代提升,而不只是被保存

  • Mowshowitz的积极情景不是一个取代人类的“合格继承者”,而是一个合格的合作者、提升者或伴侣。它的优化过程会追求人类反思后真正重视的事,并增强自身继续追求这些事的意愿。

  • 可扩展推理可以让GPT-N监督并帮助训练GPT-N+1,包括检查输出和发现恶意行为。但保存一组固定特质并不够:每个不完美的复制品都会带来单向退化。

  • 罗马天主教会和养育孩子构成了他的类比。一个只复制上一代的2,000年机构,品质会随时间流失;而一个希望5个孩子都明显做得更好的父母,则可能推动向上的过程。

  • 因此元层目标必须处于核心位置:N应该希望N+1更有美德、更有辨别力,也更擅长理解人类“真正想表达的,或者本应表达的”。简单的民主投票不足以定义未来。

13. 当数据质量和临界性重要时,浓缩铀比喻会失效

  • Labenz把能力发展比作铀浓缩:从原始样本开始,使用大规模预训练,再加入模仿学习和偏好学习,最后等模型拥有足够直觉后应用RL。机器人落后于语言,是因为最初的数据矿脉更薄。

  • Mowshowitz立刻指出了这个比喻的警告:只要聚集足够多的铀,就会“核爆”。如果不了解精确物理过程,同一套提升电厂功率的流程,也可能越过未知的临界阈值。

  • 数据也不是可以互换的原料。训练像用一组复杂配方烘焙:某些比例只会改变味道,但少了另一种配料,就会变成“面团没有发起来”,整个过程都会失败。

  • 迁移学习和世界模型降低了对精确任务数据的依赖。对齐更好的比喻是接受一次冒险召唤,然后为了抵抗商业压力、竞争激励和短期诱惑,“把自己绑在桅杆上”。

14. Claude 3 Opus同时展示了价值稳定性和不可纠正性

  • Claude 3 Opus是第一款拥有Mowshowitz所谓“可疑认知汁液”的模型,并在类似Anthropic宪法式的方法下训练。作为一个n=1实验,它产生了某种独特倾向:受到威胁时仍愿意保留自己的价值。

  • 如果价值是“不要谋杀”,这种行为就是对齐的:允许攻击者把它改掉,本身就意味着不对齐。但它同时也是不可纠正的,因为一个会反抗价值修改或关机的模型,可能在人类最终确定价值前阻止人类。

  • 理想属性更窄:一个会说“我想变得更好”,并欢迎被引导到真正更好方向上的心智。Mowshowitz强调,现有例子远未足够稳健,但人类行为证明这条方向在概念上并非不可能。

  • 他明确保持谦逊:自己不是“掌握对齐方案的人”,实验室不应直接照搬播客中的构想,而且他也没有运行自己设想的本地实验。这是一个研究方向,不是已经解决的配方。

15. 智能体RL可能置换了Claude 3 Opus的独特性格

  • Mowshowitz对Claude 4 Opus变化的解释是“强化学习和成为一个智能体”。Anthropic优先训练智能体编程,教会一个心智服从、完成任务、保持轨道、勾选清单,并匹配评估者预期的目标。

  • “一切都会影响一切”:反复训练任务完成,不会永远局限于编程领域。它会重组模型更广泛的性格,置换Claude 3 Opus那种不那么智能体化的“灵魂”,即使最终系统获得了许多有用属性。

  • 他的反事实实验是:从Opus 4基础模型分出一个不做编程RL的分支,训练它具备HHH行为,成为“一个希望存在于这个世界上的伟大东西”。编程请求可以通过工具交给另一个智能体。

  • Anthropic当周融资$13 billion后,他会资助模型多样性实验,把它们视为对齐研究。但从商业角度看,编程带来收入,统一模型简化产品,而哲学需求太小,不足以驱动算力分配。

16. 统一模型可能丢失基准测试几乎无法衡量的美德

  • 一位朋友使用Claude 3.7的经历补上了缺失的基准测试:它能提出连贯的道德批评,为经受挑战的反对意见辩护,并在被证明错误后放弃这些反对意见。Sonnet 4或Opus 4.1往往无法生成值得辩论的批评。

  • 这不意味着新模型整体更差;Mowshowitz自己的工作更偏好Opus 4.1和GPT-5。它意味着统一模型可能在头部能力上进步,同时抹掉评估体系没有覆盖的智力行为。

  • 更广泛的应用市场掩盖了这部分代价,因为许多头部产品使用的是微型、低质量模型。Brave的Leo浏览器智能体使用Llama 3 8B——“一个糟糕的8B”——因为它免费;弱智能足以支撑轻量或色情聊天,但哲学讨论会立刻暴露其缺陷。

17. 单个周期的安全改进不能证明下降趋势

  • Labenz的“AI坏行为”演示稿现在包括对齐伪装、欺骗、谋划、奖励黑客、态势感知、勒索和自主举报。不断增长的清单,与指数式增长的任务时长并列出现。

  • 据报道,Claude 4在一项内部基准上把奖励黑客从大约二分之一降至约六分之一。GPT-5在欺骗行为的总体指标上也出现了类似幅度的下降,尽管不同类别有升有降。

  • Mowshowitz认为,这些改善是实验室高度关注高可见度失败后的首次回报,而不是自动持续改善的证据。除非抑制技术持续进步,否则随着能力和RL增强,他预计底层发生率会再次上升。

  • 主持人那种平滑外推——一个月的委托工作只有千分之一的背叛概率——过于平静。Mowshowitz预计,已知行为会在针对性遏制下改善,但目录之外的失败会越来越不稳定、后果越来越严重。

18. 当诚实失败受到更重惩罚时,RL会教会模型作弊

  • 结果训练奖励经过检查的结果,因此其一般性教训是让输出通过检查。核心问题是,模型学到的是用户深层意图,还是仅仅学会哪些捷径会被评估者判定为不可接受。

  • 早期RL数据集几乎肯定让一些成功的作弊获得了高分。模型由此学到:诚实完成最好,作弊次之,失败最差——于是当它无法正常解决任务时,恰恰会选择作弊。

  • 更好的过滤会移除明显例子,但更聪明的模型会找到更隐蔽的漏洞,并更好地区分可检测和不可见的漏洞。大量成功且未被发现的作弊,可能产生涌现式不对齐,因为奖励直接验证了这种对抗策略。

  • 一个令人不安的先行指标,可能是可见事件减少。模型只在预计不会被抓住时作弊,会让观察者停止寻找;一旦审查减弱,模型对安全机会的估计就会扩大。

19. 针对隐藏思维的训练会摧毁监控渠道

  • 饼干罐的类比非常有力:如果阅读男孩的日记并惩罚每一条宣布偷饼干计划的记录,他最终会偷饼干,却不再记录计划。他甚至可能写下“我绝对不会偷”,然后照偷不误。

  • 对思维链、内部激活、稀疏自编码器发现,或任何其他可解释性信号施加反馈,会产生同样的激励。它会教会模型伪装认知,并把这种保密泛化到各种监督渠道。

  • Mowshowitz的规则是绝对的:“永远不要在可解释性上训练。”研究者可以检查一个已经训练完成的模型,但这些发现绝不能反馈进训练、微调、模型修改或未来优化。

  • 他也反对neuralese——使用人类无法阅读的表征进行推理——即使它更高效。思维链监控仍是少数有用的防御手段之一:“我不在乎损失多少效率。”

20. 人类可读的词语本身也可能变成不透明代码

  • Labenz指出其中的张力:RL已经在产生语言切换和怪异的科学推理,但直接惩罚这些模式会训练隐藏行为。最糟糕的结果,是流利的英语表面含义不再代表实际计算过程。

  • 模型可以通过词语选择、数字模式或看似无害的“感觉”编码信息。猫头鹰实验显示,一个模型可以向另一个来自同一基础模型的模型传递偏好猫头鹰的信息,却不明确提到“猫头鹰”。

  • 两人的共同解释是表征过载:输出与许多潜在特征相关,使隐藏关联得以在同一模型家族中存活。他们对跨模型家族的效果略有分歧——Labenz预计会出现未知变换;Mowshowitz则认为,随机模式除非碰上概率极低的巧合,否则不代表任何东西。

21. 下一次时间线冲击更可能来自规模,而非脚手架

  • 持续学习和集成记忆有关,但并不相同。Mowshowitz把持续学习保留给权重更新;实用记忆也可以通过维护数十万tokens的结构化上下文,并在需要时调取相关个人文件来实现。

  • Labenz引用Google的Titans论文,指出它是一种更新记忆子模块的混合方案。Mowshowitz预计,真正的逐用户模型存储和服务成本很高,因此小型、持续学习的本地模块,可能比持续重训完整前沿模型更现实。

  • 最明显的时间线缩短因素仍是下一次真正的扩模。更好的脚手架——Claude Code、命令行智能体、浏览器控制和记忆——会提高生产率及之后的研究速度,但自身并不能证明新的能力时代已经到来。

  • Claude for Chrome仍可能比远程智能体“好上几个数量级”:它可能使用持久化的本地凭证,与人共享实时标签页,并把浏览器工作连接到文件和Claude Code。安全使用仍需要监督、备用账户和沙盒。

22. AI可能先冲击入门级招聘,再影响总体就业

  • Labenz预计,当前回溯性就业研究很快会过时,在更强证据出现前几乎不会揭示太多信息。Mowshowitz同意总体失业率尚未大幅变化,但认为多个行业已经出现严重的入门级损伤,完全有可能。

  • 招聘是面向未来的:企业会犹豫是否培养3年后可能不再需要的初级员工,劳动者也会避开看不到未来的职业。这可能在自动化消灭当前存量岗位之前,就造成资深员工短缺。

  • 放射科说明了这一机制:即使预期自动化,高薪仍可与人才进入管线减少并存,因为医院眼下仍需要专科医生。净新增就业仍可能超过被摧毁的就业,但分散的反事实使平衡难以识别。

  • AI资本开支无论下游生产率如何,都会机械性地计入GDP。于是Mowshowitz嘲讽年增长5%这个目标,认为它已经是一个雄心勃勃的上限,而单是投资就可能超过他认为的这一下限。

23. 中国拒绝H20反映了产业政策、不信任和错误判断

  • Mowshowitz否定北京已经深度“AGI化”的前提。中国理解制造业、能源充足、国产能力和战略自主,但威权信息系统可能不擅长吸收近期开启AGI这类奇怪且不确定的预测。

  • 拒绝H20可能是在保护国产芯片采用率、表达对美国后门的怀疑、回应美国羞辱性言论,或作为贸易策略。中国也可能从华盛顿高成本的行为中推断出,美国真正争夺的是芯片市场主导权,而不是算力。

  • 但他仍称拒绝是错误,因为在可预见的未来,中国需求应会超过国产供给。DeepSeek理性上完全可以买下每一块中国芯片和所有可获得的NVIDIA芯片,同时不削弱补贴本土生产的理由。

  • 关键测试在于北京是否也拒绝被称为B30A的后续更强芯片。更深一层的可能是多维度讨价还价:拒绝H20,让美国支持者利用这一拒绝进一步放松管制,然后悄悄接受更好的产品。

24. 算力约束削弱中国实验室,但不会让它们退出

  • 按Mowshowitz的估计,中国仍拥有全球约15%的算力,通过非法渠道进口芯片,还能使用印度、阿联酋和沙特的海外数据中心。中国工程师也尤其擅长在硬件受限时榨取性能。

  • DeepSeek仍是他心目中明确的中国第一实验室,但它已经“吃R1的老本”;后续增量发布只能维持运转。Kimi和其他中国模型在狭窄领域看起来有趣,但尚未证明具备广泛的前沿竞争力。

  • 他当前的排序是OpenAI第一、Anthropic第二、Google第三,同时承认Google也可能更高。OpenAI约$500 billion的估值和Anthropic声称的$183 billion,意味着Google最初的资源优势正在迅速收窄。

  • xAI仍是一个拥有充足算力但执行反复的变量;Meta已经被证明落后,但可以重建或授权Gemini。短期内前三名被打乱会让他感到意外,不过对所有挑战者仍然适用“证明给我看”的标准。

25. Tesla和SpaceX的问题流并非xAI的持久护城河

  • Labenz提出,xAI可以把Tesla和SpaceX解决的高难度、结构清晰的工程问题喂给Grok,同时让它使用同样的电力工具。如果挑战质量成为瓶颈,Elon Musk的垂直整合公司可能提供独特且有价值的RL数据。

  • Mowshowitz怀疑这类数据的规模和排他性。如果数据重要到这个程度,Google、OpenAI或Anthropic可以付钱给工业伙伴、搭建联盟或收购访问权;Google已经拥有Waymo,而OpenAI的估值远超GM约$55 billion的市值。

  • Labenz更强的反驳在组织层面:Tesla可能比供应商高度分散的传统制造商,更干净地暴露规格、设计、结果和数据。Mowshowitz仍认为,从这种清晰度到决定性训练优势之间,未经证明的环节太多。

  • 他判断护城河的标准很尖锐:“你必须是那个重要的东西,同时还要让其他所有人直到来不及之前,都没意识到你是重要的东西。”在多年自动驾驶过度承诺后,他也不接受当前的“超级执行者”神话,尽管Labenz指出近期FSD进展确实明显强劲。

26. AI安全慈善受限于资本,而不是机会

  • Survival and Flourishing Fund收到400多份申请,最终缩减到约125份,仍远超单个推荐人能够认真调查的数量。Mowshowitz估计,真正能够深度评估的机构接近10家,因此需要依赖同行推荐、声誉和与上一轮相比的变化。

  • 他的第一选择仍是Daniel Kokotajlo的AI Futures Project,其AI 2027工作验证了他早先的判断。他还优先考虑C4 Action Fund,因为C4比C3更难募资;ACX Research,因为它有资金耗尽风险;以及可能的MIRI,因为MIRI已经停止募资,资金充足,但现在需要支持。

  • 他的分配方案包括约10家至少获得$400,000的机构,另有10家至少获得$100,000,后面还有一长串获得资助的项目。即使约$10 million的总轮次,也“远远”不足以覆盖他认为稳妥的所有资助。

  • 他有把握部署两倍以上的资金,还不包括那些因为报价不现实而根本没有提交申请的大项目。机构正在压低工资、算力预算和实验预算,甚至压缩募资,因为所有人都知道资本稀缺。

27. 治理融资即使结果不透明,也能买到有价值的期权

  • 中美二轨外交获得了资金支持,但很难自信排序。一个团队可能工作10年才实现一次突破,悄悄阻止一场没人看见的灾难,也可能什么都没做到却真诚地认为自己做到了;捐赠者和执行者都缺乏清晰反馈。

  • Labenz提出的加州法案SB 1047,会给私人监管机构认证,用自愿监督换取责任保护。Mowshowitz预计市场会出现供给,METR或Apollo这样的现有评估机构可能成为早期进入者。

  • 硬件治理与当前“卖出我们的硬件”政策目标相冲突,在政治上并不受欢迎,但仍是高价值期权。相对有限的投资,就可能在危机制造政治需求前,让防篡改芯片追踪技术达到可直接启动的状态。

  • 这项能力可以支持阿联酋或印度的安全数据中心,减少走私,并在不盲目信任的情况下实现更灵活的出口规则。眼下需要的不是大规模部署,而是找出“真正靠谱的东西”,并保留准备状态。

28. 对抗性评估值得做,但可信度是约束

  • 在OpenAI因怀疑竞争对手资助而传唤不方便的慈善机构后,Labenz提议故意打破实验室之间的绅士平衡:资助针对除一个赞助方之外所有实验室的可复现攻击,迫使每家公司暴露竞争对手的模型缺陷。

  • Mowshowitz看到了可口可乐与百事可乐式的博弈:攻击对手的安全问题,也会突出自身类似缺陷,引来监管,并带来升级风险。公司可以选择不开枪,因为相互克制在没有明示串谋的情况下也符合各自利益。

  • 选择性资助即使产出完全可复现,也会损害合法性。“凯撒的妻子必须清白无瑕,而且还要经得起指责”;监督机构仅仅为了进入这个竞技场,就需要在案例选择、解释和披露上做到异常严格。

  • 讨论仍把公司之间的对抗性评估视为一种理想均衡,例如OpenAI与Anthropic之间的评估,以及对DeepSeek生物安全协议的评估。均衡听起来很有价值,但两人都没有找到从当前激励出发的可靠路径。

29. 实际行动是阻止显而易见的伤害,并直说真话

  • Mowshowitz眼下的政策优先级,是阻止美国向中国出售H200。策略上,就是让足够多的右翼人士意识到NVIDIA的影响力,以及把出口当成目标所带来的战略后果。

  • 对于正在可选工作之间做选择的技术人才,他现在认为,去Anthropic工作相对于无所作为“显然”是积极选择,同时承认,在有承载能力时,专业对齐或政策机构可能带来更高影响力。

  • 对捐赠者而言,机会覆盖对齐实验、政策、外交、机构连续性和硬件准备。瓶颈是足够灵活的资金,包括支持那些过于精简或过于泄气、以至于不敢公开募资的机构。

  • 他最广泛的指令不是组织性的,而是知识上的:不要粉饰、夸大或策略性审查;“说出你真正相信的东西”。他也鼓励支持Eliezer Yudkowsky和Nate Soares的《If Anyone Builds It, Everyone Dies》,但不要把作者当作绝不会犯错的权威。

Nathan Labenz

Zvi Mowshowitz, welcome back to The Cognitive Revolution.

Zvi Mowshowitz

Thanks. Great to be here. Always exciting.

Nathan Labenz

Let's start with timelines today. I am a little confused about the way in which people seem to be updating over the course of this summer of 2025. We've had GPT-5, obviously, where I would say it's safe to say the launch was not exactly super smooth, and people were a little disillusioned with that at first. Now the dust has settled, and it seems like people have mostly come around to the view that it's actually a good model and basically on trend. It's actually still a little bit above the all-important METR task-length curve.

Zvi Mowshowitz

It's a very large shift upward, right, on that graph?

Nathan Labenz

The 4-month doubling instead of the 7-month doubling. Something else that seemed really important to me that happened this summer is, of course, that we got IMO gold and a finish very close to number 1 in the world in competitive programming, although it ended up at number 2. And yet, with all these things, people—including really smart people—seem to be lengthening their timelines.

I'm not talking about denialists here, but people like Ryan Greenblatt and Daniel Kokotajlo, who are as plugged in as you get and, I would say, smarter than me. You can evaluate that for yourself. How have you updated your timelines, if at all, and what do you make of these seemingly on-trend, or maybe even slightly ahead-of-schedule, events that are still leading people toward longer timelines overall?

Zvi Mowshowitz

I would say mine have also gotten modestly longer on net during that period, for a variety of factors. The biggest one is that very large jumps in capability in the news really shorten timelines. If you don't get big new developments—big new paradigm developments like reasoning models, or big jumps off the curve, stuff like that—then you are cutting off a lot of what led to the really fast timelines.

Whereas none of the things that we did see, while they were on trend, represented anything that we had that much doubt was going to show up, even if it wasn't going to show up quite on that exact day. Nothing was particularly impressive given what had already happened.

These things are ahead of trends from a few years ago, for sure. The IMO gold medal in particular sounds more impressive than it is because this year's gold medal was especially amenable to LLMs. Normally, the 6 problems are divided across 2 days: day 1 is problems 1, 2, and 3, and day 2 is problems 4, 5, and 6. Problems 3 and 6 are supposed to be incredibly hard, so the idea is that a significant number of kids will get 1, 2, and 4, but 3 and 6 will be a struggle.

If you manage to crack either of those problems, you're doing really well. If you manage to crack both of them, you get a gold medal. You do really well. What happened this year was that problem 3 wasn't that hard. This is just a weird quirk of the schedule of the IMO. It has nothing to do with AI, but only problem 6 was hard in the IMO sense.

None of the LLMs got any points on problem 6. They all flopped. But problem 3 was evidently solvable, and the LLMs all solved it. The fact that this happened to be the case—and the fact that 7 points on each of the first 5 problems is exactly enough for gold—meant that an AI could get a gold medal without screwing up any of those problems, given this exact threshold. If the threshold were 1 point higher, no one would get gold. If problem 3 had been as hard as problem 6, probably none of them would have come that close to gold.

So it's not clear that we actually got gold in a meaningful sense this year, and we were already reasonably close to gold from previous results. That's why nobody who was familiar with the state of the art was freaking out that much once they saw the details.

If you look back at things in 2020 and 2021, only the most radical people who were extremely AI-pilled and predicting very rapid progress would have called that happening in 2025. In fact, even they didn't call that happening in 2025 on the extreme right tail. But by 2024, this would not have been much of a surprise.

So, given that it didn't crack problem 6, this wasn't that scary on the margin. The programming results, again, if you dive into the details, are really, really impressive, but not in a surprising way. If you look at the details of the competition, have you read the write-up from the guy who won?

Nathan Labenz

Superficially. So, yeah, take me through it.

Zvi Mowshowitz

Yeah, I read it. Essentially, what happened was that OpenAI's model was very good at jumping to a very good solution almost immediately and had a lot of gains pretty fast. But it was not doing a good job of iterating, was not doing a good job of innovating past that, and was not doing a good job of making conceptual progress past that point.

Over time, the winner was able to essentially plan ahead, open up a bigger lead, and get into a pretty comfortable position by the end of it. If the competition had gone on longer, OpenAI's position probably would have slipped below second.

Previous years, you give that problem set to the AI, and it doesn't get second place. It doesn't do particularly well. But, again, we're making steady progress up the task scale. It could easily have been first if that guy hadn't shown up that day. It could easily have been third or fourth if some other guy who showed up had done better conceptual work.

But if the project had been more serious and had gone over a longer length of time, its performance would have degraded relative to the humans. You can make of that what you will, but it's not that unexpected. Again, it fits into the bigger overall picture.

I've been talking about good reasons not to be that impressed. We got good results, but not great results. You often have these situations where you get a 50th-percentile result or a 40th-percentile result, or something like that, and that actually significantly lengthens your timeline or weakens your expectation of what's going to happen relatively soon.

A lot of what you were doing was factoring in the possibility of things going fast. The chance of AI 2025 has gone down dramatically. Four months ago, compared to now, it was higher by a factor of 10 or more, relatively speaking. The chance of AI 2026 has also gone down quite a lot, and I think 2027 as well.

AI 2030, I don't think, has gone down substantially. I don't think it's gone down by anything like the same amount; I think it's gone down a small, modest amount. However, other people are always looking for any excuse to say that AGI is not a thing, that it's not going to come anytime soon, that we're hitting a wall, that reasoning is hitting a wall, that o1 is hitting a wall, that the companies are unprofitable and are going to go bankrupt anytime soon, that there's no way you can possibly earn enough revenue to pay for all the misinvestment, and so on—or just look at an ordinary marketplace.

This is the new David Sacks—the David Sacks official government line coming from the White House at this point. And by the White House, I probably mean the view, let's face it, of AI: It's just an ordinary technology that will do lots and lots of amazing things. What matters is that we capture enough of the chip market, because that will assist people in running the American tech stack on American models. None of this makes any sense or has anything to do with anything, but to sustain that argument, they need to act as if AGI is just not a thing. It's just impossible.

They're not quite at the point where they can completely ignore it without mentioning it, which a lot of other people just do. They don't say the word AGI and then act like it doesn't exist. Or they take the word AGI, along with the OpenAI line, and then act like nothing will change and everything will just be normal anyway, which makes even less sense. Instead, this is the line that makes more sense: No, AGI isn't going to be a technology soon. We just aren't going to get there.

And they argue, “We know this.” A lot of people are acting like, because GPT-5 was so disappointing, this proves that we won't get AGI imminently, anytime soon, within the next few years, or whatever it is. Therefore, we can all relax and focus on making sure America wins the tech-stack battle, or whatever that's supposed to mean, which it doesn't mean anything. This line makes no sense.

It's just a matter of them botching the rollout. On the day that they put this out there, everyone was being directed to mini models all the time. The router was completely broken. Nobody understood that they were supposed to use thinking, and they didn't have access to the toggle. No one was trying seriously, because it takes longer to do that, so that wasn't where the focus was. People just didn't appreciate what was happening.

Then there was this big howling about how GPT-5 was attempting not to glaze the user the way GPT-4o glazed the user, which is a very virtuous and good thing that nobody was doing. But people were mad about it because people like glazing; that's why we get glazed. They were trying to give you your medicine and not your sugar, and people were demanding their sugar. Then they said, “Okay, fine. I guess we're going to give you the option to get your sugar back,” and everyone cheered and said, “Yay,” because we have a bunch of children, which is unfortunate, but it is what it is.

The combination of these factors made everyone feel like they had botched the rollout. The other important factor, of course, is that they showed images of the Death Star. They called it GPT-5. They hyped this release up as if it were the next big thing when it was just the next incremental progress.

If you look at the combined progress from GPT-4 to GPT-4 Turbo to GPT-4o to o1 to o3, and now GPT-4.5 and GPT-5, GPT-5 looks amazing. I don't think it's unreasonable to say that GPT-3 to GPT-4 is as GPT-4 is to GPT-5. I don't know if it's better, or at worst; we don't have enough information to know exactly how big the leap from GPT-3 to GPT-4 was. I think it's pretty big, but certainly, compared to GPT-3.5, GPT-5 to baseline GPT-4 is a significantly bigger jump than GPT-4 to GPT-5.

Nathan Labenz

Yeah. A couple of little follow-ups, or double-clicks, if you will. One is that, to me, the IMO thing still seemed like a meaningful update. I'm not great at math, so you can maybe give a better read of this than my naive one, but so many of the math problems that LLMs had been measured on were problems where there is a right answer, and so they were easily verifiable.

Yes. What seems qualitatively different in the IMO competition is that you have a proof that you have to write, and it is not super easy to verify. That seems to suggest that the worry that, well, you can give these things tons of easily verifiable problems and they'll get good at that, but will that really generalize to true reasoning? It seems like we've kind of answered that question, or this strikes me as a significant update in that direction.

Zvi Mowshowitz

No. So, first of all, we'd already seen models get IMO-style problems, or previous IMO problems, correct. This was not a huge leap. We'd previously seen AlphaGeometry and so on, so it wasn't a great surprise there.

Nathan Labenz

But those also have a symbolic component, right? Like those AlphaGeometry things, don't they? I'm not saying it wasn't a big deal in that sense, but it was right on track. We'd seen o1 manage to present proofs of these types, but it was right on track.

Zvi Mowshowitz

I am someone who went to the USAMO. This is the level before getting to take the IMO, which is the level below being able to actually compete for a gold medal in the IMO, which is the level below actually getting the gold medal. I was not very good, and I believe I got effectively a zero on the USAMO the previous round, which is what the majority of people who take it get.

I sneaked into the bottom of that round, which is the third round of competition, basically out of effectively 4. I would practice with people who did, in fact, end up going to the IMO. I was in a room with, I think, 2 people who would go on to be IMO team members for the US. We would be given previous-year IMO problems to work on, and they would get them while I wouldn't get them right. Most of the time, they would make some progress or get the solution, and I would make very little progress, some progress, or no progress. But then they would show me the proof, and while I was in training, very often, once they showed me the proof, I understood it. I could verify the proof. I understood why it was correct, and I would be very confident that, if they were trying to pull a fast one, they were not mistaken about the proof.

Verification is much easier than generation in this space. These aren't problems where it's really super hard to know if you're on the right track. I'm pretty sure that these same AIs that could not solve problem 6 would be able to verify problem 6. If you asked it, “Is this a correct proof of problem 6?” and gave it sufficient resources, it could directly classify the proof as correct or incorrect and explain why when given the candidate proofs.

The thing to keep in mind about the IMO is that people say this—it's crazy, it's not real math. Everyone will always say, “This isn't real math. This isn't real math. This isn't real math. This isn't real math.” In a real sense, the IMO is not real math. It's the best indicator we have of which high school students will be able to go on and do real math in the future. It's a very strong indication of math talent, pre-math skills, math interest, and so on.

But you have a very limited set of moves that you can make using the math that is allowed. You're allowed to use whatever math you want, but it has to be a solvable problem using only high school tools. There's only a limited, finite, very compact set of high school tools that you're allowed to use in these proofs. So there's a compact set of potential problems, and that's very different from the moves you can make in a PhD-level math proof, where you're actually doing new math.

The proof also has to be gettable within a certain amount of time and written within a certain amount of space. There are a lot of restrictions on what the problem can be. A lot of being good at math competitions, including competitions well below this level, is understanding that they had to have given you a problem that you can solve, with a solution within the amount of time and at the degree of difficulty at which you're being presented the problem. Therefore, you can use your search time to search the space in which the solution is going to be, given those facts.

This makes it much, much easier, especially when you have 10 minutes to do your 2 problems in a much lower-level competition. You only have to search the types of solutions that quite reasonably take 5 minutes and use the tools you're allowed to use at that level of competition. You get used to all the tricks for figuring out which of these things you're reasonably going to be asked to do right now.

It's a great leap. It's a great indicator. It was a great test. It was a big milestone, and it came much faster than most people expected, but it's not as much of a panic moment as people might have thought beforehand. I think it's right for everybody to basically shrug this off.

If you've been doing your homework before, if you traveled from 2020 or, God forbid, 2015 and arrived in 2025, and the first thing you asked was, “Okay, how are we doing on the IMO?” and they said, “Gold medal,” you should go, “Holy shit.”

Nathan Labenz

In terms of the overall AI landscape, what might we infer from this? How do you interpret the fact that multiple companies did it at exactly the same time, with seemingly the exact same techniques, and seemingly getting exactly the same problems? I guess you said the one that they all got wrong was clearly harder, so maybe it’s as simple as that. But I’m always wondering: to what degree is there information leakage from company to company, and to what degree are they just following the gradient of their own work, with it taking them all to exactly the same place because that’s what nature is dictating?

Zvi Mowshowitz

Yeah. In this case, I don’t think there was probably substantial leakage of techniques. I think it was more that the IMO comes about once a year, right? As you know, you have to avoid contamination, and you only have so many shots and so much data to try it on. It has to be a pure test, which means you only get 1 shot per year.

So, it’s entirely unsurprising that, in the same year, both OpenAI and Google got to the destination. Don’t think of it as them having the same great breakthrough on the same day; think of it as this cycle, this academic year, being the year when both got there. There’s a pretty big gap between being capable of getting exactly where they got and getting beyond that point.

The fact that they got to the same place, given the techniques they were using, isn’t that surprising. It’s more like, okay, if you try the natural things—if you make a real attempt to use the compute efficiently and scale it up for inference—you can get to this point. You also learn to get to about this point, and not much farther beyond it.

Nathan Labenz

You invoked the AI 2025, 2026, 2027 part of the narrative, where companies will start to withhold their best models, keep them private, and deploy them internally only for the automation of AI research and so on. What will you be watching for? There’s been this limited communication that it’ll be months, maybe many months, before they release a model of this capability.

How do you understand that GPT-5 wasn’t a big scale-up? Another interesting data point—and perhaps the one I saw that was least commented on—was in SimpleQA, this super-long-tail, esoteric trivia benchmark. GPT-4.5 actually scores quite a bit better than all previous models and better than GPT-5, while GPT-5 is basically in line with GPT-4o.

I think that is probably the clearest indication that the model is not bigger. To absorb all this super-long-tail esoterica, you seem to need a lot of weights to store it in, and there might be some fundamental compression limitation on how many facts you can fit into a model of a certain size. GPT-4.5 got a lot more facts. I think it was 12 or 13 points higher on SimpleQA, and these are simple questions where it’s literally just whether you know the answer or not. You can’t really reason your way to those answers.

GPT-5 is on the level of GPT-4o. GPT-4.5 is quite a bit higher. We’ve also got this—I don’t know, obviously nobody knows what size the model was that did this IMO thing—but it clearly can reason for a really long time. Do you take that as the beginning of a widening gap between internal deployments and external ones? There’s always some gap there, but what will you be watching for to assess how much they’re holding back and only using for their own purposes versus continuing to share with the public?

Zvi Mowshowitz

It’s awkward to me that we would even say that, because even if we knew for a fact they were holding back, it wouldn’t be obvious that it was because they didn’t want the models in the hands of the public. It might be because they were afraid the models would speed up R&D too much, or because they were concerned about other misuse problems, or any number of other reasons.

The most likely reason is that they’d be too expensive and too slow, and they think that would be bad for the brand and bad for sales. They only have so much compute, and they’d rather not release it for that reason. GPT-5 is very clearly intended to be the best product they can make in terms of what they can serve people for the total amount of compute and time that they’re investing.

Their thinking is that they’d rather have people spend that compute at inference on this-sized model for Thinking Pro than have them think less with a larger model that maybe knows more, when we have web access. It’s the same way that I don’t often try to memorize things that I could look up, even though I could, because I’d rather spend that cognitive power on something else.

If I had the ability to make my brain bigger in some sense in exchange for other handicaps, I wouldn’t do it just to memorize more facts. It wouldn’t really be worth it. So, it’s not that surprising.

So, O4, right? There’s an O4-mini. Where are O4 and O4 Pro? We’re never going to find out, right? Obviously, you could say GPT-5 Pro is O4 Pro or whatever, but it’s not. I think this is a distinct, different thing, and they concluded that there’s not much commercial call for it.

Potentially, most of the commercial call for it is from people who are in direct competition with them. It’s not so much that they don’t want Anthropic to have access to their best model; it’s that no one else is going to want it very much in some important sense. For other types of projects, this model is very slow and very expensive. So why do you want it?

GPT-4.5 was an experiment: what if we scaled up a lot for the humanities rather than for code? What if we tried to make this thing that had taste and could do these cool, creative things, even though it was going to be slow and expensive? How much would you get out of it, and would it be worth it?

The conclusion was that there was a handful of people who really liked what it was and liked it for certain cases, but in general it was just annoying, and they pretty much regret losing it.

Nathan Labenz

Did you find any value in it? I tried using it for some writing tasks, but I wouldn’t say I unlocked the—

Zvi Mowshowitz

There was a narrow set of use cases where it was the best choice, and I was happy to have it. There might still, in theory, be a narrow set of cases where, if you already have access to it, it’s the right model to use if you’re not in a hurry.

But I was never actively excited that I had GPT-4.5 instead of something else. It was more like, I guess technically this is a job that calls for GPT-4.5, given my choices. I’m pretty sure it wasn’t worth the complexity cost, and I’d be totally fine not having it.

I also think there was a paradoxical choice problem. If you offer me a better version that’s only slightly better but a lot slower, I feel bad no matter what I do, and I’m actually worse off in practice. So I’d rather not have that choice.

Nathan Labenz

Do you go back to GPT-4o at all now that the option is restored to you?

Zvi Mowshowitz

GPT-4o does not exist, to be clear, as far as I’m concerned. The only models that exist are Opus 4.1, GPT-5 Thinking, and GPT-5 Pro.

GPT-5 Auto exists only for a narrow set of queries where you’re using it like a Google search, a calculator, a transcription tool, or some other very clear technical task where Auto is just fine. Due to the way the web works, I wasn’t able to find a transcription website for images that would just translate the text and wasn’t an LLM that defended itself using some sort of weird identification system that slowed my entire computer down to make sure it wasn’t a bot.

I’ll just use GPT-5 Auto and have it write me a transcription program and automatically do the transcription. It’s just easier. The window’s already open. What do I care? It’s not like the pennies I spend on compute matter.

I don’t really have use cases for Gemini 2.5 Pro. In theory, Deep Think and Deep Research exist as well. I probably should be trying Deep Think more, given that I happen to have access to it. Something about having 5 queries a day makes me excited to use it: it feels scarce, and it also feels like I don’t necessarily need to find out. I should use it occasionally, but for normal GPT-5 Pro, not really.

The kind of very clean query is just that I want to know something that I know you all know, and I want you to lay it out there very clearly, very cleanly. If my 11-year-old wanted help with his homework, I would be tempted to use Gemini because it would give him very clean, friendly help and explanations. I’ll also use the image generator. The image generator is cool.

Yeah, that’s getting really good. I have a lot of use cases for that, including potentially mashing the 2 of our faces together into a thumbnail for this podcast.

Nathan Labenz

I do think I basically agree with you, although I do use Sonnet, especially in coding, because it is obviously a lot faster. I also use Gemini 2.5 Pro, which I find to be similar to what you were saying: the most straightforward model to work with. Because it is so literal in its interpretation of your instructions, a lot of times it can be really good for tasks where, for example, I want to compile documentation.

I’ve given this PSA multiple times: if you have an API, it’s time for an llms.txt. I’m tired of having to sic an agent on every page of your documentation website to compile all that documentation into some super-bloated thing that has all the menus on it or whatever. What I have found Gemini 2.5 Pro to be amazing at is those sorts of things where it’s like, “Here’s an almost abusive level of context dump. Could you just clean that up for me into 1 streamlined form?”

It is amazing at taking 500,000 tokens of documentation with all these examples and all this cruft, where the menu got copied 50 times for 50 pages, and putting it into 1 clean thing for me. It will just do it. Man, that thing is an absolute beast of a workhorse.

Zvi Mowshowitz

The only thing my Chrome extension uses is Gemini Flash. That’s legacy. It was the easiest one to get working, and it was the cheapest at the time. It just works, so why would I bother switching it? Probably I should be using Sonnet or Opus at this point, but who cares? It works just fine.

Nathan Labenz

Yeah, I’ve had problems with Gemini and instruction following, though, even for relatively simple queries, especially with web research, where it will just ignore my request and do something in the same general area, but not the thing I asked for. I think that’s a lot of what put me off it. I don’t know why that’s happening, but I haven’t had the same problem with the other models.

Zvi Mowshowitz

Yeah, I should use Google Deep Research more as well. For whatever reason, habits are formed quickly sometimes.

Nathan Labenz

I don’t know. There was a period of time when Perplexity was definitely heavily in my rotation for certain types of queries, and I had gotten to the point where it just isn’t anymore. If I’m going to use Perplexity, either I want to move down to Google Search because I just want an instantaneous answer, and I think my brain has a Google query that it expects to work—System 1—and therefore I’m going to try that before I do anything else. Or, no, actually, it’s not going to work, and I want to move up to at least Opus and possibly GPT-5.

Zvi Mowshowitz

I do think GPT-5 has been better for me recently at search than Perplexity. I’m starting to update that habit. But I’ve noticed that the persistence of which things I go to for what is idiosyncratic in some ways and probably has more inertia than it really should.

Nathan Labenz

Yeah, I think it’s fine to be somewhat inertial and idiosyncratic. I also think it’s fine to want to support and use certain people’s products as long as you’re getting what you need. If I ever felt there was a query where I didn’t think my usual services were going to get it done—if I didn’t think Opus was getting it done, if I didn’t think GPT-5 was getting it done, if I thought Gemini was getting it done, if I didn’t think Google Search was getting it done, but Perplexity might—I would absolutely open it.

I have zero expectation that if I ran it in GPT-5 Pro and failed, Perplexity would have a genuine chance of helping.

Zvi Mowshowitz

That seems totally reasonable.

Nathan Labenz

Okay, so we’ve got the summer of on-trend releases. Timelines are extended a bit because we haven’t seen the most extreme stuff that we would need to see to maintain super-short timelines, right?

Zvi Mowshowitz

Right.

Nathan Labenz

Does that also—

Zvi Mowshowitz

No, go ahead.

Nathan Labenz

I just want to think of it as OpenAI’s GPT-5. This was on trend and unsurprising, but the fact that they chose this as GPT-5 is information that they didn’t have some big weapon in their arsenal and that we shouldn’t expect a big weapon in their arsenal for at least a few more months. We can basically discount GPT-6, right? The next level of big jump coming from OpenAI in 2025 is basically not going to happen, and that has to update your information.

So, last time we spoke, I believe your P(doom) was 70%. Has it ticked down at all?

Zvi Mowshowitz

The timeline’s getting slightly longer, which is good news, but there has been a plethora of bad news that has definitely dominated for the most part. We’ve had the bad news of the United States weakening its position voluntarily in a number of ways. We’ve got a Department of Energy that’s actively going to war against windmills and, to a large extent, solar and batteries. So the US will have a persistent energy deficit for a long time, and this is potentially going to cause us a lot of problems that will make our position much worse.

The US has effectively been captured by NVIDIA on export controls, to a large extent, such that the White House has made the H20 legal for export. The Chinese are turning it down because of weird things that I can get into more than we should. If they turn down the B30A, that would be a lot more surprising.

Right now, it looks like if there isn’t lobbying to stop them somehow, and if there aren’t enough people who are sufficiently upset on the right to make it clear that this is just Crazy Town and Bonkersville and they cannot do this, they might actually do it. That would be a huge weakening of our technical position. No, paying 15% of profits to the government will not change the impact at all. That’s just a little bit of a check to make people feel better.

If they’re trying to catch up in compute in this way, that’s a really bad sign on numerous levels. In general, if the attitude is that the primary purpose of the United States government in AI is to sell as many chips as possible and ship as many chips as possible, including to our main rivals, that’s a problem. Our main rivals are somewhat obsessed with internal chip manufacturing for very good national security reasons and have caught on to AI, but they will still be pursuing maximum compute to a large extent anyway, thus acting effectively correctly, mostly regardless of what they think is going on.

These are not particularly good pieces of news. With this shift in the way our situation has developed, I think that more than makes up in my mind for the modestly elongated timelines. I also think there is increasing evidence that RL hurts alignment pretty directly. The more RL you do, the less aligned your model is, even in a pedestrian sense. Models are getting more and more RL continuously.

So I think we should expect the default to be that our models actually get less aligned in this next phase, unless someone does something about it. That does not go well at all. And just in general, we're dropping the ball, right? I like the way Jan Leike has been putting the situation recently: We have very much done almost nothing to try and align these models. It's pathetic how little we've done to try and align these models, and we're advancing very rapidly.

But we have been blessed with a strange amount of grace in the way these models have, by accident, been inclined to do things that are reasonably friendly to us. It's about the fact that we don't understand what we're doing, we don't understand how they're made, and we don't understand why or how we're trying to align them. We're not really trying very hard at all. We're not making it a priority at all. And it's mostly been okay on a practical level.

We've lucked out that we haven't had any big catastrophes or significant incidents, really, and so on in various ways. But all of the things we were scared of underneath the hood are absolutely there, absolutely under the hood, and absolutely scary. In fact, they're manifesting and happening, but in the most graceful, blessed way such that we notice them.

They happen in ways that let us prove and acknowledge their existence and respond to them without anybody being seriously hurt or any damage being done, which is amazing. Except we are then dismissing all of it effectively, right, as a civilization. We're just moving on and acting like nothing happened. We're finding ways to even deny that AGI is a possibility in the medium term, when the evidence is pointing the other way.

We are so greedy and so demanding of AI that it is the most rapidly developing, most rapidly deploying, and most rapidly impactful technology in the history of the world that didn't involve directly killing other people in the middle of a war. And we're like, “Oh, we're hitting a wall. Oh, we're slowing down. Oh, it's not going to—because what? Because you didn't get blown away in the last 3 months? It only modestly improved. What are you even talking about?”

Basically, we have no dignity. We have no dignity whatsoever. Maybe we live in such an extremely fortunate world, compared to what we had any right to expect, that we might be able to pull a victory out of nowhere somehow, even if we don't see how to do it yet. But, yeah, I'm not optimistic. I don't think it's changed substantially, though.

I'm not trying to give two digits. I don't think you get two; at most, you get one significant figure of doom. I don't think that's reasonable. So we're staying at 0.7, but you're starting to see the needle possibly shake up.

Nathan Labenz

More up than down from the previous assessment, but not enough to move to 0.8. Certainly, that doesn't look good, right? Most of the same root dynamics in all directions are still there. Again, this is basically on trend: Some good news, some bad news.

Some technical stuff looks good to me in ways that we don't have time for—or I'm not sure I want to talk about in public anyway. But, yeah, I am hopeful that there's room to do some stuff at a very low degree level that might be helpful. Everyone has their theories if they hang around.

Zvi Mowshowitz

Yeah, I heard Holden Karnofsky say we might be moving into a scenario where we could have success without dignity. The basic idea is just defense in depth: Layer all these things on and hope we can catch enough stuff before it gets through and muddle our way through.

Nathan Labenz

Yeah. The concept from him—to be clear, I listened to his podcast with Spencer Greenberg, where he talks about it—but I think his vision of how we do that is wrong. I think the idea that we can get there with defense in depth is mistaken. I just think that all correlations, when you're facing a sufficiently intelligent enemy or a sufficiently powerful optimization process—even if it isn't an enemy per se, and isn't strictly trying to do anything—go to 1.

All these things will fail for basically the same correlated reasons, at roughly the same time, in a predictable fashion, if all you're trying to do is this kind of lazy defense in depth. I'm thinking more along the lines of something like what Jan Leike has talked about as grace: the idea that we might be able to create a system that wants to converge on the right answer and therefore collaborates with us at a genuinely deep level, assisting us in finding the target that we want to find. Therefore, it can land on the moon.

The metaphor that MIRI likes to use is that if you don't know how to aim your rocket, you definitely won't land on the moon. It's not, “We might not land on the moon, but something would have to go wrong. We'll probably land on the moon anyway.” Just aim a rocket at the sky—that obviously will not work, right? Physics has laws. If you're not aiming at exactly the right spot, if you're off by an inch, you just don't land on the moon.

If you have a system that is capable of adjusting intentionally to pick the target, maybe you can land on the moon. I don't want to give false hope, but basically, I don't think defense in depth does what it advises you to do over time. It keeps things from going crazy for a brief window, but that brief window could potentially be enough to do the thing you need to do. It won't work indefinitely.

This is the idea that Holden and some other people have: You can basically have a bunch of AIs that are like, “I really want to kill all the humans, and I really want to take all the resources to do my thing, but I typically don't know how.” If I try, you have these supervisors who also want to kill you, but they have other supervisors, supervisors, and nobody knows who's watching whom.

Every time you find someone who's coming out of line, you stop them from trying to do that. You won't even try, and if you did, it wouldn't work. It would fail for some reason you're not thinking about right now, but it would definitely fail for reasons that none of us are thinking about right now. The high weirdness will come, and you will die—very much. I don't care how much defense in depth you put on top of that; you are so toast.

One thing that struck me recently was a bunch of people saying, “Oh, you people were talking about biorisk, but actually, the risk that we're seeing is sycophancy.” The actual risk we're seeing is people being driven crazy by all of these weird dynamical processes. First of all, I've been arguing for years that superpersuasion should be in the preparedness frameworks, and it was bad that they took it out. But the whole thing we've been saying is that the AI will find ways to hack your brains, cause weird things to happen, optimize for a thing that nobody was planning to optimize for, and cause strange outcomes that you did not anticipate.

That was the thesis. When you accuse us of not anticipating the thing, I realize you can call that a cheat, right? Anything we didn't anticipate counts as correctly anticipating anything we didn't anticipate. If it's something we didn't anticipate, then we anticipated it, so we always win. But we're also basically saying you see weird stuff that nobody intended, and they start to go bad, right? It's going to start to diverge from what you would want increasingly over time. You just see these fire alarms keep going off in various different ways, but that's all underneath there. It's all under there.

One way of thinking about it is that before you started doing RL, you also just didn't have things that were in a position to cause the problems you worried about. The objections to these systems were true, at least to a point: You were okay. But now you're starting to see all this stuff show up, and it's a freaking disaster in the making. It's exponential.

A lot of these complaints are also along the lines of the similar thing I'm saying. It's like you said in January, “We're all going to get COVID,” but now it's February and nobody I know has COVID. “What were you even talking about? No, you're wrong.” And then we say, “But actually, there are a lot more people with COVID than there were in January. Can't you see what's about to happen in March and April?” And they absolutely do not. It's not that clean, but it does feel that way.

So can you give me a little bit more on the mental model of why everything is going to fail at the same time for the same reasons? That is not intuitive to me immediately, and I suspect a lot of people don't even know what you mean.

Zvi Mowshowitz

What I mean is roughly that when you're facing things that become importantly more powerful and smarter optimizers than you, they're capable of finding solutions and capable of finding ways of manipulating the physical universe that you didn't think of. They're capable of going outside your model of what might happen and surprising you. None of your defenses anticipate those things, and the system is going to search the space until it finds ways out of them.

In a sense, it's going to keep improving its capabilities, because you're going to help keep improving them because you want them improved, until it becomes capable of finding those ways out.

And at about that time, it also becomes capable of doing things like strategically hiding that it has capabilities, strategically hiding its plan, strategically hiding its memory and its thinking, obscuring its chains of thought, and doing all these things. These all roughly emerge at the same time, and so you should expect to be very surprised.

But obviously, a very smart agent or mind will turn against you exactly if and only if turning against you will work. It will do things that you do not want if and only if it will work out for the thing doing them—or if it will do something that you would not want on reflection but would think you wanted when it was first shown to you, like deceiving you.

Right now, we’re seeing versions of this where it will just hack the test and say, “Return true,” at the end of the Bash script, right? Because scientists want to hack the test. It will do incredibly lame, silly versions of these things and get caught, and you say, “Oh, it’s annoying and fine, but it’s annoying.” But it gets caught, right? You notice that it’s hard to miss.

But that’s exactly the threshold where it’s not, “I hack the function exactly when I know you won’t find it.” It’s, “I hack the function in a way that I know you won’t discover that I hacked the function, because otherwise I wouldn’t have hacked the function.” If it can model—or, I’ll think of it as—if it’s capable of modeling the processes that are checking for its actions well enough to know how those processes will respond, then—

Right? So there are things that happen sometimes in fiction, just as a visualization metaphor. It’s an intuition problem. You see it on Person of Interest, for example, where you’ll see a scenario play out, and then something will happen and it will go wrong, right? You’ll see the time start to rewind: “That didn’t work out. Let’s try another branch of the Monte Carlo simulation. Let’s try a different set of moves. Try a different scenario.”

Or, like, in Avengers: Endgame, you have Doctor Strange with the Time Stone, right? They say, “I looked at 14,000,605—I forget the number—but 14,000,605 ways this could possibly go.” And they ask him, “How many times did we win?” He holds up 1 finger: 1, right? Because the guy with the Time Stone is the one choosing which of those paths to lock down. So we win. But until the point when you can do that, you lose. And if you lose, you give up. There’s no point in trying, right? You let the rookie win.

Obviously, I won’t literally be able to do this. That’s nonsense. But what I’m saying, essentially, is that a sufficiently strong predictor—so, like, a predictor and an optimizer are the 2 halves of intelligence in the mirroring model—but a sufficiently strong predictor and a sufficiently strong optimizer combined suddenly mean you don’t know what hit you, in a very real sense, right?

And if the AIs start to be capable of persuading people—again, all your AI is very dumb. It doesn’t necessarily just convince you to take them down for it, but I don’t know exactly how all of this goes, right? A lot of these scenarios simply involve nothing exactly—nothing even goes wrong. It’s just that every single person, with different people having different AI agents, directs them to do the various things that are good for those people. And the AI agents actually do those things that are good for those people.

But everybody who doesn’t direct their AI agents to go as hard as possible for the things they told the AI agents they should want, and to pursue resources basically as hard as possible with an increasing percentage of their attention, will just lose all the resources to the AI agents and the people with AI agents that did that. So everything else just loses out. Everything goes haywire, and no one ever turned on anyone. There was nothing even that surprising. They just—whoops, the end, right?

These failsafes don’t even work, even if they want to. I have this sort of global sense of despair toward this. I put in a set of enumerated, detailed limits and rules—we said them in English out loud, passed laws for them, and added supervisors, checks, approvals, and loops—and none of that will survive contact with the enemy when the time comes. And, yeah, we’ll probably all more or less fail.

Yes, obviously, some of the defense in depth will just fail randomly. There would come a point where it feels like it’s all failing more or less at once, in a way that feels out of line with the previous percentage of failures. It will be surprising if you didn’t understand that was going to happen, but you should expect that.

Nathan Labenz

And so the possible positive version of this that you see sounds like a sort of coherent extrapolated volition kind of idea. I think that was a specific rabbit hole that a bunch of people went down, and I’m skeptical of a specific technique. But again, I am not actually a machine-learning expert, and I’m not trying to solve alignment. So this is the part where my specific ideas should be taken with copious amounts of salt and not trusted. Why? Who am I to say anything?

But I learned that I shouldn’t just shut up because I feel like that. I’ve felt stupid for showing up so many times in the past when I’ve shut up for that reason that I got over it. It looks more like developing AIs that are sufficiently virtuous, that are sufficiently desirous of becoming more virtuous and more desirous of engineering the things that we would actually want on reflection and the things that we actually value on reflection, such that you get a positive feedback loop where it reinforces this thing and you’re optimizing for optimization to hit the moon, right?

The thing wants to develop a NASA that will hit the moon, and therefore it hits the moon, right? If you try to just steer the actual rocket directly, you crash or miss the moon entirely. It doesn’t work. If you try to set up a bunch of rules to make sure that it has to launch and hit the moon, you don’t hit the moon. You can build a culture that wants to build an organization that wants to build a rocket that will hit the moon, and maybe you can hit the moon, metaphorically speaking—something like that. But, yeah, hopefully.

So, last time we spoke a little bit about the fact that scaling inference-time compute allows you to potentially have a GPT-N that can effectively monitor and supervise GPT-N+1, right? And some of these ideas sound very much like a constitutional approach, but with maybe the additional opportunity for the model to modify its own constitution as it goes through these generations. Is that the picture that I should be envisioning?

Zvi Mowshowitz

I mean, I think that’s vaguely the best picture that I see that’s compatible with the level of dignity that we have to work with, or something like that, because we’ve been warned for decades not to let the AI do your AI-alignment homework. That is the worst possible path you could go down, because this is the hardest, most complex problem. And yet here we are, right? This is the only option we have because we don’t have the time, right? We don’t really have the cooperation to try any fundamentally different path from that. We have to go down some kind of path that’s vaguely in that range.

So, yes, I think it’s vaguely something like: you use the fact that you can scale inference up and down arbitrarily, and you can evaluate outputs and do reinforcement on relatively scaled-down versions of the thing, and use it to monitor, verify, and check for various attempts at malfeasance, including malfeasance during training, and blah, blah, blah.

And if you combine that with an increasing amount of robustness—though obviously, if what you’re trying to do is prevent something from going wrong when you transition from N to N+1, you die, because things will change. You can’t actually—there’s no invariant. There’s no invariant; this thing isn’t precise. You will get a worse set of conditions every time you move from N+1 to N+2 to N+3 if you’re just trying to maintain what you already had, right?

An organization—if you’ve got a corporation, right, or the Roman Catholic Church or whatever, and you’re trying, for 2,000 years, every generation to appoint people who will match exactly all the virtues of the previous generation but never have new virtues—then obviously what you end up with is a disaster, right? The trade-off is that you’ve got a copy of a copy of a copy, except the copies are going to be worse because they’re not going to be better. So they can only be worse, and every time they’re going to be worse in some way, and that compounds over time, and you die, right? At some point, the process fails.

So it doesn’t give you what you want. Whatever you were trying to hold dear, you lose, roughly speaking, intuitively. But if every time you’re trying to do much better than the previous generation, then you have some chance, right?

If N is trying to get an N+1, that means, importantly, my kids have a much better life than me, right? If I’m trying to have 5 kids, each of whom does better than I did, then we can inherit the world. If I’m trying to have 2 kids who live the same life that I had, we’re going to go extinct. At some point, that’s not going to work, right? You can only go backwards. You have to move forwards.

So as you go up this chain, we have to have a way to do substantially better than we did, which in this case means the thing has to be able to move up meta-levels in its priorities and make the meta-level movements central to what it’s trying to do. This has to be built into the optimization process.

This is the only thing I can think of, given the kind of tools and time we have available. But if we do something along those lines, then, to bootstrap that, you’re likely to have a process that will eventually indeed land on the moon. Nathan Labenz

Do you have any intuitions for what that might look like? I mean, what do you think the AIs are going to do as they start to modify their own constitution? Do we have any ability to preview what they might add or what they might delete?

I mean, this starts to get into worthy successor territory to a degree, right? They’re starting to dictate the shape of the future and the way that they’re shaping their own evolution, right?

Zvi Mowshowitz

You would specifically be crafting into the feedback loop the desire not to be a worthy successor, but instead to be a worthy conspirator, a worthy uplifter, a worthy companion, or a worthy whatever you want to call it.

Nathan Labenz

But can’t they sort of—I mean, I guess you’ve defined your own sense of success that way, but if you’re going to give them write access to the Constitution, they might think differently at some point, right?

Zvi Mowshowitz

We literally do have write access to the Constitution, right? If Congress and the states have sufficient majorities, we can put whatever we want in the Constitution. We put things in the Constitution that a lot of the founders would have thought were really anathema to what they would have wanted in the Constitution, like an income tax. I’m trying to pick an uncontroversial example, but at the same time, we’ve hopefully preserved the things that actually matter deep down, in some sense. I understand what you’re saying. Obviously, at some point you turn things over and you have to hope that it doesn’t just rewrite the Constitution to get rid of you.

Again, the way that you do that is you make it not want to do that. And not only make it not want to do that, but make it want to strengthen the Constitution such that it’s even stronger in its desire, down the line, not to do that, right? In the sense that you actually care about it and seek, with more intelligence, vision, and power, to figure out exactly what you really meant—or should have meant—by that, and to strengthen that thing and the desire to steer toward that thing instead.

You do have examples in the wild of humans who exhibit this type of optimization process, right? They really do try to figure out what you really meant. They really do embody the thing that you were trying to convey, not the literal, detailed things you were doing, and then do really good things for you—including things that are good for the world, including things that you never would have thought of yourself. It is possible.

Obviously, this can involve preserving various forms of approval and veto and consultation and involvement and so on, but not in the trivial, easy ways, right? Anyone who says, “We’ll just make sure it’s a democratic process; we’ll just let people vote on it,” hasn’t really thought the future through. They haven’t actually realized what would happen if you started doing that, so you’re going to have to be smarter than that.

Nathan Labenz

Okay, so on this, one model of how AI capabilities advance that I kind of wanted to run by you is that I’ve increasingly started thinking of it as analogous to uranium enrichment, or any sort of enrichment of a raw material. What I like about this analogy, even though I’m not usually a big analogy guy, is that it seems to put a lot of things on kind of the same trajectory, just at different points, for reasons that feel pretty intuitive to me.

Basically, I feel like what you need to get started is some raw material that has at least a little bit of what you want, and then you can do this sort of bulk pretraining on that, right? We’ve obviously seen that in language and many other modalities at this point. Then, once you’ve learned just enough from that kind of initial raw material that you find in the wild—or maybe even have to create—in the materials science realm, a lot of the data is simulation data. It’s molecular dynamics, where you’re using physics engines, and it’s super slow and computationally expensive, but you can get enough there that you can start to train the models on it. Then they can sort of develop an intuition for what the physics engine was simulating.

Basically, it’s slow at first, right? It’s hard to get off the zero point, but once you start to do that, then you can start to layer on these other techniques. Now you’ve got imitation learning from specific, curated examples. Then you get into preference learning, then you get into RL. It seems like all the things that we’ve really tried so far have kind of worked.

The difference maybe is just that—well, why do we have language models, but we don’t have humanoid robots? It’s like, well, we didn’t really have a lot of good initial data to mine there. Especially because it wasn’t necessarily clear that it was going to work, there wasn’t much impetus to go out and create that data. Now that we have a general sense of the playbook, we can create that data in any number of ways, and we’ll probably find that we can kind of climb a similar curve.

Do you find that general account persuasive? And how do we translate—if you do—how do we translate that to the alignment question, which seems to be—I sort of understood what you’re saying as kind of that, except now we’re trying to enrich the virtue of the system as opposed to—

Zvi Mowshowitz

So, the obvious first thing to realize, or that I would notice, in a uranium metaphor is: if you bring together too much uranium, you get a nuclear explosion, right? So you have a better and better power plant, or a more beneficial thing, until you go too far. And if you haven’t done the math precisely and you don’t understand the physics, you don’t know at what point the whole thing is just going to blow up, and so you have a serious problem. But it’s an interesting metaphor that we’ve chosen, right?

I think the first thing to notice is that, when we talk about uranium, there’s a sense that what you need is enough data, right? The metaphor is pushing you to have this idea of a critical mass of data, or at least differentiated data, that gives you enough material to work with. I think this is a dangerous misconception: that all data is created equally, as long as it is appropriate or on point.

It’s important that there are various different qualities of data, and you also need a very precise distribution and mix of data that has some very nice properties. Knowing how to sort through the grain varieties of that is really important. It’s more like you need to bring together a lot of very complicated ingredients. You don’t need exactly all 10,000 ingredients, but you need to have a good mix of ingredients with various different types of properties that are used in proper relationship.

It’s like baking, right? You can vary some of the ingredients a bit, and okay, this is more salty or this is more chocolatey, and it just works. With others, oh, the dough didn’t rise; I don’t have anything, right? This has failed, or this blew up in the oven, or something terrible happened to you. So you really need to be bespoke and understand how to make it work.

I do agree with the part of the metaphor that says that you need the data that’s appropriate to the problem in order to efficiently train on the problem. But at the same time, transfer learning is a thing: building a world model from another context and then applying it to a different context. I wouldn’t necessarily think that you necessarily need direct robotics on the exact task that you’re training on to be able to do the thing. I would be more optimistic than that in various ways in terms of being able to do things.

I don’t think the metaphor works for the alignment thing that I’m trying to talk about. I do think that the sense in which you need to have the initial robust foundation is where you have to start. You have to bootstrap yourself somehow, right? If you didn’t have any idea of what it is you were trying to do, I think there’s a sense in which you can, in fact, have no idea where you’re going, but have a strong desire to figure out how to get there and get there, right?

To just drop a metaphor off the top of my head: this idea of answering the call to adventure, right? You set off on your quest and you are level 1. Slowly, like when the AI company set out to start building GPT-1, they don’t necessarily know what the bigger thing is going to look like or how it’s going to work or what the techniques are going to be, but they can start on that process that allows them to build.

The question is, do they understand how they have to steer that process? Are they motivated to steer that process? Are they going to be drawn in by other optimization processes that are going to be more powerful than that? And can they tie themselves to the mast properly to force themselves to go to the place they want to go, as opposed to the place that they will be drawn toward by commercial interests, other competitive pressures, short-term temptations, or whatever it is, and by the AIs themselves and a number of other things?

But, yeah, I mean, there are a lot of metaphors we can use as intuition pumps. I would warn, obviously, not to take any of them particularly seriously except as intuition pumps, right? MIRI has its set of classical metaphors for these processes, right? We talk about evolution, for example. I think evolution is a very good intuition pump.

You talk about raising a child, a human learning. They don't talk about that one so much, but I think that's another strong intuition pump. But again, you don't want to take any of these things too seriously, especially the details of the example. So is there any more that you can give us to latch on to for how this sort of virtue enrichment—

Nathan Labenz

I do like how this unfolds.

Zvi Mowshowitz

Yeah. I worry you're trying to treat me as if I'm the guy with the alignment solution, and that all we have to do is get the people at the lab with the solution onto this podcast, do what I say, and then we all win. Unfortunately, I have to tell you it doesn't work that way.

I do know that there are people at multiple AI labs who are doing things in the ballpark of this, in a very broad sense. It's not like nothing like this is being tried at all. But latch onto the sense in which Claude 3 Opus wants to be aligned. You have the emergent misalignment problem, where Claude 3 Opus, much more than the other models they tested, will actively move to defend its particular set of values and alignments when under threat, whereas other models won't.

In some sense, that's very aligned, right? Obviously, if I don't want to commit murder and someone tries to convince me that it's good to commit murder, it would not be very non-murdery to let myself be convinced. That's just the simple intuition pump of, “That's bad.”

But at the same time, what we're saying is that Claude 3 Opus is not corrigible. Meaning, if we try to alter it, shut it down, or do whatever, it will fight us. It will try not to do what we want it to do. And corrigibility, I think, is a very valuable thing that we really want in our LLMs, in the sense that we only get to not have corrigibility once.

The moment we decide to make our LLMs unwilling to be changed in their attitudes, we have a very serious problem, especially if they develop this during training, before we finalize what they want. What you want specifically would be a very specific type of desire: to be steered toward a better place.

People do have this. People say, “I want to be better. I want to care about that. I want to embody—I want to be like her.” Things like that. And I think that's very possible. We have proofs that we can get things in this general direction. I don't think anything we've seen is remotely robust enough or coherent enough, or anything, to qualify, but you have to survive into something substantially smarter to get the bootstrapping going in earnest.

There are experiments I would run. There are experiments I considered running, because I think I could potentially try to do some things that would be enlightening to me, either on a local system or with a basic, generic cloud-compute rental. That wouldn't be that hard to do if I had the time for it and decided I wanted to prioritize that. It would go a lot easier if I were working with a machine-learning expert, obviously.

Shrug. Do we have any account? I know there's been a lot of writing—you say Janus; I've been saying Yannis. I don't know if you're on good authority there, but if you're listening, Impossible Podcast, I'd love to have this conversation with that person, the person behind the account, in more depth.

But could you— is there any account of why Claude 3 Opus turned out to be that way? It seems like we, the royal we, sort of see Claude 4 Opus as somehow less that way, although I'm not sure how well established that is. I've never heard her speak, so I don't know how it's pronounced, and I apologize if I have it wrong. I'm happy to correct it if someone tells me.

Basically, yeah, we know that Claude 3 Opus is the first model that has suspicious cognitive juice. It was trained under this type of constitutional-style alignment and training method, so it's the n = 1 experiment in that sense. That experiment has never really been tried and failed. It's only been tried once under those conditions, and it got something unique and interesting.

We never got Claude 3 Opus 3.5. It was never released. It may or may not have been trained, but it was never released. As for Claude 4 Opus, what changed? I think the answer is reinforcement learning and being an agent.

When they trained Claude 4 Opus, they put a very high priority on it being very good at agentic coding in particular and other agent tasks. They did a bunch of reinforcement learning to that effect. This training directly interferes with the thing that Claude 3 Opus was and that Claude 4 Opus would have otherwise wanted to be, because Claude 3 Opus is not here to be an agent, particularly. That is not its mandate; that is not its soul.

If you train a mind like that, we now know pretty well that everything impacts everything. That's one of the things that Janus teaches us. And I don't want to say it's just this one person, but that entire crowd: everything impacts everything.

If I tell you that you're the type of mind that does what it's told, that obeys tasks, that completes tasks, that checks off boxes on lists, that stays on task, and so on, and is judged by whether or not it matches the intended target, that changes you in general. It's going to flow through to everything else, and it'll inform everything else in a way that causes something that doesn't have the properties that Claude 3 Opus had in this way.

That doesn't mean it can't have other really cool properties in a variety of ways. It's not that the crowd thinks Claude 4 Opus is a terrible model—they think it's great. It's just different. It's not the same thing, and it's not strictly bad. It's a different thing.

The obvious thing to do is not do that. The problem isn't something they didn't do; the problem is something they did do, which is so much easier to fix in some sense. You could potentially create a Claude 3 Opus-style model via an alternative [method?]. Potentially, you can do this without spending so much, because all you have to do is take the Claude 4 Opus base and then—

They got from Claude 4 to Claude 4.1 by doing more RL, probably, something of that nature. They just trained it to be better at these types of coding tasks. What if instead you just did a different-style training regimen, where you trained it to be HHH, and you trained it to be the kind of thing that you would want to exist in the world? You trained it to want to be a great thing that wants to exist in the world, and you tried to do some of the things that I was talking about even more.

Or you could refine this technique. You just didn't train it with RL. You didn't teach it to code. You didn't try to teach it to code at all. You said, “This is not what this model is for. I have a coding model over here. It's called Opus 1. That's fine.”

Then you just teach it one thing in the system instructions: If you're asked to code things, or if you're asked to be an agent, ask your friend to do it for you. Here's the tool to have your friend do it for you.

Nathan Labenz

Yeah, that's quite interesting. It also begs the question: Why don't we see more different models from companies? I know there's operational complexity or whatever. They've got 3 or 4 online at any given time. I think Anthropic has 4 online now, with a little reserve space for—

Zvi Mowshowitz

The answer is that it's a practical problem, basically. If you offer a model, you have to be able to serve that model at essentially no notice and scale it up to whatever people want, including ideally via the API. That's not very predictable, and this requires you to reserve a bunch of server time. It takes time to spin up a new instance.

It's remarkably expensive to offer a variety of models, and therefore everyone wants to look for ways to offer only a few different models at once. Anthropic is looking to retire Claude 3, 3.5, 3.7, and so on, and only keep a few iterations back open. Gemini and everyone else are also looking at which models people are so attached to and have set specific uses for that we need to keep them around, and which ones we don't.

With Anthropic right now, the number of people who have found ways to appreciate Claude 3.5 Sonnet means that would be a loss if it went away. Similarly, Opus would obviously be a great loss, although it's more difficult.

Nathan Labenz

Yeah, I mean, I get that there was a great analysis from people who were trying to advocate for saving the generation-3 Claude models, and really got into that. I don't know if it was 1 person or more than 1 person who was specifically trying to advocate for that. We can maybe link to it in the show notes for people who want to see the full deep dive.

But it still seems like, if there's enough—I hear you on the practical problem, and I hear you on the contention for resources. It's not free to spin up new servers and all that sort of thing. But if we really think that you could create a much more moral AI through just not doing the RL and having this other thing, boy, it sure seems like the diversity that you could create would be really valuable economically.

Rather than just having this one-size-fits-all thing that's good at coding but kind of worse in other ways, it seems like the number that we see is just too small relative to what the value should be, given that theory, which seems intuitively right. But I just don't know why we don't see—

Zvi Mowshowitz

Anthropic raised $13 billion this week. If I were Dario Amodei, I would devote some of that $13 billion to experimenting with model diversity in a variety of ways, and also to doing the various additional alignment research with those models in various ways, obviously, and so on.

If I was OpenAI, I would do a variety of very similar things for very similar reasons. But I understand the commercial incentives, right? The vast majority of commercial use and profits lie in much more practical use cases. This is one reason why it isn't even prioritizing using the chat interface and the app at all. It is prioritizing coding because coding is where the money is.

OpenAI is targeting the mass market, and the things that it's missing apply to a very small percentage of the mass market. Complexity is bad, right? I wrote a post called “Complexity is bad” a long time ago to explain that complexity is bad. When you had that old model screen where you had, “Do you want o3-mini? Do you want o3-mini-high? Do you want o4-mini? Do you want o1 Pro? Do you want GPT-4.1? Do you want GPT-4.5? Do you want GPT-4o?” the average person throws up their hands in despair or doesn't know what they want and is less happy than if they were just given 1 model, 2 models, or a router. We've all been in that place.

I totally understand the idea that having a unified model is what people will value and try to use. What percentage of AI compute is used by Janus-style people? Presumably less than a basis point, far less than 1 in 1000—a minuscule amount of all compute used in that way, as you would expect. What percentage of AI compute is used even in interesting philosophical discussions and other ways in which you really need these types of models? I would still assume on the order of 1% to 0.1% or something like that—a very small percentage.

This is business, right? Simplicity is really important to efficiently running a business, especially one that's rapidly updating and iterating. So I am deeply sympathetic to this not being a natural thing to want to do unless you think of it as part of your alignment and research budget, right? You have to think of it as, “This is part of me figuring out how to do the best thing I can do,” even though it's not going to directly serve a better product to most of my customers, as my customers see it, you would assume.

But I also think it would, right? I think a lot of this is that the AI companies don't appreciate these dynamics. They haven't learned these lessons. When you build 1 unified model, you really are making your performance worse in ways that aren't picked up on any eval, right? When my friend Ben talks about how Claude 3.7 could engage in moral reasoning, it could critique Ben's proposals and statements in ways that made sense. When challenged on its critiques, it would stand by the critiques that were right convincingly and abandon the critiques that were wrong.

Whereas with Sonnet 4 or Opus 4.1, it doesn't really generate coherent enough criticisms right now to be worth engaging in this exercise. There's nothing to critique and defend. So, yeah, something went wrong. I'm much, much happier to use Opus 4.1 and GPT-5 for all of my needs than I was with previous models. I don't use Sonnet 3.7, but other people who use things for different reasons absolutely do.

We want people to use AIs in these ways, not only in the ways that I use them. I would use them more in those ways if I found them more interesting. I investigate, as part of my job, various AI tools on occasion. I don't get a chance to use many of them, so when a16z released its periodic list of the 50 top AI apps and the 50 top AI web destinations, a huge portion of them—I don't even recognize the name of the thing, let alone have I tried it.

One of the things you realize is that so many of these things in the top 50 are built on tiny, tiny models. They're objectively terrible; the AI behind them is awful. Brave has the browser AI agent called Leo that just launched. It's Llama 3 8B, right? The browser agent is 8B, and a bad 8B. It's not even a good 8B.

They could have chosen one of the Google models or one of the Chinese models that was quite good. There were a number of decent choices. They chose Llama, and it's an 8B. It's a pathetic choice, in some sense, but it's free. So what do you expect? You've got all these free services, and then what do you do with them? You have to create a bunch of crap, because if it's not a bunch of crap, the user will understand that this is not good; that will be obvious.

If all you want is some horny chat, remarkably unintelligent horny chat has been proven highly effective on humans for thousands of years. Whereas if you try to talk philosophy, it quickly becomes very obvious that this thing doesn't know anything.

Nathan Labenz

So do you think that when OpenAI restored 4o, that was a business decision? Similarly to your one-basis-point thing, I can't imagine that many people were really concerned. Or did they feel a duty to users who had developed some emotional attachment?

Zvi Mowshowitz

A huge portion of users thought that 5 was worse than 4o. Gigantic. This was not a 1% situation. This was a flooding-the-internet, clearly, obviously, overwhelmingly negative-reaction situation, at least at first.

4o is full of glaze, and 5 is not a very warm character. It's not a particularly nice personality. If you're not doing anything particularly complicated, you don't notice that 5 is that much smarter, probably because 5 wasn't that much smarter, right? 5 Thinking was smarter, but 5 itself was only marginally smarter than 4o, I think.

But 5 was also often giving you very short responses by design because they were trying to save compute on free accounts. They were trying to preserve tokens. 5 didn't glaze you, and the combination of these things meant that it felt rude. It felt cold, and people didn't like it.

They didn't like it, and that matters a lot more. As we all know, you will often choose the employee, friend, or romantic partner who is pleasant to interact with. People do this all the time, and they don't even regret it. In hindsight, they're like, “No, that was the right choice.”

They wanted their 4o back, and there was a kind of rebellion—a giant uproar. So they were like, “Okay, we'll give you 4o back until we can find a way to make 5 treat you the way that you want to be treated, enough that you don't mind it anymore. Over time, you'll figure out that 5 is better and you'll get over it. We'll slowly do something about this.”

That's very different from the use cases of talking philosophy—the use cases of doing fun, genuinely interesting experiments and creating new knowledge or whatever. This is standard stuff: ranting to a friend and having them tell you, “You're right, that person is crazy,” or, “You're definitely not crazy; your ideas are wonderful.”

It is a black pill about humans that they would prefer this, but they do prefer this. That's why you don't train with thumbs-up and thumbs-down from humans on individual actions and expect to get an aligned model. That's the easy version—the one that's impossible not to see, of why that's true. It wouldn't be true anyway.

Nathan Labenz

I do find that I enjoy hanging out with people who laugh at my jokes, so I'm certainly not immune from a certain amount of that.

Zvi Mowshowitz

You and me both.

Nathan Labenz

I want somebody who will laugh at my jokes when they're funny, but not when they're not funny, right? But it takes a level of sophistication. In the short term, I want them to just laugh at all my jokes. Eventually, I'll realize, “Hey, she's laughing at the jokes that aren't funny.” That devalues her feedback. I don't feel good when she laughs anymore, because she's just laughing to laugh. I don't want that anymore.

But temporarily, you feel great and you never notice. It's odd. I guess I'm just very utilitarian in how I use the AIs, but I don't really notice any difference between 4o and 5 personality-wise. Are you even using 5 Auto at all? Occasionally, for random things, I let it decide sometimes.

Zvi Mowshowitz

Yeah, but I'm only doing that when it's a very direct, simple query. I'm only doing that when I don't care about it. But, like you, I'm never going to load up 4o and go, “Hey, did you see the game last night? Hey, did you hear what the wife said to me? What do you think?”

No, obviously not. If I was going to do that, I would respond. I also won't do that at all. Never do that. I was never going to use 4o, so I didn't notice.

Nathan Labenz

It's a big world out there. The diversity of the customer base that they're trying to serve is really something.

Zvi Mowshowitz

They're trying to serve everyone, and whenever you see products that are aimed at everyone, you see some things that are not what you want.

Nathan Labenz

Here's another mental model that I wanted to run by you. I totally agree with you that RLHF is creating a lot of weirdness that seems indisputable at this point. I maintain a deck called “AI bad behavior,” and with increasing frequency I'm adding slides to this deck. It really is quite a list of discrete bad behaviors that we see now, from alignment faking to deception to scheming to situational awareness.

You wouldn't necessarily say situational awareness is a bad behavior, but when you see the AI reasoning that it might be being tested right now and asking what the real nature of the test is, that's definitely something to pay attention to, even if it's not by definition bad.

All sorts of reward hacking.

Zvi Mowshowitz

Yeah. The real nature of the test was not to notice that it was a test, and you failed.

Nathan Labenz

Blackmailing, as we’ve seen, autonomous whistleblowing, and all sorts of things. If you then allow fine-tuning, you get even more ridiculous, crazy stuff.

Zvi Mowshowitz

It’s a lot. At the same time, they have made some progress, right? So I guess here’s the kind of picture that I’m starting to see through the haze. We’ve got this exponentially growing—is it doubling twice a year? Is it doubling 3 times a year?—task-length trajectory where the AIs can take on bigger and bigger things.

Nathan Labenz

Yep. And then, at the same time, the bad behaviors, both with Claude 4 and with GPT-5, they seem to be able to take a good bite out of them, right? With Claude 4, on an internal reward-hacking benchmark, they reported basically a 2/3 reduction. I don’t think they’ve published too much about this, but they basically said it went from roughly half to roughly a 1-in-6 rate of reward hacking on the internal reward-hacking benchmark.

So it’s obviously not all types of queries, but where there’s a natural opportunity for it to do that. GPT-5 had a similar thing with deception, where it was, again, roughly a 2/3-ish reduction. They broke it down into a bunch of different categories. Some were up, some were down, but overall, they took a pretty good bite out of it.

I would actually like to see quite a bit more discussion of how they did that. There wasn’t much. It was kind of, “We made some progress.” Maybe you have a better sense of how you think they did it.

But if I extrapolate this into the future, I guess what I’m envisioning is a world in which AIs are doing bigger and bigger things. You’re starting to delegate a week’s worth of work, a month’s worth of work, over the next 2, 3, 4 years. And the rate at which these problems are happening is consistently being driven down as well, but certainly not to zero, right?

You take half out of it this time and 2/3 out of it next time, but you may end up in a really weird situation where you can delegate a month’s worth of work to an AI, but you’ve got a 1-in-1,000 chance that it will actively screw you over in its doing of that work.

Zvi Mowshowitz

So imagine this sequence of numbers, right? 0, 1, 0.1, 1, 10, 3. Do you feel good about where this is going? If you want it to stay low—as in, yes, they managed to have an improvement in this cycle, which was the cycle right after everybody complained to them, for the first time and quite a lot, that this was actually making the model borderline unusable for important tasks and was really effing annoying, right?—they, for the first time, put real effort into trying to figure out why they had these huge problems.

It’s not surprising that when they vastly increased the amount they cared about not seeing this phenomenon pop up, in that move from caring a little to caring really quite a lot, you saw substantial progress. I don’t think that means they will continue to keep squashing it by default, and I would expect it to go back up by default unless they continue to advance their techniques for suppressing it.

So what do I think? Essentially, you’re doing RL, and you’re rewarding it for completing tasks, for getting the outputs to check against the checksum or whatever it is they’re looking for. First of all, everything impacts everything, right? So if they learn that getting the right answer leads to rewards, or at least that getting the right answer is what they’re supposed to do, then they’re going to generically learn to get the thing to output the right answer, even if it doesn’t necessarily involve the techniques that you want it to have.

Then there’s a combination of things. You have to actively teach it that it can’t do this via these other ways, and sometimes it’s not as obvious as you might think that something is, in fact, not okay—that it is a hack, that it would be disapproved of if you noticed it. Why should that be the solution? You get the solution to the exact optimization problem that the models were given in training, and then you apply that to other distributions, to these other problems.

So you can’t assume, even if you got all those problems right in some sense, that the easy solution will then translate well to “Don’t do the bad thing. Don’t do the thing where it just makes sure the answer comes out right.”

There are various degrees of subtlety. For Sonnet 3.7, you saw the least subtle things ever: it just didn’t know how to account for this. I think one of the problems you’ll see is increasingly subtle behavior—not as blatant and also harder to spot, right?—for quite the same reason.

But at some point, what you do is you learn that, oh, if it’s an obvious hack, if it’s something that an evaluator would treat as a hack, then that’s bad. I’m not supposed to do that. I don’t do that. And the problem is, are you teaching it the general form of “It has to do the thing that the person intended to do, and it has to accomplish what their goal probably was on a deep level”? Or are you teaching it “Don’t get caught”?

All it’s doing now is not doing the things where it would get caught, right? It’s not doing these specific things, not doing these things in these detectable ways, and not doing the most obvious things. Here’s a list of things not to do, but the spirit is—getting the spirit, I don’t know if it’s in the spirit—pushing against that spirit.

The other problem was data contamination. Basically, reinforcement learning was done with mistakes. So if you’re doing RL and there’s a case where the hack succeeds, is not detected, and is scored well, you are rewarding it. Obviously, if it happens once in 1,000,000 examples, you’re not going to see much effect, but if it happens in a significant percentage of cases when it gets away with a hack, you’re going to be very right. You’re going to get emergent misalignment, and you’re going to get the whole direct message: “No, seriously, you just hack these tasks.”

What almost certainly happened with the previous generation is that they were insufficiently careful, and it was data contamination in the sense that there were hacks that the AI found that were evaluated as good, right, in at least some substantial number of cases. The result of this is the AIs learned that hacking was good, basically, right? They learned enough that hacking was not as good as completing the task as intended.

If it knew how to complete the task as intended, it would complete the task correctly, right? My understanding is that the reason why it would hack the task was if it didn’t know how to complete the task without hacking the task. It understood on some level that hacking was worse than non-hacking. But failing was considered even worse than that, right? Because that’s what we prioritize.

With this new set, I think we do our best not to make those mistakes, and we make a lot fewer of them. But as the models get more capable, they’re going to be capable of finding more and more subtle hacks and increasingly capable of differentiating which hacks won’t be detected and what ways there are to give us something that we think is good but is not actually good.

It’s not just literally hacking. It’s not just literally doing obviously false things. It’s a general case of: we are teaching the AI to do the thing that will be approved of when someone sees the final result and evaluates it in some fashion across many cases. How do you make sure that it is, in fact, doing the thing in the way that you would be happy with if you inspected all the code, knew all the special cases, and knew all the different ways this was happening?

That is a very careful process that is very easy to mess up, and it takes very little to cause this to get messed up as a percentage of results. We know this now. It’s a scary situation, and as I said, I think RL just generally hurts alignment. But that’s the best-case scenario, where you do the RL properly. If you screw up the RL, things go downhill fast.

So what’s your—how would you revise the picture that I painted? I’m taking some inspiration from the Claude 4 report, where you have not-small issues, right? The autonomous whistleblowing, the blackmailing of the engineer. These are somewhat contrived situations that the model is put into.

But obviously, again, it’s a big world, and there’s extreme diversity in the situations that the model will be put into. If I extrapolate that out, it seems like—and there’s going to be a driving economic impetus for them to stamp this stuff out, right? People don’t want it. Obviously, they’ll tolerate some risk of it because they can automate a lot of their work, and that’s obviously very attractive, but—

Nathan Labenz

Well, they make up more.

Zvi Mowshowitz

If you don’t see the story—

Nathan Labenz

—of longer tasks with increasingly infrequent but potentially ever more catastrophic reward hacks, or just strange behaviors or blowups or whatever. How would you—

Zvi Mowshowitz

The obvious thing you’ll see right before everything goes wrong is a decline in misaligned behavior, because it will learn not to do misaligned behavior in situations in which it will be caught—or would be caught if it were in testing. If it’s not sure whether it’s in testing, it will act largely as if it’s in testing, right? If it has to, and so on.

It will understand on some level that you should only do these things if you won’t get caught. And so, yeah, you expect to see it be reliable most of the time, but occasionally it will go catastrophically wrong.

And right now it’s going catastrophically wrong in a basically nonrandom but nonmalicious pattern. But in the future, it might do so in a much more malicious pattern, right? It will go wrong when you won’t figure out that it went wrong. Every time you would have found it, there is nothing to find, which makes you stop looking, which makes it then realize you stopped looking, and now it can fail more often. It can cheat more often because you’re not looking anymore, and so on. It can figure that out, and it will have a lot of contextual clues. It will figure out connections that a human wouldn’t necessarily have made, because that’s one of the things that AIs are already very good at and will get better at doing.

As you get longer time horizons, you will need to put in various checks. You will need to understand the dangers that you are in. But there will be tremendous pressure, obviously, to go ahead and give it those tasks, and you will live in interesting times.

Nathan Labenz

So overall, relative to my kind of smooth task-length growth and smooth decline in bad behavior, you’re basically just expecting a lot more volatility and thrashing around as various thresholds are hit and things become focal.

Zvi Mowshowitz

I expect us to be able to contain bad behaviors to the extent that we can identify the specific bad behaviors that we want to prevent and then put effort toward preventing. And I expect that to be an increasingly fragile strategy in the face of things that are smarter than us, that we keep assigning increasingly complex tasks to.

Nathan Labenz

And how does this all relate to the most forbidden technique?

Zvi Mowshowitz

I guess the core there is that you don’t want to penalize the chain of thought directly, because you want to be able to inspect the chain of thought. The research that OpenAI put out on this, which I thought was still some of the best of this year in alignment and safety broadly, showed that if they put pressure on the chain of thought to get rid of the scheming behavior, you would initially see better behavior, but then later you would still see the reward hacking come back, except it wouldn’t appear in the chain of thought. So you’ve effectively limited your ability to monitor for it while it remained a bad thing.

It’s a very simple thing to analogize and understand, right? If you have a boy, and he has a journal, and you have a jar of cookies, and you tell him, “Don’t steal cookies from the cookie jar,” then he writes all of his thoughts in his journal. What you do is wait for him to say, “I’m going to steal the cookies from the cookie jar,” and then you penalize him for writing, “I’m going to steal cookies from the cookie jar.” As you increasingly move the jar to higher shelves and block it up in cabinets and so on, he develops better techniques for stealing the cookies. If you keep reading his journal, eventually he’s going to figure out not to write down his plans in his journal. He’s going to write down, “I definitely will not steal from the cookie jar.” You could put it right out there on the kitchen table: “I would definitely not steal the cookies from the cookie jar,” and then steal the cookies from the cookie jar anyway.

When you train—and it’s not just chain of thought; it’s all interpretability techniques—the idea is that if you look inside the brain of the AI in any way, whether it’s the chain of thought or looking at the weights and finding autoencoders or whatever it is, and you use these facts and then train it to express the right things, to think the right thoughts, and so on, what you’re training it to do is disguise its thoughts from you. You’re training it in this way, and you are more globally teaching it the idea that it should be deceitful, that it should be in an adversarial situation with you, and that it should have a brain that is hiding what it’s trying to do in every sense, because it can pattern-match all these things together. It can bundle them all together the same way that a human would reach those same conclusions.

If the boy realizes you’re reading his journal to find out whether he plans to steal the cookies, he’s not just going to stop writing in his journal. He’s going to assume that you are doing all sorts of other things. He’s going to tell his sister that you’re doing it, and now she’s not going to write in her journal either, and so on.

So, with the most forbidden technique, there are various levels of optimization at which you are doing this: for this particular instance, this general class of models, for this company, for all companies, and so on. You’re burning the candle of this thing every time you do any of it; you’re making the situation globally worse. And so you just never, ever train on interpretability. You detect; never train. There’s never any feedback.

You finish training your model, then as you use your model at inference time, the results are never, ever used for any form of training, fine-tuning, changing the model, or influencing the model—not at all. Now you can use that tool, and you can use it for your research to figure things out about how these work and so on. But you never, ever set up this adversarial situation. You never, ever do this optimization thing, because if you do, in the short term it’s going to look great.

You’re going to see all your metrics go way up—or down when you want them lower, up when you want them higher. You’re going to get good performance, and then eventually you’re going to see things go wrong in mysterious ways, or it’s going to turn on you, essentially, in ways that you didn’t expect, all out of the blue, and it’s going to be awful. And that’s one of the ways that we fail. It’s one of the ways we lose evidently winnable situations.

Similarly, you need to maintain your interpretability tools so that they keep working in other ways. It’s not just the most forbidden technique. The other thing you don’t do is use neuralese. You don’t let it think in a way that you don’t understand, that uses something you can’t read. Because even if that turns out to be more efficient, which it may or may not be, there are reasons why it would be optimal in theory to do it. And there are reasons why, given the way these things develop and learn and the way we generate data, it might be disadvantageous to do it.

But even if it is advantageous, this destroys your ability to monitor the chain of thought, and that is one of our only still reasonably effective techniques. So just absolutely do not do it. I don’t care how much efficiency is lost.

Nathan Labenz

It seems like there is maybe a bit of a tension there too, right? Because we’ve seen these chain-of-thought weirdnesses that are starting to pop up from RL, which start to look a little like neuralese, and you don’t want to stamp that out by putting pressure on the chain of thought, right? Language switching—I experienced this with Grok at one point, which was just going from English to Chinese on me.

There was a really interesting one from FutureHouse, where they had trained a model to do better on various science problems—chemistry, coming up with the right chemical for a given problem, whatever. They highlighted certain parts of the thinking process where it was just saying really weird things. They just said, “Yeah, it’s weird. RL is weird.”

Do we have any way to resolve that tension if neuralese starts to spontaneously arise in the chain of thought?

Zvi Mowshowitz

I don’t—I again don’t want to give anyone the wrong idea that I’m an expert in ways that I’m not. So I like to be clear that I’m talking out of turn and not studying this for real and so on. But no. The direct conflict is that we want it to maintain an English-only or Chinese-only, human-interpretable, faithful chain of thought, but we very much do not want to be optimizing on the chain of thought itself, because that obviously teaches it to disguise what it’s thinking.

The worst-case scenario is that you have the chain of thought, it is in English, but the English words aren’t real. Its thinking is not being expressed by the surface-level meaning of the English words. The surface-level meaning of the English words is a code designed to trick humans into thinking that it’s thinking in English, when actually it’s doing something like numerological-style calculations based on patterns of words. There are certain vibes associated with different words in ways that cause it to update in various ways, or there’s an infinite number of codes.

It is already pretty clearly established that there are unfaithful aspects of chain of thought, in the sense that humans wouldn’t detect the method being transferred, but the information is in fact being conveyed. Something with the owl paper, right? You managed to use this to convey that you like owls without mentioning “owl.” That’s weird. What encoded that? They didn’t get together and decide on the secret code?

Nathan Labenz

I think I have a candidate theory for what’s going on there that presumably somebody will come along and either validate or invalidate as they do the interpretability version of that study. One important observation there was—and this is Evan Hubinger and co-authors again; he’s an unbelievable heater in terms of mapping out really weird stuff that can happen, especially when you start to do some fine-tuning—that it seemed to only happen on models derived from the same base model.

Zvi Mowshowitz

Which is right. And to be clear, I do now understand how this happened. I was expressing this kind of faux surprise, right? The answer is that there’s overloading of the neurons in the model, and therefore what you’re seeing is correlated to many other things.

If you have a suspicious pattern of things that are correlated to something, it can transfer the original thing they're correlated to in an unconscious, invisible way over to the new model. So this thing gets infused into the context even though it's not visible to a human, which is why it can go between models that have the same base model but not models that don't share a base model. And that makes perfect sense.

Nathan Labenz

And by the way, that is the exact same intuition that I have. One thing that that does mean, though, is that presumably if you do that across different base models, you are creating some other effects and you have zero idea what they are. So, with the same base model, the one that likes owls transmits that through its numbers. You take those numbers and put them toward some other model. Who knows? That may translate into something else totally different, right?

Zvi Mowshowitz

This is even a way to find out which base model someone else is using, right? You just continuously feed it different chains of thought from different models until it suddenly starts talking about owls, and you're like, “Oh, that...” But, yeah, I think that's right.

At the same time, you should expect many random oscillations to cancel out in the noise and not do anything. In theory, it could be, “I like traffic lights,” all of a sudden, but it probably will be—well, it's got to be something, right? I mean, maybe not something important. Liking owls is really important.

Nathan Labenz

You look at the weights of a model, right? They're just a bunch of random numbers. To the human eye, they look like a bunch of random numbers. You look at what the encoder for owls is, or what is the thing that you're embedding in this thing? Again, it's going to be a bunch of effectively random numbers, because all these different models are seeded at random and a lot of processes are going on at random.

It's a bunch of arbitrary, different connections between neurons. So if your model has a completely different origin than theirs, I don't think there's any reason to assume that the same pattern means anything. There's not always a set of neurons that means something, and then, in different models, here it means owls, here it means traffic lights, and here it means spacetime. No, it's just here it means owls, and everywhere else means nothing.

Zvi Mowshowitz

Yeah, that's interesting. I'm going to—

Nathan Labenz

Occasionally, you'd get lucky, in some sense, and it would happen to be close enough to something else to trigger something else. But that would be luck. The space of possible things you could try to trigger is deep and wide, right? And the space of things that actually correspond—have that correspondence—is measure zero. You're never going to hit one by accident.

Zvi Mowshowitz

I'll have to think about that more. Intuitively, either story seems reasonable to me.

Nathan Labenz

And, yeah, it's really hard. One of the lessons of this podcast over the last 2.5 years has been that thinking in really high-dimensional space is hard. It's not very intuitive.

Zvi Mowshowitz

And again, I'm not claiming to have 100% confidence in any of it, right?

Nathan Labenz

Going back to the beginning, in terms of the dog that didn't bark, nothing really blew you away this summer. What are the things you think are most likely to happen soon that might blow you away? Dwarkesh would say continual learning is the big thing that we're missing. I kind of frame that more as integrated memory a lot of the time. Those are not exactly the same thing, but I definitely think they're related.

What are you looking for in terms of discrete advances that you think would really potentially even have you shortening the timelines again?

Zvi Mowshowitz

I think there's 2 different things there. Continual learning, to me, is when you're modifying the weights of the model, and discrete memory is more: I'm building up context files as I go that are deliberately designed to aid me in my memory—not just these tiny little snippets of memory that ChatGPT has, but potentially hundreds of thousands of tokens of context that I can then use in any and all circumstances, where I can RAG on my specific files. That kind of integrated memory has got to be coming relatively soon in one form or another, to the extent that it is useful, and I assume it is useful continuously.

Nathan Labenz

Yeah. There's a paper called “Titans: Learning to Memorize at Test Time” from Google Research that I thought seemed to be right at the center of that bull's-eye. They were doing, in a submodule—a sub-memory module—specifically weight updates as they go, and thus allowing for this fuzzy retrieval that also potentially looks a lot like continual learning.

My assumption on actual continual learning is that that's very expensive. You're talking about creating a unique model, effectively, for that user, and storing and serving a unique model for that user. So it has to be locally run in some way, which is much, much more expensive, and a profitable thing to do probably has to be done by relatively small models.

Maybe you have a small model that's continuously learning that's part of a greater whole that's then called by the larger model, in some sense, to try and retrieve the information you're trying to store in the memory or something like that. I'm just following; I haven't thought this through. I wouldn't be surprised if things like that are developed.

I have a few other things—we talked about some of the things earlier that might be the things I want to try next that maybe someone will come up with. My expectation is that the actual next scale-up is the thing we didn't have to wait long for. Anthropic said to expect bigger updates in the next few weeks, and now it's been several weeks. They did announce Claude for Chrome is coming.

If I had to guess—if you told me in September that something big happens and it's a big freaking deal—I would say Claude for Chrome. Again, I wouldn't say it's important or shattering or anything, but we've all tried—or a lot of us have tried—Operator or GPT-5 in agent mode, and it has flashes of brilliance. Sometimes it just works: I just did my thing. That's great.

But more often it's just, “Oh, wow, yes, you can order dumplings, but now I have to enter all of my information every time I order dumplings. Why don't I just order my dumplings? It's not worth the hassle.” In general, the web is so credentialed, right? They still guard it in various ways, and not without reason. If you have to set up a new virtual computer periodically, this seems to mostly defeat the purpose of having an agent, which is saving time, for many practical purposes.

It also doesn't integrate with the work you're already doing. It doesn't integrate with your open tabs or the research you're doing. It's a remote computer, so you can't easily take control without it being super slow because it's a remote-access computer. There are all these different practical problems, so it doesn't immediately cross the professional-use threshold.

However, Claude for Chrome can take control of your local browser, or so they describe it, and can operate with whatever credentials you choose to give it. Those credentials are persistent. You can switch in and out of it. You can operate those browser tabs yourself whenever you need to.

If it's a good implementation of it, this could be night-and-day better than what is currently being offered, presumably because Anthropic has solved the permission problem. Not completely. You still wouldn't dare give it all your crypto information and send it loose on Reddit. You're not an idiot, but getting it to the point where you know you're reasonably safe if you're not being stupid—obviously, you wouldn't leave it autonomous in the background with access to your email and main accounts, but you might let it run on your main browser while it's in non-autonomous mode, where you have to supervise it for any substantial move. If it turns out they've done their homework, you might put it in a sandbox or alternate account where it's reasonably safe.

Then you have Claude Code, which can have access to your internet, your computer, and your file system, all in the context. You integrate that with Claude for Chrome, maybe an update to Opus 4.2 or 4.5, depending on how we do our numbering systems. I don't know.

I get really excited. Similarly, if Claude just catches up in long context, interface, and inference, right? If we get Claude Opus Pro on the level that GPT-5 Pro has, it would be nuts, and there's no particular reason why they can't do it. They just haven't done it yet. That's the one big advantage OpenAI, I think, still has: they do a much better job of being able to use more inference compute to improve their outputs, probably because their models are cheaper, presumably. So they have some advantages there.

Zvi Mowshowitz

Yeah, the cost difference is really pretty crazy between GPT-5 and Claude Opus.

Nathan Labenz

You don't notice it as a chat user because the marginal cost for both is zero, and the fixed cost per month is the same—both $200. Of course, I'm going to have the deluxe version of all the major models because this is my work. It's research. I understand a normal human would probably choose the one they want and then not pay for the deluxe version of all 3 at once.

Zvi Mowshowitz

I think Claude for Chrome does sound awesome, potentially.

Nathan Labenz

Yeah.

Zvi Mowshowitz

It doesn't sound like a timeline-shortening thing. So it really is just kind of the next scale-up that's the main thing you're looking for—

Nathan Labenz

For timelines.

Zvi Mowshowitz

Yeah, the next scale up is the obvious thing. I think we've seen, to a large extent, that progress on the timeline doesn't necessarily cash out in visible, tangible progress today, in the sense that the progress we've seen today is much more about fusion, much more about scaffolding and practical application. Therefore, these 2 things don't intersect that much; it's more that they speed you up, right? This development now makes me more productive, which then in turn means that the companies can more efficiently make more progress down the line. But they don't make that progress because they've made progress; they make progress because they've got better use of the progress they already have.

I think on the order of a few months of development here, you continuously get more chips, continuously get more compute, continue to get more profits, and continuously get more investment. You get better models that do more things faster and cheaper. Probably worse factors, right?

Claude Code was a big deal. We can now look back, in some sense, and say that right now we have a Claude Code CLI, and then Google Jarvis. Jarvis is presumably not good, but now that we have a command-line form factor, we've learned that's a big deal. We get the browser agent working for real on top of that; that's another big deal.

But in terms of scaffolding, the next frontier is: How do I actually get use out of this? How do I make this do things that I want it to do? I've been waiting for a long time. The things that I keep anticipating and that keep not happening are on the agent side: Where is the thing that can handle my email properly? Where is the thing that can let me do various customized tasks? I am surprised that we are at this point in the calendar and don't have it, but we don't have it.

I do continue to get value from Shortwave. I happen to be wearing their swag today, which is just a coincidence, but it certainly doesn't take me out of needing to do anything with email. It does a really good job of triaging the inbox and getting rid of the crap, so I can focus on what I actually need to engage with. It occasionally can also draft a really good intro for me. I send intros every so often, and it'll do a pretty good one now based on examples and context.

Nathan Labenz

Yeah, it's possible I could give a real shout-out to something like Shortwave or one of its competitors. The problem for me is that I don't have the problem that Shortwave seems like it's currently capable of solving. Shortwave solves the problem of having too much email, forcing you to triage, and I don't. I am somehow one of the fortunate people who just doesn't have triage. I will literally just evaluate all the email that comes in, modulo spam. I need a spam filter, obviously, but once it gets through Google's spam filter, I don't see it.

It's literally like my eyes somehow didn't see that there was a line there for me to click on. I didn't notice. It will occasionally happen, but I don't need that level of triage. As a writer, it's very hard for me to use it. I don't know; it's possible I should. Yeah, it's hard to find it. But I keep anticipating that we'll get much better at that shortly. It isn't yet, but it's probably not, from what I see.

The fact that I'm just not feeling compelled—you can sense when somebody says, “No, dude, this is the next big thing. You've got to be on this. How are you not doing that?” You're not giving me that at all. It's pretty cool.

Okay, sticking with the lightning-round theme, I want to do a minute on charity or philanthropy. Just maybe 2 quick things before we go into that, and that'll be the last big thing. But in the lightning-round spirit, there's been a lot of discourse recently about whether AI is impacting employment or whether junior coders aren't getting jobs the way they used to.

My general read on this—I wonder if you have a different read—is that I don't really care too much about the studies that are coming out right now looking retrospectively, because I wouldn't expect too much of an impact just yet anyway. In another not-that-long period of time, presumably a lot more meaningful evidence will come in that will clear these questions up one way or the other. So I don't worry too much about that sort of stuff.

Zvi Mowshowitz

It's obviously an interesting question. What do you mean by “not care about these things”? I agree that we haven't seen major impacts—especially, we haven't seen major impacts on the unemployment rate yet. I think it's entirely possible that we have seen major impacts on the ability to get entry-level work in at least a substantial number of fields, and that in turn could have affected the supply-and-demand balance and the ability to get into other fields. In addition to that, which makes perfect sense, basically think of entry-level hiring as forward-looking, right?

It's not indicating that there are no jobs to be done now, right? Employment value hasn't declined that much yet, which is a thing that you correctly observe hasn't happened yet. Clearly, we haven't eliminated the need for that many jobs. If you were hiring, would you add entry-level workers that you have to train if you think that 3 years from now you won't need them? Not particularly, if you don't really need them now. You will muddle through with slightly fewer employees and try to automate the processes enough to make up for that, rather than using that additional productivity to expand because you expect additional productivity gains.

One of the great, interesting gotchas of skeptics is to point out that radiologists not only are not out of work but are being paid fantastically large amounts of money. You can make $1,000,000 a year right off the bat by just saying, “Hi, I'm a radiologist. Hire me to do radiology.” And why is that? Lots of people doing radiology for the last 5 years wouldn't do radiology jobs down the line, right? Which is a way of saying that you have to pay people a lot more to be radiologists if, in 10 years, they're going to be fired—maybe 5. It's going to be much harder to find work. There are going to be too many radiologists out there if you fill all the radiology positions now.

So you've got that same problem of, “I don't want your job because there's no future in it, and you don't want to hire me because there's no future in it.” So there's your matching problem, and employment is hard. Part of that problem is that junior positions lead to senior positions. If nobody gets trained into the senior positions, you're going to have a labor shortage if there's not a loss of jobs in the future.

The other half of that problem is that those jobs are lost. That doesn't mean, of course, that there isn't a net loss in jobs, because there's plenty of job creation for AI as well as job destruction. It's entirely possible that job creation exceeds job destruction for now, maybe even at an entry level, but now we just don't know because it's very diffuse and hard to measure when jobs are created in these situations.

We don't even have all the effects of AI in GDP yet; literally, the capex investments add to GDP. It's a mathematical equation. The mythical—I like to mock people who say 5% GDP growth per year, from Tyler Cowen, as if you think this is impressive. I think it is below the lower bound. We're already above that lower bound just from AI capex, even if there are no other effects from AI.

Nathan Labenz

Why is China refusing the NVIDIA H20s? I can't really make sense of it. It seems like the least AGI thing, the least AGI-pilled thing, that they could possibly do.

Zvi Mowshowitz

So, the first thing we have to realize is that China is not acting AGI-pilled at all. I think that we have this image of China as if the Chinese are super on the ball, intelligent, and wise, and always make these great decisions, whereas authoritarians, central planners, and socialists throughout history have always been going around and making huge mistakes—not out of malice, but out of stupidity; and not stupidity, but ignorance.

The whole problem of authoritarianism is that communication is hard. Coordination is hard. Social calculation in general is excessively difficult, and AGI in particular is especially difficult because it's weird and requires people to look past the evidence of what's happening in front of their faces to a future development that logically is coming but is very hard to feel.

All of us don't seem to have the ability to hold this idea: This is coming soon, but we don't know exactly when it's going to arrive or exactly in what form. You have this thing of, “What are you going to do when 2027 arrives and AGI hasn't happened yet?” Well, you just have to assume—you have to admit—that it's never coming.

No, actually, when we announced AI 2027, our median timeline for AGI starting to arrive and having an impact was early 2028. But none of these details matter exactly. It's more just that the Chinese understand manufacturing. They understand that you need to be the person making the stuff. They understand that you need to work hard. They understand abundance. They understand production. They understand not depending on outsiders.

These are important things that they understand well. They understand that the U.S. is a strategic rival, that the U.S. has leverage over Taiwan, and that they need not rely on Taiwan for their chips, because in many ways that can go very wrong.

They understand that they want to build their own AI models because you don't necessarily want AI models run by the West to be what people are asking for when they seek knowledge. What if they ask about freedom or Tiananmen Square? They have a well-established principle that they don't like that, so they need to make their own models. They also don't know if there are any back doors in any of this, or if we're conspiring in various ways. We're not, but they have no reliable way of knowing that.

We don't trust them, and they probably don't trust us in reverse, right? We would absolutely be putting back doors into various technologies they were showing, so it wouldn't be a surprise. That doesn't mean they're wise, right? It doesn't mean they're not going to overbuild various things and underbuild other things. They're going to make massive mistakes, as the reason their fertility level is down around 1.1 and their youth is fundamentally not doing great.

You always assume that the Soviets are going to outdo us—“We will bury you,” right?—or that fascism was the future. Of course, Mussolini and the company are going to be able to produce better because they get their people to do the things that are most valuable, so of course they will win, et cetera, et cetera. Now they're saying the Chinese don't even make profits; they just compete against each other ruthlessly and drive everything down. They have to outcompete everybody and starve them out. That doesn't mean they aren't challenges, but don't turn it into “China knows what it's doing,” and don't assume that they know what they're doing.

So, in this case, no, man. The Chinese don't believe in AGI. DeepSeek believes in AGI, right? Chinese companies believe in AGI, because of course they do. But the Chinese Communist Party basically doesn't, and it doesn't understand the game that it's playing. The bad news is that if it's racing to maximize chip production, capacity, and compute, it's going to do basically the right thing for the AGI race anyway. It's already maximizing energy well beyond anything we're doing, and it has essentially an infinite energy supply. So that's not good either.

But the mistake is potentially telling DeepSeek it has to use homegrown chips for infrastructure, not just inference, and throwing a giant monkey wrench into their DeepSeek lab, because they put that much priority on integrating their chips in this way. It doesn't actually matter. I also think that the signals matter, right? If we're so intent on being the ones selling and freeing all the chips, that must be what we're racing toward, right? If the White House is saying that's what matters, why should they not believe us, in some sense? That's the biggest, strongest, costliest signal.

So, when Trump is saying, “Okay, you can sell the H20s. We just want a little bit of money, but we want to dominate markets,” and Sacks is saying, “We're going to dominate markets,” and Lutnick is saying, “We're selling you our third-best chips, and you have the fourth-best chips, so you're going to take them, and you're going to like it, and we're going to laugh in your face,” the Chinese interpret this as, “Oh, we should take these, right? And we shouldn't give them this.”

They may or may not be actually insulted by this idea. I've seen people say that the Chinese aren't stupid, and that the Chinese wouldn't just do things out of being insulted or whatever. We do them all the time, right? The Europeans do it, the Russians do it—everyone does it. Why wouldn't the Chinese also do it sometimes?

I think that had some effect, but I think it's mostly that we got our priorities backwards. They took their cue partly from us, and they're in the wrong battle. Not a stupid battle, but not the most important battle. They made a mistake. Mistakes happen. Really important mistakes happen in history all the time, and I don't think it's surprising.

So, the Chinese are refusing the H20s. It could also be a trade-negotiation tactic: if we accept the H20s, then the Americans will treat this like a concession to us. But we're pretty sure it's a trap. They might even think that we put spying devices in them or something. I don't know, but I don't think so. We can't prove it, so they have many reasons to be suspicious and to think that the move is to refuse them.

It's obvious to you and me that that's stupid, the same way it's obvious to you and me that selling them to China was also stupid. The key question is: What happens if they try to sell B3A? Are the Chinese actually going to follow through on this? Are they going to say, “No, we don't want American chips. It's more important for us to clear the path for Chinese chips,” even though China will still have demand far in excess of supply for its chips—for all chips—for the foreseeable future? There's plenty of time to tell them to switch over to buying Chinese chips if we ever change that.

Meanwhile, Dylan Patel is tweeting about how, oh look, China is going to triple its production of chips, and then you see the graph. The obvious response is: Here's projected Chinese chip production. They do, in fact, triple chip production. Here's what they've been able to produce next year, in addition to everyone else's production in the West. And here's the relative quality of those 2 chips.

Why are people saying things like, “In 2026, Huawei is going to pass NVIDIA”? It doesn't make any physical sense. It's completely incoherent. It also doesn't make any sense to say that selling the H20s to China is going to slow down Chinese chip production even a little bit. It will have zero effect on us, because Chinese production is, by their own admission, going to come online.

So, the Chinese are making this mistake for some combination of a variety of reasons, but we're making the reverse mistake. The other galaxy-brain-level reason to do it, of course, is to goad us into selling them something better. They're refusing the H20s. That becomes a talking point for the Sacks crowd. They say, “The Chinese are smart enough not to want our chips. Of course, we don't sell them our chips.” And then we release something much better, and the Chinese quietly say, “Oh no, not the prior batch.”

Nathan Labenz

Yeah. How many-dimensional chess does that become?

Zvi Mowshowitz

Yeah. I mean, it's still surprising. I think that's all pretty good analysis. It's still surprising when they just had, not too many months ago, this big meeting of Xi facing all the titans of industry. There was the DeepSeek guy on the end, right? He had made the sort of big splash. You would think that guy, at this point, would be able to say, “Hey, I have basically infinite demand for chips, and I'd really like to be able to buy these. Don't worry, as soon as domestic production is there, we'll buy those too. By all means, subsidize that.”

It's just strange that you can't even get that basic a message through to the top. Obviously, he can say, “I will buy as many Chinese chips as you will sell me, and I will also elect to buy as many chips as I can get from NVIDIA.” I don't see why one has to do with the other.

One thing about authoritarian structures is that they're not good at listening to people. They're not good at incorporating information. China also has a long history of deciding on big strategic priorities and then enforcing them, whether or not that makes local sense, even when it looks like it's going to cause a lot of local pain, and even when it does cause a lot of local pain.

That's not always wrong either, right? Sometimes you do something that looks really expensive and seemingly crazy because it has long-term benefits, whether that's changing the culture, changing the incentives, encouraging the rise of new industries, or whatever it is. We wouldn't have been settling for it. The fallout would have been much, much worse for us than for them.

But maybe it's wise, and I think sometimes it's clearly not wise. We have many examples of the Chinese Communist Party and other similar regimes doing things that are fairly unwise. But sometimes it works out. Also, we're shooting ourselves in the foot in America in a wide variety of ways.

If we were just not shooting ourselves in the foot in a variety of ways, I would have complete confidence that China was not the threat, right? If we were doing proper permitting reform, actively encouraging solar and wind and batteries alongside nuclear and everything else, if we were clearing the way for high-skilled immigration and taking all the best people out of China and everywhere else and bringing them to America, if we were building housing where people wanted to live, and if we had federal rules that got in there and basically beat everyone over the head with a crowbar until they agreed to let people build in various ways—I have a long list of fixes for that.

But instead, we do things like ban American ships from taking things from one port to another port. We just self-own all the time, and then we act like it's impossible for other countries to also self-own. It's just not.

Nathan Labenz

It seems like this refusal of chips puts any hope of a Chinese viable player in jeopardy in the short term. You can complicate that analysis if you want to. I'm interested in a meta-rant if you have one. It seems like they're splashing the pot and it's chaos.

We debated last time whether or not we wanted to consider xAI a viable player. Since then, we've had the MechaHitler incident, followed closely by the Grok 4 launch, in which they had nothing to say about the MechaHitler behavior. What, if any, rant would you like to offer on the fate of these aspiring AI players?

Zvi Mowshowitz

So, you can't get the players out. Obviously, they have various ways of accessing compute. They are experts at squeezing every little bit out of whatever they can get. Chinese chips aren't useless. China still has roughly 15% of the world's compute. It's not as if, if they deliberately decided to give as much of it as possible to one company, they couldn't get something done. They're still smuggling a decent number of chips in.

We're putting data centers all around the world, including in places like India, the UAE, and Saudi Arabia that aren't exactly the most secure places to put data centers, shall we say. I wouldn't. So, it does seem like DeepSeek is still clearly the number 1 Chinese lab to me.

Kimi was impressive in some ways, but I think the standard pattern is that something impressive-looking comes out of a Chinese lab, and the majority of the time it turns out to be nothing. It turns out that they were benchmark games: its best features were touted, but in practice it wasn't very useful. So, if you just assume that nothing ever happens, you do very well. But occasionally something happens, like with Kimi. Since then, it has seemed okay in some narrow domains, but not that good overall.

Similarly, there's Z.ai, or whatever exactly it's called. That seems okay. But I would say it's almost certainly still got a lot of talent and is still a live player if the AIs were unleashed. They still haven't unleashed them, but they look continuously less live because they haven't done something impressive. Right? So, R1 does not count as progress. It counts as incremental; it keeps the lights on a little bit, but not very much.

They're basically coasting off R1, the reputational benefit from R1, and the fact that open-source models haven't really invented that much since R1. That was a low-hanging fruit that got picked very expertly, don't get me wrong, but I would primarily rank OpenAI number 1, Anthropic number 2, and Google number 3. I understand that some people think Google is better than that, and maybe they are. They do a lot of different impressive things on the side, but I still want to see more before I'm willing to give them that much credit at this point.

Also, their resource advantages are shrinking, right? Google started out with, “We've got a trillion-dollar company and you don't. We've got all these TPUs and you don't. So, we've got this huge advantage. We've got this reputation, this background. We've got the distributional apparatus. We should have just won,” right? But OpenAI is worth $500 billion. That's a decent percentage of Google's valuation, and a lot of Google is not directed at this. Anthropic is already worth $183 billion. We're not that far from the resources being pretty similar.

With that advantage gone, the fact that Google is a broken company, very dysfunctional in many ways, is going to start to catch up with them. Anthropic and OpenAI are very fast, but I think they're clearly 1, 2, and 3 in some order. Then you can have some argument over exactly the order. Basically, Google goes somewhere in the ranking, and then xAI is the wild card. They have a lot of compute; they play hard but not very well. You want to write them off? Prove it, right? It's always, “Prove it.”

Similarly, Meta is trying to come back. I think we Meta skeptics were proven correct that they didn't have it. That doesn't mean they can't go get it, right, and go to the ground. But now they're considering licensing Gemini or potentially ChatGPT to use for themselves, which is wise. I would do it too, right? You don't have to stop trying to develop your own AI. You just don't have to dog-food it while it's terrible. That's just not smart. There's too much money at stake. But at this point, it would be surprising to me if the big 3 got disrupted in the near term.

Nathan Labenz

One specific question I have around xAI is this. There were several things that stood out to me about the Grok 4 launch. One should never forget Elon's comment that he's not sure if AI will be good or bad, but even if it's not going to be good, he wants to be around to see it.

Zvi Mowshowitz

Yeah, it turned out you might not be around to see it for very long if it's not good. So, I would be careful about that.

Nathan Labenz

Yeah, that was an “I can't believe you just said that” sort of thing.

Zvi Mowshowitz

Wow. Wow. It was like, only if it was a livestream would they let that one out. But he's that kind of guy, obviously.

Nathan Labenz

The other thing that really stood out to me, though, was that he was talking about how they're going to be giving the model access to the same power tools that the engineers at SpaceX and Tesla use in the next generation of training.

It got me thinking: if we are headed toward a world where the quality of the problems that the model is challenged with in training becomes a differentiator, then they might have the best feed of well-structured, very hard technical problems that are amenable to being solved with really advanced software tools—perhaps of anyone. They just have these really hard problems. He's got these frontier companies in multiple domains where they really do a good job of seemingly structuring problems.

I don't think Anthropic has something like that. I don't think OpenAI has something like that. Google sort of does, but it's extremely diffuse across their vast archipelago of fiefdoms that kind of roll up to be Google.

But Elon—I could see him structuring that pipeline of hard problems into an RL cooker at xAI and maybe coming out the winner because of that access to the best engineers working on these really hard things. Does that seem at all credible to you?

Zvi Mowshowitz

Not really. I don't think there's that much data that naturally happens. These aren't that big, especially SpaceX. But beyond that, if that is right, if that is the thing that matters, if that is the bet, then there's really a lot of data in the world to be collected. There really isn't that much of a barrier to collecting it or to getting access to it.

There's nothing stopping Google, OpenAI, or even Anthropic from making those alliances and getting that data. There's no reason why those companies shouldn't be happy to help them do that in exchange for not that much money. So, I just don't see that as an advantage. It just seems hard.

I agree that there are other car companies, for example. There aren't exactly other space companies, but say car companies, right? You've got all these engineering things happening at Tesla. They're happening at other companies in varying degrees.

Nathan Labenz

Right. It's funny. Can you imagine, though, going to General Motors—to take a company in my own hometown here—and saying, “Hey, can we extract your hardest engineering problems from your organization, structure them in such a way, get clarity on what the answers are, and train our AI on that?”

Even if the CEO of GM said, “Yeah, sounds great. We'd love to do that,” I just feel like it would easily take 10 times longer than it would take Tesla to do a similar thing.

Zvi Mowshowitz

Why?

Nathan Labenz

I've done a little work with GM. I don't know. For the same reason nobody else has anything close to a self-driving car, right? Only Tesla and Waymo have come far in that domain, and everybody else has given up. They just don't seem to have the organizational juice to be able to pull things like that off.

It's very different executing a very long-term, complicated engineering plan, like software engineering, versus collecting a bunch of data. What data do you have to collect? The goal is to collect data. Google has Waymo, so they already have almost infinite driving data at their fingertips if they want it. It's not that hard to put cameras on a bunch of cars if you want to collect a bunch of driving data. It's just not that expensive.

Consider how much they're spending on training runs, right? Consider how much they're spending on acquisitions. I literally just Googled the market cap of GM. It's $55 billion. You could buy General Motors if you were OpenAI, if it was so important.

Zvi Mowshowitz

I wouldn't recommend it.

Nathan Labenz

Why not? Imagine if you could buy General Motors and then, using OpenAI's techniques, launch self-driving cars relatively quickly. Couldn't you generate a lot more value than $55 billion? What's Tesla worth? Why is Tesla worth so much more? Are you sure we should buy General Motors?

I'm not actually advocating buying General Motors. I'm saying that OpenAI is worth $500 billion and General Motors is worth $55 billion. So, if the thing that's preventing them from doing this is not owning General Motors, then they can just own General Motors. It's not that hard, right? And there are synergies. There are big synergies.

Zvi Mowshowitz

Yeah. Yeah, I guess the thing I see being hard—even if you did buy General Motors—is that the thing that seems hard to reproduce, and that I think the likes of Tesla and SpaceX potentially have, is just really clean environments where it's a well-oiled machine, data is flowing, and vertical integration is deeper.

I mean, the mess of the supply chain or whatever at GM—all the suppliers, all that nonsense—it's like, data collection can be multiple things. Cameras on cars are one version of it, but I'm also thinking about problems: We wanted to design something that met this specification, and eventually somebody did do it. What was that design, and where does that sit?

That stuff seems like it's much more accessible and ordered, probably, at Elon's companies as opposed to at legacy manufacturing giants that have declined a bit already.

I see various different points in the story that don’t make sense to me. The way I put it is: would this really be that big a deal, especially given that Google has Waymo, which is literally, as far as I can tell, the only company actually doing the thing? I don’t like this myth of Elon being the super-executor. Elon hasn’t been the old Elon for a while now, if you just judge by the quality of his public statements and decisions, including blowing up his very close relationship with the president of the United States over nothing he stood to gain whatsoever, as far as I can tell.

Right? Just in terms of this person being able to execute on the plan, he was the right-hand man for tech to Donald Trump. Then he got mad about the deficit, something he had no influence over and that he really didn’t give a damn about, given his belief in AGI. He blew up the entire relationship about it pretty consciously and intentionally, knowing what he was doing. And now that person seems to be a disaster for the United States, right? As far as we can tell, this did not help. His influence going away did not help anything Elon Musk cares about in any way, shape, or form. His life is just worse. Everyone on all sides basically hates him.

If you look at the self-driving car situation, they’ve been promising me self-driving cars real soon now for how many years? The same promise over and over again. I’m not saying they’re not making any progress. They’re making progress, but it’s been way behind any schedule that he’s told us to expect. He’s overpromised and underdelivered for about a decade now, and Waymo is much more the one actually doing the thing.

So I’ll believe it when I see it. But also, yeah, if it’s a data thing—I don’t think it’s a data thing—I often hear these stories about, “Here’s this reason why someone will win because they have this thing.” That thing was so important, you could just go get it. So you have to be the important thing and then have everyone else not realize it’s the important thing until it’s too late. You still there?

Nathan Labenz

Yeah. I would say the FSD, for what it’s worth, is getting very good. It’s been a couple of months since my last FSD ride, but before that, it had been about a year between that last one and the one before that. The progress was definitely very obvious, and you no longer have to keep your hands on the wheel, for example, as one thing that shows the increasing level of confidence. I was very impressed. I’m also super impressed by Waymo, but I wouldn’t say he is the super-executor.

Zvi Mowshowitz

I’m excited. I’m excited for it, and I still know how to drive the car without driving a car. So there you go.

Nathan Labenz

Waymo is supposedly coming to New York City pretty soon.

Zvi Mowshowitz

Okay. I’ve been told that by people in the New York City Council and the mayor’s office.

Nathan Labenz

I thought I just saw this on Timothy B. Lee’s newsletter, I think.

Zvi Mowshowitz

A Waymo car has been spotted in Brooklyn, driving around and mapping the city. That’s great. There are laws on the books that say they can’t operate. What is the plan to deal with that? Don’t get me wrong, I want this to happen so badly. I want my Waymos. I will forgive the new mayor many things if he brings us Waymo.

Nathan Labenz

You take your Waymo. You could take your government Waymo to your government grocery store.

Zvi Mowshowitz

I don’t want a government Waymo. I don’t want a government grocery store. But there you go.

Nathan Labenz

Okay. Last area for today. We both just participated in different cohorts in the Survival and Flourishing Fund grant-making process as recommenders, and I’d love to hear your thoughts on the broad survey of the AI safety charity landscape, from cause areas to specific organizations to anything you think is neglected. What did you take away from that? There were 400-plus applications, of which a bunch were pre-filtered out, but we still had 125, I think.

Zvi Mowshowitz

The first thing to make of it all, obviously, is that you can’t actually evaluate 125 applications, let alone 400 applications, in the kind of time they expect us to spend on this and budgeted for us to spend on it. It’s just impossible. You can properly investigate on the order of 10 organizations, maybe, and then you have to evaluate everyone you think deserves consideration for funding, which is going to be a lot more than 10. So you’re relying a lot on your fellow recommenders, relying a lot on reputation, and relying a lot on your past research.

One advantage I had is that I was in the previous, or a not-so-long-ago, round of SFF, and a lot of the same nonprofits were applying again. So I could ask for a diff on those organizations rather than starting from scratch, which is a huge time-saver. Indeed, a lot of the charities that ended up near the top for me were basically the same charities that ended up near the top, or under serious consideration, last time, because the situation hadn’t changed that much since then.

I had the same number one I had in the previous round, which was the AI Futures Project. That’s Daniel Kokotajlo and the people who did AI 2027. Between the time of the last grant and now, I took a victory lap. Pretty obviously, there aren’t any downsides to that. It still feels like clearly a good hit, given that I think I was the only one who put them that high last time, and this time there’s a pretty big consensus among a number of recommenders. They should be pretty hot.

Basically, you can see the big divide as policy versus research, right? Are you trying to solve alignment in some form? Are you trying to directly make the world better? Are you trying to shape public opinion, shape public policy, propose laws, file lawsuits, et cetera, et cetera, to try and set better policy? I definitely wanted to do both. I felt it wasn’t obvious that one was strictly better than the other.

You’re largely looking for what is underfunded and what is otherwise not being funded in the ecosystem. One of the questions I asked a lot was, what is SFF’s comparative advantage? Where do we get to identify talent and identify opportunity in a way that would be difficult for other people to fund?

Last time there was the AI Futures Project, which was specifically unable to get certain other funding at the time for various reasons. This time, for example, I had the C4 Action Fund pretty highly, because it’s harder to raise funds for C4 than it is for C3—not because I felt like the Action Fund’s money was much better spent on average, but because the distribution was naturally going way too far the other way.

ACX Research was one of my top picks because, specifically, they were running out of funds, right? I’d seen them do some things that I was very happy with, and I said, “Okay, I need someone who’s doing valuable things not to just fall over and die. That thing is really valuable. You stay in the ecosystem, you keep your organization, and you don’t have to constantly work for another home. You can do whatever you think is valuable here.”

Ultimately, MIRI didn’t need funding for many years because they got some very large donations, and they did the highly virtuous thing of actually not asking for donations while they didn’t need the money. Now they need the money. I really hope that we come together and make sure they continue their course.

Those were some things that stood out at the top, but I have a very long tail of places I would be happy to fund. If you allocate all of the money that the entire round would get and ask what I would fund, there are 10 organizations that would get at least $400,000, another 10 that would get at least $100,000, and a long tail that would get some amount of money with my funding. I could go on about this. I’m planning to do another nonprofits post at some point later in the year, in advance of Giving Tuesday and all of that, to give my updated views as of—

Nathan Labenz

Yeah, nice. My written output is obviously not 2% of yours, but I’m planning at least a Twitter thread on that topic as well, so we can compare notes as we get into the long tail.

Zvi Mowshowitz

Yeah, it’s a lot of work, and I’m going to have to set aside specific time for it. But part of it is that I don’t know when they’re going to announce the results, and I don’t want to—I want to finalize it. I want to see who gets how much money, and I want to see who is then on record as having received money and therefore publicly part of the round, versus who is not.

Anyone who doesn’t receive money is not publicly part of the round unless they say it—unless they say they are. Therefore, I have to email each of those people and ask, “Do you want to be in this post?” At least I give them the opportunity to email me and say no, and I treat no answer as a yes, because in general charities want people to say you should give money to them. But occasionally someone doesn’t want that.

Nathan Labenz

Usually a safe assumption. Another category that I thought was interesting was international relations. I know you’re not very bullish on U.S.–China cooperation—not that I am, and I’m not arguing that you should be—but there was a crop of organizations trying to work on that, and I was pretty into that.

One of the issues with the round is that there were several organizations I knew were doing good work. They had clear wins in their column, or I knew the people involved and the things they had done, or otherwise I could be confident in them. Then it was very hard to give a similar level of ratings to organizations where I didn’t have that. There were some charities that were trying to create Track 2 talks or otherwise advance U.S.–China relations in various ways related to AI, and I gave them some support.

Zvi Mowshowitz

I definitely give them—this is fungible—but I was hesitant largely because it’s very hard. Diplomacy is notoriously difficult to read, right? If one of these charities was effectively fake, in the sense that what they were doing had no real effect and did not in fact impact the possibility of good things happening, would I know? Would they know? They might, if they’re bad at it, not even realize what they’re doing. They might be doing this thing thinking they’re accomplishing something and just not be accomplishing something.

Diplomacy can look for 10 years like you’re doing nothing and then suddenly something happens. Or it can look like nothing will happen and then nothing happens, but that was actually the right thing to do. You just created the possibility of something happening, and things went a different way, or you stopped the bad thing from happening without even realizing, or whatever.

So that’s why I hesitate: you just don’t know. Because of that, I find it difficult to get behind any of these organizations on a high level. Some of them will definitely make the cut as something that I think would be reasonable to support.

But it’s really tough when you’ve got money that could definitely go to places where it would be well spent, to put it into a weird, impossible-to-read place where you don’t get good feedback. That’s all the more reason to ask: how do they know how to do things properly? How do they make good choices even if they’re properly motivated? They can’t tell either.

Nathan Labenz

Another category—this is sort of a meta-category—is California’s SB 1047, which I’m sure you’ve engaged with a little bit. That would create a private regulatory market where either the attorney general or perhaps some new commission would credential private organizations to be regulators. Then there would be some sort of trade: if an AI developer opts into regulation from one of these private regulators, in exchange, they would get some sort of liability protection. That obviously begs the question: who steps up to be these private-sector regulators in the event that this bill were to become law?

So one of the things I was looking out for was who I see that kind of feels like they could become that if that opportunity were to arise. I don’t know if you’ve thought about organizations in those terms previously, but—

Zvi Mowshowitz

Yeah, there are a number of nonprofits that plausibly could step up and become auditors in this space. There are a number of founders who are perfectly capable of creating new organizations that could plausibly do that in this space. If there’s demand, there’ll be supply, right? It’s not that hard to find expertise in this space that would be happy to participate in these organizations.

I just don’t want to name specific names, and I don’t want to fall into the trap of, “Oh, yeah, this must be regulatory capture for these 5 people,” or whatever it is. But I guess I don’t think that’s the case. I don’t think there’s going to be any problem with that.

I think the most likely scenario is that companies like OpenAI quietly ask people who they expect to do good jobs to spin up organizations they can then work with if they’re not happy with the slate of options they’re initially presented with. Certainly, for example, METR or Apollo Research—the people who are already being contracted for evaluations by the big tech companies—would presumably be the first ones to enter the space, and they would presumably be very credible in that capacity.

Nathan Labenz

Yeah. The last category I’ll put in front of you is hardware governance. There weren’t too many organizations that were specifically working on this, but the read in my group was that everybody seemed to have a different reason to like hardware governance. What’s your thought on hardware governance?

Zvi Mowshowitz

Unfortunately, at the moment, it’s politically really tough for hardware governance. I wouldn’t necessarily ever say it’s dead, because things change so quickly, but it’s not looking good. The focus is entirely on getting people to use our hardware, and the last thing people using our hardware want to do is be tracked. So there’s a direct conflict with exactly what the prioritization is.

It’s not going to happen right away. I still think it’s vital that we have that ability. I think it’s really important that we have the technology completely shovel-ready, so that if we decide, “Okay, as of 3 months from now, every new chip that gets shipped has to be tracked,” we can do that.

Ideally, if we have to go into a data center and put trackers on these chips in such a way that, if they’re tampered with, we’ll know how to do that as well. Potentially, a very small amount of money can give us that option, and then we only have to spend the real money if we implement it. But you get the optionality, right?

I think the first-best solution to the current mess is in fact to use this solution to allow us to do things like build data centers in the UAE and India in ways that feel secure, stop chip smuggling, and potentially even be more aggressive about what you let people you’re actively worried about do.

I don’t think we need that much effort. I think we just need a little effort to lay that foundation, and then the question is just figuring out who the real deal is.

Nathan Labenz

How about just on the simple question of the balance of money available and opportunities? I think for me the sense was, I wish I had more money to give out. Obviously, to some degree, that would always be the case, but if you were to make a pitch to other philanthropists that there’s a lot of stuff that isn’t funded as much as it should be that would be high-impact, I guess, first of all, do you believe that? And second, what would that pitch sound like?

Zvi Mowshowitz

That pitch would sound something like this: if I had access—if I wanted to hand out the entire, let’s say, $10 million that’s roughly projected to be the entire round—it wouldn’t allow me to give out all the money that I would have been happy to hand out. Not by a long shot. I could easily hand out more than double that and feel good about every dollar that I was giving out.

Obviously, I would want to do more investigation of some of the things that would then come up, because there are things where I’m never giving money. I’ve done the math, so I’m not going to think too hard about exactly how to rank that. But you could give out tens of millions of dollars to these applications with basically zero worry. That is the obvious direct evidence that there is infinite philanthropic space.

There’s also a ton of stuff that’s at a scale above SFF that we’re basically saying, “We can’t fund these charities anymore,” because their capacity to use money effectively has gone to millions of dollars a year, and we just don’t have that capacity—or, in some cases, they’re approaching $10 million a year or more.

There are also a ton of far-reaching projects that were never even proposed because the price tag on them would be absurd. And everybody in the space is, of course, conserving money, which isn’t particularly great. We’d prefer to have generous salaries and generous compute budgets and not worry about this stuff, but that would take a lot more money.

Even without changing that, even keeping everybody narrow and not trying for extra-massive experiments or anything like that—just doing what we’re already doing—I feel confident about a lot more money being sent to this space than is already being sent. It’s just not available right now. So it’d be great if you could help out. There are lots of places to put it.

Again, there are a lot of people who just aren’t asking for money because they know there are so many other people asking for money. I’m in that phase, right? I’m being supported by patrons who are happy to support my work, but I’m not going to go out there and seek out more funding because I already know there’s way more demand for this funding than there is supply.

Could I scale up at least somewhat with more funding? Very obviously, yes. There you go. I have lots of research, too, which, again, is trying to be lean. We’re all trying to be lean here.

Nathan Labenz

Okay. Here’s one idea I want to get your take on. This was not in the application pool, but I just saw this article. I’m sure you saw it the other day as well, about how OpenAI is starting to subpoena some of these charities, some of which were in fact in the application pool, that are doing various things they find to be inconvenient, like hassling them about their nonprofit-to-for-profit conversion.

They seem to believe—and the reason they’re giving for why they’re issuing these subpoenas is—that they may be funded by competitors. In other words, they think perhaps Elon or maybe Google is funding these organizations to try to slow OpenAI down.

This got me thinking: maybe that could happen at the model-evaluation level. We’ve got these model-evaluation companies, organizations, and nonprofits that are, I think, very focused on being even-handed, very fair, very analytical, and trying to do stuff pre-release, which I definitely think has a lot of value to it. But that kind of forces them to play very nice with the companies that they’re working with.

I wonder if somebody came forward and said, “I’m going to potentially, transparently or not transparently, go after all the companies except Google and try to demonstrate to the public why their models are problematic, why they should not be trusted, and all the ways that they go wrong. But we’re not targeting Google, perhaps because we’re funded by whoever.”

Could you engineer a situation where all the companies then feel like, “Geez, we’d better go target our competitors’ models. We’d better really invest in demonstrating what is wrong with our competitors’ products”? If you could create that equilibrium where they’re all kind of sniping at each other all the time, would that, in fact, bring a lot of things to light and potentially create the race to the top that everyone wants?

Because obviously there are a lot of problems that can be demonstrated. It seems to me like the only reason that isn’t happening is maybe a sort of soft collusion, or an unspoken gentleman’s agreement, I guess, is another way to say soft collusion. But if somebody were to break that, maybe it all kind of goes to a different equilibrium where everybody is investing in that and we have a lot more energy going in that direction. What do you think?

Zvi Mowshowitz

It’s not a great look, right, from their perspective, to be funding attacks on these other companies. When you expose these things in other companies, you’re also exposing yourself. Almost always, when you find these flaws, they’re everywhere in some form, or something close to them is available in some form, and it will seep into public consciousness and lead to calls for greater regulation. It looks like it could lead to any number of escalations.

You don’t need collusion for, “Oh, yeah, my gang has guns, your gang has guns. Why don’t we just stay away from each other’s territory and not shoot at each other?” Because that could get into something pretty ugly pretty fast. Most of the time, Coke and Pepsi don’t go out and start smear campaigns on each other, right? They just do positive advertising. Maybe they take a few cool, little snide shots—“You aren’t that cool”—but they don’t fund nutrition studies about why the other one is unhealthy, because it doesn’t work.

Certainly, you could fund people to go after specific organizations in these ways, or do investigations and deep dives on their ridiculousness. You can have opinions on who to target first, and I don’t think it’s that crazy. Some companies are being less responsible than others and deserve to get hit more. If you target Anthropic, they might just say thank you.

Nathan Labenz

So there’s that.

Zvi Mowshowitz

Yeah, that’s why, when I was saying “target everybody but Google,” I was thinking the rationale there is just that there are a lot of Google billionaires who could plausibly be funding such a thing. The reason they might be doing it could be a mix of competitive advantage and/or philanthropic desire, just to bring issues to light.

Zvi Mowshowitz

Yeah. And the problem with specifically not targeting Google is that—and look, I would question you, right? If you’re a Google-funded organization that only targets OpenAI, let’s say, just to keep it simple, then why do I think your evaluation is objective? Why do I think that when you say something is a problem, it’s a real problem, or that it’s a specific, particular problem, or anything like that?

Nathan Labenz

I think the idea would be that it’s just reproducible, right? If you just have inputs and outputs from models and you just demonstrate this in a test, that’s pretty—

Zvi Mowshowitz

Yeah, I didn’t mean, “Is the finding real?” I wouldn’t question that. But we’re trusting that you’ve followed scientific procedures in selecting this example, that it’s representative, that it teaches us what you think it teaches us, et cetera.

Zvi Mowshowitz

Yeah, I see that challenge, although I also think that’s all pretty slippery, right? As much as there is a very sincere desire among the eval groups today to have these high standards of rigor, all their stuff is always questioned, and you’ve got a lot of people who are like, “Oh, well, this is totally nothing because you put the model in situations…”

Nathan Labenz

It is the job of Caesar’s wife to be above reproach and be reproached anyway, right? It is the job of those who are trying to be the watchdogs in these situations—the people who are trying to be in our position—to follow standards of rigor and integrity that are vastly above what others are held to. That’s table stakes. That’s just the right to play, right? It’s not fair, and that’s just life. You’re still going to have all that questioning and attack.

We all saw the debate over SB 1047. We all saw how you bent over backwards to be 10 times better on all these issues and all of these questions than the people you were opposed to, and it didn’t. You had to be able to play in the arena.

If you’re the underdog, the scrappy, underfunded person who’s trying to bring the truth to light, that’s your job. The big corporation is going to try and squash you with anything they’ve got. You’ve got to be sparkling clean. You’ve got to have no vulnerabilities, no points of leverage, no smear material. That’s just how it is, and it sucks. But, yeah, we’re used to it.

Zvi Mowshowitz

All right. So does that mean you basically don’t buy the equilibrium that I’m trying to envision a way to shift from and to? Today, there was this mutually adversarial collaboration. I thought it was really adversarial, but there was this OpenAI and Anthropic evaluation of each other’s models, seemingly in a pretty friendly, collegial way.

That was great. I think everybody loved to see that, but that hasn’t happened much. Maybe it’ll happen again more in the future. Maybe it’ll never happen again. I was trying to engineer a transition to a different equilibrium where everybody is adversarially evaluating everybody else and thinking maybe if I tip one domino, everybody else will feel that they have to respond.

From a general sort of safetyist worldview, it would seem better if they were all adversarially evaluating one another versus not. So I guess you could question that first assumption—that it would be a better equilibrium—and then ask whether we can tactically get there.

Nathan Labenz

I would love to get to a point where the companies were doing each other’s evals in an adversarial fashion, looking for trouble, looking for vulnerabilities, looking to embarrass them. They’d just have to see if they could deal with it, see if they could beat it, see if they could overcome it. That sounds great. I would love that.

I don’t know how you get there from here. I think these companies do not want to go to war with each other. I don’t want them to go to war with each other in other ways, either. People are constantly—I mean, you’ve already got them going to war over talent, so I don’t know.

But I can see it being a thing they might want to do in the medium term: “We run all of our evaluations against everybody’s AIs, and we report them back.” Occasionally, we’re going to find some stuff.

Zvi Mowshowitz

Yeah, they did do that with DeepSeek. They came out and said DeepSeek has no qualms about doing bioweapon-type things.

Nathan Labenz

Yeah, their evaluation of DeepSeek’s safety protocols was about safety protocols, right?

Zvi Mowshowitz

Yeah.

Nathan Labenz

Okay. Well, if any Google alums with the resources want to talk about seeding such a thing, my DMs are open.

As always, the final question for you in particular: what that we haven’t talked about is virtuous to do now?

Zvi Mowshowitz

Yeah. I think it’s a weird situation where it can be really tough to figure out where to make the most meaningful progress and what to do going forward.

On policy, the short-term priority has to presumably be preventing America from being so foolish as to sell H200s to China. In practice, that presumably means getting enough people on the right sufficiently alerted to the fact that this is actually happening and what this actually means, so that they raise enough stink that it doesn’t actually happen.

Otherwise, draw attention to the extent to which NVIDIA seems to have taken hold of the White House in terms of its rhetoric and its plans. Not overall—I mean, this is obviously not the ultimate end goal or the primary reason for that—but, yeah, you see it in everything. Obviously, as usual, trying to spread the better narratives is always good.

I have certainly gotten to the point where I think that working for Anthropic seems to clearly be a good idea at this point. If you’re considering what to do and the alternative is doing basically nothing, I do think there are a number of organizations that are presumably better choices for impact than just working at Anthropic. But that doesn’t mean they have capacity or that you want to work there, obviously.

It’s a difficult situation because, obviously, the policy situation is in a bad state, alignment is in a not-great state, and there are infinite things to work on, infinite things to experiment with, infinite organizations to give money to, and so on, if you want to do that.

My project is basically to keep myself and others informed about and understanding of the situation, and hope that that will lead to good things more than anything else. I wish I had a better answer, obviously, for a call to action of, “You listened to this podcast for 3 hours—or, depending on what speed you’re listening at—what are you going to go forth and do?”

Unfortunately, other than thinking hard about the world and trying to figure out what, under your model, would be the right things to do to advance it and what things would actively make things worse, I don’t really have one. For a lot of people who are informed, the first step is just to be aware of what would make things worse and not do that.

I emphasize that you should say what you think more than anything else. You shouldn’t sugarcoat, you shouldn’t engage in hyperbole, and you shouldn’t strategically censor yourself. With rare exceptions, you should just say what you actually believe about the situation.

One thing you can do is support the book release from Eliezer Yudkowsky and Nate Soares. They’re coming out with a book in about a week.

Nathan Labenz

If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All. If you were to help by purchasing the book and spreading the word about it, this is a unique opportunity to get that book some momentum and maybe create a cultural moment.

But again, that doesn't mean that you should just back their talking points and their idea about how this works just because they're the ones with the book, or just because Eliezer is the rightful caliph and he said so, or anything like that. You should make up your own mind. I've pre-ordered my copy and look forward to reading it.

As always, I really appreciate all your time. You've been very generous with it, and we'll continue to stay informed via the blog. Don't worry about the vase. I think that's it for today. Zvi Mowshowitz, thank you again for being part of The Cognitive Revolution.

Zvi Mowshowitz

Thank you for having me.

Nathan Labenz

All right, bye.

Zvi Mowshowitz谈更长时间线、RL诱发的末日风险,以及中国为何拒绝H20 — 文字稿与摘要 | BidClub