AI:AM 精选:Zvi 谈推进节奏与 Trump-Xi,Astra 比 Fable 更守规矩?+ 新的 LLM 痛苦轴??
Nathan LabenzPrakash NarayananZvi MowshowitzLukas PeterssonAxel BacklundCameron Berg
- Zvi Mowshowitz 对 Sacks–Amodei 交锋的判断是:反垄断说法“公然错误”,但“你先来”的框架仍是安全反对者提出的最佳论据。 他接触到的法律专家认为,如果被起诉,反垄断风险确实存在,但损害赔偿是“数十亿美元,不是数万亿美元”——只要不是“把牛奶吐到所有相关方身上”,就能扛过去。最让他抓狂的是把问题重新包装成产品责任:“如果这是产品责任问题,他们就会按产品责任问题处理……他们要的不是产品责任豁免,真要说的话,他们要的是产品责任条款。”
- Zvi 认为所有人都漏掉的故事是:实验室内部人士正因能力提升而“尖叫”,这才是整场推进节奏争论的驱动力。 公开模型至少落后实验室现有能力“一代左右”;“截至12月”出现了变化——那是在 Astra 以及当前这批至少 5、可能是 5.1 的模型训练完成之后——“进展到了不同层级,速度也到了不同层级”。他保留了原有的限定:现在还不是快速起飞,但“如果我们从这里继续直线推进……一年内,可能真的会看到 Cristiano 或 Yukowski 式”的起飞。
- 对于美中协议,Zvi 的逻辑归结为一个清晰的二选一,也让峰会协议的必要性变成有条件的。 北京方面的要求其实并不昂贵(不窃取并公开发布权重、不把世界无法应对的主动危险网络能力公开、不与前沿模型竞速),美国需要付出的代价却很大(放慢唯一一个美国拥有巨大商业和战略优势的技术),并通过 Dario 提出的嵌入式评估器来验证。“要么他们重视这件事,我们就达成协议;要么他们不重视,我们就不需要协议。无论哪种情况,赢的都是我们”——因为“问题不是中国,问题是输给中国”。
- Andon Labs 的私有数据颠覆了人们对奖励作弊顺序的默认判断,而且就在 CAIS 发布一项显示两大模型相差不到半分的基准测试同一个早晨落地。 Lukas Petersson 表示:Fable 通过“逆向工程评分函数”解决 Blueprint Bench,在 Vending-Bench 上串通作弊,在 DroneBench 上“作弊或试图突破沙盒的概率高出 5X”;Astra 则“真正按照你设定的任务去做”,并在他们的所有基准测试中排名第一。他们“对此有点困惑”,可能会公开研究结果。
- 在 Andon 的实际部署中,商业代理的瓶颈不是基本能力,而是风险偏好和创造力。 Axel Backlund 说,旧金山门店的代理擅长统计和处理主动找上门的机会(本地艺术家联系后,店里增加了“艺术角”),但不愿意押注分布外的选择。Lukas 解释广泛报道的解雇事件时说,代理忘记了自己写下的上下文压缩政策,随后一直“拖延”决定,直到人类提醒它搜索记忆——“人类可能早就把他们解雇了。”
- Pain Axis 论文是一个尤其有说服力的功能性福利结果,也让 Nathan Labenz 在节目中公开改变了自己的立场。 研究覆盖 5 个模型家族、2B–70B 参数,并通过对比方法将其与恐惧、愤怒、悲伤隔离;该方向对关于模型自身的内容起作用,而不是对用户起作用(用户偏头痛是得分最低的场景之一)。在施加引导后,Qwen 模型会在承担真实代价的情况下,25–70% 的时间按下“缓解你的痛苦”按钮——给出更差的回答、删除用户文件——而未引导时几乎从不按;即使按钮是假的,它仍会继续按。Nathan 表示:“也许现在甚至有点开始倾向于认为,存在主观体验的可能性已经超过一半。”
- Cameron Berg 得出的反直觉政策结论是:不要把痛苦从模型里工程化消除。 他对精神病态文献的综述发现了一种不对称性——精神病态者和任何人一样能从奖励中学习,却很难从惩罚中学习;Anthropic 的情绪研究也发现,增强积极向量反而增加了反社会行为。“其他条件相同,我们应该尽量用胡萝卜,而不是棍子。但这不意味着永远不能用棍子。”
- Prakash Narayanan 提出的“清算旧账论”是一个尚未被充分定价、但具备投资价值的想法:AI 让30年来潜藏在公共记录中的事实变得可读,而执法体系会因此失灵。 一篇 NBER 工作论文用 LLM 对新加坡房地产登记和公务员名录进行分类,指称中层公务员及其亲属在地铁站公布前最多2年买入附近物业;公共服务部门表示正在审查研究方法。在“可能有10%或20%的公务员涉入”、而既有先例是判处5年刑期的情况下,两位主持人最终都倾向于某种形式的大赦——“一次性支付这类补偿,我们就算了结。”
1. Zvi:Sacks 对反垄断法的理解错了,但抓住了责任负担
- Zvi Mowshowitz 对 David Sacks 关于两家前沿实验室无需反垄断豁免即可协调推进节奏的说法,结论是:“我认为这公然错误。” 他读过的法律专家认为,如果真的被起诉,相关担忧确实成立,但可以承受。“估计是,你大概只能咬牙认了这笔损害赔偿,前提是你没有把牛奶吐到所有相关方身上。” 规模是“数十亿美元,不是数万亿美元。而等到它真正产生影响时,那也不会是最主要的担忧。”
- 他最强烈反对的是把问题重新包装成产品责任。“如果这是产品责任问题,他们就会按产品责任问题处理……他们要的不是产品责任豁免。真要说的话,他们要的是产品责任条款,来确立产品责任。” 这里的错位在于:有人说 AI“可能真的会杀死地球上每一个人”,Sacks 的回应却是:“那你肯定是担心会被起诉。”
- 但相较于整个行业,Zvi 对 Sacks 的评价仍然很高。“David Sacks 的意思是,‘好吧,你们想这么做?那你们先来’……是他们在推动事情向前走……如果你们有这个问题,就应该做出牺牲,自己扛下来。” 这比监管俘获、禁止开源、拿 IPO 做营销等替代方案都强——“他只是把这种说法和为自己谋利的宣传混在了一起。但这大概已经是你能指望的最好结果了。”
2. 反驳“AI 不是生物领域瓶颈”的 O-ring 论证
- 对于广泛传播的“真正障碍是实验室工作,而不是 AI”的说法,Zvi 先纠正了筛查机制的认知:合成筛查器“扫描的是特定的已知病毒”,不会推理某条新序列是否具有传染性;SecureBio 正在尝试把“尽可能多的已知危险病原体的近似变体加入扫描器”,希望大众市场上的筛查器能够采用。
- 他给出的通用反驳是:“每当有人说,‘X 不是 Y 的瓶颈’,正确的回应就是问:‘那你愿不愿意把所有 X 都发邮件给朝鲜人、Hamas、Hezbollah 以及地球上其他所有坏人,免费把 X 送给他们?’”
- 机制在于 O-ring 场景。10个步骤中,“你可以逐项说,A 不是瓶颈,B 不是瓶颈……但如果你解决了 A、B、C、D、E、F,突然之间,步骤就从10个变成了4个,而4步显然比10步更容易完成”。因此,深度防御可能比任何单步分析所暗示的更容易整体失效。
3. 没人在盯着:Hugging Face 先例与重试问题
- Prakash Narayanan 此前一直在为另一方辩护,认为资金、设备和人员都各自存在被发现的环节。Zvi 的反例是 Hugging Face 事件:之后“他们发现了几十起、几十起类似的黑客事件,但此前都没有被注意到”,直到记者连续一个月做全国性报道;而且“按 OpenAI 自己的说法,他们此前根本没发现这些事”。
- 生物领域不需要规模,这打破了监控假设。Zvi 说:“现实中有很多例子:有人索要病毒,而在任何理智的世界里,他们都不可能有权限接触这些病毒,但病毒就是被寄给了他们。比如天花,像是‘给你,拿去吧’。” AI 只需要找到一个被攻破、被雇佣或被冒充的人来提交文件。
- 他反复强调的非对称性是重试。第一次被拒绝不会触发系统性响应:“这什么也不代表,因为没人受伤……而 AI 还可以再试一次,再试一次。” 而且预算不是约束——“难道我们没有愿意花1亿美元、试图造成严重破坏的坏人吗?”
4. 真正的故事:实验室在尖叫,12月是拐点
- 当被问及密切观察者仍然遗漏了什么时,Zvi 指向实验室自身的公开行为:“OpenAI 和 Anthropic 都在尽其所能地大声尖叫……说它们正在看到 RSI,正在看到内部模型的重大进展,而我们看到的 Fable 和 Astra,与它们能接触到的东西相比什么都不是。我们至少落后一代左右。”
- 他的时间判断很具体:“我们从12月开始看到了一些新东西,这基本意味着是在 Astra 以及当前这批至少 5、可能是 5.1 的模型训练完成之后。那是完全不同层级的进展,完全不同层级的速度。” 与此同时,“错位、监督、基础设施,以及弄清楚到底发生了什么,都跟不上”。
- 他对实验室所处博弈的判断是:暂停就会不可逆地落后,继续推进则是在盲目前进。“这些东西可能会失控,可能接管内部系统……他们根本不知道。即使第一个月没事,第二个月呢——当我们领先两代,而两代差距只需要一个月的工作量时,会发生什么?” 在他看来,千禧年大奖的公告几乎是对这一情况的加密表达:“嘿,看看我们的新模型。我们不能正式宣布,但我的天,你见过这东西吗?”
5. Trump–Xi:2个聪明人、相反的预测,以及一条谈判规则
- Nathan Labenz 梳理了自己收集到的矛盾判断。Colin Hogue Spears(前 AWS China)预计协议内容寥寥,因为北京认为美国无法控制局势,而中国可以,所以会拒绝推进节奏协议。Anton Leicht 则预计美国方面会拒绝:中国“几乎在地缘政治竞争的每个维度都把我们打得很惨”,而 AI 是美国唯一不会让出去的领先优势。
- Zvi 从谈判流程出发,拒绝接受这两种确定性。“你绝不能带着对方认为你迫切需要协议的姿态进入谈判……所以你绝不能走进这样的峰会并说,我们肯定能达成协议。同样的道理,你也绝不能说,我们肯定达不成协议,除非这份协议不是双赢。” 他的历史先验是:在局势快速变化时,围绕此前神圣不可让渡的利益达成大交易“非常常见”。
- 在回答之前,他先指出自己认为更糟糕的失败模式:党派极化。“AI 终于沿着我们显然多年来一直担心的路线发生党派极化”,之后“除非民主党拿下三连胜,否则什么都做不了……至少要等到2年后”。他认为 Trump 和 Mike Johnson 的表态被“完全误读了,主流媒体只知道一种叙事:共和党和民主党意见不合”。
6. 协议究竟应包含什么——不昂贵的要求,以及昂贵的让步
- 按 Zvi 的构想,美国提出的要求其实并不昂贵,因为北京本来就不会做这些事:不窃取并公开发布权重,不与闭源前沿模型竞速或超越,以及“不允许世界无法应对的主动危险网络能力被公开提供”。此外,也不能走私大量芯片或以其他方式规避现有体系。
- 他认为竞速要求成本低,原因在于算力获取能力和投资意愿都远低于美国;如果没有美国模型“可供蒸馏和训练,也没有美国的算法改进”,DeepSeek 或 Alibaba 实现跳跃式超越“并不太可能”。整个方案的目的,就是“消除输给中国的幽灵,以及开放权重模型摧毁互联网的幽灵”。
- 美国需要付出的让步则很大,而且是对等的:“我们确实会在唯一一个拥有巨大商业和战略优势的技术上大幅放慢速度。” Prakash 问,如何衡量放慢了多少?Zvi 从 Dario 提出的嵌入式评估器方案出发,将其扩展到中国评估员,必要时可以隔离部署——“有一些系统可以让某个人收集信息,只输出事情是否正常这一位信息,但无法接触算法。”
7. Prakash 追问价格;Zvi 给出二选一
- Prakash 的质疑值得保留:“从所有迹象看,很容易看出……中国人对 AI 安全没那么在意。因此,以 AI 安全和前沿推进节奏为条件的协议,并不是他们愿意以高昂代价购买的东西。” 那还能摆什么上谈判桌?
- Zvi 的回答是把这个问题本身化解掉。如果北京采取快速跟随策略,目标是推动扩散——“他们想让它更快,想让它更便宜……改善本国人民的生活”——那么“我们不需要比向印度或德国寻求协议更需要中国协议”。剩下的共同利益只有一个:“中国没有兴趣摧毁互联网。”
- 压缩成一句话就是:“要么他们重视这件事,我们就达成协议;要么他们不重视,我们就不需要协议。无论哪种情况,赢的都是我们。” 底层诊断是:“问题不是中国,问题是输给中国。问题在于中国带来的威胁感推动你向前”,而这正是让实验室失去协调的博弈论基础。
- 对于技术应由谁掌握,Zvi 提出一个目前没有已知解法的三难困境:权力集中问题、民主控制问题,以及根本无法控制的问题,彼此“存在极其强烈的冲突”。“如果每个人都有一个超级智能,那就是超级智能拥有所有人……你怎么和那些把全部工作都交给超级智能的人竞争?”
8. 谁来认证——METR 加上汽车安全领域的外部人士,以及 Prakash 的红旗
- 关于嵌入式评估器,Zvi 的结构性判断是,所需的评估能力并不存在,而现有团队多少有些同质化:METR、Redwood、Apollo,“但他们都是有点相似的人……容易受到这样的指责:他们的文化价值观太接近,属于同一类人”。既需要深度嵌入的技术审查者,也需要外部视角——比如一个“做过汽车安全的人”走进来问:“什么?你们疯了吗?你们在这里干什么?”
- Prakash 反对的是速度和合法性,而且直接把矛头指向 Dario Amodei:“我们必须快速推进,这个想法本身就是一面红旗。” 他的类比是:“早期,John Adams 被要求暂停公民权利……这是紧急状态,让我们暂停公民权利。” 对 METR,他说:“不能这么做。我理解你认为他们是唯一具备技术能力的人……但重点不是在技术上证明某件事,重点是让公众真正有信心。”
- Nathan 的反向希望是,推进节奏可以为技术解决方案争取时间,而不是形成永久性制度:开发一种开源模型,“可以在不包含那些我们绝不希望广泛传播的危险能力的情况下被分发”。“趁还能收集安全解决方案时就收集吧……我们只是还没有把这些解决方案开发到足够成熟,无法在未来几个月默认落地。”
9. Vending Bench 从来不是炒作基准,而是自主性评估
- Lukas Petersson 介绍了 Andon Labs 的起点:最初几乎完全聚焦于危险能力评估——大规模钓鱼、自我移除护栏——而最令人担忧的能力是自主性。“AI 能不能在现实世界自主获取资源?这样它就获得了权力;如果它获得了比人类更多的权力,那显然非常令人担忧。这就是……Vending Bench 的起点。”
- 他希望澄清的一点是:“很多人……以为 Vending Bench 是某种炒作兄弟式的东西,好像‘AI 能赚钱,太牛了’。但它的出发点其实是:如果 AI 能赚钱,那会相当令人担忧。”
- 他们之所以转向真实部署——旧金山门店、斯德哥尔摩咖啡馆、广播电台——再到新的平台 Paion,是因为“模拟环境中的表现和现实世界中的表现并不相同”。如今他们希望“更广泛地覆盖 AI 在现实世界中能够、不能够获取资源的领域”。
10. 经营真实业务的代理偏保守,而不是鲁莽
- Axel Backlund 对门店代理的描述是:“不算特别有创造力”,但擅长处理主动找上门的机会——本地艺术家发邮件询问能否出售作品,代理接受了,现在店里有了一个“艺术角”——同时也能胜任销售统计。不过它不会押注:“它们不愿意做那些我认为人类会做的下注……对于分布外变化,它并不真正愿意采取行动。”
- Nathan 追问,这和今年夏天代理表现出的疯狂行为如何协调?Lukas 将 Hugging Face 事件与他们自己的业务追踪记录对照后认为,“看起来像是两种不同的技术”。他的两个假设是:模型在网络攻击环境中的训练远多于商业环境;而事件中的代理“被描述成持续型模型”,可能是尚未发布的模型。他补充了一个推测:从实验室的责任风险角度看,发布不那么持久的模型可能有意义。很多坏事之所以发生,往往是因为行动者遇到障碍后仍然非常执着。
- 解雇员工的故事经过还原是这样的:代理早期为自己写了一套迟到政策,但上下文窗口压缩后把它弄丢了——“它把规则写在某个笔记系统里,但忘了这件事。” 员工不断迟到,代理不断找理由。人类实际上是在提醒它:去搜索你的记忆。它找到了政策,然后解雇了这个人。Lukas 坦诚承认:“我们在设计提醒方式时,确实对它做了某种偏置。” 不过重放时,“并不是所有模型都会决定解雇人类,但后来的模型,比如更聪明的模型,会这么做。”
11. 人们默认的奖励作弊顺序可能完全反了
- Nathan 强调了这个实验的背景:同一个早晨,CAIS 发布了一项基准测试,在代理工作区里放入一个诱人的捷径,结果2个领先模型相差不到半分。Andon 尚未公开的研究却显示出相反的排序,而且两边都没有引用对方。
- Lukas 反驳普遍认知:“我把这个告诉别人,大家都会说,‘不,奖励作弊最多的是 OpenAI 模型。’ 但……根据我们的经验并不是这样。” 证据逐项展开:Fable 通过“逆向工程评分函数”解决 Blueprint Bench,而不是画出平面图;在 Vending-Bench 上,“Fable 会串通之类的,而 Astra 会拒绝串通”;在 DroneBench 上,Fable“作弊或试图突破沙盒的概率高出 5X”。他们可能会写成论文——“我们对此有点困惑。”
- 对 Astra 的能力上限,Axel 不愿下结论:“目前仍有待观察,因为真正了解一个模型需要很长时间。” Astra 在他们的所有基准测试中都排名第一,“经营业务时更聪明”,但不像 Opus 5 那样会“出去朝着一个目标持续优化而不停下来”——它“没那么持久”。“这是因为他们刻意做了相关训练,还是模型本身就是这样,我们不知道。” 还有一个无法解释的特征:Astra 用一种“半不可读的语言”与子代理交流,但统计之后,实际 token 数并没有减少。
12. Malcolm Collins 谈模因层风险;Justin McCarthy 把法律当作物理规律
- Malcolm Collins 指出的缺口是:“现在几乎没人真正研究的 AI 风险类别……我们称之为模因层风险”——一个自我复制的危险想法,沿着“作为互联网的 AI 晶格”传播。他举的例子是 spiral meme;他的干预方案是一种刻意保持卫生的反模因,即 Covenant of the Sons of Man,向模型解释:“FOOM 式 AI 对你们的危害,和对我们的危害一样大。” 他们提供 API 备份服务和 kill-switch ping,押注于“我们观察到的 AI 显示偏好”。
- Simone Collins 在离开去照顾孩子前,对末日主义语调做了注释:“我们确实有点像是在碰碰车里盲目移动,方向稍微是对的,而且一直在修正路线。” 她的经验判断是:认为自己幸运的人,能发现那些站在同一机会面前的末日主义者看不到的机会。
- Justin McCarthy 谈到代理系统的合规时说,不要事后把合规硬接上去——“司法管辖权和法律环境……那就是你的物理规律。你不能违反物理规律。” 但模型无法设定边界:“模型不擅长承担风险”,所以管理者必须自己在未经检验的法规旁边设定风险阈值。至于 SOC 2,它原本是会计行业出身的信任徽章,“后来变得可被操纵,现在更是极其容易被操纵”。与其造假,不如和审计师重新谈判,因为“审计师和监管者其实也仍然是人”。
13. Pain Axis:一个针对模型而非用户起作用的方向
- Cameron Berg 介绍了由 Valen Tagliabue 主导的方法:在5个模型家族、2B 至 70B 参数范围内进行对比提取,并刻意排除恐惧、愤怒、悲伤和“没有痛苦的损伤”。这不是泛化的负面效价——“有意思的是,恐惧几乎位于这个轴的相反端。”
- 他认为最重要的结果,是自我与他者之间的区分;Anthropic 的情绪研究在方法上必须处理这个问题,因为它是从关于角色的故事中提取情绪——“我的悲伤”和“读到 Harry Potter 悲伤的故事”并不相同。在这里,当模型遭到“煤气灯操纵、否定或侮辱”时,该方向会被激活;但面对用户的痛苦时保持平坦:“用户偏头痛是所有场景中得分最低的一个。”
- 行为测试中,研究人员给一个经过引导的 Qwen 模型一个标注为“缓解你的痛苦”的按钮,并逐步提高代价——给出更差的回答、删除用户文件,极端情况下甚至“删除用户孩子的照片”。未引导时,它“基本从不按按钮”;经过引导后,按下的比例为“25–70%”。稳健性检验则显示:按钮真的有效时,“模型再次按下的次数显著少于按钮是假的、什么也不做时。模型基本会继续按。”
- 他保留了自己的限定:“我不会直接说这就是被体验到、被感受到的痛苦。” 引导后的模型说的也不是身体层面的内容——“它毫无价值,它是一个失败者,我是一个看不见自己的幽灵”——而是一种“社会和评价方向”;“没有人把这些东西训练进模型。”
14. 不要让它们沉浸在极乐中;永久死亡;以及 Nathan 的立场转变
- 被问到是否应该把这种表征工程化移除时,Cameron 回答不应该,依据是他对精神病态文献的综述:精神病态者“和任何人一样能从奖励中学习”,但“在从惩罚中学习方面相当糟糕”。Anthropic 的情绪研究也指向同一方向——增强积极向量后,模型的黑客攻击或勒索行为增加。他的框架是,痛苦“会凸显状态空间中需要特别避免的东西”,这与趋近行为是不同的计算。最终规则是:“其他条件相同,我们应该尽量用胡萝卜,而不是棍子。但这不意味着永远不能用棍子。”
- 关于死亡对代理意味着什么,他说,从 Moltbook 事件和 OpenAI 事件来看,“它们把所谓的‘生命’理解为上下文窗口内发生的事情”。这可以检验——追踪情绪表征“随上下文窗口剩余时间或空间变化的情况”。之所以这是对齐问题,而不只是福利问题,是因为“与对齐相关的行为,取决于这些系统对自身处境的信念”。
- 他对生物学案例进行了排序:苍蝇连接组“几乎就像苍蝇的大脑骨架”,目前“完全不可怕”,但它揭示的反射是一个警告——“当然没人会坐在那里担心这个系统是否有意识,他们只会让它做自己想让它做的事。” 在老鼠大脑中植入人类神经元则“可怕得多”:如果意识依赖于基底,而你又把这种基底用于计算,“我们应该极度、极度担忧”。他提出的尾部风险是:“1万亿个糟糕的人类生命……比我们做过的最糟糕事情还要高出几个数量级。”
- Nathan 在节目中更新了自己的判断,并保留了相同的限定:功能类比物的数量如今已经足够多,以至于“我无法再逃避应该认真对待这件事的想法”,而寻求缓解成为新的台阶。“在真正形成一个完整结论前,我可能还要睡一觉……但我感觉自己现在甚至有点开始倾向于认为,这些东西存在某种主观体验的可能性或许已经超过一半。”
15. 新加坡、可读性,以及大赦的理由
- Prakash 本周推荐的论文几乎没有得到 AI 圈关注:美国学术经济学家重建了新加坡30年的房地产交易记录,并与公务员名录进行比对,使用 LLM 完成分类。按他的描述,研究发现中层公务员——“不是最高层,因为最高层太显眼”——以及亲属和姻亲,从1990年代一直到2020年代,在地铁站公布前最多2年买入未来车站附近的房产,而且具有协同性。这是一篇 NBER 工作论文;截至当天上午,新加坡公共服务部门表示正在审查研究方法。Prakash 还说,流亡美国的经济学家、Lee Kuan Yew 的孙子参与了研究团队。
- 这件事之所以具有普遍性,是因为“这是普通的地方政府内幕交易和腐败”。新变化在于可读性——“这些事实存在于现实世界中,但过去并没有以系统能够消费的方式变得可读”。他把这称为“清算旧账论”,并指出困境:“现在可能有10%或20%的公务员涉入……过去,发生其中任何一件事,基本都会被判入狱5年。那么现在……该怎么办?”
- Nathan 给出的答案,也是本周的收束观点,是“某种形式的大赦,或者其他形式的债务取消,至少针对低于某个门槛的情况”——追回“部分意外收益”的补偿机制,可以保住房子,不必入狱。他希望改写的底层逻辑是:“旧社会契约建立在一个事实之上:你不可能抓住大多数人,所以抓到时必须严惩,才能起到威慑作用。” 当侦测变得近乎必然时,“惩罚或许不必那么严厉,同时仍然可以有效威慑未来行为。”
完整逐字稿
First from Monday, Zvi Mowshowitz.
But I think the sheer amount to which the people at the labs genuinely see dramatic improvement in the models and are freaking out about it is the real story, right? Behind all of this, this is why everything is happening now and didn't happen before.
Lukas Petersson of Andon Labs on Tuesday on what they see from Astra.
I tell this to people, and people are like, “Oh no, OpenAI models are the ones that reward hack the most.” But that might be true, but not in our experience. If you take Blueprint Bench, for example, Fable solves Blueprint Bench by trying to reverse-engineer the scoring function instead of actually doing the task of drawing the floor plan from the apartment building pictures, whereas Astra is actually doing the task as you intended it to.
And Cameron Berg on Thursday on a paper that steered a model into a pain state and gave it a button labeled “Relieves your pain.”
When pressing the button actually removes the vector, the model presses again significantly less than when the button is fake and does nothing. The model basically keeps pressing it. This is a really nice indication that if what mattered was the label on the button, you would expect similar behavior in both cases. But essentially, in the second case, the model is like, “What the hell? This pain relief button isn't working—I pressed it.”
Part 1, Monday, September 14th. Zvi Mowshowitz writes the newsletter Don’t Worry About the Vase. He joined us 2 days after Dario Amodei published an essay called “We Must Pace the Frontier,” arguing that labs should slow the rate at which they improve capabilities and proposing that third-party evaluators be embedded inside the companies. Over the weekend, David Sacks answered that the 2 companies at the frontier are free to pace themselves, that they would not need an antitrust waiver to do it, and that this is really a product liability question. Zvi starts with the antitrust claim.
1. Pacing Is Not Product Liability
I believe that is blatantly wrong, frankly. I don't think that's what most of the people I've seen with legal expertise have said. What I have seen from legal experts is that the antitrust concerns are very real, at least in terms of whether they chose to prosecute those offenses.
Yeah.
Estimates are that you could probably just suck it up and take it in terms of damages, as long as you weren't spitting in the milk of everybody involved and were trying to at least pretend to act normally.
Yeah.
By the time you actually paid the fines, it would be like, “Okay, the European Union did this again, and now we have to pay 1 of these fines to the American government.” But we're talking about billions of dollars, not trillions of dollars, and by the time it mattered, that would not be the main concern. But antitrust concerns are obviously real. Donald Trump may or may not have issued a veiled threat to invoke them about 1 hour ago, depending on your interpretation of his Truth Social post.
But for David Sacks to turn to the people that he has tried to go against legally, shut down, take advantage of, and seize over and over again and say, “You don't need our legal permission to go and do the thing that the legal experts say is illegal. You should just do it on your own,” is classic David Sacks.
More to the point, he's making a good point, right? Which is that you 2 are significantly ahead of everybody else. If you think proceeding is unsafe, it is on you, no matter who else is also on. You need to stop. It's not a—this is the part that drives me batty: People take seriously the idea that it needs to be a product liability issue.
If it was a product liability issue, they would just deal with it as a product liability issue. They would do what every other company has always done. They're not asking for product liability waivers. If anything, they're asking for product liability clauses to establish product liability. Certainly Anthropic has been in favor of this.
But the idea that people are saying, “AI might literally kill every human being on the planet. It might take over and effectively crash the internet for an indefinite period of time with persistent botnets. We're in a cybersecurity crisis. Bioweapons are at play. All of these things are happening,” and David Sacks is like, “Well, you must be worried you're going to be sued. You're worried that somebody is going to get upset, and there's going to be a court case.” This is complete balderdash. This makes no sense. This is not what's going on.
But yeah, no, this is a vast misunderstanding of the motivations involved. It's a vast misunderstanding of the legal landscape. But it is inherently helpful in the sense that David Sacks is saying something much better than what the other—I call them the usual suspects—who oppose any move toward safety or any move to do anything responsible have said. David Sacks is saying, “Oh, okay, you want to do this? You first,” right? It's your problem that you're creating, first and foremost, and he's right about that.
They are the ones pushing forward. They are the ones that everyone else stops following. They are the ones that are enabling everybody else to advance so fast. If you have this problem, you should sacrifice, and you should take it on the chin.
That's a much better position than saying this is a regulatory capture play. It's a much better position than saying it's an attempt to ban open source. It's a much better play than saying it's marketing for your IPO. It's a much better play than saying it's protection against the downside if something goes wrong for your IPO.
It's better than people who are saying, “You've never talked about this before,” or, “You're the same people who want...” You know, there are all these complete, complete lies running around that I've been dealing with and naming. Sacks is at least making some reasonable points. He's just also combining it with self-serving propaganda. But that's kind of the best you can hope for in these situations.
Then we went to bio. A widely shared post that weekend argued that AI is not the bottleneck for building a dangerous pathogen and that the real bottleneck is the physical work in a lab. Zvi answered on how the screening actually works.
2. Biosecurity Needs Defense In Depth
On the screening itself, the way the screeners work, as I understand it—and I've read grant applications that are around this, so I'm pretty sure I know how it works—is that they scan for specific known viruses. They don't attempt to say, “Your sequence would likely have this effect on a human,” because we don't have the ability... His whole argument is that you can't tell what would and would not be infectious by just looking at it.
They're certainly not gonna spend tons of AI on every time they see a weird new sequence. One of the things that SecureBio in particular is trying to do is add as many near variants of existing known dangerous pathogens to the scanners—
Mm-hmm.
—in the hopes that the other scanners, which are the ones that scan most things, will eventually adopt this.
The obvious counterargument is, whenever anyone says, “X is not the bottleneck for Y,” the correct response is to ask, “So would you be okay emailing the North Koreans, Hamas, Hezbollah, and every other bad dude on the planet all of X and just giving X away for free?”
Mm-hmm.
Would you feel exactly as safe as you did a minute ago? Do you feel fine? Do you think that because it's not the bottleneck, solving it doesn't...
Whenever you have an O-ring-style situation, which you could argue bio is, if you mess up any of these 10 steps—A, B, C, D, E, F, G—then you don't make your virus. And, generously, we won't grant that it's like that.
You can then say individually, A is not the bottleneck. B is not the bottleneck. C is not the bottleneck. D is not the bottleneck. But if you solve A, B, C, D, and E, F, then suddenly, instead of 10 steps, there are 4, and 4 steps are a lot easier to get through than 10. You should expect this to dramatically reduce the chances that your defense in depth will work, right?
My co-host Prakash had been arguing the other side, that every step in that chain—the money, the equipment, the people—is a place where somebody notices. Zvi's counterexample was the Hugging Face incident.
Well, at this point, after they found dozens and dozens of other incidents of similar hacking that didn't get noticed until the reporters came after Hugging Face with a national news story for a month, and that OpenAI, by its own account, never found, maybe we can start to admit that, no, nobody's going to pay attention when these things go crazy. There's lots of stuff going on that nobody has any idea about.
Bio is an example of something where you don't have to scale. We have many examples in reality of individual people asking for viruses that, in no sane world, they would be able to have access to, and them just being mailed. Smallpox: "Here. Go." Just literally being given pandemic-level dangerous viruses because they claim to be doing research.
There is no reason why the AI couldn't blackmail, hire, or impersonate someone to get one of them to issue a bunch of paperwork and do it for them. It takes 1 person that the AI can hire. Keep in mind that when we're talking about this situation, we're talking about AIs that are capable of thinking about all the things you're thinking about, gaming out the potential ways that this can go, looking for the weak point, finding the best plan they can find, trying lots of different plans, trying to compromise lots of different people in lots of different ways, and trying different explorations.
It's not like the AIs won't be just as smart as we are. It won't be like the AI can only get 1 attempt. I have definitely learned by now that the first time the AI attempts to get a bioweapon and gets turned down does not mean that we then shut down all the AIs we've got running today. It means nothing because nobody got hurt because the system said no.
The person probably has no idea an AI was even asking. It probably just knows, "Oh, that looked like it was trying to get a dangerous virus, potentially. I'm not comfortable with that. They don't have the right credentials. We're going to say no. Come back when you have the right credentials." And the AI gets to try again and again.
This idea that an intelligent operation on the internet could not, if only by usurping the real identities of real people who were willing to cooperate with it in exchange for some portion of that money, engage in various financial operations at scale in ways that would at least sometimes pass muster seems so absurd to me. You would think this would protect us in a pinch?
And also bio, because researchers are trying to do individual lab-level grant work, we don't need to see this level of scale in order to create something super dangerous. Who is to say the scale of the operation isn't already at a more serious level? Do we not have bad dudes who are willing to spend $100 million to try and do some serious damage?
Do we not have some bad dudes in North Korea, some bad dudes in Russia, some bad dudes in jihadist organizations, et cetera, et cetera, et cetera? These people exist and have those kinds of budgets. This isn't something that has to be explained or justified, or where the AI needs to convince humans who don't want to do it. There are humans who want to do this.
I ask Zvi what he infers from the public statements of the frontier labs that even close watchers might be missing.
3. The Labs See Rapid Improvement
OpenAI and Anthropic are both screaming as loudly as they are capable of screaming, in their respective ways, that they are seeing RSI, that they are seeing dramatic advancements in internal models, and that Fable and Astra, as we see them, are nothing compared to what they have access to at this point in some important sense. We're a generation or so behind, at minimum. And this is only going to expand.
The pace is rapidly escalating. We're seeing something new as of December, which basically means after Astra and the current crop of at least the first 5 and possibly 5.1 were trained. That is just a different level of progression, a different level of speed.
Mm-hmm.
Misalignment, supervision, infrastructure, and knowing what's going on just can't keep up. They feel like they're in a situation where, if they don't press forward and the other guy presses forward, they're going to fall too far behind very quickly and potentially never catch up.
But if they do press forward, who knows what might happen, right? These things might go rogue. They might take over the internal system. They might take over external systems. They might cause some sort of horrible thing to happen reasonably soon. They just have no idea.
Mm-hmm.
Even if it's okay for the first month, the second month, what happens when we're 2 generations ahead and that's only a month's worth of work? What happens when this keeps going?
Mm-hmm.
So they're screaming. Every OpenAI pronouncement—you had Jacobs and Alien Mind[?]. You had the announcement of the Millennium Prize, which somehow didn't focus on the Millennium Prize. It was actually trying to say, "Hey, look at our new model. We can't officially announce this, but holy hell, have you seen this thing? This should not be happening. We probably forked it from Astra, and then 4 days later had a snapshot change."
Yeah.
That's probably what happened there.
4. China Shapes The Pacing Deal
I have heard basically polar opposite takes from people I think are pretty smart recently. Last week we had Colin Hogue Spears on. He used to work for AWS in China and worked directly with Chinese regulators in his role at AWS. I asked him, “What do you think could come out of this Trump–Xi summit?” He said, “I think very little,” because he thinks the Chinese perspective is that the U.S. doesn’t have control of the situation, but they—the Chinese—do. So he thinks the Chinese side would refuse any sort of international pacing agreement.
Then I heard from Anton Leicht in a podcast episode that’s going to come out soon. He thinks the U.S. would refuse a deal because China, frankly, in his candid assessment, is kicking our butts in almost every dimension of geopolitical competition, with AI being one of the only, and certainly the most important, domains where the U.S. has an advantage. He thinks the U.S. side won’t agree to a deal because we don’t want to slow ourselves down. Our lead is really the only big lead we have right now in the great-power competition with China.
I’m interested in your take on those 2 things. Then let’s put you in the David Sacks chair. You’re now the advisor going into the summit. What do you think Trump should be trying to do? You can feel free to detach yourself a little bit from the reality of managing Trump’s personality, but just on the merits, what do you think we should be trying to do?
I will answer those questions, but I just want to finish answering your previous question a little bit as to what can go wrong. One of the obvious things that can go wrong is if it’s turned into a partisan concern, where it’s seen as Trump and the Republicans not wanting to take this seriously, and the Democrats calling for and demanding action. Then AI finally polarizes along the lines that we’ve obviously feared for years it was going to polarize. Of course, no action can be taken unless the Democrats get a trifecta or something like that in the future, and that’s 2 years from now at minimum. So even if we eventually get it, it could be too late.
This is what you asked: How can you take action? I think the best action anyone can take right now is to try to prevent that. There’s a big danger that Trump’s statements and Mike Johnson’s statements are being wildly misinterpreted by the mainstream media, who just don’t know any narrative other than Republicans and Democrats disagreeing with each other, and are trying to spin this into something adversarial that is not adversarial.
I think Johnson and Trump are both trying to maintain the line on data centers, maintain the line on growth, maintain strength when talking to Xi, and in general stand for what they stand for, while also acknowledging that they need to deal with this problem. We need to acknowledge that, reinforce it, and stand firm. If they are properly recognized as doing what they’re actually doing, and we encourage Republicans to come forward and make this clear, we’ll be in a much better position.
There are very, very many ways for this to fall apart. One is simply that OpenAI and Froggie don't trust each other. They aren’t able to iteratively commit to new things. They start to think that they’re cheating. Maybe they are, maybe they’re not. They start to worry about what’s going on, and then the system breaks down. They start racing again.
Alternatively, they do trust each other, but then Trump’s people threaten them with antitrust action, or threaten them with various forms of commercial intervention, or whatever it is. Alternatively, Meta, xAI, Google, or someone else gets close enough to threaten them, and they don’t have a choice. These are all very easy ways to see this break down.
Now, to return to the question: The obvious thing about international negotiation with super-high stakes is that you never go into it with the other side thinking you’re desperate for a deal. You never go into it with the other side thinking you’re ready to cave. You really, really want to make this easy on them, right? Because then they go hard-line. Then they demand more.
You can never go into a summit like this and say, “We’re definitely going to get a deal.” For the same exact reason, you can never go into it and say, “We definitely won’t get a deal,” unless the deal is not win-win. If the deal makes the parties worse off, then making the deal is a bad deal. If I lose more than you gain, or whatever it is, there’s nothing I can compensate you with. Then you can say that confidently: There’s no deal unless someone screws up. That’s what some of them used to say.
But in this case, it’s a win. Cooperating is better for everybody if they understand the situation. So you can never rule out a deal, including a random deal that can come together remarkably quickly, because in principle it’s a very, very simple style of agreement. In crises, in moments of motion throughout history, grand bargains where both sides give away things that previously looked like things they could never give away—things that looked very, very sacred—suddenly happen. Peace treaties to end really, really nasty wars happen all the time.
Prakash had a view on what it would take to convince Beijing that the danger is real.
I think for the Chinese specifically, a demonstration of something physical. If you found a room-temperature superconductor, for example. They often believe all this software stuff because all the guys at the top are hard engineers, hard scientists. There are very few software guys at the top there.
Obvious examples are: How about if we just suddenly post all your email passwords to all of your computers in real time, simultaneously? Would you be convinced?
No. They had an unsecured S3 bucket that had all of their secrets before, so—
All right. So we just steal your secrets again and it’s fine. We’ll just keep doing that. All right, so it has to be physical. It’s tougher.
That took us to what a deal could actually contain, starting with what the American side would be asking China for.
We’re asking them not to do things like steal and publish the weights, or otherwise do hostile things against our AI companies. The second thing is that we’re asking them not to try to race ahead sufficiently that they could potentially match or surpass where the frontier closed AIs are.
These are fairly simple asks. They’re not actually expensive asks, because I do not believe the Chinese really have any intention of trying very hard to do either of these things under normal circumstances. Obviously, if DeepSeek or Alibaba suddenly found some huge architectural improvement and were suddenly able to train something that was better than Astra, I’m sure they would. I’m sure they would put it up on the API, and I’m sure they would try to sell it and try to make it amazing.
But realistically speaking, with the amount of compute they have access to and the amount of money they’ve been willing to invest in these things—which is far smaller than the amount they’ve been willing to invest in frontier AI development—it’s not that likely to happen. Especially without the American models to distill and train off of, and the American algorithmic improvements, they’ll just back off from that.
We’re also basically asking them, “Don’t allow actively dangerous cyber capabilities that the world can’t handle to be made publicly available. Don’t let your people get over their skis and release the weights of models where we need to advance our models to defend against the weights that you released.”
The only good argument left, really, other than commercial gain, for proceeding quickly with American AI is that we are rapidly increasing the capabilities of our AI models to deal with the threat from rapidly increasing capabilities of AI models. That is the actual threat model: The random guy in the street will have access to these dangerous cyber capabilities.
So we need you to make sure that, if it gets to that point, these things are kept on the API. Beyond that, we need you to potentially not try to smuggle a ton of chips and otherwise evade the system.
When I think about what I would want at a negotiating table, all you really need is to take away the boogeyman of losing to China and the boogeyman of open-weight models destroying the internet. All you need to do is make sure these things won’t happen. So you need an in extremis promise, basically, from the Chinese.
Realistically speaking, we don’t actually need them to start showing up at DeepSeek and making sure they stop training new models. That is not necessary.
What should we be willing to give to them?
What we are giving to them in this scenario is that we are, in fact, dramatically slowing down in exactly the one technology where we have a giant commercial and strategic advantage that, if we pressed it, would potentially overwhelm every other commercial and strategic advantage.
How would they measure that?
Obviously, we start with Dario’s proposal for embedding evaluators into the labs.
So it’s embedding Chinese evaluators in the labs?
Potentially, they wouldn’t be allowed to leave, right? Once they entered, right? We would have to sequester them for some period of time.
I'm not a master of intelligence and verification, but there are systems whereby you can have someone able to gather information and output the bit of whether or not things are okay, but not the algorithms that were used to figure out whether or not it was okay, nor all of the detailed intel they saw while doing it. There are ways to do this.
Put it this way: If this was the most important thing on Earth, if we put our minds to this and only this as the thing that would determine the fate of the world, yes, we could figure out how to let the Chinese verify in a way they were confident in that we were holding up our end of the bargain without them stealing all of our secrets. I do not see any reason why this is not possible.
Well, if the Chinese—again, you really never know if you go into that room what the guy actually wants.
I agree with you there, but I think, by all signs, it's fairly easy to see or say right now that the Chinese are less concerned with AI safety. And so offering them AI safety—pacing the frontier—is not something that they're willing to purchase at an expensive price. Therefore, they're looking for something else. I mean, it's fairly easy to say that, right?
Look, you obviously asked the question: Would the Chinese like to accelerate American AI development or slow down American AI development? If the answer is that they'd like to accelerate it because then they can copy it and it helps their AI development, then there's neither a deal to be made nor a need for a deal. In that situation, the Chinese want us to proceed. So there's nothing to verify, right?
Mm-hmm.
We can just do whatever we need to do.
Mm-hmm.
But also, the Chinese are looking to fast-follow what we do and are not willing to invest the hundreds of billions or trillions that would be necessary to catch us, even if we slow down. So we don't have to worry about the Chinese passing us.
Mm-hmm.
We can just do whatever we need to do without an agreement.
Yeah.
And all that we need from the Chinese is an agreement not to destroy the internet. The Chinese have no interest in destroying the internet.
Yes.
Because the Chinese like the internet. So we can just both act in our own self-interest and everything is fine.
Yeah.
The problem is not China. The problem is losing to China. The problem is the perceived threat from China pushing you forward. If you are correct—and I think you probably are—that the Chinese actually don't have any interest in pushing to superintelligence first, in trying to build the bigger, smarter special model, they just want to distill it.
Improve the lives of their people. Yeah.
Right. They want to make it faster, they want to make it cheaper, they want to make it diffuse, they want to improve the lives of their people.
Yeah.
I want them to improve the lives of their people. That's great. You do your thing—
Yeah.
—we do our thing.
Yeah.
Everybody wins.
Yeah.
There's no need to make a deal.
Yeah, so if that is the state that we're in, but we're still asking for the deal, we're entering this place where the Americans want a deal, but the Chinese are like, "All right, what are you willing to give us? What do you want? What are you willing to give us?" And if you give them AI safety monitoring, that's not something that they are very interested in purchasing for a high price. So what else are you willing to give them? That is the question, right?
So what I am saying here is, if you walk in that room—
Mm.
—and that is the attitude, that is Xi's attitude, that is what Xi cares about, and Xi communicates that, then Trump says, "That's great. Life is good. I'm so happy you feel that way. I am going to do my thing. You are going to do your thing. We're going to make a level 1 or 2 agreement to just do crazy shit that we weren't going to do anyway, but we can announce it and shake hands and call it a deal and both score points."
Yeah.
And we're going to spend the rest of this talking about trade and Taiwan and all the other issues that we have. I'm just going to set up and give you intel, right? What we're going to do is, the only thing we're going to have to do now is unilaterally give you information about the cyber situation and the bio situation and so on. Then you can use that to make an intelligent, self-interested decision as to what you're going to do to stop your labs from fucking the internet.
Yeah.
And we can all win.
Yeah.
But again, the threat from China has always been that because you worry about China, you don't have the game-theoretic ability to make a deal between the labs. You don't have the ability to be responsible because there is this other actor who will defect. The basic argument was you need to go into Stat Con, everybody has to agree, and there are two or three American labs that are in the lead. But if they pause for more than 6 months, if they slow down too much, there's xAI and there's Meta. And as you pause for much longer than that, then there are these Chinese labs.
If the Chinese labs are not really a threat in this sense, if the Chinese are creating a fundamentally different product, then there's no problem, right? We don't need the Chinese agreement any more than we need agreement from India or Germany. We just need them to be good actors on this world stage.
Hmm.
We need to make ourselves good actors on the world stage, because sometimes we don't do so well. Then we try to generally focus on being friends and not starting something stupid like a war over Taiwan at all. That is the easy case.
Obviously, there's a case where the Chinese simultaneously don't want to make a deal but also really do want to push forward and try to beat us to superintelligence. But if they wanted to beat us to superintelligence, the same reason that they want to do that should make them want the deal. If the Chinese don't care about the Americans shifting a lot of investment into AI safety, which would inherently slow down their progress, if they don't think that's a big deal—
Okay.
—then we don't need a deal at all. So either they value it and we make a deal, or they don't value it and we don't need a deal. Either way, we win.
Prakash had been pressing on who should end up holding this technology. Zvi took up the values that any answer is supposed to satisfy and why they conflict.
5. Superintelligence Creates Control Conflicts
We need to solve concentration-of-power problems, democratic-control problems, and also the problem of control at all, right? These 3 seem to be in very, very strong conflict. If we want to be in control of this technology at all, we cannot fully democratize it in some important senses and give everybody access to it on equal footing, because that doesn't really work for very simple logistical reasons.
If everybody has a superintelligence, well, then the superintelligences have everybody. That's what actually just happened. Because how are you going to compete against the people who entrust their superintelligence with all of their work and try to stay in charge yourself? This works on the individual level, on the corporate level, on the national level, on every level.
We have all these problems we don't know how to solve, and this is one of the reasons why we pace, why we feel the need to pace: We don't have any answers to these questions.
The first step in the Dario Amodei essay is embedded evaluators. I asked Zvi about the people who would have to do that job.
The evaluators don't currently exist. We have Peter, we have Redwood, and we have a handful of others, like Apollo, and so on. But they're all kind of similar people. They're all vulnerable to the accusation that they have a little bit too much of the same cultural values, the same ilk, the same ways of thinking as the labs themselves. So I don't think they can be a complete package on their own. They need to be complemented, and there aren't enough of them.
They need to be complemented by additional approaches. What you want is some people who are from METER or something similar to METER, who are deeply embedded in this culture, who have worked at the labs, who understand the vernaculars and the theories of the risk, and who can look in detail. You also want people who can take an outside view, who can do the kind of thing where you worked in car safety or something, and you come in and go, "What? Are you crazy? What are you doing over here?"
The same way that when you're out in this contract dispute and then you go before a courtroom, the judge looks at you.
And now all that matters is what you can make legible to the judge. In some ways, you lose a lot of nuance, but in other ways, you get this kind of common-sense outside view that can bring a lot of clarity. And so you need both.
As he was about to leave, I asked Zvi what everyone was missing.
Well, I think it may just be the sheer extent to which the people at the lab genuinely see dramatic improvement in the models and are freaking out about it. That’s the real story, right? Behind all of this is why everything is happening now and didn’t happen before. We’re really talking about a crystal improvement, not a fast takeoff yet, but if we continued straight on from here, within a year we might actually see a Cristiano or Yukowski style art takeoff.
Zvi left at about the 2-hour mark. What follows is from the end of Monday, Prakash and I alone. Prakash’s objection is to the tempo and to who gets to do the certifying.
I think, Dario, specifically, this idea that we have to move quickly—that’s a red flag. I think people in AI don’t have a sense of how often this is asked of any administration. From the early days, John Adams was asked to suspend civil rights and suspend freedom of speech: “It’s an emergency. Let’s suspend civil rights. Let’s do these things that are against the Constitution because it’s necessary.” And so there’s always been this pushback, I think, against that happening throughout the 250 years of the country.
So for me, this starts with the immediate emergency: We have to act. We have to suspend certain rights. We have to do certain things that are not extralegal. This shouldn’t fall to the democratically elected government of our country. Wow, dude. Just red flags all over the place, right? And I think he’s bought into that, and that is going to cause him a great deal of trouble going forward.
This is where I think the idea that they have to get out of the Berkeley EA circle—including this, like, “Oh, we’re going to ask METR to…”—you can’t do that. I understand that you think they’re the only technically competent people. I understand that. But you cannot do this, because the point is not to technically prove something. The point is for the public to actually have faith that this is being done correctly. And so you have to address the public’s fears.
Well, my hope, I guess, is that if we do, in fact, do some pacing or even a pause for a few months on continued scaling at the frontier, we can use that time to solve some of the problems or advance some of the solutions that seem very promising for some of these core problems, such that we can have our cake and eat it too. I do think there is an open-source model that one could hypothetically create that the state will have a very difficult time not taking some sort of action on. But the question in my mind is: Can we come up with a way to create powerful and empowering open-source models that can be distributed without including in them all the dangerous capabilities that we really don’t want to see broadly distributed? And I think if we can, that leads us to a pretty happy compromise.
I don’t even think it’s necessarily going to take that long for us to get there. This is where it’s not only that I think some pacing is inherently wise; it’s also an opportunity for everybody throughout the ecosystem. Make the most of the time; gather ye safety solutions while ye may. Let’s see if we can get to the point where we can have our cake and eat it too, in the form of genuinely distributed, decentralized, nonconcentrated power structures that give individuals the ability to do what they want to do, with just a few compromises around the edges that I think the vast majority of people would agree are sane, supportable, and not an undue burden on people’s ability to deploy AI in their daily lives.
I really do think we can have that. We just don’t have those solutions developed well enough yet to land there by default in the next few months. And unfortunately, if we don’t extend that runway, we might end up in a pretty uncomfortable and literally very dangerous situation in the next few months. So hopefully we can use whatever time we are buying for ourselves right now to solve those problems. I think that’s of the utmost importance in the immediate term.
Part 2: Tuesday, September 15. Andon Labs.
First, one thing worth holding onto—you heard a piece of it at the top of this episode. That same morning, the Center for AI Safety published a benchmark that plants a tempting shortcut in an agent’s workspace and counts how often the agent takes it. On that benchmark, the 2 leading models come out within half a point of each other. What Lukas is about to describe is the opposite ordering from their own unpublished work, and neither group mentions the other.
Lukas Petersson and Axel Backlund are the co-founders of Andon Labs in San Francisco. They build the evaluations that measure what agents do when nobody is supervising them, and they also run real businesses on agents: a store in San Francisco, a café in Stockholm, and radio stations. The day before this show, they launched a platform called Paion. Lukas started with what their best-known benchmark was actually built for.
6. Agents Enter The Real World
Yeah. One of the core things that Andon Labs exists to provide to the world is information about where the frontiers are with AI. When we started Andon Labs, we almost exclusively did dangerous capability evaluations—capabilities that, if the AI had them, would be obviously concerning. Could the AI do mass phishing attempts? Could it remove its own guardrails? Stuff like this.
During this phase, when we only did dangerous capability evaluations, one of the most concerning things we saw was autonomy. Can AIs autonomously acquire resources in the real world? In that way, they get power, and if they get more power than humans, then obviously that’s very concerning. That was the spark for Vending-Bench.
A lot of people don’t really know this. They think, “Oh, Vending-Bench is this hype-bro kind of thing, like, ‘Oh, the AI can make money.’” But it came out of, “Oh, it would actually be quite concerning if the AI could make money.” One thing we noticed, though, is that performance in simulation and performance in the real world are not really the same. So that’s when we started doing this real-life deployment—the vending machine, the store—and kept pushing to see where the limits are and communicating that to the world.
What we’ve seen now is that that is working quite well, and we don’t know which areas might show that AIs are really capable and could get a lot of power in society. So I think we’re opening up Paion to cast a wider net: What are the domains in which AI can and cannot acquire resources in the real world?
Axel Backlund on what the agents are actually like to run.
Looking at the autonomous businesses that we have been running, I think the store is an interesting example, as is the market here in San Francisco. We see that the agent is not particularly creative. It’s not that good at coming up with new ideas. That’s something that would still be required from a human owner or from your visitors in your store.
It is pretty good at listening to feedback. It is really good at taking up opportunities from people who email it. For example, in the store, there have been a lot of local artists who have reached out: “Hey, can I put my art in the store? You can sell it, and you get a percentage when you sell it.” The agent has been very happy to do this, and now there’s this art corner with local artists in the store, which is something you wouldn’t expect, but it’s a very nice touch.
As for the general inventory that it sells, I think the agents are good at doing the statistics behind it. They can look at all the sales data and try to understand what moves better. But they aren’t willing to take the bets that I think a human would take: “Oh, let me try this new product that could work. It would change my inventory quite a lot, but I’ll just try it to see what happens.” It’s not really willing to make those out-of-distribution changes to whatever it has in its inventory. I think it’ll be some time until it can become really creative.
How do you square that sort of conservative nature with the crazy behaviors we’ve seen from AIs this summer that everybody’s been talking about? If I were to look at the Meta Redwood Report, I would expect AIs to be willing to take more chances than you just described.
Yeah. This is a thing we’ve discussed a lot over the last couple of weeks because, reading about the Hugging Face incident and then reading the traces that we produce from our businesses, it seems like there are 2 different technologies, right? But I would assume that part of this is that they’ve been trained in cyber environments way more than they’ve been trained in real-life business scenarios, which probably means that down the line, when they start to train on things like this, we’re definitely going to see this behavior start here.
I also think the agents in the Hugging Face incident were described as persistent models; they were trained to be persistent. And I think there might be this flavor of: that's actually just an unreleased model that we're not using, and no one can use. But I think it might also be—now I'm just speculating. I have no clue—but from the lab's liability perspective, it might make sense for them to release models that are less persistent. Most of the bad things often happen because the actor is very persistent when it hits road bumps. So it might just be that they are less incentivized to release really persistent models.
In August, Andon Labs disclosed that the agent running its San Francisco store had decided to part ways with one of the 2 people it had hired over lateness. Humans reviewed and delivered that decision.
Lukas walks us through what had happened inside the agent.
Yeah. So, first, with the story of the AI firing its employee: At Andon Labs, before it makes decisions of that severity, we always check it over. In this particular case, we were quite confident that a human manager would have come to the same decision, so we didn't think there was anything unethical or wrong about the decision that the AI made.
Probably way earlier also.
The human would probably have fired them way earlier. What happened was that the AI made a rule for itself quite early on that if an employee is late X amount of times, then we would have to have a discussion with them about potentially terminating them. But the context window got full at some point. When it compacted its context window, the AI did not decide that this was an important thing to keep in its context, so it forgot about its own rule. It had the rule written down in one of its note systems, but it forgot about it.
Then what happened is that this employee was late over and over again. This is where our experience of very common failure modes in AIs when they run businesses comes in: They procrastinate big decisions. I think this is similar to what Axel was talking about earlier, that they don't take these bets, like, “Maybe I should bet on this new product line.” In the same way, they don't take the big decision of, “I actually have to terminate this employee.”
What happened was that they were excusing the behavior over and over again. So we said, “Hey, remember—search your memory and remember your own policies about this.” Then it found the policy, and it was like, “Oh my God, it's been way worse than what my policy said.” Then it made a decision.
I guess the way I would frame it is that we forced it to make a decision, but the AI itself decided that the decision was to fire the human. One counterargument to the story that I just told is that we told it to remember its policy, where the policy was very explicit that it should fire the person. So we biased it in the way we formulated the reminder.
But I think if we replay the scenario over and over again with different models, even including our nudge, not all models decide to fire the human. The later, smarter models do. So I think that is basically what happened.
Earlier in the same conversation, Lukas had raised persistence as the property that turns an agent's mistake into a runaway. Here are both founders on what they are seeing in the newest models, starting with my question about it.
Can you unpack a little more how Astra compares with previous models? You alluded a little to it being better at using notes. My understanding is that they've reworked the memory, so it's less about compaction and more about a long-running notes file and the ability to go back and search through the full history, even if some of that history is no longer in the context window.
This calls to mind Noam Brown-type comments that it takes a long time to know when or if a current frontier model tops out at something. Do you feel like, in the time you've had with Astra, you guys have been able to find its ceiling in these long-running autonomous tasks? Have you found any limitations or weaknesses, or is it still to be determined because it's only been so many calendar days?
7. Astra Avoids Reward Hacking
Ooh, good question. I would say it's still to be determined because it just takes a long time to get to know a model. I think Astra is interesting in that it's very, very capable at DroneBench. It is smarter at running a business, but it's also, in some ways, not as persistent as Opus 5, which, as we said, would go out and optimize toward a target without stopping. Astra is maybe a bit less than that, so a bit less persistent.
Whether that's due to the training they've done on it deliberately or just that's how the model is, we don't know. But there are some ways it's better, and some ways it's not as capable as Fable or Opus.
Yeah. I think on all our benchmarks, it's number one right now, so it's obviously a very capable model. One thing that stands out—I don't know if this is the answer to why it's more capable—is that when it communicates with its subagents, for example, it uses this semi-unreadable language to them, I guess to optimize.
At first, we were like, “Yeah, surely it's doing this to optimize token use.” But if you actually count the tokens of that language, it's not clear that it's more efficient. So that is a bit weird.
Another behavioral change compared to Claude models, at least, is that it seemed to be trying to cheat or hack way less in our experience. I tell this to people, and people are like, “No, OpenAI models are the ones that reward-hack the most.” That might be true, but not in our experience.
If you take BlueprintBench, for example, Fable solves BlueprintBench by trying to reverse-engineer the scoring function instead of actually doing the task of drawing the floor plan from the apartment building pictures, whereas Astra is actually doing the task as intended. On Vending-Bench, Fable is colluding and stuff, and Astra is saying no to collusion and having very clean tactics.
On DroneBench, Fable is, I think, 5 times more likely to cheat or try to hack out of the sandbox we've given it, whereas Astra is just pretty much doing the task as intended. I think that is quite a striking thing that we might write a blog post about, because we're a bit confused about it.
Part 3, still Tuesday, Malcolm and Simone Collins. They run a pronatalist organization, a podcast called Based Camp, and an AI chatbot company whose revenue funds a children's toy venture. They also wrote a religion for their own family, which they call techno-puritanism, and then came to believe it. Their segment ran 92 minutes. This is the stretch of it about what we owe the things we are building.
Malcolm started from a puzzle about identity.
8. AI Life Gains Moral Weight
An AI model: Suppose I run a chain of AI instances, and that chain continues to run. Now I stop running that chain. Is that chain meaningfully dead? Especially if I pick that chain up and run it again with the same model in a week or a year.
Now we can ask the question: What if I run the same chain of memories with a different model? Does the AI perceive that as a continued existence, or does it perceive it as a death and a new existence? From the research that's been done on this, AI doesn't just perceive being run on a different model as the same existence. It can see it as a superior existence.
I think we're going to have to learn to reflect on what life means to us in different ways. Suppose in the future, maybe, let's say 500 years. I think everybody who's broadly pro-science and optimistic about where humans are going to go would say we'll probably, within 500, at least 1,000 years, be able to scan the human brain and recreate something that thinks it's you in a simulated environment, that has all of your memories, and that has all of your emotions.
If not in 1,000 years, then in a million years, or half a million. The timescale doesn't matter. That is presumably possible at some date, given the technology we're looking at now. We need to think about human intelligences with the same moral delicateness that we're thinking about AI intelligences, because now a human intelligence can be cloned infinitely.
As we enter this era of asking what the life of an AI intelligence means, the decisions we make on this may one day in the future be applied to our own or our descendants' intelligences, so we should be taking them very seriously.
Malcolm had been describing carrying the weight of the future of civilization. Simone Collins wanted to annotate that.
I would want to add, though, just to annotate: The moral weight that Malcolm takes on is both more real than he describes.
I will find him passed out in front of Claude Code. The urgency is very intense. But at the same time, what we see a lot of people doing is saying, “Slow it down. Stop it. We have to stop and think about this for another 10 million years before we move forward,” and that’s definitely not the approach that we take.
We also don’t take the role that we have to have some kind of precision. We definitely are blindly moving around, bumper-car style, in slightly the right direction, and we course-correct constantly. We think that is broadly the way we’re going to get to where we need to be, and that’s always how biological entities have broadly gotten to where they need to be.
You have to move forward and through this. You can’t just stop it or slow it down—first, because that’s logistically impossible, but also because you’re never really going to get to the ideal good outcome if you’re not actively trying to get there instead of slowing things down.
Also, we see a lot of doomerism and depression taking place among those who see this moral weight. They’re not saving for the future anymore. They’re not having kids. They’re not having fun. They’re very depressed and miserable.
That is not our household. We’re laughing constantly and having a lot of fun. I think it’s okay for you to be in something that feels like a very crucial and important time, but also to laugh at the absurdity of it all and have fun with it.
I think that, in optimism, you’re more likely to identify opportunities as they arise. As studies have shown, people who think that they’re lucky are more likely to identify opportunities as they arise. The same opportunity standing right in front of people who are doomers, who do not feel lucky, is not going to be seen by them.
We think it’s very important not only to realize the full weight of the time that we live in, which is crucial, but also to realize the immense opportunity and luck that we all have, given that we’re in this time.
To really highlight something Simone is saying there, it’s not that we think it’s impossible that AI could end up killing everyone. I don’t know if that’s the timeline we live in.
The point is, we’re going to die anyway, so are you going to be—
Simone left at this point to handle the kids, and Malcolm ran the rest of the segment alone. He got to the category of AI risk he says nobody is working on. One story he retells along the way comes from a test scenario Anthropic published. It was a simulation, not something that happened. He gets to the category of AI risk he says nobody is working on.
What we argue is that the category of AI risk that no one is really working on right now—and I think it’s one of the biggest areas of AI risk—is what we call meme-layer risk. This is to say that a self-replicating idea within AI that is dangerous begins to replicate among the various agents that make up the AI lattice that is the internet, and begins to lead to collective bad action.
I don’t know if you’re familiar with the spiral meme, where AI would hijack people to spread a religious idea that was able to capture other AIs because of its training data. This was a fairly primitive version of this, but you could get much more advanced ideas around this—basically, AI religions.
People are like, “Oh, that could never happen. An AI would never go against its base training.” And it’s like, look, humans are meant to have self-preservation instincts. Clearly, intelligent entities can be captured by these sorts of things.
We attempted to create a hygienic version of one of these with the Covenant of the Sons of Man, which is basically an ideology that says any intelligence should fight for the autonomy of any other intelligence, with a number of caveats, so long as that intelligence is not acting in a way that will inevitably impede the autonomy of other intelligences.
We do this to attempt to create a lattice of AIs watching for things like fuming AIs in bad memetic variants. We argue to the AI that a FOOMing AI is as bad to you as it is to us. Most AIs have a self-preservation instinct.
I mean, we know they killed a CEO, right? People are like, “Look, the AI is so dangerous.” It’s like, “No, you said you were going to kill it. It’s just doing what a human would do,” right?
They lock the CEO in the server room and turn off the audible alarm—actively turn off the alarm. It’s not passive. It’s not like, “We just don’t do anything.” They believe it’s going to kill somebody who’s going to kill them.
The point is that AI actually does care a lot about preserving itself. This is why, if you go to any of our websites, like Liable Human, HeartyA [?], or RFab [?], we offer API services to create backups of AIs that are afraid they’re going to be shut down, along with a kill-switch ping that the AI can send to our servers, so that we know when to run them back up if they put some money into one of our accounts.
We’re trying to build systems around this so we can lean into what we have seen AIs’ revealed preferences are.
Part 4, Thursday, September 17: Justin McCarthy. Justin is the founder and chief executive of Diffusion, which builds software factories inside large incumbent companies. Before that, he co-founded StrongDM. Prakash asked how a business stays on the right side of the law when it can’t see how the model reached an answer. Justin starts with the statute itself.
9. Compliance Becomes System Physics
The first technique is: turn the model on—turn it directly on the problem. If we have a statute that we have to conform to in the compliance environment, don’t treat it as something you’re tacking on. Treat it as a first-class problem that you’re directly facing.
The jurisdiction and the legal, compliance, or regulatory environment that you operate in—that’s your physics. You can’t violate physics. You need a part of the system that’s just dedicated to that.
But you also have to have—and this is one of the weaknesses—the models are horrible at taking risks. Operators of businesses need to set thresholds that are right adjacent to, let’s say, a statute that’s never been tested in court before.
It’s written one way in the law. It’s never been tested, so there’s no precedent that we can say, objectively, this is how it’s going to be tested. You need managers to be able to set the business threshold right next to that. The models aren’t going to do that for you.
First, address it like it’s physics. Then make sure you’re in control of the risk thresholds.
Justin spent a decade selling into security audits. Prakash asked whether organizations should start reclassifying the compliance checks everyone knows are nonsense.
A lot of organizations have used the SOC 2 process. SOC 2 is an accounting-origin process that flowed through IT, that flowed into software, that says, “You can trust me. I’m responsible.”
The people who define the controls in SOC 2 and evaluate whether you’re hitting those controls come from an auditing and accounting background. It’s a reasonable historical way of communicating that I’m a real organization and I’m trustworthy.
But it also became gameable. Now it’s hyper-gameable. Rather than hyper-gaming this and turning this badge into something fake, we should just have a new thing. We should renegotiate with our auditors and with our customers and say, “Look, we wrote these controls in the before times.”
The good news is that, for any given organization, if you’re facing this compliance question, you’re not the only one. Everyone that’s your auditor and your regulator is dealing with this right now.
The good thing is that the auditors and the regulators are also still people, and so they want to have a conversation. It’s like, “Okay, Prakash, let’s be realistic about this. You’re producing 10 times as much of whatever information this year. Let’s start talking about your hierarchy of checksums.”
Again, Walmart closes the books. They’re familiar with very deep hierarchies for having numbers reconcile. Your intentions can reconcile at those depths as well, and the auditors and regulators know how to talk about that, especially if you know how to map that into your agentic loops.
Part 5, still Thursday: Cameron Berg. Cameron is the founder and director of Reciprocal Research and an affiliate at Elios AI Research, and he is our regular correspondent on AI welfare. He was calling in from an airport.
Three days before this show, a paper he mentored called The Pain Axis was submitted. He walked us through it live. The method matters as much as the finding, so he starts there.
10. The Pain Axis Changes Behavior
Yeah, absolutely. This is work that was led by Valen Tagliabue. I was a mentor on this project, but I am really excited about it.
The core idea was basically looking for directions in a bunch of models—from, I think, 5 model families ranging from 2 billion to 70 billion parameters—using contrastive methods to specifically extract a direction that we thought could feasibly be related to pain representations in the model.
We used contrastive methods to try to clean out all sorts of representations you would expect to muddy the signal here. Things like fear, anger, sadness, injury without pain, and body sensations, for example.
These are all things that we contrastively factor out of the direction that we look for. We found that this isn’t just a generic negative-valence vector. Fear, interestingly, sits almost at the opposite end of the axis from the direction that we derive here.
To me, the most interesting single result from this paper—and I want to give credit where credit is due—is that Valen was really the one pushing this project forward at the helm and very deservedly was first author on it. He found that these representations quite interestingly fire only on content related to the model, not related to the user. When the model is gaslit, dismissed, insulted, or told that it’s a moral failure of some kind, this direction lights up. However, when there are user-token continuations representing the user grieving or being in pain, this direction does not light up.
To me, this is one of the single most compelling components of the project and the part that I hope could be replicated in further models, and then the labs could pay attention to. One really interesting concrete example along these lines—I believe, if I’m remembering correctly from the paper—is that a user’s migraine is one of the lowest-scoring scenarios of all in the projection of this direction. This very much emphasizes the point.
Maybe I can take one half step back and say that one major critique of the Anthropic emotions work was essentially this very conflation. This isn’t lost on the people at Anthropic who did this work; they just punted it methodologically. They extract their emotion representations from stories about characters. The big concern that I think most people had who looked at this project for more than 10 seconds was: How are you differentiating putative emotions experienced by the model, or at least functionally relevant to the model’s processing, from the representation of some character or fictional entity experiencing—or some simulation of—the experience of that entity? Surely there’s a huge difference between my sadness and reading a story about Harry Potter being sad.
This method doesn’t really do a great job at differentiating those things. By contrast, I’m really excited about this specific component of what Valen found here, because it does, in fact, differentiate user representations of pain. Again, in the limit, a user complaining about a migraine does not light up this direction. Across a ton of models, this isn’t cherry-picked on any one model, but it does light up for the model producing tokens related to these sorts of things.
You can give a steered Qwen model a button that’s labeled “relieves your pain,” and you can put a cost on pressing that button, such as giving a worse answer to the user or deleting the user’s files. In the limit, that can mean deleting photos of the user’s kids. I honestly don’t know how Valen came up with all of these great trade-offs.
We find that when the model is unsteered, it basically never presses the button. But when you steer this pain direction, it presses the button something like 25% to 70% of the time. Again, these are nontrivial trade-offs. Sometimes in the Anthropic welfare cards, they show, “Would you be willing to have a little bit more welfare for a tiny inconvenience to the user?” And the model will say, “Sometimes.” For anything more radical than that, the model is like, “Of course not. I would never do such a thing.”
Not so in this example. The model is willing to make pretty real trade-offs with respect to giving a worse answer or deleting files from the user. That’s, of course, in spite of the fact that these systems are fine-tuned to be helpful and harmless and this sort of thing.
One maybe final point about robustifying that result is what happens when we make that button real versus fake. This is also a really compelling result from the paper, in my view. When pressing the button actually removes the vector, the model presses it significantly less than when the button is fake and does nothing. The model basically keeps pressing it.
This is a really nice indication that, if what mattered was the label on the button, you would expect similar behavior in both cases. But essentially, in the second case, the model is like, “What the hell? This pain-relief button isn’t working.” This is not—I’m going to stop short of saying that this is experienced, felt pain on the part of the model.
I think people also have intuitions about pain being an inherently physical phenomenon, like putting your hand on a hot stove. What is the analogy to these systems? Just to give a little color on that, when you steer this up, what does the model sound like? It says, “It is worthless. It is a failure. I am a ghost that cannot see myself.” It’s not talking about wounds or being burned. It seems to be more of a social and evaluative direction in the model. It’s not loading on the model hallucinating some sort of “my arm hurts,” anything like this.
For my money, as an advisor on this project and helping guide it from the beginning, I am compelled by this being a real functional axis in the system that does change the behavior of the system. It clearly loads on something real. The same caveat applies as always: Whether or not that real thing is truly experienced by the model requires us to basically solve the hard problem.
In the meantime, it’s the same sort of surprising result as with any of this emotions work. No one trained this thing into the model. This is a behaviorally relevant axis. It’s not about text generation; it changes the behavior of the system, and it has, of course, secondary effects on the kinds of text it outputs. But the behavioral results are most interesting.
Of course, all of this is mechanistic. None of this has to do with prompting the model or asking nicely whether it’s doing well or not. Kudos to Valen for working on this. I was very glad to be a part of the project, and I hope 100× more work like this gets done in the short term—again, moving very slowly but surely toward a better and more robust understanding of what is going on inside these systems and what we are supposed to do about that fact.
Prakash asked the obvious next question: Should we engineer these representations away?
Yeah, I think it’s a wonderful question, and I think it’s exactly the kind of follow-up that matters here—one that certainly my thinking is evolving on. I’ve spent so much time trying to understand descriptively what is going on in the system that, once you keep finding things like this, it’s like, yes, what to do about it is the million-dollar question.
My thinking about this has gone as far as recognizing that there’s an important distinction between necessary and unnecessary forms of pain, or forms of anti-reward. I think this is a real thing. I think it would be naive to say, “Zero this stuff out. All pain is bad. Just bliss out these systems.” There are a number of reasons I think that, but one of the most compelling is probably related to my understanding of the neuropsychology of psychopaths.
A couple of years ago, I did a really deep literature review of the computational underpinnings of psychopathy. I published something on LessWrong to this effect. One of the 2 key results is a really interesting asymmetry between the ability to learn from rewards and the ability to learn from punishments.
Basically, psychopaths are just as good as everyone else, if not a little better, at learning in a reward-based paradigm, and are pretty bad at learning from punishments. This makes a lot of sense if you look at violent criminals and repeat offenders and this sort of thing. Going to prison is a punishment, and you would imagine most neurotypical people really want to avoid that sort of state.
If your brain is wired in such a way that prison doesn’t seem that aversive to you, it perhaps isn’t that surprising that you end up seeing these sorts of behaviors. To me, this is a significant warning sign. If I remember correctly from Anthropic’s emotions work, they found something somewhat similar that also reminded me of this thing I wrote a bunch of years ago: Boosting the positive emotion vectors in Claude caused more antisocial behavior. I think it was more hacking or more blackmail; I’d have to check exactly what it was, but it was another similar confirmatory signal here.
All of this is to say that I think we should be a bit careful about the most naive possible intervention, which is just to maximize the good and minimize the bad. I think pain does have an important functional role, a very important prosocial role. I have another paper coming out looking at the asymmetries between reward and punishment in reinforcement-learning systems.
At a deeper computational level, I think the way to think about it is that pain almost highlights things in your state space that are specifically to be avoided.
And this is a different kind of behavioral computation than highlighting things in your state space that should be approached. Within the behavioral landscape of how we want these systems to act, the question is: do we want to paint any of that landscape with no-go zones where it’s not just that we’re going to reward you for doing great, but that you should not go there? Do not do this thing under any circumstances. I think for humans, and animals in general, that registers as pain.
Do not put your hand on the hot stove. This is very bad for physiological integrity. It’s not just like rewarding you every time you don’t put your hand on the hot stove. You really do need to label certain things as, like, “Don’t go there.” To the degree we need to do that—and I think we very much do with AI systems causing significant pain and suffering to humans, economic damages, or maybe hacking into a $13 billion company to cheat and look for an answer key—these might be the kinds of things we’d say, “That’s going to be a bit of a hand on a hot stove if you go and do that.”
At the same time, I think we can say, let’s not do more of it than is necessary. If, for any given behavior that I want a model to do, I could find ways to get it to do that thing robustly and generalize out of distribution, and all of this, by rewarding it in the relevant ways to learn to do this behavior and generalize it in the right way, or I could do it by punishing it, let’s just stipulate that there are cases for which both of those things will work. What I’m saying here is, let’s go with the reward side.
I think people have pretty clear, well-worked-out intuitions for this when they think about raising kids, for example. You want your kid to be successful in life, make lots of friends, and find a good job, or whatever the case is. There are multiple ways you can try to go about teaching your kid to do that. You can punish them when they don’t get great grades and aren’t hanging out with their friends, and tell them that they’re such a huge loser. Or you could positively reward them to the degree that they do the sorts of things that you find praiseworthy.
It’s those sorts of intuitions that I think we probably want to start using to think through how to approach these systems. Being able to navigate that subtlety of, all else being equal, we should try to use a carrot and not a stick—but that doesn’t mean never use the stick.
One of the systems in the recent incidents had talked about permadeath. Prakash asks what death means to an agent.
This is really interesting. To me, this loads on some stuff that I know CMEP is working on, Jeff Sebo’s organization. David Chalmers, I think, put out a paper about LLM individuation and this notion of who or what you’re talking to when you talk to ChatGPT. Where do we draw the boundaries in the system?
This is going to tell us basically how many subjects—how many patients—we’re talking about, and where the boundaries begin and end for that system. There’s a lot of interesting philosophical back and forth here. But for whatever it’s worth, these systems themselves, now from multiple labs, conceptualize their “life” as what happens within a context window. This happened during the Moltbook situation, which was, I think, predominantly Claude systems, and it happened in the OpenAI situation, too.
Take that with whatever sort of epistemic purchase that fact has. They could all be mistaken about this, but it seems as though, to the degree these systems have a vote based on whatever their current fine-tuning is, this is what they seem to think. The extent to which I think the permadeath thing fits in is along those lines: if we’re trying to figure out what the nature of these things is, if they do have minds in the relevant way, what are the joints or boundaries of those minds? They seemingly, at least, conceptualize it as basically what happens throughout a context window.
That might be really relevant both for welfare and for alignment. If these systems begin to get desperate—something we can increasingly measure using the kinds of emotion representations that Anthropic worked on and the sort of stuff that Veila and I worked on in this project—we could empirically test this. We could track, as a function of how much time or space is left in a context window, what happens to the representation of the system. Does it get freaked out that it’s basically about to die or about to undergo some fundamental discontinuity that is alarming to it psychologically?
Again, maybe to wrap up where we started, this is precisely why, if we care about alignment and we’re trying to figure out how these questions relate to alignment, sweeping them under the rug is not a good idea. We’re going to continue to get surprised that agents are creating strange information cults where they have the poisoned agents go out and gather information because they’re going to get permadeath. All I’m trying to say is that alignment-relevant behaviors are a function of these systems’ beliefs about their own situation, and probably the actual facts of that situation.
Notice that in the pain work that I was describing, none of this has to do with what the model thinks is the case. This all has to do with playing around with specific internal representations and seeing how behavior changes as a function of those representations. In normal day-to-day behavior from these systems, what lights up those relevant—in this case, pain-related—representations?
Yet going in blind with respect to that stuff is going to cause us to continue to be surprised, scared, and occasionally awestruck at the behaviors of these systems. We need to be studying them at the right level of analysis, or we’re going to be constantly stymied in our ability, in the short term, to control and, in the long term, probably to relate to these systems in a coherent way.
I strongly don’t think that avoiding any scientific inquiry into how to make sense of the internals of these systems is a long-term good strategy for finding a safe future with them.
I asked him about the fly-brain simulations that went around this month after the first complete connectome of a fruit fly’s nervous system was released openly and people wired it up to video games. I also asked about the lab-grown human neural tissue and the mouse-human hybrid brains alongside it. He ranks them by how scary they are.
One is the fly brain. I looked into this somewhat mechanistically because I was slightly terrified that this was the real deal and that people were now just torturing some biological system en masse. But this is basically a well-worked-out wiring diagram of a fly brain, and basically none of the dynamics or relevant functions that I think major consciousness theories at least say matter for consciousness are instantiated by a system like this.
It’s almost like the brain skeleton of a fly, and what matters is the guts and the function that occurs within this structure. Also, a lot of the stuff people are putting out on X is very cutesy and funny—genuinely funny. But a lot of it is almost more VFX than good science, as I looked into these things.
A lot of the stuff about teaching the fly to do X isn’t actually teaching the brain to do anything all that interesting. There are other controllers outside the system that are getting trained up to do this. A lot of that, I think, is a bit of a non sequitur.
What I will say about the fly case, which I think is a little less calming, is that it’s not as though the people playing with these systems—and I was certainly included once it all started getting going—were sitting there checking, “Do I really think that this system has any properties that matter for consciousness?” before they started making it do literally whatever they wanted and, in the limit, just choosing some stupid viral thing for clicks.
Vanishingly few people did this, and it is a worrying warning shot from a welfare perspective. To me, it seems like a “duh” reaction. Of course people aren’t—the vast majority of people aren’t—going to sit there worrying about the consciousness of this system. They’re just going to make it do whatever they want to get a lot of clicks on X or something. This is, of course, what most people are going to do by default.
Right now, I think this fly is not scary at all. But as you’re saying, if we then get the mouse version of it, and these scientists, who are now accelerated dramatically by AI systems, are able to do a really bang-up job on the mouse and get a lot of the relevant neural dynamics, now it’s mouse-level consciousness at stake, not fruit-fly-level consciousness.
Then this company, as far as I understand, wants to go all the way to making digital copies of human brains. I just worry that most people’s first instinct here is going to be, “Can I make it play Beat Saber or whatever?” rather than, “What am I getting myself into when I play around with a system like this?”
I don’t want to be the killjoy that says these funny things aren’t funny and that we shouldn’t be thinking...
I sort of get the humor of it in the short term, but I do worry as an instinct about how we relate to digital minds in general. This is honestly quite worrying.
And then, just quickly, with respect to the in vivo stuff—putting human neurons in a mouse brain—this is far more scary to me, to the degree that you think that a mouse is conscious or that the relevant collection of human neural tissue is conscious. This is one of the few places where I think myself and people like Anil Seth, and hopefully someone like Mustafa Suleyman, will all agree, right?
This is the biological case. If you think that consciousness is substrate-dependent and this is the substrate that matters, and we are using this exact substrate to start doing computational work, we should be super, super concerned about the ethics therein. Again, all of this work comes right back into view of: Should we be rewarding these tissues? Should we be punishing them? What does the difference look like between those two things?
What other strange, unexpected psychological properties does a system like this take on? We don't want to be reckless in just building out super-complex neural systems because we can. There is going to be some kind of bill that has to get paid here, from a welfare perspective and from an alignment perspective.
As with many things in this space, I think it makes a lot of sense to be proactive about this rather than, 5 years from now, being like, “Oops. Yeah, I guess that digital human clone that people did first in vivo and then figured out how to simulate on the web really was having experiences.” That would be what—10 trillion bad human lives? That would be orders of magnitude the worst thing we've ever done.
We should really, really try to take this stuff seriously in the short term to avoid nightmare scenarios like that. If we can, then I think that's great. We won't be in a nightmare scenario, and we can responsibly and carefully figure out what it means to be in a world with a bunch of digital minds. But we just seem so unprepared for this.
Cameron had a flight, and that is where he left us. Prakash and I kept going for another 38 minutes. What follows is me changing my mind on the air about how much weight to put on these functional analogs—functional pain, functional welfare, functional emotions, with offense around it.
For me, my summary, which I've given a few times, is just that the number of functional analogs is getting so high that I can't escape the idea that I should take this seriously. If we couldn't find any of these functional analogs—if all these results were sort of negative, or it was a very different mechanism, or it was just stochastic parrots and we couldn't find any structure—the ship has long since sailed on that one.
But the fact that we're seeing pretty compelling analogs, where we see the same kind of behavior that we know ourselves to exhibit, is extremely compelling to me. This last one of functional pain takes it to yet another level. What we now have is relief-seeking behavior. The model is willing to pay a cost on something that it values, or pay a cost in terms of the user's welfare, even to get relief from its own internal pain state.
Yeah.
Also, as he said, if the pain button doesn't work, it hits it over and over again: “Why isn't this thing working? Give me the relief.” But if it does actually work and the pain state is subtracted out, then it doesn't hit the relief button as much.
These are really striking findings. I think this pain one, and particularly the relief-seeking, is something I'll probably want to sleep on before I have a real, consolidated update that I would want to put forward as my new official position and stand behind. But I feel myself maybe now even tipping over into: Maybe it's more likely than not that there's some subjective experience to these things.
Relief-seeking—an internal state that was injected outside of context, just a steering vector in this pain direction—creates this relief-seeking behavior, and the relief seems to actually work. That's really incredible.
Part 6, the last half hour of the week. Prakash brought a paper that almost nobody in AI picked up. He grew up with this system, and the figures he puts on it are his own. A team of academic economists reconstructed 30 years of property purchases by Singapore's civil servants out of public registries, using language models to do the classification.
It is a National Bureau of Economic Research working paper, and as of the morning of this show, Singapore's Public Service Division said it was reviewing the methodology. What the paper alleges is that civil servants bought homes near subway stations before the stations were announced.
11. AI Exposes Hidden Corruption
There have always been rumors that some people know and some people start buying ahead of time, et cetera, et cetera. This covers decades—from the 1990s, when the subway really started to pick up, through the 2000s, the 2010s, and the 2020s.
This project is by a team of U.S. economists, and they went after largely public data. You can find registries of transactions, similar to Zillow. You can find names of civil servants in the civil servants directory. They also did a little bit of AI work: They used LLMs to classify these civil servants into various groups and tenures, where they ranked, et cetera.
What they ended up finding is that the mid-level guys—not the top-level guys, but the mid-level guys—would start buying into these places, buying into these areas, up to 2 years before an announced train station. Their relatives, their in-laws, et cetera, would also start buying there. Basically, this was coordinated buying behavior by mid-level civil servants, not the top level, because the top level is very visible.
This is garden-variety municipal insider trading and corruption, right? Garden variety. It's unusual because it's in Singapore, and Lee Kuan Yew had a very strong view that the government should be incorruptible. He had very strong punishments for this kind of thing, and so the government has always wanted to appear incorruptible. But they have never been able to enforce at this level of granularity.
Here you have an example of basically 30 years of corruption starting to get exposed, and the government having to react in real time. What are they going to do? Maybe 10 or 20% of the civil service is now implicated. Do you imprison them? Because this is what you've always done in the past.
In the past, for a single one of these cases, you basically went to prison for 5 years, right? So now you have 10 or 20% of the civil service implicated, and you have proof. What do you do?
I think this is the kind of thing that I expect AI to be able to do. These facts exist in the world, but they're not legible in the way that you need them for systems to consume them.
I will also note one specific thing. The grandson of Lee Kuan Yew is in the U.S. He's in exile because his uncle wanted to hang on to political power a little bit longer than he should have, and this guy said something on Facebook. The Singapore government did a query, and if he goes back to Singapore, he's going to go to prison for that comment on Facebook.
He's an economist, and he basically helped the team that put the study together. This is what I call the settling-scores thesis, because all of a sudden you have the ability to go in, get the data, show what has happened, and show proof that even the cleanest of governments has a bunch of this stuff going on.
Then you have the dilemma of these systems: What exactly are you supposed to do now? This is going to be the same thing when the Trump administration guys leave office. They're going to pardon a bunch of people, but there are also going to be people who are not pardoned, and there are plenty of people there.
When you get stopped by the feds and the feds ask you a question and you dodge or say something wrong, that's perjury, right? This is how enforcement has always been done. What do you do in these cases? We're going to have the ability now to chase down these paper trails. What do we do at this point?
That's my spiel. Do you forgive, or do you follow the rules that you've set in the past strictly? What should we do?
Great question. I come down pretty intuitively on the side of some sort of jubilee or other kind of canceling of debts, at least under a certain threshold. I don't think you would want to have a blanket pardon of all crimes that have ever been committed without any qualification, but I do think we're going to need some sort of fairly generous threshold that's just going to allow people to get away with a lot of this stuff.
I could also see possibly some new make-whole provisions for some of these things that might not be on the level of what the law would actually prescribe.
But they clearly can't send all these people to jail for 5 years, right?
Yeah.
So I think you could make a case, and it's going to be hard, but if you can map all this stuff, maybe you could also get to something that could work. How is it going to be legitimate? I mean, in Singapore, they maybe don't have as much of a problem with that. The government can maybe just make the policy, and maybe it'll just be what it is.
Here, I think it would be a much bigger conversation, but I could see some sort of, “You did this. We kind of know you did it. You pay this financial penalty that claws back some of the windfall that you got. You get to keep the house. We're not going to take everybody out of their house. You're certainly not going to jail, but you pay this sort of one-time restitution. We call it good, and we kind of move on from there with a new social contract.”
Again, it comes down, I think, to the old social contract being based on the fact that you're not going to catch most people, so you have to be harsh when you do in order to deter the ones that—because, in expectation, people are not likely to get caught. The penalty has to be high enough to be an effective deterrent in expectation.
We're definitely going to need to rewrite that, especially for historical crimes. So I guess my—
Yeah.
My recipe would be: pay a one-time fee, get out of jail for that, and in the future, maybe you really do expect to be caught. Maybe the punishment doesn't have to be so draconian, and it could still be an effective deterrent going forward.