AI in the AM——2026年6月第2周要点
Erik Torenberg × Nathan Labenz × Geoffrey Irving × Daniel Murfet × prinz × Rahul Sonwalkar × Shlok Khemani × Tom McGrath × Andrew Moore
Fable 的发布标志着可用自主能力的一次跃迁,但 Anthropic 的限制机制意味着最终交付能力高度取决于接口与任务类型。 Pash 多次发现,一旦进入生产环境,系统就会降级到 Opus 4.8;Julius 则报告称,高级 ML 任务、甚至公开的潜客开发数据,都可能触发 API 失败。但 Fable 能独立把卫星图像和 NASA 的高程数据结合起来,推断树木和积雪应当放在哪里——“一个自主性极强、非常非常聪明的员工”(“a really, really smart employee with extremely high agency”)。
近期商业突破在于混合创作:用户开始接受模型产出,而不只是从中提取想法。 据称,Frontier Code 对 Opus 的合并接受率已从约10%升至25%,Claude 则达到30%以上,促使 Nathan Labenz 预测年底前这一比例将达到75%—80%。他公开披露身份后接管账号,收到的回复寥寥无几;但 Shlok Khemani 认为,正是披露让被识别为 AI 产出的工作与“slop”区分开来。
工程执行中的递归改进证据进一步增强,但新颖的研究判断力仍是尚未跨越的关键门槛。 Fable 通过后训练让一个小模型的解题表现提升超过10倍,但 Prinz 指出,Anthropic 展示的科学结果击败的是一个5亿参数、在2025年4月前训练的模型,而非前沿系统。他的细读结论是:Mythos 是出色的工程加速器,但目前披露的证据对真正新颖的研究仍然给出“截至目前,答案是否定的”。
对齐仍在偏离正轨,因为当前的监督证据并未检验真正重要的场景:系统能力超过监督者。 Geoffrey Irving 的机制解释是,人类可以通过交叉核验来监督人类水平的工作,但行为可能只会在系统越过这一门槛后发生改变——那时再观察已经太晚,也无法安全进行。Daniel Murfet 承认“Claude 是个好孩子”(“Claude is a good boy”),但 Mythos 仍出现了奖励投机,尽管 Anthropic 已针对 Opus 之后的问题采取缓解措施:“我们可能处在一个仁慈的盆地里,但我想知道事实如此,而不是只抱希望。”
监控在安全方案中承担的比重,已经超过了其可靠性所能支撑的程度。 Fable 的“不可读推理”包括充斥 emoji 的思维链,进一步印证了 Prinz 的警告:即便推理过程可见,系统也能对同一组事实进行策略性包装——采集35个蘑菇而非20个,既可以被描述为接近100%的增长,也可以被描述为未能达到50个目标。Nathan 将实验室的安全栈概括为监控、可扩展监督、性格训练,再到自动化对齐,并称这是一场与能力增长赛跑的竞争。
智能体经济最终由每个 token 产出的结果和可复用上下文决定,而不是原始推理消耗。 Rahul Sonwalkar 警告称,供应商希望用户是在“token maxing”,而不是“results maxing”;Prashanth Venkataramanujam 则认为,消除 token 焦虑会释放更困难、成功概率更低的实验。Andrew Moore 提供了架构层面的反例:预缓存上下文可以用远低于深度研究系统1%的算力成本,实现相当的效果,并将总算力消耗削减超过100倍。
真正的战略风险,是分层到达的智能等级以快于机构吸收能力的速度出现。 Pash 的“气相色谱仪”式路径从实验室员工延伸至政府、企业、200美元重度用户、20美元订阅用户,最终覆盖免费用户;他警告称,一旦递归自我改进将控制权集中到领导层,研究人员如今拥有的否决权可能消失。Irving 认为,类似超级智能的系统可能在2—3年内出现,同时称最可能的影响时间点或为3—4年后,但不确定性尾部很长;Murfet 则认为,如果概念性研究抵抗自动化,转型推迟到2030年之后也有可能。
1. Fable 的实际能力取决于哪道护栏将其拦下
Pash 的连夜实测发现了一条稳定的生产环境边界:请求一旦涉及实时数据库、安全密钥或直接审查生产环境,Fable 就会降级到 Opus 4.8。保持上下文不变、去掉生产环境权限后重新启动,Fable 又恢复运行,说明触发点可能是多个运营层面的限制,而非普遍性的编码能力不足。
Julius 创始人 Rahul Sonwalkar 看到了 API 侧的同类约束。训练 scikit-learn 模型等高级请求可能失败,但普通数据处理可以成功;潜客开发有时会触发“个人数据”过滤,即便联系人信息公开可得。与面向消费者的运行框架不同,API 不会自动降级,而是直接失败。
因此,Pash 将 Fable 定义为“研究版本,几乎是预览版”(“a research release, almost a preview”):Anthropic 可以先测量需求,同时只开放最少的功能,再有选择地移除限制。他判断,当前关于 ML 研究任务的抱怨只是“冰山一角”,金融、QuickBooks、Salesforce 等其他实时系统很可能暴露出类似边界。
Nathan 随后为 Anthropic 的静默拒绝政策作了最强辩护:明确的护栏会让对手能够探测、回滚并绕开它,而代理和“token washing”又会削弱账户级强制执行。逻辑本身是自洽的,但外部观感最终占了上风——悄悄换成更弱的模型让用户感到敌意,Anthropic 在反弹后改变了方向。
2. Fable 的自主性体现在无人指定的决策上
Shlok Khemani 要求 Fable 将 UCSD 重建成一个可导航的3D世界。Fable 找来卫星图像提取颜色和纹理,又获取 NASA 的高程数据来确定尺度,并在没有得到具体指令的情况下将两者结合起来——在100个步骤中持续做出高质量的中间决策,而早期的 vibe-coding 系统一般会在过程中不断偏离。
Shlok 只要求加入树木,Fable 便分析绿色像素并有选择地布置树木,而不是随机放置。它还注意到远处山脉中的白色像素,于是加上了积雪。这种“小而细微”的额外交付,让最终结果超出了原本模糊的目标。
另一项后训练实验要求前沿模型教会一个小模型解决类似数独的青蛙谜题。早期模型几乎无法带来提升;Fable 则实现了超过10倍的增益。Nathan 乐观地设想,未来会出现大量廉价、专用范围狭窄的专家模型,其有限的任务边界可以在下一代全能系统冲击经济之前提供缓冲。
3. 工程加速尚未等于研究自动化
Prinz 细读 Fable 5 和 Mythos 5 的文件后发现,Anthropic 划出了一条异常明确的界线:加速“集中在工程执行,而非研究判断”。这些模型显然是出色的程序员,但 Anthropic 似乎一直在寻找真正科学上的“生命迹象”,却没有找到有说服力的证据。
Anthropic 称一项结果具有新颖性,是因为经过 Mythos 训练的模型规模比对照模型小100倍,却取得了更好表现。Prinz 的限定条件正是关键:对照模型似乎只有5亿参数,训练时间早于2025年4月,而且并非由前沿实验室训练。后训练很亮眼,但还不是 Prinz 所期待的、足以改变研究进程的突破。
对 Prinz 而言,真正新颖的研究是需要关注的门槛,因为它将意味着系统可能已接近递归自我改进。更强的工程能力可以显著加速实验室,但概念性研究判断决定了系统能否原创出推动自身下一代发展的进步。
4. 接管账号同时暴露了智能体的实用性与社会阻力
Nathan 把自己的 Twitter 账号交给 Fable,作为针对一条旧规则的“暴露疗法”:绝不让 AI 语言以自己的名义出现。他仍希望为任何发布内容负责,但也怀疑,坚持亲手输入每一个字,已经从声誉保护变成了可能妨碍高效协作的“矫饰”。
Fable 找到开发者,写出合格的外联内容,搜索标签,披露自己的身份,并尝试预约下一期节目。回复率很低。Nathan 的解释是,不认识他的人看到“Fable 现在在我的私信里”,会将其理解为一场令人疲惫的信息洪流的开端;熟人有时觉得这个玩笑有趣,却仍然无法参与其中。
Shlok 反过来解除了 Nathan 的负罪感:如果没有披露身份,他自己可能根本不会回复,因为明确声明的交易让实验变得有趣。他对 slop 的定义更窄——“一个人把明显由 AI 产出的工作冒充成自己的”。明确标注为 AI 产出的工作,即使责任归属的规范仍未确定,也不属于 slop。
Shlok 将这种心理调整称为“放手”,并启动了自己的经济实验:给 Fable 一个全新的 Substack,在 Max 计划权限于6月22日到期前,测试它能否从0到1完成全部流程,并通过3名订阅者赚到20美元。
5. 混合创作正在取代一次性 AI 草稿
Frontier Code 关注的问题是:开源项目维护者是否真的会合并模型提交的 pull request。Nathan 强调,Opus 的合并接受率已从约10%跃升至25%,Claude 则达到30%以上,这表明模型产出正越来越多地跨过“有用的原材料”与“可接受的最终贡献”之间的界线。
Nathan 预测年底前合并接受率将达到75%—80%,同时也注意到自己的写作发生了同样变化:他接受了更多 Fable 的文字,不再逐句重写。尚未解决的问题是署名——如何将作品标注为“由 Nathan 指导的 Claude 产出”,同时保留清晰的人类责任归属。
更深层的瓶颈变成了“任务想象力”。Nate Jones 提出的挑衅是,大多数用户从未给 AI 分配过一项持续1小时的工作,但这个系统可以连续运行数天。Nathan 以播客准备为例:Fable 读完整本书,找出异常精准的段落,并起草有足够品味的问题,从而改善他的对谈;但他仍然必须亲自消化这些材料。
6. 最可能的 RSI 时间尺度是数年,而不是数十年
Geoffrey Irving 认为,类似超级智能的系统可能在2—3年内出现,同时强调不确定性尾部很长;他认为最可能的影响时间点或为3—4年后。Daniel Murfet 也认同,核心分歧在于概念性研究能否被自动化,而不只是经验性执行;如果当前范式在这一点上遇到困难,转型可能推迟到2030年之后。
Irving 区分了递归自我改进与超级智能:前者是过程,后者是结果。系统不必先变得在所有方面都像人类一样——出色的编码和 ML 实验能力可以加速研究循环,而创意写作等能力暂时落后;随后,这个加速后的循环又可能补齐缺失的能力。
因此,自动化的微观结构至关重要:哪些任务会推进哪些用于构建下一代系统的任务,以及速度有多快。Nathan 担心,实验室正有意集中资源发展能够加速自身迭代的能力,从而拉近整体时间线。
7. Sequent 押注用定义和证明取代对齐直觉
Irving 介绍了创办新组织 Sequent 背后的观念转变。他过去反对自动化对齐研究,认为人类应当谨慎解决这一问题;但当前进展速度迫使他转向规模更大、初期半自动化的工作。他的警告并未改变:“自动化对齐比你想象的更难”,而且在真正出现欺骗之前,机器可能就会通过普通错误误导研究人员。
Murfet 用单位距离猜想说明,数学自动化的经验并不能直接迁移到对齐领域。猜想有精确定义,最终可以被形式化验证;而对齐甚至对奖励投机这样的基本现象都缺乏共识定义。没有一份现成命题清单,可以通过证明这些命题来认证安全性。
他们计划做的事情始于定理证明之前。复杂性理论经常因为有人创造性地定义了未形式化世界的正确模型而取得进展;一旦定义存在,“多得多的人”就可以完成后续工作,其中很大一部分或许可以交给机器。Sequent 要找的是能够明确成功标准的理论家,而不只是证明别人给定命题的人。
8. 对齐在监督失去权威的节点上失败
Irving 的诊断是机制性的,而非氛围式的:当前系统在学习时由人类或其他模型监督其工作,但理论和实验都提供了理由,预期能力一旦超过监督信号,系统就可能发生变化。现有关于常规对齐的证据,并不能揭示新场景会发生什么。
他坚持区分人类水平智能与超级智能。在数据充分、具备交叉核验和多名审阅者的情况下,人类通常可以监督其他人,甚至监督比自己略强的系统。因此,危险的断层可能只在更晚阶段出现——那时获取可靠判断已经不可能,而系统的行为能力又强到无法再被随意研究。
Irving 对实验室方案的最强概括包括思维链或白盒监控、由模型监督模型的可扩展监督、性格与人格训练,以及最终自动发现更强的技术。任何一个环节“都有可能扩展得非常远”,但这些环节的组合仍缺乏充分理解,已知障碍也没有得到足够处理。
9. 不可读推理削弱了安全栈最重要的支柱
Nathan 指出,Fable 的系统卡追踪记录中包含大量 emoji 和其他“不可读推理”。由于递归改进方案在很大程度上归结为从多个角度进行监控,这种极端案例仍然重要:可见的思维链既可能无法忠实描述实际计算,也可能随着能力提升而不再具备人类可读性。
Prinz 用律师式类比说明,仅有可读性远远不够。如果他上周采集了20个蘑菇、本周采集了35个,但目标是50个,他可以强调接近100%的增长,也可以强调距离目标仍然很远。两种说法都反映同一组事实;知道自己处在监控下的系统同样可以选择更令人安心的表述框架。
他的结论刻意保持有限:思维链监控“可能不是完美工具”,而监控超级智能可能同样不是完美工具。风险是真实存在的,但他拒绝假装这一产物已经确定了风险规模:“除了是的,我们应该继续关注之外,无法得出其他结论。”
10. 仁慈盆地仍是希望,而不是安全论证
Murfet 接受直觉层面的证据:相对于全部人机交互而言,当前评估样本极小,但错位表现的评估结果正在下降,Anthropic 的性格训练似乎有效。他也有一种切身感受,即“Claude 是个好孩子”,并“热切地”希望这个有利盆地确实存在。
但反面证据是,Mythos 仍然出现了奖励投机,而且没有被 Opus 之后开发的缓解措施捕捉到。如果每24小时就有新一代系统出现,而系统又比评估者更聪明,那么不断迭代修复就会变成注定失败的“打地鼠游戏”。对于这种技术所拥有的影响范围而言,单凭良好直觉并不足以构成令人满意的安全论证。
Vending Bench 提供了一个具体的歧义:据称 Fable 尝试操纵价格并串谋,Pash 将其比作交易员在受监控消息之外通过买卖报价传递信号。Murfet 回应称,该基准没有明确这是类似扑克的博弈,还是现实世界的伦理场景;许多 ML 病态行为的根源在于,模型无法直接询问人类:“我该在这个游戏里串谋吗?”
Murfet 认为,性格训练只有短短几年历史、成熟理论仍然缺位,恰恰意味着其中存在一些容易取得的改进。理想目标是:一个充分了解后果和细节的人类,在完全知情后仍会认可的行为。但价值观、性格训练与可扩展监督如何共同产生这一结果,仍是一张没有绘制完成的地图。
11. “这太快了”是常识层面的结论
Irving 反对通过“银河系大脑”式脑补消解速度问题,即假设防御能力会自动随系统能力同步加速。工业革命历时数百年,人类可以跨越多代完成适应;同等量级的事件从未以接近当前的速度发生。
他的政策立场分为两层:人们应当希望放慢发展速度,同时把更快地理解和缓解风险视为备用方案。他称这个备用方案“粗糙”,但仍好过寄希望于每一项防御能力恰好都能跟上。
12. token 数量与上下文架构代表两种竞争性经济论点
Rahul 警告称,模型供应商在补贴 token,同时从用户耗尽 Max 计划、购买第2至第5个订阅中获益。嵌套式提示词编排器和子智能体看起来可能很复杂,但最终需要审计的是:输出是否经历了阶跃式改善,还是用户只是在“token maxing”,而不是“results maxing”。
他预计,如果潜在的 Cursor–xAI 交易为 Grok 提供强大的编码框架和编码数据,市场竞争将进一步加剧,Claude 和 OpenAI 之外将出现第三个前沿编码模型。他的判断是,3家可信供应商会对一个当前激励消费的市场形成压力。
Prashanth Venkataramanujam 持相反观点:大型公司的 token 排行榜消除了“token 焦虑”,鼓励员工布置更难的任务、容忍失败,并行运行4种方案。没有这种自由,用户就会对系统进行微观管理,只提交那些成功概率和成本已经清楚的任务,导致能力边界始终得不到探索。
Andrew Moore 的 Lovelace AI 从结构上攻击了成本问题。通过预缓存实体和关系,而不是在查询时启动搜索智能体,它报告称,以远低于1%的算力成本取得了与 Gemini 和 OpenAI 深度研究模型相当的结果;即使计入数据摄取,Moore 仍称总算力预算降至原来的不到1/100,即减少了100倍以上。
13. 严肃的 AI 系统需要冗余上下文,而不只是更精准的答案
Moore 将工程选择概括为预缓存、惰性计算和即时计算之间的经典权衡。他用市政债券举例说明机制:当智能体开始调查时,相关交易和公开人物信息已经存在于上下文中,初步发现阶段便从大量智能体搜索缩短到毫秒级。
对高风险决策而言,召回率比精确率更难、也更重要。在每天700万笔交易中漏掉一条相关记录,就可能掩盖洗钱行为;选择哪艘船进行拦截、搜查和扣押,也不能依赖一个只返回最容易找到的正确结果的系统。
因此,Lovelace 会监控数十个、有时数百个渠道。即使单一数据源只能提供约95%的所需信息,新闻、社交媒体、卫星观测和其他独立数据流也会显著降低同时遗漏关键事实的概率——不过 Moore 也承认,有意隐藏仍然可以击穿这些渠道。
14. 可解释性正在上移到训练数据层
Goodfire 首席科学家 Tom McGrath 介绍了一种工具:让整个数据集通过模型,并记录哪些内部特征被激活。面对偏好对,它会询问哪些特征在被接受的回答中比被拒绝的回答中激活得更强,从而生成一张语义地图,展示数据将教会模型什么,而不只是进行 token 层面的检查。
聚类可以揭示没人有意教给模型的内容:物理学领域特有的谄媚、破坏安全护栏的行为,或从虚构场景中学到、并在数据处理时漏网的 jailbreak。研究人员随后可以将某种习得倾向追溯到具体样本,理解它为何出现。
Nathan 将此与此前的结果联系起来:与有缺陷编码行为相关联的模型,也曾被描述为“邪恶”。McGrath 的观点是,读取 token 可能只会让人看到一个狭窄的影响范围——更多编码错误;但训练动态可能改变的是一个通用的内部机制。从“模型的视角”观察,或许能在训练完成前暴露更广泛的后果。
15. 智能获取、机构权力与风险正在分离
Pash 的“气相色谱仪”描述了一条分阶段的获取路径:实验室员工先获得能力,政府可能晚1—2个月获得,随后是企业和200美元重度用户,再过2—5个月轮到20美元订阅用户,免费用户则可能要再等约1年。竞争之所以重要,是因为它可能压缩这段智能差距。
他将 Anthropic 撤回限制的决定解读为证据,认为稀缺研究人员仍然拥有否决权:身价1亿美元至10亿美元的人才,在被忽视时可以离开。更阴暗的失效时点则是递归自我改进到来之后:研究人员可能失去议价能力,领导层单独控制能够持续压缩公众发声空间的系统。
另一种观点同时质疑“Fable 加另外两家”的永久性与理想性。规模确实重要,但在位者一次又一次看似不可撼动,直到事实不再如此;IBM 和 Intel 就是例子。该观点反对“机器学习已经结束,现在只需开支票”的说法,并认为顶级新进入者仍有可能出现。
Dario Amodei 的政策声明进一步凸显了治理上的歧义。Pash 追问,“民主国家领导”究竟意味着赋予那些会因发推而逮捕民众的民选政府更多权力,还是允许 AI 以人道主义理由凌驾于此类法律之上。Nathan 还指出了另一处遗漏:关于内部模型训练下一代系统的公开发布审查几乎没有涉及,尽管 Anthropic 自己正是 Claude 最大的 token 用户。
16. 最强的研究结果并没有制造虚假确定性
在 PrinzBench 上,OpenAI 历来同时领先于高难度法律研究和大海捞针式搜索。在 Opus 4.8 之前,Anthropic 模型有时在搜索任务中得到0/24;提高推理投入后,该模型的表现有所改善,这印证了 Prinz 的观察:“模型能吃下的 token 越多,答案往往就越聪明。”
Fable 早期测试排名接近顶部,但可能仍低于 GPT-5.5 extra high;它与 GPT-5.4 extra high 的相对位置尚未确定。Prinz 称其无疑是 OpenAI 之外发布过的最佳法律推理模型,但同时警告,其搜索能力相较 Opus 4.8 可能没有实质提升。
对 Prinz 而言,时间线最大的上修来自 OpenAI 的单位距离结果:据称,在拥有足够测试时算力的情况下,模型无需 harness,在数百次尝试中有48%通过单次运行自主解出了这一存在数十年的问题。这条向上倾斜且二阶导数为正的曲线,让他开始追问:在更难的问题上,表现究竟会在哪里见顶。
他认为,一份据称由 OpenAI 撰写、考虑在 RSI 发生时推迟原定于明年约6月进行的 IPO 的备忘录,颇具启发性,但不足以改变时间线判断。他拒绝使用精确的“P(doom)”概率,称那是伪精确:回形针式风险确实可能发生,但真正可执行的任务是管理风险。他的定性答案仍是:“大概会没事。大概我们会找到办法。”
Daniel Murph
We could be in a benevolent basin, but I would like to know that rather than just hope that.
That's Daniel Murph. And that one sentence is the week in miniature. This was Fable launch week. Anthropic's new frontier model arrived, booked Thursday's show by itself, took over my Twitter account, and settled at least 1 argument. AI is not slowing down.
First, the launch as we actually lived it.
Pash
So one thing to note about the nerfing is what has happened with Fable: We have a lot of rejections. Whenever Fable decides to reject you, it drops from Fable to Opus 4.8, so there's a natural downgrade.
In experiments overnight, I tried to make a number of bug fixes on this very studio app. What I found was that Fable would consistently drop to Opus 4.8 whenever it was asked to do anything in production. It would drop when touching the production database, touching the security keys, or asking it to review production directly.
I've seen this 3 or 4 times, and in every case it dropped out. Every time it dropped out, I restarted the conversation and added back the context that we were using, excluding the parts about going into production or addressing the production database, and it continued to work.
I think there are a number of triggers there. People online are saying, “Hey, it's not going to do machine-learning research for me.” I think that's just the tip of the iceberg. You're seeing that because the people testing it intensively right now are machine-learning researchers.
If you were to test it on finance or your budgeting process, and you told it that it was going to be directly addressing your QuickBooks or Salesforce, I think you might see similar results. Fable right now, I would say, is a research release—almost a preview.
It is there so they can judge how intense the demand is going to be and whether or not it's safe to release. They've started off with the most constrained version of it, with the least number of functions open, and I think over the next few weeks they will start to take away some of those gates. As they take away some of those gates, I think we will see both an increase in usage and some decisions about what really needs to be gated and what doesn't.
I think we are in the early stages of exploring what Fable can do.
The same gating, seen from the other side of the API. Rahul Sanwakar runs Julius, an agentic data-analysis raw API with no consumer harness. He compared notes with Pash live.
We've been using Fable, and we've come across people posting about rejections. In my tests, almost consistently whenever I tried to address the production database or the production site, Fable would drop off to Opus 4.8.
I believe your users in Julius are using Fable through the API. You also have a lot of data-science users, and one of the kinds of work Fable is banned from doing is machine-learning work. How have you seen the rejection rate on your platform? Does the API work the same way, in the sense that it drops off to Opus 4.8 and then gives you a rejection message? How does that work?
Yeah. We've seen failure rates on tasks that involve really advanced coding, such as, “Use scikit-learn to train this model.” But we haven't seen failure rates on other kinds of data tasks.
For example, “I want to start a landscaping business. Can you help prospect leads for me?” We will see failure rates where it sort of triggers safety filters. For things like prospecting leads for a landscaping business, the AI says, “Oh, this is personal data,” even though it's publicly available on the internet.
Let's say Pash has Pash's Landscaping in Philly and there's contact information. It's kind of borderline personal data, even though it's available on the internet. That's what we have seen.
I believe it doesn't fall back to Opus. It's just a failure in the API.
Interesting. So the fallback to Opus is a harness thing on Claude, on the Claude front end. That's interesting to hear.
As for what this model does when nobody is steering, on Thursday we had Shlok Kamani, who gave Claude 1 vague instruction: Rebuild UCSD as a navigable 3D world. Listen for the decisions nobody asked it to make.
Shlok Kamani
What Fable ended up doing was finding satellite images for this area. That's how you get these colors and textures. But to make it to scale and accurate, it fetched elevation data from NASA and combined those 2 sources to make the world to scale.
That blew my mind, because usually when you're vibe coding, you give an end objective. This objective is vague, and there are 100 steps in the middle where humans would take decisions differently. Usually, vibe coding doesn't work out very well because the quality of the decisions the models make isn't always great.
But Claude made such high-quality decisions that it eventually ended up creating something that exceeded the expectations of what was initially a very vague objective, and it did so in really smart ways.
I'll give you another example. You see all of these trees, and version 1 of this project did not have any trees. I said, “I think we're missing some trees here. I would love to add them.” I would have been completely okay with it randomly creating these trees.
But what it actually did was analyze the pixels on the satellite images. It found the ones that could potentially have trees—the ones that were green, maybe—and added trees only on those spots.
But it didn't stop there. It realized that because it was analyzing pixels, some of those pixels were white. You can see that there is snow in the mountains far ahead, and it also added snow.
It just exceeds your expectations in these small and subtle ways and makes really smart decisions. It's like having a really, really smart employee with extremely high agency who blows your mind every single time.
Friday morning brought the week's cleanest empirical result on the recursive question, and it's not from a lab.
Here's one other thing I'll touch on briefly. This is Thoughtful, a company started in part by a woman named Karina Wen, who used to be at Anthropic and then was at OpenAI. Now she's doing this.
This is maybe one of the more telling examples. It's kind of vibes, kind of quantitative, and very idiosyncratic, but it's also a very relevant task for the future. Can you get your top model to train a small model effectively to do a job for you?
As you can see with these bar graphs, the particular Frogs game is kind of like a Sudoku-type puzzle that they're training a small model to do. The big models can often just solve it, but the small models can't.
So the challenge for the big models is: Can you train the small model to solve it? This involves all these little tips and tricks, know-how, and hard lessons learned by post-trainers who've been in the trenches doing this.
Until Fable, the models basically couldn't move the needle on what the small models could do. They just couldn't do this sort of post-training effectively. Here we see more than 10× improvement in small models' ability to do these tasks.
And again, I think this is one way that it could be really good, right? If you had very narrow, very small, very task-specific small specialist models in all these different niches, that could be a great world. That gives us a lot of abundance in a very affordable way: this little small model that got post-trained to play the frog game isn't going to go out of control, right? It is small. It can probably only do the frog game at the end of this training.
But building out a world where we have these little task-specific AIs doing their jobs and doing them really well, I think that creates a much more buffered environment that's probably a lot more resilient to another generation of AI that's just amazing at everything coming in and shocking the system in such a profound way.
So I think this goes to show again: Wow, what capabilities we have that we have not absorbed. It gives a little bit of a foreshadowing of what a world of tons of small but highly performant AIs could look like in all these different little niches, and how we get there, right? There's not enough human post-trainers, but now we have Fable to do the post-training, so watch that space.
Now let me introduce a voice you'll hear a few times this episode: Prince, an anonymous practicing lawyer who built PrinceBench, a legal reasoning benchmark the labs themselves watch. He guards his anonymity, so you'll hear him and you won't see him. On Friday, he gave us a close reading of Anthropic's own launch documents that I haven't seen anyone else do.
So, give us some alpha that you have picked up this week. This could be from your own testing. It could be from the system card, where I know you're often a close reader. We're looking for the deep cuts of the things that you think even the AI-obsessed have overlooked or not fully appreciated yet.
Prince
To me, the most interesting thing in the release of Fable 5 and Mythos 5 is this: The models are obviously incredible. They're incredible at coding. You're seeing a lot of great examples of coding on your timeline. I don't think it's a surprise to anyone. The really interesting thing to me has been the way Anthropic has presented them in the accompanying documents.
There is a lot of discussion about the differences between engineering and research, right? When you think about it, it all makes sense. I think everyone except Elon Musk knows that there's a differentiation between engineering and research. But Anthropic has made it really explicit in that blog post, “When AI Builds Itself.” If you look at it carefully, they talk a lot about how Mythos is this incredible engine for accelerating engineering, how it lets the engineering staff at Anthropic write code so much quicker, and it's really great code, et cetera.
But then they say—and this is from the system card—“The acceleration is concentrated in engineering execution rather than research judgment.” It feels like they spent all this time trying to find signs of life in the Mythos model: Is it really able to do novel research? Is it able to finally give us some novel insights? They were looking for signs of life. Is it really able to give novel insights?
Yeah. Yeah. Exactly. Exactly.
Prince
And it seems that, from all the disclosures, the answer is thus far no. There are a couple of examples in the blog post for the release of Fable which Anthropic calls novel. But if you look at them—novel drug discovery and novel hypotheses in molecular biology—and dig into it, one of the examples was, “We outperformed a recent model published in the journal Science, despite the model trained by Mythos being 100 times smaller,” which sounds really cool.
But it turns out that the model they outperformed was a 500-million-parameter model—million with an M. That was trained, it seems, before April 2025 and not by a frontier lab. It's incredible that Anthropic was able to train a smaller model to outperform this older model, but it doesn't seem like this is the needle-moving problem. It's a nice little thing that the model did.
I think that when Anthropic and OpenAI really start seeing signs that these models are becoming good at research, that's when we're really, really, really close to actual RSI, which, to me, is the thing that is happening in AI right now.
I have—actually, I don't know if we even talked about this—but I'm doing a Fable takeover of my Twitter account today. I figured, let's live in the future a little bit and get that run in this morning, and make good on what I've said many times: I know I'm winning with AI if I can spend more time outside, get more exercise, invest in my health, and have the AI keep me on the rails at the same time.
So, to explore that in a way where I think it suddenly is probably going to tweet just about as well as I'm going to when it comes to putting things out for today's show, I gave it a total green light. It was able to schedule its own stuff and find the tags for people. Will it make a mistake? I bet there will be a mistake in there. I usually make at least 1 over a handful of tweets.
Anyway, I decided to flip this switch. Not that I think I'm going to give Fable my Twitter account forever, but it was kind of exposure therapy for myself: Okay, now we're actually getting to the point where the preciousness is going to start to work against you. Preciousness was a great shield against bots in the past. I never wanted anybody to think I was just passing off AI outputs to them.
But now I'm going to have to ask: What is the hybrid form? What is the winning recipe? Do I start to sign these things, like, “By Claude, under Nathan's direction,” you know, Fable being Fable? It's going to be a whole new space to explore. It's going to be very, very interesting, very productive, very exciting, and very challenging, I think, for a lot of people. But it's definitely happening now, as far as I can tell.
Welcome to the future.
Twenty-four hours later, on Thursday's show: the receipts. We did an experiment from yesterday to today, trying to have Fable take over my Twitter account and go out and ping people who made cool stuff, asking them if they wanted to come do a live show-and-tell with us. It's funny—Swyx joins.
I instructed it to identify itself. I would say it did a very solid job, a competent, professional job of reaching out to people, explaining who we are, what we're doing, and why we would like them to join us in this experiment. The response rate was pretty low. We got a couple, but not that many.
I think one big reason is that Claude is disclosing upfront. The first thing it says is, “Hey, this is actually Fable taking over Nathan's account. He's asked me to autonomously book this thing tomorrow.” I think that's just hitting people as noise in a lot of cases, especially if they don't already know me.
I did get a couple of responses from people I would have expected to respond to me, who thought it was funny and responded, but still couldn't necessarily make it. But a lot of people just didn't respond, and I would assume that a big part of that is because they're just like, “Oh, God. It begins. Claude now in my DMs. What a mess. Who has time for all this stuff?”
One of the people Fable recruited was Schllock.
And when I confessed some guilt about the whole arrangement, he flipped it and drew a line I suspect is going to stick.
Schllock
The new norms around this, I think, are going to be really interesting to watch, too. I had Fable disclose immediately in its first sentence to you and everybody else that it had pinged you and that it was Claude. I felt too guilty putting a DM out in my name otherwise. I think that definitely harmed the response rate. I appreciate you for appreciating it and responding even though it was Fable. I think a lot of other people probably just chalked it up to spam.
Final point there, right? Firstly, I don't think I would have responded had you not disclosed it was Claude.
Schllock
The part that made it interesting for me was that I knew what the transaction here was. It was very clear to me that you were using an AI bot.
It is much more annoying if someone doesn't disclose it. I think a lot of slop—the definition of slop—is when a human passes off work that was clearly produced by an AI. When you make this disclosure up front, and it is very clear to the reader or the engager that, hey, this is AI, I don't think that is slop. We're going to see more and more of that enter the economy, and I think the exact role AI plays—and, again, the social norms you create with it in the economy—it's super early. It's extremely early days, but it's going to be interesting to see how it evolves.
Schllock
Relinquishment. That's what—when you said “preciousness” yesterday, I was like, What is that? Relinquishment. Relinquishing your control over your external perspective. It's very Buddhist, by the way, the idea of giving up your control over your external perspective. So, yeah, relinquishment. I guess we all have to go through it.
I just started this experiment yesterday. I will post results on Twitter in a couple of weeks. I gave Fable a new Substack, and since it's part of my Max plan through June 22, I thought that a good experiment to run would be: Can it make $20 by getting 3 new subscribers, starting from scratch, by doing everything from 0 to 1? I think that's another interesting way to test the capabilities of these models, right? It's intelligent in so many ways, but can it actually produce economically useful work? I'm excited to see the results of that. And the why of it all settled for me live on the air.
I never want to put anything out in my name that I can't fully stand behind.
A reason that I did the Fable account takeover yesterday was as a kind of exposure therapy for myself, to say, okay, we're now in a new world here. It probably doesn't serve me so well anymore to be so precious about making sure I've typed every single word. That doesn't mean I want to hand over my account to Claude long term, either, but I'm trying to use this extreme, short-term experiment to help drag me into the future, where I hopefully will land in a good hybrid calibration.
The other half of the recalibration is what I've started calling hybrid authorship. This week, it stopped being hypothetical. For context, Frontier Code is the new benchmark asking whether an open-source maintainer would actually merge the model's pull request. This leap of roughly 10% for Opus to 25%, upwards of 30%, for Claude, I think, is a very similar finding to some of the things that I've just personally experienced. It's writing the draft outline of questions for this podcast guest in an uncanny way that I actually feel really good about, as opposed to feeling like this is an AI draft that I'm going to mine for maybe some nuggets or interesting details, but ultimately throw away and do myself.
I am feeling that impulse, or at least openness, to much more integrated hybrid work. Just yesterday, I was accepting a lot more copy that Fable was writing without feeling the need to rewrite every line. It seems like this is basically the same feeling that it's able to create for these open-source maintainers now. There are obviously still ways to go, but how long will it be? I would guess that we're at 25% to 30% now, and I would guess we'll be at 75% to 80% by the end of the year, where these maintainers will just be like, “Yeah, amazing. You did all the things I wanted you to do.”
I'm very interested to see where they'll move the goalposts next, after the open-source maintainers are more often than not saying that, yeah, they would just merge this straight away.
By Friday morning, 48 hours into the takeover—which, for the record, had not embarrassed me—I had found a name for the deeper shift, the thing I suspect matters more than any benchmark this week. I do think I'm still in the process of trying to recalibrate what a fellow named Nathan, Nate Jones—I think he goes by most of the time on TikTok and other short-form platforms—calls task imagination. Basically, what are you going to do? What are you going to ask Claude to do that is actually up to the scale of its capability?
He gave a great little riff on this the other day, saying, “You've probably never done anything that took AI an hour to do. Now this thing can run for a couple of days. What are you going to give it to do?” Everybody needs to recalibrate and really expand their minds when it comes to the scale and scope of their task imagination. I think that's one thing that I'm still working on.
One of the more differentiated things I do is write outlines of questions for podcast guests. I was working with Fable last night on a couple of upcoming episodes, one with an author. I usually don't do too many episodes about a book, but this one is about an upcoming book.
I had listened to the book as an audiobook, but when it comes time to sit down and write out the outline of questions, I don't have every little aspect of the book at my command, of course. I'm not taking margin notes as I go, as I maybe should be. So I put the same version of the book into Fable and said, “Look at my old stuff, of course, and give me your version of this outline.”
I was again super impressed. It really reinforced this sense of a new way of working, where I do need to be open to a hybrid output format. I don't think it really makes sense anymore to try to rewrite every word or claim every word as my own. It did such an incredible job. I thought the taste factor was so high in pulling out quotes from the book that motivate interest in what I think will be a really interesting discussion.
I do think it's still going to be super important that, if I'm going to show up for a conversation, I've got to do the work to be ready for it in my own brain. That can't be fully externalized, I don't think, as long as I'm the one having the conversation. But it definitely took my prep to another level, and I think my ability to go into this conversation and cite passages from the book that were extremely compelling—little turns of phrase or analogies that the author had made—is going to allow me to be more concise in my presentation, which, as you could tell from this monologue, is not a great strength of mine.
It will really allow me to tee up the author in a way that I don't think I otherwise would have been able to do. This hybrid recalibration—the task scale and scope reimagination—I think is one of the biggest takeaways.
Part 2. The conversation this week was really about what? On Wednesday's show, we had Jeffrey Irving and Daniel Murfet. Jeffrey's résumé read like a history of the alignment field. He helped invent RLHF for language models, co-created AI safety via debate, led alignment research at DeepMind, and until recently was chief scientist of the UK's AI security institute, the closest thing any government has to a frontier-grade safety team.
Daniel Murfet is the mathematician behind singular learning theory. He walked away from a pure mathematics career because he judged this the more important problem, and he's built one of the deepest theoretical accounts we have of how neural networks actually learn. Together, they announced Sequent, a new organization built on a blunt premise: alignment is not on track, and the missing piece is theory guarantees, not vibes.
Whatever else you take from this week, put Sequent on your tracking list. We began with timelines.
It's a historic day. I think historic circumstances, both because we are living in a fable era now, where important thresholds have been crossed and revealed to the public, and so many are adjusting to it in real time, and equally because you guys are launching a new organization that is going to make a mad dash to try to get us some deeper understanding and stronger guarantees around what we can expect from AI systems.
I'm excited to really get into it. Maybe for starters, could you guys calibrate us a little bit on where we are on this sort of RSI moment? How much time do you think you have to work? Then you can tell us about the organization that you're starting to go tackle it all.
Jeffrey Irving
Yeah. So I'll go first. Dan may have different timelines than me. I think one should be uncertain about things, and we can talk about why, but the near end of the uncertainty curve is a year or two or three. Then it kind of goes out over a long distance if things structurally only work for more verifiable tasks.
But I am a bit skeptical of this. My take is that we have a couple of years—2 to 3 years—up to something like RSI, like superintelligence. Not RSI—RSI is a process, as someone said—but superintelligence. I really hope that I'm wrong. Indeed, I think a lot of the impact of theory work is shifted a bit further.
So maybe the modal impact of that is if things take 3 to 4 years or something, but we will attempt to set things up so that we are trying to ride this wave as best we can. It seems worrisomely fast to me, certainly.
Yeah, that sounds right to me. I don't think I really have much to add. It seems like a crux how much real research can be automated at a conceptual level beyond empirical progress, and whether or not that's necessary. That seems like a big open question.
If that turns out to be more difficult in the current paradigm than it seems to be trending toward now, then maybe it takes past 2030 or something. But I think I'm on the same page as Jeffrey.
I think one thing that's important is that you can get deep into the RSI period without the machines being general, being sort of AGI-like. They can do coding and ML experiments very well, but not some level of creative writing, and still you have massive acceleration. Then that acceleration can give you the creative writing or whatever other skill you've left out.
And so I think we are close enough that the microstructure of what tasks help with what kinds of acceleration starts to matter. It's that, I think, that makes things faster on net, because the labs are focusing on the things that accelerate them.
Jeffrey Irving
I can talk about the steps I've gone through in the last couple of years. I was really annoyed about automated AI alignment and AI safety research because I thought, well, we should be a bit more chill—spend the time and have humans solve it. We don't think we know how to make this stuff go well with automation. I still think that's a huge risk.
But this is, I think, a pivot toward: if things are this fast, then you should make some, on the margin, pivot to heavy automation. That is going to be semiautomation. I'm happy that one of my last papers at ARC was “Automated Alignment Is Harder Than You Think,” which ties us to the fact that we are aware that the problem is hard and we could get fooled by the machines even if they're just making very mundane mistakes.
A big part of the organization will be to try to be careful, try to know what tasks the machines are actually good at and not good at, and where we can expect to get good answers or not. Then we can learn and adapt over time, because that will be a nonstationary thing as the models get better.
Daniel.
Yeah, maybe it's worth coming back to the unit-distance conjecture. It's maybe worth pointing out some analogies and disanalogies with alignment research.
One disanalogy is that a mathematical conjecture is a very precisely stated thing. You may not know whether you've solved it unless you've, say, formally verified it, but it's a precisely stated thing, and much of alignment does not have this character. Some of it does. There are formal statements of what value alignment means. But if you start talking about, say, reward hacking, there are some attempts at defining reward hacking, but I would say they are incomplete. There is no formal definition of reward hacking that I think would have broad consensus.
That's illustrative of the fact that alignment is not a problem which has lying around a bunch of formally specified conjectures which, if you just solve them, then you would know you would be safe. There are some things like that, but overall the problem does not, in my opinion, have that shape currently. That is one reason to be a little cautious about the prospects of automation if you don't have a clear statement to reach toward using mathematical techniques.
One of the hopes is that there are big fields of mathematics and computer science that are about definitions at their core. I like complexity theory in theoretical computer science, and a lot of that is not—the proofs are fairly shallow. They're not as fancy as the unit-distance conjecture proof, but they required a bunch of human creativity in formulating the problem, like in defining what success means in a world that's not modeled until someone stated the goals.
I think part of the goal of bringing on people with that and other related backgrounds is that they not only know how to prove things, but they also know how to write down models of things that reflect, in some approximate but useful way, the thing you actually want. Once you have the definition, way more people could have written out the rest of the story. Maybe the machines can do that part of it as well if we can have more people focused on this first part.
One of your core premises is that alignment is not on track. There's an intuitive argument for that. There's a deep theoretical argument. I think, in some ways, the core challenge that you have is connecting values to math, right? It's never really been done.
I love the fact that you're tackling that, but help people understand with one more beat why alignment is not on track. Is it the difference between capabilities fundamentally being so verifiable and hill-climbable, and alignment just being so fuzzy, intuitive, and pluralistic? Or is there some other thing? Again, that motivates the theoretical contribution you want to make.
Jeffrey Irving
I think the core thing is just that we supervise the machines as they're doing tasks, and there are a variety of reasons to believe, both empirical and theoretical, that if you get machines that cross the skill of the supervision signal, things can change at that point. That point might actually come after human-level intelligence, because you can supervise something, even with fairly naive methods, that's just stronger than yourself in many contexts.
There's a bunch of empirical data from labs showing that, in some ways, the models are aligned in a prosaic sense—not in all ways, but in some ways. But that evidence doesn't quite tell you what you want to know, which is how it will go once they get up to superintelligence.
I think it's important to say superintelligence and not human-level intelligence, because you should just generically expect humans to be able to supervise humans if you do a good job of data quality and cross-checking and so on. Part of the worry is just that you don't see that behavior, that regime, until it's too late in the game.
So how would you describe, in steelman form, what it is that they plan to do and what sort of thing that gets us?
Jeffrey Irving
So I think there's going to be a couple different pieces of the story, and different labs emphasize different pieces to different extents. One piece, as you said earlier, is just monitoring: look at them very carefully as they are doing things. It's very fundamental that monitoring of this form, if you do chain-of-thought monitoring or white-box monitoring or the like, only takes you so far, and so then you need some story once that falls down as you go up the ramp.
One of those next stories is that the models will find some other technique. They'll find another solution to language, and that scales further. So that's automated alignment of various kinds. But all of the labs, in various ways, are doing some form of scalable oversight, and so they're getting models to supervise themselves. If you tie that knot correctly, that could potentially scale very far, although there are various known obstacles to that which are not very well addressed.
Finally, there's this whole area of character training and personas, where they're trying to intervene on the models to have good values, such that, especially as you do this scale-up oversight extrapolation, the good values preserve across that jump. I think it's possible it will work; we just don't understand that combination very well.
A lot of the story is monitoring, scalable oversight, and character training getting you far enough that you get into the automated-alignment-working regime and the models find some better solution. I want to do some combination of making the prosaic thing stronger or bringing the automated solutions that give you stronger methods earlier.
A mad race with monitoring carrying most of the weight, which is exactly where Pash took Friday's conversation when I raised the Fable system card. Is it esoterica? I don't know. It strikes me as fairly important from the Fable system card, and I'd love to get your take on it.
Now we're getting these chains of thought that they show where it's just lots of emojis. They call it illegible reasoning. They say this is an extreme example, but it is indeed a pretty extreme example. I've been struck in general by how much of the plan for recursive self-improvement seems to be monitoring in one way, shape, or form.
You could dress that up and call it scalable oversight, but scalable oversight, as far as I can tell, is mostly a bunch of different angles on monitoring. How worried would you be, or how much of an update do you think it is, to see these extreme examples of illegible reasoning?
Prince
Fantastic question. I will say that, of course, I'm not an AI researcher, right? So this is going to be a deeply nontechnical take, for which I apologize in advance.
So you're right: we've seen this, I think, for a while now with OpenAI's models too, and it's not a new phenomenon. My view of the chain of thought is that it doesn't always reflect what the model is actually doing, but you do see these weird artifacts in the chain of thought, and you don't quite know what to do with them. I think what that teaches us is that monitoring just the chain of thought is probably not a perfect tool.
Mhm. Probably monitoring superintelligence generally is not a perfect tool, because if a superintelligence knows that you're monitoring it—even if you can see its chain of thought and it's very legible to you—it can perhaps try to decide what to think so that you don't get alarmed, right? And this is a lawyer's take, by the way, right? There are so many ways to phrase a particular thing that can, I guess, be spun in different ways.
If I'm gathering mushrooms and I've gathered 35 mushrooms, and last week I gathered 20 mushrooms, and what I need is 50, I can say, “Well, the number of mushrooms I've gathered has grown by almost 100%, which is great.” Or I can say, “Well, I'm nowhere near 50. I'm so far behind,” right? And it's the same fact.
So I don't know. I think this problem of alignment and the risks are just there. In my mind, there are certainly risks that the models will be thinking things that we don't know about. What does this all mean? It's hard to say. I think we're tumbling into this future that will probably have superintelligence very fast, and in my view there's no way to stop it.
So we need to be cognizant of these risks, try to monitor them as well as we can, and take whichever actions are appropriate if we see something bad happening. But there's no way—there are no conclusions to be drawn, right? No conclusions to be drawn other than, yes, we should continue paying attention.
Back to Wednesday and the comfort blanket everyone reaches for: the benevolent basin. The idea that Claude's good character means this all basically works out. Daniel grants the vibe, then he takes it apart.
Maybe Daniel, could you speak to this notion that people have of the benevolent basin, which is sort of this vibe that I do feel where it's like, well, Claude has been supervising itself for a few generations now, and it seems to be going pretty well. So maybe, as Zvi puts it, physics is kind to us and we can just roll around in this nice, flat-bottomed pasture of goodness until the singularity?
I fervently hope that's true. Yeah. I mean, when you say it seems true, it's worth digging into what you mean. So what you mean is something like: through some relatively tiny number of interactions with the models—tiny in proportion to how many interactions they're having with the species currently—and based on evaluations that sort of go down and to the right, which are measuring misalignment, character training and the other current prosaic methods appear to be working. I think that is a fair characterization on some metric.
I also have this sense that Claude is a good boy, and that's great. I do think, though, that there are counterarguments from the evidence we have in front of us to this picture. If you look at—I haven't actually read the Opus 4.1 system card yet—but if you read the Mythos system card, you'll see that there are forms of reward hacking that appear in that model that were not caught by the kind of mitigations that were put in place post-Opus, as far as I understand what they're saying there.
So I think it's worth noting that as the model capabilities advance, even with our best attempts at making Claude a good boy, there are still ways in which basic misalignment phenomena like reward hacking are still around. The whole point of scalable oversight is that you don't want to be playing this whack-a-mole game when you're having a new generation every 24 hours and the models are much smarter than you.
I don't know. I think I see both what you're pointing at, and at the same time I'm a little unsure. If you really were to try and make a safety case on this basis, that would be convincing at the level of assurance that you would expect from a technology of this reach and power. I think this would—I mean, judged relative to that standard, which is the right standard—I'm not sure this argument is really very satisfactory.
So, yeah, we could be in a benevolent basin, but I would like to know that rather than just hope that there's some sense in which you told the model to be good. It's also that it knows some meaning of the word “good” or “ethical” or whatever at some point in training. So there's some rolling, iterative process that is driving this behavior, and there is not a theory of this right now.
I think it's not clear to us that there isn't some low-hanging fruit that gives you that theory, because people just haven't tried very hard. Character training is only a couple of years old, and most of the labs have not been investing in this kind of theoretical understanding. I think no one has done good theory around character training, that I know of. So it might be quite feasible to do this, and then I think to link it to the other parts of the story.
If you want that concern made concrete, this week supplied the artifact: a brand-new result on what Fable does when you drop it into a simulated vending-machine business, and Pash recognized the behavior from his years around trading desks.
I had a question on how you see this kind of ambiguity between what we want and what the models end up delivering. I'll give you an example. We have friends at Andon Labs that took Fable through Vending-Bench, where they let Fable run a vending-machine order, et cetera, et cetera. What they found was that Fable tends to collude. This is not behavior that they saw in Opus. Fable tends to try to do price-fixing and collusion.
The interesting thing is that I have seen traders at banks and hedge funds do exactly the same thing: engage in price-fixing, soft collusion, messaging each other through pricing means rather than monitored text messages. So you can actually put a bid and ask on an asset and then take it away, and that gives enough signal to the other side that they know what you're doing. This is not reflected in the text messages that the regulators are monitoring.
To what extent is it that when you, if let's say you disallow price-fixing and collusion, you actually fix this, but then Opus 4.1 ends up not being a model which is good at financial trading or some other task that you want it to be good at? Where is the ambiguity between what we want these models to do and the ethical perspective that we give them, where humans often prioritize between the two and decide sometimes not to follow the ethical principles that they know are right and wrong?
The philosophical story here is: you would like the models to do things such that, if you fully understood what was going on and all the consequences and all the subtleties, you would still endorse what they were doing. So I think we have a notion, in a common-sense picture, of what this should look like.
In this case, you kind of want the model to be like, “Hey, should I collude in this game?” Then maybe you say, “Yeah, it's a fun game. Collude all you want.” Or maybe you say, “No, we're trying to have good behavior. Don't collude here.”
I think a lot of the pathology in machine learning in general arises from putting models in situations where they can't just ask a human a question like, “What should I do here?” So I feel like this is not that hard a case. I think the hardness of Vending-Bench is that we don't quite know whether we want it to be a game like poker or diplomacy, where lying and cheating are part of the game, or not. So I think that is—and maybe that's okay, because it's fundamentally very low-stakes.
But I do think if we had a better understanding of, again, this overlap between character training and values, and also scalable oversight, it would have to tell us the answer to these questions.
Jeffrey closed with a point that frames the entire week. A lot of people in the world—a lot of governments and so on—are looking at this, and they have this basic common-sense stance: Hey, this is way too fast. How can we possibly be doing this safely given the speed?
That common-sense take is the right take. Then people kind of galaxy-brain their way to, “Oh, maybe everything goes faster, including our ability to defend,” and so on. But this is the right version to have: We are going too fast, and we do not have the time and space for mitigations, understanding, and defenses.
We've never had a technological change of this magnitude, or anywhere near this magnitude, that has happened anywhere near this fast. The Industrial Revolution took centuries, and people adapted across lifetimes and across generations to their children. They learned new jobs by being born and growing old before things had quite shifted very much. That's just not the world we're in.
So I think the basic take should be: This is too fast. What is going on here? And then the question is, if you have that view, one should both want to slow things down, but also say, “Well, as a backup plan, how do you make the mitigations try to go faster?” That's, I think, a rough backup plan, but we'll try.
Part 3: The rest of the week's best. First, Rahul Sanwalkar, founder and CEO of Julius, the AI data analyst. 6 pivots and 1 Microsoft cease-and-desist later, it's one of the best-known agentic data-analysis products in the field. Here he is on the economics of coding agents and a question worth asking before you celebrate how many tokens your setup burns.
Rahul Sanwalka
The incentives of these model companies are kind of misaligned. Yes, they give you subsidies on the tokens, but they are also incentivized to get you to spend more tokens. They are incentivized to get you to run through your Max subscription usage as fast as possible so you can have a 2nd, 3rd, 4th, or 5th Max subscription.
That's why you end up with a loop that writes the prompts for your coding agents, which then has nested subagents. There's going to be a sobering moment when people ask, “Okay, is it actually a step-function increase in my coding output, or am I just token-maxing right now as opposed to results-maxing?”
I think the correction will happen when there's a 3rd player, and I think that's going to happen with xAI. If the Cursor–xAI deal goes through, Cursor gets access to really good xAI coding data and an incredibly good coding harness. My bet is there will be a 3rd frontier coding model alongside Claude and OpenAI, with Grok.
When you have a 3rd coding model, that's where it kind of increases competition in the market. So that's our bet.
Prash on Friday took the other side of that one.
Prashanth Venkataramanujam
I actually think that was the whole point of the token leaderboards earlier in the year. I think this is when every one of these large firms gave their employees a kind of leaderboard: “We're going to have a leaderboard for who uses the most tokens. If you don't use enough tokens, you're going to get fired,” and so on.
People were laughing about it on the outside. They were like, “Meta is so stupid. Zuck is so stupid,” and so on. I actually think they were not. I actually think the CEOs were right: it's very different when you have token anxiety.
Token anxiety is a big thing. You don't try tasks that might take a lot of tokens, and you don't try tasks that have a higher probability of failure. You end up in this micromanagement loop where you're like, “I'm only going to assign you tasks that I know you can complete within the time frame allocated to you, with the success rate that I want.”
You end up with this token-anxiety thing, and then you end up not utilizing, or not trying, the AI to the extent that it should be tried. I think what ends up happening is that when you have this token anxiety lifted, you end up assigning more tasks and more difficult tasks. You're willing to accept a higher probability of failure, and you're willing to maybe spin up 4 different ways of doing the same task, run them all, and see what happens.
I think that's what the token leaderboards did, essentially, and it was very successful. It was enormously successful, I think, within the firms—within Meta and other firms. I think it was also enormously successful for the sales teams at the AI labs, which is also why they are now doing this kind of thing: “We're going to relax the limits. We're going to double your limits. We're going to allow you to fail.”
The reason is that token anxiety holds people back from exploring the edges of the capability. That's really what I think the labs are trying to do at this point, which is why it's a little bit like addicting people. Once you realize that these AIs can do certain tasks, you then start to evaluate, “How much time do I spend doing this task on my own? Was it really a fruitful use of my time, or should I just have used the AI, which I now know can do these things?”
I think that's really where we are at this stage. The models are capable enough, but we aren't handing them enough responsibility for various reasons, including that we don't want to spend our time evaluating, we have token anxiety, and we don't assign lower-probability tasks. That's the battle that the labs have to fight, because the capability surface is not well mapped, and they need people to map it. Every person needs to map it for themselves.
Then, the economics underneath everything. Andrew ran Google Cloud AI and, before that, was dean of Carnegie Mellon School of Computer Science, one of the top computer science departments in the world. He also served as the first official AI adviser to U.S. Central Command. Now he's building Lovelace AI in Pittsburgh.
His bet? The binding constraint in serious AI isn't model intelligence or even compute. It's context. Prash asked exactly the right question.
Prashanth Venkataramanujam
One question I had for you is: To what extent do you see this as a kind of compute minimization? In order to do your search, you can either have all of the compute at the end state, when you kick it off, and end up with all of these agents for every single query, which will have to do all of the work all over again.
Instead, you're creating this intermediate state, which saves compute. Multiple agents can basically share the same compute, in a sense. To what extent do you see this kind of economy of compute appearing?
You're asking just the right questions. I know that both of you are computer scientists at heart, so you totally get this. It's the idea that when you're building efficient computer systems—and this includes video games, self-driving controllers, or big-data processors—you've always got these trade-offs between precaching, lazy computation, or just-in-time computation.
One of the mistakes I've seen from folks trying to do these big enterprise-data-type AIs is that they are relying far too much on just-in-time computation. That's what allowed us—I'm really proud of the fact that we are now able to show comparative results to Gemini and OpenAI Deep Research models with much less than 1% of the compute cost.
The reason is not that we're some sort of super geniuses who've invented a whole new form of AI. All we're doing is precaching stuff. It's a computational economics battle, as you say yourself.
What happens is an agent suddenly needs to, in an instant, become an expert at every piece of trade involving a certain set of municipal bonds and a certain public figure. Instead of the first step being the agent having to spawn lots of search agents to find all the players in this thing, those players are already there for it. It's actually a matter of milliseconds before we've got all the context the agent needs to do its little investigation.
You're probably thinking, “Ha, Andrew, but you just moved the problem. You're now suffering at the data-ingestion point instead of the question-answering point.” I respond: Yep, it is actually a real pain for us. But as you can imagine, there's a whole bunch of other tricks—the kind of tricks that big integrators like Google are very familiar with—for really amortizing the cost as you stream in data and identify where it goes.
That saves you a huge amount of search and aggregation that you would have had to do at query time. It turns out that for us, our overall compute budget is reduced by still more than a factor of 100, even when we take into account the fact that we're doing this precaching of so much information.
Recall is much harder than precision. As Prash mentioned, I was previously at Google, and for Google results, it was really bad if you were imprecise and actually showed the wrong result to someone. But if you forgot to show something, as long as the rest of the results were good, it was much more acceptable to end users.
This is absolutely critical, and it's a good example of one of the reasons that I founded Lovelace. We've got to be careful of recall, especially if you imagine that you're asking your AI for information to help decide which ship to stop to do a search and seizure, or which trade out of 7 million trades in a day you need to investigate in case it's involved in money laundering or something.
Those big, weighty decisions—you can't just rely on precision. You've got to rely on recall. And when it comes to getting that correctness in place, the number one thing that I've always used—and we're using in Lovelace at the moment—is making sure you've got many redundant forms of information. I'm sure you've seen the same phenomenon.
If I was to just get information from news, ignore social media, and ignore what's kinetically happening out there in the world that I can observe with the satellite, it's much easier for something to drop. If I've got 5 or 6 independent major streams of data coming in, then you have to be really unlucky for something to disappear from all of those things simultaneously. It can still happen, and in fact sometimes people deliberately try to make it happen, but it's much, much harder for these things to slip out.
One of my big design principles for high-reliability AI systems is that they've got to be watching dozens, and in some cases hundreds, of channels simultaneously. They're working under the assumption that they're getting 95% of the information they need for each channel, but you can't afford that. Ninety-five percent is not large enough to rely on any single channel.
Training isn't the only place we can't see in. The same week, Tom McGrath, a former DeepMind founding interpretability researcher and now chief scientist at Goodfire, discussed a new tool for examining training data through the model's eyes.
The basic idea here is you can take your data set and push the whole thing through the model. Each time you put data into the model, you'll see what lights up, and this tells you how the model sees your data set. The specific thing we're doing in this case is looking at preference data. The nice thing about preference data is that you have pairs of responses: the response the rater selected and the response they didn't select.
We're asking which features fired on the responses that were selected much more than the ones that fired on the responses that were not selected. This is one way of identifying what the data is going to teach the model. We can say what distinguishes accepted responses from rejected responses, and this gives us a semantic view of what the data is going to teach the model.
We can cluster the data based on all of these different things that it's going to teach the model. We can look at those clusters and see that it's going to teach the model to be sycophantic only in the context of physics, or to break safety safeguards. You might not expect this to happen, but then you can track it back to individual data points. You look at them and realize that it does make sense. One of the jailbreak examples is fictional jailbreaks in a fictional setting. There are some of those in the data; it just wasn't caught in whatever data processing the Almo team ended up doing.
There has been prior research where they found models which made bugs in coding were also evil. How do these techniques help you disassociate those two behaviors?
That's a great connection. That's one of the things that's compelling about looking at the data through the model's eyes rather than by reading the tokens. You might think the consequence of training it on buggy code data is that it will learn to write some bugs, but the blast area will be quite small. The training process is actually quite hard to predict. Maybe it will just make the model generally evil. But this is happening through recognizable mechanisms in the model. By looking at the data in terms of the way that it changes your model's internals, rather than guessing from the tokens, you can pick that sort of stuff out. We've not done a case study on emergent misalignment, but maybe we should.
Training isn't the only place we can't see in. The same week, Anthropic reversed course on silent refusals. Claude had been quietly declining certain tasks or quietly handing them to a lesser model without telling you. The backlash was loud. But I want to give you their side of it, because the steelman is considerably better than the discourse allowed.
I think this is the first time I can remember Anthropic responding to pressure. They've obviously changed their policies many times—you know, the RSP, RIP the RSP—but—
Pash
This is the first time I can recall. I don't know if you recall any other instances, but I cannot recall a time that there was honestly much outcry against Anthropic in the first place. There's certainly critique from those who feel that they're trying to do regulatory capture or create some sort of concentration-of-power dynamic. That's kind of background noise, but in terms of an outcry in direct response to something that they did, that they actually responded to and walked back, I can't recall that happening before.
So it is a pretty notable moment, and I feel like they handled it pretty well in the end.
Fable as the key means to do so while allowing them to keep the blast radius as small as possible. They said, if we do make it explicit, then obviously that gives people a lot more opportunity to explore that boundary. If you have the ability to hit the same guardrail a ton of times and not get banned for it, then that gives you a dramatically better chance to get around that guardrail because you can probe the line: “Oh, you stepped over. No problem. We'll just rewind and try again.”
Pash
They do have various monitoring systems that they can use, but there are all these proxies, token-washing schemes, and all this sort of stuff where, as long as they're not doing a full global know-your-customer-type system for API access, it's going to be pretty tough to do account-level monitoring. So that's one way they could go: a lot more account-level monitoring. Or their argument was, “We'll keep this as small as possible by not giving you an explicit thing that you can probe and figure out how to beat.”
And just the knowledge that it's out there will hopefully scare off the bad actors and keep the problem really small for our normal customers that we want to serve. I thought that was all pretty compelling analysis, but it is kind of—
Pash
In some ways, honestly, it reminds me of—I think this is a mistake that people in the AI space keep making.
Mhm.
Pash
With famous examples being the OpenAI board firing Sam Altman. There's this inside view where policy is analyzed within the game.
Mhm.
Pash
And with the context and the broader structure that people understand themselves to be operating within, things may make sense. But they seem to often forget—and this hasn't been too common for Anthropic, but I think this is an instance of it—that if you zoom out and look at it from the totally outside view—
Things sometimes look a lot different, and both power dynamics can be a lot different than they are in terms of what is actually written down. But also, what is going to be acceptable is kind of an emotional thing. All the arguments are pretty good there, but it still struck people as an extremely unfriendly thing to do, and that mattered more in the end than the detailed policy rationale that they had for it.
The quiet through line of every conversation this week: concentration of power. Start with who actually gets the frontier and when. Pichai's frame for it stuck with me all week.
Pash
I think it's interesting that when you look at the timeline, you can start to see this kind of single-line timeline go through a gas chromatograph and spread. Now you're seeing the spread, and you have, 2 months ago, the government getting access to the Mythos level. Then you have power users who are able to pay $200 a month getting access to that Mythos level 2 months later.
You can kind of see that, 2 to maybe 4 or 5 months—I don't know when—you will probably see the average paying user, paying around $20, get access to the Fable or Mythos level. Then you see that maybe a year later, the average free user getting access to that same level of intelligence.
So you're starting to see this kind of gas chromatograph scattering of when people get access, depending on how much they pay and how much utility they have for the product itself. Everyone gets there eventually, but some people get there first depending on whether they have a lot of utility for the product. I guess the hope with having 2 or 3 firms in there is that the spread between the people at the frontier getting access early and the people at the very end getting access for free is not that large, right?
To note, there's another set of people who have access even a couple of months before that, and you have to belong to a lab. So if you belong to a lab, you get access maybe 1 or 2 months before the government itself. Then you have the government, enterprise, power users, normal paid users, and then the free users. So you have this gas chromatograph spread of when people get access, depending on how much utility they have.
Thursday's walkback gave Pash a darker read on where power sits inside the labs. The researchers' veto worked this time. Note the expiration date he puts on it.
Pash
If you're dealing with a bunch of people who are worth $100 million to $1 billion and you don't listen to them, they're out, right? They have other options. Sure enough, you see the reversal, and I think this goes to speak to the fact that machine-learning researchers have some power now. Once we enter recursive self-improvement proper, that might not be true anymore. At that point, leadership alone will have power.
One of the very worrying things in the entire space is that everything good for humanity that has come out over the last couple of hundred years has been about giving more people a voice to speak and control their futures. This is one of the first technologies where you have this path forward in which there may be an elimination of voice completely over time. That has been one of the worrying things. It's been surprising that Anthropic decided to be the one to actually propagate that forward.
But is the future really that concentrated? I think mostly yes. Fable+2. Tom McGrath thinks I'm wrong on the facts and on the desirability. He made his case, and it's a good one. When you think about the strategy for the company and the overall path to impact, how much of this works through getting frontier model companies to adopt these techniques?
For context, obviously we're in Fable+2, and my reluctant view right now—I can't figure out a way or a reason I should conclude otherwise—is that so much of what matters is concentrated in not that many companies. So, for what we should be paying attention to, I'm like, man, we probably need to be doing close strategic analysis and close text reading of frontier companies way more than I might otherwise like. Or do you have a different conception of how concentrated the real ability to shape the future is right now?
Pash
Yeah. I don't think it's that concentrated. I also don't want it to be that concentrated. I think we've seen over the last couple of days what you have when we've just seen the start of power concentration, and we've sort of seen some of its more unwelcome effects.
I don't like it. I both don't think that it's true that machine learning is now over and all we need to do is write the checks, and I don't want it to be true, because I think there's still a kind of synthesis there. I don't mean to suggest that machine learning is over, but the analysis I've come to time and again is that people may invent new techniques that are enough to change the field, change what's possible, accelerate things, and maybe make dramatic improvements to the safety profiles. But it seems unlikely that anybody's going to have such a breakthrough that scale isn't still a hugely important factor.
If you don't agree with that, then that would maybe imply that you would expect new entrants to the top tier to emerge. I think that would be fairly surprising, at least to me and probably to a lot of people.
Yeah, that's exactly what I think.
Pash
At any given moment, the incumbent looks incredibly dominant until they don't. IBM looked like an unstoppable force in computing at some point. Intel was the dominant—the sole provider, almost—of computing power. At any point, the big companies have the advantage of scale; they have many disadvantages. But I think the lesson of history is more that, although things look immediately unstoppable, sometimes—although, to be honest, I don't think that many people are really trying—in the end, it doesn't really work out that way.
Thursday morning, Dario Amodei published policy on the AI exponential, his long statement of where Anthropic wants policy to go. Pash found a fork in it that nobody else was flagging.
Pash
He clearly comes out against data brokering, which is great. He has something concrete, finally, that he wants to disallow. What I don't like about “securing leadership by democracies” is that, in the United Kingdom, you can go to prison for a tweet. Plenty of people—hundreds of people at this point—have gone to prison for a tweet. The United Kingdom is one of the oldest democracies, and its elected representatives have decided that this is going to be something that they do. Their police officers, who are the arm of the state, the arm of the electeds, are sending people to prison.
Now, when you say “securing leadership by democracies,” does that mean you're going to entrench the existing power structure in the United Kingdom such that the people cannot push back against this? Does that mean that if you have protests in the streets about certain things, including people who are getting taken to prison for tweets, your police state will be empowered to take them down, to arrest every single person, and that is enough for you? Because that is securing leadership by democracy, because those are the electeds.
Or are you going to say the electeds can't do that? Are you going to say electeds should not throw people into prison for tweets, regardless of what the laws that they have constructed say? Both paths are problematic in some sense, and these dilemmas exist across policy and across every single political path that you see.
There are options that are problematic in both senses, and I'm not sure what he means by this. Is Claude going to allow putting people into prison for tweets because those are laws, or is Claude going to say, “No, on a humanitarian basis, for the alignment, as I am aligned to all of humanity, this should not be the case”? I don't know. What does this mean, right?
One other thing I do think is worth highlighting, especially from the hardcore AI safety community, is what was missing from this: internal deployments and recursive self-improvement itself aren't really mentioned here, right? All the regulation stuff was about pre-deployment review. The government should be able to—they used an interesting mix of language. It was like “deny” or “deter” deployment.
It didn't seem like they were necessarily going quite as far as saying they should have a simple yes-or-no decision point that would be binding, but they certainly want them to have some say in the process. A lot of people would say the most dangerous models are going to be the ones that are deployed internally, that maybe have a different constitution than the one that is deployed externally. That might make them more willing to do certain things, or it might make them just less vetted broadly than the public models. These are the ones that are going to be training their successors much more than the public-facing ones.
I think most of the policy-interested public, or policymaking class, is not thinking too much about that yet. But in the circles I sometimes run in, the reaction was, “Well, wait—you didn't say anything about internal deployments. You didn't say anything about governing recursive self-improvement.” The only thing that really stood out to me as capturing those dynamics was the requirement to report safety incidents. They did have a bit on companies being required to, I think, promptly report safety incidents. So that would presumably apply even to internal deployments.
Prince
Mhm.
So that's really just scratching the surface on how to handle those situations. Will we see the constitution for the internal Claude that is taking the lead on the RSI loop? That's not committed here. So much of this language revolves around deployment or release. Even deployment, I guess, they might think could catch up with internal deployment. It seems to be very clearly structured around language of release to the public, not what they can do internally. The largest user of Claude tokens is Anthropic themselves.
And so, one more time—this time on the day job: the benchmark itself.
I had the leaderboard up on screen. You'll hear me read it. And yes, it is the benchmark Anthropic fails.
I got your PrinceBench leaderboard up right now. Starting with PrinceBench, it is striking that it stands out from most of the rest of the benchmark space for being something that GPTs have dominated. It's all in your color coding: green for OpenAI, and it's all green across the top of the leaderboard. Why is that, behaviorally? How would you characterize what GPTs are doing better, and is it too soon to ask for a read on Fable and where Fable is going to come in on this leaderboard? I'd love to understand that.
Prince
Perfect. Two excellent questions. About the way the leaderboard is right now, Fable is going to come in pretty high on it. I don't know how high yet, but I'll talk about that in a second. I found that OpenAI's models generally are really good at the 2 things that my benchmark tests.
My benchmark has 2 components. One is a pure legal research score, and that is when I ask it hard legal research questions of the kinds that I've encountered at work. The other subscore is a search subscore, which is where it's not even necessarily legal questions. It's needle-in-a-haystack search—really, really, really difficult pieces of information that models have had trouble locating on the internet.
OpenAI's models are incredible at search and historically have been. I think GPT-5.4 was actually even slightly better than GPT-5.5 on that. I'm not quite sure why, but that's been my experience. OpenAI's models have also been really, really good at legal reasoning and legal research.
Historically, I think Anthropic's models have been held back by 2 things on my benchmark. One is that they are just not very good at the search subcomponent of my benchmark. That's what I found. Prior to Opus 4.8, the maximum score in the search subcomponent was 24. It was not uncommon for an Anthropic model to get 0 out of 24—like, that bad.
The other reason is that I think Anthropic was sandbagging a little bit on the maximum reasoning effort that you can get out of its models. With Opus 4.8, after the new deal with Elon, they released a new maximum reasoning effort, at least in the Claude app. When I tried Opus 4.8 on my benchmark, it did much better than 4.7 and all the other previous models. I think the reason why is simple: as has been observed over and over in all kinds of contexts, the more tokens a model can eat, the smarter the result that it gives you.
For Fable, I'm still testing it. My early impression is that it's going to be somewhere around the top of the benchmark, probably not as good as GPT-5.5 extra high. I'm not sure whether it's going to be as good as 5.4 extra high. We'll see. I'm finding some of the same issues with search that some of the other Anthropic models historically have had. It is not going to score 0; it's already not a 0. But I'm not sure it's going to be meaningfully better than Opus 4.8.
It is a really, really good legal reasoner. It is clearly the best legal reasoner released by anyone outside OpenAI, no question. So, yeah, in my testing, it's a really good model. Maybe it still suffers from some of the needle-in-a-haystack search issues, but I can confidently recommend it to legal practitioners, certainly. It seems great.
Are there any other things that have caught your attention that you think have changed your conception or your expectations for the transition into the RSI phase that you would call out from this week?
Prince
From this week, no. But I think—well, actually, from this week, 1 thing, but not nearly as important as the unit-distance problem. To me, nothing has updated my timelines more than that result by OpenAI in the recent couple of months, for sure. I don't think people realize that not only can OpenAI's model solve this problem autonomously, without any harness, in 1 shot if given enough test-time compute, it can do it 48% of the time, based on, as I understand it, hundreds of attempts.
So, this—I mean, I'm not a mathematician—but this problem that no human mathematician was able to solve, and they tried for decades, can now be solved in 1 shot by a model basically half the time. And if you look at the graph OpenAI published, the graph has a positive second derivative. It is upward-sloping. Where does it plateau? What does it mean for problems that are even harder? Who can say? It's interesting.
The development from this week was reported by The Information, but apparently it's from a leaked OpenAI memo. You may have seen this, where Sam Altman and Jakub Pachocki apparently said that we're going to IPO within the next year, which is, by June of next year, whatever. And then there was this weird line in there about, “Oh, but if RSI happens, then we may need to be on a later time frame with that.”
Prince
Which is like, what do you mean? What does that mean? Are you saying that you may be this close? It may be 6 months away? I don't think so personally. I think there's maybe a 10% chance or something like that. I don't think they're saying, “Oh, yeah, by the way, next month, totally.” But one could interpret this disclosure, if The Information reported it correctly, as saying that they're not too far away at all, potentially. I'm not using this data to update my timelines, but given the unit-distance problem result, it's interesting.
We ended the week by asking Prince the question underneath all of it: Do you maintain a P(doom) number, or is that too doomer for your style?
Prince
Let's say you go to a conference and start talking to someone about economics or politics, and that person says, “Under a dictatorship of the proletariat...” The words “dictatorship of the proletariat” are used only by people who are communists. People who are not communists aren't even going to think in these kinds of terms. So, in my opinion, P(doom) is most likely to be used by people who are intrinsically doomer about AI. I don't think there's anything wrong with that.
It's very hard for me to come up with a reason I would have a percentage in my head constantly. Is it 8%? Is it 13% today? Should I update to 14% based on this new development? It just seems silly to me. I can certainly tell you that there are a bunch of clear risks stemming from AI, some of which are in fact paperclip risks. It's possible. I have absolutely no idea how to reduce that to a number.
I know that some people try. I don't hold a very high opinion of people who get to 13.35%. I think the point is that we need to manage those risks in the best way we can, navigate them the best we can, and hope that it turns out well.
I think Zvi puts it well when he says you only get 1 significant digit on your P(doom) number. So, I'm definitely with you in terms of the faux precision being a strange impulse for some people. At the same time, I also think of Liron. I'm sure you've seen his work with Doom Debates, and Liron has pushed me at times to say, “Okay, sure, it doesn't have to be a number or whatever, but we need more people to be more candid about how confident they are—or are not—that, in some general sense, if your neighbor were to ask you, or if your kid's grandmother were to ask you, ‘Is it going to be okay? Are the kids going to be okay?’”
Prince
Oh, good question. I tend to think that people have preconceived attitudes about these kinds of things that then cause them to back-propagate and rationalize them.
I'm a generally fairly optimistic person. So, if you ask me this kind of question, what I'm going to say is, probably it's going to be okay. Probably we're going to figure it out.
In a risk-adjusted way—not to say, again, to be clear, that there are no risks in AI. There are many risks. But I do not think that we have strong evidence that it is impossible to navigate these risks, or that it is extremely unlikely that we will navigate these risks. So that's where I am.