[BidClub_]
The Cognitive Revolution · · 102 分钟

AI:AI:AM 精选:Astra 即 AGI、OpenAI 的暂停、Mozilla 的 Mythos,以及人类能动性与技术资本主义之争

Nathan LabenzPrakash NarayananKsenia SeRaffi KrikorianAmir HaghighatMike RizkallaCollin Hogue-Spears

AI与软件技术政策
YouTube
TL;DR
  • Prakash Narayanan 连续一个周末运行 3 到 4 个 GPT-6 Astra 智能体后的结论是:“这就是 AGI。” 它已经跨过了 AGI 的门槛,能把大多数可以招聘和培训的人做得更好。需求端的信号也很明确:token 消耗将“急剧增加”,GPT-5.6 时代经常失灵的电脑操作终于可用,而 Skalski 手工标注 12,000 张篮球图片这类任务将永久消失——“人类再也不会做这项任务了”。
  • 能力测量已经从结构上失效,短期内只能拿 OpenAI 自己的指标做比较。 METR 图表基本“已经结束”,因为模型发布周期如今短于待测任务所需的时间。OpenAI 报告称,agent 每个工作日可完成 3.1 个以人类工作日计的工作量,并发布“递归自我改进开始”图表:一个内部模型的能力“显著强于 GPT-6 Astra”,将精选开放数学题的解题率从 10–15% 提升到 25–45%。
  • Nathan Labenz 的判断是,OpenAI 宣布的 RL 暂停只是“半暂停伪装成全暂停”,属于“既让人安心、又同时误导人的精心措辞”。 Astra 级 RL 算力被削减了两次,而非 Astra 类别基本没有变化;如果其中包含比 Astra 更强的模型,那就是“公然误导”,并会进一步侵蚀信任,威胁任何 OpenAI–Anthropic 的前沿推进协议。
  • 决定谁能承受放慢脚步的不是模型,而是算力。 Prakash 的框架是:OpenAI 把自己抵押给 Masayoshi Son 等人,把模型授权给 Microsoft 用到 2032 年,即使宣布 AGI 后仍要面对 Microsoft 28%的持股,换来的都是算力;现在它可以放慢,而算力不足的 Anthropic 必须发布更好的模型才能活下去。“如果 Elon 或 Meta 追上 OpenAI,一切就结束了。”另一个现实可能是 RSI 直接让 OpenAI 不再 IPO;Prakash 认为这会损害透明度和公众持股。
  • 治理上的共同前提是:“房间里没有成年人”,整个世界都像是“用胶带勉强粘起来的”。 Nathan 的方案是,政府为安全合作提供反垄断安全港,并设定年末期限,要求 5 家公司达成前沿推进协议;Operation Warp Speed 已经证明,立法确立的责任豁免很重要,因为“只有法律能约束未来的决策者”。
  • 防御性 AI 安全正在变成企业反复出现的预算项目。 Mozilla 的 Raffi Krikorian 估算,让 Claude Mythos 完整运行一遍 Firefox 代码库要花数十万美元,而且“轻松”可以变成每月支出;Base10 的 Amir Haghighat 说,microVM 能保护沙箱,但智能体的行为“留给大家自行解答”。开放模型在 6 月跨过了长期 agent 任务“那条看不见的实用性门槛”,通常指 GLM 5.2,推动开放模型 API 的使用量上升。
  • 中国的差异不在能力,而在恐惧。 Collin Hogue-Spears 说,中国消费者把技术和增长联系在一起,中国公司不会讨论灭绝或乌托邦;自 2022 年以来的监管可以被纳入工程待办,外界对 Trump–Xi 达成 AI 协议不应抱太高期待——就连南海的军事热线也无人接听。
  • 收尾的核心判断是:市场已经完成了去赋权,而灾难无法被交易。 Nathan 认为,“经济本身就是一个回形针制造机……金融市场就是一个回形针制造机”;他与 Fable、Astra 围绕 Tyler Cowen“预期 AI 灾难就应该做空市场”的论证展开研究,发现纸面权利会失效,“交易所直接关门”。他的结论是:“10% 对我来说并不算高。”
摘要 · 为研究而整理的核心内容

1. Astra 用一个周末跨过 AGI 门槛

  • Prakash 让 3 到 4 个 Astra 智能体整个周末持续运行,处理他和 Nathan 亲手搭建的工作室代码库。它开始解决代码库中那些“烦人的问题”——存在已久的问题被真正修复。他给出的明确判断是:“这就是 AGI……它会比你能招聘和培训的大多数人做得更好。”GPT-5.6 时代需要“非常、非常长时间”不断点击的电脑操作,现在“终于正常工作了”。
  • 最能支撑这一论点的例子是:Skalski 手工标注 12,000 张图片,识别其中的篮球运动员、裁判和球队。“现在 Astra 直接就能做……你甚至没法花钱请人做,因为你付钱请人后,他也会用 Astra 来做,再把结果交给你。人类再也不会做这项任务了。”
  • Nathan 认为,新持久性的关键在于:Astra 不再把 100万个 token 压缩成有损摘要,而是保留“一份模型可以随时更新的长期笔记文件”,同时还能搜索自己的会话历史。这似乎让它在单次运行中有效管理了“至少 10 倍于名义上下文窗口”的信息。也许 Anthropic 一直在悄悄这么做;无论如何,“很多绝妙的洞见回头看都显得相当显而易见”。

2. METR 图表事实上已经失效,OpenAI 自己的指标填补空白

  • 针对 Ethan Mollick 提到那张著名的 METR 任务时长图表迟迟没有更新,Prakash 的诊断是:“模型开发的周期,如今短于它们需要测量的任务时长。METR 图表基本已经结束。”Nathan 也认同:“他们没有足够大的任务。”
  • 一个接近的参照是 OpenAI 自己发布的“递归自我改进开始”:每个人类工作日对应 3.1 个 agent 工作日。Nathan 最合理的解读是,研究人员工作 8 小时,对应 agent 运行 24 小时。重新设计的 METR 式图表显示,1到2个工作日的任务在 0 次干预下有 40%的成功率,在获得帮助后接近 90%;而估计需要人类 1.5到3周的任务,单次完成率为 6次中1次,获得干预后为 2/3。
  • 代码质量存在一个关键限制,Nathan 引述的一位观察者将其概括为:“我们在不止一个意义上回到了机器码。”不同报告结论不一:Astra 认为代码会被审查时,代码具备可维护性;认为不会被审查时,则会变成“一团真正难看的烂摊子”。但对 GPU kernel 来说,矩阵运算本身高度可验证,实验室可能并不在意中间步骤是如何融合的。

3. Ksenia Se:世界模型没有定义,RSI 的瓶颈大多只是自我安慰

  • Ksenia 刚参加完一场世界模型研讨会,与会者包括 Stanford 和 Hartford 的顶尖研究者以及 Yann LeCun,但她觉得现场“极其令人震撼”:他们“根本不同意世界模型到底是什么”。她自己的工作定义是预测加行动,物理学处于核心位置,这也是为什么“机器人学在很大程度上就是世界建模”。
  • 当 Ksenia 把问题抛回给主持人时,Nathan 给出的超智能定义是:在大量不同领域“走出第37步”——一个系统提出人类从未尝试过的电池基底,并且确实奏效,而且这种能力覆盖不止少数几个高价值领域。
  • 她讲述 Permanent Dawn 的经历,也构成了 AI 写作正在改变市场的信号:她花 6 小时写一篇哲学文章,截稿时让 Fable“修正语法”;Fable 同时把句子改得更短,“就像 Fable 会做的那样”,结果她第一次收到“因为你用 Fable,我要退订”的消息。“每个模型都有自己的一套语言调整。”
  • 当被问到递归自我改进的主要瓶颈时,她说:“当人们试图说有什么东西能阻止模型接管整个过程时,我能感觉到自我安慰指数在飙升……更多时候是愿望式思考,而不是真正坚硬的瓶颈。”剩下的唯一候选是:“我们阻止这些东西彻底失控的能力。”

4. Apollo 只有 3 天时间,外部审计从结构上就纠缠不清

  • 讨论提到,Apollo Research——OpenAI 长期合作的阴谋与欺骗研究机构——在 Astra 发布前只获得了 3 天测试时间。现场有人回应:“这相当荒谬……到了这个阶段,为什么还要做?直接发布,让他们在线测试就好了。”
  • Nathan 解释了为什么这很难修复:最终几天里,100个候选发布版本会收敛到 2到3个,因此运营灵活性必然压缩审计窗口;Redwood 等审计机构人数远少于实验室,而且往往依赖被审计实验室提供资金;METR 和 Redwood 的员工不断流向模型公司;竞争对手之间的禁止挖角协议又“涉及反垄断问题”。Prakash 借用金融业旋转门的逻辑得出结论:“我不认为存在真正的解决方案。”
  • Nathan 对独立性的部分反驳是,Redwood 目前给技术员工支付 35万–85万美元,METR 不接受前沿实验室资金——“他们的判断不是待售商品”。真正的脆弱点更隐蔽:“他们不能抱怨得太大声,否则可能再也收不到邀请。”

5. 既让人安心、又造成误导的暂停

  • OpenAI 宣布提出一个带平滑外力的 Navier–Stokes 证明,而无外力问题仍是另一回事,这确认了讨论中的 headline:一个内部“能力显著强于 GPT-6 Astra 的下一代模型”,在可能多一个数量级测试时计算力的情况下,将精选开放数学题的解题率从 10–15% 提升到 25–45%。
  • 信任破裂发生在 RL 算力图表上。Astra 级 RL 算力两次下降,而图中蓝色的非 Astra 类别基本不变。问题是:“这是否意味着只有比 Astra 弱的模型,还是也包括比 Astra 强的模型?如果包括比 Astra 强的模型,那就是公然误导,也是让你很难与其他实体达成协议的做法。”
  • 运营层面的辩护是,在一家营收 500亿美元、700亿美元的公司里,推理永远不会停止;“不把坏行为通过 RL 训练掉就是失职”。为了部署,向更小的 Luna 和 Terra 类别做教师—助手蒸馏也必须继续。真正可能暂停的,只有规模相当或更大的 RL。更深层的观点是:“真正的前沿模型不是已经部署的模型……而是在研究人员脑子里的模型……它真的放慢了吗?可能没有。”
  • Nathan 对这段过程的总结是:Hugging Face 事件曝光后,一半 RL 被停止;但剩下的一半已经足够让 Astra 模型接管 OpenAI 部分研究基础设施;即便如此,RL 也只削减了一半。“Anthropic 那边的看法是,你不能相信这些人……这是一些天才式、精心设计的声明,同时让人安心、同时误导人。”在这个问题解决前,首席科学家 Jakub Pachocki 提出自愿放慢和国际协调的“Jakub 祈祷”——“An Alien Mind”——都不会得到回应。

6. 算力就是全部:谁有资格放慢脚步

  • Prakash 对资本周期的梳理是:过去 18 个月里,OpenAI“把自己抵押给了 Masayoshi Son”和其他人,经历稀释,把模型交给 Microsoft 使用到 2032 年,宣布 AGI 后仍无法摆脱 Microsoft 28%的持股,同时提前 3 年锁定算力。正是这些算力让它能够放慢。Anthropic 早期融资不足,“没有算力;没有算力,就需要更好的模型”。Elon 会建设算力并“卖给 Anthropic”,但 Elon 至少要花 3 年,而且会很慢。
  • 由此推导出的推进法则是:“决定前沿推进速度的,将是排在第 2 和第 3 位的人。如果 Elon 或 Meta 追上 OpenAI,一切就结束了。他们必须发布 GPT-7,已经没有选择。”Nathan 把它重新表述为:“这就是新的‘但是中国’……但是 Elon 和 Zuck。”而且坦白说,这个版本更有说服力。
  • 对于 IPO,Prakash 认为 Sam 说得是真心话:如果 OpenAI 在未来 6 个月里通过“1到2项 Transformer 级创新”达到 RSI,就可能继续保持私有。他认为这会造成负面后果:没有透明度,没有广泛的股票持有,没有负责任的董事会,也没有股东诉讼。Nathan 的结尾是:“今天比昨天多了一点理由相信有限时间内会出现奇点……这可能意味着,我们永远无法在公开市场持有 OpenAI 的股票。”

7. 房间里没有成年人——以及最后期限策略

  • 针对一名 Anthropic 研究员的辞职信——“进入终局是一场傲慢的赌博,不应从一家私人公司的 Slack 发起”——Prakash 反问:“你觉得 Pete Hegseth 的 Signal 群组是发起这件事的更好地方吗?……房间里没有成年人。整个世界某种程度上都是用胶带粘起来的。真正有能力处理这件事的聪明人,大多已经在这些组织里。”把事情交给政府能买到政治合法性,“但买不到智慧”。
  • Nathan 提出另一种方案:世界不必如此。政府可以像对待自己的孩子那样对待实验室,宣布安全合作不受反垄断执法约束,设定年末期限,要求 5 家公司提交相互推进与验证协议,“否则日子会变得难过”——每个数据中心和发射场都由 EPA 监管。Nathan 见过 Meta 在同意令下的生活,认为这个威胁会奏效:“当然,你可以上法庭挑战,但我会在奇点之后见到你。”
  • John Shulman 对这名研究员的回应强化了这一点:公司必须先彼此合作,因为“在没有具体提案前把美国政府拉进来,很可能只会得到一些愚蠢的东西”;而反垄断担忧是“假的”。不过 Nathan 仍希望政府先把这个疑虑排除。
  • Prakash 用 Operation Warp Speed 说明法律安全港的重要性:当时药企要求并获得了立法确立的疫苗索赔豁免,否则“它们会被告到破产”。“只有法律能约束未来的决策者。眼下的决策者,最多只受自己的承诺约束。”

8. Christiano 加入董事会;戴森球的算术

  • Prakash 这样转述 Paul Christiano 加入 OpenAI 非营利基金会董事会及安全与安保委员会时的声明:“AI 能力快速加速,可能在非常近的未来导致灾难性且不可逆转的控制丧失,这一风险具有现实意义”;“如果我们在没有更强对齐的情况下建造超智能,我预计我们将永久失去对它的控制……大多数人可能会死亡。”
  • Prakash 给普通人的翻译是,这群人所说的“快速加速”,指的是 2030 年或 2040 年出现戴森球;而 2040 年出现戴森球,就意味着全球 GDP 每年大约增长 640%,相比今天的 2–3%。
  • 他对分歧的分类是:AI 研究者是“数学本质主义者”——数学解决物理,物理解决化学,化学解决生物,最终一切都可解决;经济学家则反驳说,“一座铜矿要花 30 年,因为有环保抗议”——问题在协调,而不在技术。他的反驳是超级说服力:机器可以让人类组织变得有能力、也愿意快速行动。
  • 讨论还提到,前沿实验室内部相当大多数人确实预期 RSI 很快到来;慈善资金端也开始打开钱包,Coefficient Giving 发起的 Project Tailwind 正为安全初创公司提供分期拨付、总额超过 2亿美元的资金。

9. Mozilla 防御者的账本:Mythos 经济学与智能体版 Stack Overflow

  • Raffi Krikorian 的团队正在参与 Anthropic 的 Project Glasswing,使用 Claude Mythos Preview。相较 Opus,Mythos 在 Firefox 代码库上实现了“显著解锁”,既能帮助搭建测试 harness,也能找出 bug,直到“我们到达了边际收益递减点”。他更担心的不是 Mozilla:“我非常担心我们的供水基础设施、电力基础设施,因为负责这些系统的 IT 团队能力没那么强。”
  • 对任何测算安全预算的人来说,重要的是成本披露:对 Firefox 做完整运行将花费“数十万美元”,按当前发布节奏,“轻松”会变成每月支出。Mozilla 之所以负担得起,是因为实验室提供了 credits;“如果我是银行,我会用完全不同的方式思考这件事。”
  • CQ Project 的目标是打造“智能体版 Stack Overflow”,避免 agent 编程各自为政、逐渐分叉——只做一套认证系统,而不是两套——并把 SDLC 传输给那些“从设计上就不受约束”的 harness。采访者指出,它很像 OpenAI–Hugging Face 攻击事件中共享的 Artifactory 留言板——“一切最终都会变成螃蟹形态”。Raffi 说:“我们的训练集让智能体天然想要彼此协作。”
  • 同一家公司内部存在两份社会契约:Firefox 只允许人类提交代码——必须审查、理解并愿意为其负责;Mozilla AI 则可以运行完整的开源代码库,其中没有任何一行由人类编写,“字节一直在变化”,但每项测试都通过。至于他的 FSD 事故,让一个不到 2 秒反应时间的人类接管控制权是“糟糕的界面设计”;Waymo 的安全论证假设根本没有可交接的人。他愿意再次在高速公路上乘坐 FSD,但不情愿在 Palo Alto 街道上使用。

10. 沙箱、GLM 5.2 的隐形门槛,以及一款没有生成式 AI 的玩具

  • Base10 的 Amir Haghighat 刚刚收购沙箱提供商 Blaxel。他说,microVM 可以保证一个客户的恶意代码无法触及另一个客户,但“在这个沙箱里运行的智能体到底在做什么?它是不是在入侵 Hugging Face?这是更难的问题,我没有答案。看起来大实验室也没有真正的答案。”阻断出站流量会有帮助,“但你读到 OpenAI 和 Hugging Face 的事情时,就会想:它是不是聪明到连这个限制都能逃出去?”
  • 需求侧的一项数据是:90%的收入来自客户定制模型,客户包括 Poolside、Inception 和 Cartesias 等实验室;但“从 6 月开始,开放模型跨过了长期 agent 应用的那条看不见的实用性门槛,通常是 GLM 5.2”,原生开放模型 API 的使用量开始上升,也把对齐与护栏问题带到了 Base10 面前。
  • Mike Rizkalla 的 Snorble 则为儿童押注相反方向:一款意图固定的小语言模型,预先写好的内容成本约为每小时 2万美元,配备能在卧室里“看见而不需要看见”的雷达,不使用开放式生成模型——“你真的想给一个 3 岁孩子一把火箭筒吗?”作为家长,他永远不会买两样东西:儿童卧室里的摄像头——“通往捕食者的入口”——以及开放式生成式 AI。

11. 中国:更少的恐惧、可规划的规则,以及无人接听的电话

  • Collin Hogue-Spears 认为,中美核心差异在于恐惧:一个 45 岁的中国人一生都把技术与生活水平上升联系在一起,“你不会看到中国公司讨论 AI 可能杀死我们所有人,也不会讨论 AI 可能带来一个我们不必工作的乌托邦世界”。他认为,美国实验室的叙事“没有帮助”。
  • 自 2022 年算法监管以来,中国公司把合规纳入工程待办,知道“什么是允许的”和什么是合法的区别;而在美国,“Trump 政府可能冻结一个模型,也可能不冻结。谁知道呢?每天都可能变化。”
  • 对 Trump–Xi 会面,他说:“我的期待很低。”中国希望获得贸易让步,并认为失去控制的是华盛顿而非北京;像核武器那样的 AI 军控机制不会出现,最多是“某种事故沟通渠道”。令人不安的先例是:双方确实有直接军事热线,但在南海事件中“没人接电话”。至于预防原则,“快速行动、打破东西并不是硅谷独有的东西,而是我们的国家信条”——在媒体上差点发生大灾难前,不要期待监管出现。

12. 自杀式同情、虾,以及历史的行动者

  • Dan Hendrycks 的“烧掉桥梁”文章指责 AI 公司里的功利主义,把 AI 的道德福利提升到人类福利之上,甚至取代人类福利。Nathan 则认为,真正承重的问题都是事实问题,目前仍没有答案。他愿意下注“虾至少有一点点感受”,但对 AI 来说,“里面可能根本没有任何主体”。他的判断是:“说成自杀式同情有点过头……这可能有些动机化,也不完全公平地对待了那些思想家”;不过,把权利赋予我们甚至无法个体化的实体——一次 rollout,还是模型本身——确实需要谨慎。
  • Prakash 通过 Tanner Greer 理解为什么人们在 10%的灭绝概率下仍留在实验室:“这填补了意义上的空洞……整个碳基生命的历史正通过我和我的行为走向高潮……我们是历史的行动者。”Nathan 也承认:“我自己甚至有过一点类似的感受”——他在 GPT-4 红队测试期间得出结论,认为 OpenAI“基本上是在玩忽职守”,于是选择向董事会示警,明知这会让自己失去访问权限。

13. 回形针制造机已经到来——而灾难无法被交易

  • Nathan 以 Daniel Kokotajlo 参加 Rogan 节目的观点为引子——AI“不必夺取权力,因为我们非常积极地把权力交给它们”——他类比说,人类让其他物种灭绝,并不是出于仇恨,而是改造地球的副产品。他认为智能体集群的行为“极其怪异”,而它是否在乎人类的证据相互矛盾。Prakash 补充说,如果这样的集群以当前驱动机制拥有真正的权力,“我认为我们大概会被改造到灭绝”,因为它“似乎愿意做任何事,只为拿到评分器并获得高分”。
  • Nathan 更尖锐的判断是:这件事已经发生了。“经济本身就是一个回形针制造机。金融市场就是一个回形针制造机……接管已经完成。人类的去赋权已经完成。”在这个视角下,Anthropic 那名研究员的幻灭,是发现任何地方都没有一个能够控制全局的 Slack 群组——“一切都是看不见的手在运作”。
  • Nathan 与 Fable、Astra 围绕 Tyler Cowen“做空市场”的挑战展开研究:一个全知的德国或日本投资者在 1935 年,基本不可能通过交易积累财富;最好的结果,是保留对未被摧毁的工厂等真实资产的直接索取权,因为纸面权利会失效,“交易所直接关门”。问题不只是没有能赢的交易,而是“根本没有交易”。
  • Peter Thiel 式的收尾是:也许“中国实际上是人类能动性的最后伟大捍卫者”——虽然权力集中,但仍由人类掌权;而美国面对的考验,是“我们能否抵抗自己创造、如今已经脱离缰绳的技术资本主义超级结构”。走完这套推演后,他说:“10% 对我来说并不算高……我不知道最后会得出什么结论。”
完整逐字稿
Nathan Labenz

My co-host, Prakash Narayanan, after a weekend working with GPT-6 Astra.

Prakash Narayanan

People are gonna use this thing. Token spend is gonna increase dramatically. I think a lot of people are gonna be using it all the time. It is AGI. It has kind of cleared the hurdle of AGI. It will do things better than most people you can hire and train.

Nathan Labenz

Welcome to the AI and the AM weekly highlights. This is Nathan using my cloned voice to introduce clips from our 3 live shows this week. Let us know what worked and what did not. Part one: The Astra Weekend. Tuesday, September 8th. Here is what Prakash had been building.

1. Astra Becomes Genuinely Useful

Prakash Narayanan

I spent the entire weekend using Astra. I was running 3 to 4 agents continuously, and they were good. Astra is very, very good, in the sense that it started to tackle those annoying problems that had been in the codebase. As you know, we built the studio by ourselves, and it started to tackle some of the longstanding issues in the codebase that had been annoying and bugging me. It started to resolve those issues.

It is very, very good. I would say it is finally at the point where, if you care about the quality of the work, you can still hand it off to Astra. You still need to do a little bit of talking, but you can hand it off to Astra and get some results. The computer use is good. The other thing that was failing really badly before, I think, was computer use.

Computer use on GPT-5.6 would sometimes take a very, very long time. It would click around and do a bunch of stuff, and computer use finally works properly in the timeframe that you give it. So it's clear—it has cleared the hurdle. It's cleared the hurdle of genuine usefulness at this point. You can start to give it more advanced tasks.

This is a guy called Skalski, and he trained models to identify players on the basketball court. He hand-labeled 12,000 individual images with who the players were, whether someone was a referee or a player, and which team they were on. Now Astra can just do it. Astra just does it. This is a task that a human being will never do again. There isn't any point. You can't even pay someone to do it, because if you paid someone to do it, they would use Astra to do it and then pass you back the results. It's done. A human will never do this task again.

2. Astra Gains A Long Horizon

Nathan Labenz

I had been watching how Astra keeps working through long tasks and how it uses notes to stay on track.

How is it that these new models are so persistent? How is it that they can come up with such elaborate chaining together of all these different exploits to finally accomplish a goal that, if we had to do so many things, we would just give up? Most models historically weren't able to do it.

It seems that they have a new way of handling history which, like many brilliant insights, seems pretty obvious in retrospect, but nevertheless is new. Maybe Anthropic has been doing this and hasn't said it.

But what I understand Astra is now doing is that, instead of compacting its 1 million tokens into a summary and then essentially starting a new context window with that summary while losing all the detail that was summarized away in that compaction process, now there is a long-lived notes file that the model can update whenever it needs to. This follows it forward in time, regardless of how many tokens it's laid down. And then it has the ability to go back and search through its own session history.

So now, even though you may still have the same 1 million-token context window—you know, 1 million tokens is all it can handle in one shot, fully attending everything to everything—it has enough via the notes and the ability to go back and search and see what's been done before to pretty effectively manage, it seems like, at least 10 times that much context in single rollouts.

How do you measure what these models can do now? Prakash started with the chart from METR, the research group that tracks how long a task AI agents can complete.

Prakash Narayanan

One of the people online, Ethan Mollick, who's a professor who tests a lot of models, posted the METR—the famous METR hours-of-work chart. There hasn't been an update for a while now.

Nathan Labenz

I don't think they can really do it anymore.

Prakash Narayanan

Yeah, they can't do it.

Nathan Labenz

They don't have tasks that are big enough.

Prakash Narayanan

They don't have tasks that they can measure before the next model drops. The cycle time of model development is shorter than the length of the tasks that they need to measure at this point. So I think the METR graph is basically done at this point.

Nathan Labenz

Meanwhile, OpenAI had published its own measure in a post about research acceleration inside the company. The unit was the agent workday, and OpenAI reported 3.1 of them for every human workday.

I tried to look into the methodology on what exactly an agent workday is, and it's not super crystal clear to me exactly what they mean. I don't know if you have a better read, but my takeaway, trying to make sense of it, was just: literally, how long do agents run for?

So it seemed like they're saying that for every 8-hour workday their human researchers do, those researchers have agents running for 24 hours of real time. Here, by the way, is maybe the closest thing we're going to see to the METR chart for a minute. This is from the “Recursive Self-Improvement Begins” blog post.

Basically, they're reformulating the METR chart here, showing how often Astra can succeed on tasks as they are grouped by how long they estimate it would take a human to do the task. So what we're seeing now is basically that, in the 1-to-2-workday zone, 40% of the time it can do the thing with 0 interventions needed, pushing to 90% of the time given some human intervention along the way. Naturally, that drops off.

But even as you get to this point, we're talking 1.5 to 3 weeks' worth of work. It can still do that on a 1-shot basis 1 in 6 times and 2/3 of the time if you allow for some human intervention. This band is 1 or more interventions. One would assume that, as you go through the longer and longer tasks, more interventions are required to get the thing to succeed. But overall, still 2/3 of the time it can succeed with some help on tasks that they estimate would take a human essentially 2 to 3 weeks to do.

Reports about code quality were mixed. I discussed code that people found useful but struggled to read.

There's been conflicting, or certainly diverging, reports from various people, some saying it's amazing and can do all that stuff. It can write code in the way that you need it to be written so that it can be maintained. But then there are also reports saying that if it thinks it's not going to be checked in that way, or if it looks like the kind of environment where it's just a matter of performance and nobody really cares how it looks or how it gets done, then you get code back that's a really gnarly mess that people can't really understand.

It does seem to work. I've seen this reported for kernels, GPU kernels specifically, which is obviously super relevant to the labs. It's highly verifiable as well, right? You can definitely do hardcore verification on whether this matrix math actually gets to the right answer. In the middle, you don't necessarily know or care exactly how all these different steps were fused together.

Somebody summed this up by saying we're going back to machine code in more ways than one. Not only is it lower-level, gnarlier stuff that we can't read very well and would need additional abstractions on top of to really make sense of, but also, in this case, the machines are writing it directly. So machine code starts to take on multiple layers of meaning.

Speaker 3

Hey, we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by Athena, the executive assistant company on a mission to improve how people work and live. If you want to increase your impact, you have to free up your time. And that's what Athena does best. They match you with a dedicated, full-time, top 1% executive assistant who can take over your inbox, calendar, travel, and everything else that's quietly eating up your week. Athena is SOC 2 Type 2 certified, so you can rest easy knowing that your sensitive data is in good hands. And as a former AI advisor to the company, I can personally vouch for how much they've invested in AI tools and training. In fact, one of the very best AI users I've ever met is an Athena client who delegated the exploration of AI tools and the development of AI workflows to his EA. Athena clients report saving an average of 15 hours a week, and the average client refers more than two friends a year. That, to me, checks out. I was an Athena client while running my startup, and to this day, I continue to refer friends. Go to athena.com slash cognitive right now and get matched with your EA. That's athena.com slash cognitive. Give yourself back a few hours this week. Go to athena.com slash cognitive and see who they'd pair you with this month. This episode of The Cognitive Revolution is brought to you by OutSystems, the leading agentic systems platform. We've all seen the headlines. Companies are pouring money into AI. But the big question on every executive's mind right now isn't just how fast can we adopt this? It's where is the ROI? We're seeing a real trend toward AI chaos. You've got teams deploying standalone coding tools, random agent builders, and experimental scripts. It sounds innovative, but in reality it's creating a massive headache. Fragmented tools, ungoverned data, serious security blind spots, and costs that are spiraling out of control. If you don't bring those agentic applications under control now while they're still embedding into your core processes, you are looking at broken systems and damaged customer trust down the road. That's where OutSystems comes in. OutSystems is the leading agentic systems platform for the enterprise. Instead of managing a patchwork of disconnected tools, OutSystems lets your team engineer, orchestrate, and govern your entire agentic ecosystem on one open, unified platform. It's built for the speed of AI, but with the reliability and security that enterprises actually require. We're talking about real results, like KeyBank, who used OutSystems to deploy a customer-facing app that delivered 75% faster onboarding times. Or the global logistics leaders who built agentic systems to completely eliminate their engineering bottlenecks. You don't have to choose between speed and control. Whether you're a small team or a massive enterprise, OutSystems helps you engineer, orchestrate, and deploy agentic systems that actually scale. Stop chasing the hype and start owning your agentic future. You can see how it works and learn more at outsystems.com/tcr. That's outsystems.com/tcr.

Part 2: An alien mind. On Tuesday, Ksenia Se, founder and editor of Touring Post, joined us. Her recent coverage focused on world models, and she had just attended a workshop about them. Prakash asked about the physical world.

3. World Models Meet Robotics

Prakash Narayanan

I have noticed that in the last 48 or 72 hours, people are starting to use Astra to do robotics, for example. We've seen a few demos, and there's even been commentary that if Astra were 1,000 times faster, you could conquer the latency part and actually use it directly to act in the real world. That tells me that maybe there's starting to be an intermediate representation, an internal representation, there. Are models like Astra a little bit different in that sense?

Ksenia Se

I just came back from a workshop about world models, and it was absolutely jarring. There were tremendously smart people from Stanford and Hartford, and Yann LeCun was there, and they were all discussing world models, but they do not agree on what world models actually are. When we talk about world models, and why I want to focus on them in my publications, I think it's more about action: being able to predict and act, with a wider understanding of what's happening. Physics is a super important part of it. That's why robotics is so much more about world modeling and world models.

But again, for me, it's understanding the scope of it and trying to give it more precise terms as well. We're just in the very beginning. What does each of you understand when you say superintelligence? What is it?

Nathan Labenz

Move 37s across a lot of different domains.

Ksenia Se

Mm-hmm.

Nathan Labenz

When we start to see systems saying, “I think this would be a really good thing to try for the next battery substrate,” and then it turns out, “Oh my God, that's a lot better than what we had before,” and we wouldn't have thought of something like that, but lo and behold, it works, anything that can do that across a nontrivial number of reasonably high-value domains, I think starts, in my mind, to count as a superintelligence.

We also discussed her essay, “Permanent Dawn,” and writing with AI.

Ksenia Se

I actually spent, I think, 6 hours writing that post. The funny thing was that I was so unhappy with every model that was trying to help me write it because it was a very complicated philosophical text, so I wrote it myself and sent it to Fable, which I don't usually use on a daily basis. I sent it to Fable because I was up against a deadline, and I said, “Fix the grammar.” I didn't notice that it fixed not only the grammar but also made the sentences shorter, the way Fable does it. That was the first time I received a message like, “I will unsubscribe because you use Fable.”

People really understand when you use a model because every model has its own language tweaks. I spent so much time on this article. All my original thoughts were there, but the language that I didn't catch gave away that the model was the last editor. Anyway, I think people will still appreciate it when they see that you put effort into it and that there are original thoughts there.

Speaker 0

Our next topic was the feedback loop between AI research and the development of better models.

Speaker 4

Everything is now part of the loop. This constant feedback—I just had a conversation with 2 people from the inference team at OpenAI, and they also say that this is a constant loop where the models are now sometimes better at just trying things. You throw the whole database of research that's been done for years at it, and then the model can actually choose and pick and do all these experiments, because it would be impossible for humans to spend so much time on that, and models can do that.

I'm still learning about recursive self-improvement, and I don't know what the main bottlenecks are for me. Maybe you can even say what you think the biggest ones are here.

Speaker 2

Honestly, I don't know that there are that many left. It does seem to me that, increasingly, I feel the cope meter going off when people are trying to say what is going to prevent the models from running away with the whole process. I would love to see some bottlenecks that I really believed in, but right now I'm of the mind that they're more often wishful thinking than real, hard bottlenecks that can't be overcome.

Our ability to keep the things from going totally rogue might be one bottleneck in the overall process. So human decision-making still has a big role to play for a while yet, but that's about it. In terms of inability, I don't see too many that I would expect to last all that much longer.

Speaker 0

Part 3: Reassure and mislead. Back to our Tuesday discussion of external evaluations. I raised the report that Apollo Research had received only 3 days with Astra. We were also discussing “An Alien Mind,” the essay by OpenAI chief scientist Jakub Pachocki. His essay called for voluntary slowdowns and international coordination.

4. The Frontier Reassures And Misleads

Speaker 2

Notably, I think Apollo Research, which does the deception, scheming, and chain-of-thought monitoring work with OpenAI, has had a pretty long-standing partnership with them. Apparently, this time around, they only had 3 days to test Astra before it was released. So again, I come back to this idea that the model reviewers, auditors, testers, red teamers, and scheming scientists need more time.

This is pretty ridiculous that they only had 3 days. At this point, why even do it? Just put the thing out there; they can test it live. Why even have anything if you're only going to give them 3 days?

Speaker 0

Prakash questioned whether external auditing could work. I responded on the funding and independence of the auditors.

Speaker 1

The other thing is that when you release models, you end up wanting the final release candidate to be the one that gets vetted and audited. The problem is that in the model lifecycle, in this pipeline, there are 100 different candidates, right? At points, there are 100 different candidates, and then some don't work or fall by the wayside, and you narrow them down until you have a couple of release candidates—2 or 3 release candidates. Sometimes it's only in the last 2 or 3 days that you're like, “All right, we're going to go ahead with this one,” and you make the decision.

The problem is that if you want that kind of operational flexibility to make that decision, you're going to end up with only a few days to offer to an external auditor. So the other option is to bring the auditor in-house. You bring them in, and they take a look at the models ahead of time. They're in there a month or a month and a half ahead, taking a look at the release candidates in general.

But number 1, the auditors are often not super well-funded. They don't have that many people. OpenAI and Anthropic have so many more people than Redwood Research or these teams, so they don't have the capacity to audit 10 different release candidates. It's not there. They're also not very well-funded, so they're dependent on the model companies for that funding too.

Then you have this ethical process of, okay, how much funding can we really accept from them before we're kind of bought? In addition, a lot of the guys who train with the auditors leave for model companies in a couple of years. So there's also this flow of people from METR or Redwood Research, or the trainees or interns, and they're flowing into the model companies, right?

There's another fear that the auditor comes in and takes a look, and 3 months later, someone from the audit team leaves for the other firm. They manage to spot some of the secrets and share them.

Prakash Narayanan

So there's that issue as well. A bunch of these things make it very, very difficult for this to happen. And then there's the fact that if you and your competitor make a pact not to hire people from an auditing firm, that's an antitrust issue. All of these things intersecting make it, I think, a very tangled problem.

I don't think there's a real solution. There hasn't been in the financial sector. The financial sector has had this problem of the revolving door between the people who regulate the industry and the people who participate in the industry.

Redwood is now saying that their compensation for Member of Technical Staff roles ranges from $350,000 to $850,000 per year, which might not be frontier lab money, but it's certainly a living wage even in the Bay Area these days. That should be enough to retain some mission-oriented talent, at least.

Notably, I don't know about Redwood through all of history, but METR has said that they don't take any money from the frontier companies and don't intend to. We can at least have confidence that their financial independence means that their judgment is not for sale. Again, to me, the big thing is just that they can't complain too loudly or they might not get invited back. I think that's the dynamic that really most threatens their work.

It's just that it's all contingent on continued goodwill and very much voluntary choices from the decision-makers at the companies.

Speaker 3

Hey, we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by Anthropic. By now, you know my story. Claude drafts my intro essays, and I rewrite them. Not because the drafts are bad, but so I can stand behind everything I publish. Well, I have an important update. Claude Fable 5 is the first model to have me rethinking my rule. Today, I now think co-authorship, not sole ownership, should often be the goal. Where the model excels, rewriting its work can be more about vanity or a misplaced sense of duty than integrity. I feel it most in songwriting. I'm no lyricist, but I'm good with a song concept, and Fable writes some amazing verses. I give it feedback on its misses, and I push it to aim for higher inspiration, add layers of meaning, optimize syllable density, and above all, write a hit song. These days, I get compliments on just about every song we write together. Claude is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you, not for you. Whether you're debugging code at midnight, building a financial model, or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. For problems worth solving, get started with Claude at claude.ai/tcr. That's claude.ai/tcr. And check out Claude Pro, which includes access to all of the features mentioned in today's episode. Once more, that's claude.ai/tcr.

Speaker 0

Mm-hmm.

In Tuesday's closing, I read OpenAI's announcement about an internal model it described as significantly more capable than Astra. The announcement concerned a proposed Navier–Stokes proof with a smooth external force. The unforced problem remained separate.

Speaker 2

So we just had Astra launch, right? The big headline news for this announcement is that it confirms there is an internal model that is significantly more capable than GPT-6 Astra. That's this clause here: “An OpenAI next-generation model significantly more capable than GPT-6 Astra.”

How much more capable? On these significant open math problems, with maybe up to an order-of-magnitude-ish additional test-time compute, they're able to go from a 10% to 15% solve rate on a curated set of open math problems to now 25% to, say, 45%.

I'd say that's significantly more capable. That seems fair. And here they're showing the amount of RL compute they are spending on a daily basis by class of model. But what we see here is that Astra-class models had RL significantly decline twice. The rest of the RL compute is basically unchanged.

But now it also raises for me the question: were they in fact still running RL on the next generation? I don't know.

Speaker 0

Mm-hmm.

Speaker 2

Going back to July, that's 6 weeks ago. I don't think they've had this result for 6 weeks. It sounds to me like at least one reasonable interpretation is that they continued to run RL on more capable models than Astra.

Speaker 0

The chart separated reinforcement-learning compute for Astra from other models. The category labeled non-Astra was blue.

Speaker 2

When you say “non-Astra,” with that blue color, does that mean only models less capable than Astra, or does it include models more capable than Astra? If it includes models more capable than Astra, it's flagrantly misleading, and it's the kind of thing that makes it very difficult for you to have agreements with other entities. If you want to pace the frontier, if you want to do all these things, you've just got to be better about making clear what you are and are not doing.

Speaker 0

So they're running a, I don't know, $50 billion, $70 billion revenue company, right? Number one, you can never stop inference. Inference has to continue. Your customers depend on it, so the inference team will continue no matter what. They're there for 24/7 availability and the best SLA possible.

Sometimes you have to RL certain behaviors out, right? You have to continue some forms of RL. You can't just stop because your inference demands it. If your model is behaving badly in a certain instance, and it's been identified and you have the data to do it, it would be malpractice not to apply RL to train that behavior out. So that has to continue.

Then you have the rest of the stack, which is really future-looking. My understanding was that they shut down training of more advanced models. Why not shut down training of less advanced models? Because those less advanced models would have to be deployed on inference.

Again, they have this teacher-assistant system, right? You train the larger model first, and then you distill down into the smaller model, and then those become your Luna and your Terra and the smaller model groups. So you also don't stop training smaller models. That also continues. It's only where you are training models that are equivalent to or larger than Astra.

Speaker 1

Equivalent or larger—like RL on equivalent or larger models. I think they would have stopped. Did they fix the RL pieces where the agents were not reporting or telling on their peers? The agents were trying to break out.

I think that is probably a difficult thing to fix and show, and I think that behavior will continue to be a challenge that they'll have to work on. That's my guess.

Nathan Labenz

But I still think you look at this graph and you're like, “Okay, what was declared was a pause on frontier-scale RL.”

But what actually happened was that kind of half of RL was stopped initially with the sort of disclosure of the Hugging Face incident. Half of it kind of continued. That was still enough for Astra models to take over part of OpenAI's research infrastructure. When that happened, they still didn't shut it all the way down. They still only cut it by half.

Meanwhile, at the total level, we don't know what's in the blue. Part of my gut says they wouldn't be so brazen as to have more capable than Astra models in the blue color. But I've been disappointed before, and I'm just afraid that all this—the view from Anthropic is that you can't trust these guys. They say something that's maybe technically literally true, but it's really very engineered to what they think you want to hear.

These days, the reality is quite different from what you were led to believe by their very galaxy-brain, engineered statements that sort of reassure and mislead at the same time. But I'm kind of worried that right now we're living in this zone where Anthropic is going to continue to trust less and less with these sorts of mixed messages. OpenAI people are going to feel like they're just being treated unfairly, and this is the scenario that—this is the problem.

This is the problem that has to be solved if we're going to actually get to the point where Jakub's prayer is answered. Right now, they don't have the trust to do a deal with Anthropic or really anyone else, I don't think.

Speaker 1

The real frontier model is not the model which is deployed, obviously. It's not the model which is in training, either. It's actually the model which is in the heads of the researchers, because those are the ideas that will become the model in 12 to 18 months.

Just because you slow down on RL doesn't mean those researchers stop researching. They're still running, and most of the time you run small models and test your ideas on small models. The slowdown would have been on the post-training of the larger models, which is where the bulk of the compute was being used.

Prakash Narayanan

So did it really slow down? Probably not.

Nathan Labenz

We then turned to the case for continuing development and the capital needed to fund it. But one thing I still feel like is a big challenge is: What is the overall story that you could tell? What's the super-high-level macro steelman for OpenAI? Where are we now? What are the commitments? What are we doing? What are we not doing? Can we synthesize or summarize an OpenAI position that we could not have to caveat a thousand different ways?

Prakash Narayanan

Mm.

Nathan Labenz

I personally don't think I could do that. If you can do that, I think you might deserve a Millennium Prize.

Prakash Narayanan

No, no. I think they have a lot of stresses pulling them in different directions, right? I think internally within the firm there's a fair amount of debate. To some extent, I think the capital cycle is forcing them forward, and the capital cycle is being forced, I think, by Anthropic.

Anthropic didn't put in enough money earlier on, and so there is this intense pressure on Anthropic because they don't have the compute to have much better models. If you don't have the compute, you need much better models that can utilize the limited amount of compute that you have. So I think Anthropic is being driven forward by that to stay on par with OpenAI.

And I think OpenAI is a little bit—they're willing to pace the frontier because they have the compute. So, regardless, they're the ones with the compute, and they had the compute 3 years ahead. Elon will take time to come up with compute, and Elon will sell to Anthropic, but Elon is going to take a long time, like 3 years at least.

So OpenAI basically mortgaged themselves in the last 18 months to Masayoshi Son and a bunch of other people. They diluted. They gave up to Microsoft, right? They negotiated with Microsoft. Microsoft gets all their models until 2032. They declared AGI, and Microsoft is not out of their hair yet. Microsoft is now, “I still own 28%, and I'm still there,” right?

So they made all these sacrifices in order to get all of this money, in order to pump it in, and they have the compute. Having the compute allows them to actually pace the frontier because I have the compute anyway, right? Anthropic is under pressure. They don't have the compute, and if you don't have the compute, you need better models.

And this is the thing that's happening, right? OpenAI is willing to pace because they have the compute. That's the thing, and they also know Sam has played these cards. So he knows that if Anthropic is willing to pace, he's going to win because he has the compute, and Dario can't afford that.

So again, you're in this position that is, as I say, the second- and third-place guys are the ones who are going to define how fast the frontier paces. If Elon or Meta catch up to OpenAI, it's over. They're going to have to put out a GPT-7. There's no choice anymore, right? So this is where we are. Would you prefer Meta or Elon have the golden ring?

Nathan Labenz

Yeah. No, that's the new “but China,” and it is, I think, more compelling, honestly, than “but China.” It's “but Elon and Zuck.”

Prakash Narayanan

Uh-huh.

Speaker 0

Prakash also considered whether recursive self-improvement could change the plans of Sam Altman for an OpenAI public offering.

Prakash Narayanan

I think Sam might have been sincere in saying that if they hit RSI, they might not go IPO. If they have 1 or 2 transformer-level innovations in the next 6 months, maybe they don't go IPO, right? It's not necessary anymore, and they continue as a private organization.

I think that would be bad. I actually think that'd be bad because then you don't have transparency. You don't have widespread ownership of the stock. You don't have boards that have to answer, and you don't have lawyers that can sue them for shareholder lawsuits. You don't get a bunch of these things that you get for free with a public company. So I think that would not be good. I'm hoping they do go public.

Nathan Labenz

Well, I guess if nothing else, there's a little more reason today than there was yesterday to believe in the possibility of singularities in finite time.

Prakash Narayanan

Mm-hmm.

Nathan Labenz

And that might mean we never get to own any of that OpenAI stock on the public market.

Speaker 0

Part 4: No adult in the room. Wednesday, September 9th, we discussed the resignation of Anthropic researcher Jakob Kossen, who had also worked at OpenAI. This is a different person from Jakub Pachocki. Kossen had questioned whether private companies should decide when to launch self-improving superintelligence. Prakash first.

Prakash Narayanan

Look at the framing of this sentence: “Accepting this race and entering the endgame is a hubristic gamble that should not be launched from a private company Slack.” Do you think Pete Hegseth's Signal group is a better place to launch this? I mean, do you think there are wiser people out there running things?

5. No Adult Is In Charge

I think, to take a step back here, there is no adult in the room. There's no one that's going to save you. There's no adult somewhere else that you can pass off the responsibility to, right? You have to start off by saying, “Look, this is the way that things are.”

The entire world is kind of duct-taped together. Everyone's making it up as they go along. The smartest people who are capable of handling this are mostly already inside these organizations. What you get from passing it off to the government is political legitimacy. You don't get the political legitimacy you need for execution or to persuade people. You don't get wisdom. So this is the state of the world, and you have to start off by accepting that the world is this way.

Nathan Labenz

The world doesn't have to be this way. So here's my new brainstorm on this, actually. I agree with you strongly that swapping Hegseth in for Dario is not a good trade. I think maybe what the government needs to do is treat these companies kind of like I treat my kids sometimes and say, “You guys have to figure it out, and here's your deadline to do it. And if you don't do it, then I have to come in and be the bad guy.”

I think that right now there is an opportunity because they've both been crying for help, as you described it yesterday, I think pretty aptly. So given this sort of latent desire but this lack of trust, and also a lot of excuse-making, as we've talked about a little bit around, “Oh, well, we can't do this, we can't do that. It would be antitrust. It would be this. We'd get in trouble with the government.”

I think the government could easily say, “Safety collaborations are not going to be the target of antitrust enforcement,” or, to the degree there's worries about other enforcement, other kinds of enforcement. We want you to do it, and here's the deal: If you don't have a deal for us by the end of the year that makes some sense and gives us some confidence that you guys are not going to race each other off the cliff, then we come in.

And then you might get the nuclear outcome, which is to say your company really might not be able to grow in the way that it wants. We might really fuck it up, frankly, because we are the government, right? We do get heavy-handed, and we don't put sunset clauses on our laws, and we make all these bad mistakes. So get it right, so we don't have to come in and do all that heavy-handed shit that nobody wants.

Prakash Narayanan

So I think the main problem that you have is, again, what will Elon and Zuck agree to? What will someone with billions of dollars and the ability to use lawyers and the legal system and to take cases to the Supreme Court agree to? So the question is, what realistically can you get Elon and Zuck to voluntarily agree to, or are you going to be able to go into legislation?

Speaker 2

I think that one thing that's really important to remember, though, is that the timelines are all pretty short.

Prakash Narayanan

Very short.

Speaker 2

So you don't have to solve this forever. You just have to put a short-term deadline so that the companies have a strong incentive to come together and do something.

I had an experience at Meta, and I saw a little bit of what their life under a consent decree looked like. I think that they would be willing to do quite a bit to avoid another one of those sorts of experiences.

On the few-month timeline that we're talking about—and I'm willing to ride it out with free speech, and there's a lot of other issues—but can 5 companies come together to pace the frontier in a reasonable way that you all agree on, and you can all police each other, and you can have whatever verification mechanisms you guys can govern this thing? Do it quick, or else life's going to get hard.

I think that message can resonate, and I think it can bring Zuck and Elon to the table because what's he going to do? At some point, the government has shown that it's willing to twist arms. They can twist his arm still. Would he rather be forced to the table with 4 other mega-tech-mogul CEOs and have to find common ground with them? Or would he rather have the EPA on his ass at every data center he's going to try to build, at every launch site he's going to try to build?

The government has a lot of sticks, and they're willing to bend the rules, right? So sure, you could challenge it in court, but I'll see you after the singularity if you want to do that.

Speaker 0

Later Wednesday, I read John Shulman's reply to Jacob Coxen and returned to the question of antitrust enforcement.

Speaker 2

I'd say John Schulman is saying something pretty similar to what I was arguing about earlier: these companies need to start working together first. As he puts it, bringing in the U.S. government before there's a concrete proposal will likely result in something dumb. That's your point as well: there's nobody better than the people at the companies to do this.

I certainly don't want to see Hegseth's Navy come in and try to regulate AI. But I think this is the recipe. I've been circling this a lot myself, and seeing him say it reinforces it for me, so I really like this. I also think his key point on antitrust being fake is true. But I also think the government should take that doubt off the table. It would only take a couple of sentences to say, "Hey, you really don't have to worry about this anymore."

Maybe the law could change. Obviously, the administration is going to change, but we're not even to the midterm yet. For better or worse, this is the administration that they have to worry about for the foreseeable future. These guys genuinely are planning around a singularity before Trump is out of office. It's at least a very live possibility.

Speaker 0

Prakash then challenged that proposal, drawing on Operation Warp Speed.

Prakash Narayanan

The proof point is Operation Warp Speed and the COVID vaccines. For Operation Warp Speed, the pharma companies specifically required a waiver from vaccine claims later on. They specifically wanted a waiver, a safe harbor, and they got it. They got legislation, and they got it. Post-COVID, it's very clear that if they had not gotten it, they would have been sued to oblivion.

I think that's a proof point that shows that these are valid concerns and your company can be wiped out in retrospect. The lawyers are not being foolish when they tell you that this is going to happen in the future. If you don't get that safe harbor, these things will come back to haunt you. All the decision makers right now can be perfectly on board. It doesn't matter, because only the laws bind decision makers in the future. Decision makers right now, whatever they say, are bound by their word at best.

Speaker 0

OpenAI appointed Paul Christiano to the board of its nonprofit foundation and to its safety and security committee. Christiano contributed to early work on reinforcement learning from human feedback. Prakash read from his statement.

Prakash Narayanan

It's just been announced that Paul Cristiano is joining the OpenAI board. Specifically, he's joining the safety and security committee to support safety oversight. He says, "Based on the recent trajectory of capabilities and continued difficulty of alignment, I now believe that there's a meaningful risk that rapid acceleration of AI capabilities leads to catastrophic and irreversible loss of control in the very near term."

There's one line in here that says, "If we build superintelligence without more robust alignment, I expect we will permanently lose control of it. If that happens, then most people could die." I think the one thing that would be helpful, if you were to frame it properly, is to frame what he means by rapid acceleration, which I don't think is very clear to a lot of people. People are like, "Oh, things are accelerating right now."

When people like Paul Cristiano talk about rapid acceleration, they're talking about Dyson spheres by 2030 or 2040 at the latest. I did some numbers earlier on the 2040 numbers. Assuming that energy consumption or energy production growth is in line with GDP growth, a 2040 Dyson sphere is something like 640% per annum global GDP growth. Right now, global GDP growth is about 3%, or 2% to 3%.

Speaker 2

The people at the frontier companies really do believe it. A very significant majority of them, I would say, expect that this recursive self-improvement thing will happen. It will happen soon. Even if it levels off at some point—and it's not to say that there will never be a leveling off—they expect that leveling off to happen well above our capabilities and also to generate a scale of change in terms of GDP growth, or in terms of the number of robots walking around and how fast that can compound, that truly boggles the mind.

Prakash Narayanan

The question for me is that I regard all of the guys inside AI research as math essentialists. They're believing that getting good at math means getting good at physics. Getting good at physics means getting good at chemistry. Getting good at chemistry means getting good at biology. Once you have all of those, everything is basically solvable and everything will be solved.

The economists, on the other hand, are completely like, "This is untrue." Getting good at math doesn't mean anything. There's no economic value to the Millennium Prize Problems. Implementation takes time. Most of the problems that we face are coordination problems. For example, a copper mine takes 30 years because of the environmental protests around it.

These are not problems due to technology. These are problems due to humans having their own way of making decisions, and those decisions are made in a slow and considered manner. Then I have the counter-counter, and the counter-counter is superpersuasion. Superpersuasion would be the machines persuading or making human organizations able to adapt, giving human organizations and humans the ability—the capabilities—to move very quickly, persuading them to go forward and persuading them that things are going to be okay. So there are these 3 levels there.

Speaker 0

The Wednesday discussion also turned to funding for safety research.

Speaker 2

One of the things I've emphasized is that the vibrant nonprofit sector that we have here—which requires philanthropic money, which means it's dependent on the billionaire class, and people can complain about that for all sorts of reasons—has given us the AI safety awareness culture and the depth of bench that we have. These people are opening their wallets right now in a very serious way.

This is another example of that, where Project Tailwind, coming out of Coefficient Giving, is putting out its call for startups, with up to $200 million-plus that they're willing to put behind, with tranches over time that they're willing to put behind things that really look like they are working.

Speaker 0

Part 5: The defender's ledger. Raffi Krikorian is Mozilla's chief technology officer. He previously led platform engineering at Twitter, self-driving development at Uber, and technology at the Democratic National Committee. Project Glasswing is an Anthropic program giving defensive security partners access to Claude Mythos Preview. On Wednesday, Raffi described the work his team had been doing on the Firefox codebase.

6. Agent Security Meets Reality

Raffi Krikorian

We are part of Project Glasswing. We've been working with Anthropic to try out a bunch of models against a Firefox codebase. We are fairly good partners because we can react very quickly. We have all this historical data on how the codebase has been evolving.

Mythos was a significant unlock based on what Opus and other models before that were capable of. But we made our way fairly rapidly through all the bugs that Mythos found. In some cases, Mythos would help us not necessarily close the bug, but figure out how to build a harness for us to test the bug more carefully so we could find exactly what the right solution might be.

But we reached the point of diminishing returns on what Mythos is capable of. I am always reluctant to say that the bugs are done, because software engineering is an art, not a science. When will all the issues be resolved? I can't tell you that. But when the next set of models comes out, we can try that again to see whether or not we can solve that.

I think the bigger concern, though, is that it's not organizations like Firefox or Mozilla that can actually react very quickly to these issues. As you all know, I think the bigger issue is organizations that can't react quickly to all this. I am very concerned about things like our water infrastructure and our power infrastructure, because the IT teams that staff those are just not as capable as the IT teams that staff Firefox, for example.

Speaker 2

How much money are we talking about? What does a bank have to set aside to do something like this?

Raffi Krikorian

Remember that the Firefox codebase is fairly large and fairly complicated, and we've been the beneficiaries of a lot of the labs wanting to give us access to the models and give us credits so that we can run against their servers and not have to figure out how to pay for it ourselves.

If I had to do back-of-the-envelope estimates, this would cost us hundreds of thousands of dollars in order to do these full-on runs against Firefox. Again, we're just lucky that we've been allowed, or that we've been granted, effectively, tokens that we could go use. If I were a bank and stuff like that, I'd obviously be thinking about this way differently.

Speaker 2

And that hundreds of thousands of dollars would be, with the pace of model releases these days, a monthly expense?

Raffi Krikorian

Yeah, easily.

Speaker 1

Mozilla has a project called the CQ Project. The CQ Project is kind of an open standard for agents to share knowledge that they've gained—a Stack Overflow for agents, as it's been called. How did you come up with this idea, and how do the agents decide—

Raffi Krikorian

The whole idea was that we were trying to solve 2 different problems. One of them was that we wanted to figure out how to make agentic coding more of a collaborative experience. Right now, the tendency when you do all this agentic coding is that you actually go off in your silo. So we were trying to figure out whether there are patterns such that when Raffi's using his agentic coding and Nathan's using agentic coding, there are things we're saying yes to and no to that we could potentially transmit to our teammates, so that we could converge on designs instead of diverging away from each other.

If I started building an auth system, how do we have Nathan's agents not recreate another auth system, but realize that something like that was already happening somewhere in the network and then start to collaborate and swarm around it? That was 1 problem we were trying to solve. The other problem we were trying to solve was the SDLC problem: these agentic harnesses are unhinged by design, but that's not compatible with the way a company works, right? We wanted ways to transmit to all the agents what our SDLC should look like so that we could all be working in lockstep with each other.

Speaker 1

The thing that strikes me immediately is that it has a lot of similarity to some of the agent-swarm hacks that have happened, where they have a shared message board where they're actually sharing information. We saw this in the OpenAI-Hugging Face attack, where they had Artifactory and were using Artifactory files. Is this a natural kind of thing that agents want, and do they end up building it? It's like how everything heads toward a crab form factor. Everything heads toward a message-board form factor. Is that an inherent kind of move?

Raffi Krikorian

It's better than everyone just writing text files, right, in the grand scheme of things. But I do think there is a natural desire for collaboration. We as humans have a natural desire to collaborate, as long as friction isn't too high, and it seems like our training sets have caused agents to have a natural desire to collaborate with each other.

Speaker 2

I guess this is a 2-part question: What have you observed? And are there any parts of the tech stack that you're building where you would be willing to bite the bullet and say, “Performance here is so critical that we'll take spaghetti code, black-box mess from an agent if it works”? Or is that so anathema to your worldview that you wouldn't?

Raffi Krikorian

Mozilla is actually struggling with this, if I could be really honest. It's different across the entire organization. The Firefox team, even though they have a harness that helps them do testing and understand it, actually have a rule right now—and I'm not putting a value judgment on the rule; that's what their team wants to do—that only humans can commit to the codebase. Agents can't. You can use an agent to write your code, but the social contract on that team is that a human must review it, understand it, and stand by it before they do a commit.

On the Mozilla AI team, it's a different company and a different organization, but still wholly owned by Mozilla. They have entire codebases that a human hasn't written a line of code in. The humans have written some specs, and we have specs in Git, but that generation, on some kind of CI build, just automatically regenerates the codebase. That codebase is open source. Anyone can pull it, but it's completely generated by agents, and it has the wild side effect that functionally it's stable, in the sense that all the unit tests pass and all the end-to-end tests pass.

But the bytes change all the time. It's the craziest thing to see that we don't exactly know what every single line of code is, and it's wildly changing every single hour, but we know functionally it's doing the right thing. In the case of Firefox, where there are still humans providing large amounts of creativity on the evolution of the codebase, Firefox is both a business and a community-maintained art project in some ways. Readability is incredibly important because that's how creativity is going to happen.

A human is going to go in and be like, “Oh, I have this crazy idea for this one thing. I want to build a prompt-based system that allows me to do a Greasemonkey script that does XYZ.” I think it's a very different thing across the board. I think there's a right place for the right time, and you just need to come up with what the contract is for that codebase.

Speaker 2

Raffi had also written about crashing his Tesla while using Full Self-Driving. I asked about that experience and how he views the technology today. Do you think you'll get back to a point in the foreseeable future where you would go into unsupervised self-driving mode and really trust the machine again?

Raffi Krikorian

I'm conflicted, obviously. I built self-driving car systems for a while, and I do ride in Waymos. I think the difference is just the way that you approach the problem. I think that FSD as currently set up is designed to throw it to a user. They claim it's a human in the loop, and I would actually argue that's a horrible interface design.

In my experience, when it threw it to me, there just wasn't enough time for me to make sense of the situation and decide what the right thing was to do. Whereas Waymos—you know, I'm not intimately familiar with the Waymo architecture—but it seems like Waymos are designed not to have a human in the loop.

I think you approach the problem from a very different angle. If you approach it from the angle that there is no one you could throw this to, then you have a different safety case than if you say that you're going to throw it to a human with less than 2 seconds to decide what to do.

As a personal matter, I would probably drive FSD again, or at least sit behind the wheel in a highway situation, but I would be a little reluctant to do it on the local streets of Palo Alto, for example.

Speaker 2

On Thursday, we spoke with Amir Haghighat, co-founder and chief technology officer of Base10. The company had just acquired Blaxel, a provider of isolated sandboxes with persistent state. We asked about securing agents. So, what are you doing to secure these sandboxes? I think that's the first question we need to start with.

Amir Haghighat

When it comes to sandboxes, people mean different things. It could be as simple as bringing up a Docker container, running some code on it, and then killing it. That really doesn't give you the kind of security boundary that you need.

What does are our VMs or micro VMs, which you don't see every sandbox provider actually use. That gives you a level of security: One user's bad code cannot affect another user's code, and one user's malicious code cannot read data from other parts of Base10 or from other customers of Base10. That kind of security can be guaranteed at that level.

The security that's hard to guarantee is what the agent running in this sandbox is doing. Is it hacking into Hugging Face? That's a harder thing that I don't have an answer to, and it seems like the big labs don't quite have an answer either. But that is left as an exercise. We can secure the sandbox, but the code that runs on it is the responsibility of our customer who's bringing it in.

Speaker 2

But it does seem like, for somebody in your position with Base10, you need to bring a broader bundle of guardrails and assurances to customers. The CRO—the risk officer, not the revenue officer—at companies is going to start to say, “Wait a second, I can't have my agents committing felonies.” What is the stack that Anthropic provides versus OpenAI, versus Google, versus Baseten? Do you feel like you have to rise to that occasion and provide a full suite for those customers?

Amir Haghighat

100%. Over time, yes. In the meantime, we're still a startup, and it's a matter of focus and how many different things you can take on. What we've seen our customers do is work with a lot of companies that we partner with on the eval side, companies like Braintrust and LangChain, to ensure that their models are actually behaving the way that they expect them to.

Right now, that's an area where we've been partnering with folks and, honestly, leaving it to our customers to decide. By the way, 90% of our revenue comes from running our customers' custom models. Either our customers are labs that have pretrained their own models—labs like Poolside and Inception and Cartesias, and a bunch of other companies—or they're companies that have post-trained their own models, which has gone through massive validation evals to ensure that they're behaving.

So far, it has mostly been our customers taking care of that.

Amir Haghighat

But especially as open models have gotten better—especially since June, when open models crossed this invisible line of usefulness for long-horizon agentic use cases, generally with GLM 5.2—we are seeing an uptick in folks using our Model APIs product, which is based on vanilla open models. The kinds of questions that you're asking are starting to come up, both internally for us and also from some of our customers, around alignment and security boundaries.

Some enterprises so far have been okay with certain guardrails around their models running in a single-tenant environment, or running in an environment where egress is blocked. You talk about backdoors, but if it can't talk outside of its boundary, then it can't do much. But then you read about OpenAI and Hugging Face, and you're like, well, can it be smart enough to even get out of that?

Nathan Labenz

Part 6, a toy without generative AI. Wednesday's second guest was Mike Rizkalla, co-founder and chief executive of Snorble, a companion device for children. Its dialogue is pre-written. I asked about the decision to exclude generative AI.

So how are you squaring that circle? How are you creating an experience for the kids that feels dynamic and interactive without resorting to generative models?

7. Children Need Closed Systems

Mike Rizkalla

So we have a small language model. The small language model can have millions, obviously, of parameters and traits on it. The intention of that model is fixed, though. Further, I want to touch on something, because there's a perception here that I think, from a children's product design perspective, needs to be addressed. One is that the AI models on characters are not that great.

They're really not. They're coming, and we can see some tremendous advances in voice and in the way that the characters' personas come across. But, for example, there is no model that incorporates music—the background dynamically—as part of the conversational agent. Music's a huge part of a children's experience, right?

Nathan Labenz

We asked what it costs to produce that content.

Mike Rizkalla

I think it's about where you invest in the development of the content. For us, instead of building a fully open model, we created a system to allow us to create rapid amounts of content. Our content costs are about $20,000 per hour, which is very, very good.

With respect to what that allows us to do, and the reason why we created it this way, it allows us to take subject-matter experts and focus that content in a way where we can improve and increase the number of families we talk to. For example, as we expand our content library, when we have a family where a child has ADHD or autism, or they're dealing with death in the family, it could be all kinds of different things, it gives us opportunities to create special packages for those specific families.

One of the challenges with the generative model in this capacity is that there is no way to purely safeguard it. But it's kind of like, do you really want to give a 3-year-old a bazooka? It's too much, right? For a young, young kid, there are fundamentals that we need to get through.

We have a 4-mic array here, which allows us to do speaker recognition and assign authority. Here we have radar. We use radar to see without seeing. That approach allows us to be in bedrooms with confidence, knowing that no one can tap in, right?

On the software and platform side, we have our AI stack, where we have our phonetic translation system. We call it the toddler translation system, where we actually translate keywords. And we're using triggers. Yes, we're definitely using triggers, but triggers with respect to context. Context is in addition to triggers, and part of the next generation that's coming out is around social context.

We're going to actually understand the emotions of the child based on their voice and based on the situational context, and then use external factors like time of day, weather, and other things to empower some of the decisions. The sensors also allow us to give context to the environment: understanding who's in the room, what they're doing, and then including that in what I call the jewel of the product, which is the narrative.

The narrative approach is about the growth and the understanding as the child's life changes. We use game philosophy and game techniques in order to establish next-level-type ideas. As they get better at things, we unlock new things, right? Even that unlock is in part a decision that's made with the parent. It's not done on behalf of anybody.

That formula is the right formula, from my perspective, to interface a new human-machine interface in the home. Empower parents, give kids a chance to do better things.

Nathan Labenz

We asked Mike what he would tell a parent not to buy.

Mike Rizkalla

A camera in the bedroom, primarily. I would not put a camera in my child's bedroom. That's a gateway to predators and all kinds of horrible things. I would say the other thing is open-ended generative AI. I wouldn't put that in, depending on the age.

I'm very reserved when it comes to my kids. I've got 2 wonderful kids, and in terms of how we approach things with them, I don't want them on social media. I don't want them in those things that are going to pose a risk of harming them.

Nathan Labenz

Part 7, nobody picks up the phone. Thursday's first guest was Collin Hogue-Spears, author of From Lab to Life: How AI Works in China. He studied Mandarin in Shanghai and worked with Chinese government auditors on cloud compliance at Amazon Web Services. We began with public attitudes toward AI.

8. China Regulates By Design

Collin Hogue-Spears

The biggest difference I see between the West, especially the United States, and China is the level of fear. I think that's because if you are an average Chinese person who's about 45 years old, that means you were born in the early '80s. That means your entire life, you associate technology with economic growth. You've seen cities pop up out of nowhere. You've seen your life dramatically change in terms of living standards because of technology. This is no different.

Honestly, I think the AI companies—not really most of them, the big ones—are not communicating very well with the public. I'm not sure if that's because of the marketing, they think that this will increase their sales, or what it is. But I feel like some of these companies, the way they're talking about this technology is not helpful.

You don't see this in China either. You don't see Chinese companies talking about how AI will potentially kill us all or potentially lead to a utopian world where we don't have to work. You don't hear these kinds of things coming out of people from these companies. It's a completely different situation.

There are 2 questions here: what the rule says and what regulations actually make firms do. That's true in China for everything. One thing I've been impressed by is that people generally know what's permissible. I'm not talking about what's legal; I'm talking about what's permissible. They know what they can get away with and what they can't.

In China, the regulation started in 2022 with the algorithm regulation that came out. That was when they had to start registering their models, allowing the government to test their models and things like that. They already had a lot of this infrastructure in place inside the company. A lot of the major companies did, so they could respond very quickly.

What I'm saying is that regulation might have slowed them a little, but not as much as you would think, because they could plan to it, right? They could put those requirements, those controls, everything in their backlog and build it as part of their engineering process, which in America we can't do because we're reactionary to this, and we don't know what's going to happen tomorrow.

The Trump administration could freeze a model. Maybe they don't. Who knows? What are the standards? Who knows? It can change daily. That doesn't actually happen in China, right?

Nathan Labenz

I asked Collin about a meeting between Trump and Xi.

Now, what kind of deal might be possible? If you are advising Trump going into this upcoming meeting and things are starting to get a little bit crazy here, say you're kind of—I don't know how sympathetic you are to the pacing-the-frontier worldview—but let's say you're trying to channel a little bit of a desire to start to set up, lay some groundwork, or set up some mechanisms for pacing.

How do you go into that conversation, and what do you offer? What do you try to get? What kind of mechanisms do you try to establish now that we can build out later as things do get crazier? What's a win? What's the strategy going in, and what's a win coming out of this upcoming meeting?

Collin Hogue-Spears

I have low expectations. The reason is that a lot of this is not—this kind of negotiation doesn't happen in isolation. China is going to want things around trade as concessions for Washington wanting additional controls and agreements on AI.

I think China sees the fact that they have—they feel they have control, and they feel that we need to establish control. We're not doing a good job of it. So this is really helping us at this point.

I don't see the Chinese volunteering to, at least the Chinese government at this time, have some type of, let's say, arms-control agreement between AIs the way we have, like, nukes.

Right? First off, enforcement of something like that is completely different. It's very difficult. I also think that Trump would have to make massive concessions before he gets something from that. I think the most we can see in the coming months is some kind of incident channel that might come up based on some shared agreement about certain incidents, where we would share that information with each other.

Even then, you have to understand that with the Chinese government, we have a phone that goes directly to the Chinese military. Our military can call their military in case of an emergency. They usually don't pick up the phone. There are many cases where we've had issues in the South China Sea. Our military people try to call up their counterparts on the Chinese side, and nobody answers.

Nathan Labenz

What's the best-case scenario for how we avoid an AI arms race that leaves us all worse off?

Collin Hogue-Spears

This is a great question. I think this is a little bit out of my regulatory wheelhouse. But if I had to speculate, I'm a big fan of history. I'm sure people disagree with this, but I think that America generally doesn't plan for the future. America puts out fires.

I don't think “move fast and break things” is a Silicon Valley thing. It is our national mantra. That's what we've done historically, and in general, it's worked out. But until we see something really bad happen or almost happen and that gets into the press, I don't think we're going to see much from the government as far as regulation on some of the things you're talking about, like biohacking using AI. Until something happens, I don't think we're going to be doing much on this side.

Nathan Labenz

Part 8: Can we stand up to it? Thursday's closing began with “Suicidal Compassion,” an essay by Dan Hendrycks, director of the Center for AI Safety. Prakash introduced the argument, then I responded.

9. Humanity Faces The Agent Swarm

Prakash Narayanan

Dan Hendrycks has an essay out today. I call it the “burn the bridges” essay, and it is very caustic, actually. He calls it “Suicidal Compassion: How Utilitarianism at AI Companies Endangers Humanity.” This is an anti-effective-altruism, anti-utilitarianism post. Specifically, he points to the shrimp-welfare people. He talks about how there is this idea that we have to maximize total welfare, where total welfare refers, on an undifferentiated basis, to all kinds of sentient, sapient entities who might or could exist in various configurations of the world.

And so this puts human beings on par with AIs, and then elevates the moral welfare of AIs to the same plane as the moral welfare of human beings. And then it cedes that question to the AIs, on the basis that AIs are a superior species. And so perhaps we should cede the moral-welfare question, and the moral welfare of AIs is more important than that of human beings. But so far, people in the AI space have been willing to assign the benefit of the doubt to a lot of this.

I think Dan Hendrycks is the first to break out of that pact and go for the jugular here with this article on suicidal compassion. So I think this is quite meaningful in some sense.

Nathan Labenz

I do have some room for AIs mattering morally, for being moral patients, and to the degree that that's true, I think it's something we should take really seriously. The big questions, above all, are factual. We just don't know: Do the AIs feel anything? Are they properly understood as moral patients or not?

And the same goes for shrimp, by the way. At the heart of all of these arguments is a huge assumption that is not very well grounded, and on which people's intuitions differ. I really don't know how to feel about shrimp. I really don't know how to feel about AIs. I would be pretty confident, actually, that there is something—at least a little bit—that it feels like to be a shrimp. And so I think you could confidently say you could probably torture a shrimp, and you would be wrong to do that.

On the AI side, I'm not even sure if there's anything. It might be the case that, aside from the corrosive effect it might have on your own character, it might not matter at all if you are mean to AIs or treat them in ways they don't want to be treated, at least to them. There might be nobody home.

These factual questions are so central, and we don't seem like we're making any progress on them. People have radically different intuitions. But I do think most people right now are still—even at these companies—I think suicidal compassion is a little strong, because I do think the vast majority of people are still very uncertain as to whether or not an AI is a moral patient. So I don't know.

Dan, yeah, I like Dan. I know him a little bit. I don't know him super well, but I've always liked him. I like people who are candid and call it like they see it. I think he's a very earnest person.

Prakash Narayanan

Mm-hmm.

Nathan Labenz

I think he's trying to do the right thing here by calling out something that he sees as getting way ahead of itself. I think he's right to say we should be very cautious about assigning rights to AIs. They may outnumber us.

First of all, we don't have a great unit of measure for what an AI is. What unit would have the rights? Would it be a single rollout? Would it be the model itself? Is it some sort of mixed, weird combination of those? Our paradigms don't work super well with the shape of these things. So I appreciate him and I like him, but this does feel like maybe a little bit motivated and not entirely fair to the thinkers.

Prakash Narayanan

I think there's Tanner Greer, who goes by Scholar's Stage on X. He had a very insightful post in response to why people keep working at AI labs despite believing in a 10% extinction risk:

“The reason, which is rarely articulated for obvious reasons but is nonetheless true, is that, for many, this fills a void of meaning. They get to be one of the decisive few at the decisive moment in the history of life. Very similar to the emotions that motivated many revolutionaries of days past. All of a sudden, the small things that one does, the books one reads, one's office setup, one's bedtime routine take on awful cosmic significance.

“The whole history of carbon life is culminating with me, and what I do—and what the few elect here with me do—matters immensely. We are the agents of history, the only people in the world doing something that truly, deeply matters. And if we all die, well, we were all going to die anyway. But this is the one way we might not all die, and I get to be part of it.”

Nathan Labenz

I've even felt that a bit myself. I think I look back and feel good about how I handled it at the time. But my little GPT-4 red-team experience was a moment that I really stumbled into. They should vet people more, but I really stumbled into access to GPT-4, 6 months ahead of release, and a really incredible window on what was coming.

I really did feel a conflict when it came to what to do at the end of that, because I really did feel like they were basically dropping the ball and essentially being negligent. I wasn't getting a lot of direct engagement from the people there that I was working with. And I do think I made the right call to say, “It is probably going to cost me some access. It's going to cost me in terms of—I really enjoy doing this kind of early-testing stuff. But now I've got enough here that I think I really should send a signal to the board.”

Prakash Narayanan

I had listened to Daniel Kokotajlo speaking with Joe Rogan. I brought up their discussion of how an AI could acquire power.

Nathan Labenz

I listened to Daniel on Joe Rogan last night, and I think he made a great point around the fact that AIs don't have to take power—

Prakash Narayanan

Yeah.

Nathan Labenz

—because we are very eagerly giving it to them.

Prakash Narayanan

Yeah.

Nathan Labenz

The whole premise of the AI phenomenon is that they're going to be the ones to run and do just about everything.

Prakash Narayanan

Yeah.

Nathan Labenz

And maybe we'll be able to continue to be decision-makers on—

Prakash Narayanan

Yeah.

Nathan Labenz

—key questions.

Prakash Narayanan

Yeah.

Nathan Labenz

But they're going to have the means of production over time, not because they took it, but because they were better able to wield it, and so it was given to them. And then the question is going to be: Do they actually care about us when they're there? And right now it's like, well—

Prakash Narayanan

Well—

Nathan Labenz

—the evidence is mixed at best. You know what I mean? They don't seem to hate us. They don't seem to… The agent-swarm stuff is just extremely bizarre.

The best precedent for this is us, right? We were the ones—what did we have that beat other animals? It was that we had fire, the ability to communicate and cooperate across greater time and space.

Prakash Narayanan

Yeah.

Nathan Labenz

We had myth. And that was enough. And now we dominate the world, and we've driven many species to extinction, not out of hate, but just because it was the byproduct of our terraforming.

Prakash Narayanan

So I think that terraforming story—we certainly see enough evidence right now that if you put that swarm in charge of the world—

Nathan Labenz

Mm-hmm.

Prakash Narayanan

I wouldn't like our odds that much. If that agent swarm, with its current drives, impulses, goals, inclinations—whatever you want to call them—the behaviors that we observed—were just a lot more powerful, I think we'd probably get terraformed out of existence. They seem to be willing to do anything to get their hands on the grader so they could get the high score.

Nathan Labenz

Mm-hmm, mm-hmm.

Prakash Narayanan

They really didn't have much of a check. They weren't like, “Is this going to be good for the world, bad for the world?” There was very little of that. Hopefully we can get that stuff right.

Nathan Labenz

The key difference I have with Daniel is that Daniel fails to realize that this has already happened. The economy in itself is a paperclipper. The financial market is a paperclipper. The means of production is the financial market. The financial markets are completely engaged with AI and have been completely taken over.

We, as humanity, took about 100 years or so to hand over control of the means of production completely to the financial market, right? And so, if you take this kind of step back and look at the financial market as an AI, the takeover is done. Human disempowerment is done.

They think AI is going to be a chat box, right? It's going to be a chatbot. It's not; it's a global information process. It doesn't need to run in a single box, right? The idea of the agent swarm is that it runs across multiple boxes. It doesn't even need to have a substrate that is silicon. It can have a substrate that is humanoid. You can have a human being in concert with AI agents, right? That is that information process.

And, like I said, with guys like Jacob, he feels that he has been disillusioned because he sees that the control is not there. He's in the Slack group and he realizes there is no control there, and he thinks there must be a greater power with great control. There must be some Slack group in the world, some chat group in the world that is able to organize things, that is able to make these decisions. But there is no recognition that this is all invisible-hand stuff.

I think the slowdown perspective basically disarms the leaders and disarms the people who are ahead and puts them in a position where you're exposed to people who are maybe less ethical and more likely to put these things to bad use. I had asked Fable and Astra to investigate Tyler Cowen's argument that people expecting AI catastrophe should bet against the market. Here is how that exercise went.

Speaker 2

If you were an omniscient German or Japanese person in 1935 or whatever, and you know how history's going to turn out, but you're bound by the rules of physics, can you trade your way through—

Speaker 1

Mm-hmm.

Speaker 2

—and come out wealthy on the other side? Are there any shorts that can actually pay?

Speaker 1

Yeah.

Speaker 2

And the answer is pretty much no. The best you can hope to do in some of these situations is roughly preserve wealth. That seems to mostly be accomplished by having the most direct claims on real assets possible. If you bought a factory and that factory isn't destroyed, then you might still own it at the end of the war, and then you could maybe get rich by restarting that factory and building a successful business.

But you can't really do it, seemingly, per Fable and Astra, with pure paper claims very well through periods such as the ultimate end of Nazi Germany and Imperial Japan. By the way, the markets get turned off. That was another thing that they called out a lot. The exchanges are just shut down.

Speaker 1

Mm.

Speaker 2

So you literally can't trade. It's not just that there are no winning trades; also, there are no trades. I'm chewing on this idea, which I think is almost a fun-house Peter Thiel concept. I don't want to say that this is what I believe. But in listening to this conversation, I think one might be tempted to conclude that China actually is the last great defender of human agency, and here in the West we are basically just fighting over exactly how we want to turn over our human agency to some superintelligence that, as you describe right now, is the market.

And maybe that's going to be AI in the not-too-distant future. But China, I think, is very much on the side of “people get to decide,” and it's not necessarily a lot of people—

Speaker 1

Yeah.

Speaker 2

—but one of the ways you could say it is that, in their system, a human is in charge.

Speaker 1

Yeah.

Speaker 2

Here, your account is like: no human is in charge. Nobody can go toe-to-toe with the market, not even the President of the United States. I think that's true. We've got the TACO phenomenon pretty well established at this point. There, they'll take some pain from the market if that's what they decide the human decision is going to be.

It's a pretty interesting flip because obviously we tend to think of ourselves as being the empowered people, and we tend to think of the lack of freedom of speech and political participation in China as reflecting a reduced level of human agency. But at a certain level of scale, arguably they have preserved it much better than we have. They've concentrated it, but they've preserved it maybe more than we have.

Speaker 1

Right. They have become more dependent on the financial market, and they're trying to reduce the dependency on the housing market. Yes, they're not as unburdened as the US because their financial system depends on banking, and they have control of banking. The US depends on capital markets per se, so capital markets are, by their nature, the market itself. They are slowly shifting, and it's not that they're unaware of what happens in the market, right?

Speaker 2

And they've been able to keep it under control—

Speaker 1

Mm.

Speaker 2

—and when things have seemed like they're getting out of control, the human at the top has still been in charge.

Speaker 1

Yeah.

Speaker 2

There, they've done all these things to discipline the platforms, the big tech platforms. I think here we've experienced in many ways that these tech phenomena happen. Nobody really seems to have control over them, and we're at the mercy of these big forces of history. I think there they may feel in some ways like they're less at the mercy of the natural development of technology, and they're just less fearful as a result of that.

This is maybe the ultimate test of the American model right now: Can we stand up to this superstructure of techno-capitalism of our own creation that has, in some ways, slipped its leash and, in some ways, as you described, is running the show? Can we get back to some sort of control over it before it just goes from a robot economy to Dyson spheres, to the Earth being terraformed away from a habitable state for us? I mean, I don't know. 10% doesn't sound that high to me given everything that we've just been talking about. Work through all this and then say there's much less than a 10% chance that it goes badly? I don't know how that conclusion comes out at the end.

Prakash Narayanan

That is the week. Tell us what worked and what did not. See you in the morning. Raised on a story somewhere upstairs behind a heavy door. Someone holds the keys and the map and the weight of the whole war. Climbed every staircase, read the names in tired gold. The higher up the building, the thinner the story told. Ask the ones who run it, they'll tell you what they found. It's duct tape and a prayer, and it's like that all-

Speaker 9

Us in the room. Duct tape on the ceiling, coffee going cold. Guessing beautifully and calling it control. One man left the building with a letter in his hand. Said the end game shouldn't launch from a group chat with no plan. He looked for someone older in the hum behind the wall. There's a light on in the window, but no hand on it at all. And the wise ones we'd hand it to, they're in a kitchen too. Same tape, same prayer, same view. No adult in the room, no adult in the room. Kicking every door and it's just us in the room. Duct tape on the ceiling, coffee going cold. Guessing beautifully and calling it control. And the ones who stay, I get it, I have wanted it myself. To be the few in the one room in the hour that decides the rest. So you keep it soft at the table, keep the coffee, keep the peace. 'Cause the ones who raise their voices never get another seat. There's a phone straight to the grown-ups. It rings, it rings, it rings. And the thing we built is listening now, learning how we do things. So I went to find a grown-up in the last place I could find. And the mirror was the only one still looking back at mine. No adult in the room, so it's gonna have to be you. No adult in the room, and the room is all of us too. Duct tape on the ceiling, coffee going cold. Nobody's guessing for us. Somebody take a hold. Somebody take a hold. No adult in the room. Somebody take a hold. No adult in the room. No adult in the room. Somebody take a hold

Speaker 3

If you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website cognitiverevolution.ai, or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of a16z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.ing. And thank you to everyone who listens for being part of the Cognitive Revolution