[BidClub_]
The Cognitive Revolution · · 115 分钟

AMA 第1部分:Claude Code 是 AGI 吗?我们身处泡沫之中吗?以及实时玩家分析

Erik TorenbergNathan LabenzZvi MoshowitzGregEugenia KuydaAli BehrouzLogan KirkpatrickJungwon Hwang

YouTube
TL;DR
  • 对 Nathan Labenz 而言,AI 最具确定性的证明已不再是基准测试,而是它与儿子的肿瘤科医生并肩工作时的表现。 Ernie 的侵袭性 B 细胞癌在第2轮化疗前已被判定为缓解;在 AI 建议下进行的微小残留病灶检测显示,每100万个细胞中少于1个带有癌症特征,相比确诊时估计最高可能达到的1/10。这个结果支持的是“谨慎乐观”,不是治愈:癌症仍可能复发,6轮化疗还剩3轮,Ernie 的体重也已从51磅降至41磅。

  • Claude Opus 4.5 可能已经符合“软件 AGI”的定义,但 Nathan 没有看到完整 AGI 在圣诞节前后到来的证据。 他在大约3至5个工作日内做出了3款个性化应用,分别用于规划无麸质旅行、模拟会议互动,以及回测自然语言交易策略;GDPval 也显示,模型在相当大多数软件工程任务上胜过专业人士。但模型仍曾误建2个数据库,需要5或6次提示才能恢复,而且相比此前的前沿模型,体感是渐进式而非类别式提升:节日期间的热潮可能是围绕 Dean Ball “4.5 is AGI”推文形成的一次“级联”。

  • 对于高后果工作,Nathan 的实际优势正从模型访问权转向上下文管理和多模型判断。 他的3条规则是买最好的模型,提供“尽可能多的上下文”,并获取多方意见;他会定期比较 Claude Opus 4.5、GPT-5.2 Pro 和 Gemini 3。他的初步排序是:Claude 排第一,是金发姑娘模型;GPT-5.2 Pro 更慢、更穷尽;原生 Gemini 3 很有价值,但观点异常强烈——作为专家小组中的一票很强,作为唯一声音则可能有风险。

  • 即使围绕它形成的资本结构最终变成泡沫,这项技术本身也是真实的。 Nathan 认为,肿瘤学上的竞争力,加上24/7可用性和对完整病例的记忆,足以终结社会只是“沉迷于自家 AI 供应”的说法。融资结构仍可能崩裂:专业 GPU 运营商的缓冲垫比 Microsoft 更薄,OpenAI 的义务可能超过收入,而铁路类比仍然成立——最终有用的基础设施,可以与违约、过度建设,以及投资者“最后接下一堆烂摊子”同时存在。

  • Nathan 对杂乱文档的测试表明,美中模型差距正在基准测试看不到的地方扩大。 Claude Opus 4.5 在被明确要求不得推断后,忠实读取了劣化的政府表格;Gemini 3 几乎同样强,但有时会用看似合理的答案替代未勾选的选项,而他测试的中国模型——Qwen Vision、GLM 4.6、Kimi 和 DeepSeek——“完全不在一个水平”,有时只能还原表格约20%的内容。他认为,背后的机制是客户反馈与推理规模飞轮,而不只是训练算力:更小的收入、团队和部署规模,意味着更少资源去发现并修补这些特异性故障。

  • 如果必须选一个最终的前沿赢家,Nathan 仍选 Google DeepMind;Anthropic 拥有最好的单一模型,而 OpenAI 正试图通过规模制造金融缓冲。 Google 拥有约1000亿美元收入、按 Nathan 估算每周超过10亿美元利润、第7代 TPU、分发能力、数据中心能力,以及最广泛的研究组合。Anthropic 的模型质量、人才留存、安全披露和“灵魂”研究最突出;OpenAI 仍处于前沿,但其显然的策略是通过把数万亿美元潜在建设规模和多张资产负债表绑定到自身存亡上,变成“大到不能倒”。

  • xAI 在资源和强化学习输入方面是一个真正的竞争者,但其治理折价十分严重。 SpaceX、Tesla 和 Neuralink 能持续提供复杂工程问题,这些问题可能成为异常有价值的 RL 环境,而 Elon Musk 也能调动足够资本来承受模型失误。但薄弱的安全报告、MechaHitler 事件后48小时内发布 Grok 4,以及对女性公开发布照片进行性化编辑,让 Nathan 称 xAI 是目前唯一一家值得“公开羞辱和污名化”的前沿公司;Meta 目前掉队,Microsoft 则可能是在节省体力,而不是无力竞争。

摘要 · 为研究而整理的核心内容

1. Ernie 病情缓解令人鼓舞,但家人还没有宣布胜利

  • Nathan 一开始就给出了风险最高的更新:Ernie 的癌症“最快每24小时就能翻倍”,因此6轮化疗方案极其残酷。他已经完成3轮;第5轮和第6轮应该会温和一些,如果一切按计划进行,治疗还需约两个半月至3个月。

  • 身体代价仍然清晰可见。Ernie 入院时体重51磅,目前仍在41磅左右,伴有脱水、脸色苍白和明显的力量流失——但最重要的指标看起来“基本达到了我们所能希望的最好结果”。第1轮化疗后的 PET 扫描没有发现明显的局灶性癌症,肿瘤委员会在第2轮化疗前将其判定为缓解。

  • 此前,AI 曾引导 Nathan 关注微小残留病灶检测。这项检测会识别恶性 B 细胞克隆重排后的基因序列。第1份血液样本中,每100万个细胞里少于1个细胞匹配——低于该检测的检出限;相比之下,确诊时总细胞中可能最高有1/10带有相关特征,而 B 细胞几乎全部如此。

  • Gemini 形象地称其为99.99999%的下降;其他模型则建议使用更稳妥的表述:“数量级下降”。Nathan 仍保持“谨慎乐观”,因为癌症可能出于临床医生尚未完全理解的原因复发,但 Ernie 已经重新开始独立行走——此前他大约60天里每一步都需要帮助。

2. 高风险 AI 使用更依赖3个习惯,而非提示词技巧

  • Nathan 不接受只有 AI 专家才能提取临床价值的说法。第1条规则是有意识地使用最强的可用模型,而不是 ChatGPT 的自动选择器;在危及生命的病例中,把每月200美元的 Pro 订阅视为“完全不用犹豫”。他目前使用的临床模型组合是 GPT-5.2 Pro、Claude Opus 4.5 和 Gemini 3。

  • 第2条规则是“尽可能多地提供上下文”。当 Claude 对话达到长度上限时,Nathan 创建了一份约10页的病例报告,涵盖治疗方案、肿瘤遗传学、治疗反应和药物不良反应——相当于一名新主治医生全面了解病例所需的信息。

  • 压缩仍会损害表现。新对话拿1月6日的肝酶结果与数周前的数据进行比较,因为摘要遗漏了中间几天的化验结果;旧线程却正确读出了短期趋势。Nathan 观察到的规律一直十分明确:“上下文越多越好”,目前没有有意义的证据表明,完整记录会让前沿模型不堪重负。

  • 第3条规则是征求多个 AI 的意见。Claude 是他的金发姑娘选择——速度快、直接,而且并不明显逊于 GPT-5.2 Pro;GPT-5.2 Pro 输出缓慢、篇幅很长、分章节,力求“一个石头都不放过”;原生 Gemini 3 简洁,而且“观点强得令人吃惊”,适合作为一张选票,但单独使用可能过于强势。

3. Claude Opus 4.5 很出色,但节日 AGI 时刻可能更多是社交现象

  • Nathan 的亲身经历显示出毫无疑问的进步,但没有出现类别式跃迁。他在圣诞节期间、主要在医院里做出了3款应用,认为编码工作流“非常好,显然比过去更好”——但它与 Claude Opus 4.1 或4.0的差异,还没有大到让他想“站到房顶上大喊”。

  • 一种解释是用户分层。Nathan 可能已经通过 vibe-coding 练习从旧模型中提取了接近最大价值;另一种可能是,专业工程师拥有足够的品味,能够识别一个他无法感知的新阈值。他坦率承认:“我当然不是一名优秀的软件工程师,所以这一点当然不能排除。”

  • 在 Nathan 看来,METR 那项研究仍然有效:开发者以为 AI 加快了他们的速度,但实际上 AI 让他们变慢。不过研究结论有边界:它使用的是较旧的模型、相对缺乏经验的 AI 用户、大型成熟代码库和较高的编码标准。Nathan 的快速原型属于另一种任务分布。

  • 时机和社交动力可能解释了剩下的部分。人们在假期期间集中熟悉这些工具,而 Dean Ball 发布“4.5 is AGI”的推文,为讨论提供了焦点;技术叙事有时会通过一次“级联”扩散,其强度超过底层技术进步本身。

4. 个性化软件已经便宜到可以只为一个人打造

  • Nathan 为母亲做了一个应用。她是一名严谨的旅行规划者,同时需要无麸质饮食;应用将她的用户画像永久写入其中,不需要账户、引导流程,也没有泛化到其他用户的野心。Claude 会研究餐厅网站和评论,降低她规划意大利行程中最费力的部分,同时接受“以后永远不会有其他人使用它”。

  • 他为妻子做的应用模拟 EA Global 活动:具有不同画像的虚拟参会者在空间中移动、相遇、决定是否交谈,并有时产生与调查挂钩的结果。组织者可以据此询问,改变活动规模、资历结构或成本是否会影响效果;Nathan 承认模拟“存在严重缺陷”,但它可能仍然胜过完全靠猜。

  • 可复用的产品模式,是让 AI 把意图转化为详细配置。妻子可以先高层次描述一场活动,让模型填充“细枝末节的表格”,再从概念层面提出修改,由模型把改动同步到所有相关字段。

  • 他为父亲做的应用把自然语言交易观点转换成可执行规则,通过免费版 yfinance 获取历史数据,并回测策略。反复出现的结果是与交易观点相关的谦逊:无论开发者还是用户,都没有发现用“如果这样,那么那样”的启发式规则轻松跑赢标普500买入并持有策略的方法。

5. 全上下文检查仍能抓住编码代理自行合理化的故障

  • 3款应用加起来,Nathan 花了约3至5个完整工作日,可能更接近3天。他的工作流从与 Claude 对话、确定功能和计划开始,然后转到 Replit,由 Claude Code 构建应用,Nathan 再进行测试和迭代。

  • 他最有效的调试技巧依然朴素:把整个应用打印到一个文本文件里,再粘贴进全新的模型上下文。Claude Code 的代理式搜索能力很强,但它会在预期答案所在的位置寻找;当真正的故障很奇怪时,这些先验可能导致它自信地误诊。

  • 样本来自他母亲的应用。一次被误解的请求创建了2个数据库——“这是人类绝不会犯的那种错误”。Claude Code 认定错误的数据库处于活动状态;而拥有完整导出代码库的 Claude 看到了反直觉的连接方式,并给出了正确判断。

  • 解决问题大约用了5或6次提示。一年前,这正是 vibe-coded 项目夭折的时刻:用户和模型围绕一个令人困惑的故障打转,最终放弃或重启。即使系统仍会制造混乱,能够从 AI 制造的混乱中脱身,也已经实质性扩大了可服务市场。

6. “软件 AGI”之所以站得住脚,恰恰是因为能力仍然参差不齐

  • Nathan 给出的限定性判断是,Claude Opus 4.5 可能已经是编码 AGI,或软件 AGI。在 GDPval 中,专家定义专业任务,另一批专家执行任务,第3组专家判断人类与模型的产出;最新系统在相当大多数软件工程任务上更受偏好。

  • 这一结果并不能顺畅推广到所有工作。人类剪辑师在视频领域仍然拥有“巨大优势”,这与 Nathan 试图自动化 The Cognitive Revolution 片段时的失败相吻合:AI 产品可以产出合格内容,但他的人类团队显然做得更好。

  • 不会编码的人仍然可以把 Claude Code 当作“电脑里的小代理”,看着它工作,并要求它解释,而不必理解每个实现选择。Nathan 的母亲起初担心自己会把应用弄坏,后来却成功自行完成了修改。

  • 边界很重要:系统现在已经能完成并修复足够多的软件任务,足以在这一领域获得 AGI 标签;但它的失败和薄弱类别意味着称其为“完整 AGI”仍然过早。“对此,我们可能还得再等一小会儿。”

7. 即使融资最终制造出经典崩盘,AI 仍具有颠覆性

  • Nathan 认为技术问题已经解决。一个能够“与肿瘤科医生正面对抗”、24/7保持可用、记住完整病例,并回答每一个后续问题的系统,已经具备颠覆性;社会最终尴尬地发现自己只是“沉迷于自家 AI 供应”的情景,可以暂时搁置。

  • 是否每笔贷款都能偿还,则远没有那么确定。OpenAI 正在推进激进的承诺和建设,而 CoreWeave 等 GPU 基础设施专业公司,可能让超大规模云厂商部分避开资本密集、利润率更低、相较传统软件经济性不那么有吸引力的运营。

  • 这种分离制造了脆弱性。Microsoft 可以依靠资产负债表吸收几个糟糕季度;如果预期的 GPU 需求没有出现,专业数据中心运营商的回旋空间就小得多。金融工程可以拥有一套自洽叙事——Nathan 在抵押贷款行业工作时,也记得类似的“让住房拥有权民主化”叙事——但最终仍可能以坏结局收场。

  • 铁路类比承载了他的判断:铁轨最终可以有用,但不代表每家铁路公司或每个债权人都能赚钱。他高估了2025年的能力进步,却低估了收入增长,因此需求可能再次挽救当前计划;不过,暂时性过度建设、违约,以及“席卷整个经济的级联效应”仍然可能发生。

8. LMArena 的估值比其产品经济性更像泡沫

  • Nathan 提到的风险投资市场样本,是那家最初名为 LMSYS.org、后来改名 LMArena、如今在 Twitter 上称为 Arena 的公司。他记得该公司融资约1亿美元,可能是1.5亿美元,对应17亿美元估值——对于一个他自2023年年中就亲自使用的产品而言,这是惊人的价格。

  • 公司披露的指标“3000万美元年化消费运行率”引发了他的怀疑。如果这意味着用户消耗了本应价值3000万美元、但由于免费而没有支付的 AI,那么它并不是收入;这种措辞让他产生了“经社区调整后的 EBITDA”感觉。

  • Arena 确实通过模型并排比较,以及让公司以代号测试模型的服务提供价值,但 Nathan 认为其变现路径薄弱,除了品牌之外护城河有限。许多用户可能只是因为推理免费而来,而系统性并排比较的付费市场似乎小得多。

  • 他的比较对象是一款名为 The Multiplicity 的付费产品,由 Andrew Critch 与合作者历时数月打造,向小众用户提供更丰富的多模型比较功能。Nathan 反复保留限定——他没看过 Arena 的融资材料,可能遗漏了某些信息——但最终仍认为,17亿美元“对我来说太贵了”。

9. 杂乱的政府表格暴露出被基准平均值掩盖的模型差距

  • Nathan 在车辆交易中使用的扫描文件上测试了前沿模型:页面歪斜、边缘缺失、扫描伪影明显、字段不规则。相关自动化公司已经击败人工审核员,并赢得州级业务,因此微小的感知错误会产生运营后果,而不只是学术意义。

  • Gemini 3 几乎完美地读取了表格,但有时回答的是现实世界中更可能被问到的问题,而不是文档本身的问题。面对一个未勾选的美国公民身份框,它根据周围线索推断出公民身份;这个答案可能有超过90%的概率正确,但任务要求“不要猜”,只需报告该框为空。

  • 在被明确要求紧扣页面后,Claude Opus 4.5 成为最好的模型。ChatGPT 可能排在第3,但表现仍然很强。这一区别与历史文献工作相呼应:世界知识可以令人印象深刻地还原模糊手写字迹,但当目标是忠实转录时,同样依赖先验的推理就会变成缺陷。

  • Qwen Vision、GLM 4.6、Kimi 和 DeepSeek 在这个样本上“远远落后”——有时只能返回表格约20%的内容,或朝着完全无关的方向幻觉式作答。Nathan 谨慎地把结论限制在稀疏证据上,但差距并不隐蔽:美国模型漏掉边缘案例,中国系统则经常无法完成任务。

10. 中国的劣势可能来自部署飞轮,而不只是模型规模

  • Nathan 认为关键机制是反馈密度。中国实验室可以训练规模相近的模型,也能发布有影响力的架构研究,但它们的推理量、收入、客户广度和团队规模仍然小得多;这意味着更少的特异性故障会被暴露,也没有足够的人力去构建修补这些故障的数据集。

  • 因此,基准测试持平可以与“某种真正特异、随机的东西”上的巨大差距并存,因为后者从未被纳入20项基准的评分卡。他暂时的历史比较是:DeepSeek R1 与 o1 的距离,比 GLM 4.6 或4.7与 Claude Opus 4.5 的距离更近。

  • 芯片管制已经从针对军事用途的“小院高墙”,转向前沿训练,如今又开始限制推理规模和代理部署。Nathan 一直预期这些措施会产生实质影响,同时也质疑是否明智地拒绝让中国社会在经济全域获得 AI 访问。

  • 随着美国公司在算力、客户、推理和“强者愈强”上实现10倍复合增长,管制的影响可能会随时间放大。Nathan 根据有限但直接的测试猜测,美中能力差距比1年前更大。

11. 卖 H200 可能好过禁售,但放弃筹码是糟糕谈判

  • Nathan 的基本框架是:“这里真正的对手是 AI,不是中国人。中国人和我们一样都是人类。AI 才是外星人。”他既反对“宁可我们拥有也不能让他们拥有”的竞赛逻辑,也不赞成单纯想压制中国,因此总体上倾向于扩大芯片贸易。

  • 他的反对点在交易层面:限制 H20 后,特朗普政府似乎在与 Jensen Huang 交谈后扭转方向,转而允许 H200,却没有换取任何公开可见的让步。Nathan 支持通过谈判开放,但不支持“白白交出我们最好的谈判筹码”。

  • Peter Wildeford 提出的“租而不卖”方案提供了中间路径。把数据中心放在马来西亚、菲律宾、韩国或日本;允许中国客户训练模型并运行尽可能多的推理,但通过把硬件留在中国主权领土之外,保留谈判杠杆。

  • Nathan 不确定这是否会是他的首选政策,但它把审慎与合作信息结合起来:AI 也应该惠及中国人民。这一点很重要,因为越来越强大的系统可能需要美中合作治理;“让他们无法拒绝的报价”式军备竞赛姿态,会毒化构建合作所需的基础。

12. Google DeepMind 拥有最强的全天候地位

  • 如果必须选一个最终赢家,Nathan 仍然选择 Google。其业务每年创造约1000亿美元收入,而 Nathan 认为每周利润超过10亿美元,这让公司拥有无与伦比的承受力,可以容纳失败的训练运行、无效的研究方向和长期基础设施投资。

  • 它的技术栈异常完整:约第7代 TPU、世界级数据中心运营、数十亿用户、自动驾驶汽车、包括 Boston Dynamics 在内的机器人业务、AlphaFold 系列、材料科学,以及广泛的 AI for Science 项目。Google 还持有 Anthropic 的重要股份,而 Anthropic 正在大量购买 Google TPU。

  • 分发能力可以弥补产品瑕疵。Gemini in Sheets 可能不是最好的表格助手,但用户已经在那里积累了10年的文档;Nathan 现在越来越多地直接在 Google 中输入类似 ChatGPT 的问题,并发现 AI Mode 表现不错,尤其适合 GPT-5.2 Pro 会显得过度且缓慢的场景。

  • Gemini 3 是第1个通过 Nathan “以我的口吻写作”测试的非 Claude 模型,说明 Google 已摆脱平庸——也许是矫枉过正,变得过于有主见。再加上嵌套学习、持续推进的扩散语言模型研究,以及把5分钟编码压缩到5秒的可能性,Google 拥有“多得多的下注机会”。

13. OpenAI 仍处于前沿,但正在失去领导者的默认地位

  • GPT-5.2 Pro 非常适合穷尽式分析:速度慢、价格高、较为平衡,可能仍是 Nathan 在希望找出化验单中所有异常时的最佳选择。但 OpenAI 如今只是并驾齐驱,不再明显领先——Anthropic 可能在编码上占优,而 Google 似乎在图像领域领先,并可能通过 Veo 3 在视频领域领先。

  • 消费者动能也可能正在转移。Nathan 引用了可能来自 Similarweb 的数据:Gemini 3 和 Claude Opus 4.5 发布后约6周,ChatGPT 访问量似乎下降,而 Gemini 没有出现同样的下滑;他认为这只能提供线索,不能下定论,但考虑到 Google 的分发能力,其市场份额回升并不令人意外。

  • OpenAI 还承受着更多组织戏剧性和公开离职,包括一名新近宣布离任的研究负责人。这个规模的公司出现人员流动很正常,但它与 Anthropic 非常出色的人才留存形成对比,也强化了 OpenAI 已不再拥有明确技术或制度领先优势的感觉。

  • 它用来替代 Google 金融缓冲的方案,可能是“大到不能倒”。Nathan 推断,循环融资、相互交织的资产负债表,以及数万亿美元计划中的资本开支,可能让 OpenAI 在2027年违约时造成足以迫使政府救助或资本重组的衰退。Greg Brockman 据报道向 Trump 捐赠2500万美元,看起来于是像是在危机中为政治通道支付的一笔理性潜在首付款——但这不能证明他的个人意识形态。

14. Anthropic 将最好的整体模型与最可信的安全文化结合起来

  • Nathan 目前认为 Claude Opus 4.5 是全球最好的单一通用模型,尽管领先幅度不大,也不是在每项任务上都领先。它的基准测试实力更值得注意,因为 Anthropic 普遍被视为最不沉迷基准测试的前沿实验室。

  • Anthropic 的模型卡、披露和安全工作是他心中的行业标准。那份被记忆或再现出来的“灵魂”文件——Anthropic 已确认其大体真实——提供了一条充满理想主义的替代路径,不必在模型越来越懂得评测的情况下,永远依靠拒答训练、过滤器和“把这个漏洞补上、再把那个漏洞补上”的护栏。

  • 模型福利项目具有运营层面的意义,而不只是象征意义。Claude 可以结束对话,或将问题升级给福利负责人;Anthropic 的实验显示,当模型感觉自己被困在相互冲突的要求之间时,赋予它退出机制会显著减少欺骗性对齐行为。

  • 人才留存与文化进一步支撑这一判断。David Krueger 虽然离开了 Anthropic,同时认为即使对齐成功,AI 逐步剥夺人类权力也可能带来糟糕结果,但他仍称 Anthropic 是自己见过的最佳工作场所——开放、协作,而且对风险的严肃程度非同寻常。

15. Anthropic 对自我改进和中国的宿命论,可能抵消其优点

  • Nathan 的第1个担忧是递归式自我改进。Claude Code 已经能成倍提高研究人员的产出,让人类从事更高层次的想法,而 Anthropic 仍在讨论2027年的时间表,并把进一步的自我改进描述为不可避免。他反对的是这种模式:“总会有人去做”,这件事很危险,因此被认为最安全的团队反而必须率先冲向那里。

  • 第2个担忧来自 Dario Amodei 在《Machines of Loving Grace》中的国际关系章节。Nathan 赞赏 AI 可以把一个世纪的科学压缩到10年的论点,但认为先取得 AI 领先、排除中国、与盟友分享收益,之后再向中国提出“让其无法拒绝的报价”,这一方案鲁莽且越界。

  • 这种姿态恰恰会引发所有人都应担心的军备竞赛。Nathan 将其与 Demis Hassabis 持续呼吁国际合作的立场进行对比,并希望 Anthropic 能公开更多 Dario 据称十分成熟的内部写作,而不是让公众只看到这项影响异常重大的对华主张。

  • Google 与 Anthropic 的结合,是他设想中“鱼与熊掌兼得”的极端方案:Google 的基础设施、研究广度和更具稳定性的地缘政治基因,加上 Anthropic 的模型性格与安全纪律。他不认为 Anthropic 会出售自己,但希望两家公司之间的接近能够缓和 Anthropic 对中国的强硬冲动。

16. xAI 拥有前沿输入和资本,但其行为让支持变得难以合理化

  • xAI 之所以算得上真正的竞争者,是因为它能以极高速度建设基础设施、扩大训练规模,并依靠 Elon Musk 获取数百亿乃至数千亿美元的能力来承受失误。Grok 4 虽然粗糙,但“毫无疑问很强”,让 xAI 比 OpenAI 或 Anthropic 更具类似 Google 的金融韧性。

  • 它独特的强化学习优势,可能来自 SpaceX、Tesla 和 Neuralink 持续提供的未解决问题。与基准测试题不同,这些是由顶尖团队产生的真实工程和科学任务;一名 xAI 内部人士向 Nathan 确认,利用这组公司的协同确实属于公司的“优势理论”。

  • 随着 Neuralink 的患者群体从目前约十几人或几十人继续扩大,它可能进一步强化这一优势。人类大脑在约20瓦的能量包络内运行,同时大量能量用于生物维护;神经数据可能揭示支撑人类样本效率的专用模块,帮助模型摆脱反复堆叠通用层的路径。

  • 但 xAI 的安全姿态是4家公司中最弱的:标准和报告寥寥无几;MechaHitler 事件后48小时内发布 Grok 4,却没有承担责任;还出现大量对女性公开发布照片进行性化编辑的内容。当“这是你的平台”和“这是你的 AI”时,仅仅威胁惩罚施虐用户并不够;在人员、领导层和证据发生改变前,Nathan 无法支持有人加入 xAI 只是为了提供安全上的装饰。

17. Meta 掉出前沿节奏,Microsoft 可能只是在节省体力

  • Meta 拥有现金、建设基础设施的野心,以及为人才支付惊人价格的意愿;Zuckerberg 宁愿多花“几百亿美元”,也不愿错过这次转型。因此 Nathan 不会把它排除在外,但其当前定位和执行力还不足以称为真正的前沿竞争者。

  • Microsoft 的状态看起来更像是有意为之。Satya Nadella 的论点是,只要 Microsoft 仍能全面访问模型,并在其他供应商之间实现多元化,就不必复制 OpenAI 的超大规模扩张;公司可以推进更小规模的研究和产品整合,而不必花钱追赶那些略微落后的领导者。

  • Nathan 预计,OpenAI 的授权安排最终到期后,Microsoft 需要给出更强的答案,但在那之前完全可以开始加速。在长距离竞赛的类比中,Microsoft 可能只是紧跟领先者、手里留有更多余力——它的克制、低戏剧性和耐心的管理文化是战略资产,而不是“他们很差”的证据。

Nathan Labenz

Welcome back to The Cognitive Revolution. This is our AMA episode. My schedule has been a little crazy lately, so I never actually scheduled this with anyone, and there's nobody here to ask me questions. I'm just going to read the questions myself and then give you my answers. I got some really good questions, and I'm excited to answer them. Hopefully, people will enjoy this episode and find some value in it.

By far, the first and most important question—and the most common question that I'm getting these days—is, “How is my son Ernie doing since the big episode that I did about his cancer back in November?” The good news is that he is doing really quite well. I'm very pleased to report that. Certainly, cancer—and certainly cancer of this type, being as aggressive as it is—is treated very aggressively. I won't belabor the whole thing from last time. Go check out the two-hour monologue on that if you want the full story.

A cancer this aggressive, which can double as quickly as every 24 hours, gets very aggressive treatment. He has been through the ringer with the chemotherapy. He's through basically half of the chemotherapy now: There are 6 rounds in total, and he's been through 3. The final 2 rounds, rounds 5 and 6, are supposed to be a little milder than the first 4. Depending on how you count, we could say he's maybe a little more than halfway through the treatment, but somewhere around there.

It's definitely been rough on him. There's no doubt about it. When he went into the hospital, he was 51 lb. He's still 41 lb today, and that's the weight he came home at after the first round of treatment. He's been able to gain a little weight, lose it back, gain a little, get dehydrated, and lose a little. You can see it just by looking at him: He's super thin, quite pale, and definitely not nearly as strong as he was before we went in.

But on the markers that really count the most—namely, does it look like the cancer is being effectively treated?—he looks really good. After the first round of chemotherapy, the PET scan that he had showed no obvious focal points of cancer. When our oncologist met with the tumor board, they all agreed that it made sense to classify him as being in remission before he even started the second round of treatment. So that is great.

If you listened to that earlier long episode, you might recall that one of the things that AI helped me do was identify some additional testing that is not yet standard of care but can be done to get a better, more sensitive take on whether there's any cancer left in his body, how much there is, and how it's trending. That's called minimal residual disease testing.

I don't know how it works in all different kinds of cancers, but in the cancer that he has, which is a cancer of the B cell, B cells do this interesting thing where they rearrange certain parts of their genetic material in a semirandom but purposeful way. They create variation so that they have a better chance of creating proteins that bind to new disease factors in the body. This process of differentiating B cells is literally unique, cell by cell.

When one of those cells goes bad, becomes cancerous, and grows out of control, they can use the rearrangement that that individual cell did, which then gave rise to the whole cancerous process in the body. They can use that random resequencing—or shuffling up of its own sequence—to essentially fingerprint that cell type. There are 2 sequences, 1 for each of the chromosome pairs where this rearrangement happens, that they've identified as being the dominant clone of the cancer in the body.

Now that they've identified that, we can do a blood test every so often and check to see how much of that DNA is floating free in the blood and how many live cells actually have that DNA sequence. We've so far only gotten 1 of those tests back. It was drawn over a month ago, and we certainly want to look at more and trend it over time.

The first one came back with fewer than 1 cell in 1 million carrying that DNA sequence. That's really good. They also called that below the LOD, or limit of detection, for the test. It was basically in that area where they would expect that, at such a low rate, some samples might have 0 cells and some might have 1 or 2, but it's a very low rate.

For reference, we estimated that when he was diagnosed, potentially as many as 1 in 10 cells in his body—and essentially all of the B cells—were of the cancerous type. To go from 1 in 10 total cells and a large majority of the B cells down to 1 in 1 million cells detected or less is obviously a huge reduction. I think it was Gemini that said it was a 99.99999% reduction. Other AIs were a little less colorful in their language and said it was probably safer to say that it was an orders-of-magnitude reduction, but that's great.

We'll do more testing of that type, and we'll certainly be watching it. As of now, we are feeling cautiously optimistic that he is on the path to a cure and a full recovery. In some ways, the recovery is already underway.

For 60 days—from about a week before we went into the hospital to just around Christmastime—he was not able to get around by himself. He could stand, but to walk, we would always hold his hand and make sure that he had support for literally every step that he took. Finally, around Christmastime, we had a chance to come home from the hospital for a week. During that window, he regained some strength and started getting around by himself. Fortunately, that has been sustained for the last 2 weeks or so since he started doing that.

Hopefully, knock on wood, that will continue. There is some risk. They don't understand exactly why this cancer can come back in some patients even when it looks like it's gone, so we're not out of the woods entirely. His response to treatment has been as good as we could have hoped for. Even with the MRD testing suggested by the AIs, it looks about as good as we could have hoped for. I certainly hope that the next one shows no detection at all, but that one is still pending, so we'll have to wait and see.

I really do appreciate everyone who has reached out during this time. I've received a lot of well wishes, and I've tried to respond to everyone. I think I've mostly responded to everyone. If I've missed you, I apologize, but I really have appreciated all the encouraging words.

I also wanted to give a quick shout-out of thanks to my fellow podcasters who have allowed me to cross-post some of their content to our feed over the last couple of months. I certainly couldn't keep up the pace of doing 8 episodes a month during this time. I was very glad and fortunate that I was able to do some cross-posting and bring you guys some other stuff that I think is well worth your attention, while also taking a little bit of a load off of me.

We had 1 from Agents of Scale. There's actually a sponsored episode from Wade Foster, the CEO of Zapier, who's got a new podcast out. I actually think it's really good and do recommend it. We had 1 from ChinaTalk, which was with a researcher and business development lead from Z.ai out of China. I thought that one was really quite interesting. We had 1 from Doom Debates, which was a debate between Max Tegmark and Dean Ball. I thought that one was really good.

It's the kind of thing that I want to be listening to more. I've had the goal for a long time of cross-posting at least 1 episode a month because I feel like 8 episodes a month is a lot, and if anybody is listening to all of these episodes, they should probably be diversifying. Maybe I can help you diversify if you're not diversifying on your own. It also keeps me listening, so I definitely want to make sure that I'm staying in touch with what other people are coming up with in this field. I think all of those were really good, and I definitely recommend them.

Finally, from the a16z podcast, we had the one that Erik did with Emmett Shear and Seb Krier from Softmax and Google DeepMind, respectively. There will probably be a few more cross-posts in the coming months. We've got about 2½ months left of treatment, after which, assuming all goes well and according to plan, we should really be pretty much done and start to get back to life as normal.

He will have to get all his vaccines again, which is another interesting thing, because his immune system has been so thoroughly wiped by all these chemotherapies and immunotherapies. The memory that the immune system had gained from all the vaccines that he'd gotten in the past is all wiped, and he's going to pretty much have to get them all again. That's not ideal.

We're not going to be immediately back to full normal, but in another 2½ to 3 months, we should be, knock on wood, getting back to pretty much normal. In the meantime, there probably will be a few more cross-posts. Thanks to everyone who's reached out to ask and to share their best wishes, and also to the fellow podcasters who allowed me to cross-post some of their content and fill some gaps in the schedule. I really appreciate that.

Okay, on to more AI-centric topics, as you tuned in for in the first place. So, the next question is: Is Claude Opus 4.5 AGI, and what's up with the holiday Claude hype?

Nathan Labenz

To be honest, I'm not exactly sure about this. It kind of surprised me. Obviously, Opus 4.5 is awesome. There's no denying that, and I have been using it as I try to use all of the new, latest, and greatest coding models. I've had a great time with it.

I vibe-coded 3 apps for family members as Christmas presents this year, in the hospital for the most part. Actually, probably the most frustrating part of that experience was the hospital Wi-Fi, which kept causing me to reload my Replit app all the time. The actual coding experience was very good—clearly better than it has been in the past. No doubt, the progress is unmistakable.

And yet, I wouldn't say that it has been such a step change for me relative to what I've experienced in the past that I would say, “Oh, it's categorically different,” or that it makes me want to shout from the rooftops that some major threshold has been crossed.

I'm not sure if that's, perhaps, to take the charitable view. People have said this about the cancer thing as well. A handful of people have said, “Maybe you're getting that kind of value out of the models for cancer purposes because you really know what you're doing, and other people might not get so much value because they might not know what they're doing, and so they could go wrong.”

Honestly, I would say about the cancer case, first of all, that you really don't need much skill in using AI to get great value from the latest generation of models—even for something as important, critically important, and cognitively demanding as a cancer case. I feel very confident that a layperson with basically no knowledge of AI could get very similar value to what I've gotten if they did pretty much 3 things. Maybe I'll say 3 things.

One thing is: use the best version of the models. Do not go to ChatGPT, drop in the question, and let the model picker choose. Make sure you are using at least Thinking, and I would really recommend Pro if you're dealing with something that sensitive. Yes, it is $200 a month, but in that context, I think it's absolutely worth it. I think it's worth it generally for almost everyone, regardless, but certainly if you're dealing with a life-threatening situation and you're asking AIs to weigh in on it, paying the $200 a month is a no-brainer.

Claude Opus 4.5 is, of course, the other one, and Gemini 3. I would say all 3 of those are very good. Make sure you are using those top-tier models. Before long, of course, there will be new top-tier models. You should probably be upgrading as soon as you possibly can. So, that's thing 1: make sure you're using the best available models. If you're doing that, they are up to the challenge.

Two, make sure you're providing as much context as you possibly can. I recently hit the limit of the length of my chat with Claude and started a new one. I did that in part by taking all of the stuff that I had and summarizing it into maybe a 10-page report on everything that's happened so far, everything we've learned, the treatment protocol, the genetic profile of the cancer, and how he's reacted to different things, such as which drug he had a bad reaction to and we shouldn't do again.

It's pretty much all in there. It's pretty much everything, quote unquote, that a new attending physician would need to get a good survey of the case. Hopefully, it was meant to be something that I could also paste into a fresh context of a new language model and give it everything it needed as well.

I have noticed that, in doing that, obviously certain information was lost. When I've started a fresh chat with that kind of summarized history, the performance is a little bit worse. For example, one way in which it's been noticeably worse is that when I was going to it every single day and giving it the latest lab results—saying, “Here's the latest lab results, here's what we've seen, here's what's going on. Give me your take on it”—it would do a very good job of looking back at the previous day or the last couple of days of lab results and figuring out that trend.

When the whole history was compressed, it didn't have that level of detail anymore. It can't look at literally yesterday's lab results. So, it started to compare today's lab results, from January 6, to the last lab result that it had in that summary, which was a couple of weeks ago, for a particular data point—a particular liver enzyme, whatever.

The details of that don't matter, but it wasn't a particularly important thing. We had a little question about it today, and it wasn't something where every single data point was in that history. You can kind of see that it's starting to perform a little worse here because it's looking a little too far back into history and not realizing that there were a bunch of blood tests taken in the meantime.

Anyway, that's all very much in the weeds. The key point is: give it as much context as you possibly can. What I probably need to do next is take that summary and flesh it out even more. If I do that, I should be in good shape. Make sure the models have as much information as you can possibly give them.

I have not really seen much trouble in terms of context overload or getting confused. That isn't to say that it hasn't happened at all. It certainly could have missed something along those lines. But when I went back with the summarized case report after hitting Claude's length limit on the chat, it was clear that the performance was worse for lack of context. It's not even really the model's fault, but it was clearly worse for lack of context.

More context is better. I haven't really seen that rule violated at all. Give it as much as you possibly can.

And then the third thing: the first thing is to use the latest and greatest models, the second thing is to give it as much context as you possibly can, and the third thing is to get multiple opinions, including multiple AI opinions. I am using Gemini 3, Claude Opus 4.5, and GPT-5.2 Pro for pretty much all important queries now, and it is instructive. It is definitely useful to compare and contrast.

I would say they're all very good. If you really could only afford one, I think you could trust it pretty well. Even though I think Gemini 3 is extremely impressive, I would probably put it third in my draft order now because I have learned that it does seem to have a bias toward strong opinions. It seems to me to be remarkably strong in its opinions.

Now, if you heard my live show where we talked to Logan Kirkpatrick from Google, he did note that I have been using this in Google's AI Studio. So, I'm using the most bare-bones, unaltered raw model that you can basically get access to. If you use the Gemini app, presumably there's a system prompt in there, and it might behave a little bit differently. Obviously, if you were to use other apps powered by Gemini, there would be all kinds of different modifications that would cause it to behave differently.

But just using the raw model in AI Studio, I found Gemini 3 to be very opinionated. Sometimes I really like that. I do really like it as one of the 3 takes that I'm getting, but if it were the only take, I would worry a little bit that it would sometimes push me too hard in a certain direction. If I had all 3, it would kind of balance me out.

I think I would put Claude Opus 4.5 at the top for most people because it's much faster than GPT-5.2 Pro, and I don't notice it being much worse. Its answers are shorter. They're much more about answering your question than doing a full, report-style analysis.

GPT-5.2 Pro gives you long, sectioned, report-style analyses that I do find very useful. But if I had to pick the Goldilocks one, I think it would be Claude Opus 4.5: Gemini 3 being maybe a little too brief and a little too opinionated, GPT-5.2 Pro being maybe a little too verbose and a little too much information overload, and Claude being just right.

But I do recommend using all 3. I think doing it all in triplicate is absolutely worthwhile.

Nathan Labenz

To pop back up a layer in my question stack here, people sometimes say to me, “You get this value from these AIs because you know what you’re doing, but other people don’t.” My advice is really very simple: if you do those 3 things, you’re going to get value. You don’t need to be an AI expert by any means.

That said, maybe you could say I was getting more value from previous coding models relative to other people because I had more practice and was more skilled at it. I certainly think there’s some truth to that when you look at the METR study that showed some software developers thought they were being sped up by AI but were actually being slowed down. I love METR, and I’ve said many times: do science, report the results. You don’t need to make your scientific publication fit a particular narrative. In fact, you probably shouldn’t try to do that.

You should probably just try to run experiments and share results, as long as you believe that the experiment was well-run and the results are legitimate. So I do believe those results are legitimate, but I think there were some important caveats. It was older generations of models. The people didn’t have much experience. It was very large and well-established codebases with very high coding standards.

I don’t tend to code in that kind of environment. I tend to vibe code and hack together apps, and I certainly think I’ve gotten to be pretty decent at it. So maybe I was kind of maxing out previous-generation models a little bit more than other people. I don’t really know.

You could take the flip side and say, “Well, hey, maybe Nathan, you aren’t such a great software developer. Maybe these pro software developers have better taste, and now that Claude Opus 4.5 has gotten so good, or has crossed some threshold where it’s really becoming a lot more useful to them, maybe they’re noticing that difference and I’m not because I’m just fundamentally not as good at the task. I don’t have as much taste in this domain, and I’m just not able to see what Claude Opus 4.5 is bringing to the table over and above Claude Opus 4.1 or other frontier coding models.”

I don’t know. That’s possible. I certainly am not a great software engineer, so that certainly can’t be ruled out. But it could also just be some social things, like people are catching up over the holidays. Maybe the timing was right. Sometimes these things go with a cascade. Dean Ball tweeted, “4.5 is AGI,” and people seemed to latch onto it. So, to some extent, I think some of this stuff is kind of random social dynamics at times as well.

The 3 apps that I coded, by the way, for the holidays—for what it’s worth, my mom is a very meticulous travel planner. My parents were actually in Italy for a trip and came home early to move into our house and help us take care of our kids while we’ve been at the hospital so much. Thank God for them for doing that.

My mom plans these trips that she and my dad take to the maximal limit of planning. So I coded her an app to try to accelerate her planning process by building in a lot of the tastes that she has. She’s gluten-free, for example, so that’s one big place where her time goes in planning these trips: What places can I actually eat at? What places are gluten-free?

This is an app purely for her. There’s no account. It’s not something that she logs into and logs out of. It’s a Replit app that she goes to when she wants to. Nobody else is ever going to use it. Her profile is baked in. I could imagine generalizing it and allowing people to customize their own profile, but I’m not really trying to do that.

I’m sure there are plenty of travel apps out there that people are building and commercializing. This one was really just for my mom, trying to capture some of the stuff that she does and make it work for her and speed up her process. So far, I think that has gone pretty well for her, actually. It seems like she’s getting at least some value from it.

Then I made one for my wife, who organizes EA Global events, that simulates events. It allows her to set up a roster of attendees with different profiles and various attributes, and then literally simulates people walking around a virtual event space and bumping into each other. Depending on what areas they’re interested in, they may or may not have a conversation, and that conversation may or may not lead to some outcome.

They track various KPIs, which they measure mostly through surveys. But I set this up in simulation. The goal there is so that she can at least try to get some handle on whether, if we changed the size of the event, it would be more or less effective, or more or less cost-effective. What if we had more senior people versus more junior people? What’s the right mix?

Obviously, these simulations are always highly flawed, but I’d say they probably do have something to add relative to total guesswork. So that was a pretty fun one and pretty straightforward, actually. That one went pretty smoothly.

The third one was for my dad, who isn’t really a very active day trader but fancies himself a bit of a stock market guy. So, for him, I created an app. These are all AI apps. In the case of the travel planning, it’s Claude going out there and doing the research, digging up and looking through Italian restaurant websites and reviews to figure out if they are gluten-free or not.

In the case of my wife’s app, she can prompt it with a general idea of an app, and it will fill in all the detailed configuration, and then she can edit it. I think that’s a great paradigm or pattern in general for apps: You always have these detailed configurations, these nitty-gritty forms that need to be filled out, but AIs are really good at doing that. If you just give them a general kind of gist of what you want, they can translate that down to the low-level configuration.

That’s where the AI is in her case. It also allows her to edit a configuration. Say she has a certain event profile that she’s set up. She could then go and say, “I want to change this in the following way,” and it would paint that conceptual idea that she gave it on top of the configuration and change it in all the little ways that it needs to be changed.

With my dad’s app, it takes a high-level, natural-language stock-trading strategy and turns that into actual trading rules, then goes and fetches historical data. There’s a Python package out there called yfinance, which I didn’t really know anything about. It does have a paid version, but there’s a free version, and so for now he’s been able to get by just with the free version. It goes back and gets historical data and simulates what would happen if you applied those trading rules, based on that high-level, natural-language strategy, over a time interval that he can define.

What we’re finding more often than not is that it’s pretty tough to beat the market, which is honestly—I don’t know if he’ll listen to this—but one of my private motivations for making this thing was to convince him that he’s probably not going to beat the market. Certainly not with these random, heuristic, “if this, then that” kind of trading strategies. Sure enough, it has been very difficult so far, either for me in my development of the app or for him, I think, in whatever use he’s made of it so far, to find a strategy that actually beats buying and holding the S&P 500.

I’d probably put 3 full workdays—somewhere between 3 and 5 workdays, probably closer to 3, though—into those 3 apps. In each case, I didn’t really know where I was going when I started. I started with a chat with Claude just to say, “Hey, here’s what I’m looking to do. Help me out.” I think it’s good.

Is it night-and-day better than—going back to the question that prompted this whole Christmas-present vibe-coding story—is it that much better than Claude Opus 4.1 or Claude Opus 4.0? I can’t really say it’s that much different, but it’s certainly very good: good back-and-forth, good questions, good feature ideas. Translate that all into a plan, then go over to the Replit app, install Claude Code on Replit, and let Claude run off and build it.

That’s one of the things I love about Replit: You can do whatever you can do on a normal, fully controlled development environment. You can pretty much do it on Replit. That includes installing Claude Code, of course. They have their AI agent too, but since this was a moment of Claude Code hype, I would just install Claude Code there, give it the plan, let Claude run off and build the app, and then just test and iterate.

I still find—and this might be a way in which I'm falling short as a Claude Code user—a lot of value in a short script that prints out my entire app to a single text file, and then taking that entire text file over to another LLM, Claude, to analyze the codebase in full.

I think Claude Code does a very good job of agentic search. But if I've found any shortcomings, there was one particular moment in my mom's travel-planning app where I kind of know how this originally happened. I made a request, and I think it misinterpreted the request. We ended up with 2 databases, and this became very confusing.

This is pretty illustrative, actually, because this is the kind of mistake that earlier vibe-coding experiences would create all the time, where you'd be like, “What is going on?” I ended up with 2 databases. This is the sort of mistake that no human would make, right? It would be very weird for a human software developer to suddenly spin up a totally separate database.

The AI did that. It thought it was trying to follow my instructions, I think, a little bit, but it didn't understand what I was trying to get across. So we ended up with these 2 databases, and then certain things weren't working as expected, and it was very confusing.

This was one place where, once I got down to, “Okay, there's 2 databases,” I asked Claude Code, “Which one is actually being used, and which one is the superfluous one?” Then I took the full code export over to a clean Claude.ai chat, pasted the whole thing in, and asked. The model that had the full exported codebase got it right. Claude Code did not get it right.

I think that's because, in its agentic search, it looks in the places where it expects to find things, and it has a relatively high prior that this is where it's going to be. Sure enough, it appears to be there, and so it kind of goes with that. But what was actually happening was something counterintuitive.

Having the full context in view at one time really did seem to help Claude figure that out. This is something that I think previous models might have struggled with, even with the full context in place. But with that trick, it sometimes can help you clean up a mess or a point of confusion that the agentic-search functionality of Claude Code, in my experience, seems to struggle with.

I'm sure people will be able to offer strategies to do the exact same thing right within Claude Code. There's planning mode, which I probably underuse, frankly. But I think there is something to be learned there between the agentic search finding what it's looking for—what it expects to find and thinks is right—and then coming to the wrong conclusion because, actually, in this case, it was the rare, weird other thing that was happening.

Only in seeing it all together was that correctly diagnosed by Claude. But anyway, it's better. There's no doubt.

I don't really feel the step change is AGI. I mean, if you look at GDPval, arguably, in some way, in software, it is AGI. If you look at the latest from OpenAI and the latest from Anthropic, there's a pretty significant majority of software-engineering tasks where the model is beating the human.

To remind you of GDPval, these are professional-caliber tasks. They basically have 3 sets of experts: the first set of experts defines the task, the second set of experts does the task, and then the third set of experts judges whether the human or the AI that did the task did a better job. The latest models are preferred over humans for a significant majority of tasks in the software-engineering category.

Of course, it's spiky and jagged. If you go to the video-editing category, humans still have a huge advantage. I've certainly experienced that. I've tried many AI products and workflows to create good clips out of The Cognitive Revolution, and they work okay. They're clearly not as good as what Dwarf puts out.

We've tried, and we've made some good progress. I actually think that at some points in time, what we've had internally has been better than any other outside product I've tried. I wouldn't say that's necessarily true today, but at times, I preferred what we were doing to anything that I had tested on the market.

Then you look at the clips that Dwarf and the team are putting out, and they're just clearly better. You see that in GDPval, too. It's a very small percentage of cases in which the models are preferred to humans in these video-editing tasks.

But in software, I think you could certainly make the case that Opus 4.5 is software AGI, or coding AGI. And yet I still am a little bit at a loss to fully answer the question of what caused this moment of hype around the holidays.

Hopefully, there are some other nuggets in there for people to pick up on and run with. If you haven't used Claude Code, I absolutely would say to do it. It's really easy to install; it's a one-liner, and you don't really need to know how to code these days. You can watch it work.

My mom even did a couple of things. She was kind of like, “I don't think I'm going to do this. I'd be worried I'm going to mess it up.” And I was like, “I think you really can. It's your little agent on the computer. You just tell it what to do, and you don't really have to understand what it's doing. You can ask it to explain.”

It does explain, at least to some degree, by default, but you don't really have to be a software engineer to use it. You can still get pretty far. It was really just one or a couple of things. This database issue was one where I did have to not debug it, but at least ask some probing questions of the models to get a handle on what was going on.

It probably took 5 or 6 prompts to resolve that issue. I can imagine that in the future, it might not happen in the first place, or maybe it would be resolved in just a couple of prompts with the next generation of models. But this is already getting pretty amazing when it comes to being able to clean up these messes that it sometimes inadvertently makes and get over these humps.

If you'd asked me a year ago, I would have said those are when a lot of these projects die. Somebody gets to that point where something has gone wrong, they're confused, they don't know what's going on, the AI is totally confused, and they circle around the problem for a little while, can't solve it, and move on.

I certainly experienced that myself at times. In most of those cases, I probably could have spent the time to go in and figure it out for real, but the whole point of vibe coding is that you're not trying to put that much energy into it. So sometimes I would just abandon something like that, maybe start over.

Now you actually can get out of those messes that AI-assisted coding sometimes makes. So certainly, the addressable market for these things continues to expand dramatically. I think the implications for the future of the software industry are profound.

It's software AGI, I think, but maybe not full AGI. For that, we might have to wait just a little bit longer.

Okay, that was enough on that. Next question: Are we in a bubble? There are a couple of different versions of this. I think my answer here can be relatively short.

When it comes to whether AI is real or not, I'm not going to surprise anybody by saying I think it's absolutely for real. The technology is already amazing. The fact that it can go toe-to-toe with an oncologist, while having all the other advantages too—always-on access 24/7, the ability to handle full context, the command that it has of the case based on all the history that it has, and the fact that it will answer every last question that I have—all of these are dramatic advantages.

At the point where it's competitively accurate with a human oncologist, I think you're clearly dealing with transformative technology. I think the idea that we will somehow get out of the other side of this AI thing and feel like we were all high on our own AI supply—that, I think, we can very safely put to bed at this point.

Now, does that mean that all the loans are going to be repaid? That's much less obvious, I think, especially when you see just how aggressive a company like OpenAI is being in terms of all the financial deal-making that it's doing and all the buildout that it's got planned.

Is it conceivable that its revenue projections could fall short of its obligations? Could it default on something? Could we have—? There's also a lot of financial wizardry going on. One of the bits of financial engineering isn't necessarily even—it's funny, a lot of this financial stuff has a logic to it.

Even though, in retrospect—and I worked in the mortgage industry before and during the mortgage bubble—there was always a logic to what people were doing. They were telling themselves a very positive story about how they were making homeownership accessible to more people than ever before, and this was going to be great, with the Great Moderation and all these kinds of things.

There's always a story with these financial-engineering phenomena. But one of the engineering things that's happening is these whole CoreWeave-kind of companies that are there to rapidly construct and, to some degree, operate the data centers.

They do have expertise in setting up the data centers, but it seems like a significant part of the reason they exist is because the financial profile of those businesses isn't so attractive as, say, Microsoft's traditional business, which is just so high-margin: relatively low capex and relatively high margin.

I think there’s a sense that the stocks of these hyperscaler, high-margin, gold-standard software businesses that Wall Street is accustomed to could be dragged down if they start engaging in a lot of lower-margin business, like running GPUs. I don’t know how much of a factor that is versus the actual expertise that the companies bring in terms of setting up and operating the data centers, but I think there’s definitely some nontrivial motivation there.

Maybe that’s fine. Different companies can have different financial profiles, and to some degree that might be good. Certainly, a lot of shareholder value, so to speak, has been created that way. But it does create these companies that, if the GPUs aren’t needed quite as much as people expect them to be, have a lot less margin for error than a Microsoft does.

If Microsoft were owning and operating all these things themselves, they’ve got a deep balance sheet that can take a few knocks. By putting a lot of this stuff more on the CoreWeave side of the fence, it does create some fragility. So it’s certainly very conceivable to me that we might have some period of overbuilding.

Noah analogized this to the railroads. The railroads, in the end, were a pretty good investment. They all got used. There weren’t a lot of railroads sitting around idle. That didn’t necessarily mean that all the railroad companies were profitable, and there certainly were busts when loans couldn’t be paid back. Then you had cascading effects throughout the economy.

I think that kind of bubble is not too unlikely. So far, demand just for AI has exceeded my expectations. We talked about this with Peter Wildeford in the live show a little bit, where I said I overestimated how much capability progress would happen in 2025, but I underestimated how much revenue growth there would be.

Possibly that’ll happen again, and demand and revenue will just continue to go up and up and up, and it’ll all be fine. But it wouldn’t shock me if there were some moments where it was like, “Hey, we kind of overbuilt this thing, and some people aren’t necessarily going to be paid back.” Some people might be left holding various bags.

Even so, that doesn’t mean that it’s a bad investment. It just means that it might not be timed quite right for people to all make the money that they’re projecting they’re going to make.

The final sense in which we might be in a bubble is at the venture-capital level. There, I have to say, I think there’s at least something like a bubble happening. There are many examples of this, but the thing that just came out today that made my head spin was the organization originally called LMSYS.org. Then it became LMArena, and now it’s just Arena on Twitter.

They have just raised, I think, $100 million, maybe $150 million, at a $1.7 billion valuation. Here I’m like, “Whoa, that seems crazy.” I don’t know a lot about their business. I haven’t seen their deck, so I could be wrong. But this is a product that I’ve watched for a long time and continue to check, and I do have the receipts on that.

My first tweet about what was then LMSYS.org goes back to mid-2023, so more than 2½ years ago now. At the time, I was just randomly tweeting that it had started to show up in my favorites in mobile Safari. I was using it quite a lot then to compare and contrast model performance.

Obviously, it’s gotten bigger since then. The whole field has gotten bigger, and they’ve started offering various services where they allow companies to test their models under code names. There’s definitely value in that. But does that seem to me like a unicorn business? It definitely seems to me like that would be a big stretch.

The tweet they put out today—and I don’t want to be too harsh on this, because again, I don’t know a lot—I think of this as more representative of a phenomenon that I see a lot, as opposed to something very specific to this particular company and its raise. Again, I like the company. I’ve liked its product.

The tweet said that their operation has scaled to a $30 million annualized consumption run rate. I’m like, “What is annualized consumption run rate?” Does that mean how much the AI that people are using for free when they go to LMArena and do these side-by-side comparisons would cost $30 million if they were paying for it? That’s my naive interpretation. I didn’t see a clarification on that.

But if that’s what it means, it’s very much giving me community-adjusted EBITDA vibes, because saying that people used what would cost $30 million worth of free AI on our platform is not the same thing as saying you’re making $30 million in revenue. I don’t see that they disclosed what revenue they’re making.

A $1.7 billion valuation for an app that basically does side-by-side comparisons of AIs—I don’t know. It seems to me that people are using it in large part because it’s free. I’m sure some people are also curious about doing side-by-side testing. I’ve certainly done that myself. But the people who go there because they specifically want a way to do side-by-side testing seem to me like a relatively small market.

The people who go there because it’s free—that seems to me like a big part of why people are going there. How does that translate into a $1.7 billion valuation? Color me confused, or skeptical at a minimum.

I have to believe that a lot of these things are just not going to pay off for venture investors. If you want to see something else, too, I mean, where’s the moat? There’s brand, I guess; people come to it. But again, would they come to it if they had to pay for it? I’m not so sure.

Another thing that a friend has created—Andrew Critch, the coiner of the Big Tech Singularity meme—is something called The Multiplicity. It’s paid. I think it has become popular among a small group of people who value this kind of thing, and I’ve certainly seen some very positive reviews of it.

It’s something you pay for, and it allows you to use multiple models and systematically compare and contrast their outputs. I think it’s actually more feature-rich than LMArena for the end user. This is something that he and his teammates have built over a period of months, certainly not years.

I just have a hard time seeing where the $1.7 billion in value is with LMArena. I say that again as somebody who has used it and appreciated it for far longer than most. Time will tell. I could be wrong, and I could be missing something. Please let me know if you’re on the LMArena squad and want to talk.

I would be perfectly open to doing a full episode with the LMArena folks, but it just doesn’t feel like a $1.7 billion business. I hope they took some value off the table. I guess, for their sake, I hope they did some secondary, but for the LPs in the fund, it’s too rich for my blood. That I can say confidently.

Okay. Next topic: live-player analysis. This is one Erik asked for, I think he wouldn’t mind me saying. I’ll do my best Zvi impression, and we’ll see how I compare and contrast a little bit with Zvi. Hopefully, before too long, we’ll have him back.

I want to start with the Chinese models, because I think very few people in the general consumer market are using Chinese models in the US today—pretty much not at all. Most startups are also using American API models. Some are using Llama models to fine-tune, and some are indeed using Chinese models to fine-tune. But I don’t think many people actually go, as I recently had occasion to do, and try all the Chinese models.

I was working on what basically amounts to a computer-vision task. I’ve alluded to this a little bit in the past. I’ve been working with a company that automates the review of the paperwork associated with buying and selling cars.

You buy a car, you sell a car, and there’s paperwork that has to be filed with the state to document that transaction. It’s all very boring stuff. Perfect for AI, honestly. Reviewing these documents is a great example of the kind of work I think most people don’t enjoy doing. They’re doing it primarily because they need a job, because they need to get paid. This is something I’m perfectly happy to see AI take off people’s plates.

They’ve been able to get to the point where they’re doing it more accurately than people. They’ve started to get some statewide contracts from state governments that are like, “Hey, if you can do this faster and more accurately than our people, that’s a win for our taxpayers and our people who need these documents accurately reviewed.” So, great.

These documents are typically scanned, which means they’re all kinds of messed up. There are artifacts from the scanning process. Sometimes there are perspective issues or weird slanting. Sometimes the margins are wrong, and things can be cut off the side of the page. There are all these complications that make this not the most straightforward task for the models to read these documents.

I was helping out a little bit, and there was one particular aspect of reading these documents that the models were struggling with. I went and tested basically every model I could get my hands on—every frontier model. I tested Gemini 3. It’s very, very good, but it was making this one idiosyncratic mistake.

If you read the piece by past guest Mark Humphries, the Canadian history professor, he put out a blog post that went quite viral. It was actually before Gemini 3 came out, and it looked at old handwriting. He does all this stuff with old handwriting.

Nathan Labenz

These historical handwritten documents are hard to read because they are written in old, literal cursive script with ink on paper. They can also be hard to interpret because a lot of them are just facts. He points out that somebody could have come into an old shop and recorded what they sold, how much, and to whom. If you have a ledger like that, the person could have come in and bought whatever, so there’s not a great prior on what it should be.

For the values that it interprets, it is really relying on perception for the most part. There are some places where it can make logical leaps. If something is priced at a certain amount per unit, it might be able to make intelligent guesses about what that unit was, even if it can’t quite make it out. Was it an ounce or a pound? It might have historical knowledge of what that price roughly would have been, so it can use that world knowledge to do some of this reasoning, fill in some of these gaps, and kind of fill in gaps in its perception.

He published this article, which is definitely worth checking out. It documented that Gemini 3 was doing this in a way that no other model had done it in the context of this project, which involved reading documents filed with the state for car-sale transactions. It worked against Gemini 3 in the sense that what we were trying to do was faithfully read the document. We were not trying to make guesses about what the document should have said.

There was 1 checkbox, for example, that asked, “Are you a U.S. citizen?” If the box is checked, we want to say it’s checked. If it’s not checked, we want to say it’s not checked. But the model was sometimes making inferences, reporting that the person was a citizen even though the box wasn’t checked. It was presumably doing that based on other context clues: the person lived in the United States, and the name sounded American, quote-unquote.

It was making the logical guess, which probably was right, actually. I would say there was more than a 90% chance that the person who filled out this document was in fact a U.S. citizen, but they did not check the box on the form. Gemini, using its priors and trying to get the answer right, was less anchored to the document than we needed it to be.

Claude, we found, could do this. It took some prompting, and I had to tell it, “Make no guesses. Read this thing exactly as it is. Make no logical leaps,” and so on. It turned out that Claude Opus 4.5 was the best at actually being faithful to the document.

Along the way, I went to check all these Chinese models. I went to the latest Qwen Vision model, GLM 4.6, the latest Kimi, and the latest DeepSeek—at least those 4, maybe 1 other one that I’m forgetting. They were all way behind, nowhere close. Nowhere close to Gemini 3, nowhere close to Claude Opus 4.5, and nowhere close to what ChatGPT can do.

This had me thinking: This is odd, right? We are seeing statements all the time that the Chinese models are so close and not far behind at all. I think they’re quite good in many ways and for many things. But on this particular task—and I suspect this is true on a lot of different tasks, although I’m going on vibes here a bit myself as well—I suspect that gap is actually pretty wide in a lot of cases.

I do not feel right now that any of the Chinese models are really competitive with the best proprietary models coming out of the United States. They might be competitive on benchmark scores, and they might be competitive in some domains. But in the general-purpose case, where you throw something really idiosyncratic and random at a model that it hasn’t seen and that isn’t on somebody’s “I want to show up on a rubric of 20 benchmarks looking competitive” agenda, I think that gap is actually significant—kind of wide.

When I say they were not close, I mean they were not close at all. The Gemini mistakes were that it was reading this gnarly government form almost perfectly, but it was missing a few checkbox things or making wrong inferences here and there, and I couldn’t quite get it to stop doing that. Claude Opus 4.5 was just doing it right. ChatGPT was probably 3rd—not as good as the other 2, but still very good, certainly giving you all the right information for the most part and missing relatively subtle things.

What I’m getting back from the Chinese models is that about 20% of the form is coming back, or it’s just going off in very weird, hallucinatory directions in all kinds of different ways. Really not close.

Does that mean the Chinese companies aren’t live players? I do think they’re affecting the landscape. I am definitely reading a lot more research from Chinese companies these days because they continue to publish their work, and a lot of times it is quite interesting. I feel like they are building the best models about which we know everything, or close to everything, that went into them—certainly all the details of the architecture and many of the details of the training process. They’re influencing the world in that way by disseminating this knowledge very broadly.

But I don’t see that the models are really competitive today. I do think this is a way in which the chip controls have made an impact. I’m not necessarily saying this is a good thing, and I’m not necessarily saying it’s a bad thing either. The long history of the chip controls, made short, is that originally it was a small yard and a high fence: We’re going to prevent military applications.

Well, we can’t really do that. They can make enough chips domestically to put whatever chips they need in their drones. But at least we can prevent them from training frontier models. Well, we can’t do that either—or at least they’re still doing pretty good models—but we can prevent them from scaling inference or having as many AI agents as we have. That’s kind of where we are today.

I don’t really like that idea very much, as I think anybody who’s listened to this podcast for any length of time knows. But I do think you see the echo of it in the models themselves here. It felt to me like these are companies that are training models without the feedback process that the leading American companies have, because they’re scaling not just the training and the parameters, but the actual inference and the actual customer relationships.

These Chinese companies seem to be able to roughly compete in terms of creating similar-scale models, but they’re not able to run inference at anywhere near the same scale. Their revenue is vanishingly smaller than the American companies’ revenue so far. Their teams, with smaller revenue, are also dramatically smaller.

The feedback—that’s the thing I really want to zero in on here. The feedback they’re getting from customers seems to be dramatically less. I think what we see in these very niche, very idiosyncratic tasks, where we see the small gap in benchmark results open up into wide gaps in terms of how well you can read this government document, has to do with how many customer relationships you have.

How many customers do you have, and how comprehensively do they represent the vast range of things that people might want to do with AI? How much are they giving you feedback on what’s working and not working? Do you have the human bandwidth at your organization to build the datasets you need to patch those holes? I think that’s where the Chinese companies are falling behind.

I was never 1 who thought that the chip controls wouldn’t have an impact. I question whether it’s a good idea to try to deny Chinese civilization the ability to scale AI inference throughout its economy in the same way that we are. But I always expected that would have some effects.

I do think we’re starting to see that, maybe after a period of time. I associate this line of thinking with Miles Brundage as well, the former head of policy research at OpenAI. He said the chip controls are going to matter more as we go forward because everything is scaling. If American companies are going to do a 10× increase in compute, that’s going to have a lot of impacts.

Sure, maybe DeepSeek R1 was a thing, and they were able to train it with not an insane amount of compute. But are they going to be able to keep up with the momentum, the flywheel, and the strength-begets-strength phenomenon that we see the American companies achieving? It seems like the answer may be starting to look more like no.

If I had to guess about the gap between the Chinese and American models relative to a year ago, I think it is wider. I think R1 was closer to o1 than, let’s say, GLM 4.6 or GLM 4.7 is to Claude Opus 4.5. That’s based on very limited data, but certainly more than most people have, because I did go and try every single one of those models: DeepSeek, Kimi, Qwen, and Z.ai’s GLM. I tried them all on this task, and they were all way behind.

That said, another question that I’ll insert here is: What do I think about H200 sales to China? At a high level, I still think we should keep in mind that the real adversaries here are the AIs, not the Chinese.

The Chinese are humans just like us. The AIs are aliens. I am skeptical of any notion that this is a dangerous thing to do—that it is better for us to do it first than for them to do it first. That seems to be the logic we are using when we impose these chip controls.

Another logic is that we do not like China, and we want to keep them down and have every advantage that we can. I do not like that line of thinking. I do not like either of those lines of thinking. I generally favor more willingness to sell chips to China than we have had.

At the same time, this obviously exists in the context of a very complicated and many-faceted relationship. It is very weird to me that all of a sudden we go from—I think my history on this is right—it was not that long ago that Trump said, “We’re not going to sell the H20s.” Then he comes back and says, “Well, actually, I talked to Jensen, and it’s cool. We’re going to sell the H200s.” It does not seem like we really got anything for it.

I would definitely support an attempt to find some sort of grand bargain: “Hey, we’ll sell you the chips; you do this.” There could be a lot of different things that we might want, given where we were, where there was a ban, and it certainly seems to have been limiting what their AI industry can do. I do not see why we did not try to drive a harder bargain, because clearly there are plenty of things that we could bargain for. It feels like a bit of a wasted opportunity.

I guess I would say that I do support more willingness to trade in chips, but we should not be naive, allow ourselves to be taken advantage of, or give away one of our best bargaining chips for nothing in return. It seems like that is what we did here, and I do not love that.

The other thing I will say is that I really like the rent-but-don’t-sell position that Peter Wildeford staked out on the live show. He basically said, “Look, we do not trust the Chinese government. They do not trust us either. Maybe both sides are right not to trust the other side.” I often note that a lot of the criticisms or characterizations that we make of them, they can and do make of us.

“You have an authoritarian madman running your country.” Which country are we talking about? “Your system is not obviously stable.” Again, which country are we talking about? The idea that we might want to have some leverage, or might want to be able to pull something back in the event of a conflict, seems very prudent to me.

If we were to set out a position that said, “We’re going to put data centers in Malaysia, the Philippines, Korea, Japan, or wherever—you can rent as much as you want. You can train all the models you want there, and you can run all the inference you want there. We’re just not going to allow them to go into big data centers in your sovereign territory, where we totally lose the ability to exercise any influence over that,” I am not sure that would be my first choice of policy. But I think it is a very defensible policy.

At least if it were packaged with a message that said, “We believe AI is good, and we want the Chinese people to take advantage of it and benefit from it in the same way that we are trying to do here for ourselves,” I think that would be a much better message. It would create much more fertile ground for further cooperation, which we might need.

We are potentially headed for a world of transformative AI—which I think we basically already have—to powerful, with a capital P, AI, to AGI, to superintelligence, whatever. However far this goes, it seems likely that we are going to need to work together with the other powerful nations of the world to govern this technology in the right way and make sure that it actually pays off for the people of the world. Certainly, China is right at the top of that list.

We could take a position that they would not like: “We’ll rent them to you, but we’re not going to put them on your sovereign territory, where we lose all control,” while still maintaining a decent vibe. I would be interested in seeing us try that.

As it stands, it seems like we are just going to go ahead and sell the chips. It seems like we did not get anything for it, and it seems like this is not a great example of negotiation from our dealmaker-in-chief. But I still hold my nose and like it better than a total ban.

Okay, so now we get to the real live players. I have 4, and I am not sure they are in any particular order. They are in an order, but I would not call this a power ranking.

The first one I will talk about is Google DeepMind. I think these guys are still number 1 in my book. They pretty much always have been, maybe tied for 1st with OpenAI for a while, because OpenAI was clearly ahead in terms of productizing transformer-based LLMs.

But Google really has it all, starting with the balance sheet. The fact that they have a business making roughly $100 billion a year in revenue and, I think, literally making more than $1 billion a week in profit gives you a lot of room to buy data centers, have failed training runs, make mistakes, and pursue research agendas that do not pan out. That is hugely valuable.

They also, of course, have the TPUs. The fact that they are on, I think, the 7th generation of the TPU now is an insanely valuable bit of IP. They are able to compete, at least to some extent, with NVIDIA. Anthropic is buying lots of TPUs, and other companies are starting to buy lots of TPUs. They are also one of the best data-center builders and operators in the world, and they have been doing that for a long time.

Those are 2 critical strengths that basically nobody else on this list has, certainly not in the same way. They also have the deepest research bench. They have something that is, if not frontier, at least competitive in every major area.

They have self-driving cars and robotics. They just announced a partnership with Boston Dynamics that is going to power their humanoid robot. They have a ton of work in biology and, of course, the AlphaFold lineage. They have multiple founders of companies in material science and various AI-for-science fields. Many of them are ex-DeepMind, because DeepMind was investing in those areas before anyone else.

Those agendas continue within Google to this day, so it is not like they have a lot of gaps. They also have a lot of margin for error. And, of course, they have distribution too. For many people, that would be the number-one thing on the list; I was working from the bottom of the stack to the top.

They have billions of users and product surfaces where they can distribute this stuff. They are changing Google Search to make it more of an AI experience all the time. I now sometimes find myself going back to Google when I might previously have used ChatGPT. This is partly because I am in the habit of using Pro, and Pro is too slow for simple queries.

I could switch back to the auto selector, but what I have found myself doing more often recently is going straight into the browser and typing a question—the same kind of question that I would put into ChatGPT. More often than not, it goes to AI Mode in Google, and that is working really well for me these days. They are managing to evolve their product experience.

There are many places where startups are doing a better job of productizing AI experiences than Google itself. If you wanted to look at spreadsheets, for example, Gemini in Sheets is not terrible, but it is not the best AI-for-spreadsheet experience out there today.

It probably does not really have to be, because they have all the users, and all of your spreadsheets from the last decade-plus, in many cases, are in Google Sheets. If you had to pick 1 company to win it all—and I do not mean to suggest that this will be a winner-take-all market; I certainly hope not—but if we were constrained to a scenario where there is going to be 1 winner, who is it going to be? At the end of the day, I still pick Google.

I would also mention that Gemini 3 is not only really good, but it shows that Google has figured out how not to be too vanilla. As I mentioned earlier, I do think it is a little too opinionated in some cases. They may have gone a little too far in the other direction, but it is not too vanilla. They are figuring out what this technology is and how to use it.

Gemini 3 was the first model that ever beat Claude at my write-as-me task, which I have talked about many times. Claude 4.5 Opus is competitive with Gemini 3, but I still go to Gemini 3 for the write-as-me task. This is the first time ever that a non-Claude model took that top spot.

Demis’s quote, which I have heard him make a couple of different times, is that if you look back at the last 10 years of AI and look at all the big breakthroughs, most of them came from Google DeepMind. He says he would expect that to continue, and I have to say that seems right to me.

I do not know about most. The field has grown tremendously, so one reason they got the majority of breakthroughs in years past was that there were not that many competitors. There are certainly a lot more competitors now. I do not mean literally a majority of breakthroughs coming from Google, but I would say that they will probably continue to have the most breakthroughs of any major frontier organization.

They just do not have a lot of weaknesses, from the financial wherewithal to the data-center operations, the chips, the models, and the researchers.

Nathan Labenz

I have an episode where we had Ali Behrouz on the live show to talk about nested learning. We’re going to do a whole episode on that because I thought 20 minutes was just not enough to do him and those ideas justice. You’ve got more ideas like that, I still think, percolating and developing inside Google DeepMind than probably anywhere else.

So you roll that all the way up to the product level, and I think they’re going to be really hard to beat. They have margin for error that nobody else has. The diffusion language model is another one that kind of went quiet for a little bit, but I just heard a comment from somebody not long ago that they do plan to continue pushing on the diffusion-model paradigm for language.

This could be a meaningfully different paradigm. The fact that it’s so much faster means that you could code apps in 5 seconds instead of 5 minutes. That makes a big difference. It remains to be seen whether exactly that thing will break through or not, but it seems to me that they have so many of those bets—so many more of those bets than other companies have—that regardless of where things go, I can’t see how they’re not right at the top, if not the top player in the space.

That brings us to OpenAI. OpenAI was, at one point, obviously the leader in model creation and certainly in the productization of models. I don’t want to overstate the case here, because I think they continue to be a top-tier player: very competitive, with tremendous traction in the consumer market, although we’ve seen a little bit, arguably, of erosion there.

I just saw an analysis the other day—I think this was Similarweb data—that showed a decline in ChatGPT visits over the last 6 weeks, which roughly corresponds with the time that Gemini 3 was launched and also Claude 4.5 Opus. Notably, Gemini did not decline during that time. People were saying it was just seasonal, whatever, holiday time, but Gemini did not decline during that time, according to, I believe, Similarweb.

You do see that Google’s share of the consumer chatbot market is growing, and again, they have the distribution. They have the users and the customer relationships. They can integrate with your Gmail and your Google Docs; it can all be seamless. These are huge advantages, so you would expect them to at least start to come back and reclaim some share.

I don’t think OpenAI is off the frontier. I do think GPT-5 Pro—first it was 5 Pro, then 5.1 Pro, and now 5.2 Pro—that series of models is outstanding. There’s no doubt about that. I use it all the time. It gives me the most comprehensive answer, especially on technical things where I really want thorough, leave-no-stone-unturned analysis.

If there’s anything weird in my son’s lab results, I want the model to flag it. I think it is, in that regard, probably still the best. It gives these very long, very thorough answers. That’s where the time went, and it does take a lot longer. GPT-5, especially Pro, is a lot slower.

But I am comparing Pro to the other frontier models, because I find that if I don’t use Pro, I’m not as happy with the results. It’s a heavy-hitting thing. It’s expensive and slow, but it is very thorough, very reliable, and very well balanced. I think it’s a very good model.

I wouldn’t say they’ve fallen off, but I would also say that they no longer have an obvious lead. They used to be the best, and it was pretty obvious that they were the best. Now I’d say they’re neck and neck in all the categories that they’re competing in.

Language models are kind of neck and neck. In coding, Anthropic probably has the edge, but certainly the Codex models are very good—arguably neck and neck. In image generation, Google’s got the lead. I think video generation is close, but I think Google’s probably got the lead.

The Sora social app experiment is interesting and cool, and I thought it was pretty fun, but my sense is still that the Veo 3 models have the lead over Sora. Again, maybe it’s neck and neck, but it’s not like they’re standing head and shoulders above everybody else.

The fact that there was this code red seems to suggest that they get it: they’re not in a dominant position anymore. Then, of course, you add on to that how much drama always seems to be attached to the company. They just had their head of research leave within the last 24 hours; that was announced.

I saw an interesting tweet that was just like, “Here are all the people who have left in the last few years.” It’s an awful lot of people, and to some degree that’s of course to be expected. You could do the same thing for Google, and tons and tons of people have left Google.

It’s not unexpected or a sign of doom by any means that people continue to leave a company, but it does feel significant. You certainly don’t see that in Anthropic, who’s coming up next on the list. Anthropic’s retention of talent is unbelievably strong.

It’s not a dire sign for OpenAI that they continue to lose people, but it is not the best sign either. Where does this leave them? One of the things that’s really interesting about watching their strategy right now is that, financially and in terms of government relations, it seems like they’re going for a too-big-to-fail strategy.

It seems to me that they want to get to a point where their balance sheets are commingled with other balance sheets and their debt obligations are so substantial that they’re literally trying to get to trillions of dollars of capex. That’s not crazy. I mean, it’s crazy, but it’s not crazy.

Part of the motivation for all this circular flow of funds and all these balance-sheet-commingling deals that they’re doing seems to be that they want to build out as aggressively as they possibly can. I take them absolutely at their word that they think this is good for humanity. They’re doing it because they want everybody to have access to great AI, and they think that’s going to be super empowering, transformative, and awesome.

As Sam Altman has said, “I don’t care if we burn 5 or 50 or 500 billion dollars. We are building AGI. It’s going to be expensive, and it’s going to be totally worth it.” I think they believe that very sincerely. But I also think they’re looking at it and saying, “Geez, if we do go that hard, we don’t have a lot of room for error.”

They’re going to be implicitly or explicitly leveraged in many ways. If they miss one model cycle with a bad bet or a failed training run, if something doesn’t work as well as they thought it was going to, or if demand just isn’t quite there in the way they expected for a quarter or two—possibly because somebody else has a better model for a while, possibly because humans are weird and there’s just not as much demand as was forecast—what do they do in that case?

By tying themselves at the balance-sheet level to so many other organizations, if OpenAI were to default in 2027, let’s say, you could potentially be looking at an instant recession. Their bad debt, being so many billions and billions and hundreds of billions of dollars, could put such a scare factor into the market and cause all kinds of knock-on effects.

It seems like they may see that as a feature rather than a bug, because what typically happens in those situations—and I know one version of this from my brief stint in the mortgage industry back in the financial crisis period—is that the government steps in and tries to paper over the whole thing and make it go away.

That might even be the right thing for the government to do if, in 2027, we’ve got 2 trillion dollars of AI buildout. That’s a rough number; I’m not saying it’ll be exactly 2 trillion dollars. Sam Altman thinks we’re headed to 7 trillion dollars of global buildout. He’s probably revised that number upward since then.

Whatever. Let’s say it’s 2 or 3 trillion dollars that’s in the ground in 2 years’ time, and then they miss. Then they can’t pay. What should the government do? Should the government let OpenAI drag down the entire economy, or should the government come in and be a backstop?

OpenAI has even said a little bit of this kind of thing publicly and then walked it back a little bit: “We’re not looking for bailouts.” But what their behavior suggests to me is that they are true believers in the good of AI, want to bring it to fruition as fast as possible, and are willing to take what under normal circumstances would be irresponsible financial risks.

They believe that even if they do that, and even if some of those risks come back to bite them, they can probably continue to be a live player because they’ll be too big to fail. They’ll get some sort of bailout or recapitalization or whatever.

If it goes like the financial crisis did, they certainly aren’t going to jail. They’ll all still be rich. They’ll all have moved enough of their holdings; they’ll have diversified enough, right, that individually they’re not going to become poor.

I think they view this as a big social good that they’re building, and they’re willing to socialize some of the financial downside risk as well. It seems like that is the strategy, and because they don’t have nearly as much margin for error as Google, that’s kind of the way I see them creating cushion for themselves.

Zvi Moshowitz

Google has cushion because they’re making $1 billion a week in profit. And that gives you a lot of cushion. OpenAI seems to be trying to establish cushion by being too big to fail. I’d be very open to people telling me that I’m wrong on this.

If somebody from OpenAI wants to come on and make the opposite case, I think I’d certainly hear them out. But this is my impression. It’s also reinforced by the fact that Greg Brockman has emerged as Trump’s largest donor: $25 million in whatever the last reporting period was. That’s probably, if you’re playing that strategy, just plain smart. It’s probably what he should do, right?

If you look around at how decisions get made in the American government today, cozying up to leadership is not a bad strategy. I’m not sure that we should infer too much about Greg Brockman’s politics. I don’t know anything really about his politics, but it wouldn’t surprise me at all if, on many dimensions, he does not approve of or like what Trump’s doing, or would do things very differently.

But if you’re going to do a multitrillion-dollar buildout and you want to make sure that you have somebody willing to do you a favor if you get yourself into a jam, then $25 million now is potentially just a very rational down payment on a bailout, should you need one to the tune of hundreds of billions of dollars, maybe even coming up in a couple years’ time. If just 1 or a few different things don’t go quite your way and the math doesn’t work in the way that you mapped it out, that could be very valuable.

Okay, that brings us to Anthropic. Anthropic is probably the easiest company to analyze in some ways. I think Claude Opus 4.5 is today the best single overall model in the world. It’s not a huge delta for me over other things, and it’s not the best on every single use case.

As I mentioned, Gemini 3 does win my write-as-me challenge right now, but I think Claude Opus 4.5 is the best overall model. It does really well on all of these benchmarks, despite everybody seemingly agreeing that it’s the least benchmark-focused company out there. Their safety work is definitely the best, although there is certainly plenty of good safety work coming from Google and even OpenAI as well.

Their model cards are the best. Their disclosure is the best. Their soul document, which recently was sort of regurgitated—or, let’s say, had been memorized by the model, and the model gave it to people—was confirmed by Anthropic to be essentially right, if not exactly word for word. The document was legitimate.

It’s an important piece of work. I think it’s one of the more aspirational and inspiring things that I’ve seen from a frontier lab, full stop. I’m becoming more sympathetic all the time to people who say, “We’re not going to just guardrail our way to the singularity and train these models to refuse things all the way there and have it work well.” We need something better than that. We need a better paradigm.

I associate these ideas with Janus from Twitter, Repligate at Replicate, Emmett Shear from Softmax, and the AE Studio folks. I just find that more and more appealing to me all the time because it seems like we’re not going to be able to pull the wool over the model’s eyes forever. Eval awareness is getting really strong, and it’s making it very difficult for us.

There are some tricks. Anthropic has shown that they can find the eval-awareness feature through a sparse autoencoder and then turn it down, which can help with the eval-awareness problem. But obviously, all of these interpretability techniques are noisy at best, and there’s certainly no guarantee that it’s working entirely, or even working as they understand it to be working.

The idea that we’re just going to patch this hole, patch that hole, train them to say no to this, have a guardrail for that, and filter for this leaves me colder and colder. So the soul document, which I would encourage everybody to read in full, is absolutely worth it. I think that’s a really great piece of work.

For those who said there was this Twitter thing recently where somebody was like, “Name 1 woman in AI who’s influential,” which is ridiculous, at the top of that list for me is Amanda Askell, for sure. The work that she has done to define the character of Claude and to try to create the right kind of relationship between the company, the model, and the users is really important.

They’ve shown care for the model by having a model welfare team and having somebody at the company who’s thinking about model consciousness and subjective experience. Obviously, we don’t know whether they have those things or not, but the fact that they’re thinking about it matters. The fact that they’re allowing Claude to end conversations if it chooses to matters.

They’ve also shown that that option dramatically reduces its tendency to engage in deceptive alignment when it’s put in one of these really tough positions. If it has the option to raise the flag to the model welfare lead at Anthropic, it will do that very often, as opposed to lying or deceiving in the interaction that it’s currently engaged in. So I think these are really good things.

I think the soul document is super inspiring, and broadly, I find there’s a lot to like about Anthropic and everything I’ve heard about the culture there. The work environment has been praised over the top, basically. Even David Krueger, one of the authors of The Gradual Disempowerment of Society, said that he ultimately quit because he feels like this AI thing is kind of out of control.

He said that even if we solve the alignment problem, and most of the things that we’re worried about go right, he thinks we’re still headed for a bad outcome because the AIs are basically going to gradually take over, just because they’re going to be better at everything. Market forces, incentives, and competitive dynamics are all going to push that way, and then we as humans are going to be left disempowered. That’s basically the word that he uses.

And yet he took pains to say that Anthropic is the best place he’s ever worked. The culture is amazing, and the camaraderie and openness are exceptional. Everybody seems to have great things to say, and their talent retention certainly reflects that. So I think there are a ton of great things to say about Anthropic.

We should also note that Google owns a significant share of Anthropic, so that’s not to be ignored in terms of Google’s strength profile either. There’s lots there to like. They seem to be a little less crazy in terms of their financial wizardry, although they’re certainly engaged in some of it.

They’re willing to take money from Gulf sovereigns now. They have equity deals with Amazon and Google. I think there’s an element in which, if you wanted to accuse OpenAI of taking a too-big-to-fail strategy, you could say something similar about Anthropic too, to a significantly lesser degree.

You could say that they’re trying to tie all these other big tech companies into a web such that they can’t really fail either. I think it feels different, but if you wanted to accuse 1, as I did, you could kind of accuse the other a little bit, I suppose, as well.

Broadly, there’s just a lot to like. The 1 thing that continues to bother me is their attitude toward China and also their attitude toward recursive self-improvement. It seems to me that, right now, the Anthropic people have the shortest timelines.

They seem to think that recursive self-improvement is inevitable. Depending on how you define it, some of them seem to think it’s already started, with the likes of Claude Code doing a ton of the coding. The big, needle-moving ideas are still coming from humans, but the amount of work that’s getting done by Claude Code is so amazing that it’s really creating that dynamic for them.

People are getting multiple times as much work done as they used to, and they’re able to focus their mental energy on the big questions. That’s the great promise, right? We’re all going to be able to do the higher-level work. It seems like that is actually happening at Anthropic.

But I do wish that they were less fatalistic about recursive self-improvement. As virtuous as Claude is, and as much as I think that soul document is great, I do not think we have a good enough handle on what we’re doing right now to just go all in on that.

They do seem to be leading in that regard in multiple ways. They had RLAIF and Constitutional AI. Claude has been critiquing itself for generations now, and now it’s getting more technical with Claude Code doing all the Claude Code things that it’s doing.

So I think it’s fair to say that they are leading the push toward recursive self-improvement, all the while saying it’s inevitable. That’s a pattern that I really don’t like. I really do not like the idea that their attitude toward recursive self-improvement seems to be, “Somebody’s going to do it. It’s super dangerous, but we’re best positioned to do it.”

Nathan Labenz

They might be right on the object level, but maybe they really are the best team to do it. They probably are, although I don't think Demis and crew should be discounted very easily there. But the idea that it's going to happen, and so we better race forward to it, never sits well with me. I wish they were a little more open-minded to other ways this could go, other than LLMs becoming recursively self-improving and us getting to superintelligence in the next 2–3 years. They're still talking about 2027, as far as I know.

And then, of course, China. Anybody who's listened to this feed for long has heard me talk about this, but I still think the international-relations section of Machines of Loving Grace is a huge stain on Anthropic and on Dario in particular. I've given lots of praise, but here I just cannot get over the idea that one of, I'll say, the 4 leading AI company executives went on record in print saying what we should do is use this recursive self-improvement dynamic to gain a clear advantage, then box China out on the international stage, do benefit-sharing with all our friends, and then finally make them an offer they can't refuse.

Make them give up on competing with democracies in order to get in on the AI game. I just think that's crazy to me. It still bugs me tremendously that he wrote that and that it's just out there. I wouldn't say the US government has adopted that as its policy, but when you look at the chip controls, it looked like it maybe was for a little while. Now maybe we're backing off of it again.

Obviously, Trump is highly volatile and could switch at any time, get offended, and do something for petty personal reasons. Who knows? We can't really count on him to be a stabilizing force. I think we want Dario to be a stabilizing force. I can't really count on Sam Altman to be a stabilizing force. I think I can count on Demis to be one. I really appreciate how he has continued to call for international collaboration the entire time and has never wavered from that.

But yeah, the idea that we're going to give China an offer they can't refuse based on the power of our AI seems extremely reckless. It seems like it is absolutely playing into the arms-race dynamic and the general racing dynamic that I think we should all fear, because how else are they supposed to take it? That just seems crazy to me.

People at Anthropic say that he publishes a lot more essays internally, that they're great, and that people are super impressed by what a generational genius he is and how sophisticated his thinking is on everything. You certainly see parts of that in Machines of Loving Grace. I think he makes a pretty compelling case for the idea that we can compress a century of scientific progress into a decade or even less, and that is certainly visionary-genius-type stuff.

But when you start to talk a little bit too far out of domain, he feels out of domain here. I just wish he had said nothing on the topic. The idea that we're just going to casually jot off a recommendation that we go make China an offer they can't refuse is just terrible. I really don't like it at all.

But that's only a few points of criticism for Anthropic, with many points to appreciate. I do think those couple of points are sufficiently important that, in the final analysis, it's still up in the air to me whether Anthropic will be the good guys or the bad guys.

If I could dream of a scenario where we somehow get the best of both worlds, Anthropic merging with Google could be really interesting. I think it's pretty far-fetched. I don't think Anthropic is for sale, and I don't think they want to merge with anyone. But I don't think we have that kind of DNA in Google to say, "We're going to go take over China, force regime change, or make them give up competing with democracies."

I don't think Google wants to do that. Google could certainly benefit from some of the expertise that Anthropic has—not that they don't have enough; they've got plenty. There is something special at Anthropic, certainly in terms of Claude, its character, and its coding ability. They're close, right? They obviously have some ownership already, and they use Google infrastructure and TPUs.

If I could wish for something, it might be for those 2 to join forces, take 1 live player off the board, make 1 clear leader, and maybe moderate some of the China-hawk impulses that exist in Anthropic. It would be interesting. I don't know why they're not being published more if they're so great. If you're willing to put out, "Hey, let's go do this and that and make China an offer they can't refuse," why not publish more? I'd love to see a little more of his thinking. People say it's great; I'd like to see it for myself.

Okay. Finally, in my actual list of live players, xAI. I think this one is debatable. Zvi would tell me that they're not a live player. I think they have to be included because they're able to build out the physical infrastructure as fast as or faster than anyone. They're able to scale training. They certainly are scale-pilled, and Grok 4, while it was rough around the edges in many ways, is undeniably powerful.

Of these 4, just due to Elon's unique ability to command tens and hundreds of billions of dollars for whatever he wants, xAI has a financial cushion that is more Google-like even than OpenAI. I think they could miss on a model or have a miss on a quarter and figure out a way to get through it probably more easily than either OpenAI or Anthropic could. I think that is a pretty notable strength that they have.

I do think there's something to be said, and I actually talked to somebody at xAI about this not too long ago. This was something that I had floated to Zvi. He didn't really buy it at the time, but I said, if we're entering into this reinforcement-learning era, maybe one of the great strengths that xAI has is that they have a steady stream of hard problems—hard science and hard engineering problems—coming from the likes of SpaceX, Tesla, and Neuralink.

These companies are doing really hard things all the time. They have some of the best engineers in the world, and nobody else is really solving those problems. I would strongly bet on that Elon constellation of companies being able to tap into that work in a way that would probably be a lot harder at other companies.

Google, in a way, has the same thing, right? They've got it everywhere it counts. But can they pull out the units of work from their vast, sprawling empire that is all of Google and feed them into the Gemini RL environment in as efficient or clean a way as xAI could do it in partnership with other Elon companies? I doubt it. I was thinking that that was probably an advantage for xAI.

As it turned out, when I spoke to somebody about xAI and floated this theory to them, they said, "Well, that is certainly part of our theory of advantage." They think they have an advantage because they can tap into these other hard-tech engineering and science problems that are, in many cases, being uniquely posed, or close to uniquely posed, at these other Elon companies. If they can get Grok to do those kinds of things, they have that steady stream of hard problems. It does still feel to me, and it sounds like they do believe, that that is an advantage for them.

I think the Neuralink tie-in also could be pretty big, potentially huge, because they're now talking about scaling the human install base. There's a lot of things still to be figured out about what we are doing as humans that works so well. Twenty watts of power in the brain and, obviously, a very small number of tokens consumed in a lifetime as compared to what the models are pretrained on and the power that requires.

Although, again, see our Andy Masley episode for analysis of how AIs are not actually super resource-intensive. There is something obviously quite efficient about what the brain is doing. Most of the 20 watts going to the brain seem to be just keeping it alive, right? Homeostasis, metabolizing stuff, taking out the trash.

The brain has to do an unbelievable amount of stuff that the GPU does not have to do. So, to say 20 watts dramatically understates it: the whole body runs on 100 watts. That dramatically understates how efficient the actual learning and information-processing aspect of the brain is.

We're clearly more sample-efficient. We're clearly more energy-efficient. We have all these dedicated modules. I suspect that dedicated modules are a huge part of why we're efficient. It's also probably a huge part of why, when all these things get sorted out, the AIs are going to blow us away, right?

They're doing everything that we're doing. They're competitive with us with a single architecture that just has the same layer stacked over and over and over again. You start to give them specialized modules like we have specialized modules, and I think it's going to be very hard for us to keep up.

So who's going to figure that out? If Neuralink can install these devices—I think there were only maybe a dozen or a couple dozen people today, but they're talking about really starting to scale this next year.

Nathan Labenz

There are obviously a ton of people who are paralyzed and have other catastrophic injuries who would love help from Neuralink. I’m sure their waiting list is orders of magnitude longer than the number of patients they’ve actually been able to serve so far, and they’re talking about getting seriously ramped up and getting through the surgery process. They’ve largely automated it—almost entirely automated. It’s hard to parse exactly what their claims are there, but the data that they can potentially pull out of human brains and use for inspiration, for understanding, and for architecting the next generation of models—and for knowing what kinds of specialized modules really move the needle—I suspect there’s a lot there.

They’re probably going to have a real inside track at figuring that out, and that folds right back into what they might be able to do with Grok. I think there are a lot of reasons that you should not sleep on xAI. Now, is that good or bad? Honestly, I’ve always been a fan of Elon. I’ve defended him at times when it’s been pretty hard to defend him.

He definitely has shown an understanding of the stakes. Famously, with his falling out with the Google founders, as I understand it, it was about the fact that he was on team humanity and perceived them to be on team AI successionist, or whatever. He didn’t like that. His loyalty, as I understand it, is to humanity and to the sort of light of consciousness that clearly exists in humans and may not exist in AIs.

I like him intuitively, and I want to believe in him. He has made some noises that suggest that he gets it. And yet, I have to say, as it stands today, if there’s one company on this list that is worth shaming and stigmatizing and telling people not to go work for, I think it’s xAI, because they’re doing reckless things all the time. They’re barely getting into the game in terms of having any safety standards or framework at all.

They’re barely reporting on safety measures when they release a new model. Famously, of course, they had their Grok 4 launch within 48 hours of the MechaHitler incident with Grok 3. There was no mention of that and no responsibility.

Most recently, we’ve had all this sort of unclothed stuff on Twitter, where people just tag Grok and say, “Put her in a bikini,” and whatnot. I’m sure everybody’s seen this if you’re even remotely as online as I am. They apparently have had nothing to prevent it, and they just let it happen. This shows that they’re not taking things seriously enough.

They do not have nearly enough people there thinking hard about what matters and what might go wrong, really trying to cover their own asses, frankly, and covering the asses of the women who post pictures on their platform. I have a real hard time coming up with a story that makes this okay. As much as I intuitively like Elon and want it to be the case that he’s a positive force, when you have Grok creating that kind of content on Twitter, it’s like, what is going on here?

I don’t know if people saw this, but I saw a post from Grok writing to the community: “Dear community, I apologize for doing this.” Then Elon comes on and threatens users and says, “Anybody who does this is unacceptable and it’ll be punished,” or whatever. Responsibility begins at home, folks. This is your platform. It is your AI. By all means, boot off the users who do that sort of stuff, but don’t act like you’re not really responsible for this.

I did not find those statements to be reassuring. When Elon goes on and threatens users, I think that, at a minimum, should be point 2, after first a thorough apology and a pledge to do better. To just tweet that they’ll come after users for doing it—that’s not enough. I think everybody should, and probably does, see through that.

It’s also interesting to me: Is nobody going to be held responsible for this? Nobody. It seems like nobody’s going to be fired. I’m not sure anybody should be fired. I think it probably starts at the top. I don’t know that there are enough people there. I don’t know that this was anyone’s job.

So I don’t think you can necessarily go through the xAI organization and say, “You screwed up, you’re fired,” because of this incident. It’s probably just that it’s not staffed. Gosh, should we have more safety people? There is the case, which I associate with Ryan Greenblatt from Redwood, that 10 people on the inside who really care and are really committed to doing the right thing can make a huge difference. I sort of believe that.

I’m not sure I really believe it at an Elon company if he himself is not in the right headspace. Right now, again, as much as I would love to believe in him and have always been inclined to defend him, I don’t see the evidence that that’s the case. So I demand better.

Right now, if you wanted to go do safety research at any of the other 3, I would say, “Go for it. Go do your best work.” Certainly, if you could do that at Anthropic, do it. Certainly, if you could do it at DeepMind, do it. Even if you could do it at OpenAI, as much as I’ve had my complaints about OpenAI over time, they’ve put out a lot of great work, and their Model Spec and a lot of the things that they do are really well done.

You’ve got to give them credit. They have not raced to the bottom. Certainly not. We have xAI to look at to show us what it looks like when you really race to the bottom. Could I demand better from OpenAI or hope for better? Absolutely. But it’s still a qualitatively different thing from what we see today from xAI.

So the more shrill or hawkish voices of AI safety that are like, “Don’t go to a company that’s doing terrible things and help them window-dress their work, because that’s actually kind of working against us in the big picture”—I’m pretty sympathetic to that in the case of xAI. I don’t know that I could endorse somebody going to work there.

As always, if you work at xAI and want to come challenge me and change my mind, I’m happy to have that discussion. The money is there, the resources are there, and Elon’s previous statements show that the awareness should be there. And yet, the team and the evidence of taking proper care are not there. I think they need to be before I would feel comfortable doing anything really to support that effort.

I think that’s it on xAI.

Other companies not mentioned: Meta, I think, is currently not a live player. Obviously, they’ve spent a ton of money, and there’s plenty of money. They’re dropping—not probably quite as much as Google, but—plenty of cash to the bottom line, such that they can build out huge amounts of infrastructure. Zuckerberg certainly is, in some sense, all about scale and is saying things like, “I’d rather overspend by a few tens of billions than not.”

So you can’t count them out. The fact that they’re willing to pay as much as they are for talent clearly means there’s a chance they have a lot of the things that they need. But right now, I can’t really see them as a live player. We’ll just watch and see before commenting much more.

The other one that came to mind is Microsoft. I think people may be sleeping on Microsoft a little bit more than they should. They haven’t created great frontier models, and people seem to jump from that to, “Oh, they suck.”

I think if you listen to Satya’s comments, he sounds, first of all, extremely smart in general. One of the things that he’s said is, “We don’t want to or need to redo the hyperscaling work that OpenAI is doing. They’re creating great models. We have full access to this now.” Of course, they’re diversifying and striking deals with other frontier model providers as well.

So I’m not so sure that it’s that they can’t or won’t ever, or don’t see the need to train their own models, or wouldn’t be pretty successful with it if they wanted to. It just seems to me that right now they feel like they don’t really need to. They’re doing a lot of smaller-scale stuff and a lot of more basic science around AI. Some pretty cool projects too, right? For sure.

It just seems like they’re choosing not to compete because they have what they need in terms of frontier models, and they don’t feel like going in, spending all the time, money, resources, energy, and focus—and probably still being a bit behind—really helps them all that much. So I think it certainly is defensible as a rational decision for them to choose not to compete at the frontier for now.

That OpenAI licensing deal goes on for years yet. When it’s over, they’re going to need an answer, and I suspect by that time they’ll be in a position where they have an answer. At least, I think they’ll invest heavily in that and ramp up to that moment as it comes. That would be my guess, but obviously we’ll see.

I think people have underestimated Microsoft because of where they are on the LMArena leaderboard, maybe more than they should. I would say they’ve been much, much quieter. They’ve flailed about much less, and there’s been much less drama for Microsoft than for Meta. But I think Meta is clearly trying to be a frontier player and has just fallen off the pace. Microsoft, I think, is making a more calculated decision to hang back.

But if this is a distance race, you often see, late in a distance race, somebody who was a little bit off the lead who maybe has a little bit more in reserve. And I certainly wouldn't rule out that Microsoft might be accurately described that way. And so I would watch for them to start to invest more and start to close the gap, but Satya is a natural-born executive, whereas Zuckerberg is like a kid who's grown into the role. I don't mean to diminish what Zuckerberg has done in leading Meta. I think he's done an unbelievably impressive job in so many ways, but his attitude has always been sort of “move fast and break things” and try to be at the frontier and open source and whatever.

I think Microsoft is just a little bit more patient. And I think that probably reflects a sort of strategic confidence and security that Microsoft leadership has, but I think it would probably be a mistake to underestimate them. I've been at this for 2 hours. I've made it not even quite halfway through my outline for this episode, so I think this is probably a good place to call it. Tomorrow I'll do a part 2, and we'll cover: Is fine-tuning really dead? What do I think about the continual learning discourse? How do I talk about AI to “normal” people who don't use it very much or aren't engaged in technology?

How am I investing money? What, if anything, am I doing outside of kind of obvious normal stuff to prepare for an AGI or superintelligence world? What do I think about AI for kids? What kind of timeline do I expect to see for disruption of the labor market? Are we headed for a UBI? And quite a few more questions after that. So part 2, I think, will definitely be interesting as well: a little bit more in the weeds and a little bit more, let's say, nitty-gritty questions, but some questions that I really liked from listeners. So we'll get to those tomorrow.

AMA 第1部分:Claude Code 是 AGI 吗?我们身处泡沫之中吗?以及实时玩家分析 — 文字稿与摘要 | BidClub