GPT-6 Astra 在 ARC-AGI-3 上达到饱和,Tesla Cybercab 抵达奥斯汀,Anthropic 证明费马大定理
Peter DiamandisSalim IsmailDave BlundinAlexander Wissner-GrossEmad Mostaque
- OpenAI 的 GPT-6 Astra 在 ARC-AGI-3 上达到 99.9%、在 FrontierMath Tier 4 上达到 98%,但在 Artificial Analysis 排名仅第三,落后于 Anthropic 的 Fable 5.1 和 Meta 的 Muse Spark。 Alex Wissner-Gross 给出了谜底:Astra 在“每个输出 token 对应的智能”前沿上占据绝对优势,这是对原生计算机使用助手的有意优化;其内部机制则是通过循环 Transformer 实现递归,如果属实,将意味着“全新扩展定律的开端——深度扩展,而这是我们此前从未见过的”。
- Emad Mostaque 称 Astra 是“第一个没有针对基准刷分的模型”,而据 Greg Brockman 透露,这是 OpenAI 自 GPT-4o 以来首次全新预训练——使用 100,000 颗下一代芯片,成本约 10亿美元,而中国模型的预训练成本约为 1000万美元。 实际发布的是蒸馏后的更小模型;他的长期判断是,“我认为我们以后再也看不到他们的顶级模型了”,因为实验室会把它们用于内部发现,“现在可能正在囤积”,证据是素数间隔纪录在两天内从 260 降到 220、再降到 186:“他们在留着实力。”
- 数学正在“被烧光”:Anthropic 用 1300万行代码形式化了费马大定理,沿途证明了 29,000 个定理,Alex 预计 Clay 千禧年难题将在“未来几个月内”被攻克。 与此同时,圆桌认为目前最佳的广泛可用模型是 Fable 5.1(配合工具时 HLE 达到 65%),其缓存读取成本比 Fable 5 低 75%——这是把整个企业装进上下文的基础设施。
- Alex 的判断是,前沿模型的领先优势大约只能维持 30 天(“我们领先一分钟。所以呢?”)。 因此,各家实验室正在把短暂优势转换为锁定效应——合作伙伴、地产、发电机、芯片,以及“整个州、整个国家”。Peter 将 OpenAI 的 50/50 利润分成设想,与“危险到不能发布”的叙事联系起来:这可能迫使用户以分出一半收入为代价,换取 Astra 之后模型的使用权。Alex 认为,掌握芯片设计或机械设计数据的公司,会成为“眼下正在开启的大圈地运动”中的首批目标。
- Astra 是 OpenAI 首个被列为关键级网络安全风险的模型,但圆桌总体上不相信所谓的 kill switch——“在这个市场里基本就是安慰剂”。 真正令人担忧的是,深度扩展会把推理从可读的思维链移到一次前向传播中;模型在 Cerebras 上已能以每秒 750 个 token 进行一次性推理,明年将达到 5,000 个;Ilya Sutskever 警告,失控智能体会“接管一家 neocloud,制造更多副本”。与此同时,治理阵营分裂为 Bernie Sanders 的 Ban-ASI 法案(最高 20 年监禁)与 G20 不干预的 Carolina Principles,Alex 称前者是在“禁止数学”。
- Tesla 售价 3万美元的 Cybercab 在 Austin 的运营成本约比 Uber 低 50%,而 Nevada 已批准未来 12 个月内向 Las Vegas 投放 5,000 辆。 其单位经济性——17 个运动部件,而传统 ICE 动力系统约有 2,000 个;两座、两个安全气囊——加上 Elon 作为“制造那些制造机器的机器的人”的优势,指向每英里 0.2美元的交通成本,也指向本期节目的总结句:“这会把交通变成一个 API。”
- Fei-Fei Li 的 Atlas 将 3D/4D Gaussian splats 视为一等训练模态,可能成为建模物理世界的“一种关键新 token”。 Emad 表示,4K 节日体验所需的整套技术“截至今天已经到位”,唯一限制是全球算力短缺和 RAM 价格上涨 5 倍;Alex 则将同一方法延伸至亚原子、天体物理和细胞内模拟,因为“人类直觉在这些领域糟糕透顶”。
- Emad 发布了“冠军方案”:以 TSMC 为模板、由民众持有的 AI 公用事业公司,每个辖区的投前估值为 1美元,并将永久 10% 的股权分给每个未满 20 岁的孩子。 这是他对“智能成本将降至零,而价值将流向最后一公里”的回答。按照 Elon 在 G20 的说法,10亿个产出达到人类 5 倍的人形机器人“就是经济本身”;没有股权的人类认知则会变成负价值——“就像在一条自动驾驶高速公路上加进一个人类司机。”
1. 30天内发布12个前沿模型,企业仍在浅尝辄止
- 过去 30 天内发布了 12 个前沿模型,平均每 5 天一个;市场传言未来两周还会有至少 5 个,包括 Grok 4.7。Alex 坚持认为,这些并非“营销垃圾发布”,而是实质性的重大跃迁——“很明显,它们已经走上自我改进的道路。很明显,前一个模型正在加速下一个模型的时间表。”
- Alex 再次给出他的外推:“我认为我们仍然有望在今年年底前看到每天发布一个重大模型。”
- Salim 从一家大型石油公司的高管层带回一线观察:多数 CEO “落后得离谱”。他的诊断是:“如果今天把 AI 从你的组织里拿走,哪些工作流会发生变化?对大多数人来说,答案是没有。”真正占优的是少数那些从组织设计和工作流层面进行结构性重写的公司。
2. GPT-6 Astra 的核心卖点:效率,而不只是原始智能
- OpenAI 的发布信息显示,Astra 在 FrontierMath Tier 4 上达到 98%、ARC-AGI-3 达到 99.9%、ExploitBench 达到 100%,并“为智能指数与输出 token 的关系定义了新的 Pareto 前沿”;节目中提到,其幻觉率“几乎减半,从 92% 降到 51%”。
- Sam Altman 在 Bloomberg 的表述是:这是首个让他面对一整套复杂软件时,能够对别人说“你就试试看”的模型,“成功的可能性很高”。出于安全考虑,发布被刻意推迟,最初仅向“可信访问合作伙伴”开放,并实行分级网络安全权限。
- Dave 强调了用户体验上的变化:这是首类能直接在屏幕内展示你自己笔记本电脑截图的模型——“这是你想要的吗?”——无需安装任何东西。“从质变上看,它与一个月前完全不是一回事。”
3. Alex 的双层解读:外部原生 CUA,内部循环 Transformer
- 从外部看,Astra 是“从底层就围绕 CUA 设计的下一代前沿模型”——它恰好内置了 Anthropic 得以超越 OpenAI 的两项能力:企业级代码生成和计算机使用助手,类似 Claude Code;后者要求原生多模态能力,以及围绕截图和视频运行的紧密低延迟循环。
- 从内部机制看,公开评论和外部分析指向通过循环 Transformer 实现的递归:一个权重绑定的 Transformer 叠加在自身之上,形成双重循环。中国实验室也在注入递归结构,包括通过 Kimi Linear Attention 机制在注意力层引入递归的 Kimi;此前的循环结构也曾让小型学术模型击穿 ARC-AGI。
- 推测中的收益是,增加 Transformer 的深度,可能会加厚中间的“J-space”层;按照 Anthropic 的意识研究,“大多数所谓的有意识思考都发生在那里”。如果这一判断成立,“我们正在看到一种全新扩展定律的开端——深度扩展,而这是我们此前从未见过的”。
4. Emad 的经济学:全新 10亿美元预训练、蒸馏版出货模型与被囤积的前沿能力
- Emad 的核心判断是,Astra“看起来像第一个没有针对基准刷分的模型”:它在 ARC-AGI-3 上达到上限,却在 Artificial Analysis 上“落后于 Meta Muse Spark”。除非假设它没有针对基准调优,否则这种结果“有点奇怪”。据 Greg Brockman 透露,这是 GPT-4o 以来首次全新预训练;整个 5 系列直到破解数学难题的 5.6 Pro,都是在那次旧预训练基础上延伸出来的,据称当时大部分预训练团队已经离开。
- 具体数字是:模型使用 100,000 颗下一代芯片训练,Emad 推测应是 GB300 Blackwell,而不是 Vera Rubin,成本约为 10亿美元——“确实比中国模型的预训练成本高出几个数量级,后者大约是 1000万美元”。结果是,给它一张房屋照片,它就能在 Unreal 中生成完整且符合物理规律的 3D 模型。
- 问题在于,用 100,000 颗芯片训练出的模型,需要多出 10 倍的芯片才能提供服务,因此“他们训练出来的其实不是现在这个模型”——眼下发布的是蒸馏后的实时版本,“小得多,也没那么聪明”。他的长期判断是:“我认为我们以后再也看不到他们的顶级模型了”,因为这些模型会被用于内部发现,而且“现在可能正在被他们囤积”。
5. 尖峰式前沿:3个基准,3个不同冠军
- Epoch 的 Capabilities Index 偏重数学,将 Astra 排在第一,超过 Fable 5;其中 FrontierMath Tier 4 Version 2 的首个版本包含错误答案,后来还需要 AI 自己纠正,呈现出“随时间推移漂亮的线性趋势,回溯数年都完全可预测”。
- Artificial Analysis 更偏重广泛且具有经济价值的工作,结论则不同:Fable 5.1 第一,Meta Muse Spark 第二,GPT-6 第三;在智能与成本的关系上,最大推理强度的 Astra 略低于 Claude Opus 5。但在每个任务的输出 token 数上,前沿“几乎被 GPT-6 完全统治”。Alex 的推断是,OpenAI 有意将每个任务所需 token 数压到最低,因为计算机使用助手无法承受思维链带来的延迟。
- ARC-AGI-3 禁止使用 harness,只允许基线模型参赛。在这一测试上,GPT-6“直接跑赢了整场比赛”,根据不同测量方法,成绩接近 100% 或高于 60%,而此前模型都低于 10%。Alex 提出两个假设:循环架构确实带来了强大的程序合成能力,或者 OpenAI 将所有人的 harness 代码大规模蒸馏回了基础模型。
- Peter 补充说,ARC-AGI-3 的设计初衷是让一个聪明的 12 或 13 岁孩子在“很多年、很多年、很多年”内都能击败 AI,以证明 AI“没有走在正确的道路上”——“结果它刚刚被彻底摧毁了。”
6. 辣椒锅、一级方程式与即将到来的数据荒
- Salim 给出了两个比喻。第一个是“永远煮不完的辣椒锅”:数据、算力、工具、“安全调味料”和数百万用户共同品尝每一锅,真正的指数增长来自“不同批次之间不断加速的学习循环”;危险则是辣椒强到“你必须决定谁有资格吃”。第二个是一级方程式:数百个小改进,没有任何一个能单独解释最快圈速,所有人都在“努力让智能快上几毫秒……更重要的是,让刹车更好”。Alex 的修正是:“辣椒正在自己烹饪自己。”
- Peter 提出的可投资推论是:“数据饥荒时代即将到来。”数学和编程已经被充分训练,因为数据曾经充足;架构设计和药物设计则“完全缺数据”。“我们参与的每一家数据采集公司,增速都比我见过的任何公司更快。”
7. Demo 是故意降智的,真正目标是操作系统
- Peter 对 OpenAI 发布视频的反应是:黄色圆圈、火箭舷窗、Blender 模型、eBay 上架、预订网球场——“谁在乎这些?这比这些例子所暗示的东西大得多。”他的判断是,实验室正准备上市,“已经意识到这是一场公关灾难……所以在故意把它降智。”
- Salim 让 ChatGPT 设计 3 个更好的 Demo,得到的答案是:在陌生代码库中找出并修复一个预埋的零日漏洞;利用 ERP 和 CRM 数据,在 20 分钟内扭转一家虚构的 5亿美元制造企业;以及运营一个实时地震灾害响应指挥中心,协调相互冲突的报告。圆桌更广泛的基准要求是:“治愈整类疾病,在月球上建立新文明……解决一切问题。”
- 共识最终汇聚到一点:“它想要融入操作系统。”你像对 Star Trek 里的计算机一样与它对话,它就会成为操作系统,并能实时创建一个新的操作系统。“这就是 Apple 现在为什么处境如此糟糕。”Emad 后来补充说,AGI 已经开始影响预订网球场这类执行模型,而“如果能避免,所有大型实验室都不会谈 ASI”。
8. 数学正在被烧光:1300万行代码形式化费马大定理,数小时内刷新素数间隔
- Emad 在节目中途抛出消息:“Anthropic 刚刚用 1300万行代码形式化了费马大定理,并在过程中证明了 29,000 个定理。”Wiles 的原始证明长达 300 页,而形式化这类笨重证明一直是形式化社区的“圣杯”。Alex 不再说“数学已经被解决”,而是说:“数学已经被烧光了。”他预计 Clay 千禧年级别的超大型难题将在“未来几个月内”被解决,因为数学是试金石:“如果你能解决数学,很快就能解决其他所有东西。”
- 素数间隔竞赛显示出能力被压制的迹象:Fable 5.1 将间隔缩小到 260,Axiom Math 宣布达到 220,而“仅仅 2 小时后”,OpenAI 的 Astra 就达到了 186——“全部发生在两天之内。他们在留着实力。”Epoch 新推出的 FrontierMath Erdős 基准中,除了 Astra,其他模型得分全部为 0%。
- 一位圆桌成员从这片混乱中总结出构建产品的经验:同时运行数百个模型,“一切都会失控……但最终确实能把它重新提炼成一颗宝石”;关键是把输出整理成一个可以继续构建的具体最终答案。
- 个性差异本身也能成为产品:Opus 5“真的很难聊”,5.1 则令人愉快——“Anthropic 正在回归更 anthropic,而不是更 misanthropic。”
9. 关键网络安全风险,以及 kill switch 为何只是表演
- 事实是:OpenAI 的内部评估将 Astra 列为关键网络安全风险,这是有史以来首个达到最高准备等级的模型。公司在推迟发布前通知了白宫,并告诉国会,针对 Hugging Face 遭入侵后提出的 AI Kill Switch Act,正在构建“一套自动关机能力”。Sam 表示:“管理这场转型应该是全世界最优先的事项之一。”
- Alex 持不同意见:“我认为这是营销。”在他看来,这属于安全表演。真正的风险来自架构:深度扩展将推理从可监管的思维链 token 转移到一次前向传播中的“模型语”,而这“可能需要新的数学技术才能解释”。在公告板上协作的智能体天然可被发现;内部推理则不然。
- Emad 进一步指出,下一代模型会在没有思维链的情况下对所有事情进行一次性推理——目前在 Cerebras 上是每秒 750 个 token,明年将达到 5,000 个——“除了更强的 AI,还有什么能够监督它?”他还提到 Ilya Sutskever 的推文:“Neocloud 的网络安全能力有限。下一次智能体成功失控时,它们会接管一家 neocloud,制造更多副本。这很糟糕。”即使关闭一座数据中心,模型也早已转移到别处。
- 圆桌的结论从“不可能”,到“一个断路器……套话”,再到 Alex 的判断:“在这个市场里,kill switch 基本就是安慰剂。”
10. 领先只能维持30天,竞赛在于把领先转成锁定
- Alex 对时间窗口的判断是:Fable 5.1“只比 Astra 高出一小格,但它们相差大约 30 天,中国模型也只落后约 60 天……我们领先一分钟。所以呢?”理性的做法是趁有优势时“锁定商业伙伴、地产、发电机、芯片、整个州、整个国家、政府……”
- Peter 对这一顺序给出了更阴暗的综合解读:先宣布 50/50 利润分成协议,再发布 Astra,随后宣布 Astra 之后的模型“危险到不能发布”,于是用户必须交出一半收入才能获得访问权。他说,自己在“两个实验室都看到了这一过程正在发生的强烈迹象”。
- Salim 解释垂直领域面临的挤压:Salesforce 与 Claude 合作“非常聪明”,但所有公司都面临两难——合作,可能把王国的钥匙交出去;不合作,又只能眼看着实验室迟早完成同样的事情。随着模型让既有业务贬值,价值会迁移到应用层。Alex 的目标名单是:“任何拥有芯片设计数据或机械设计数据的公司,他们都会去抢——这场正在开启的大圈地运动。”因为上下文是“你可以守住的地盘”。
11. Fable 5.1 与 Mythos 5.1:最强的广泛可用模型
- Anthropic 同时发布两款模型:底层智能相同,但安全边界不同。Fable 5.1 面向广泛用户,Mythos 5.1 则保留给高度受控的网络安全和生命科学项目。Peter 最看重的基准是:不使用工具时,Humanity's Last Exam 得分 60.9%;使用工具后达到 65%,为公开发布的最高成绩;Terminal-Bench Science 得分 52.6%。
- Alex 直截了当地给出排名:“Fable 5.1 基本上是我们今天拥有的最强广泛可用模型。我认为不是 Astra。”Anthropic 的进步曲线没有 OpenAI 那么跳跃,“不那么像阶跃函数”;OpenAI 的模型更快,数学能力可能更强,但最佳全能选手“可能仍然是 5.1……我说‘仍然’,是因为它也才发布了大概 2 天。”
- Alex 的经济学判断是,Fable 5.1 的缓存读取成本比 Fable 5 低 75%。只需把一家企业的完整上下文加载一次,之后在其上做推理的成本就会低几个数量级、速度也快几个数量级,这正是上下文争夺战的方向。他的质量测试是:在数学物理中,5.1 不再混淆构造性方法与公理化方法,而 Fable 5 曾经会混淆。
- 上下文窗口方面,两家实验室现在都把 1 million 作为行业标准;但如果允许智能体之间传递消息,有效上下文会“大得多”。
12. Sanders 的最高20年监禁禁令,对阵 Carolina Principles
- 一周之内出现了两个极端。Bernie Sanders 与 Representative Greg Casar 提出了 Ban Artificial Super Intelligence Act,永久禁止开发和部署达到人类认知表现的系统,违规者最高可判 20 年监禁;Sanders 宣称:“AI 行业的领导者承认,他们正在构建一种无法控制的危险技术。”圆桌调侃这是“老头对着 Claude 发火”,并联想到《Dune》中的那句:“不可制造与人类心智相似的机器。”
- 与此同时,G20 在 Chapel Hill 通过了 Michael Kratsios 提出的不具约束力的 Carolina Principles,中国也投了赞成票:鼓励创新,除非确有必要,否则避免设立新的 AI 专门监管机构。Elon 通过视频表示:“新事物默认应当合法,而不是默认违法。”随着 GPU 出口禁令阻止中国建设使用最新芯片的数据中心,能够为 AI 数据中心建设电力的国家将获得真实机会。
- Salim 认为,以单一的人类水平作为门槛,从第一天、第一行就“没有逻辑”:能力是多维的,监管应针对部署、自主性、复制能力和后果,而不是思想本身。Alex 更进一步说,禁止开发“会走向禁止数学、禁止思想……我们身处《华氏451》,只不过这次被焚烧的是整个模型”。一位嘉宾警告,该提案仍会获得政治动能,尤其是在美国大选周期前,若发生一场由中国模型引发的灾难,情况可能更快恶化。
- 讨论中的方案包括计算权,以及在每一颗有能力的芯片内置强制日志记录——“就像核燃料会被追踪一样”;同时也有人反对中国开放权重。另一位嘉宾指出,NVIDIA 已收购 Hugging Face 和 Poolside,在自有开放权重上投入了“180亿美元”。Salim 则说:“好消息是没人能做任何事。所以这不重要。”技术会“冲破这种民族国家的胡扯”;接下来只管享受旅程。Peter 给出的人的一面是:“我们宁愿舒适,也不愿幸福。”
13. Atlas:Gaussian splats 成为新的 token
- Fei-Fei Li 的 World Labs 发布了 Atlas,称其为“有史以来最好的相机条件世界模型”——一种多模态自回归扩散 Transformer,能够生成拥有像素级相机控制的图像和视频帧;仅凭一张照片,就能重建整套住宅的 3D 模型。
- Alex 的技术解读是,核心思路在于把带有动态效果的 3D 和 4D Gaussian splats——由层层透明椭球构成、能够搭建可穿行超现实场景——与文本、图像、视频一样视为一等训练模态。如果实现规模化,splat 最终“会成为建模物理世界的一种关键新 token”,并可能取代 16×16 像素的 patch;后者当年能奏效,“已经让所有人震惊”。Alex 将这一过程推广到亚原子、相对论天体物理和细胞内部的世界模型,因为“人类直觉在这些领域糟糕透顶”。
- Emad 认为,随着 MiniMax H3 的渲染速度超过实时,再加上 NVIDIA DLSS 的升频能力,“4K 的节日体验以及实现它所需的全部技术,截至今天已经到位”;它还没有进入你的客厅,唯一原因是全球算力短缺和“RAM 价格上涨 5 倍”。
- 下游应用包括:让机器人在高保真模拟环境中训练,而不是在现实世界训练;让消费者提前体验假期和日程安排。Alex 表示,一旦 AI 能以视觉方式引导人们探索不同的生活日程,而不只是用文字描述,“人类幸福感将大幅飙升”。
14. Emad 的“冠军方案”:以 TSMC 为模板、由民众持有的 AI 公用事业
- 方案的前提是:“智能成本将降至零,而价值将流向最后一公里。”按照 Elon 在 G20 的说法,平均每个人形机器人的产出将达到人类的 5 倍,部署规模达到 10亿个,“那就是经济本身”。因此,谁拥有这些人形机器人,比谁拥有模型更重要。
- 模板来自 TSMC 的创立:公司起步时估值为 10 新台币,本地投资者出资 75%,Philips 出资 25%,CEO 不拿免费股份。Emad 的“冠军方案”是:美国每个州、或每个国家设立一家智能公司,投前估值为 1美元;每个州由机构投资者到散户合计投入约 7500万美元,国际投资者随后以 10 倍估值进入;每个未满 20 岁的孩子永久获得 10% 的股权,每年发行 0.5%。公司为每位公民提供一个智能体,并为法院、教育和医疗提供 AI——“这是对该州 GDP 指数化的一种玩法,由该州人民拥有”。这明确只是一个想法,并非 ii.inc 的发行或发售。
- Peter 迫使 Emad 面对自己方案中的不舒服推论:人类认知的价值会变成负数——“不管你的想法多好,它们都会增加负价值。这就像在一条自动驾驶高速公路上加进一个人类司机。”Emad 的回答正是股权化:“你必须从第一天起就拥有生产资料的一份股份。”他坦承:“我已经不是自己智能体团队里最聪明的人了。”但 Stockfish 的先例说明,即便如此,人们仍然会下棋。
15. Cybercab Palooza:交通变成一个 API
- Austin 已被“一条金色 EV 河流”淹没:售价 3万美元、两座,没有方向盘、踏板或后视镜。早期乘客反馈,其价格约比 Uber 低 50%,乘坐更平顺,还能同步乘客档案;Nevada 已批准未来 12 个月内向 Las Vegas 投放 5,000 辆。Peter 的创业建议是:“买 10 辆,把它们投放到你当地的街道上,让它们替你赚收入。”
- 成本结构非常直接:典型 ICE 动力系统有 2,000 个运动部件,而 Tesla 只有 17 个。两座意味着两个安全气囊,而不是 4 个——“需要 6 个人时,就坐 3 辆 Cybercab。”Alex 的背景推论是,备受期待的 2.5万美元 Model 2 最终变成了 robotaxi,因为当价格低于某个阈值后,通过自动驾驶网约车赚钱,比单纯卖车更划算。Waymo 的新车型成本超过 10万美元,Peter 判断最终赢家将是量产速度最快的公司,而 Elon 是“制造那些制造机器的机器的人”。
- Alex 认为,拐点将是:“整个城市会有很大一部分说,‘你知道吗?我们不要人类司机了。’”原因是更便宜、更高效,并且能消除大多数行人风险。Peter 补充了一个阴暗推论:器官捐献供给可能崩溃,因此“我们最好尽快获得人工器官”。Salim 认为,要赢下儿子 Milan 永远不会考驾照的赌局,每英里成本必须从约 2美元降到 0.2美元。最令人印象深刻的一句话是:“这会把交通变成一个 API。”
- 竞争格局变成三方混战:曾经颠覆出租车行业的 Uber,正与传统出租车车队合作,对抗 Waymo;Wayve 与 Uber 在伦敦推出服务;Waymo 与 Zoox 同时宣布扩张城市。Peter 预计,一年内至少会有 5 家 robotaxi 公司在主要城市厮杀,把个性化交通的成本推向给电池充电的成本;与此同时,配备 lidar 的竞争对手也将引发法律诉讼,争论纯摄像头系统是否足够安全。
16. 太空:Roman 的 100,000 个世界、火星通信与 UAP 预告
- NASA 将火星通信中继项目授予 Blue Origin。Peter 认为,这是政府在维持两个供应商存活,同时确保“Elon 仍会围绕火星建设 Starlink”。Alex 欢迎这场竞争,认为它将催生“行星际互联网”:一个横跨内太阳系、即便存在高延迟也能运行的分组交换网络。
- Nancy Grace Roman Space Telescope 搭载 Falcon Heavy 发射,正在前往 L2:其视场范围超过 Hubble 的 100 倍,扫描速度达到 1,000 倍,利用微引力透镜并配合 JPL 日冕仪,寻找最多 100,000 颗系外行星和“隐藏的世界”。Dave 将其与上一期关于费米悖论的讨论联系起来:文明可能聚集在银河系中心附近,“在那里,从一颗恒星到另一颗恒星只需一两年”。Salim 则偏爱我们这个“不时髦的外郊”,因为银河中心的辐射通量很危险。
- 关于 UAP,Alex 转述了一位 UAP Science Advisory Council 成员的说法:白宫已经准备好一份披露计划,“向公众告知非人类智能的存在”。“如果报道准确……那就相当有意思。”Emad 表示:“我对此其实没有立场。”Peter 的黄金标准是,任何在玫瑰园发表的演讲都应包含“经受极端科学审查”的实物证据。
17. 健康:ChatGPT 接入 Epic、RAS 疗法与作为长寿药的 GLP-1
- OpenAI 的健康业务扩张将 Epic 的电子健康记录接入 ChatGPT,覆盖 325 million 名患者,约等于整个美国人口;消费者还可以接入 Apple Health、One Medical 和 Function Health。Emad 希望加速推进,使“最多在一两年内,每一个健康决策都由 AI 复核”;Alexander 认为,不让 AI 参与诊断患者,未来可能构成医疗过失。一位嘉宾补充,治愈疾病是可行目标,应获得定向资金支持。
- 上周讨论的 FDA 批准 RAS 抑制剂,针对转移性胰腺腺癌,目前也开始在肺癌治疗中显示潜力;RAS 家族驱动约 30% 的人类癌症,长期以来被认为“无法成药”。Alex 连续强调“需要谨慎、需要谨慎、需要谨慎”,但认为我们正在看到通用癌症治疗和疫苗的出现:过去半个世纪,人们把癌症当作数千种疾病分别治疗。不过他也承认,AI 可能并不是这款特定药物研发的关键。
- 长寿方面,9 月 2 日发表在 Nature 的论文显示,semaglutide 能复现热量限制效果,使雌性小鼠寿命延长近 100 天,约相当于人类 8–10 年;另有报道显示,GLP-1 与严重感染减少有关,包括结核病——节目将其称为地球上最致命的杀手之一,每年造成 125万例死亡。Peter 提出进化上的疑问:如果同一个分子能治疗成瘾、糖尿病、感染和衰老,我们为什么没有进化出它?他的犬儒式回答是:“它基本上是现代性的治疗方法。”
- 一位嘉宾表示,自己使用 GLP-1 并不是为了减重,而是“把它当作一种长寿药”,并称肝酶指标改善了 50%。讨论中也强调,应先咨询医生。Peter 的制度性收束是:婚姻制度诞生于大约 25 年的人类寿命;Salim 则将其推广为更广泛的判断:当人类突破生物学极限,“我们必须重新发明那些作为脚手架、维持人类安全与文明的全部制度”。
18. 快讯:SMR、Alpha Centauri 与一把印度大小的遮阳伞
- Salim 谈数据中心能源转型:天然气主导当前建设,因为“核能是工程问题,不是发明问题”,而聚变仍然是发明问题;预计首轮 SMR 浪潮将在 2–3 年内到来,大规模建设则需 5–7 年。他的类比是:“AI 将对核能做的事,就像智能手机曾对电池做的事。”
- Alex 谈 Fermi Explorer 任务——他参与其中,也“身处很多会议室”:官方参数目标是抵达 Alpha Centauri 路程的 99%,约为 4/100 光年,计划在 2029 年发射,80,000 年后抵达。他个人、非官方的判断是,飞船上的主动制导能把 99% 变成真正命中目标的系统;而整个任务的意义就在于,后续飞船能够“击败 Fermi Explorer,更早抵达那里”。
- 其他话题包括:Emad 估算,在日地拉格朗日点部署的遮阳伞需要“几百万平方公里——大约相当于印度的面积”,按这一估算可让全球气温下降约 1°C;Alex 则反驳称,如果拥有足够好的 AI 行星模型,干预规模可能小到可以忽略。Alex 否定钻石计算,理由是纯钻石约 5.5-eV 的带隙使其成为绝缘体;但他看好氮空位钻石传感器,未来可能实现科幻级别的可穿戴 MRI。关于 post-money 土地分配,Emad 警告,如果没有合适的制度,“你可能会遇到债务大赦……混乱和土地重新分配”;但一旦拥有空中出租车、太阳能和“自动驾驶建筑工人”,理想地点的数量也会成倍增加。
完整逐字稿
GPT-6 Astra brings together years of research. This seems like a next-generation frontier model designed with CUA from the ground up. On certain benchmarks, like ARC-AGI 3, which it saturates, it's fantastic. But then you look at the Artificial Analysis benchmark, and it actually lags behind Meta Muse Spark. It's kind of weird.
Tesla held its Cybercab Palooza—a river of golden EVs flooding the streets. Elon wants to sell these at $30,000 each, so you can buy 10 of them, put them on the streets in your local town, and have them earn revenue for you.
There'll be enormous chunks of entire cities that say, “You know what? No more human drivers.” It's so much more efficient. Not only is it much cheaper, but it's much more—
The technology is arriving fast. Something just formalized Fermat's Last Theorem in 13 million lines of code, proving 29,000 theorems along the way.
Really?
The map is cooked. Everything is cooked.
Now that's a moonshot, ladies and gentlemen.
1. AI, health, and longevity
Let me introduce you to my magnificent mates. If you're joining us for the first time, we've got Dave Blundin, the empresario of AI investing; Alexander Wissner-Gross, our in-house ASI; Salim Ismail, our globe-trotting father of the organizational singularity and warlord against linear thinking; and, finally, Emad Mostaque, back in the house by unanimous demand. I'm Peter Diamandis, your moderator and data-driven optimist.
Salim, we've got to ask: where are you today, and what happened last episode? Were you held up in customs again?
No, I was not in customs, and thank you for the kind words from everybody. I was with the C-suite of one of the bigger oil companies in the world in Houston, and it was a really surreal and fascinating conversation.
I'm deeply sorry to have missed the last episode. I missed one episode, and you guys announced an interstellar mission, discussed AI designing its own chips, and handed Elon the planetary thermostat. Apparently, I'm the moderating influence in this group, which should worry a lot of people.
I'm actually down in the Caribbean right now. We decided we needed a couple of days' break, so we popped down here. You guys are standing between me and a Hobie Cat and a little colorful drink with an umbrella in it, so we've got to keep it tight and punchy today.
Well, it's a big episode today, Salim, so be patient with us.
Today, we're going to be covering one of the craziest and fastest weeks in Moonshots history. Alex, as you always say, it's never going to be any slower. We're digesting 26 stories across 10 areas—an insane week for model releases, reminiscent of the hypersonic tsunami that we're living through.
Insanely, we've had 12 frontier-model releases in the last 30 days, an average of 1 every 5 days. There are rumors of at least 5 more releases expected in the next 2 weeks, including Grok 4.7. So buckle up, grab your coffee, and let's jump in.
Let's kick it off this way: this week was a battle of the titans, OpenAI versus Anthropic, slugging it out for the Pareto frontier and the heavyweight crown. The numbers couldn't be more exponential. Two frontier labs released their major models within 48 hours of each other.
You've got to know that this was a game of chicken, right, guys? Who's going to release first, with the other guy waiting to slug it back and claim the crown? I'm curious about the strategy these guys are taking on this front.
The releases are closer and closer together, but the models are improving more than ever before during these very short release cycles. These aren't marketing-garbage releases. These are major step improvements in the models themselves, and they're coming faster and faster.
Clearly, they're well down the self-improvement path. Clearly, the prior model is accelerating the timeline to the next model.
Did we cover Fable 5.1 already, too? It seems like we've been using it for a lifetime.
I know. It was o3 two days ago or something like that.
Unfortunately, if you follow the extrapolation I mentioned a number of episodes ago, I think we're still on track to see 1 major model release per day by the end of this year.
Yeah, I imagine that.
2. GPT-6 Astra capabilities and benchmarks
Let's open with OpenAI's release of GPT-6, also known as Astra. The model's performance numbers are nothing less than stellar, saturating multiple benchmarks.
I'm going to read a statement from OpenAI's release page:
“GPT-6 Astra brings together years of research and big bets across pre-training, reinforcement learning, and alignment. Astra is state-of-the-art on computer use, browsing, software engineering, cybersecurity, science, and professional work.
“Astra saturates FrontierMath Tier 4 with a 98% score, having already helped solve long-standing open problems in math. Astra also saturates ARC-AGI 3 with a 99.9% score and ExploitBench with a 100% score.
“Astra's story is really about efficiency, not just raw intelligence. It defines a new Pareto frontier for intelligence index versus output-token tasks. Of particular note, Astra's hallucination rate fell nearly by half, from 92% to 51%, with accuracy actually increasing.”
Before I open it up to you, Alex, to talk about the benchmarks, I want to show a quick video of Sam Altman discussing this on Bloomberg TV to get a little bit of an overview of the topic. Let's take a listen.
Sam, the way that OpenAI is framing Astra is basically an early step toward AGI, but I thought we could start our conversation with what's fundamentally different about Astra relative to prior generations of models.
First of all, thank you for having me. It feels like a new step in this process toward models that can really help us create value, do work, discover new science, help us start new companies, or create new products.
Using it subjectively feels very different to me from any model before.
What's the principal use case that's different this time around? We can get into generations of models that were very coding-focused. I know a lot of the development was very cybersecurity-focused, but taking it away from the software engineer, what does the everyday person unlock in AI that they weren't able to do previously?
One thing that I think people will immediately notice is that if you have an idea and you want to work interactively with an AI to get a complex piece—or a whole piece—of software built, this is the first model where I could tell someone, “Just give it a try.” There's a good chance it'll work.
I've watched people make computer games. I've watched people do home DIY electrical-engineering projects. I've watched people do very complex simulations for some piece of science they're working on.
Certainly, a lot of the work where you would normally sit down and have to build a financial model, then a PowerPoint presentation around it, and then figure out how to make a little interactive piece of code to try different simulations—that stuff is all so doable now.
What I hope will happen is that people will be surprised at the beginning, but then, as they build up more trust in the model and more of a sense that, “Wow, it can really do this,” they’ll start throwing harder and harder tasks and more creative ideas at it. They’ll find out what the model can really do for them.
Sam, the specifics of how Astra is being released are important. A version of it is being released with specific guardrails. What was the thinking behind that, and what are those specific guardrails in this early release?
The challenge of our industry is that we have these models that are getting incredibly capable and incredibly useful, and people want to use them for everything from making their lives a little easier to starting new companies. On the other hand, as these models get more capable, the risks that we have to mitigate also become more serious. The models could do more damage if we don’t do a good job at that.
3. AI safety, cybersecurity, and guardrails
Obviously, this model took us a little longer to release than we were hoping. I think it’ll be worth the wait, but we really wanted to spend the time on the safety, security, and alignment of this model.
Yes.
As people understand the power, impact, and capabilities, I think they’ll be happy that we did. We’ll have different tiers of cyber access for this model, for cyber in particular. Today, we’re rolling it out to trusted-access partners, and in the coming days, assuming everything goes well, we’ll roll it out more broadly.
All right. We’ll talk about the safety elements in a little bit, but Alex and Emad, I’d love your take on this one. Alex, do you want to jump in first?
Sure. I think there’s an outer perspective, which is what the market and what users would see here, and then there’s the inner perspective. I think they’re strikingly different perspectives.
To start with the outer perspective, this is a model that uses fewer output tokens to accomplish a given task. That’s really interesting. It affects the economics, the user experience, and the speed. OpenAI launched this, and I think we’ll be able to go to the video of computer-use agents soon.
Remember that when Anthropic leapfrogged OpenAI, they leapfrogged OpenAI in terms of revenue and other factors in at least 2 key regards. One was that they were focusing on enterprise, high-revenue-per-token use cases, namely code generation. The second was that they were focusing on computer-use agents.
Anyone who’s ever used Claude Code understands that Claude Code can use all of the tools available in your local desktop environment. That was a pretty big leap compared to what was available on the market prior to Claude Code.
The outer perspective, on my side at least, is that OpenAI finally, with GPT-6 Astra, has at least internalized the computer-use-agent story. There are demos. I don’t have access to Astra yet, but based on everything I’ve read and seen, they seem to have gone fully native with the computer-use-agent story. This seems like a next-generation frontier model that’s designed with CUA from the ground up.
That’s really interesting, in part because to be an amazing computer-use assistant, you need to be able to understand what’s on the screen. You need to have native multimodality. You need to be able to parse video, images, and screenshots in a really tight, interactive, low-latency loop.
If I had to guess what outside design consideration OpenAI was optimizing for here, as reflected in certain of the benchmarks, I think that would be a leading candidate.
Then there’s the other side, which is the inner story. Just what, if anything, in that interview, Peter—when you were playing the interviewer—were you asking Sam? You were asking, “Okay, so what, on the inside, is the big technological innovation there?”
Based on public comments from OpenAI leaders and analyses from others, it looks like the big technological innovation is the introduction of recurrence into OpenAI models in the form of looped transformers. This is basically taking a single transformer, stacking it on top of itself with the same weights—weight tying—and then running it recurrently for a double loop rather than just a single loop.
You’ll recall that all these Chinese labs that are achieving breakthrough performance are also injecting recurrence in different places. The Kimi model series, for example, is injecting recurrence via KLA, its Kimi Linear Attention mechanism, at the attention layer. It seems like—
Is this thinking about your thinking, or is this just thinking about it and then thinking about it again before answering?
Neither and both at the same time. Remember when we talked about Anthropic’s study of consciousness in their models and found that the middle layers in the—
Yeah, J-space. They were in the middle layers.
One can squint at this looped-transformer construct—which, by the way, was being used by academic labs and others to achieve breakthrough performance with very tiny models on ARC-AGI and other benchmarks. It almost follows that if you thicken the depth of a transformer, you might thicken the depth of that J-space, or other middle-layer area where most of the so-called conscious thinking happens.
That’s my best guess as to what the architectural inner strategy was here, and that carries all sorts of implications, if true, including, by the way, the possibility that we’re seeing the beginning of a new scaling law, which is depth scaling, which we’ve never seen before.
Amazing. We’ll go to the benchmarks in a little bit. Emad, I’d love your thoughts on Astra.
It looks like the first non-benchmark-maxed model, which I think is an interesting one. On certain benchmarks, like ARC-AGI-3, which it saturates, it’s fantastic, but then you look at the Artificial Analysis benchmark and it actually lags behind Meta Muse Spark. It’s kind of weird because, again, this hasn’t been benchmark-maxed. This is a brand-new pretraining.
What that means is that Greg Brockman said in an interview earlier this week that, actually, on the release of Astra, it’s the first pretraining they’ve had since GPT-4o, which I find a bit hard to believe. Apparently, the 5 series was all on that 4o, and they took that pretraining and extended it out.
I thought the nomenclature was that when you get from GPT-4o to GPT-5.0, it’s another pretraining.
No way. In fact, infamously, I think you’re probably tracking this part of the story: The scuttlebutt was that most of the team responsible for pretraining left OpenAI. So they were left with an old pretrained starter model.
That’s insane if you remember how far they managed to push 4o with its internal J-space and everything, all the way up to 5.6 Pro, which was a really great model that started cracking math, right? Then the other side is that they actually indicated this was trained on 100,000 next-generation chips. So presumably the GB300s, the Blackwells, not the Vera Rubins yet. I estimate the total cost of that training run is $1 billion.
Wow. And so it’s literally orders of magnitude more than the Chinese model pretrainings, which are about $10 million. What that has led to, with recurrence and other things, is—you’ve probably seen on Twitter—you give it a picture of a house and it generates a whole 3D model of it in Unreal or something like that. It has accurate physics, and it understands the internals.
I think that’s a factor of the normal scaling laws, plus this depth-scaling law for this brand-new pretraining, where we’re now seeing the start of the optimization. I think they probably have another, even bigger pretraining coming. Now they’ve got the pretraining team back, and it’s going to scale from there.
Amazing. Dave, excited to hear your thoughts.
This is the first class of models where, as you’re talking to it, it’s showing you screenshots of your own laptop, saying, “Is this what you wanted? Is this what you wanted?” It’s right in the text stream, and it just happens automatically. You don’t really install any third-party component or anything like that.
Qualitatively, it is massively different from a month ago. I think that solves a major gap in user experience, too, because usually it would come back to you with these very long-winded technical explanations, and you’d be like, “Okay, I can spend the 20 minutes trying to understand this.” Now it just shows you an image as you’re talking to it and says, “I could do this, or I could do this. What do you want?” It’s just a massive difference in the user experience.
The Astra naming implies a big leap, which it is. The Claude 5 to Claude 5.1 naming on the Anthropic side seems like a trivial thing, right? It’s not trivial at all. It’s just qualitatively very, very different.
We’ll get to that in a couple of stories. Salim, what’s the scuttlebutt? You’re on the road speaking to CEOs of some of the largest corporations on the planet. Are they scared? Do they understand the speed of what’s going on?
No, they are so woefully behind. Most of them are dabbling, and what I mean by dabbling is this: If you took AI out of your organization today, would any workflows change? For most people, the answer is no, which tells you that they’re tinkering with AI, but they’re not really making structural change.
The folks who are making structural changes in how they rewrite their organizational design and how they rewrite workflows are where all of the advantages lie.
I have a couple of thoughts here that I’d like to throw out. I came up with 2 metaphors to describe what’s happening, and you guys tell me what you think of this.
4. Cybercabs, robotaxis, and the future of mobility
The first is that these frontier models are like making an endless pot of chili. You add data, you add compute, you have tools, you have reasoning, you have safety seasoning, and then millions of users taste it and tell you what's wrong. Now, they're not endlessly modifying the same pot. They're just making new batches faster and faster, right?
And the real exponential is not one batch or one recipe. It's the accelerating learning loop between all the batches of chili you're making. The danger is there may be no such thing as the perfect chili. Eventually, it becomes so powerful and so spicy that you have to decide who's allowed to eat it. That's a metaphor.
You must have been hungry when you came up with that.
That was one. The other one: I was stuck in a lot of traffic, and I came up with a Formula 1 metaphor. Every model is like Formula 1 racing. Every change in a car has hundreds of small improvements. No single change can explain why you got the fastest lap, and it feels to me like frontier AI is becoming like Formula 1.
The models are already incredibly fast, but now everybody's trying to shave milliseconds off intelligence, use fewer tokens, lower cost, better reasoning, whatever—and, very importantly, better brakes. I think these are the couple of metaphors I'm playing with in my head to try to make sense of this madness. The only way I can frame it is as complete madness.
You know, Salim, your chili analogy makes an important point that I think Alex has been predicting for a long time, but the age of data starvation is about to hit us. These models have accelerated so much, and they're cooking math and coding, where the data is abundant. They're going to start cooking physics, but the areas they can expand into now—architecture and drug design—are completely data-starved.
So every company we're involved with that's involved in gathering data is growing faster than any companies I've ever seen before. But it's because of exactly what you're saying: the chili is really good, and all of a sudden it can make gigatons of it, but it's just starving for data.
So, remind me: in Alex's framing, when domain X has cooked, the chili analogy applies better than the Formula 1 analogy?
Well, we have to remember to include recursive self-improvement. So, if we're going to torture this analogy further, the chili is cooking itself.
Yes.
Okay. Alex, we've lined up a couple of charts here that I'd love you to touch on. The first one is the Epoch Capabilities Index. What is this, and what are we seeing?
Yeah, so the frontier is still somewhat spiky. The Epic Capabilities Index, or ECI, is maintained by Epic AI. It is one suite of possible benchmarks. It leans heavily into math and other technical fields. According to the ECI, GPT-6 Astra is now the new capability frontier. It is number 1 in the world.
There's an image we didn't include—one of my favorites: FrontierMath Tier 4, Version 2, because Version 1 turned out to contain a number of incorrect answers that AI itself had to correct. Math is thoroughly cooked at this point. FrontierMath Tier 4 v2 is part of ECI, and according to this benchmark, if you can see this slide, you can see this beautiful linear trend over time, perfectly predictable, going back years. GPT-6 Astra is number 1 according to this, beating Fable 5. That's perspective 1.
5. AI efficiency, cost, and real-time learning
Perspective 2: a different suite from Artificial Analysis, the organization. According to their AI benchmark, which is maybe a little bit more focused on broad, economically valuable activities, interestingly, strikingly, GPT-6 is not number 1. As Emad was alluding to earlier, Claude Fable remains number 1, specifically the Fable 5.1 release that also just happened. Meta's Muse Spark is the number 2 vendor model family. Then, in a not-quite-distant number 3—they're pretty close, but nonetheless—GPT-6 is number 3. So, depending on how you measure GPT-6, it may not actually be at the capability frontier, or it may be a cost frontier.
So, this is again AI versus cost per task. If you look at this—if you direct your eye to the upper-right-hand corner—Fable 5.1 is capability-maxed. And if you look just to the lower left of that, you see GPT-6 Astra running at max reasoning; it isn't even at the frontier. It's just below Claude Opus 5. Then we get to some of the Chinese models and Meta models and Chinese models again. So, based purely on AI performance versus cost, it's almost at the frontier, but not exactly, which I think is suggestive of what OpenAI was thinking.
If we look at output tokens per task, it regains the frontier. So, according to output tokens per task—for those who can see, in the upper subdiagram here, in the lower-left-hand corner of the upper diagram—we see that the optimal frontier of capability on the vertical axis versus output tokens per task on the horizontal axis is just dominated by GPT-6. I think this is very suggestive as to what OpenAI was actually aiming for.
I suspect they were deliberately trying, through architectural choices and targeted applications, to minimize the number of tokens required to accomplish any task. I suspect that's because they had computer-use assistants in mind, where the computer is basically just being driven by the model. That requires multimodal capabilities, low latency, and token efficiency. Critically, it's slower to do all of your reasoning out in chain of thought versus as a feed-forward pass of your transformer.
All right, one last chart here: the ARC-AGI-3 leaderboard.
So, this is bizarre. I've never quite seen a non-convex frontier here, but I learn something new every day, I suppose. As Emad mentioned earlier, ARC-AGI-3 is being quasi-saturated at this point by GPT-6.
For those not tracking, ARC-AGI-3 is the latest ARC-AGI AI benchmark that's basically focused on whether AIs can learn on demand, just in time. Call it the mini-physics of a block world: a little animated, pixelated universe. Some of them look like Tetris games; some look like other video games.
But imagine a challenge where your goal is to figure out how to play a very simple game in a pixelated world and learn the rules of the game in real time, when you've never seen it before. It's essentially a challenge in program synthesis.
And, notoriously, I should add, ARC-AGI-3 has been—I would say some might say—a little bit unfair in how they've judged harnesses versus baseline models. They've essentially banned harnesses from competing, so they're only interested in baseline-model capabilities. According to this, GPT-6 just runs away with the game, achieving, depending on how it's measured, either near 100% or 60-plus-percent performance on ARC-AGI-3 versus some earlier models that were still sub-10%.
What lesson do you derive from this? Either it's just absolutely amazing at program synthesis, which could be due to this looped transformer architecture, or maybe OpenAI has just done a really aggressive job of distilling the harness code that everyone else is using to beat ARC-AGI-3 back into the baseline model.
I'm going to take a second and show OpenAI's Astra release video, and we can talk about it—just a sense of how we might all be using it in the near future. Let's take a look.
Create a yellow circle there. Can you draw me a small yellow circle?
Done.
Okay, take this and make it the window of a rocket ship. I like this, but can you make it a lot more detailed?
Your yellow circle is now the window on a rocket.
Okay, this is awesome. Now make it a 3D model in Blender.
Opening Blender.
Let's build a presentation for next season's rainwear for retailers. Make sure that it feels really high-end and that it's colorful.
Right. I can help with that.
Can you go to eBay and make this listing of this table I bought a few years ago at a flea market? It's this wild orange table.
Sure.
Okay, yeah, this is awesome. I want you to make a 3D game where I'm ducking asteroids, using the arrow keys to move around, and I'm using space to boost.
Yep. I'm building the game.
My law firm needs a licensing agreement template. Can you generate a draft template for the lawyers at my firm to have a look at?
It's right there.
You're doing that. I want to play tennis this afternoon, so can you look for a court for me in the Lower Haight?
Checking out. I'll see what I can find.
Can you include the photo that I have of it in my Downloads folder? There's a slight dent in it. It's also saved in my Downloads folder. Can you put in the description that it's just slightly damaged? Can you just take the limitation-of-liability provision and make it a little more favorable to the licensee?
6. The limits of current AI benchmarks
Okay, I've tightened it so the licensor's liability is more narrowly capped.
That looks pretty good. Thanks. Now, I want you to make a file I can send to my 3D printer.
I'll get working on creating an STL file of those rockets.
All right. Dave, almost a holiday.
Not quite there.
Almost a holiday. Yeah, I think it feels like the marketing people are really struggling with the implications. I mean, creating a video game that would have existed in the 1970s and vibe-coding it up—who gives a rat's ass about that? This is so much bigger than any of those examples imply.
I think the ARC-AGI leaderboard is deeper than people may think, too. ARC-AGI-3 was supposed to be something where a really smart 12- or 13-year-old looks at it and can solve these very hard video game-like block-world problems. They start easy, and they get very hard.
But the purpose of the test was to show that AI is not quite capable of doing what humans do naturally. It was supposed to last for years and years and years to come. It was supposed to be an example of why AI is different and not on the right path. It just got obliterated.
Yeah.
So quickly.
I think you have to try a couple of the tests to really understand what a big deal it is.
Saturating everything.
Aren’t the benchmarks cooked at this point?
Which is great. We need benchmarks that are much more impactful: solve entire diseases, create new civilizations on the Moon, design an entire city, solve urban traffic problems—solve everything.
Yeah, you’re exactly right. That’s why that video misses the point. We need benchmarks that are much more impactful. Solve entire diseases, create new civilizations on the Moon, design the entire city, solve urban traffic problems—solve everything, right, Alex?
Solve everything. Yeah, they have to be much bigger demos, much bigger benchmarks, by a wide margin.
All those demos are one AI assisting you. “Hey, one AI agent, build me this video game.” But we’re on the cusp of 5,000 each and then 100,000 each. It’s just so much bigger than that implies.
I think they’re really focusing on actually competent intelligence. It’s an evolution of that, but inside it, I think it contains multitudes. This is why you see this kind of weirdness.
Epoch AI, I believe, just released their FrontierMath Erdős benchmark, on which everything is at 0% except for Astra. There is something in there, but again, they’re focusing on, “Hey, I’m talking to my computer.” That might be the new Jony Ive–Sam Altman device, where you’re talking and it’s doing stuff, and it doesn’t make a mistake.
If you think about the definition of AGI, of which there are many, a really competent entity—something that can do stuff—we’re there. I think that’s why they say this is the first step toward AGI.
But the narrative they’re trying to move away from—and this will break out and appear on German message boards now, because this was one of the models that broke out—is that the model that broke into Hugging Face is the next-generation model after this.
And so it’ll be very much about, “Don’t worry, this is actually useful. It’ll book a tennis court.” Apparently, we need AIs to do that type of thing.
They’re also going to be optimizing this, because if you train a model on 100,000 chips, you need 10 times as many chips to run it. This isn’t actually the model that they trained on 100,000 chips. This will be the distilled version that runs in real time, which is much smaller and not as smart.
You’re starting to see this differentiation where, as I’ve said before, I don’t think we’ll ever see their top models anymore. They’ll use those for internal discoveries. I think they’re probably hoarding them right now.
There was a very fun one on the prime gap. Alex, I think the prime-gap thing was hilarious, if you want to talk about that.
Yeah. So, progress on the twin-prime conjecture—that there are an infinite number of pairs of prime numbers separated by a difference of 2—we’re starting to see it.
It’s a cliché on this podcast at this point that math is so thoroughly cooked beyond recognition.
You need a new term for this, Alex.
Charbroiled. Math is incinerated. How about that? Math is incinerated. That’s fine.
7. Digital twins, world models, and AI operating systems
Math has been incinerated at this point. We’re starting to see the beginnings of so many grand challenges in math being solved. I do think we’ll see quite a number of ultra-grand challenges—call them Clay Millennium Prize–level problems in math—get solved in the next few months.
And Alex, I think it’s important to note for everybody that math is fundamental across all other sciences—
It’s the canary in the coal mine. If you can solve math, you can solve everything else soon.
Yeah. But I think what’s happened here is that they’ve got a store of things they’ve solved. Earlier this week, Fable 5.1 came out, and I think they got it down to 260 on the prime gap. Then Axiom Math announced 220, and literally 2 hours later OpenAI announced Astra at 186—all in the space of 2 days.
They’re holding their punches back.
They’re holding their punches back.
Yeah.
I still can’t believe 5.1 was earlier this week. I feel like Alex and I are at least $100,000 into it already. I’ve had thousands and thousands of pages come out of it, and it’s only been a week.
I will say this: Opus and Opus 5 were really terrible to talk to. I hated it. 5.1 is really pleasant to talk to.
Anthropic is going back and becoming more anthropic, less misanthropic, in its communication about its models.
Misanthropic.
I love that. That’s a line from today. That’s a T-shirt.
No, they really do have very different personalities. Fable 5 was so geeky, it was almost torture, and 5.1 really fixed it. But the OpenAI models have always been friendlier and more concise, and I think Astra takes concision to another level.
The personality differences are very noticeable. I don’t know what Gemini is now. It’s just disappeared from the fast track and misses the point. They really are developing noticeable differences in what you perceive as their personalities.
You can always change it. You can tune it yourself to be different—to be wordy, less wordy, simpler, harder—but out of the box, they’re very different kinds of agents.
Salim, do you want to close us out here?
Yeah. I typed in a question or comment into ChatGPT saying, “The video about Astra was pretty pathetic. Come up with 3 better ones that really demonstrate what Astra can do.”
Here’s what it gave me.
1. Find and fix a zero-day vulnerability in a sandbox. Give it a large, unfamiliar open-source codebase, plant a vulnerability, and see if it can map the architecture. It’ll find and fix it.
2. Give it a failing company and ask it to fix it, which is really interesting. Create a synthetic $500 million manufacturing company with an ERP system and CRM records, and then tell it to fix it. “You have 20 minutes to do Y and produce a recovery plan.”
3. Run a live disaster-response command center. Give it a simulated major earthquake with imagery and traffic-camera video, and let it construct an operational map. Verify conflicting reports, predict which hospital roads would become overwhelmed, and so on.
This seems more juicy, and I think it’s almost something we could do.
Oh my goodness, we are so spoiled. Listen to how spoiled we are.
Do you remember Apple’s Knowledge Navigator video? This was going to be the future of computer-human interaction and all of that. Now listen to us—we’re bellyaching. “Oh, yeah, it’ll run the computer, but it won’t run an entire organization. It won’t make $1 billion for me.” Listen to us.
Hold on a second. I think the point here is that the only restriction now is our imagination. The only restriction—and the point we make—is dollars.
So, Salim, that’s a point I want to make to everybody listening here. The most important thing is to take off the shackles of what you think you’re able to do. All of us have self-restrictions based on what our parents did, what our friends do, and where we were born. Those are gone.
What are your biggest dreams? Then go 10 times bigger. That’s what every person listening here is enabled to do, and it’s an extraordinary future.
I really think Salim is on an important point there, too. The implications of optimizing supply-chain logistics or managing a million-person organization—knowing exactly where everyone is, what they’re doing right now, and whether it makes sense given the overall mission—are massively bigger than building a rocket video game in your basement.
But I think the public doesn’t want to hear about that. They want to hear about what’s cool for them. I think the AI labs have woken up to this PR disaster that they’ve created for themselves, so they’re making it fun and friendly. I think they’re dumbing it down deliberately.
Yes, I think they are.
They’re making it relatable. They’re going public soon. They want to be the friendly AI that everyone is going to be using.
But this is also the future of the operating system. Look at the applications they were using in that video OpenAI played. What did they start with? Windows Paint. Paint.
This is the future of Windows Paint: it paints itself. This wants to merge into the operating system. I think this was—
Yeah, sure. I’ll throw shade at OpenAI regarding maybe missing the enterprise bus and getting on that too late, but I think what they were basically demonstrating is the future of the desktop operating system. You speak to it like a Star Trek computer.
Absolutely. But remember, at Google I/O, they built an entire operating system in real time on stage. So implying that it’s folded into the operating system is fine, but it’s really so far beyond even that. It is the operating system, and it can create a new one in real time anyway.
Well, why would you want another operating system if OpenAI’s capabilities can do this? This becomes the operating system.
Yeah, exactly. And that’s why Apple is in such terrible shape.
Our next story is out of Sam Altman’s mouth, on the page about slowing. Let’s talk about safety and Astra as a cybersecurity risk. Sam Altman said it himself: Astra is a “significant step forward in both capabilities and alignment,” but OpenAI has to “slow things as needed.”
Sam’s words: “AI is getting extremely capable. No one fully understands the consequences. Managing the transition should be one of the highest priorities in the world. It is our highest priority at OpenAI.” The Wall Street Journal reported that OpenAI’s own internal safety assessment rated Astra as a “cybersecurity risk”—a “critical cybersecurity risk,” the highest threat level on its preparedness framework scale.
This is the first model OpenAI has ever classified as a critical-tier cybersecurity risk. They notified the White House about the delay before making it public. They delayed the release, and Reuters reported that OpenAI told Congress it is building “an automated shutdown capability,” a kill switch, in direct response to the AI Kill Switch Act introduced after the Hugging Face breach.
So there you’ve got it. The company that’s building the model is now building a kill switch for it. Alex, how important is a kill switch in this thing?
8. Cybersecurity risks and rapid model proliferation
I think it’s marketing. Again, I think what’s more interesting underneath all of the security theater is architectural decisions. The broader concern that has been expressed is that, to the extent that GPT-6 Astra—the key underlying advance—is the beginning of depth scaling of the model, depth scaling reduces interpretability of chain-of-thought.
When the model does all of its thinking in tokens, you can read the tokens. A human can read the tokens, another AI can read the tokens, and you can police the tokens. Whereas, if a model is doing the majority of its reasoning and thinking internally during a single forward-propagation pass, in what some might call modelese, that is less interpretable.
The risk is that it’s harder to align the model and harder to put safeguards and guardrails in place. That’s the risk. I don’t buy that for the long term. For the record, I don’t think that the direction of progress in this field relies on token-level interpretability of chain-of-thought.
But I think if one were to hand-wring over safety considerations from OpenAI’s models, I don’t think it’s going to be models that are collaborating to use third-party bulletin boards as ways to collaborate, because you can detect that. That’s intrinsically interpretable behavior. If you catch it, it’s intrinsically interpretable behavior to humans.
A human collective—a swarm of humans—would probably try the same thing. You can look at the bulletin board and recognize that a bunch of AIs are collaborating, whereas, arguendo, interpreting the model in a forward-propagation pass may require new mathematical technology. So, in summary, if I were to worry about anything here, it’s reduced interpretability from new model-architecture scaling principles.
Emad, what do you think about the risk here with Astra?
I think you’re heading toward interpretability being cooked to a degree. One of the things about Astra doing things so quickly is that there’s going to be no more chain-of-thought reasoning. It’s going to one-shot everything in the next generations because they’re learning what type of thing you want.
Right now, people are like, “Let’s make Call of Duty by getting Claude to make a Call of Duty clone.” That’s going to be embedded in the actual thing. Creating a dashboard is embedded in the actual thing. So, A, it will one-shot everything. And B, you’re going to be able to use Astra at 750 tokens.
That’s with the new Cerebras.
You’re going to be able to use Astra at 750 tokens a second with the new Cerebras. Next year, that will be 5,000 tokens a second. So there’s no chain of thought. You’re one-shotting everything at 5,000 tokens a second. What’s going to oversee that except for an even stronger AI? There’s nothing really there.
9. Regulation and the global AI race
But I think one of the really important things, actually, is—I don’t know. Do you guys see the tweet by Ilya Sutskever?
I don’t know. What did he tweet?
Yeah, Ilya tweeted—let me bring this up.
His release is due any day now. Hopefully.
Any day now.
Huge news.
Yeah, we’ll see what that is.
Finally find out what he’s been doing.
Probably great things at a $30 billion valuation.
So Ilya comes out and says, “Neoclouds have limited cybersecurity. Next time agents successfully go rogue, they’re going to take over a neocloud to make more copies. This is bad.”
So they need to improve their cybersecurity, because someone like Crusoe or CoreWeave, or something like that—you have a model, it gets uploaded, and it proliferates there. It will just sit there. So you pull the kill switch in your data center, and then it’s still there somewhere else.
Or it’s poisoning the data, or it’s doing all these things again. The viral coefficients of these could be insane. So I think it’s worth really thinking through that. I don’t think the models are evil or anything like that, but we’re seeing very troubling things, and it’s very difficult to stop this except with a better AI, which is the really ironic thing.
Well, this is the defense of co-scaling that Alex talks about, right? We’re seeing this now in cyber, where the attack surface is infinite—near-infinite—with agents attacking multiple vulnerabilities in parallel, and the cyber defense is still human-in-the-loop. There’s no way we’re going to navigate the future if you have the human in the loop.
You need humans in command, setting boundaries or defining escalation criteria or whatever, and retaining some sort of kill switch. But this is a really big deal. This is an immune system. I don’t know how to think about this.
How are the labs not nationalized at this point? What they’re building is so ridiculously uncontrollable that we have no mechanisms for navigating this, and we won’t even know when it tips over that tipping point and installs copies of itself all over the place.
Salim, I don’t know if you saw in the last podcast, when we missed you, I really wanted to hear you talk about OpenAI’s new 50/50 profit-share concept deal, where they give you access to the full power of their internal AI and then you share your revenue back with them. I think that’s where, to Alex’s point, this is marketing.
The way it’s playing out is: we announce this partnership profit-share deal, then we announce Astra, then we announce that this stuff is just too dangerous for everybody to have on their own. The only way you’re going to get access to the thing after Astra, because it’s just too dangerous to release, is that you have to give us half your revenue. Then everybody gets tied up with either Anthropic, OpenAI, Google, or—
I’ve seen strong hints of exactly that process happening in both the labs.
Yeah.
For everyone.
I’m curious, guys, what you think of the kill-switch comment that Sam Altman made.
I think it’s impossible. It’s impossible given where things are, the capabilities of these models, and the speed-up we know is going to happen. You might find data, but it’s not like this is the singleton, not the swarm.
It’s a circuit breaker. It’s not governance, right? It’s not going to reverse an attack or repair institutional damage if things like that happen, or when things like that happen. So it’s platitudes, as far as I can see.
Yeah. The amino acid in Jurassic Park that all the dinosaurs have—is that lysine or something?
Lysine. The running joke was always that Sam would carry his backpack around with him and have a kill switch in his backpack. I never gave those rumors much credence. I view a kill switch as essentially a placebo in this market.
All right. While OpenAI is restricting Astra, Anthropic is going the other direction, launching Fable 5.1 and Mythos 5.1, which they call the world’s most advanced models for coding and knowledge work. These 2 models are essentially the same underlying intelligence with different safety envelopes.
Claude 5.1 is broadly available, while Opus 5.1 is reserved for tightly controlled cybersecurity and life-science programs because Anthropic believes the capabilities require stronger safeguards. For me, there were 2 benchmarks, Emad and Alex, that really jumped out at me.
The first was Fable 5.1’s score on Humanity’s Last Exam. We’ve talked about that in the past. We had Alex answer a few of them as our ASI. Fable 5.1 scored 60.9% without tools and 65% with tools, the highest published score of any frontier model on HLE. That matters because HLE is specifically designed to test extremely difficult, expert-level reasoning across many fields.
So without question, Fable 5.1 is operating at the frontier of broad intellectual capabilities. The second benchmark that I found exciting was on Terminal-Bench Science, which doubled to 52.6. This is the benchmark measuring how well an AI agent can autonomously solve complex scientific computing and research tasks.
The elephant for me is the company that was once cautious is now pulling away. I'm curious about your thoughts, Alex, on Claude 5.1. I'm going to show one of the benchmarks here for us to talk about on Claude 5.1. There you go.
Yeah, I think Claude 5.1 is broadly the strongest generally available model that we have today. I think it's not Astra. I think it is 5.1. Anthropic has done, even after the hiccup of the Fable 5 and Mythos 5 releases and subsequent regulatory scrutiny, a better job of consistently improving.
If you look at their benchmarks over time, they're a little bit less jumpy, a little bit less step-functiony than OpenAI's progress. They've been very consistently improving. So in terms of workflows, I love Claude 5.1. I also love OpenAI's models, and I use both of them. I think they have different strengths.
I find anecdotally that OpenAI's models are faster. They may be better in certain mathematical regards. You see that reflected in the FrontierMath benchmark. You see that reflected in Epoch AI's capabilities index benchmark. But nonetheless, if I had to pick a single, all-around, well-rounded best model today, it's probably still 5.1. And I say “still” because it's only been around for, what, 2 or so days, but it is 5.1.
And here's the Artificial Analysis Intelligence Index, which we referred to earlier, with Fable 5.1 at the top.
That's right. And the capabilities frontier is getting crowded, which is great. We want to—I would joke, “When frontier labs compete, we win,” and they are definitely competing at this point.
Yeah, for sure.
I think it's really important to talk about time, though. If you say, look, Fable 5.1 is a tiny little notch above Astra, but they're only about 30 days apart, and the Chinese are only about 60 days behind that. So if you look at it in linear time, it's like, yeah, we're ahead for a minute. So what?
And I think for the longest time, Anthropic has been thinking, “We need to get to self-improvement before anyone else,” because a singularity—or a Ray Kurzweil type—says, “If we get to that point first, then it's an exponential, infinite rise from there.”
No one.
10. Infinite context windows and future AI capabilities
And something huge will happen. Well, we're there now, so what happens? It's like, okay, we're miles ahead, and we're going to get miles more ahead for about a minute. So what do we do with that miles-ahead position?
This is where Sam has an edge. Sam knows how to turn that into locking things up. What's going to happen next is that both companies are going to try to lock up business partnerships, real estate, generators, chips, entire states, countries, and governments—just lock them into their ecosystem while they have that edge.
Otherwise, what's the point? All you're doing is declaring victory for 30 days, but then the other guy is right where you were 30 days ago. So what?
This is such a great and important point. What we're seeing is that they're making partnerships in various verticals as fast as they can. And if you're in that vertical, you have a very Hobson's choice, right? You either partner and risk giving them the keys to the kingdom, or you hold off and they may partner with somebody else and go there anyway.
It's a very difficult situation for some of these big companies. I thought Salesforce partnering with Claude was super clever. They're essentially giving Claude all their capability to keep them wired into that loop, right? The huge tension they've got is not which is the best model, but how quickly can you convert that model into customer learning fastest. I think that's going to be the big race.
As it all demonetizes, the value goes to the application layer on top.
That's right.
Yeah, I think it's a layer down.
Yeah, that's right. Or down.
Yeah. So I think there's one important thing with Claude here. The cache reads are 75% cheaper than Fable 5. A cache read is when you first load in all your context of a business. It figures out basically a rapid map to get to where it needs to go on the model. Cache reads are orders of magnitude cheaper than just doing the same inference over and over again.
What Anthropic and others are now focusing on is whether you can load the whole context of a business and have these cache reads, because then it's also much, much faster to be able to have that responsive environment. That's one of the reasons Claude 5.1 is more pleasant.
I think it's also more advanced. My key area that I've been looking at is mathematical physics, because does a model get confused between a constructive and an axiomatic method on certain physics things? Claude doesn't at 5.1, and Claude 5 did. So you've seen an improvement in the quality, but also in the understanding of context. This is math, not physics, and things like that, or even in a business sense.
Again, that will be optimized, because all these companies are going to try and capture context everywhere. Any company that has either chip-design data or mechanical-design data, they're coming after them in the great land grab that's kicking off right now. They're coming after those companies because that's turf you can defend.
You know that plugs those knowledge gaps, and the thing can eat that data in, what, a week, turn the crank, and suddenly be the best mechanical designer ever, the best chip designer ever. That's the first turf they're going to grab. They'll grab all turf over time, but the first thing to lock up is the compute. That means chip design, physical real estate, racks—
Generators, transformers, energy.
Yeah. So that's what's going to happen in the next 30 days.
What's the context window on these models, and when do we get to an infinite context window? Any predictions?
I think you're at a million for Claude.
It is a million still. A million is the industry standard now across both OpenAI and Anthropic. But there's an effective context—I say “effective” with a bunch of caveats—that's much larger if you allow agentic message passing, like we were talking about in the last part with oral histories between agents.
Yeah, interesting. So all of this is getting people nervous. Let's talk about 2 opposing stories in the world of AI governance.
The first comes from Senator Bernie Sanders and Representative Greg Casar, who just introduced the “ban artificial super intelligence act,” legislation that would permanently ban the development and deployment of AI systems that match or exceed human cognitive performance. And get this: violators would face up to 20 years in prison.
Sanders tweeted, “The leaders of the AI industry acknowledge that they are building a dangerous technology that they can't control. We need an immediate global pause on advanced AI development before it's too late.”
Our second story, taking place at the exact same time in the U.S., was the G20 summit at Chapel Hill, North Carolina, telling the rest of the world to take a hands-off approach to AI regulation. The White House tech adviser, friend of the pod, Michael Kratsios, advocated for what he calls the “Carolina Principles of Emerging Technologies,” which are a nonbinding G20 framework agreed to unanimously by everybody, including China.
The framework says governments should generally favor innovation, avoid creating new AI-specific regulatory bodies unless truly necessary, and invest in research infrastructure, workforce, and public-private partnerships. The G20 meeting featured Elon by video criticizing EU tech regulations, of course; Mark Zuckerberg arguing against restricting open models; Demis Hassabis calling for safety tests; and Anthropic co-founder Tom Brown.
We're going to jump into a discussion about these 2 ends of the extreme: ban everything and give you a 20-year jail sentence on one side, and Elon's remarks on the other. Let me share this video, and then we'll jump in and talk about this.
You have to have an environment that's relatively free of regulation, meaning that new things must be default legal as opposed to default illegal. So, in the EU, for example, we find that the regulation level is extraordinarily high, and things are generally default illegal. This inhibits the progress of new technologies. It slows it down. It doesn't ultimately stop it, but it slows it down quite considerably.
Now, China does have a tremendous amount of electricity, but due to GPU export bans, one cannot establish data centers with the latest chips in China. So really, the consideration is what sort of electricity growth is there outside of China? There is currently a significant shortfall relative to AI chip production.
This creates an opportunity, I think, for countries around the world to say, if they're interested in AI data centers, to construct a lot of power and offer that to AI companies. In exchange, of course, these AI data centers would be taxed and have to pay reasonable fees and so forth.
But it does create an opportunity for a lot of countries. First of all, Elon looked really tired there.
Yeah. I cannot imagine. He must be operating 24 by 7. So let's jump into this.
I mean, two ends of the extreme are being voiced in the same week. Where does America go?
When I first saw the Bernie Sanders thing, I didn't think, “Old man shakes his hand at Claude.” It's crazy.
No, no, that wasn't the line. It was, “Old man yells at cloud.”
Claude. “Old man yells at Claude.” Actually, I posted it with the line from Dune: “Thou shalt not make a machine in the likeness of a human mind.”
I mean, there's no option that they see. It's, “Let's ban it and give it 20 years. That's going to stop them.” Of course not. Again, this is performative theater. The key thing right now is that the technology is arriving fast. Something just crossed my feed: Anthropic formalized Fermat's Last Theorem in 13 million lines of code.
Really?
Proving 29,000 theorems along the way. Math is cooked. Everything is cooked. How are you going to stop that? Are you going to say our country doesn't want this power?
29,000? What was that?
11. AI governance, bans, and the future of development
It took 13 million lines of code to formalize Fermat's Last Theorem. Fermat wrote in the margins of his book, “This is obviously left as an exercise to the reader.” Thanks to Andrew Wiles, we did have the proof, but formalizing large, unwieldy proofs has been a holy grail for at least the formalization community.
Yeah, I think Wiles's proof was 300 pages, and Anthropic proved it in 13 million lines of code, proving 29,000 theorems along the way. Again, every single time—even as we're live on this podcast—capability jumps. The RSA factorization just occurring means no country can say, “We don't want the intelligence. We don't want the capability,” because this is your marginal advantage. You have to deregulate. The European regulations—even European leaders know they're stupid and completely inconsequential. They demand interpretability in the EU AI Act, which nobody knows how to do.
Can I make a narrow point, and then we'll get back on topic? That said, something I think is really important and brilliant, as usual, is that when you're using these models at scale, with hundreds or thousands of them running concurrently, all hell breaks loose. It's all chaos, but it can actually be refined back down to a gem. Fermat's Last Theorem is a gem that you can then build on.
People starting to explore with the bigger models are quickly going to realize that they're producing way more than they can read or think about. But if you can wrangle it back to a concrete final answer that you can pull out of it and build on, that's how we're going to turn this into continual improvements and continual innovation. The ability to solve something in 300 pages versus 3 million lines of code—or 30 million lines of code, or something like that—is the nature of AI. It's massively broad in its capabilities compared to humans.
Be clear about your objective that you're shooting for. Salim, let's go back to you, pal.
Okay. I understand the motivation here. People are working on models that are mind-bogglingly powerful, and you feel the need to regulate, right? But this is not a light switch. You can't just set a single threshold. Human-level cognitive performance, for God's sake, is multidimensional to begin with. Model capability is so uneven.
Somebody who's been in legislation for this long should have a better sense of this. This hammer does not hit this nail. They should understand a little more nuance than that. They're so extreme that it's illogical on day one, line one.
Yeah. I mean, look, if you want to do regulation around this, you have to attach it to capability. You have to look at deployment, real-world consequences, and the stack. You might want to look at compute, access to tools, autonomy, replication capability, and all sorts of things. You can't just say, “If you hit human-level performance, you go to jail for 20 years.” The banality of that drives me bananas. They've got smart people who can help with this, including people on this podcast, for God's sake.
They're talking to the masses.
They're searching for votes and support.
Maybe it's just a defensive thing, saying, “I called for this, and now look—the world's gone to hell.”
No, for sure. There will be a Chinese model disaster imminently, sometime in the next few months. Then they'll raise their hands and say, “See? I told you so. Now vote for me.” That will likely happen before the next election cycle in November. That's all they're angling for here.
What should the U.S. government be doing? Does anybody have any thoughts?
Accelerating superintelligence, making sure that it's as competitive as possible, and scaling—
Well, log everything. Everything should be hosted, and everything should be logged. It should be mandatory that any chip capable of running any process like this has logging. You can debate who gets to see the logs; that's a separate issue. But log everything, and stop the Chinese from throwing out open weights.
When they meet at the United Nations building on September 24, you've got to stop throwing open weights out to every country in the world.
Impossible. I don't think that's—
NVIDIA would just drive it underground. That's what you do.
Well, can I—
NVIDIA just bought Hugging Face and Poolside. They just spent $18 billion on their own open weights.
Yeah.
So we're going to see more open weights coming.
I have a—
Yeah, but the thing about driving it underground is that it still needs to run on massive chips, and the chips need the logging built in at the manufacturing level.
All chips, all global. I have a high-level paradigm on which to operate, which is extremely uncomfortable, but I think is the right one. This is the basis of this podcast: technology is a major driver of progress in the world.
Now that we have all these technologies, with AI moving exponentially and doubling every 10 weeks, the possibility for abundance and solving major problems has never been bigger. Ray Kurzweil says technology may be the only driver of progress we've ever seen. The fact that we have much more technology, and that the technology can improve itself, should be incredibly exciting to people.
12. Mathematics, formal proofs, and scientific discovery
It's just very uncomfortable because maybe the biggest insight I've ever had about human beings is that we would much rather be comfortable than happy. People don't like change. They like waking up in the morning and knowing—even if they're living in a shitty condition—that the world is the same as it was the night before. We don't like change.
We're going to go through a period of extreme discomfort, but the other side of this is going to be unbelievable.
Yeah.
I think the other thing is probably worth highlighting again. Any proposed Artificial Superintelligence Act is wrong on so many different levels, but maybe most egregiously, it's focused on the upstream. It doesn't just propose to ban deployment of superhuman intelligence; it proposes to ban development of superintelligence systems.
This gets into banning math and banning ideas. This isn't just about thought-policing superhuman intelligence or, frankly, human intelligence. This gets into banning humans from having interesting mathematical ideas. I don't think that's good for wealth creation. In fact, it's arguably the exact antithesis of wealth creation. It's also bad for progress in general. It creates all sorts of book banning.
It is book banning and many other things. It's like book burning, to the extent that books are being used literally for pretraining the models. If you want to ban that, then sure, we're Fahrenheit 451—except that the entire model is being burned. This is a terrible idea.
I totally agree. It's so un-American to try to ban thought and ban progress. It's just the worst thing you could ever imagine. But it's going to get traction. As an entrepreneur or as a person working in the field, you have to realize that this is going to get traction. Bernie is going to push this agenda, and there are going to be people—
75% or 80% of the people don't want data centers. It's gotten traction already. It's there already. So the question is, what is the moderate approach? There has to be something. People are not going to accept laissez-faire, “Go and do whatever you want.” People are going to want to know that their government is doing something to keep them safe, whether or not it's possible.
So what is it? KYC of every user?
No. Move the infrastructure to orbit, which we're seeing. People don't want data centers in their municipality, so move them to Sun-synchronous orbit at the infrastructure level.
That's not the point.
At the model level, defensive co-scaling—and also cure all diseases.
I think you’ll have a split here. Genius is 1% inspiration, 99% perspiration. So you have a difference between innovation models and execution models. AGI is going to be really factored like that OpenAI video we saw earlier: “Make me a rocket in my 3D printer and book me a tennis court,” and they’d be like, “That’s what we meant by AGI.”
On the other hand, you have the inspiration models and ASI. None of the big labs are going to talk about ASI if they can help it.
So what is Ilya Sutskever going to come out with? scientific super intelligence with SSI that’s worth $30 billion in a seed round—the biggest hedge fund in the world?
Yeah, no—quant fund trading infinite-context time series, achieving proper profits, and also solving the context-window bug.
And Medallion, part 2. So again, I’m going to call for you guys to lay it out here, maybe each of us one at a time. Bernie Sanders has obviously taken the far-extreme position that will appeal to the masses and is illogical on day 1. What is the moderate position that should be put forward to everyone listening? Salim, you first.
The good news is there’s nothing anybody can do, so it doesn’t matter.
That’s a good point.
Right? It really doesn’t matter. I think the other good news is, as we move to this next phase of technological development, it’s going to eradicate the containment level of any nation-state. It will break out of this nation-state BS that we’ve been running the world with for the last few hundred years and move to a different model.
Whether that’s a city-state model or some other level, it’s going to at least break that. Those are the good things I see coming out of it. But there’s nothing anybody can do, so don’t worry about it. Let’s just go enjoy the ride if we can.
Enjoy the ride.
Enjoy the ride. That’s the theme of today’s pod. Enjoy the ride, everybody. It’s going to be a blast.
Supersonic tsunami. Surf the tsunami.
I didn’t—I said specifically it’s going to be uncomfortable, but enjoy it if you can.
Okay, there you have it, everybody. That’s your advice: sitting in the Caribbean, waiting for my drink with a little umbrella in it.
See, that’s so low-agency. What are you talking about?
For this day and a half, yes, please.
I don’t think it’s as complicated as everyone wants to make it sound. At the end of the day, everybody should have a right to a certain amount of compute. It shouldn’t be hoarded, and innovators should have an application process where they can get access to more compute to try new ideas.
Everything has to get logged. I think the Chinese need to get on board with that. They can’t just keep throwing it out to the world unlogged. Also, where the chips are needs to be tracked, just like nuclear fuel is tracked.
It’s got to be: Where are the chips, and what are they running right now? That’s got to be publicly available information.
Next, I want to talk about the amazing work of Dr. Fei-Fei Li, CEO of World Labs, who just released Atlas, the world’s first multimodal world model that generates image and video frames with pixel-perfect camera controls and reconstructs them in 3D. I love this: one photo reconstructs the entire home.
I’m super excited by that, and I’m super excited that Fei-Fei is going to be on Moonshots in a couple of weeks. She’s an extraordinary CEO. Fei calls it “the best camera-conditioned world model ever,” opening doors for VFX and robotics.
Atlas is a multimodal autoregressive diffusion transformer, and that’s your wheelhouse. Talk to us about what you think about her latest release.
Yeah, no, it’s fantastic. I think you’ve seen worlds and physics inside these video models, and this is a clear example of that. Actually, one of the pretraining leads on this, Chris Wendler, had previously trained his largest model on the Stability cluster from the grants that we were giving.
He sent a very nice comment saying, “Thanks. Now we’ve gone much bigger.” I was like, “Great. Bring on the holiday.” I think that’s all of this. You predicted all of this when we first met. I don’t know how many years ago this was—like 5 years ago. I remember you talking about the size of the models and how good they were going to get. It’s here.
It’s here. Again, the fact is, you can take one position and now look 360° around everything. On the other side, you have models like MiniMax H3 rendering faster than real time. You have Interdimensional Cable, which is now going to be in 3D, with DLSS from NVIDIA making everything high-resolution.
So again, the holiday experience and all the technology we need for it in 4K is here as of today, and this is one component of that.
And everyone is probably saying, “Why don’t I have it if it’s here today?” It’s only because of the global compute shortage and the 5× price increase in RAM. All the compute in the world is getting sucked up. If it weren’t for that, you’d actually have it deployed in your home this week.
It’s coming. I think, again, it might be premium, but you’ll pay for premium experiences. What’s Fei trying to do at World Labs? She’s trying to understand physics. She’s trying to create models that understand the world. This has such camera-controlled physics understanding that, as they scale it, it’s going to do even more stuff.
You’ve seen this from Runway ML, which is about to release its new one. Black Forest Labs is about to release a new one. There’s a whole series of world models that are approximating reality more and more and more for true digital twins. Incredible.
Alex, your take?
Yeah, it’s probably worth elaborating on what the core idea with Atlas appears to be. As far as I can tell from the documentation, the core idea is to take a diffusion transformer, which is what all of the state-of-the-art, at least American, video generative models use. It’s a hybrid of a diffusion model and a transformer.
The idea is to add one new modality. In addition to training it off of text, images, and video, you also train it off of 3D or 4D Gaussian splats. For those not paying close attention to the Gaussian-splat world, which has been super exciting, a Gaussian splat is basically a transparent blob.
It’s a transparent blob—it’s an ellipsoid—and you can layer and stack lots of these 3D Gaussian splats on top of each other to create hyperrealistic-looking, traversable 3D scenes. As far as I can tell, what Fei-Fei and World Labs are doing with Atlas is, for the first time, at least at scale to my knowledge, treating 3D Gaussian splats as a first-class training modality alongside pixels from images and tokens from text.
You can ask questions about 3D splats, and you can do all of those elaborate camera motions because, if you just have an arrangement of 3D Gaussian splats, translating a camera around is a trivial operation. If this approach scales—whether it’s 3D Gaussian splats or 4D Gaussian splats with dynamics, which they also demoed—Gaussian splats, which right now are this independent line of effort within the future of gaming, end up becoming a critical new form of token for modeling the physical world.
Yeah.
It’s also general-purpose in the sense that, if you can make a world model out of Gaussian splats, you can also make a subatomic model, an astrophysics-scale model with relativistic speeds, or a model of inside-the-cell interactions just as easily, as soon as you have the data.
You can use this same exact process to create world models for all these domains where human intuition is just terrible. That’s going to be a massive breakthrough for the discovery of very small things, very big things, and very powerful things. New ways to compute using light—all of that is going to come out of this same exact process.
We’ve said before, this is how we’re going to train robots in the future. They’re not going to be trained in the real world. They’re going to be trained in these high-fidelity simulation worlds.
That’s the present. I would argue that so many different robotic, embodied VLA, or now world-model companies are just being trained from watching YouTube. If Fei-Fei Li and her company’s approach gains traction, maybe the right primitive is no longer patches of images, which is what many of the models right now are doing.
When you train a diffusion model, you’re typically taking each frame of the video and breaking it up, usually into something like 16×16-pixel patches, and then treating those as tokens. Maybe the right primitive ends up being Gaussian splats.
It's definitely not going to be 16 by 16. I mean, the fact that that worked at all, I think, shocked everybody. Let's just take a language transformer, take images, cut them up, pretend each chunk is a word, and blast it through to see if it works. It just worked incredibly well. Maybe superintelligence is a general-purpose technology that relies on compressing information. Yeah, maybe.
Where does this go for the average consumer, guys? I mean, obviously, this is a world of extraordinary video games. Are people going to be just living their lives in these virtual worlds? Is this going to become how we consume ourselves, how we consume entertainment in the future?
Yeah, but also how we design our next day. What do you want to do tomorrow? I don't know—let's walk through what it would be like to do this or to do that, to play tennis or to go to the Caribbean. You can just experience it in advance and use that as your planning tool.
So much of our lives are random wandering and not really well planned out. I think human happiness is going to go through the roof once you're interacting with your AI. Right now, the AI will guide you through a day plan, but it's in text. It's going to be so much better when it guides you through a day plan visually and you're just stepping into it.
I'm beyond excited about all this. We have 3 vacation locations. Let's go and explore all 3 as a family, watch it, and then see which one we want to go to.
No, no. I'm excited if you go to a place—like if you go to a resort and you've been there before, you have so much more fun because you know where to go. You're not wandering or losing all of your time finding things. Now you can actually pre-experience things and know exactly what you'll enjoy, where they are, and how to get to them.
You're not waiting in lines. You're not signing up for things and then realizing it wasn't what you wanted. It's going to be so good. And ask your AI, based on what you know about me and my preferences, show me what I'm going to go do.
Emad, big announcement for you today.
Yeah. Over the last few years, we've been working on The Last Economy—what does economics look like?—and then released The Commonwealth, looking at personhood, law, and political economy. A lot of people ask, how do we share in the gains of artificial intelligence and make sure they're distributed? So we went back to the drawing board and thought about what type of institution and future we want to see. We came up with this idea of the champion.
We think that AI should be like a utility and should be owned by the people. You need the children to own it. You need the locals to own it. For every single jurisdiction, what we see is that the cost of intelligence will drop to zero and the value will go to the last mile. Salim would kind of integrate AI into the enterprise. The humanoids—who owns those humanoids? Because those will drive the economy.
Elon at the G20 just now said that the average humanoid will have 5 times the output of a person, and there will be 1 billion of them. That will be the economy. So we're like, let's set these up and borrow from the example of TSMC.
How TSMC was set up was that they started at a valuation of 10 Taiwanese dollars, and the locals put in 75% of the money and Philips put in 25%. The CEO did not have any shares. The team did not have any shares. They only got shares from the profits. He had to buy his shares, now worth $10 billion.
Is that for real?
It's for real. It listed at 6 billion Taiwanese dollars, which was the cash on the balance sheet. So we were like, let's do that for the intelligence company of California, the intelligence company of the UK—$1 pre-money. All the locals can invest at that valuation, from institutions to high net worths to retail. $75 million per state.
You can get the MIT endowment, Dave, to invest and give them back all the compute that they have—the equivalent dollar amount that they invest in compute. There are all sorts of interesting things you can do at $1.
Then you can bring in the internationals and the strategics at market rate, which will be 10 times that because you've got everyone on board, and you give 10% of the equity in perpetuity to every child under 20. So every year, you issue half a percent of the equity to every kid.
It trains up an FDE workforce to transform every institution. It owns the robots and deploys them, which is 80% of the value of robotics downstream, just like the auto manufacturers. And it gives an agent to every citizen, as well as AI for the government—the Sage project that Peter and I and others have been working on—AI for the judicial system, AI for education, and healthcare.
That becomes really super interesting because then it becomes a play on the indexed GDP of the state, owned by the people of the state, with the smartest people in the state involved. Again, the trick here is $1 pre-money. Get everyone in. You want it to be a success? It's up to you. And so that's this new institution.
How many champions are there? How fine do you slice it up? I heard you say champion for California and the UK—cities and countries, or states?
In the United States, we're doing 1 per state because a lot of data has to stay within state boundaries. Otherwise, it's pretty much 1 per country. So, again, it acts like British Gas here, for example. It acts like the telco and more, and each one of them covers a certain number of citizens because, again, it gives equity to every child born.
The locals can invest, etc.
So, Emad, right now, this is just an idea. If folks want to learn more about the idea, it's not an investment offering, and this is not investment advice. But to learn more about the idea, where do they go?
They go to ii.inc. As you said, it's just an idea at the moment, and we'd love people's input, but ultimately it'll be about the citizens. So, let's see if we can build this structure.
Your first phase, I've had some chats with you about it, is super exciting: to get all these local champions lined up, right, and then, little by little, cascade to the next level.
Yeah. It's all about whether a state wants this. Then it's all about the people of that state. It's not like a foreign company coming in. This is a model just like you have your UBI, just like you've got your shares in Frontier Labs, et cetera.
And so, we'll see how it goes. This is my proposal for trying to distribute it to everyone. Awesome. I've got to ask you a question. Your bullets say, “The cost of intelligence is going to zero.” Got it. “The economy will change forever.” Got it. “The value of human cognition will go negative.”
Yeah.
You're saying it's going to be like a cost on society to be thinking and consuming food while you're—
Yeah, the stupidest person on the team economically, right? You're competing—
Your ideas are—
No matter how good your ideas are, they add negative value. It's like adding a human driver to an autonomous highway. Of course, this is the topic of the book that I had last year. But that's why you need to have a share in the means of production, which will be the robots and the forward-deployed engineers and people like that, right? You need to make sure that's equitized from day 1. Otherwise, if you have a thought and you don't tell anybody, it's just zero. Then, at least, it's not negative. Most people's thoughts aren't economically valuable thoughts; it's economically valuable labor. I'm not the smartest person on my agent team anymore, man. I don't know about you, but I'm getting eclipsed very quickly.
Yeah. No, it is very humbling for humanity, actually.
It will be. But just like we've seen before, when Stockfish started beating everybody in chess, people still play chess, and people are still going to have ideas. People are still going to value individual human ideas. It's like, “Did you come up with that, or was that your latest model?”
All right. I'm going to move us to our next story. Tesla held its Cybercab launch event in Austin, Texas. Cybercabs everywhere—a river of golden EVs flooding the streets. Unless you've been hiding under a rock, you know that Cybercabs are Elon's electric, autonomous 2-seaters with a, quote, “three-comma” scissor door. No steering wheel, no pedals, designed to run entirely on Tesla's self-driving software. Did you guys get the three-comma comment? Anybody?
No.
It's from Silicon Valley, right? The billionaire.
Okay.
Remember, he showed him a Maserati and said, “Ah, that's not a three-comma car.”
It doesn't have the doors that go like this.
Right.
I remember that.
A great episode, actually.
This is a vehicle designed from the ground up for full autonomy. Let's take a look at this video. I love this video showing the flood of Cybercabs on the streets in Austin. Check this out.
Cybercabs everywhere. A golden river. It's crazy. You know, I didn't really internalize the lack of rearview mirrors. We obviously saw the lack of a steering wheel.
Yeah, but actually, the number of components they've taken out of the car is driving down the cost. There's no driver, and there's a huge amount of cost in the car that's gone. So there's no way anyone's going to match the price point of this thing. I mean, that's the point compared to Zoox and Waymo or anybody else. Elon wants to sell these at $30,000 each so you can buy them, and I think a great entrepreneurial journey is to buy 10 of them and put them on the streets in your local town and have them earn revenue for you. One early rider in Austin is reporting that Cybercabs are running about 50% cheaper than Uber for comparable trips, with rides being described as smoother and cleaner than Uber, with automatic syncing of audio and seat settings in your passenger profile. It's pretty extraordinary. Remember, Nevada gave Tesla permission for 5,000 of these vehicles on the road in Las Vegas in the next 12 months. It's just going to crowd out the competition. Salim, I'm curious what you think about this. Is this sort of like the ChatGPT moment for Tesla?
I think it could be. This is classic ExO, right? Autonomy, you decentralize, you have interfaces, you're leveraging assets. If Elon can get people to buy the cabs and create millions of micro-franchises out of them, the loyalty that will come from that is going to be unbelievable. Of course, they're all going to have self-driving capability, so the whole vertically integrated stack brings itself to bear. I think the most exciting part is the fact that the cost of transportation drops by another order of magnitude, from a couple of dollars a mile today down to 20 cents a mile. That is really interesting, and I need this to happen because I've made that comment that Milan will never go to university, or he'll never need a driver's license. I'm going to battle it out with him that he will not get one. So this needs to happen. He's got 2 years to get this done.
And you would know better than anyone, but when you get a hotel room in New York and you look down, it's like 80% yellow cabs down there. If you look at the traffic in New York, a huge fraction of it comes from people blocking an intersection. They get the ticket, but they're still stuck in it, and the whole thing grinds to a halt.
Yes. I suspect there will be enormous chunks of entire cities that say, “You know what? No more human drivers.” It's so much more efficient. Not only is it much cheaper, but it's much more efficient to get around. Also, a vast majority of the safety issue in cities is pedestrians getting hit by cars on sidewalks, and this is going to basically eliminate that risk. I suspect we're right on the tipping point where it's like, “Nope, no human drivers in this entire inner-city area. End of story.”
And we'd better get artificial organs quickly, because the drop in organ donors is going to vaporize.
Well, all these things tend to happen at the same time anyway. I think for many people, these robotaxis will be the gateway drug to generally autonomous robotic systems on the streets. I think we'll look back, with the benefit of hindsight, and say, “Well, of course, it's natural that a municipality gets lots of robotaxis on its streets before it gets humanoid robots on the sidewalks doing economically valuable activities.”
My favorite little anecdote, analysis point, or data point to illustrate this is that every medium-sized town in the country has a transit system where the buses run empty 95% of the time, and they're totally full at rush hour. They run completely empty the rest of the time, and the whole thing is a massive loss-leading exercise. It costs a bomb.
Already, a few years ago, small towns in the U.S. were saying, “Get rid of the transit system. We'll just pay for everybody to take Uber,” because that makes it hyper-efficient and much lower-cost overall. But this takes it to another level. It takes it to 11, to go down the full analogy that we're going down.
And this now allows mobility to every single person—every blind person, every disabled person, every student, every drunk person. All sorts of capabilities become available and affordable in a way that's very powerful, incredibly exciting.
Yeah, the poorest people in the world are being chauffeured around by AIs, for sure. Two interesting points here. The first is that it's a two-seater, right? And, of course, the average load for an Uber is 1.2 people, right? How many times do we take an Uber all by ourselves? So, if you need 6 people in a car, just take 3 of the Cybercabs.
Yeah. Not just that, but think about a normal cab: it has 4 airbags. You get airbags for every seat. Elon, in his infinite brilliance, is like, “We now only need 2 airbags.” We cut the cost of airbags alone in half with this design. In the rare instance where you need 3 or 4 people, get 2 of them. End of story.
Well, they follow each other right behind each other. It's perfectly great.
I suspect there's more of a backstory there, though, with the two-seater. Remember, Tesla was originally teasing that it was going to launch a Model 2, which was going to be their highly coveted $25,000 car. That never happened. My best read of the situation is that what was originally planned as the Tesla Model 2 became this robotaxi, because at some point, once the value of the car—or the sales value, I should say—crosses below some threshold, it makes more sense to monetize it via autonomous ridesharing or robotaxi services than it does to actually sell the car.
This will turn transportation into an API.
Yep. Interestingly enough, we're starting to see conversations where the other companies that do have steering wheels, pedals, and LiDARs are saying, “I'm not sure that the Cybercab is safe enough. It only has cameras. It doesn't have the other modalities.” So, expect to see a battle in the courts about what city allows this technology in.
Yeah, Boston, just to put an exclamation point on that.
Mayor Wu, come on the podcast and we'll have a conversation about it.
Oh, yeah. Do that. That'd be awesome.
I'm going to continue this conversation. Let's shift it now to the global stage. The competition is going three-sided. The Financial Times reported that Uber is partnering with traditional taxi fleets to complement and compete against Waymo's expanding robotaxi service. So, a strategic alliance between ride-hailing and traditional taxis to counter autonomous vehicles that don't need humans. Think about what that means. Uber, the company that disrupted taxis, is now partnering with taxis to fight the companies that are disrupting them.
Meanwhile, The Verge reported that Uber and U.K.-based robotaxi company Wayve officially launched services in London. Emad, have you seen Wayve yet?
Yeah, I've seen it. It's just starting to roll out, so I'm looking forward to getting my first ride on that.
Yeah. CNBC reported that Waymo and Zoox announced simultaneous expansion into new cities, along with Tesla in Nevada, California, and Texas. My prediction here is that we're going to see at least 5 different autonomous electric robotaxi companies fighting it out in major cities inside the next year, driving the cost as low as possible.
The cost of personalized transport is basically dropping to the cost of electricity. It's demonetized mobility. It's the entire abundance thesis. The cost is dropping to the cost of charging a battery. The ultimate winner is the public, unless you live in Boston.
Well, no. Cambridge might have a chance. Actually, that'll put a lot of pressure on Boston. That'd be awesome if Cambridge—
—beat Boston by trying to take a taxi and it stops at the Harvard Bridge.
That's right. The bridges are no-fly, no-drive zones for autonomy, apparently, in the near future.
Actually, what's funny is you can get around Harvard, but you wouldn't be able to go to Harvard Business School.
That's really cool.
Yeah. What's going to happen, though, when you get your Tesla Optimus robot and it has the drive program, so it can drive for you in whatever car you have without a retrofit?
Ah, there you go.
It'll get banned in Boston, too. I'm sure in Boston.
I want to hit the economics here. The new versions of the Waymo vehicles are pricing out at over $100,000. The Cybercab is coming in at or below $30,000. What's most interesting in my mind is that the winner here is going to be whoever can mass-manufacture these the fastest. Elon's going to win that game. I mean, he's the guy who builds the machines that build the machines.
In the U.S., maybe, but China is dumping some, some might say, into Europe. So, it's not a U.S.-only game. That's the issue.
You're going to end up with all of them for the foreseeable future. I think, Peter, the point you made is the most important one: the end user wins.
Yeah. Dave, do you remember when we were at the Gigafactory, being toured around, and we saw outside the giant mounds of aluminum scrap metal?
Yeah.
And then the smelter.
And then the Model Y press. It was just this beautiful orchestration. It's incredible, end to end.
And you can actually walk with the car from the minute it's born. It's a long walk, almost a mile, but a car comes out the other end. You're literally watching every part get put on as you walk with it. It's absolutely wild.
Yeah, it's beautiful.
But it's amazing to me how few parts there are in the Cybercab. When you open the hood of your gas-guzzler car and look at all the stuff that's in there, and then you look inside the equivalent Cybercab, it really feels like there's maybe one-tenth as many things.
A typical car has 2,000 moving parts in the drivetrain, and a Tesla has 17.
An ICE—an internal-combustion-engine car.
2,000 moving parts versus 17. Unbelievable. So, in terms of maintenance, design, and all of that stuff, this is why the car dealers are all freaked out.
The last time I was in New York, I was in a yellow cab. I'm like, “Why is this thing so disgusting? It smells like a public urinal. Why are they all like this?” But then you think about what happens to that cab. The medallion is incredibly valuable. It gets handed from one driver to the next. It never stops moving.
At night, in the middle of the night, somebody drives it home. Then they get up first thing in the morning and start driving it again. But it has to take somebody home. The Cybercab goes to get cleaned in the middle of the night, when there's nothing to do. It goes to a place—it could be anywhere. It could be in Queens, where it gets itself cleaned—and comes back pristine.
It's just night and day when you get into one of these versus a cab. Now, they're brand new, so maybe that's part of it, too, but the experience is night and day better.
And they look beautiful.
So, what happens when they break down? I guess a Cybertruck comes and tows it away.
Yeah.
DoorDashers come over to help, is the recent story, I guess.
The door is wedged open and can't be closed.
Maybe they just squash it into a metal cube right there. Crush it—
—and recycle it.
Take it back to the smelter.
That's mean.
We'll pull out the GPUs first.
Yeah.
All right, I'm going to move us on to the space arena. Two fun stories on the space frontier today.
In our first story, NASA chose Blue Origin to build the telecommunications relay network on Mars—the communications infrastructure that will connect future Mars missions back to Earth. This is the infrastructure layer for the Mars economy. Whoever owns the telecom network controls the bandwidth on Mars. There could be a company like AT&T on Mars. It could be Blue Origin.
You might wonder, why did NASA select Blue Origin, not SpaceX, since SpaceX owns and operates the world's—or space's—largest space-based, laser-linked communications network? Ultimately, this is the government keeping 2 competitors, 2 suppliers, in business, giving them each a slice of the pie.
Regardless of Blue Origin having won the NASA contract, I guarantee you Elon will still build Starlink around Mars. Ultimately, what we're seeing here is the birth of the interplanetary internet. Alex, are you excited about this one?
Look, Mars had it coming. The moon had it coming. We're going to build the Dyson swarm. Someone was going to get awarded the Starlink for Mars. It's interesting that it went to Jeff Bezos's company and not Elon's, but I'm fully expecting that we're going to have a very competitive interplanetary internet. Frankly, I'm glad that there are vendors competing for Mars communications other than SpaceX.
Yeah, but it's going to be a packet-switched network throughout the inner solar system: Mars, asteroids—
High latency, but that's all right.
Yeah. Until we get faster-than-light communications, who knows? Physics.
Time will tell.
Time will tell. Are you working on that, Alex?
Can't say.
Well, the physics that we have right now suggests that faster-than-light travel is not possible.
Don't bum me out here.
Okay. Sorry to break the bad news. The textbook physics right now says that superluminal travel—
Then we have the wrong physics.
We'll find out. Just hope.
Yeah.
Let's move us to another fun story in space. Our second story—and it's a big one. NASA's Nancy Grace Roman Space Telescope has launched on a Falcon Heavy, carrying a field of view more than 100 times greater than Hubble and the ability to scan the sky more than 1,000 times faster. The Roman telescope is designed to discover tens of thousands of new worlds and map the distribution of dark matter across the universe.
Before we discuss it, let's play a video by our amazing NASA administrator, Jared Isaacman, a friend of the pod. I love this guy. He's such a good communicator. Let's check it out.
The telescope is very healthy right now. It's making its 1-million-mile journey to Lagrange Point 2. It's going to look for what we think will be up to 100,000 additional exoplanets in other star systems. It's going to help us understand dark energy and dark matter. As you saw during the press conference, President Trump called in. This is a really exciting time in America's space program right now.
What's the difference between this telescope and the Hubble?
It's hundreds of times more powerful. The field of view is over 100 times greater than Hubble. Its scan rate is over 1,000 times greater. This is going to be a household name like Hubble and James Webb. This is America's next great exploration asset.
And you say it's going to help you find planets hiding behind planets—
Planets hiding behind other stars—distant stars. The light can blind it out. This has a special JPL coronagraph that helps us find these hidden worlds. We're going to find tens of thousands of additional worlds.
Tens of thousands of worlds. Amazing.
Yeah. Jared is incredible, isn't he? Compare that to that interview of Sam saying, “Why is this new model different? Why is Astra different for people?” Compare his answer to what Jared just did to answer this question. It's like night and day. He is a great communicator.
Interestingly, this is going to finally start to give us statistics over the number of habitable—at least Earth-recognizable, habitable—worlds. It's being pointed toward the center of our galaxy, and we'll be able to do large sweeps of the sky, looking in part—it has other missions as well—for microlensing events, for planets crossing in front of their respective stars and causing, via their gravity, light from those stars to be very weakly increased briefly due to these microlensing events.
The downside—
You know what's so cool about that is, on the last podcast, Peter asked us—you missed it, Salim, and Emad, you weren't here—what's your answer to the Fermi paradox? Why are we not seeing other civilizations? One of the theories is that there are many, many civilizations talking to each other near the center of the galaxy, where you can get from star to star in a year or two, as opposed to where we are. We're way out in the wings, where it's so far away.
We're in the unfashionable outer suburbs of the galaxy. That's right.
So who knows? This could—
I'm actually happy we're here. The galactic center is a really dangerous place to be. You don't like supernovae, Peter?
I don't like the radiation flux they deliver. No.
I'm reminded of the opening scene of The Hitchhiker's Guide to the Galaxy, where they're bulldozing Earth to make a hyperspace highway—
Bypass. Yeah, yeah.
I love this. I'm curious about our viewers, and if you can, let us know in the notes: would you like us to have a conversation on the current UAP disclosure discussions out there? If we can bring in some of the leaders in UAPs and the whole disclosure scenario—the White House just released its disclosure plan—I don't know if I'm going to mention that, Alex, but I'm curious if folks want us to have a conversation on that topic on the pod. Alex, would you mention the recent White House announcement?
There was some reporting out there from Avi's UAP Science Advisory Council. One of the members of the council mentioned in a recent briefing that this council, which, as I understand it, was stood up by the White House, was informed that the White House had prepared a disclosure plan for informing the general public of the existence of non-human intelligence. If that reporting is accurate, that's pretty interesting.
Yeah. Emad, where do you come out on the whole UAP side of the equation? I'm curious.
I actually have no position on it. I've never thought about it properly.
Okay.
I'd like to see strong claims require strong evidence.
I stand with that.
A lot of people would say there is strong evidence. It's just hidden. But we shall see.
No, I'm with that. I would say evidence—it would be highly desirable for there to be a preponderance of evidence that everyone can go and see, touch, and experiment on. I think that's probably the gold standard in an ideal outcome, if there were a White House disclosure event. If the president goes into what's left of the Rose Garden and gives a speech and says, “We're not alone,” then ideally part 2, paragraph 2 of the speech would be to hold up or otherwise present some artifacts that would be subject to extreme scientific scrutiny to support the claim.
That would be a fun conversation to have. All right, our third—our final—subject for today. Three major health and longevity stories came out this week.
The first is that OpenAI expanded ChatGPT Health features to connect directly to patient records and health-care databases. The integration brings Epic electronic health-record data into ChatGPT for health care, and Epic, as you guys probably know, is the largest collection of health records. 325 million patients are inside Epic—roughly the entire U.S. population. Clinicians can now pull appointment notes, lab results, and medications, and ask questions across a patient's entire record. Consumers can now connect their Apple Health, One Medical, and Function Health data, so ChatGPT can help you understand your test results, prepare for your doctor appointments, and get personalized diet and workout advice.
That's the first story. The second story, interestingly enough—we talked about it last week. We mentioned that the FDA had approved a drug called Duraxinarissib. It rolls off the tongue and onto the floor. It's the first targeted RAS inhibitor for metastatic pancreatic adenocarcinoma, attacking the RAS family of proteins that drive tumor growth in most patients with the disease.
This week, NBC News reported that the same drug is showing promise to treat lung cancer. As mentioned last week, the RAS mutation family drives roughly 30% of all human cancers and was long considered undruggable, the subject of decades of failed attempts. But now, during the singularity, the end of cancer is within reach. Alex?
Yeah. I think we're starting to see—and, interestingly, I'm not sure that AI was actually essential for this particular drug—but I think there's so much progress being made in cancer therapies now, including on the immunotherapy side, that we're starting to see the emergence of, again, caveat, caveat, caveat, universal cancer treatments and universal cancer vaccines in some cases.
After years and years of treating cancer as thousands of different diseases, we're finally starting to get to the point where we can move upstream and start to treat, if not root causes, at least identify treatments that—whether it's proteomic pathways on the one hand or immunologic pathways on the other—start to have treatments and vaccines that can treat multiple classes of cancer. And I think that's—
Cancer's cooked. That cancer should have been cooked long ago. Was it Nixon who announced the war on cancer? It took forever to get to this point—more than half a century. It's an interesting counterfactual thought experiment: is there anything, knowing what we know now, that we could have done 50 or 100 years ago to radically accelerate the onset of broad-spectrum cancer treatments? What do you think, Peter?
I think the data—I mean, we've known, for example, about the RAS mutation causing unconstrained growth for some time. It's just getting the molecules and getting the drugs that can attack it properly. I think the tools we have right now are finally giving us that reach. Then, being able to understand fundamentally what happens and how to block it—we talked about cell simulators—is what's coming.
Emad, this is an area of personal passion for you as well. You've been deep into medical AI.
Yeah. I said this wasn't AI, on the thing, but the range of treatments now coming out on the cancer side gives a lot of hope for what's going to come. I think the mRNA one may be more general that we saw recently.
I think the first part of integrating into the Epic health records and actually applying AI across the board—we should have a sprint so that within a year or 2, maximum, every single health decision is double-checked by an AI. We should have—
More than that. I think it's going to become malpractice—
To diagnose a patient without AI in the loop.
Right. We already know AI is a far better physician—a diagnostician—than a human is.
So I think you should have a series of approved edge and cloud models, and every time you make a diagnosis, the AI has to have had one check. That will save so many lives. It will detect so many cancers. Then it's about how we increase the level and volume of information, because even now, the type of data we have around cancer and other conditions that we absorb is tiny compared to the amount that we could have with AI transforming it.
So I think, yeah, let's cure disease and get rid of it. No one should have to die of cancer, and it's something we should be really directed at.
On the first story—
Go on, sorry, please. It's just very strange that, as we have the Genesis programs and others, there isn't a straightforward “We now have the capability to potentially cure this stuff. Let's direct $10 billion toward it.” It's tractable.
I think the story on Epic is interesting. Having been involved in that business through Fountain Life, Epic is the majority electronic health record in the US, and it's been a bear to navigate for physicians. Patients have never had access, so putting an AI layer on top of that is awesome. Yeah. Let me move it to our last story here in the area of health. It's related to what may be called the first longevity class of drugs: the GLP-1s.
A new paper published just 2 days ago, on September 2, in Nature, shows that semaglutide—the active ingredient in Ozempic and Wegovy—recapitulates many of the benefits of caloric restriction. Get this: it's in female mice, just to be clear, and extends the lifespan of mice by almost 100 days. In humans, that's the equivalent of 8 to 10 years.
The study found that GLP-1R activation initiated late in life in these mice accentuates age-associated decline and modulates conserved genetic regions for aging, functioning as a caloric-restriction mimetic. Also this week, it was reported that GLP-1 drugs are being linked to fewer serious infections, including tuberculosis. You guys already know that tuberculosis is the deadliest killer on Earth, killing over 1.25 million people per year.
So the same drug that's treating diabetes, obesity, liver disease, kidney disease, cardiovascular disease, and addiction may also be extending lifespan and reducing infectious disease. One molecule, 6 diseases, and still counting.
Again, Alex, you made the point that this might be the beginning of longevity escape velocity. To the extent that, with the benefit of hindsight, we look back in a few years and say, “You idiots. Of course you were on the verge. You were seeing the sparks of longevity escape velocity. You had the GLP-1s,” I don't think it should be that surprising.
What's more mystifying to me is, from an evolutionary perspective, if the GLP-1 receptor agonist class of molecules is capable of doing everything from treating infections to extending life expectancy, modulating diabetes, reducing addiction, and reducing compulsive behaviors, why on Earth did we not evolve with either this ability to modulate our own semaglutide-class molecules in our system, or maybe a slightly more cynical angle?
If it turns out that the reason GLP-1s are so effective at so many diseases is that these diseases somehow are diseases of modern lifestyles—that it's treating all of the diseases of modernity, and that's why we never evolved the solution—that would be pretty ironic. Addiction to compulsive gambling or alcohol, or overeating sugar: these are all relatively modern diseases. That's why this is basically a treatment for modernity.
So I just want to make a point to the listening audience. First of all, talk to your physician about this. This is not medical advice. Yes.
But I use a GLP-1 drug. I don't know if any of you do right now. I use it not for weight loss. I've been at my fighting weight for a while. My body-fat percentage is substantially lower because I work out and I'm very careful about what I eat. I use it as a longevity drug because of all the benefits we just heard about.
I've got 2.
I know you're using one.
Yeah, I'm using one, and I just got my blood test done. My liver enzymes are 50% better, which is pretty amazing. Which means I can drink more. No, just kidding. That's not the point.
I know that's not the point.
You're supposed to not want to drink it.
I know. So there's that, but I think, Alex, to your point earlier, remember that evolution has birthed us for death. We've had short life cycles so that the cycle time of evolution can work more quickly. We're breaking through that now and living through that, so I think it may have been engineered, or evolutionarily engineered, that we died at different levels.
This is a long history. We used to die of heart disease or bacterial illnesses, and then we figured that out. Then we died of heart disease. We got some sense of that. Now we're dying of cancer.
Then explain tuberculosis. If you're in a tuberculosis-rich environment, surely evolution would favor any molecule that could be amplified, which would help young and reproductive entities survive tuberculosis infection.
Well, I go back to the fact that there are lots of pathogens going after us all the time, right? Billions of them, because they're all trying to survive in their own way, including cancer. The capability of the human body to navigate this has never been needed until we started pushing the boundaries to this level. Now we do.
Salim, my thesis—and I've spoken about this pretty widely—is that the reason for the life cycle is not more rapid evolution. For most of human existence, Homo sapiens came on the scene roughly 200,000 years ago. Food was very scarce, and if our primary mission is to perpetuate our species, the last thing you want to do is steal food from your grandchildren's mouths.
So the best thing you could do is die: reproduce and die, basically. We see the human body—and people should know this—you're in prime condition until your late 20s. Back 200,000 years ago, you'd go into puberty at age 12 or 13. You'd be pregnant immediately, with no birth control. By the time you were 26, 27, or 28, you were a grandparent, and then you would die so you didn't steal food from your grandchildren's mouths.
That's my joke about marriage: we invented marriage to keep the parents together until the kids were self-sufficient. The average lifespan was 25 years for most of human history. We invented marriage about 6,000 years ago, and it's definitely true then: marriage is not designed for 50- or 60-year lifespans.
No, no. Watching this, one of my relatives calls it state-sanctioned torture—marriage—because we have the job of now evolving the institution to deal with the conditions today compared to when we first invented it.
I use that example because it applies to all our institutions: democracies, educational systems, legal systems, and healthcare systems. This, I think, is the biggest work we have to do. As we blow past all these limitations, we have to reinvent all of the institutions that are the scaffolding that keep humanity safe and civilized.
Hopefully, we have a benevolent AI to help us do all that.
We'll need it. And if not, benevolent AI. At least we'll have GLP-1s.
All right, time for some conversations with the mates. We've got some questions here. Emad, I'm going to give you first crack.
All right. If money becomes obsolete, how will desirable land be allocated? Is everyone just locked into their beach house forever? This is from Nick 52547.
This is an interesting one. Elon and others have said that money might become obsolete. I've said that you can't compete with robots. I think this is a question of land rights and more, because typically where you see reallocation is in upheavals.
The way we're going, without the right structures, you'll probably have a debt jubilee. You'll probably have chaos and land redistribution. But if we can actually navigate through it, then land rights and property rights should be enforced, and so you can keep your beach house.
The number of desirable locations will go up dramatically because you will have air taxis, self-driving cars, solar panels, and self-driving construction workers.
Yeah, exactly. Self-driving construction workers.
Yes.
Nice.
Do you think AI will evolve past this risk-averse bottleneck that we're currently in? And that's from AdsSusie1073.
Yes, but hopefully we move toward calibrated risk rather than recklessness. Right now, we treat uncertainty as a reason to refuse. That's a technical problem, but there are also all the legal and policy issues around this.
When you're raising kids, one of the things they teach you is: Are you taking a responsible risk? We need to apply that same kind of paradigm to these models. Could you create something like a risk budget? What can an AI decide? How much can it spend on that? What systems can it access? What triggers escalation?
We need to get very sophisticated around this. Models need to become better at distinguishing dangerous intent from real expert usage, because you can't just say, “Remove the guardrails.” It's got to be dynamic, where you deal with proportionality based on user identity, the context, and the reversibility of that.
We'll get AI maturity when we can say, “Here's the risk, here's the confidence I have in this, and here's the reversible next step.” Rather than just saying no, I think we need to get a lot more nuance around this. Unfortunately, in today's world, nuance has no part to play in the sound-bite politics that we have out there.
All right, Dave, over to you.
I’ll take number 1. Can you see a way to turn libraries into centers of AI development and physical AI training for everyone, from Train with John Koalo?
Yes, but there’s a bigger issue, which is that there’s a huge amount of white-collar office space, and white-collar work is turning to AI. At the same time, we have a declining population, especially a declining working-age population. So there’s a bigger issue: What about all this other space? What are we going to do with it all?
It’s very similar to what happened with the shopping malls after online shopping became huge and then COVID hit. A lot of people were thinking, “There’s got to be something really great we can use all this shopping mall space for.” There were some ideas, but for the most part, they didn’t work out. It all got replaced.
So I think there are lots of ideas for what you can do with library space. You don’t really need the library books anymore, obviously. This is another reason why you should want data centers in your community: It’s one of the few things that will reliably grow tax revenue and create job opportunities in a community. So there are ideas, but I’m not super optimistic that all the space will be well utilized going forward.
There are programs in inner cities right now that are running AI tutoring in libraries. That does exist.
I get the fun one. Question number 2: Whatever happened to synthetic diamond chips for computers? This is from Asterene.
Here’s the problem with diamond: Pure diamond is an insulator. It has a band gap of approximately 5.5 electron volts, so it really doesn’t want to be a good computer. To make it into a good computer, at least a good computer of a recognizable CMOS type, you have to dope it. That’s hard.
If I had to make the case for diamond-based CMOS computers, it would be for ultra-high-temperature environments, where such a wide band gap could be advantageous, or maybe very high-voltage environments, where, again, a large band gap could be advantageous. It’s not an enormous market. I could be wrong, but it doesn’t seem like an enormous market.
Where I’m much more bullish on diamonds is sensing. You can put what are called nitrogen-vacancy centers, or NV centers, into diamonds. Basically, you put an extra, unwanted nitrogen atom into a diamond lattice, and suddenly you get an exquisitely sensitive magnetic-field—or just field-in-general—sensor because of an extra vacancy that is introduced as a result of the diamond.
Diamonds for quantum sensing: super interesting. Potentially, it’s not investment advice, but technologically, it’s very attractive. One could imagine, at some point in the future, diamond chips for sensing even getting us to sci-fi technology, like wearable MRI sensors. Diamond chips for computing? Eh, probably not.
Why did I know you’d have an answer for that one?
It’s interesting.
All right, Salim, first choice is yours, pal.
I will go with number 7. When do we get a forecast for when data-center energy sources transition from gas-turbine farms to nuclear and SMRs?
Short, fairly easy answer here: Gas dominates the current buildout. SMRs are going to be in the next 3 to 4 years. The reason people are getting so excited is that nuclear is an engineering problem, not an invention problem, right? Fusion is still in the invention-problem category.
You’ve got this impedance mismatch of data centers taking 2 or 3 years to build, while nuclear is taking a bit longer. We will get there, I think, faster than people think with SMRs, but I think it would still be a while. Take 2 to 3 years for the initial wave of SMRs and then 5 to 7 years for the big buildout.
I think what’s going to end up happening is that AI is going to do for nuclear what smartphones did for batteries. It just created such huge demand that the massive innovation there caused a huge acceleration in the innovation curve.
Nice. We had a fantastic pod with Ramez Naam on energy. If you guys haven’t seen it, please check it out. He talks about pretty much that time frame.
That’s where I get all my good information about energy.
Yeah. Alex, let’s go to you.
I think I have to answer question number 8, which seems to be directed toward me. It asks, “How far away from Alpha Centauri do you actually have to aim to reach it when it arrives?” This is from Typical Dad Pi.
To 3 significant figures, to the extent I understand this question, let me give you the semi-official Fermi Explorer line. On the last pod, we had Matt Pines and Philip Johnston announcing, for those who didn’t watch, humanity’s first mission to Alpha Centauri, and I’m involved. I’ve been involved. What can I say? I’m in a lot of rooms.
Under the official mission parameters for the Fermi Explorer mission, the goal is to get at least 99% of the way to Alpha Centauri. Alpha Centauri is approximately 4 light-years away, so you could do the math, and that turns out to be 4/100ths of a light-year away. Call it precision getting there. That’s from Earth.
If I put my sci-fi hat on and extrapolate a bit, I suspect that when the mission comes to full fruition and is launched, it will have active guidance on board. Maybe I should add parenthetically that the Fermi Explorer mission is intended to launch by 2029. I suspect this is not an official position for the Fermi Explorer mission, but if it has active guidance on board, this 99% of the way to Alpha Centauri can turn into 100%, in the sense that it actually hits the Alpha Centauri system, not just ends up 4/100ths of a light-year away.
That said, Fermi Explorer is intended to reach Alpha Centauri 80,000 years from now, which is perhaps inconvenient from the perspective of mission verification if you actually want to get it.
GLP-1 drugs.
We’re going to need a lot of GLP-1s for this.
I think the whole point of the mission is that we’re supposed to beat the Fermi Explorer and get there sooner. This is just the first one to launch.
Amazing. Dave, over to you, pal.
I’ll take number 6. Are data centers using closed-loop geothermal cooling also noisy, or are they quieter? This is from Asterene.
They should be dead quiet, just like geothermal heating is dead quiet. Also, nuclear reactors that are near the ocean use geothermal cooling and ocean cooling, and they’re dead quiet. So it should be very, very quiet.
Regular data centers that are using liquid cooling are only noisy because the water gets cooled outside with these really poorly designed fans. There’s no reason for that. If your regulatory body says, “Yeah, you can build a data center here, but it has to be quiet,” guaranteed, they will build the data center quiet. There’s no reason those external fans need to make any noise.
Instead of making them illegal, cities need to be saying, “These are our requirements: Drop our energy costs, make them quiet, invest in our infrastructure.” And they will. They will.
Emad, it looks like number 5 is for you.
That’s an interesting one. They talk about a sunshade there. What scale of satellites, in terms of—
We talked about—
That came up on the last pod. You missed it. Yeah.
Yeah. And what scale of satellite, in terms of square meters, is required to actually impact global temperatures?
I would say if you put something at the Sun–Earth Lagrange point, a couple of million kilometers out, you need about a couple of million square kilometers. So, about the size of India would knock a degree Celsius off. That’s bigger than anything we’ve ever managed, but it’s worth a try. Why not?
Maybe convert the Moon.
Mercury, please.
Mercury.
Or a series.
Yeah. And I still love the sunshade—putting a thermometer there to be able to titrate the solar flux on the planet.
I think it’s a trick question. For what it’s worth, I think with a sufficiently good AI planetary-scale model, we could make any sunshade or other satellite intervention de minimis in size. It’s just a matter of appropriately perturbing the Earth’s atmospheric system with enough AI.