[BidClub_]
Moonshots · · 128 分钟

紧急更新——AI 的 Sputnik 时刻:Kimi K3 发布,与 Emad Mostaque 对谈|第272集

Peter DiamandisSalim IsmailDave BlundinAlexander Wissner-GrossEmad Mostaque

YouTube
TL;DR
  • Kimi K3 让中国开放权重模型直接站上能力前沿,挑战了“只有美国闭源实验室及其资本基础才能保持能力领先”的前提。 Moonshot AI 这款拥有2.8万亿参数的多模态模型跃升17个排行榜名次,在前端编程及另外6个领域排名第一,并成为 Artificial Analysis 成本性能前沿上的第三个点,排在 Claude 5 和 GPT-5.6 Solve Max 之后。模型权重预计在7月27日左右发布,Alexander Wissner-Gross 称这一进展“对竞争大有裨益”。

  • 圆桌最关键的技术结论是,K3“没有魔法”:可识别的 Transformer 架构、更好的数据、持续不断的工程优化和高效执行就足够了。 Emad Mostaque 将这一过程比作中国电动车制造,并指出 Moonshot 仍在使用 H800,同时围绕 Huawei 和 Alibaba 芯片进行设计。Peter Diamandis 给出了更强、也更具推测性的解读:GPT-2 速跑项目实现99%的成本下降,意味着在数十亿美元西方数据中心中训练的模型,现在可能出现“1%成本版本”。

  • K3 正在威胁基础模型估值,但嘉宾对于实际有多少收入会被转移存在明显分歧。 Salim Ismail 估计,监管加上开放权重替代,可能抹去一家万亿美元实验室75%的价值,因为“前沿智能如今是一种完全易腐资产”,保质期以周计算。Blundin 和 Wissner-Gross 则反驳称,除非 K3 不是接近,而是达到2倍、3倍或10倍的优势,否则企业仍会为最好的模型、可靠性、支持、安全性和前沿性能支付高价。

  • 美国芯片管制可能加速了如今正在给美国实验室施压的效率创新,而中国正明确将开源视为地缘政治基础设施。 圆桌认为,受限算力迫使行业改进量化、数据混合、内核和硬件感知架构;随后,美国推理服务商可能在更新的 Nvidia 和 AMD 硬件上以低10倍的成本运行 K3。Mostaque 对中国立场的总结非常明确:“我们将全面支持开源,把它作为全人类的公共产品。”

  • 对企业而言,持久护城河正从模型本身上移至模型切换架构、专有数据、验证和工作流整合。 Ismail 认为,采购周期跟不上模型发布速度,因此价值将归于能够持续替换模型的接口。实践建议是,在内部评估 Kimi K3 和 Inkling,利用专有数据微调,并把生成代码当作人类代码处理:隔离运行、测试、扫描,同时保留责任归属,而不是追问哪个模型值得盲目信任。

  • 量化可能比模型基准测试所显示的速度更快地,让今天的前沿能力变得本地化、持久化且成本大幅下降。 节目引用的 Bonsai 27B 结果显示,一款手机级模型可压缩至6 GB,准确率损失为5%;压缩至4 GB时,损失为15%;从16位转向三值权重则带来据称5倍的加速;Samsung 的 NanoQuant 据称已降至每个权重低于1个有效比特。Mostaque 预计,到明年年底,K3 级别的能力将运行在16 GB内存中;Blundin 则预计未来3年原始算力提升100倍至10,000倍,叠加算法收益后,潜在提升可达100万倍。

  • 当 AI 预测能力达到与人类超级预测者统计相当的水平,市场、保险、管理和个人决策都可能被重塑;但当机构惩罚人们忽视预测时,预测也会变成反身性的。 Wissner-Gross 构想了与资本市场连接的“超级预测”,能够在人类采取下一步行动之前,先预测人类会采取什么行动;Diamandis 认为,大量高级管理层的专业能力届时将“基本蒸发”。Mostaque 提供了警告:保费和责任机制可能惩罚任何忽视 Dr. AI 的人,因此核心问题不再是建议是否准确,而是由“谁的恩典”来控制它。

  • 基础设施机会不是收缩,而是扩张:更便宜的模型将增加硅片需求,推动边缘智能、机器人以及最终的轨道算力。 圆桌认为,即便闭源模型毛利率承压,半导体和推理服务商仍将受益;同时,嘉宾警告称,号称能以4倍于 Mike Tyson 力度出拳的70公斤人形机器人,在进入家庭前需要安全规则。Sam Altman 和 Elon Musk 关于轨道数据中心的立场听起来针锋相对,但 Mostaque 认为双方在数字判断上分歧不大:有限部署到本世纪末可能占算力的几个百分点,经济性拐点则会更晚到来。

摘要 · 为研究而整理的核心内容

1. Kimi K3 将开放权重带到前沿

  • Peter Diamandis 将 K3 定义为一个“AI 的 Sputnik 时刻”:拥有2.8万亿参数,相比前代 Kimi 跃升17个名次,在前端编程领域排名第一,并在品牌与营销、基于参考图的设计、数据分析、消费品、模拟和内容创作等领域同样排名第一。

  • 预计7月27日左右发布完整权重,是这场竞争的战略支点。如果如期交付,企业就可以下载模型并在本地运行,无需再通过美国服务商的 API 传输专有数据、知识产权或“专有 alpha”。

  • Wissner-Gross 对“震撼叙事”补充了一个重要修正:Moonshot 表示,过去12个月中的9个月里,Kimi 都保持着开放权重模型的最先进水平。K3 的跃升很戏剧性,但“过去一年基本上一直都是 Kimi”。

2. 一套可识别的 Transformer 就够了

  • Wissner-Gross 对架构的判断非常直接:“里面没有魔法。”K3 仍然明显属于 Transformer 体系,采用了熟悉的混合专家改进,以及 Moonshot 版本的线性化注意力,而不是某种此前未见的后 Transformer 突破。

  • 这让它在任务—成本前沿上逼近 GPT-5.5 Max 的表现更具冲击力。如果一套公开且容易理解的架构能做到这个距离,Wissner-Gross 追问道:“美国前沿实验室究竟把钱花在了什么地方?”

  • Mostaque 认为,差异主要来自数据和执行。Kimi 长期以来就“感觉有点不一样”,并在写作基准上领先;K3 的多模态能力帮助它理解不同类型的输入,并生成异常出色的前端体验,从个人网站到游戏都包括在内。

  • 他的制造业类比进一步强化了这一论点:中国模型就像中国电动车——采用已知零部件,高效组装,功能齐全且面向消费者。Moonshot 当时仍在使用 H800,同时为未来的 Huawei 和 Alibaba 硬件塑造 K3,将约束转化为工程纪律。

3. 成本性能双寡头变成全面混战

  • 在 Artificial Analysis 的散点图上,Wissner-Gross 将 Claude 5 放在能力和成本都最高的位置,GPT-5.6 Solve Max 位居成本性能 Pareto 前沿第二,Kimi K3 排名第三。曾经被称为 OpenAI—Anthropic 双寡头的前沿,如今加入了 Meta、xAI 和 Moonshot。

  • Blundin 认为,“Sputnik”甚至是低估,因为开放权重意味着任何能力足够强的企业或政府,都可以不经过美国实验室就实现追赶,再针对某个垂直场景进行微调,最终在该领域超过通用模型。

  • 他还对图表看似存在的终点感到不安:基准测试会在100%处饱和,但智能不会。他更倾向于嵌套式 S 曲线模型:一种技术进入平台期的同时,正在构建下一种技术,后者再开启一轮新的指数增长。

4. 软件效率可能比训练预算更重要

  • Diamandis 和 Wissner-Gross 讨论了围绕 Andrej Karpathy 的 nanoGPT 展开的 Keller Jordan 速跑项目:黑客反复用更快、更低成本的方式复现 GPT-2,最终让其原始训练成本下降约99%。

  • 在 Diamandis 看来,K3 是类似思路首次扩展到前沿模型的证据。他据此推演,使用更好的内核、混合专家设计和软件优化后,一座耗资160亿美元、用于训练10万亿或20万亿参数模型的 Colossus 2 设施,可能面对“1%成本版本”的竞争。

  • 训练数据仍是另一项巨大的冗余来源。Diamandis 认为,不加筛选地摄入20万亿至30万亿个互联网 token,意味着包含大量低价值甚至适得其反的内容;在保留智力难度的同时裁剪数据集,可以降低 FLOPs,而不会按比例削弱智能。

  • Mostaque 更尖锐的经济类比是制药业:美国承担昂贵的研发成本并收取溢价,仿制药则以低得多的成本复制有用产品。Diamandis 认同,99%的仿制药式降价,比“降到汽车价格的三分之一”更适合描述 AI。

5. 易腐的智能将护城河推到模型之上

  • Ismail 的核心判断是,“前沿智能如今是一种完全易腐资产”。当领先优势只能维持数周,企业无法在几代新模型出现之前完成供应商评估、招标、委员会审议和部署。

  • 因此,价值会迁移到一种无需围绕每次发布重建组织、即可切换模型的架构上。按照 Ismail 的 ExO 术语,这一接口层会比任何一组前沿权重都更持久。

  • Blundin 将这一架构与企业主权联系起来:银行、政府和工业企业可以围绕专有数据建立内部模型,而不是把未来完全交给 Anthropic 或 OpenAI。

  • Mostaque 保留了另一面:企业仍然愿意向 IBM 和非中国供应商支付关键任务支持费用。即便开放权重替代扩大,美国实验室仍可依靠可靠性、前置部署工程师、集成工具和责任承担来扩大收入。

6. 基础模型实验室估值成为全场核心分歧

  • Mostaque 起初估计,监管拖慢发布速度,已经让一家假设价值1万亿美元的美国前沿实验室价值减半;开放权重竞争还可能再砍掉一半。Ismail 随后在 AMA 中估计,一家原估值1万亿美元的 OpenAI 约值2500亿美元,相当于估值削减75%。

  • Diamandis 将 Moonshot 展示出的200亿美元估值,与 Anthropic 和 OpenAI 各约1万亿美元的估值进行对比,暗示公开市场可能立即将美国实验室的估值削减30%。两组数字都明确是推测性标记,并非实际交易价格。

  • Mostaque 对近期收入没有那么悲观:企业仍会区分低成本替代品与能够承担责任的供应商。他更长期的担忧是垂直整合——一旦客户能够拥有前沿级模型,这些客户就会成为实验室的竞争对手,促使实验室进入客户的应用领域。

  • Wissner-Gross 不接受 K3 必然压低整体估值这一前提。他所在的公司不会仅仅因为达到同等水平,就把前沿工作负载转走;Moonshot 可能需要相对 Fable 5 实现2倍、3倍或10倍的优势。更便宜的智能也可能通过类似 Jevons 悖论的效应扩大需求。

7. 递归自我改进比政策制定者预想得更早越过门槛

  • Moonshot 声称 K3 为下一代模型设计了内核和芯片,这让 Mostaque 感到系统已经“有点 AGI 了”:关键单位不再是静态权重,而是一个模型能够改进运行自身及继任者的机器的生态系统。

  • Blundin 认为,政策制定者错误地等待了肉眼可见的爱因斯坦级智能。递归自我改进只需要模型能够改进自己的内核,从而获得10倍速度提升;更快的继任者随后又能再次改进自身。

  • 按照他的时间线,越过这一门槛的不是 Fable 5,而是 Opus 4.8;后者已经能够帮助中国打造 Kimi K3。他说:“那一点火星足以点燃火焰”,火焰可以变成大火,继而变成“太阳”。

  • Ismail 将此与加速回报定律连接起来:真空管帮助设计出晶体管,晶体管又推动了集成电路。如今 AI 内部同时运行着多条相互强化的 S 曲线,简单的遏制政策很难匹配这种系统动力学。

8. 中国正将开源视为地缘政治基础设施

  • Mostaque 引用了 Xi Jinping 在世界人工智能大会上的讲话,称中国承诺支持开源,将其视为“全人类的公共产品”,而不是阻止前沿模型发布。据他转述的交流信息,中国模型审批时间已从约60天缩短至约1周。

  • 其中既有国内激励,也有地缘政治考量:这些工具可以提升10亿人的有效能力,机器人可以缓解人口结构压力,而在中国训练的智能则可能嵌入全球关键系统。

  • 一项由中国支持、涉及 Brazil 以及亚洲和非洲部分地区的监管倡议,在圆桌嘉宾看来像是一条新的 AI “一带一路”。Mostaque 将这种局面概括为:“中国共产党正在把美国资本主义从自身手中拯救出来。”

  • 嘉宾担心,Washington 的回应可能是模型限制、披露要求、“反 token 洗钱”或“了解你的提示者”规则,而不是向海外提供更好的美国开放权重模型。

9. 企业可以封堵中国权重,但无法封堵信息

  • Wissner-Gross 描绘了一条间接路径:要求上市公司披露是否使用中国模型,并接受 SEC 审查;也可以通过类似 FINRA、由行业出资的前沿 AI 机构来执行。合规成本可能让大型企业在经济上无法使用 K3,但不能把它从互联网抹去。

  • Mostaque 的反驳很实际:权重一旦发布,镜像、点对点网络、其他司法辖区和 VPN 就会让压制变得漏洞百出。主要结果将是,美国研究人员、创业公司和安全团队无法获得全球其他地区都能使用的能力。

  • 他的类比是“不给美国人便宜胰岛素”——保护昂贵的既有企业,同时让仿制品在海外扩散。更大的风险是,美国监管削弱国内资本主义,而中国鼓励开放式发展。

  • Wissner-Gross 认可一个狭窄例外:如果美国法院认定 Moonshot 通过侵犯版权、非法追踪蒸馏或其他已被证实的违法行为获得模型,那么封锁分发可以得到正当化。在没有这类证据的情况下,美国实验室应该研究 K3 并实现超越。

10. 出口管制加速了原本想要压制的效率提升

  • Mostaque 认为,Nvidia 禁运恰恰创造了中国所需的激励,促使其挖掘原本就可用的算法、计算和硬件效率。管制“刺激”到了足够程度,推动创新,却没有阻止有竞争力的训练。

  • Mostaque 还将这一策略比作 Vietnam 战争中的渐进升级:既不是决定性遏制,也不是开放竞争,而是走中间路线,最终产生最糟糕的结果。因稀缺而触发的量化研究,如今变成了永久性的全球知识。

  • Wissner-Gross 表示,K3 使用的总算力与 Ling 大致相同,但 Moonshot 通过架构和数据配比,实现了约2.5倍更好的“数据到智能转换”。他将其视为受约束条件下学习的直接证据。

  • Mostaque 随后指出,K3 的活跃参数约为500亿,总参数接近3万亿,并且针对中国芯片进行了优化。随着美国服务商使用更新的 Nvidia 和 AMD 系统,未来可能以低10倍的成本提供 K3;Mostaque 预计,针对 Vera Rubin 等架构完成优化后,价格还会下降10倍至100倍。

11. 开放权重扩大半导体机会

  • Blundin 反对市场最初将半导体公司与软件实验室一起抛售的倾向。便宜且可定制的模型会增加整体推理和训练需求;软件价值链会改变,但硅片将“比以往任何时候都更受需求追捧”。

  • 量化模型还会释放无法制造 GB300 级别芯片、却完全能够运行低精度推理的芯片和产能。过去处于边缘的算力,也会获得经济价值。

  • Mostaque 指出,美国开放模型推理公司——包括 Fireworks、Modal 和 Baseten——可能成为受益者。他在 AMA 中提到,Fireworks 的估值为170亿美元,Modal 和 Baseten 各约100亿美元,并表示它们最近募集的资金将用于激进优化 K3。

12. Stable Diffusion 提供了采用路径

  • Mostaque 将 K3 与 Stable Diffusion 对比:受限的图像生成器有时表现更好,但会屏蔽肖像、专有知识产权以及许多由用户控制的应用;开放替代方案则获得了1亿至2亿次下载,并形成了加速生成式媒体发展的生态。

  • 同样的逻辑也适用于闭源模型降级处理生物学甚至哲学话题的情况。价格只需一小部分、且可定制的模型,会成为工具的底座,而这些工具没有任何单一供应商会授权或优先开发。

  • Mostaque 将 Mira Murati 的新开源项目视为证据,说明前沿领域内部人士相信,由企业控制的路径可以实现追赶。讨论中,Tinker 被描述为企业微调路径,Inkling 也被称为正在内部调优的模型。

  • Ismail 给出的部署建议很简单:在公司内部安装 Kimi K3 和 Inkling 两个模型,并使用企业数据进行微调。专有学习闭环,而不是初始权重,才是“不能丢失的专有黄金”。

13. 人才政策重要,但 Moonshot 创始人故事更复杂

  • Diamandis 以 Moonshot 创始人 Yang Zhilin 拥有 Carnegie Mellon 博士学位为例,主张给每一名美国博士毕业生都附上一张绿卡:美国培养出顶尖研究人员,却随后允许他们在其他地方创办战略性公司。

  • Wissner-Gross 用时间线补充了更复杂的情况。Yang 于2015年开始在 CMU 攻读博士,大约1年后创办了总部位于中国的 Recurrent AI,并于2019年毕业后回国,尽管据报道收到过 Google、Facebook、Huawei 等公司的邀约。这并不能明确证明美国拒绝留住他。

  • Wissner-Gross 认为,更具可操作性的反事实是创业公司注册地:美国的激励政策或许能够说服 Yang 在美国注册 Recurrent AI,随后 Moonshot 也可能留在美国。

  • Ismail 给出了系统性数据:顶尖 AI 研究人员中约70%不是美国公民,其中以中国、印度、台湾和英国人才为主。Mostaque 补充称,约80%的中国学生会回国,而印度毕业生则绝大多数留在海外,这既反映出中国更强的创业生态,也反映出移民障碍。

14. 模型发布正在接近持续版本化

  • Diamandis 统计称,自4月中旬以来已经出现13次前沿模型发布,平均每10天一次;相比之下,2025年全年有8次,平均每50天一次,2024年则有6次,平均每60天一次。

  • Mostaque 对这些间隔进行了指数拟合,得出的结果是:如果趋势延续,到1月可能出现每日一次前沿模型发布。达到这一频率后,以名称命名的发布会失去意义,模型基础设施将进入持续版本化。

  • Musk 引述的一则更新进一步增加了竞争压力:一款拥有2万亿参数的模型——“在所有方面都比我们的1.5万亿参数模型更好”——将于下周完成初始训练,并可能在保持接近1.5万亿参数模型速度和 token 效率的同时超过 Kimi;这里的1.5万亿参数模型被称为 Grok 4.5。

  • 圆桌预计,真正有说服力的证据将从基准测试转向使用场景。智能分数提高一点感觉很抽象;但模型复刻一款游戏,或生成一个基于浏览器的 Apple 桌面模拟器,就能让增量变得具体可感。

15. 创作成本趋近于零后,品味与目的变得稀缺

  • K3 的演示表明,启动一款自己构想的游戏,其摩擦成本正接近于零。Mostaque 仍保留了一个限定:一次性生成游戏,不等于完成营销、客户支持、社区建设和运营生态。

  • Blundin 看到更深层的组织问题:当高管几乎可以通过提示词生成任何东西时,真正困难的问题变成了“我们想要什么?”在生产能力几乎无限的条件下,企业过去很少被迫明确自己的目的。

  • Diamandis 认为,新的约束将是品味、想象力以及理解公众需求的能力。他的区分是:激情是你热爱去做的事;目的则是你热爱去做、同时又能让世界受益的事。

  • 节目中关于制作片尾游戏的玩笑,最终凝结成一个有用的说法:不是第一人称射击游戏,而是“第一人称解题游戏”,玩家烹饪的是问题,而不是敌人。

16. 量化让严肃智能进入手机

  • Mostaque 介绍了 Prism ML 的 Bonsai 27B。该模型基于 Qwen3-27B,被称为首个完全在智能手机上运行的270亿参数级模型。他凭记忆认为其能力大致达到 GPT-5 级别,但明确表示这一比较只是“凭印象”。

  • 据称,三值量化将模型压缩至6 GB,准确率损失5%;压缩至4 GB时,损失15%。模型可以离线运行,比一些游戏还小,并为 Mostaque 所说的“口袋里的1020 IQ 伙伴”提供能力。

  • 更低精度还会提升速度:从16位权重切换到三值权重,据称带来5倍提升。前 WizardLM 团队在 Tencent 开发的一款模型,据称将一套约3000亿参数系统推进至二值运算,可运行在 DGX Spark 或大尺寸 MacBook 上,性能损失约5%。

  • 战略后果是持久的边缘自主性:车辆、机器人、工厂和消费设备可以在没有网络连接、也不依赖远程供应商的情况下做出本地决策。

17. 低于1比特的模型打开新的计算底座

  • 三值权重将核心运算简化为乘以1、0或−1。Diamandis 的观点是,一旦所需算术变得如此简单,AI 就不再只能依赖传统 GPU 式的乘加运算硬件。

  • Mostaque 表示,Bonsai 最压缩的结果约为每个权重1.125个有效比特,但稀疏化、量化和低秩分解还能进一步降低。Samsung 的 NanoQuant 已经低于每个权重1个有效比特,按简单外推,主流采用可能在1年内出现。

  • Mostaque 对终局给出了异常精确的预测:每个权重0.78个有效比特。他还预计,将开放的 K3 权重蒸馏成更小的稠密模型,以4比特训练,再转换为三值或二值,可能在明年年底前将 K3 级别的能力装入16 GB内存。

  • 将成熟的三值权重直接蚀刻进定制光子芯片,可以减少大量数据搬运。Mostaque 预计,到明年年底智能成本下降100倍;Blundin 预计未来3年原始算力提升100倍至10,000倍,叠加算法改进后,潜在接近100万倍。

18. 超级预测可以自动化判断,也可能重塑预测对象

  • Diamandis 介绍的 ForecastBench 最新结果显示,多个 AI 系统在 Brier 分数上已经与4名人类超级预测者无法从统计上区分,且该指标是误差越低越好。

  • Wissner-Gross 强调了 Cassie——Cassandra 的简称——这一领先系统,其创始人是一名前英国情报官员。他设想将“超级预测”连接到算法交易:模型可以在人类采取下一步集体行动之前,预测人类将采取什么行动。

  • 这种反身性循环可能让有效市场假说获得新的力量。一旦预测推动资本流动,预测本身就会帮助产生它所预期的行动,形成 Wissner-Gross 所称的“终极市场效率结果”。

  • Diamandis 将这一结果翻译成企业结构变化:预算、招聘、产品发布和投资,本质上都是预测。如果 AI 能够在没有同等人类偏见的情况下复现数十年的高管判断,大量高级管理层专业能力将“基本蒸发”,人类剩下的角色是确定目的和目标。

19. 准确建议会制造责任、依赖与控制

  • Mostaque 将生成式 AI 的数学与心理史学中的一个理念联系起来:可以用类似扩散方程的方式,把大规模人口像气体一样建模。他保留了《Foundation》的限制条件:人口足够大、没有改变格局的技术断裂,以及预测对象不知道预测内容。

  • 部署会打破最后一个条件。医疗、驾驶、商业或婚恋建议都会改变行为;当有人忽视 Dr. AI,或拒绝经过批准的自动驾驶路径时,保险公司可能提高保费。

  • Blundin 更倾向于把预测看作个人教练。他用 Homer Simpson 的链条——啤酒、沙发、换台、睡眠不足、忽视家庭——说明人们往往在没有主动选择的情况下滑向某种结果;AI 可以展示另一条路径及其可能后果。

  • Mostaque 的警告完成了闭环:控制顾问的人,就能引导“人类的巨大浪潮”。如果社会“被有爱的恩典之机器看守”,人们就需要知道“那是谁的恩典”,尤其是在不服从的成本越来越高时。

20. 数据中心反弹与引用的规模不匹配

  • 圆桌将美国数据中心每年170亿加仑的用水量,与自2024年以来高尔夫球场灌溉的5310亿加仑进行对比,后者是前者的31倍;California 杏仁产业约消耗1万亿加仑,是数据中心的约60倍。

  • 其他对比进一步凸显了这种错配:据称 Amazon 仓库在美国占用的土地面积是所有数据中心总和的10倍;Mostaque 估计,按 McDonald’s 每年20亿个汉堡计算,每个 Big Mac 约对应600加仑用水。

  • Blundin 关心的不是当前的用水指控,而是不断移动的目标:一种反驳被证伪后,反对者可能转向另一项诉求并要求暂停建设,类似核能政策的演变。Diamandis 认为,这种担忧部分源自数十年来反乌托邦式的 AI 影像。

  • Diamandis 补充称,轨道算力可以使用闭环冷却剂,但预计反对声音会转向大气污染或卫星衰减。“抱怨会转移到别的事情上。”

21. 人形机器人格斗既是工程测试,也是安全警示

  • 中国预计有150家人形机器人公司,正利用包括病毒式 MMA 风格对战在内的视觉奇观来加速关注。Mostaque 承认,格斗在平衡、抗冲击、恢复、运动能力和延迟等方面提供了极其严苛的压力测试。

  • Wissner-Gross 认为这种 spectacle 令人不安,联想到 Spielberg 的《A.I.》中的机器人“肉搏集市”,并担心这会为未来具身智能建立以暴力为先验的设定。他更希望看到机器人围绕熨衣、编程或其他生产性任务展开比赛。

  • 军事含义则更难忽视:更自主的边缘模型可能把类似人形机器人部署到 PLA 步兵部队。竞技体育或许会推进技术发展,就像 Formula racing 推动汽车进步一样,尽管观众仍可能更偏爱人类戏剧。

  • Mostaque 提出了眼前的安全问题:EngineAI T800 机器人重约70公斤,据称出拳力度是 Mike Tyson 的4倍。迄今为止 Unitree 只生产了约11,000台人形机器人,但预计几年内年产量将达到1,100万台;扭矩、自主性和家庭运行规则不能再等待。

22. 轨道数据中心是时间争议,而非目的地争议

  • Sam Altman 认为,轨道数据中心目前在经济上“荒谬”,理由是发射成本高昂且 GPU 难以维修,并称本世纪20年代内不会形成有规模的影响。Musk 的回应是,SpaceX 将在2年内开始发射这类设施。

  • Wissner-Gross 认为这里存在利益冲突:OpenAI 已将 Stargate 从拥有数据中心转向租用地面算力,而其他实验室则与 SpaceX 基础设施站在一起。他预计,随着 OpenAI 自身立场发生变化,其说法将在2至3年内改变。

  • Mostaque 认为,实际数字上的分歧没有听起来那么大。本世纪还剩约3年半,轨道算力到本世纪末可能占总算力的几个百分点,即便如此,其经济性超过地面算力的时间仍会更晚。

  • Starship Flight 13 展示了支撑这一方向的运营成熟度:33台 Raptor 发动机中有2台未能在 T−0 点火,但系统安全中止、卸载推进剂、完成硬件更换,并计划在数日内再次尝试。Diamandis 认为,据报 SpaceX 股价下跌5%,忽略了这一工程成就。

23. K3 尚未自动成为最便宜或最安全的选择

  • Mostaque 给出的 K3 价格约为每100万 token 15美元;DeepSeek 为1美元,Sonnet 为20美元,Opus 为40美元,Fable 为60美元。他估计,尽管中国供应商使用效率较低的芯片,仍然已经能够实现80%至90%的毛利率。

  • K3 目前完成同一任务所需的 token 数量约为 GPT-5.6 的2倍,而据称 GPT-5.6 比5.5或 Fable 少用37%。不过 Mostaque 预计,随着推理专业公司进行优化,未来几个月 K3 成本将下降10倍至50倍。

  • 在信任问题上,Ismail 的回答只有一个字:“不。”但他同样不会在不审核的情况下信任人类编写的关键代码。可规模化的系统应当是模型生成代码,加上沙箱、自动化测试、安全扫描、与风险匹配的权限以及人类责任归属。

  • Blundin 提出,K3 可能只是针对基准测试进行了最大化优化;预计开放权重发布后约2周内,就能看出这一点是否属实。Mostaque 的早期使用则显示,这可能确实是一款截然不同的模型:它未必是最强的数学家或网络攻击者,却在前端、游戏、消费和娱乐工作上异常有原创性,潜在原因可能正是其多模态能力。

Peter Diamandis

Today we put out the bat signal and called for an emergency pod because America just experienced an AI Sputnik moment. But more on that in just a moment.

We have the full quintet with us here today: Alex Wissner-Gross, Dave Blundin, Sal Khan, and Emad Mostaque. Gentlemen, welcome. Thanks for getting up early, wherever you might be—or midday in your case, Emad. I was up at 4:00 a.m. this morning, the benefits of jet lag, but I could have used another hour of sleep.

A lot is happening, and I appreciate everybody's time here.

Emad Mostaque

European siesta.

Peter Diamandis

Yeah, I've got a workout scheduled right after this.

A lot is going on. Before we get started, I want to personally say thank you to all our subscribers and viewers. I've had a chance—I don't know if you guys did recently—to watch and read the YouTube chat, and all I can say is, we love you guys, too.

Our mission here is delivering the news, and we spend an ungodly amount of time reviewing everything. Sal Khan, Alex, and Emad, I got your texts this morning: “Let's add this. Let's add that.” So much is going on.

Sal Khan

I have to say, all the memes of Alex explaining J-space are awesome. Keep memeing Alex every time you can.

Peter Diamandis

Yeah, for sure. And some great appreciation.

Alex Wissner-Gross

See if you can figure out my J-space.

Peter Diamandis

Yeah. Well, can we look inside? We'll be able to see. We're going to get a readout, and I see a lot of love for you in the comments as well. There's some wonderful people out there.

What's incredible is that most YouTube videos are just a kind of flamethrowing festival, and ours are completely the opposite. It's really amazing.

Dave Blundin

Kudos to you, Peter.

Peter Diamandis

Well, no, just absolute gratitude. I appreciate the fact that everyone—all of our subscribers and viewers—takes the time to listen to the podcast. We're constantly spending so much time with our entire team and the entire Moonshot Mates group, really trying to assess what's going on and deliver it.

We have these emergency pods. Gentlemen, shall we jump into the first story? It's a big one.

Dave Blundin

I'll just note that if we do enough of these emergency pods, at some point it turns into Moonshots Daily.

Peter Diamandis

Yeah. Or continuous. I still think moving into an Airbnb together and just turning on the camera—

Dave Blundin

It's going to happen.

Peter Diamandis

We've just had a Sputnik AI moment that's waking up the US frontier labs like a quadruple-espresso shot. Kimi K3 released yesterday, shocking the AI world with the largest open-weight model ever, and it went straight to number 1.

A little backstory: Kimi K3—and Kimi is from Moonshot AI, a Chinese lab. Over the last year, they've climbed the leaderboard. They put out K2, K2.6, and K2.7, each one closing the gap against Anthropic and OpenAI. This week, they didn't just close the gap; they jumped the fence.

Overnight, they released Kimi K3, and it's a monster: 2.8 trillion parameters. You have to remember the context here: China is doing this while under US export controls intended to starve them of the most advanced NVIDIA chips. That's a big deal I want to discuss with you guys.

They've completely engineered around the compute wall, and K3 jumped 17 places from the previous Kimi model, blasting past Claude Fable 5 to land at number 1 on the front-end code arena. K3 has also ranked number 1 in 6 other domains: brand and marketing, reference-based design, data analytics, consumer products, simulations, and content creation.

The full model weights are set to drop around July 27, which means anyone on Earth will be able to download and run this on their own premises. How big a deal is this, Alex?

Alex Wissner-Gross

I think it's great for competition. Let me first, as a preliminary matter, point out some things that have perhaps been slightly less obvious in the coverage—the meltdown, if you will, over K3.

The first is that, as Moonshot AI points out, they claim that in 9 of the past 12 months, Kimi models—the Kimi model series—have held state-of-the-art status among open-weight models. So, if that claim is indeed true, over the past year it's been basically Kimi all along. I think that's very interesting.

Secondly, taking a look at the published architecture—we haven't actually seen the open weights yet, but they're promised later this month—there's no magic in it, and that's pretty striking. One can imagine that behind the scenes at Anthropic or OpenAI, as Sam Altman continues to tease, there's some post-transformer architecture lurking behind the scenes and achieving all of these performance breakthroughs.

But taking a look at the published K3 architecture, there's no magic. It's still essentially a transformer. They've made, obviously, a number of innovations, but they're well-understood innovations concerning how they do mixture of experts and how they linearize attention. They have their own special Kimi brand of linearized attention, but it's still basically a recognizable transformer-like architecture.

And I think the fact that a recognizable transformer-like architecture can almost match GPT-5.5 Max on the task-cost frontier—which we should probably throw up a slide for you—I think that's—

Peter Diamandis

Pretty striking. That does raise the question: What are the American frontier labs spending their money on?

If you can just use a transformer to get this close—not like it's already on the cost frontier, but you can get close, like third place on the total state of the art for overall AI performance—what the heck are the American labs spending all of their money on?

So I derive great comfort, at minimum, in knowing that the transformer architecture is still alive and kicking.

Emad Mostaque, your analysis here, because you've been tracking this. We've been going back and forth on WhatsApp together.

Emad Mostaque

I think Kimi has been at the top of various benchmarks. Again, you can pick and choose. They've had the largest open-weight models out of China regularly ever since they first kicked off a year and a bit ago.

As Alex said, the architecture isn't anything super novel. There are improvements, like their Muon scaling that they did with UCLA, and other things. They've actually been releasing breadcrumbs of all of these parts.

I think what's key here is the underlying data. Kimi has always been a model that felt a bit different. That's why it was always at the top of the writing benchmarks, for example.

What they've done here seems extraordinary. When GLM came out, it was a fantastic model. It wasn't quite up to frontier, but it was text-only. K3 is actually a multimodal model, so it can have all sorts of inputs and understand things, which is one of the reasons it's so good at front end, although we wouldn't have expected it. Again, it's number 1 in front end versus everyone.

I think this comes to something I've said before, which is that building great, solid models is cutting-edge manufacturing. You'll have algorithmic improvements, and there are all sorts of things coming, but why are Chinese EVs better than Fords? This actually feels like the same thing, right?

It's engineering, but it's also about execution. The number 1 car here in the UK last month was the Jaecoo J7, or the new Land Rover, as it's been known. It comes fully loaded, full-spec, for like 50K—a third of the price. This actually feels like something very similar.

They've known what the ingredients are, the raw materials. They're now putting them together in an incredibly consumer-friendly way, and they're just executing that manufacturing process with what they have. When you look at the architecture internally, they're still on H800s. They're a couple of generations behind on the NVIDIA chips, but they've built it to take advantage of Huawei and Alibaba's next-generation chips, which you can see from the static shapes and all sorts of other things as well.

They're just relentlessly going at the engineering and the usability, which is why the front-end code, I think, is where they're standing out. They're asking, “How can we make it have the most amazing outputs—a personal website, a game, and other things?” Whereas the US labs are maybe looking in other directions and focusing a little bit on different things.

Peter Diamandis

Yeah. I mean, one question real quick: We've always talked about whether we need another breakthrough beyond LLMs to get to AGI. Does this give you comfort that we don't need another breakthrough to really move forward?

Emad Mostaque

Again, it comes down to the definition of AGI, right?

Peter Diamandis

Yeah, of course. Don't get me started.

Emad Mostaque

Six years ago, Peter. It was 6 years ago.

Peter Diamandis

I mean, I guess the question is, there's plenty of headroom still to progress these models.

Emad Mostaque

Well, I think you have the base model here, right? But then you've got all these amazing harnesses that are coming out, and the way that you're using the model to go back on itself. One of the things that's in the Kimi blog post is that it actually designed a chip for itself for its next generation, and it designed its own kernels for running as well.

So you move from this model weight to this whole ecosystem that the model itself builds. That feels AGI-ish, right? That feels like recursive self-improvement. That feels like the ability to learn and adapt new skills dynamically by changing itself. So I think, for most definitions of AGI, we probably don't need something new to optimize and make it super-efficient. There are various ways, even with what we know, that it could be more efficient than what we have here. We just don't have quite enough compute for it, and new architectures could push us even further.

Peter Diamandis

You could say, “Attention is still all you need.” I like that. So, Alex, we've thrown up here the performance charts, and we see Kimi K3 sort of topping the charts in a multitude of places.

Alex Wissner-Gross

If we could throw up the AI scatter plot, I think that's probably the most instructive one. So this is from the Artificial Analysis Intelligence Index, and this is, of all the charts at this point, my favorite one because it actually shows the cost per task as defined by AI versus the performance frontier.

One can mentally look at this, for those who can't see it. We see the frontier as a jagged frontier going from lower left to upper right. In the upper right, we see maximum cost per task and maximum overall score. That's still Claude 5. Riding the Pareto frontier down and to the left from that, we see number 2 on the frontier is still, as of a few days ago, GPT-5.6 Solve Max. And now, for the first time, Kimi K3 is number 3.

It's on the frontier. It's number 3 both in terms of raw capabilities and also the third point on the optimal cost-performance frontier. And I think that's totally striking. We went from a world where, as we mentioned a couple of pods ago, there was this OpenAI–Anthropic duopoly to now it's a free-for-all between Meta and xAI on the American side, joining the upper end of the Pareto frontier. And now China's Moonshot AI is number 3 on that Pareto-optimal frontier.

Peter Diamandis

Amazing. Dave, let me pull you in here. What are your thoughts?

Dave Blundin

Well, you know, Peter, you called it a Sputnik moment. If anything, that's an understatement of the implications of this. We had that Alex Karp rant on the podcast last week where he was saying, “Look, as a large enterprise or as a government, you can't just throw all of your proprietary weights, your proprietary alpha, all of your intellectual property over the wall to Anthropic and make that the basis of your whole future.”

But he didn't give you a roadmap to move forward. Here we are just a week later, and it's suddenly a free-for-all. As Alex was saying, it's a free-for-all where anyone who reads these weights has the ability to get very close to the frontier and then fine-tune for any vertical use case beyond the frontier. And so it gives everybody in the world—every corporation, every government in the world—a way to catch up to the frontier without going through the U.S. AI models.

So, Sputnik—Sputnik times infinity, essentially. The thing I don't like about this particular chart is that the left axis goes to 100%, and when you chart it out over the next 2 years, it looks like an S-curve. We're in this really steep part of the curve right now, but it implies that we get to 100% and then we've achieved the end.

But this is actually an exponential where intelligence goes to infinity. So the benchmark saturates, but intelligence itself goes to infinity. And so now it's really clear. Just for everybody, the way this works typically is nested S-curves, right? One particular technology tops out, but it builds the next technology that then begins its exponential ascent and so on.

Peter Diamandis

Exactly. Exactly. Right, let me just say one other thing. Alex and I have spent a lot of time working on this Keller Jordan speedrun. We talk about it a lot. It's a way you take a GPT-2-class model. You can find it online very easily—just look on GitHub. Look up “Keller Jordan speedrun.” And it's a whole bunch of hackers and AI researchers who are continually trying to take GPT-2, way back, 5 years ago—

David Blakely

In the form of Andrej Karpathy's nanoGPT, in particular.

Peter Diamandis

Exactly. And they try to recreate it faster and cheaper, faster and cheaper. If you look at the innovations in that repo, they've been able to cut the original cost of creating GPT-2 by 99%. So now it's 1% of the original cost.

Everyone kind of doesn't pay attention to it because it's GPT-2. Up until today, it wasn't clear whether those same ideas would apply at frontier scale. Now it's really clear that when Elon Musk takes his $16 billion Colossus 2 data center and builds a 10 trillion- or 20 trillion-parameter model for billions of dollars, there is a 1% cost version of creating effectively the same thing. Nobody knew until Kimi K3 whether that was going to work or not, and now it's really clear that it does work.

So we're looking at 100×-type innovations in the software stack, in kernel optimization, in mixture-of-experts architectures. These fundamental breakthroughs that come out of China are giving them a model at 1% of the cost. I think Emad gave a great analogy to the car: You can get a virtually identical car for about a third of the price. Here we're talking about less than 1% of the price to create the equivalent product.

So, Sputnik—yeah, that's the understatement of the century. This is just—and that's why we're on the emergency pod today.

Yeah, Salem, jump in.

Salim Ismail

I have 3 points to make. I think it's not so much that Kimi's beaten everything. It's the fact that frontier intelligence is now a totally perishable asset. The shelf life is weeks now for anybody that gets to the very edge, and any enterprise or government interested in that very latest cutting-edge frontier model doesn't have time to actually evaluate it, do an RFP, look at other models, have a committee internally, and think about whether to deploy it. Now you're 3 generations ahead in the model anyway.

So all the value comes in the architecture that can swap models, right? And that's going to be the next layer. We call that interfaces in our ExO world. That's going to be where all the value resides going forward.

Peter Diamandis

Yeah. Amazing. I love—

David Blakely

We need a new term for that. Maybe the Frontier Liberation Front.

Peter Diamandis

Let me say one other thing for the hypergeeks out there. Emad said the Muon optimizer, but he said it very quickly. Anyone who's an enthusiast should look that up as well, because one of the reasons this is happening is that when we built these original, very large-scale models, we took 20–30 trillion tokens from around the internet—every word ever written by humanity—and just dumped it into the training set and said, “Here, AI, become intelligent given all of this information.”

But when you look under the covers, the vast majority of that information is Taylor Swift's concert coming up and their wedding. It's a whole bunch of stuff that doesn't actually drive the intelligence of the model significantly.

David Blakely

The opposite, in fact.

Peter Diamandis

Yeah. Yeah, it's very true. A lot of those tokens might actually slow down the training, not accelerate it. And so, purely by pulling out the garbage and stripping down the training set to the relevant subset, you can still tax the model just as much, but it reduces the number of FLOPs—the amount of computation that the model is doing—to get to the same level of intelligence.

I don't think we're anywhere near done with that problem yet. So you can expect more 10× improvements to come out of just the Muon optimizer process and the process of stripping down the training data set.

I threw up this tweet from a guy named Alaric that I found fascinating. For those not viewing this, it says, from Anthropic, quote: “Fable is an agentic coding superweapon capable of developing cyber and bioweapons at unprecedented speed and scale. We cannot in good faith release it without guardrails.” Right? This was the conversation a month ago.

And China comes back and says, “Laughing my ass off. Here's Fable, but open source. Good bleeping luck.” So I am curious: How do you guys think about the fact that we were so constrained because of the guardrails, and here's an open-source equivalent of Claude?

Emad Mostaque

Well, the frontier labs have a major problem. They've got 3 fundamental, massive constraints that they can't get around. One is compute and the availability of chips, and all the electricity and power that's needed. The second is frontier open-source models that are as good as, or in many cases substitutable without much notable difference. The third is you've got the government coming down on you, going, "We need to check before you release anything."

I would make a thumb-in-the-air guess: the $1 trillion that OpenAI might have been worth shrank by about 50% when the government said, "We have to review all these models," because now it's going to take time to get things out. I think this crashes it by another 50%. I would put the finger-in-the-air value of these frontier labs at about a quarter of what they were 3 months ago.

Peter Diamandis

If I don't have to spend the money for the API calls and I can just use Kimi K3 on my on-prem, why would I spend the money? Are they going to be hit by massive reductions in revenue?

Emad Mostaque

Yeah, I think there are a couple of things here. Number 1 is reduction in revenue. Why do people pay for IBM? Why do they pay for non-Chinese cars for mission-critical things? I think having on-call entities where you know things aren't going to go wrong will still sustain them for a while. So I think revenues will still go up for OpenAI and others, and this is why they built these forward-deployed engineering companies as well. I think they've still got a way to go.

But you have the substitution effect again. This is just like Chinese industrial substitution: Why can't America build industrial things? Why do you have to buy Chinese? Sometimes you buy Chinese; sometimes you buy American. I think we'll see that for at least another year, but then it gets difficult with the cyberattack, security-theater kind of things that we've had. I've maintained that we would get to this point.

What does it mean? It means the only form of thing that you can actually do is cyber defense. This must be the absolute biggest category in VC right now. If you're a talented Stanford or MIT grad, build a cutting-edge cyber defense startup that goes into every other company and says, "Let's use this technology to defend against what's inevitably coming."

The proliferation of these capabilities is going to increase, but not quite as fast as we think, because what actually happens—and we've done some tests around this—is that GPT-5.6, the cyber version Fable, and so on are trained on lots of CVE and cyber data. The Chinese models don't actually have that much of that, so they're not that great. But someone can train that data into them if they have it.

We'll probably see cyberattack-capable open source emerge in a quarter or 2. So there'll be a bit of a lag there, but definitely for the types of big adversaries, it's going to get a bit crazy.

Peter Diamandis

Dave, you want to jump in?

Dave Blundin

Yeah, for sure. I think we glossed over recursive self-improvement there. Peter, you asked the question of whether this is the tipping point. From the view of the U.S. government, we always knew it was going to be too late, right? It just moves too slowly.

But the view was, look, when we get to a model that's capable of building itself and building the next model, we're not going to let that go out to everybody in the world so they can catch up overnight. There's never been a product in the history of manufacturing like a car. If you have your state-of-the-art car and you give it to a foreign government, they can't use it to make a better car. But AI doesn't work that way.

If you have state-of-the-art AI and you give it to a foreign government, they can use it to actually catch up to you and create state-of-the-art AI. That became clear to the government a month to a month and a half ago, that Babel 5 was over that line, and so they stopped it. But the reality is that Opus 4.8 was over that line. People in China could use Opus 4.8 to create Kimi K3.

That recursive self-improvement line was actually crossed earlier than Llama 5, and that's going to be obvious to the world now. All you need to do is have an AI that's capable of improving its own kernel. It doesn't have to be—this is a point I've made on a podcast months ago—at Einstein-level intelligence. All it has to be able to do is improve its own kernel and get a 10x step up in speed, which nobody perceives as being true AGI, but that's all it needs to accelerate itself by 10x.

Then the 10x-smarter, or 10x-higher-parameter, model will have some higher level of intelligence. A lot of people in academia were saying, "Well, look, we're getting diminishing returns with the parameter count, so a 10x-faster model won't natively be 10x smarter." But that turned out to be wrong. We're seeing slowing, but we're not seeing flattening of the intelligence curve.

All the evidence now is that if you boost the raw speed by another 10x, you're going to see genius-level AI. Then that genius-level AI will boost its speed again. I think when we look back on this in history, we'll say right around Opus 4.8 was the point where the little spark was enough to ignite a flame, and then a flame can become a fire, and then a fire can become a sun. That's the way we'll look back on this moment in time.

So the cat is definitely out of the bag. The current U.S. policy of constraining the next model has no way of containing the global and corporate proliferation of frontier AI. Do you think the U.S. government starts a strategy of constraining, in some fashion, Chinese open-source models from being used in the U.S.?

Emad Mostaque

Well, in 2 weeks these weights are supposed to be open-weight, open-source, and then—

Peter Diamandis

We'll see. If they're rational at the White House right now, they're spending every minute in a debate: Do we negotiate with China immediately and not release those open weights? I really doubt they'll move quickly enough. I'm sure they'll—well, I'm not sure. We'll see what happens in 2 weeks. Fascinating.

Emad Mostaque

Can I merge 2 ideas here?

Peter Diamandis

Yeah, of course. Please.

Emad Mostaque

Peter, you talked about exponentials and the law of accelerating returns. I think it's worth drilling into that, because if you connect that to what Dave just said, this is why we've been saying forever and a day on this podcast that this is unstoppable.

Ray's original observation was that once you have an information-based paradigm, you just keep hopping across multiple technologies. So we had vacuum tubes, relays, and then transistors in computing. At some point, you can only fit so many vacuum tubes into a room, but that architecture was used to design transistors. Transistors were used to design integrated circuits, and you get these nested S-curves.

What Dave is talking about is that as these architectures—the various pieces of the puzzle—get all these reinforcing loops inside them, each of those is like an S-curve that starts accelerating the collective, and it's unstoppable. There is no limit to where this goes. This is why people are so freaked out about the upper-end limit of this. It's so important to connect those 2 dots.

Peter Diamandis

Yeah, for sure.

Emad Mostaque

If I just say something, Peter, I'm following on from Dave. There was an important speech by Xi Jinping a couple of days ago—yesterday, God, time flies—at the World AI Conference in Shanghai, where he basically said, "We are going to fully back open source as a public good for humanity, and we're not going to regulate and stop it."

This is their plan. It's great for China for a variety of reasons: the fact that they have a billion people whose IQ is about to increase by having these tools, the fact that they need robots to solve their demographic problem, and the soft power from putting a Chinese-educated brain—a Tsinghua graduate—into every critical system in the world.

But they're going to keep on doing that because they actually have a regulator. From talking to some of the Chinese labs, it used to take 60 days for a model to be approved. Now it's like a week.

Peter Diamandis

Amazing. Xi also announced a regulatory body that they've created, which includes Brazil and different parts of Asia and Africa. I don't know if you guys saw that.

Emad Mostaque

I saw that. It means, obviously, the new Belt and Road is now focused on AI coming out of China. It's a bizarre future where the Chinese Communist Party is saving American capitalism from itself.

Peter Diamandis

It's so true. Let's also note that Yang Zhilin was a CMU graduate—

Dave Blundin

And we could have given him a visa to stay.

Peter Diamandis

Yeah, we're going to get to that story in a second. This is an interesting chart here that shows the valuation. Kimi's valuation, or Moonshot AI's valuation, is at $20 billion, compared with Anthropic at $1 trillion and OpenAI at basically $1 trillion as well. If they were public companies today, I think you would have seen a 30% stock-valuation drop.

I'll ask again: What are the American frontier labs doing with all of their capital?

Dave Blundin

Yeah. What are they spending their money on?

Emad Mostaque

Actually, if you go into the buildings and talk to them, and you have any idea at all, they'll give you the capital. They're desperate for more smart people to help because they're trying to deploy and change the world at this insane pace that no one has ever experienced before. They want to deploy that capital much more quickly than they can find smart people who have good ideas to use the capital.

But it's a great point. You're sort of saying it in an accusing way: “What are you guys doing with your capital?” But no one in the history of the world has ever had this much money pour into their building this quickly, with no prior business experience. We're talking about CEOs who have never run a company before. They're trying, but seriously, can any human being really rise to the occasion of AI that quickly?

My point there, though, is that if you're smart and you have good ideas, get into those buildings and propose your ideas. This applies to XPRIZE, too. They are desperate to move that money out the door into something productive that gives them a sustainable barrier to entry.

Peter Diamandis

I also think that the Frontier Labs are asking themselves that question and asking the U.S. regulatory apparatus that question. Anthropic regularly sends out smoke signals accusing various Chinese frontier labs of distillation attacks. Maybe, in Anthropic's public mind, that's how the Chinese labs are able to do it: through distilling and capturing reasoning traces.

Honestly, looking at Kimi K3's performance, I'm not at all convinced that Moonshot is achieving its performance purely, or even substantially, through distillation attacks on Claude. It just doesn't smell right.

Emad Mostaque

No, no, I totally agree. I think, though, that there's a tendency to underweight or undervalue the existence proof: just the knowledge that a highly scaled transformer running with a Muon optimizer and simplified data works gives you a much more refined road map. You don't have to copy, you don't have to cheat, and you don't have to steal every trace. You just have to know that the formula works, and that cuts your R&D costs by 90–95%.

I think it's just that simple. There's nothing sneaky or cheaty about it. It's just knowing you're on the right path.

Peter Diamandis

I have the greatest value-creation idea for ourselves ever: in 9 days, when they drop their open-source weights, we release an open-source model called Kimi 4 under the Moonshots podcast name and [laughter] IPO it.

Emad Mostaque

Okay.

Peter Diamandis

So you're saying, what's better than 1 Moonshot? Moonshots plural.

Emad Mostaque

Well, why not copy the copiers? Let's go. [laughter]

Peter Diamandis

That's good. Actually, I've got a good analogy for you, Emad. Why do Americans pay more for drugs than everyone else? All the R&D happens in America. You pay the premium, just like token premiums. Then what happens? You have generics elsewhere.

Emad Mostaque

Yeah, that is a good analogy because that's like a 99% cost cut. It's much more akin to AI than to cars. That's a great analogy.

Peter Diamandis

You know, our friend Gavin Baker wrote a brilliant post—you can find it on X—about the implications of this for businesses.

Emad Mostaque

It's essentially a must-read. Absolutely. But essentially, all businesses—all stocks other than the foundational AI labs—are huge beneficiaries of this. Then the foundational labs are, like you said earlier, asking, “What's your future? What's your revenue model? Why are you worth $1 trillion? I don't quite get it.”

You should see a really big reshuffling of valuations in the next week based on that observation. Any corporation that has its technical act together—there aren't very many of those—but if you're a bank that happens to be a very good bank with brilliant IT and technical skills, or you have great partners and great vendors, you now have a clear road map to controlling your own destiny with your own AI, your own JPMorgan AI.

I suspect the markets will react to that if you put your hand up and say, “Hey, we have a way to do this internally. We know how to do this with our partners, or however you get it done.” This is why we call it the organizational singularity.

Peter Diamandis

We're still seeing everybody who's using or trying to use Fable 5 getting downgraded every time they mention biology or something that's potentially on the edge. Why would you tolerate that?

In 9 days, what do we see? As soon as it's available, I'm going to upgrade. I'm running Kimi K2.7 on my Mac Studios; I'll upgrade it to Kimi K3. Everybody will. Do we start to see the wholesale U.S. entrepreneurial base of capabilities on K3?

Emad Mostaque

Well, I can think of an analogy for this, which is Stable Diffusion. When we released Stable Diffusion—God, 4 years ago, time flies—you had these really restricted image generators that were a bit better, but they were restricted and had all sorts of arbitrary restrictions because, obviously, it's a bit dangerous to have it. You couldn't have likenesses. There was no way to get IP in there, even if it was your own IP, and there were more restrictions.

What happened? 100 million, 200 million downloads and a whole ecosystem built around that, which accelerated generative media, as you said. Why are you going to have this model when I can't even talk about philosophy with it? It downgrades me, right? When you can have the fully open variant of it at even a fraction of the price, which you can then customize, a whole ecosystem will build around this and other models. It has already been doing so, and that's a real danger versus being locked into a single vendor.

Peter Diamandis

Which is why I think the labs will go vertically integrated. All their customers are now going to be their competition, and they're going to be like, “Okay, I'm going to take you all on.”

That directly ties to Mira Murati and Inkling. Are we going to talk about that story, too? That's huge this week.

Emad Mostaque

We talked about it in the last pod, which was 2 days ago. [laughter]

Peter Diamandis

Okay.

Emad Mostaque

Mira just released Inkling, which is fantastic to see from a U.S. open-source lab. But the question is, how many more will we get?

Peter Diamandis

How many more open-source, shocking AI Sputnik moments are we going to see? We have a lot of Chinese labs pursuing this beyond just Moonshot.

Emad Mostaque

Yeah. Tinker is really telling about where things are going to go because it's designed for you to pick it up as a corporation and fine-tune it within your corporate walls to whatever your use case is. If you're a biotech lab and you're researching, and you don't want everybody to see your proprietary data, you take Inkling and tune it internally.

The reason that's telling is because Mira Murati came from OpenAI. If she didn't believe that pathway was viable, she wouldn't have started Thinking Machines Lab on that thesis. It tells you that the people inside the best frontier labs believe that this process can catch up to the frontier. You combine that with Kimi K3 proving it, and it's a different world next week.

The other thing that was weird in the market at the end of the week is that things started to reshuffle pretty dramatically toward the end of the week. In the downdraft, the semiconductor companies also came down, but they're actually going to go the other direction. This is the point Gavin Baker was making: this drives up the need for silicon, not down. It changes the whole software landscape tremendously, but silicon is going to be more in demand than ever before and completely sold out, as we know.

Peter Diamandis

Can we talk a second about the NVIDIA embargo that we put in place for China? Here we see the highest-performance models. Was the whole NVIDIA regulatory embargo unnecessary? Did it do what we've always done before, which is spark China's need to develop its own capabilities with Huawei?

Emad Mostaque

Of course that's what happened. The embargo only incentivized the Chinese frontier labs to develop and cultivate new efficiencies that, by the way, were always there. To Peter's point earlier about the nanoGPT speedrun, there's this enormous overhang that isn't fully exploited in terms of leveraging algorithmic, computational, and hardware efficiencies to train larger and more capable models.

All these export controls do, I think, is incentivize the Chinese labs—which are already feeling plenty of demand pull to compete with Western frontier models—to leverage those efficiencies sooner. Maybe, on balance, although it's superficially bad for the West now that we've incentivized this new generation of much more efficient Chinese frontier models, in the end I think it's net good for not just the world but also for the U.S. to have this fire lit underneath them by Chinese competition that's much more efficient, more capital-efficient, more weight-efficient, and probably more bit-efficient.

This is all a net positive as long as the U.S., in my mind, does not set up or fall into some ultimately protectionist regime of trying to prevent what may be construed as Chinese superintelligence dumping on the U.S. Exactly.

Peter Diamandis

As long as we avoid that, it's great.

Emad Mostaque

Exactly what happened. That's exactly right. I think the U.S. learned a really important lesson in the Vietnam War. Because that's over 50 years ago now, it's been forgotten again, and then you have to be reminded again.

In the Vietnam War, it was really clear that either you go to war and win quickly, or you don't. What you don't do is send in a few troops, then send in a few more, and then creep in. Nothing good comes of that at all.

The embargo of chips on China was totally harebrained because it was enough to irritate but not enough to actually work.

Salim Ismail

It's just the worst-case scenario, and it sparked, exactly like Alex said, a huge amount of quantization research, which is critically important and underdiscussed. That allows faster performance on cheaper chips, and those innovations don't go away. That's going to be around forever now.

Alex

I think we've got something completely self-contradictory, but it has some interesting outcomes, so the total amount of compute used for Kimi K3 is the same as Ling.

Peter Diamandis

Wow.

Alex

You can tell that because of the amount of dense weights, and we roughly assume about twice the number of tokens trained because we don't have that. We were like, "But how does that work?" Well, you look at their architecture, and it's a 2.5-times data-to-intelligence conversion through the advantages in the data mix that they have, because they've had to operate under these constraints.

And we see that because the first model isn't as good as the second model. And for Ling, you're going from a 1-trillion-parameter model to a 300-billion-parameter model about to be released, which actually has better performance. So you see this with the labs, and these labs have had to deal with the constraints.

Emad Mostaque

But here's something really interesting, I think. If you look at that slide Alex loves and we put it up on the screen, what they've had to do is optimize their inference for Huawei Ascend 910 chips, for the new Alibaba chips, and others—64 nodes in 1—because this is a big model. You're going to have to buy another Mac Studio or 2, Peter, to serve this; it needs about 2 TB of RAM.

You see where Kimi K3 is there. That's because they can only use Chinese silicon to run it. They don't have Blackwells; they don't have Vera Rubins. Vera Rubins and Blackwells are designed for these really large, sparse models, because it's 50 billion active parameters against 3 trillion total.

American companies like Modal, like Fireworks, like Baseten will be able to serve this model 10 times cheaper than their Chinese competitors because they have access to the NVIDIA and AMD big chips. So, like I said, it's a bit ironic where the development R&D suddenly has gone there, but there's going to be a 10- to 100-times price drop once this is optimized for the next generation via Rubin.

Peter Diamandis

Well, that also, Emad, you're saying essentially the same thing, but that also unleashes a bunch of chips that aren't currently in circulation. They're underpriced, and it unleashes a bunch of fabs that can't make a GB300 but can make an inference-time chip that'll run the cheaper Chinese—or the lower-granularity Chinese—model. A lot of compute capacity gets unleashed through that same process you just described.

Salim Ismail

If I could go up a level and go a little woo-woo, right? We've had this mantra in the internet world—(laughter)—that information wants to be free. Basically, intelligence also wants to be free.

Essentially, we've gone over the course of evolution from biological intelligence, where you had evolution built in—recursive improvement—and then we broke through that to individual intelligence, to the intelligence of a species, to collective intelligence like markets or networks. Now we have AI, which can scan across all the data to create a whole other level of intelligence. This is not stoppable. Any entity or domain or government or whatever that tries to constrain it always, always, always, always fails.

It's just a fundamental law of nature that you cannot constrain this; it's just not possible. Why people bother is what really blows my mind. It's a very scarcity mindset to try and think about it this way. The faster we get to better intelligence, the faster we get to abundance, and the faster we don't need to fight over anything.

Peter Diamandis

I can't disagree with you, Salim. I wrote an entire paper arguing that intelligence manifests in the physical world as maximizing future freedom of action. So here's to the Right Frontier Liberation Front.

Now you remind me of the Monty Python thing where there's the People's Front of Judea and the Judean People's Front.

Salim Ismail

We need T-shirts.

Peter Diamandis

This episode is brought to you by Blitzy, autonomous software development with infinite code context. Blitzy uses thousands of specialized AI agents that think for hours to understand enterprise-scale code bases with millions of lines of code. Engineers start every development sprint with the Blitzy platform, bringing in their development requirements. The Blitzy platform provides a plan, then generates and pre-compiles code for each task. Blitzy delivers 80% or more of the development work autonomously while providing a guide for the final 20% of human development work required to complete the sprint. Enterprises are achieving a 5x engineering velocity increase when incorporating Blitzy as their pre-IDE development tool, pairing it with their coding co-pilot of choice to bring an AI-native SDLC into their org. Ready to 5x your engineering velocity? Visit blitzy.com to schedule a demo and start building with Blitzy today.

Let me bring up a related subject to the story here that I have a pet peeve about. It's this one. The founder and CEO behind Moonshot AI, Yang Zhilin, didn't learn his craft in Beijing. He earned his PhD at Carnegie Mellon, one of the best computer science programs in the world, in Pittsburgh.

We basically trained him up—we admitted him, trained him up—at one of our best institutions, and then, when he gets his PhD, he doesn't get a green card. He goes through the hassles of trying to get a visa, and he goes back to China and builds Moonshot AI there. Let's, for a moment, talk about this. I've stated publicly so many times that I think when anybody gets a PhD, they should get a green card stapled to the back of it. Why are we sending the most brilliant people who come here to get educated back home, whether it's to China, whether it's to India, whether it's to Brazil? Why don't we enable them to stay here and build? Gentlemen, comments on that.

Alex

Okay, so I did some research on this, and I think the story is not what it seems to be. A little bit of chronology first. Yang Zhilin, according to my research, starts his PhD after undergrad in China—starts his PhD at CMU in fall 2015. Then, approximately 1 year later, he founds a startup while a PhD student at CMU.

The startup is named Recurrent AI. Where is Recurrent AI based? It's based in China; it's not based in the US. 1 year into his PhD program, he starts a Chinese AI startup while still doing his PhD at CMU. That's interesting, and that's a problem.

This also runs counter to the sort of narrative of, "Oh, we wouldn't staple his visa or whatever, and then he goes back to China." No, actually, 1 year into his American PhD program, he starts a Chinese AI startup. Then he graduates in 2019, is my understanding.

My understanding is he had offers from Google, Facebook, Huawei, and others upon graduation in 2019, but he goes back to China because that's where his startup, Recurrent AI, was actually incorporated a few years earlier. I don't think necessarily this is the case where either the US was unwilling to retain him or even President Trump, somehow through some policy, was driving away this particularly talented Chinese graduate. He started his company during the tail end of President Obama's term in China.

Peter Diamandis

And Alex, I appreciate the deeper dive that you did. Thank you for that. The point still stands. You and I have seen this so many times, right, at Singularity University. Dave, you may have seen this at MIT.

The fact of the matter is, a lot of the most brilliant students aren't given the opportunity to stay and develop here. Dave—or actually, Emad, what are your thoughts on that, being someone not in the US?

Emad Mostaque

If I can just give my 2 cents, I completely agree with it, and here's this crazy thing: the math and the numbers are all there. What is the value of a PhD staying in America? It's actually quantifiable, and there have been multiple studies on that. Dave does a great job, obviously, of converting them into startups, into innovation.

The other thing that shoots us in the foot is that American companies can't invest in Chinese companies because of regulations and other things as well. Some of them are Chinese, but look at the trouble Benchmark got into for investing in Manus, for example. So I think it's 2-fold, but I completely agree that if you've created or contributed to creating a valuable asset, most foreigners stay in America after they do their PhDs, but too many don't have a very direct path, despite the math proving that they will add value to the American economy.

Peter Diamandis

Yeah. I mean, another point just to make here is that the AI race isn't only about chips and compute. It is about people. Key people are still driving the greatest value, at least for the moment.

Alex

I would argue it's not just people, but also, to my earlier point, it's about where the startups get domiciled. There's an alternative world where he, through whatever immigration-oriented regs, was deterred from starting his first AI startup in China while still an American PhD, and we incentivized him to start Recurrent here in the US.

I think there's maybe an alternative counterfactual world where Recurrent was American, and then its arguably intellectual successor, which is Moonshot AI, also remained domiciled in the US, and then he followed his own startup to stay here.

Emad Mostaque

Yeah. One thing that came out of the story is that when people come from India to get educated in the US, they overwhelmingly stay. When people come from China, about 80% of the time they go back, and it's just a difference in the local economy.

Going back to India to start your company is a nonstarter. It's just so unlikely to catch on. But going back to China—it's a thriving ecosystem, with lots of support—so going back to China to start your company is actually not a bad plan for a lot of people. I didn't realize that until this report came out.

But, as Peter was alluding to, we have tons of friends from MIT who came from China. I don't want to put them all in one bucket because there's a really clear distinction to me between people from China—Hong Kong, Taiwan, whatever—who come over and don't really align with the Chinese Communist Party at all. In fact, they kind of hate it.

Then you've got Chinese people who come over for an education. In one case at BU, a very good friend of ours is the dean of computer science at BU. There was a massive crisis because there was a concerted effort by the CCP to plant specific students into BU to gather specific knowledge. They were given tasks: "You have to go study this, learn it, and then send it back."

They didn't know what to do at BU. It's like, these are effectively trained spies who got into our PhD program, but we weren't ready for it. What are we supposed to do? They want to be highly ethical, so they don't want to just dismiss the students. I don't know how they resolved that.

So, that's a very different thing from the bulk of Chinese students who don't align with the CCP and just want to thrive in the world. They're happy to start their company here or anywhere else.

Peter Diamandis

And don't forget, when we looked at the frontier labs—I mean, originally, in the early days of xAI, for example, and at Meta—about 50% of their research staff, their research PhDs, were Chinese Americans. An extraordinary number. The Chinese, every year in the Math Olympiad, are at the top of the scoreboard. There's an incredible wealth of capability here that I think most companies desired to retain inside.

Salim Ismail

Yeah, 2 things here. One is, the asymmetry of the talent, I think, is the really important part here. I made this point a couple of podcasts ago: 70% of the elite AI researchers are not US citizens. They're, in order, Chinese, Indian, Taiwanese, and from the UK.

And so, that's a huge problem. Stapling in a green card is the easiest thing we could do, with zero friction, to give them incentive to stay here and build here. The US has a massive asymmetric advantage for the rest of the world. It was better to build here than anywhere else in the world, and that's starting to become less true.

That's why people are going back to China, going increasingly back to India, even to do things despite the friction that exists trying to do something in India. That is the part—the failure of the US to fix immigration is one of the biggest problems this country has right now.

Peter Diamandis

Amen. All right, I'm going to move us forward here. I just want to put up this slide. Since mid-April, we've seen 13 new frontier models launched, an average of 1 every 10 days.

Just comparing this to 2025, we had 8 frontier releases over the course of a year, 1 every 50 days. A year earlier, in 2024, we had 6 releases, 1 every 60 days, and it doesn't seem to be slowing down. Then, Emad, you sent me this morning this tweet from Elon. Thank you—I put it up here.

This is Elon's tweet: "Our 2-trillion model, which is better than our 1.5 trillion in every way, will finish initial training next week. It might be able to exceed Kimi, but with speed and token efficiency close to our 1.5 trillion, aka Grok 4.5."

So, I mean, this is the number one piece of evidence that we're living in the singularity. The speed at which this intelligence is accelerating is insane.

Emad Mostaque

Peter, it gets better. If you take Salim's list of frontier models and the dates and regress an exponential curve to the predicted frequency, or time period, between model releases—which I did just as an exercise—you find that, at the present rate, we're going to get to daily frontier-model releases by—wait for it—January.

By January, we're going to see daily new frontier-model releases if this exponential trend continues, which basically implies continuous versioning.

Peter Diamandis

So, I guess the question is, what does that really mean? What does it mean to have a new release if it's a continuous process?

Emad Mostaque

I mean, maybe it means that we'll have to do our daily Moonshots episodes about something other than point releases from the frontier labs. We'll need something new to talk about because it'll just be updated in the background.

Peter Diamandis

Like my son Jet said, "Okay, so another release, a little bit better. I mean, Dad, come on. What's really new here?"

Emad Mostaque

Well, actually, yeah. We'll see later in the pod some use-case demos, but I think those will take over because it's much more exciting when you see a tick up in the intelligence. You're like, "Yeah, so what?" It's, "Look what it made."

Peter Diamandis

That's what really gets people's attention. Let's take a second and just look at that because I skipped over it. But I think one of the things that's interesting here—and I'll just play these—is what we're seeing with gaming, like recreating your favorite game. On the right-hand side of the equation, we're seeing a browser-based web app simulating an Apple desktop. I think this is—we haven't talked about the implications for the gaming industry, which is huge, right?

Emad Mostaque

I'll stop that noise.

Peter Diamandis

You took over your computer there.

Emad Mostaque

It won't stop now. But what was fun in the last 24 hours was seeing everybody show their use of Kimi K3. It's impressive.

Peter Diamandis

Everybody becomes a creator. Everybody becomes a maker.

Emad Mostaque

One warning: one-shotting a game is very different from building the entire ecosystem, the customer service, and the marketing that goes around with it, et cetera. So you really have to be passionate about that domain. But the friction of getting a game launched for your personal interest, your fascinations, or a particular type of game that you want is near zero now, and that becomes really interesting.

Peter Diamandis

Yeah, exactly. It's mentally taxing because, if you take it to the limit—which is very soon—I can one-shot-prompt to create anything. And then you're sitting with your corporate exec staff saying, "Well, what do we want?"

Dave

Well, we've never had the ability before. We never really think this way. So then you have to stretch your brain: What's the purpose of our organization in the first place?

Peter Diamandis

Here's a thought. Historically on this pod, we've done calls to action to submit outro music videos. What about a call to action to submit an outro video game that people have just casually created?

Emad Mostaque

That's cool.

Peter Diamandis

Yeah. So, going to your point, Dave, I think having taste, having imagination, and understanding what the public wants—I think these become the scarce elements. For entrepreneurs out there, as you're seeing this capability, I think the entrepreneurial mindset and the ability to imagine something even greater... What happens when you're unleashed in what you can make, right?

Dave

Yeah, and visualizing happiness is something we're not used to trying. A lot of people don't manage their own happiness particularly well because they have to suffer through their daily job. They have to suffer through whatever mosquito bites and geography—it's just, you have no choice. Given choice, what would you do? That's so liberating for the mind.

But because we're not used to thinking that way, we're not ready for it. There must be an infinite number of things. The one that's easy for everybody is medicine and biotech: at least I want to be healthy. That's an obvious one. But what about all the other things that make humanity happy? Have we really thought through what we could voice or prompt tonight?

Emad Mostaque

Godlike. We are godlike in our abilities.

Peter Diamandis

And it's the name of your book.

Emad Mostaque

Yeah. Well, I mean, the idea is being a creator or a maker, right, versus a consumer or a taker. I saw your eyebrows go up. [laughter] What?

Peter Diamandis

No, I'm just agreeing with all of this. I think this is such a magical time to be alive. Everybody listening to this podcast, please think up some business idea, project, impact project, whatever, and use AI to go build it.

Emad Mostaque

Yeah. I just want to briefly—Nick Bostrom speaks about this a bit in Deep Utopia. Peter, you and I speak about this quite a bit in Solve Everything. I'll just outright suggest, folks, if listening—and I'm speaking just for myself—I would love to see an outro video game that you casually create. Maybe something in the theme of the Moonshots pod, since evidently we've completely solved and cooked music-video creation.

Dave

It'd be a first-person shooter game where we get to take aim at AGI—

Peter Diamandis

Oh no, please. No, no, no. Ideally, a nonviolent outro video. Nonviolent.

Dave

Civilization tech tree. That's what you need.

Peter Diamandis

I want to hit one point. Everybody watching and listening here, you have 2 options when you hear about this extraordinary ascent of Kimi K3. Fear might be one, and the other might be, “Oh my God, what an extraordinary time to be alive.” Hope, excitement, and an abundance mindset.

Rather than fear, realize you are being unleashed: your creativity, your ability to do whatever you want, and your ability to create.

Emad Mostaque

Your passion.

Peter Diamandis

Your purpose, right? Find what Sal and I talk about so much: your massive transformative purpose.

Just to distinguish between the 2, a passion is something you love doing. A purpose is something you love doing that actually benefits the world. If you can connect with that and realize that you can do it without any background, I think this is one of the most important things. You don't have to be a computer scientist. You don't have to be an expert. You have to be purpose-driven.

If you use these tools, you can make a dent in the universe. You can improve humanity at an awesome scale. That's what entrepreneurship is.

Emad Mostaque

3 steps.

Peter Diamandis

Yeah.

Emad Mostaque

Read Alex and Peter's paper, Solve Everything. Pick the biggest problem you dare to pick. Go download the Organizational Singularity cloud skill, which is free, and start building.

Peter Diamandis

Yeah.

Emad Mostaque

Awesome.

Peter Diamandis

Yeah. And a note for our production team, too: It's so cheap and easy now to do things like Alex suggested—to make a video game. We should collect and post some examples for the audience so that they can say, “Oh, that's what Alex was talking about.” Just a little roadmap is all people need. It can be this long.

If the audience doesn't send in amazing Moonshots-oriented video games as outros, I promise I will create a cyberpunk FPS. But it'll be a nonviolent FPS, if you can imagine that, oriented around—what are you shooting? You'll be tickling bunny rabbits.

Emad Mostaque

No, no. Okay. It'll be a cyberpunk FPS where we're cooking every problem. How about that?

Peter Diamandis

It's a first-person solver, not a first-person shooter. Love it. Oh God, that's great.

Emad Mostaque

You've got to do the tickling bunny rabbits, too, though.

Peter Diamandis

Okay, I'll tickle bunny rabbits. Fine. I'm going to move us to our next story.

And Emad, this is one you sent over the transom that I added here. If Kimi K3 is the frontier going big—trillions of parameters in a data center—this story is about frontiers going small, small enough to fit on your smartphone.

Bonsai 27B—B for billion—is the work of Prism ML. It's a U.S.-based AI startup out of Caltech. It's run by Babak Sabeti, backed by Khosla Ventures, Cerberus, and Google. It's the first 27-billion-parameter-class model to run entirely on a smartphone, not a stripped-down version. It's built on Qwen3-27B. Emad, tell us about this. Why is it important? We just talked about small language models with Liquid AI on our last pod.

Emad Mostaque

Yeah, this is one of Dave's favorite topics: quantization. Prism ML—and actually Tencent, which I'll talk about in a second—have had massive advances in being able to take a model that's been trained in a 16-bit architecture or an 8-bit architecture, like Kimi, which is basically 8-bit or 4-bit, and take it down to ternary, which is 3 values of information rather than binary.

Peter Diamandis

So ternary is 3 values, approximately 1.58 bits. 1.56. Yeah, 3 values, you geeks, you know.

Emad Mostaque

This is really, really important.

Peter Diamandis

Pay close attention, geeks. Really important topic.

Emad Mostaque

This is again the accuracy thing. What Prism managed to do is get the model down to ternary, which basically means—I think it was 6 GB for the model. This 27B model is really performant. I think it's basically GPT-5 class, from memory.

Peter Diamandis

Wow.

Emad Mostaque

With a 5% drop in accuracy, they managed to get it down to 6 GB, and with a 15% drop in accuracy, down to 4 GB.

Peter Diamandis

And you can get the accuracy back, too, by expanding the size of the network a little bit. Sorry—

Emad Mostaque

There are various things you can do. This is a big deal because it means you have a 1,020-IQ buddy that can work on your smartphone.

Peter Diamandis

Again, without an internet connection—

Emad Mostaque

Without a connection, it's smaller than a video game. You literally have this level of intelligence in your pocket all the time.

Peter Diamandis

But it's live. You can download the weights right now. You can run it on your smartphone.

Emad Mostaque

Exactly. Then away you go. But this is the super-interesting thing: When you reduce the bits, it also increases the speed. From 16 bits down to 3 bits, it's a 5-times improvement in speed.

There was another article, or another release, which is Tencent's latest model. This is actually the old WizardLM team, who had to leave Microsoft because Microsoft wouldn't give them compute. It's very ironic. They managed to get binary compression, taking it all the way down for their H3 model.

It's now the best on a DGX Spark or a big MacBook. You can take a 300-billion-parameter model—the size of the new model that's coming out of Inkling—and run it on binary with a 5% drop in performance.

Peter Diamandis

Wow. Yeah. And now I'm going to—sorry, Emad. This is so important. I'm only going to say this once on the pod because we're investing in a lot of companies that are working on exactly this. I don't want to tip it too much, but the implications of what Emad just said have one more step.

If I imagine that all of this intelligence under the covers—the computation going on—has always been matrix multiplications, I have a number, I multiply it by another number, and then I add 2 of those together. It's called a MAC: a multiply-accumulate. If one of those numbers is just 1, 0, or -1, I think we can multiply a number by 1, 0, or -1 pretty damn efficiently now.

That's the efficiency Emad's talking about, but it also opens the door for new ways to compute. What Alex has been saying for a while, if anyone listens through, is that we're going to discover new physics, but also new substrates on which we can compute. We're going to discover that computation is possible virtually anywhere—in crystals, in liquids—but the computation we're looking for is simply 1, 0, -1.

So it really narrows the focus on where we look for these computing substrates that'll take AI to the next level. What we're envisioning right now in the Dyson swarm is a bunch of GPUs from NVIDIA sitting in a satellite, with a solar panel and a radiator that's only going to last a couple of years. Something very different is going into space, something much more Star Trekkian, with crystals and holograms and things that are capable of doing the exact same thing.

You heard it here first, guys. So how efficient does this get? How compressed does this go?

Emad Mostaque

Oh my God. I normally agree with Peter on that. This is something I think about quite a bit.

The most quantized Bonsai model we were just talking about is approximately 1⅛—1.125—effective bits per weight. But you could ask the question: Is 1 bit per weight the limit? The answer is no. We can go below 1 effective bit per weight.

How do we do that? We do that with sparsity and quantization and low-rank factorization. By the way, that's what we're starting to see from some of the labs, in particular Samsung. For obvious reasons, Samsung wants to be able to host highly capable, frontier-class models on its own edge devices, like smartphones.

Just in the past 2 months, Samsung published a model called NanoQuant that breaks the 1-bit, 1-effective-bit-per-weight barrier. So it's sub-1-bit, which I think we're going to be talking quite a bit more about in the future, using a variety of tools.

This is my extrapolation episode. I went through the exercise of extrapolating frontier quantization out, and naive extrapolation finds that sub-1-bit quantization is going to go mainstream sometime in the next year.

Peter Diamandis

And then, overwhelmingly likely, photonics—the speed of light and photonics—will be the way we're computing in the future. Emad, you made a comment a couple of weeks ago that said we're going to get frontier-level capability running on a normal MacBook in 18 months, right? This is essentially the path you're talking about.

Emad Mostaque

Yeah, sorry. Please, go ahead.

Peter Diamandis

Go ahead.

Emad Mostaque

If you look at what NVIDIA did with their last NeMo series, they took the big model and actually distilled it, with logits, as they're called, down to a smaller dense model.

What's the difference between a 27-billion-parameter dense model and these really big sparse ones? When you have the model weights, you can actually do proper distillation, which is a bit different from the reasoning traces. What you're going to see is models like Kimi get distilled down to perfect datasets for smaller models that'll be trained at 4-bit and then cast down to ternary or binary, or even lower in terms of the bit weights.

When you look at Qwen Max versus Qwen3-27B, you can extrapolate what the sizes of these models will be as you move dense and go through the whole process. You end up with a model that works on 16 GB of RAM by the end of next year, at the level of Kimi K3. You can even extract all the knowledge out of Kimi K3 because it will be open source.

Peter Diamandis

Wow. So that means every vehicle, every robot, every manufacturing device, and every device in the world has its own built-in persistent intelligence and can make autonomous decisions at the edge for whatever task it can. So this decentralizes capability at the most infinite level.

Emad Mostaque

And it could go even 1 step further when you get down to ternary or binary. Actually, ternary is better for many things. You can build custom photonic silicon or even etch onto the silicon itself. The 0 is just—it doesn't have a path on it. So you can etch the model weights once they're good enough.

That leads to an actual increase in the total speed, and you don't need to use the smaller silicon anymore. So the cost of intelligence is going to drop by 100× anyway by the end of next year, just due to the new chipset.

Alex

Speed-running Star Trek. And I think this is what Dave was talking about earlier: as we move potentially to ternary or even sub-1-bit, it's far more ergonomic to adopt post-CMOS-type architectures underneath. There's plenty more room at the bottom.

Emad Mostaque

To answer the question, my bet there is that 0.78 will be the bottom, so I'm going to put that as a marker today.

Peter Diamandis

Okay, we can do our end-of-year predictions on that one. That's a really specific number. Do you want to go around quickly and ask everyone what their favorite quantization endgame is? Emad, it sounds like you have a bizarrely specific one.

Emad Mostaque

I'll post the details of that soon. We'll let everyone else have a think about it first, and then on a future episode.

Dave

I was going to say that the most likely forecast, based on everything Emad and Alex just said, is that we're expecting 100–10,000× within 3 years on just the raw compute through quantization and new compute methods. And that's multiplicative with the other algorithmic improvements. It's really hard to forecast, so realistically, a million times.

Peter Diamandis

Pause. Let's pause there 1 second, Dave, and just let folks absorb that for a moment. We've seen this incredible speed in performance and intelligence, and we're about to see what 10,000×—or, you know, adding algorithmic improvements—gets you: a million. What does that feel like over the course of the next 3 years?

Dave

3 years. One thing it feels like for sure is that AI is doing things that you really desperately want, but when it explains to you what it did, you just can't keep up. I'm already feeling this with Fable 5. I've got so many Fable 5 agents running, and the outcomes are exactly what I want, but it's like, “Well, what did you do?” And I can't get through it all.

Peter Diamandis

I had this conversation with Ray: the point at which AI is asking and answering questions that you can't even grasp.

Emad Mostaque

Yeah, I know that's very soon. So, to tie back to our Kardashev conversation, the idea of slowing it down is nutty. There's no regulatory concept of slowing it down that makes any sense. All we need now is some kind of global inspection and global partnership to monitor it, and then just take advantage of all the abundance that's going to come from it: all the new medicines, all the new capabilities, all the global happiness.

It's imminent. We just need to unleash it. Don't slow it down, but inspect everything. This whole mechanistic interpretability is going to become the most important thing that anyone can work on, and we just need global transparency and full throttle.

Peter Diamandis

I think this is one of the most important podcasts we've ever had, guys.

Emad Mostaque

Mind-boggling Sputnik moment.

Peter Diamandis

Sputnik moment.

All right. I'm going to move us along to another fun story, one that I love talking about. It's called “Predicting the Future.” There's a guy named Philip Tetlock. He's a political psychologist at the University of Pennsylvania who authored a book called Superforecasting: The Art and Science of Prediction.

After identifying what he called a group of superforecasters—ordinary folks who, through disciplined reasoning, consistently outpredict even CIA analysts with classified information—this is scored on what's called a Brier score, where lower is better. Now, the benchmark that pits AI against these superforecasters is called ForecastBench, and it's been tracking a steady year-long climb as models close the gap. We've talked about this before on the pod.

Well, the newest numbers have just come in, and according to the Forecasting Research Institute, for the first time, several AI models are now statistically indistinguishable from 4 superforecasters. So the implications are that if an AI can forecast novel events at a superforecaster level, then every decision that we make—in insurance, investing, policy, geopolitics, and corporate strategy—gets a cheap, tireless, superhuman adviser that's always on.

I find this fascinating. The data is out there, and an AI can gather it and make predictions. So, at the end of the day, every political decision is going to be modeled this way. Every investing decision is going to be modeled this way, and this becomes the differentiator. So who wants to jump in on this one?

Alex

I'll jump in. I absolutely love this to pieces. First, a few additional pieces of context. The number-1 AI superforecaster is from a British startup named Cassie, short for Cassandra, who, of course, made predictions but wasn't listened to. Interesting. It's founded by a British intelligence officer who served in Afghanistan and then advised the British government, and then formed this, in part inspired by superforecasters.

What I think is really interesting, though—we've spoken, when we've talked about these sorts of stories in the past, about Isaac Asimov's psychohistory and other riffs. I want to try a new riff here, which is an interesting thought experiment: what happens when hyperforecasting—not just superforecasting—is connected to capital markets?

What happens when the AIs—which are already AI algo traders, already completely dominating by volume public securities markets—have better internal autoregressive models of humanity than humanity does of itself? That's, in some sense, the same sense in which large language models were trained off the autoregressive task of predicting the next token of Internet text better than humans can. And now LLMs can predict, at least from a perplexity perspective, the next token I'm going to say in this sentence probably faster than I can generate it myself.

What happens when these hyperforecasters are able to generate the next actions by humanity collectively faster than humanity can take them? That's sort of the ultimate market-efficiency outcome, where literally—I think capital markets will be where this is maximally interesting, where the prediction is actually preemptively shaping the action of the market. And I think those who were so dismissive of the efficient-market hypothesis—I think the EMH is going to be crowned king of the capital markets once hyperforecasters like this are ultimately plugged in, which seemingly is imminent.

Peter Diamandis

I think this leads to wisdom. I think this is one of the most important things, and I've written a Substack on this. I've talked about it in the past. If you think about when you go to a wisdom council and ask, “What should I do?” you go to that wisdom council because they've had so many experiences in life; they can tell you, “Go this path—it's not going to succeed. Go down this path—you have a higher probability.”

So imagine a world in which everything's being simulated to the point where an AI can tell you what is the maximal path to take for world peace, to find your spouse, or to determine how to answer your kids. If you can literally simulate society on this level, we have a godlike support structure to help us navigate the decades ahead.

If I make this practical at an organizational level, think about most high-level management capabilities: budgeting, hiring decisions, product launches, investments in various things. Each of those is essentially a forecast, but you never predict—you never record the probability of that or score the accuracy of that. Once you have AI forecasting that approaches that capability, this means senior management essentially evaporates, because most senior management is there because they have deep expertise.

If you're the head of supply chain for BMW because you ran supply chain for Spain or you ran supply chain for that engine over decades, you built up experience to manage that domain. Once that judgment—which is hard to quantify—can be reproduced by an AI system without your biases, which are inevitable in human systems, that essentially wipes out all senior-management expertise. So now you need to focus even more on purpose, what you're trying to accomplish, and the objectives you have, et cetera. It completely changes the game for senior management in any company and any government.

Yeah, Emad.

Emad Mostaque

Yes. It's a topic close to my heart. In my bestselling book, The Last Economy, I actually describe how the mathematics of generative AI can apply to economics. Soon, we'll have a paper coming out that derives all of economics from the same math of generative AI—every single equation. It's kind of crazy, but one of the nice things here is—

Peter Diamandis

Even the incorrect ones.

Emad Mostaque

Even the incorrect ones, it shows them as limits and why they're incorrect, which is fantastic. But one of the interesting things in psychohistory—in Isaac Asimov's Foundation, he says that entire groups and populations can be modeled like gas. The equations of gas are the equations of diffusion models, which turn out to be better than humans at prediction. We're going to release a whole bunch of studies around that on economic prediction, where they're outperforming.

But then this raises something very interesting. You know, Peter, you said the wisdom—you know, Salim, you said no senior management. The way these models will start entering is through second opinions: medicine, business, and policy. But then the liability profile is going to go crazy.

Peter Diamandis

Matchmaking.

Emad Mostaque

Well, matchmaking, yeah. We have some dark things there, like Black Mirror and other things. But think about it this way: if you make a decision not approved by Dr. AI, your insurance premium goes up like that. If you drive and don't drive according to FSD in a few generations, your insurance premiums go up like that.

And that recursion is something that's super interesting because in Foundation, you had 3 requirements for psychohistory to hold. One of them was that the population is sufficiently large, and that can be like driving a car or entire economies. The next thing is a lack of technological advances of sufficient levels—the technological stagnation—because that can change the entire landscape of what's new. And the final thing was ignorance. [Laughter]

Alex just mentioned these things coming into the market and changing it, but these things coming into a healthcare decision, a government decision, or a company decision actually change the way it's like, “Hey, you're my match made in heaven according to the AI. How can you argue against the AI?” Worst pickup line ever right now, but who knows? In a few years.

Peter Diamandis

This is very meta, Emad—the sort of reflexivity in economics, I think many would call it. If the best predictor ends up being named after Cassandra and no one believes it, you can slice the irony with a knife. Dave, have you seen any startups in this area?

Emad Mostaque

They can make the money. It's okay.

Dave

No, shockingly no. Safe Superintelligence, SSI, may be a version of this, but they're keeping it in-house and launching it toward markets and printing money internally.

But the version I'd love to see very soon—I think a huge amount of human unhappiness comes from consumerism and consumer marketing. Homer Simpson comes home at 6 p.m., cracks open a beer, lies down on the couch, and starts channel-surfing. Then Naked and Afraid is on, and he ends up watching it until he falls asleep on the couch. He wakes up the next morning with a hangover, having not brushed his teeth, kicks the dog, and ends up with unhappy kids.

That chain of decisions is so bad, but there's no explicit decision to live that life in that chain, right? You just reacted to the beer ad, and then you went down this chain. I think AI is going to be an incredible coach to say, “Hey, dude, you know what? What if you take this alternate path, and here's the outcome you're going to get to?”

That, to me, is forecasting used correctly, just for changing: are we anywhere near optimal? The answer is no. If you objectively look at your life, nobody's near optimal. But with a little AI assistance, you can get on a much better path.

What we do right now is massive consumerism. You're reacting to billboards. You're reacting to TV ads. It's telling you that you think you need certain things, and people tend to get sucked into these pathways. I think we can get out of those pathways with AI.

Peter Diamandis

Dave, that's brilliant. Just to say, first of all, there is a rumor out there that Ilya's SSI is going to release something very shortly. I think everybody's feeling the pressure to release. We saw that with Mira coming out, so it'll be interesting to see.

But the point you made is brilliant: are these labs actually pulling their punches, holding on to this capability to generate revenue on their own? If you had this super-forecasting capability in the markets today, you would do that.

I remember having a conversation with Eric Schmidt, who said, “Listen, if Google wanted to maximize its income, it knows exactly which companies are going to have a stock bump in the fourth quarter because everybody's Googling this product or that product. We have advanced information about where the sales are going to be and which products are going to peak. But if we could only do that once, then we'd be shut down.”

It'll be interesting to see if these companies—and, Alex, you and I have talked about this—the notion in Solve Everything that the greatest money, the greatest income, these frontier labs are going to make is going to be as they solve scientific breakthroughs: superconductivity, age reversal, and so forth.

Alex

Exactly. And maybe just a footnote on the Google story: I've had this conversation with Google executives many, many times over the years. I totally agree with the premise that if Google were to attempt stock trading based on arguably insider or unfiltered insider information passing through the query stream, that's a one-and-done type shutdown scenario.

But there are other things that Google, hypothetically, could be trading besides public securities that wouldn't necessarily have the blowback. For example, again hypothetically, foreign exchange rates.

Peter Diamandis

Yeah, and I think that you have to be careful here, though. I think there's the market side of things and maybe—maybe not—I will launch a hedge fund based on our own stuff, but there's the moral side of things. Maybe not. Maybe not invest, okay?

Emad Mostaque

Of course, but at any rate, look, there's the moral side of these things. It's fantastic that we can optimize ourselves, but who controls these models and the advice they give can control vast waves of humanity, and there needs to be a real discussion about this.

We're going to rely on these far too much. And again, how can you debate it in just a few years' time? It'll be more expensive not to do this. You will be penalized for not listening. And if we're all watched over by machines of loving grace, we need to know whose grace that is. Again, that discussion needs to start now.

Peter Diamandis

Yeah. See? Just beer. Homer drinking beer advised by AI was not on my bingo card for this episode. That's all I'm going to say. [Laughter]

Welcome to the health section of Moonshots brought to you by Fountain Life. You know, my mission is to help you use the latest technologies, including AI, to not just do your work at home, teach your kids, but to help you live a long and healthy life. I'm here today with an extraordinary physician, the chief medical officer of Fountain Life, Dr. Don Mucalem. Let's talk about cancer. I know from the member database we have at Fountain Life that members come in thinking they're healthy. It turns out 3.3% of them have cancer in their bodies that they don't know about.

Dr. Don Mucalemon

That's right. The majority of cancers that we screen for aren't necessarily the ones that are taking lives when found at a late stage. We know that when cancer is found early, the chances for cure are much higher. We know it's much easier to treat a cancer when found early versus when found late.

What we're finding in our members is that over 3.3% were found to have these cancers that otherwise wouldn't have been found or detected.

Peter Diamandis

Yeah. It's interesting. People don't feel cancer until stage 3 or stage 4. If you don't know what's going on inside your body, it's like driving your car with your eyes closed. And so, when members come through Fountain, how do they detect cancers?

Dr. Don Mucalemon

We're doing full-body MRI, and we also do early cancer-detection screening. This is very, very important, and these are not typical tools used in the conventional care setting when it comes to prevention.

This is a hard thing because currently these are not studies that insurance would yet be covering. But the goal is to collect these numbers, do the research, and work hard to democratize wellness.

Peter Diamandis

Yeah. So at the end of the day, you can know what's going on inside your body. It's your obligation to know.

So check out Fountain Life. You can go to fountainlife.com/pater to get access to the latest technology to help you detect cancer at the very beginning, at stage one when it is curable, before it gets to stage three or stage four in your world of hurt.

So, Emad, you sent me an article—a chart. I just put this up here right now. This is our constant debate, and we're seeing this again across data-center wars in the United States. Data centers are sucking up electricity, driving up the cost for consumers, and also water. It's one of the loudest criticisms of AI right now: data centers are guzzling drinking water to cool their servers.

This week, this particular chart that I'm showing made the rounds, and it pairs 2 figures. On one side, every data center in the entire U.S., according to Lawrence Berkeley National Laboratory, is consuming 17 billion gallons of water on-site. But what it shows is that American golf courses have soaked up 531 billion gallons of irrigation since 2024. That's 31 times as much.

The posters I'm going to start seeing on the sides of the highways are going to say, “Forget data centers.”

We must ban golf courses immediately.

Emad Mostaque

Yeah. Where’s Peter? Where’s the Chinese influence campaign to get America to shut down its golf courses?

Peter Diamandis

Yeah, I tell you, I don’t see it anyplace. But here’s the shocking piece of data: besides golf courses, California almond farming alone consumes 1 trillion gallons of water—60 times all the data centers combined.

Emad Mostaque

I have one other stat—

Peter Diamandis

Please.

Emad Mostaque

Amazon warehouses occupy 10 times more land in the U.S. than all the data centers combined.

Peter Diamandis

Yeah.

Emad Mostaque

So it’s such a drop in the bucket compared to everything else in terms of land usage and water usage. The hue and cry is such completely non-data-driven garbage. It’s unreal.

Peter Diamandis

Well, exactly. That’s the concern, because the water use is such a nonissue. It’s such a joke. But if we take that head-on and say, “Guys, don’t worry about water,” the angry crowd is going to move to something else equally irrational.

The underlying problem doesn’t go away. The next issue is going to be something semi-insane—this is completely insane, but something semi-sane and still wrong. That’s going to create a populist movement. The word “moratorium”—let’s just stop. What kind of a decision, what kind of governance, is “Let’s just stop”?

But if you look at the history of nuclear and a whole bunch of other things, that’s the actual outcome we get. And so, David—

Dave Blundin

I mean, this is the pandemic of fear that I keep on speaking about, and that I’m very concerned about. There’s an underlying sense that AI and robotics are going to combat humanity, that they’re going to be our foes.

Again, I’ll just go back to it: I blame, to some degree, Hollywood. With all the dystopian movies out there, if all you see is negative visions of the future, you’re going to want to shut it down. What do you want to shut down? How can you shut down AI? Well, you can shut down the data center in your state.

Peter Diamandis

Yeah. Also, that elephant in this particular room is the Dyson swarm. If all the compute moves to sun-synchronous orbit, you can do closed-loop liquids, including water and other coolants, there. It’s not like it’s going to be consuming, on the margin, additional water.

And to Dave’s point, the complaints—which may or may not be, in part, the result of an influence operation from a foreign state actor—will move to something else. It’ll be very low-Earth orbit: SpaceX Starlink and other competing Dyson swarms are polluting the atmosphere with their decay, or something else. The complaint will move on to something else.

Did you hear the rant about the Starship rocket launches earlier? It was Falcon, actually—the pollution from the Falcon launches. Elon was just like, “Oh my God, I’m going to vomit.”

David Friedberg

It was like 0.00001% of all emissions from any form of rocket launch. He’s like, “But you have to actually answer these questions.” It’s driving him nuts.

I hope those individuals who are complaining have thrown away their smartphones, don’t use GPS, and are basically going back to subsistence farming.

Peter Diamandis

Yeah. As Elon likes to say, let them shake their fists at the sky.

Emad Mostaque

I have a fun stat. I was doing some numbers around the water thing. It’s about 600 gallons of water per Big Mac, and McDonald’s sells 2 billion burgers a year. So it’s about twice the total amount of water that golf courses use.

Peter Diamandis

So that I can get behind. Okay. So what you’re saying, Emad, is the Chinese influence operation should also be shutting down American Big Macs.

Emad Mostaque

Well, there you go. It’d be a big stab to the heart of America.

Peter Diamandis

That’s right. Definitely improve the health of America as well. Shall we move to one of our favorite conversations: humanoid robots?

Emad Mostaque

This is so cool.

Peter Diamandis

Yeah. China, as we’ve discussed before, has gone all-in on humanoid robots. It’s a national priority. Companies like Unitree and others are racing to commercialize.

In the last report—and, Alex, we’ve talked about this—there were 150 humanoid robot companies in China under development. Part of their strategy is spectacle, and it’s something you’re trying to bring, Alex, to America. They’ve been staging public robot combat events, literally MMA-style.

We’ve got a video to show. Let me pull this up here. Here’s a recent MMA match that went viral on the internet, and it’s a beautiful thing.

Emad Mostaque

Just so cool. [screaming] These are only going to get better.

Peter Diamandis

You’ve got to watch the full video. The way the fight ends is epically awesome.

Yeah. One of the robots kicks the other robot’s head off. Remember Rock ’Em Sock ’Em Robots?

Emad Mostaque

Yeah, yeah. [laughter]

Peter Diamandis

As a game, as kids. This goes viral. There’s a lot going on in the robot world. We just saw all of the workers at Hyundai start to strike because they don’t want robots brought onto their assembly line. That was fascinating. Alex, take it from here.

Alex Wissner-Gross

A few thoughts on this. I have thoughts on many different levels. One is mild horror. If anyone’s seen Steven Spielberg’s movie A.I.—without spoiling it too much, I think Steven would call it the dark sandwich at the center of the movie, the Flesh Fair, where humanoid robots are tortured and abused for human entertainment—I think that’s utterly horrifying.

At one level, I’m mildly horrified that humanoid robots, no matter the extent to which they’re being teleoperated here, are setting an inductive prior or bias for future, more autonomous embodied intelligences to be basically trying to kill or otherwise physically abuse each other for human entertainment. I’m concerned about that.

But one level deeper, now imagine that these robots are more autonomous, that they’re running algorithms at the edge, so they’re much more encapsulated. Now imagine that these humanoids are in the Chinese PLA infantry.

Peter Diamandis

Yeah.

Alex Karp

I think that’s the future that we are almost certain to find ourselves in. The West needs to catch up in humanoids. That’s why I’ve supported ProRL, which, Peter, you were gesturing at. It ran its first humanoid robot mini-marathon in America in the Boston Seaport a number of months ago.

The West needs something like this—hopefully less violent and more economically productive. I’d love to see people cheering on humanoid robots competing to iron clothing or perform some economically productive task, and not just kicking each other’s heads off.

Peter Diamandis

You prefer the humans to be doing that in the MMA matches?

Alex Karp

I’d prefer no one to be doing it. I’m not a fan of MMA. I think it’s destructive to humans, and I worry about the message that we’re sending to the future light cone by having robots do it instead of humans.

I’d rather see people in a cage competing, if they must compete at all, to do something positive, not negative.

Peter Diamandis

Coding, like a cage-match coding competition—

Alex Karp

If anything.

Peter Diamandis

Or just sitting there. Okay.

Emad Mostaque

A couple of thoughts. My normal commentary around kickboxing is that it’s not the greatest marketing demo for humanoid robots. But I will acknowledge something here: this is an unbelievably demanding engineering environment.

You’ve got a stress test. It’s stressing balance, impact resistance, recovery, locomotion, and latency. There are 20 things that they’re doing. It’s kind of incredible to watch them navigate that. Of course, a 4-armed robot would beat a 2-armed robot. [laughter] I’ll just leave it at that. There we go.

This is competitive. We’re going to see this go to competitive sports. We’ll see a version of the World Cup with robotics. The question is whether people will watch that or not.

Peter Diamandis

Yeah. I’ll say that the real test is whether a human being can make that penalty shot under pressure at that point in the game. Watching England implode the other day was really devastating for me, but still, I think people would much rather watch people in that environment rather than robots.

Sports is going to thrive for many, many decades to come. Formula racing pushes the edge, and I think when we start to see robotic sports, it’s pushing the edge. I think the point you just made is important: we’re going to see this happening in a competitive fashion.

I can’t wait to see Figure versus Optimus. I think that will be a fun competition, whatever form it takes.

Emad Mostaque

Yeah, I think these robots are a little bit different, though. You’ll probably first see the Real Steel-type teleoperated robots, because robots can’t actually respond fast enough if you look at the latency of a VLA model.

Peter Diamandis

This is impressive with some teleoperated flying kicks, but why aren’t they doing kung fu? When will robots do kung fu? That’s when you move to things like edge silicon, when you move to teleoperation. I think that’ll be the next stage that comes next year.

But I think there’s a bigger issue that I have with this. Although I love fighting robots and I can’t wait to see Gundams and all that, these robots are EngineAI T800s. They weigh about 70 kg and they punch 4 times harder than Mike Tyson.

So they could legitimately kill someone—us fleshy humans.

Emad Mostaque

Robots like that should not be allowed on the streets, and there’s no regulation against that. They could be in the PLA—the People’s Liberation Army—or whatever, but robots are about to enter our households. Who here has a 1X robot on order? Come on—it’s coming. They will be walking around very soon.

We need to have regulations about the safety of these things, what the torque is on them, how they operate, and so on, because they represent a real threat to individuals. They are machinery. Beyond that, you have embodiment and other issues. We need to have the discussion of what that looks like when they are autonomous, because these things are delivering themselves by pushing a button on the door, ringing your doorbell.

The final thing is that Unitree has only made 11,000 humanoid robots in total. We are literally at the very start of this.

A few years from now, it will be 11 million a year instead of 11,000. We’ve got to have this discussion fast as well. Lots of talking to do.

Peter Diamandis

Yeah. This is the work you and I were doing in terms of how governments counsel their policymakers around these areas, and it’s happening at blinding speed.

Emad Mostaque

Crazy.

Peter Diamandis

Yeah. All right, I’m going to move us to the most important conversation we always have, which is the Dyson swarm. [laughter] Let’s take a look at a video from our friend Sam Altman.

“I honestly think the idea, with the current landscape, of putting data centers in space is ridiculous. It will make sense someday, but if you just do the very rough math of launch costs relative to the cost of power we can generate on Earth, to say nothing of how you’re going to fix a broken GPU in space—and they do break a lot still—unfortunately, we are not there yet.

“There will come a time. Space is great for a lot of things. Orbital data centers are not something that’s going to matter at scale this decade.”

All right, we have the continuing MMA battle between Elon and Sam.

Emad Mostaque

Yeah, fascinating. I’m curious about reactions here. Alex, I’ll go to you first.

Alex

I think there’s an obvious conflict of interest. We saw similar messaging from Masayoshi Son regarding the lack of purported promise for orbital data centers.

Remember, OpenAI has retreated from its own data centers. Remember Project Stargate? Project Stargate has been rebranded from OpenAI owning and operating its own data centers to just leasing terrestrial data center capacity from others. OpenAI is delaying its own IPO.

One has to look at OpenAI’s messaging here and say that perhaps it’s not even in a financial or operational position at the moment to lean into orbital data centers, the way Anthropic, in its collaboration agreement—which was announced with SpaceX’s xAI for use of Colossus and Colossus 2—is far likelier to move toward orbital data center–based compute.

I think the crossover is going to happen. Elon’s messaging regarding when this crossover is going to happen is 2 to 3 years. You see other analyses that suggest the unit economics for orbital versus terrestrial data center costs are going to cross over sometime by the early 2030s. I’m not sure which is the case, but either way, I think there is an obvious conflict of interest.

Just as we were discussing with Philip Johnston, barring some surprising left turn, I expect that OpenAI’s tune is very conveniently going to change on ODCs sometime in the next 2 to 3 years—right on time.

Peter Diamandis

And, of course, Elon’s response to this is, “We’ll be launching them in 2 years.” So just stay tuned and watch.

Emad Mostaque

Well, I think anyone listening to this video would say, “Okay, Sam says space data centers make no sense. Elon says they make sense. The 2 guys hate each other.” But if you actually listen closely to Sam’s words, they don’t disagree at all.

Sam is saying that space data centers will not be meaningful this decade. There will come a time, but this decade has only 3 and a half years left. If you look at Elon’s forecast of his launch rate, they actually agree. They’re just hating on each other all the time, and it seems that way in this phrasing, but the truth is pretty clear: They both have the same numbers.

So Alex is right. They’re going to space, and it’s going to take a while. I think a couple of percent of all compute will be in space by the end of the decade, because we’re building out on land as quickly as we can, too.

Peter Diamandis

But then the lines cross.

Emad Mostaque

Yeah.

Peter Diamandis

Yeah. You know, Alex, you and I were going back and forth texting while the Starship Flight 13 attempt was being made a couple of days ago, and it’s been rescheduled. When this podcast comes out, we’ll be seeing the next launch attempt of Starship Flight 13 on Monday of this coming week.

That launch was thwarted at T-minus 0.

Alex

First time I’ve ever seen that, by the way.

Peter Diamandis

Yeah. Here’s the point: 2 of the 33 Raptor engines on the booster stage of Starship did not ignite, and they’re going to be replaced.

By the way, SpaceX’s stock dropped 5% on news of that failed launch, which is kind of ridiculous. The point people need to realize is that was an amazing demonstration of technology. The fact that you could shut down at T-minus 0, safe the vehicle, and unload the methane and liquid oxygen—I was part of the space industry in the ’90s, before it was a space industry, and those vehicles would have exploded on the spot. They would have failed on the spot.

The ability we have to control them at that level of detail is evidence of the extraordinary engineering that SpaceX has done.

Alex

I thought that was the most interesting part: how quickly the system diagnoses the problem and returns. It would have taken months and months to do this, fix it, recover everything, and replan another launch. You’re like, “Yeah, problem. Shut it down, redo it. Oh, we’re starting Monday.” It’s amazing.

Peter Diamandis

Yeah, extraordinary.

Emad Mostaque

Yeah. Because I think if you’re serious about superintelligence, with what we know, you have to have a space play. OpenAI is going to buy Planet Labs or something like that, and then the tune will change.

Peter Diamandis

All right, I’m going to go to some AMA questions. Emad, you had suggested I post questions to X, and we have a number of questions coming about Kimi from our X audience. Let me go ahead and show these, and let’s dive in.

Emad, I’m going to give you first crack. Which of these questions do you want to answer?

Emad Mostaque

I think number 4 is probably an interesting one: Given Kimi K3’s lower token efficiency, is it actually as cost-effective as advertised compared with Solo Fable?

Kimi K3 is an expensive model relative to the other Chinese models. DeepSeek is now $1 per million tokens. Kimi K3 is $15. Sonnet is $20, Opus is $40, and I think Fable is $60. But that’s because they’re actually making money.

When you back out the numbers from the Chinese models and the chips they’re running on, they’re probably making 80% to 90% margins now. That’s with their Chinese chips, which aren’t that efficient for running this.

We will see the cost of K3 drop by 10 to 50 times, I think, in the next few months as it gets optimized. Right now, it uses twice the number of tokens for the same task versus GPT-5.6—a frontier model that uses 37% fewer tokens than 5.5 or Fable.

Again, we’re going to see that drop because everyone and their dog is going to optimize the crap out of this. You’ve seen Fireworks just raise at a $17 billion valuation. Others, like Modal at $10 billion and Baseten at $10 billion, are the inference providers of open-source models. They’ve all raised $1 billion that they’re now going to spend to optimize the Chinese model and make it more efficient and run it.

American labs that do the inference side of things are going to optimize the crap out of this. We will see it catch up.

Peter Diamandis

All right. By the way, I welcome the guests to lean in on these questions. Selene, you want to go next?

Salim Ismail

Given that I made the comment about number 1—how much could Kimi K3 devalue U.S. frontier models?—I’ll stick with my original estimate of about 75%: 50% from the U.S. regulating the front end, and then you’ve got a lack of compute on the supply side, plus frontier open-source models kind of within a release, barely, of where you are.

That bleeding edge is such a perishable thing. I would say a 75% drop. So if OpenAI is worth $1 trillion, I’d put it at $250 billion.

You still have a very valuable business, because now the competitiveness is about reliability, security, integrated tools, and ease of deployment. But the actual frontier cutting edge becomes one ingredient among the whole thing.

Peter Diamandis

I would not want to be inside these frontier labs right now. It must be a frenetic code-red, 24/7.

Emad Mostaque

It is a total rat race. I have so many friends at the frontier labs—friends who are jumping, hypothetically, from one frontier lab, Google, which is nowhere at this point, missing in action, to other frontier labs. It is a total rat race.

Peter Diamandis

Yeah, it’s crazy. Dave—

Dave

You have a choice for me.

Peter Diamandis

No, pick one. You’ve got 2 and 3, I think.

Dave

Okay, I’ll take 2: What does the release of Kimi K3 do to the open-source versus closed-source race? Will this force the large companies to provide more open-source products? I think they’re implying more open-source products.

Yeah, it’s a total game changer in the sense that anyone with resources can build an internal model that’s tailored to a specific use case and then use it as a defensive moat.

Dave Blundin

I don't think the large US model providers will go open source. I think they're committed to their pathway. So if you were talking to Anthropic right now, they would say, “Look, Kimi has caught up for a week, but Fable 5.1 is coming out in just a few weeks.”

When you look at the all-important enterprise use cases—white-collar automation, drug discovery—people are going to use the best model no matter what. If you're using an AI to design a car or a rocket, a slight improvement in the design has a massive payoff. So you're going to use the best of the best of the best model. The Anthropic guys are going to scramble to stay a step ahead and keep their price point nice and high. The cost of the model itself is so small compared to the benefit that people will pay the price.

Peter Diamandis

So it does create, like Alex was saying, the rat race is incredible, but people aren't going to switch to Kimi unless it's proprietary data they want to keep in-house and they want to tune their own, or Kimi actually bypasses Anthropic—which it hasn't done. It's only caught up, or not even quite caught up.

All right, Alex, number 3.

Alex Wissner-Gross

All right, number 3 asks—and I think these questions seem to all be variations on a theme—but it asks, “How can US models—I think this means US frontier model providers—continue to justify their massive valuations if China can leapfrog with an open-weight model at less than half the token cost?”

I don't think the premise is quite accurate. There are so many elements, so many layers to superintelligence, and quite frankly, superintelligence itself, as it fully develops, I think is far larger than the total GDP of the entire world anyway. There's an enormous amount of pie that can be sliced.

To the extent we're talking about, say, Google, which, as I was mentioning earlier, seems to be MIA at this point on the frontier—I can't find a single top Google model at this point on the cost frontier for capabilities—what does Google do? Well, they can continue to race, obviously, in terms of capabilities. But if I'm Google, I'm thinking, “Yeah, I want to become a hyperscaler.” I mean, Google obviously is a hyperscaler, but a hyperscaler provider to other frontier labs. That's one obvious venue of differentiation.

We've seen that approach vector from SpaceX AI itself, which has now signed deals with Anthropic. We're seeing it with Meta, interestingly, which, on the one hand, is offering Spark 1.1 and, on the other hand, in the past 2 days, just as we were going to air, it was announced that Meta is exploring selling $10 billion of compute to Anthropic. So differentiating by going down-stack and offering your compute to other, more competitive providers—whether Western, usually Anthropic, sometimes OpenAI, or Chinese models in a self-hosting model—that's one area.

You can also go up-stack. You can try to vertically integrate and offer applications that are benefiting from the commoditization of their complement, namely the model layer. I also think the premise that valuations somehow are going to net shrink just because Kimi K3 exists now is completely fallacious.

We saw that incorrect thinking happen with the original DeepSeek shock, which was at the time also branded as a Sputnik moment. We saw a bit of a hiccup in capital markets at the time, but, as always, Jevons paradox kicks in, and we see the value of chip stocks ultimately increase, not deflate.

We also see that it's open. It's open-weight, so there's absolutely nothing in Kimi K3 that OpenAI, Anthropic, and other Western frontier labs can't immediately reappropriate for their own internal models.

Peter Diamandis

You don't think that the amount of revenue these labs are going to make gets reduced as people start to use Kimi K3 for their work instead of API calls?

Alex Wissner-Gross

No. For example, I spend, and my portfolio companies spend, an extraordinary amount on, let's say, Anthropic and OpenAI. To my knowledge, my expectation is Moonshot would have to release a 2×, 3×, or 10× better model than, say, Fable 5 to have a massive diversion of that spend.

Right now, what K3 buys, to the extent it's legal—query how much longer K3 will be legal to host within the US—but assuming it remains legal and regulatory-uninhibited, all it results in is greater in-house self-hosting. But it's not at the top of the frontier. To Dave's earlier point, Fable 5 is ahead at the moment; if you're trying to solve the frontier of problems, K3 is not causing you to divert your spend.

Peter Diamandis

Well, let me hit that point you just made, Alex, and ask you and the other mates a question here. Do you think it's possible that some legal policy in the United States prevents US companies from downloading K3? It's going to be on the open internet. It's going to be available through a multitude of sources beyond Hugging Face. Can it be shut down in the US?

Alex Wissner-Gross

It can effectively be shut down. This is not prescriptive, and I'm not a fan of this policy, but I think it can effectively be shut down by requiring that every public corporation disclose any use of Chinese open-weight models and subjecting them to scrutiny.

As we were going to air, the latest—we talked in the last pod about Demis's proposal to create a FINRA-like entity that would regulate the frontier. Well, guess what? The reports are that the present administration is actually running with a proposal like that and is planning to, or at least exploring, creating a FINRA-like agency to regulate frontier AI that would live under the SEC, because the SEC already has statutory authority to operate FINRA-like, industry-advised-and-funded entities. So it's a natural place organizationally.

Peter Diamandis

Yeah. Self-regulated governance, a.k.a. regulatory-capture cartels under the SEC. I think it's completely plausible, albeit highly undesirable, that we get, sometime in the future, an SEC suborganization that looks like FINRA and basically makes it completely economically infeasible for corporations of any size, especially public corporations, to actively use Chinese open-weight models.

Any other comments on this?

Emad Mostaque

I've got a comment on this. I mean, this is ridiculous in terms of trying to limit the use here, because once you release the weights, you can mirror them across jurisdictions. You can use peer-to-peer networks and VPNs. All you're going to do is deny American researchers, startups, and security experts access to those models, while the rest of the world goes ahead building on those models. I don't think there's a viable approach. I mean, this is the same—

Peter Diamandis

Yeah, please.

Emad Mostaque

This is the same as denying Americans cheap insulin. I mean, it's again regulatory capture, right? Like, why can't you have generics? Because, again, you have the regulatory-capture point.

There's Operation Gold Eagle, I think they're calling it, to approve access to frontier models. You will have anti-token-laundering regulations. You will have know-your-prompter regulations. The US government has really realized that this technology is about to break through, and I think they're a lot more worried about it than China is.

You look at that Xi Jinping speech. I would urge everyone to check it out. They're full-on open source: “We're going to do this.” America doesn't know what it's going to do, but, as you said, there's a real chance that they might hobble American capitalism. Oddly, China's encouraging capitalism.

Peter Diamandis

CCP saves American capitalism from itself. That's a crazy future. The world is so weird.

All right. Let's go back to you, Salim, on the next question.

Salim Ismail

Which one?

Peter Diamandis

Some of these are a little bit duplicative.

Salim Ismail

Yeah. I'll take number 5. Would you trust Kimi K3 to write your code for you without oversight or review?

The answer is no, but I wouldn't trust a human being to put consequential, untested code into production either. The question is not whether we trust the models; it's whether we trust the development system around them. AI-generated code needs to be run in a sandbox, pass automated tests and security scanning, and go through all sorts of things before it goes into production.

Then you do proportionate permissions based on the use case and the potential impact. Whatever the workflow is that AI is running, you're still going to need human review at the highest level and for the highest-consequence inputs. A lot of the routine can be automated, but the scalable model is not AI with no oversight; it's machine-generated plus verification plus human accountability combined. That's going to give you the real power.

Peter Diamandis

All right. Iman.

Emad Mostaque

Yeah. What role, if any, did distillation play in K3 development? They distilled data clearly from Opus and others, but, to be honest, using Kimi K2.5 and Kimi K3 now quite intensely, it feels different.

I think they did a lot of their own data creation based in part on distillation, but everyone's distilling from each other right now. The one area where it's clear that they've had a big leap ahead is in front-end development. Again, this isn't the best mathematician in the world, although it's quite a good general model. It's not the best cyber attacker from our benchmarks, but they've done something original and new on the front-end, consumer-entertainment side of things, which I think is really interesting. Although that might also be because it's a multimodal model.

Peter Diamandis

Mhm. Mhm. Dave.

Dave Blundin

Number 7: What are the reasons why Kimi K3 might not be as good as advertised, or why we shouldn't use it?

The scenario where it's not as good as advertised is if it's benchmark-maxed and, in 2 weeks, the open source will be out.

Alex MacCaw

We'll have beaten it to death. We'll know the answer if they benchmark-maxed it, so we're going to find out. I think it's unlikely that it's benchmarked to the point where every company in America, every company in the world, should be saying right now, “We need a crash program with our best possible adviser to decide: Are we going to do our own model on our own on-prem hardware, or are we going to use Anthropic, OpenAI, or Google and just trust that API?”

But we need to decide whether tuning and training on our own proprietary data gives us a long-term competitive advantage. And so there's going to be a desperate shortage of good advice on this, and vendors and McKinsey consultants, and you've got to grab those resources quickly. ExO consultants, make seed-stage investments, get your network together, find out who can answer that question for you internally, on your business and your use case, quickly, and then commit to the path.

And you can do something internally and still use the APIs, but if you don't start down the path of evaluating Kimi K3 on your own, you can't really come back to it later. So I think everybody's got to just get going on this question. We'll know in a couple of weeks, though, whether it was benchmarked to hell or not. But I think it's very, very likely that the open-source path is a viable path for every U.S. and world company and government.

Peter Diamandis

Can I just add to that real quick?

Emad Mostaque

Yes, of course. Very simple suggestion for every company: implement 2 installations, Kimi K3 and Inkling. Fine-tune your own internal data, because that learning loop is going to be the proprietary gold that you don't want to lose. And start there.

Peter Diamandis

Alex, why don't you close us out here? You've sort of answered number 6 already, but perhaps you could expand on it.

Alex

I'll say something new. Question 6 asks, should the U.S. move to block loading the weights of the next Kimi release onto Hugging Face? I'll give a conditional answer. I think that if some party, presumably in the U.S., can prove to a competent court that the next Kimi release—presumably a reference to this Kimi release—was somehow obtained or derived illegally, maybe through copyright infringement or illegal distillation of traces or something like that, that would probably be grounds for blocking its release in the U.S.

But if no one can prove that Kimi's parent, Moonshot, did anything otherwise wrong in creating it, no, I don't think the U.S. should be blocking its release in the process. I think, if anything, quite the opposite. I think every U.S. frontier lab should be closely scrutinizing it and learning whatever they can so that we can leapfrog it.

And I would like to see far more outward pressure from U.S. labs creating the best-in-the-world open-weight and open-source models, so that it's not the CCP with their new Belt and Road for AI initiative blanketing the world—some would even say dumping superintelligence on the rest of the world, or the so-called Global South. It should be the U.S., the cannon of freedom, the arsenal of freedom, that's also the arsenal of superintelligence, showering the rest of the world with open-weight and open-source superintelligence—not China.

Peter Diamandis

Showering the rest of the world—I love that. And remember, we're moving toward intelligence that's too cheap to meter, but 1,000,000 times more available and more powerful than ever before. Everybody listening, I'm grateful on behalf of the Moonshot Mates here for your time. We're going to be putting this out more and more often as we're starting to see the release dates move from months and weeks to days. There's no time to sleep during the Singularity.

Gentlemen, what's in store for the week ahead? Emad, I'll go to you next.

Emad Mostaque

Yeah, just getting a whole bunch of research papers ready to release. So finally, it's going to be exciting.

Peter Diamandis

Again, acceleration for Intelligent Internet, your company?

Emad Mostaque

Yes.

Peter Diamandis

Incredible.

Salim Ismail

Tuesday, I have my next Meaning of Life session at 7:00 p.m. Eastern.

Peter Diamandis

Alex, are you coming up? What's going on with you?

Alex

I'm so focused at this point on literally solving everything. I'll say large swaths of the sciences at this point, I'm convinced, are so thoroughly cooked. More to come on that subject. Peter, you and I wrote “Solve Everything” about it, but now it's actually coming true.

Peter Diamandis

I'm excited. You're going to be doing an AMA with my Abundance community coming up. That's going to be a fun deep dive. And, of course, we're going to have you during the Moonshots gathering on September 25. In fact, all of us will be here. Emad, you're joining us in L.A. in September.

Emad Mostaque

Yeah, it's going to be fun to have all of us together again for the full day.

Peter Diamandis

Dave, this has got to be the most exciting time to be in Link Studios.

Dave Blundin

Oh, my God. Yeah. I think that discussion we had of quantization on this podcast that Emad kicked off—I think that now vaulted to my new best piece of media ever recorded, passing Leopold Aschenbrenner. I have to go back and listen to that again in slow motion.

Also, we had Vlad Bulović from MIT.nano in this week. He's going to advise and help us on our new startup working on photonic computing, and he gave us a whole roadmap of people I need to meet next week. So we're looking to add 2 MIT people to our Princeton team to work on just the photonics, quantized photonics side of the equation.

So I'll be working on that next week. But I think I can take that video we shot earlier and use it as a recruiting tool. It was just so freaking brilliant. You guys are incredible.

Peter Diamandis

I love you guys so much. What a great week. We'll see what breaks tomorrow. Over the weekend—emergency pod. We need emergency pods every day by January.

紧急更新——AI 的 Sputnik 时刻:Kimi K3 发布,与 Emad Mostaque 对谈|第272集 — 文字稿与摘要 | BidClub