AI 专家回应:Elon 的 Grok 4 如今登顶 AI——这将改变一切,与 Emad、Salim 和 Dave 对谈 #182
Peter Diamandis × Emad Mostaque × Salim Ismail × Dave Blundin
Grok 4 已将学术基准推近饱和,竞争焦点正从原始智能转向可用的自主执行能力。 它在 AIME 2025 上拿到 100%,Grok 4 Heavy 在 Humanity’s Last Exam 上达到 44.4%,高于 Gemini 2.5 的 26.9% 和 o3 的 21%。Emad Mostaque 的关键区分是:模型“在推理,但还不会规划”——规划、记忆和智能体整合将成为下一阶段的瓶颈。
xAI 的领先既是模型故事,也是基础设施与执行力故事。 公司成立仅 28 个月,就在 Elon Musk 实现运营 100,000 块 H100 这一看似不可能的承诺后,扩张到据称 340,000 块 GPU;嘉宾估算已部署硬件价值接近 100亿美元,并称 xAI 的目标是 100万块 GPU。Dave Blundin 回忆,专家曾说这种规模不可能保持系统一致性,随后只能感叹:“天哪,他真做到了。”
训练成本结构已经反转,围绕合成、结构化推理数据的新飞轮正在形成。 Mostaque 表示,后训练过去只占算力消耗约 1%,DeepSeek 时升至 10%,如今已大致与预训练持平,部分原因是前沿模型能够为下一代模型生成数据。但这并不保证 Grok 会“更理智”——模式坍塌仍可能发生;Peter Diamandis 则表示,模型竞争正从纯粹堆算力转向“工程与质量挑战”。
即使高端智能可能在高价值工作流中卖出更高价格,推理成本也正在快速商品化。 Grok 4 的报价为每100万输入token 3美元、每100万输出token 15美元;Mostaque 估算,“100万个非常好的词”成本约20美元,并预计同等智能的成本每年可下降 5–10 倍,最终可能降至每100万词 1美元。Diamandis 认为,能够压缩工程时间的边际能力提升,仍会让开发者支付显著更高的价格;Mostaque 则怀疑 300美元的 SuperGrok Heavy 套餐是为转化企业客户而设的亏损引流产品。
企业价值初期将来自增强能力、降低错误率和吸收数据,而不是立刻全面替代人类。 Emad 表示,Arc Institute 正在数百万份实验日志和 CRISPR 工作流上测试 Grok 4;Peter 引用了大致的医学研究结果:单独使用 AI 的表现超过医生,也超过医生加 AI 的组合。Mostaque 仍强调,受责任风险限制,全面替代“还远得很”;Salim Ismail 则强调,AI 能处理人类无法整合的扫描和传感器数据。
编码、游戏和视频暴露出同一个机会:生成能力先到位,规划、反馈和分发仍未解决。 xAI 展示了一款 4 小时生成的第一人称游戏,并称专用编码模型将在数周内推出;发布演讲者预计,第一款真正优秀的 AI 游戏和第一部值得观看的 AI 电影将在明年出现,而今年可能就会有半小时可看的 AI 电视内容。Mostaque 对在位企业的警告是:更低的制作成本对公司有利,但“对行业里的从业者来说,这很糟糕”。
如果前沿模型最终收敛到同一个能力平台,持久瓶颈就会转向接口、智能体、芯片获取和分发。 Mostaque 预计 Grok 5 将协调 60 到 6,000 个智能体,调用专业工具,表现得像一个“把事情办成、而且不睡觉”的远程员工。嘉宾讨论了 Google 约 300万块芯片、xAI 和 Meta 的百万芯片目标,以及整个行业的供应约束;Liquid AI 等高效边缘模型则可能提供日常智能层。
1. Grok 4 能以研究生水平推理,但仍不会规划
发布视频将 Grok 4 描述为在学术问题上“所有学科都超过博士水平,没有例外”,但同时提醒,它可能缺乏常识,也尚未发明新技术或发现新物理。发布演讲者认为发明能力可能会在今年晚些时候出现,并表示如果明年还没有出现,他会“感到震惊”。
Blundin 称这是类似钢铁侠 JARVIS 的“黄金时刻”:助手可以造出战甲,但人类仍必须决定“你要如何拯救世界”。它能解决极其困难的问题,却不会独立判断应该造什么,以及为什么要造。
Mostaque 进一步明确了这一限制:“我认为它在推理,但目前还不会规划。”他将下一轮模型规模与约 1 ronnaFLOP、即 10²⁷ FLOPs 联系起来,并表示模型能力应继续在算力与数据的共同推动下提升。
Diamandis 问道,在所有学科都具备研究生水平的能力,是否已经足以称为 AGI,并表示:“我们已经在毫无察觉的情况下通过了图灵测试。我们会不会也在毫无察觉的情况下通过 AGI?”Mostaque 回应称,社会会迅速适应强大工具,但完整的智能体能力仍需要更多基础模块。
2. 基准测试趋于饱和,前沿从回答问题转向发现问题
AIME 2025 给出了最清晰的信号:Grok 4 得分 100%。Diamandis 感叹:“你真的快把基准测试做完了。”但他同时指出,算力、数据和算法仍在进步,质量正越来越成为差异化因素。
Humanity’s Last Exam 包含 2,700 道题,覆盖领域极广;据估计,最优秀的人类在自己理解的领域内大约能拿到 5%,最高或许也只有 10%。Mostaque 举例说:“计算 Lie 群 G₂ 的分类空间的约化 12 维自旋配边。”这个例子说明,没有任何个人能与模型的知识广度匹敌。Blundin 还给出了一个来自五维引力理论的例子。
已发布的对比结果是:o3 得分 21%,Grok 4 为 25.4%,Gemini 2.5 为 26.9%,Grok 4 Heavy 为 44.4%。当被问及何时能达到 100% 时,Mostaque 回答“最多两年”,很可能就在明年。Diamandis 对测量问题感到不安:AI 可能很快就会提出并回答人类无法理解的问题,人类将无法判断它究竟在以多快的速度进步。
嘉宾对未来的最大猜想是新科学。Mostaque 提到,模型可能从数学推进到物理、化学和生物学,并称如果今年或最迟明年年底出现新物理学发现,他不会感到意外。Blundin 转述了 Alex Wissner-Gross 的猜想:解决量子隐形传态相关问题,或许能揭示宇宙中的其他智能;这只是推测,并非 Grok 4 已被证明具备的能力。
3. xAI 将极端规模转化为工程优势
Diamandis 提到 xAI 成立于 2023 年 3 月,并称公司在 28 个月内成为第一名模型。他回忆,Musk 曾承诺在当年夏末前部署 100,000 块 H100,当时听众都说“不可能”——但他最终做到了。
据称目前的集群包含 340,000 块 GPU,每块成本约 30,000 美元或更高,粗略计算价值约 100亿美元。Diamandis 还表示,xAI 计划在年底前部署 100万块 GPU。嘉宾将这一需求与每天约 10亿美元流入 AI,以及 Jensen Huang 预计到 2030 年行业年投资达 1万亿美元联系起来。
Blundin 描述了这场技术突围:专家认为如此庞大的集群不可能维持“幂律和一致性”。Musk 回到第一性原理,重新设计芯片连接,并让系统真正运行起来,最终引发行业感叹:“天哪,他真做到了。”
Mostaque 回忆,2022 年 Amazon 为他的团队搭建了一套 4,000 块 A100 的系统,当时是公开超算中速度第十快的系统;随着规模扩张,数百块芯片被烧毁。如今,硬件扩张和模型扩张都已经变成可以直接攻克的工程问题。
4. 后训练如今与预训练同等重要
Mostaque 将早期预训练描述为把互联网快照扔进“巨型超级计算机搅拌机”,产出的模型像“没喝咖啡、衣冠不整的研究生”。随后进行的强化学习清理只消耗约 1% 的算力;DeepSeek 将这一比例推高到 10%,而 Grok 4 的发布演示称,后训练投入如今已与预训练持平。
其运作机制是数据飞轮:前沿模型生成结构化推理轨迹和训练材料,供下一代模型使用,从而降低对无差别抓取互联网数据的依赖。Peter 的更广泛结论是,优秀模型之间的差异如今更多来自更好的数据、算法和工程质量,而不是简单地“把一切都扔进锅里”。
当被问及这是否会让 Grok 更理智时,Mostaque 诚实地回答:“但愿如此。”更多后训练并不会消除模式坍塌或潜空间故障,只是让模型获得了比原始 Reddit 数据和互联网规模数据更结构化的课程。
已公布的价格是每100万输入token 3美元、每100万输出token 15美元,并标注上下文长度为 56,000 token;API 讨论中则另行提到 256K 上下文。Blundin 提醒,许多宣传中的上下文窗口并不能完全使用,但如果大窗口确实可用,模型就能在一次处理中并行处理约 100 本书的信息。
5. 廉价智能与高价订阅可以共存
Mostaque 将 Grok 4 的成本与 Claude 4 Sonnet 和 o3 作比较,同时称 Grok 4 的表现优于两者。按每个token约 0.7 个词计算,他估算“100万个聪明的优质词”成本约为 20美元。
仅依靠 Vera Rubin 硬件,他预计明年推理成本就会下降 3–4 倍;再叠加算法进步,同等智能的成本每年可能下降 5–10 倍。他给出的终点是:“100万个惊艳的词,只要 1美元。”
Diamandis 认为,SuperGrok Heavy 每月 300美元仍可能对代码、机械设计及其他关键工作具有吸引力。如果边际能力提升能够节省昂贵的工程时间,或把错误答案变成正确答案,买家就可能接受再高 10 倍的价格,即使竞争产品正越来越商品化。
Mostaque 猜测,高级套餐正在亏钱,类似 OpenAI 此前对其 Pro 档位的描述。他将其视为面向企业的亏损引流产品:先解决“界面问题”,通过 Andrej Karpathy 所称的上下文工程把团队数据接入模型,再向企业销售;对企业而言,每位高级知识工作者每月 300美元微不足道。
6. 企业采用从研究和医疗增强开始
Mostaque 表示,Arc Institute 这家同属投资组合、也是领先的生物医学研究中心,已经在研究工作流中测试 Grok 4。模型可以筛过数百万份实验日志,并在“瞬间”选出有前景的假设,包括 CRISPR 研究。他补充说,xAI 的企业业务仅启动了 2 个月,Grok 将通过超大规模云厂商提供。
Mostaque 预计,医疗领域会先走“增强优先”路线:降低错误率、改善结果,最终才会替代专业人士。Diamandis 认为监管将是主要障碍,而 Mostaque 强调,全面替代所涉及的责任风险“还远得很”。
嘉宾引用了一项 Google 医疗 AI 研究的大致结果。Diamandis 描述称,医生单独诊断的准确率约为 80%,医生加 AI 的“半人马”组合接近 87%,AI 单独使用则在 90% 出头。Mostaque 记得医生单独使用的数字是 70%,并提到 Daniel Kraft 估计医生约 30% 的诊断是错误的。这些数字只是大致呈现,并非一个已经确定的单一结果。
Mostaque 强调,人类偏见也会影响输出。Ismail 更广泛的观点是,问题不只是击败一个医生:现有扫描和传感器已经产生了超出个人吸收能力的信息,而 AI 可以纳入那些“过去根本不可能进入诊断”的数据。
7. 生成式娱乐扩大供给,但注意力和分发仍然稀缺
发布视频展示了一款据称在 4 小时内完成的第一人称射击游戏。发布演讲者强调,核心游戏逻辑并不是唯一难点:获取纹理、文件和其他资产,过去一直是制作具有视觉说服力的作品时的约束。
发布演讲者预计,AI 将生成艺术内容,把它映射到 3D 模型上,并通过 Unreal 或 Unity 等引擎生成可执行文件,可能就在今年,最迟明年。时间判断是:明年出现一款真正优秀的 AI 电子游戏,今年出现半小时可看的电视内容,明年出现一部值得观看的 AI 电影。
Diamandis 预计娱乐内容会出现激进碎片化:游戏可能每 4 小时迭代一次,朋友们也许会观看同一部电影的个性化版本,并拥有不同结局。他预计互动媒体会比被动观看增长更快,包括能够直接回应观众的角色和声音。
Mostaque 反驳称,人类的注意力不会增长。他引用的数字是:游戏市场约 4,500亿美元,电影市场约 700亿美元;电子游戏市场已从约 1,700亿美元增长到 5,000亿美元,Metacritic 平均分则从 69% 升至 74%;电影行业增长幅度小得多,IMDb 平均评分约为 6.3。
Mostaque 认为,大众共享的头部故事仍会存在,因为“分发、分发、还是分发”依旧决定触达范围。更低的成本对公司有利,也能帮助个人创作者讲出更丰富的故事,但“对行业里的从业者来说,这很糟糕”。
8. 编码的下一层抽象是上下文,而不是更多手写代码
发布演讲者表示,xAI 最近训练了一个兼顾速度与智能的专用编码模型,预计将在“数周内”发布。Mostaque 称现有 Grok 模型已经能写出整洁代码,并预计专用模型会进一步提升。
Diamandis 重新提到 Mostaque 早先“5 年内不再需要程序员”的预测,这一说法曾招致印度用户的辱骂邮件。随后他问,顶尖程序员是否只是会因此写出多 100 倍的代码。Mostaque 回答称,下一种角色将是负责指挥系统的“真正优秀的上下文工程师”。
Mostaque 的逻辑是,代码之所以成为中间语言,是因为编译器无法理解人类意图的复杂性。他提到 Cursor 在一年内实现 5亿美元营收,Anthropic 的营收约为 40亿美元,其中可能约三分之二与代码相关。
剩下的弱点是编排。Diamandis 表示,模型已经能生成模块和仪表盘,有时甚至会提前想到他尚未考虑的需求,但创意项目仍需要规划、协调、多智能体系统,以及更紧密的 UI 反馈闭环。
9. Grok 5 的竞争将围绕智能体、世界模型和稀缺算力展开
发布演讲者表示,xAI 将在 3–4 周内开始用超过 100,000 块 GB200 训练视频模型。Mostaque 将其与团队首次训练的最先进视频模型作对比:当时使用了 700 块 H100;如今领先的视频项目通常使用约 2,000–4,000 块芯片。
Mostaque 称视频模型是“世界模型”:学习视觉变化的同时,也会学到物理规律的表征,从而支持 3D 资产、模拟世界和自动驾驶训练数据。他预计 xAI 的系统初期会是独立模型,但语言、图像和视频系统最终可能会合并为一个模型。
他设想的 Grok 5 是一个多智能体系统,拥有“60、600 或 6,000 个”工作者、一个世界模型、系统连接能力,以及调用 Maya、物理模拟器和 Lean 的能力。规模足够大后,它就会变成一个“极其多才多艺的工作人员”;人类与 Grok 5 或 Grok 6 的交互,可能看起来就像一次 Zoom 通话。
Mostaque 预计 Gemini 3、GPT-5 和 Grok 4 级别的系统将处于大致相同的智能平台上。差异化因素将变成一个类似远程同事的接口:给它发消息、交代任务,在它不确定时收到进度确认,并让它运行超过目前所称的 7 小时任务时长。他偏好的 AGI 定义是“真正有用的智能”(actually useful intelligence),能够“把事情办成,而且不睡觉”(just gets the job done and it doesn’t sleep)。
资本并不是首要约束。嘉宾提到 Google 约 300万块芯片、xAI 和 Meta 的百万芯片目标、OpenAI 的资本以及 Stargate 规模的基础设施,还有 Amazon Trainium 对 Anthropic 的支持。下一次跃迁将是 1,000万块芯片,而全世界据估计只有 2,000万块;因此获取、封装和供应能力将成为决定性因素。
Diamandis 注意到,Apple 尚未进入讨论。Blundin 表示,Apple 控制着 TSMC 约三分之一、用于其 M2 和 M3 产品线的制造产能,可能成为数据中心领域的重要玩家。Mostaque 认为,模型一旦足够好,就会开始像公用事业,因此高效边缘模型的重要性将持续上升。
Nvidia 仍是默认选择——“买 Nvidia 不会被解雇”(you don’t get fired getting Nvidia)——但买家会在任何能买到的地方采购有能力的芯片,因为虚拟工作人员的成本远低于人类团队。Diamandis 表示,Liquid AI 的边缘模型可以运行在 M3 级芯片和汽车硬件上,据称效率约为暴力运行 Transformer 的 100 倍;它们可能负责提供日常智能,而稀缺的前沿系统则处理“天才级”工作。
How impressive is Grok 4 for you?
If you look at the AIME benchmark, which is an advanced math quiz, Grok 4 scored 100% on it.
You’re literally running out of benchmarks.
It’s got to be driving Google nuts that Elon got this done in 28 months from a cold start.
When he said he was going to put this huge cluster together, every AI expert in the world said, “You cannot get power laws and coherence at that scale. You just can’t do it.” Every AI expert was like, “Oh, God dang, he did it.”
The amount of compute and resources, again, are going exponential. Now, it’s the real quality that differentiates the top models from each other.
My big question is, where do we go from here?
Now, that’s a moonshot, ladies and gentlemen. Everybody, welcome to Moonshots. An episode of WTF just happened in tech this week. Special episode today following the release of Grok 4. It’s large language model release month, with an extraordinary string of new models coming up. I’m here with my Moonshots mates, Dave Blundin, the head of Link Ventures; Salim Ismail, the CEO of OpenExO; and a special guest to help us dissect all of this, Emad Mostaque, the founder of Intelligent Internet.
Guys, it was a pretty epic day yesterday. Good to see you all. Pleasure to have you.
Yeah, likewise.
This is our special Grok 4 edition. Emad, you’re in London, yes?
Yep.
Fantastic. Salim, where on the planet are you, buddy?
New York.
Okay. Dave’s in Boston. I’m in Santa Monica. All right, let’s get going.
So, just to jump in. The goal here is to dissect what happened yesterday blow-by-blow, what’s Grok 4 all about, and just to shadow what’s coming. We’ve got a few new model releases coming with Gemini 3, GPT-5, and probably a few others, so let’s kick it off with this video.
Grok 4 is postgraduate, like PhD level in everything—better than PhD, but most PhDs would fail. So it’s better said—I mean, at least with respect to academic questions—Grok 4 is better than PhD level in every subject, no exceptions.
Now, this doesn’t mean that it may lack common sense, and it has not yet invented new technologies or discovered new physics, but that is just a matter of time.
Mhm.
I think it may discover new technologies as early as later this year. I would be shocked if it had not done so next year.
All right, Dave, you want to take the first bite?
Yeah, it’s awesome. This is actually a golden moment in time because it is an absolutely brilliant assistant that can do almost anything you want it to do. But, like Elon said, it’s not reasoning yet, so it’s not coming up with the fundamental answer of what we should build and why. That’s still in the hands of the creator—the human operator.
This moment in time is really, really golden. It feels just like an Iron Man movie where you’ve got JARVIS. JARVIS will build the suit for you; you have to decide how you’re going to save the world. It’s a really, really fun time to be using these brand-new models. Like you said, there’ll be 3 of them in the next month or so. This is the first round, and he’s dead right: the PhD-level solution is all measured in the benchmarks. We’ll get into those in a minute.
It does virtually anything. The capabilities are mind-blowing, but it doesn’t decide what to do and why.
I would love your take on this, Emad. You’ve been plugged into this world intimately for a while. How impressive is Grok 4 for you?
I think it is very impressive. Picking up what Dave said, I think it is reasoning, but it’s not planning as yet. When we got to this ronnaflop level—I think that’s the term, around 10²⁷ FLOPs—would we continue to see improvements? Part of that is the compute, and part of that is the data, as we’ll get to later. The answer is yes.
Again, like Elon said, getting above postgraduate level in every subject means it can now execute and it can reason. It doesn’t have planning yet.
So, I mean, isn’t that AGI? Isn’t that sort of the definition of AGI? We passed through the Turing test without noticing. Are we going to pass through AGI without noticing, too?
It’s like hedonic adaptation. You’re like, “Of course, it’s fine.” But already, if you want to get a job done, it will do the job for you—summarizing a book, writing a summary of something, translating, and so on. Life is just the same so far because you haven’t got that final step that Dave mentioned.
There are a few extra bits that we need for full agentic capability above that, but we’re nearly there because we have that final building block now with this next level of model.
Yeah.
Where it’s reliable—
The distinction, by the way, is that it is reasoning. It has to be to solve these really hard PhD-level problems, but it’s not planning. A ronnaflop is 10²⁷ FLOPs, so that’s the scale of these algorithms.
That was the level the AI Act said they wanted to ban, by the way. So this would be the first banned model.
Yeah, that’s a great point—the first.
I think one of the things that’s happening is the absolute beauty of capitalism, where you’ve got big juggernaut companies fighting it out for supremacy, taking massive risks, choosing design paths, taking huge gambles, and really, really going for it. I think it’s magical to watch this happening.
I love this tweet from Sawyer Merritt. It says, “xAI was founded in March of 2023. Just 28 months later, it’s now the number-one model in the world, verified by independent testing.” Incredible achievement.
It is insanely fast compared to everything else that’s being built. I remember in May, 2 years ago, when Elon was first raising money. I had a chance to sit in on an investor pitch in the first round for xAI, and he said, “I’m going to have 100,000 H100 GPUs operating by the end of the summer.” Everybody was like, “No freaking way.” And he did just that. He hasn’t slowed down.
Here we see in the Artificial Analysis Intelligence Index that Grok 3 was placing fifth or sixth, while Grok 4 leaps to the front of the line. Emad, are we going to continue seeing this leapfrogging each other? Is there no end in sight?
It’s getting very difficult because, if you look at the AIME benchmark, which is an advanced math quiz, Grok 4 scored 100% on it.
You’re literally running out of benchmarks in order to do that. The amount of compute and resources, again, are going exponential because you need to squeeze that out, as well as have good data and good algorithms.
Before, you could just chuck everything into a pot and slush it around. Now, it’s the real quality that differentiates the top models from each other. It’s become more of an engineering and quality challenge than just a brute-force challenge.
Insane.
Can I pause you for a second?
Please?
Okay, so I’ve got a problem. I would suggest that if I’m trying to answer that problem or get a solution to it, I could go to any of these, and they’re going to give me roughly the same answer. Yes.
So, we’re at a point where the new step is—I want to get into the details of Grok to figure out why it is so radically different from any of the others. That’s where I think the fun will come.
The funny thing is, we’re basically going to Einstein and asking him to summarize a poem for us. There’s such a massive level of intelligence, and the utilization for the general public is minuscule.
All right, let’s look at what’s next. Grok outperforms the highest-level test, Humanity’s Last Exam. Up until now, o3 was at 21%, Grok 4 was at 25.4%, and Gemini 2.5 was at 26.9%. Then Grok 4 Heavy comes in at 44.4%.
We were talking about this a little bit earlier. Can you speak to Humanity’s Last Exam for us?
Yeah, this was developed by Scale AI and a few others to have an exam that even the most polymathic people in the world would find difficult. They estimated that some of the smartest people in the world would score maybe 5% on it, with a maximum of 10%.
The top models at the time, probably 6 or 9 months ago, scored 8%. Now you have a qualitative leap above that, to the 44% level. I think it’s interesting because, as Salim was referring to, what are these models for? They’re at this super-genius level.
It’s like having a mega liberal arts program. The next step is going to be to have really useful people in the workforce on one stream, and then the other stream will be to take the subcomponents of this and push them up to superhuman reasoning, discovering new things at a level that we could never have had before. I think this is one of the indications of that because, again, I tried to read some of the questions. I didn’t even understand the questions.
Examples. I literally just gave a presentation on this yesterday, so I have it right in front of me.
Humanity’s Last Exam: 2,700 questions. When the slide says, “For reference, humans can score 5%,” that means the very best humans in any given domain can score 5% within just the domain they understand. And I’ll tell you why. Here’s an example question: “Compute the reduced 12th-dimensional spin bordism of the classifying space of the Lie group G₂.” And then it goes on from there. Most people can’t even understand one word of that.
Exactly. Here’s another one: “Take a 5-dimensional gravitational theory compactified on a circle down to a 4-dimensional vacuum.”
So, yeah, these are the hardest questions, and that’s why this exam is supposed to last for a long time. A 44% score is just way outside the range of human ability because nobody has that broad knowledge spanning all these topics.
So, how far—how long before we hit 100% here, too? Any bets?
2 years max, I would say. Probably next year.
You know, there was a conversation years ago about AI getting to a point where you can’t understand the questions it’s asking and answering. We’re not far from that. At some point, we’re unable to measure how rapidly it’s advancing. That becomes a little bit frightening.
It’s got to be driving Google nuts that Elon got this done in 28 months from a cold start.
Absolutely, largely because Elon is phenomenal at large-scale manufacturing, large-scale organizational management, and people working at 4:00 or 5:00 a.m., sleeping in tents on the factory floor. That’s his wheelhouse. That’s Tesla, that’s SpaceX, and because all the intellectual property was more or less open-sourced by the research community at Google and Meta, he was able to pick up all that brilliant thinking and just plow it into implementation. It’s also small teams, right? It’s not large. I mean, Google’s a massive organization.
Yeah.
I think there’s something else here, though. Remember, we talked about this last time when Grok 3 came out, right? When he said he was going to put this huge cluster together, every AI expert in the world said, “You cannot get power laws and coherence at that scale. You just can’t do it.” And he went right back to first principles, created new kinds of connections between the chips and whatever, and did it. Every AI expert was like, “Oh, goddang, he did it.”
This is the incredible ability he has to go into a domain with a beginner’s mind, go to first principles, and just reengineer the heck out of it to achieve massive performance. I think this is an indication of that. My big question is, as you mentioned earlier, Dave, where do we go from here? What does it mean to have 50% versus 44% on this test?
Yeah.
I think if I can just give it a little bit of context about 2022—
Amazon built us the 10th-fastest public supercomputer in the world: 4,000 A100s.
2022.
And that was 2022. That was the 10th-fastest in the world.
Wow.
Of any supercomputer that we were training on. And there was an instance where literally hundreds of the chips melted because of the scaling. Now they’ve managed, by turning this into an engineering problem, to scale the hardware but also the inside of the model, which I think is really important.
The reason it’s above PhD level in each of these areas is that it was a computation-scale problem.
And so what happens is that if you could scale a liberal arts person all the way up to postgraduate level in everything, you would. Then you specialize down, and you look at some of these things.
And Salim’s question there—
Just for reference, everybody, the xAI cluster now has 340,000 GPUs.
About $30,000 or more each.
Yeah, do the math.
$10 billion.
A lot. I mean, this is why we’re seeing $1 billion a day going into AI and why Jensen said there’ll be $1 trillion a year by 2030. And it’s not slowing down.
So, here’s another image from the little conversation Elon had yesterday. These are the benchmarks his team put up. I don’t know if you want to hit on any of these, Emad, Dave, or Salim. Any favorites for you?
My favorite one is AIME 2025 at 100%. You’re done. GPQA—these are all hard benchmarks.
I think Elon would want to go to 110%. He likes 11 as 11.
But the only one I don’t recognize is on the bottom right. Emad, do you know what that is? USAMO 2025?
I think it’s the U.S. Mathematical Olympiad.
Oh, right. Of course.
So, it’s about to happen. Again, these are novel, hard benchmarks—effectively all of them—and they’re being saturated because, ultimately, AI can reason about mathematics and science better than we can. Again, it can’t plan just yet. It doesn’t have the same memory capacity, and the building blocks haven’t been put together. But it’s already superhuman narrow capability in many narrow areas. So it’s inevitable, I think, what happens next.
You know, we glossed over his quote there: “Discover new physics.” It wouldn’t surprise me if it’s this year, certainly no later than the end of next year. Alex Wissner-Gross has been having a field day with that all day.
I bet.
First of all, what does it mean to discover new physics? That’s pretty interesting by itself.
Well, I mean, Alex has been saying we’re going to solve all of math, and then physics comes next. Chemistry and biology follow quickly. I mean, this is the most exciting thing for me. Will these models literally unwrap the principles of the universe before us, right here, right now, during our lives in the next 5 or 10 years?
Yeah. Alex has been saying that.
Well, the one that’s really fun to think about is the quantum teleportation, Peter, that you brought up at one of our enterprise meetings.
So, how do you reconcile the fact that 2 entangled particles can be infinitely far apart—
Yet still communicating in real time—
With the fact that the speed of light can’t be transcended? Alex’s speculation is that if we can solve physics in the next 1, 2, or 3 years, and it turns out that you can communicate using quantum teleportation, we instantly discover all these other intelligences around the universe.
Yeah, we’ve just been listening at the wrong frequency with the wrong codecs.
These are the key takeaways. I’m going to just read these out loud, and we can talk about them. They spent just as much on fine-tuning and training the AI after the initial phase as they did on pretraining. So that’s a big change. Emad, you want to dissect that for us?
Yeah. It used to be that everything was basically: you take a snapshot of the internet, and then you put it into this giant supercomputer mixer, and it figures out all the connections—the latent spaces—to guess the next word. Then you had this very weird AI that came out that was a little bit crazy. It’s like a disheveled graduate student without his coffee. Then you had to tidy him up with reinforcement learning. That was the post-training, and that was 1% of the compute.
Then, with DeepSeek, it was 10% of the compute. Now it’s moved to equal because they figured out how to chain reasoning traces. In fact, I think part of what they did—we’ve seen this with other labs—is they used their frontier model to make data for the next frontier model.
Having large amounts of compute to create your own training data in a structured manner allows you to take that latent-space landscape and make it smarter and smarter and smarter, just like your brain adapts as you learn more reasoning and see more and more things. Rather than having to have these massive scrapes of the internet or whatever, it’s more and more structured data making up these models, which are making them smarter reasoners.
So, the 50% additional compute dedicated to the fine-tuning—does that mean we have a more sane version of Grok?
Fingers crossed. It doesn’t necessarily mean that, because you can still get all sorts of mode collapse within it if the latent space goes awry. But probably, because again, you’re training it just on a certain field of things as opposed to Reddit and other things.
In terms of order, I'd say this is probably like 100 million each. So it probably adds up to 1 Meta AI researcher.
A new unit of measure in the AI world. That's funny. So, let's comment on the cost here: $3 per million tokens, $15 per million output tokens, and long context windows of 56,000 tokens. How does that measure up, Dave, in your mind?
Well, it's pretty normal these days. It's a longer context. A lot of the claimed context windows aren't real. Under the covers, the dimension of the neural net is much smaller than the claimed context window, so I suspect, at this scale, that this is the true dimension of the network, but I don't really know. We'll have to dig in over the next couple of days and find out.
What it means is that you can feed in 100 books' worth of information concurrently. It instantly digests all that knowledge and then gives you an intelligent answer based on all of that information in 1 pass. So, it's just the next step in what's been going up sequentially from model to model to model.
Emad, do you expect we're going to be constantly reducing the price per token? Is this a demonetizing curve for a while to come?
100%. The cost of this is about the same as the cost of Claude 4 Sonnet, Anthropic's second model, or o3, but it's better than both. It's about 0.7 words per token, to give you an idea. So, the cost of 1 million very good words that are smart is $20.
Mhm.
But next year, with Vera Rubin, the next-generation chip they're going to put in there, just from the hardware it'll be 3 to 4 times cheaper, and they'll probably figure out some more stuff around that. So, for equivalent intelligence, the cost probably drops by around 5 to 10 times a year. So it'll be a buck for 1 million amazing words. It's hard to believe the most powerful technology in the world is diminishing in cost. It's crazy.
I want to put a comparator here, though. This is amazing: We could put hundreds of our books into the thing, and it would hold all of that in real time, as Dave said. But let's note that a single human cell has several billion operations going on in it at any point in time, right? So we're multiple orders of magnitude from modeling 1 cell. We've got a long way to go to try and model life or get to really big, big, big things.
Let's talk about SuperGrok Heavy. I've got to love Elon's terminology, right? We've got Falcon Heavy, and now we've got SuperGrok Heavy. He loves his terms, and I love them, too, actually. It made me smile when I saw that.
Why heavy, by the way? Is there a reason for that?
Falcons. The Elonverse.
Yeah. No, I mean, Falcon Heavy was able to have 3 boosters to launch a heavier payload to orbit. So why not talk about heavier capacity? In reality, Falcon Heavy had multiple boosters, and this has multiple agents.
The next one will be Grok Starship.
It'll be BFG. BFG. Yes.
The price point here sets a new high bar. That's going to scare a lot of people. I say the same thing I said last time: Try it. Burn the $300 for 1 month. You can turn off the subscription, but you've got to try it to know what you're missing or not missing. A lot of the use cases, the day-to-day use cases, won't matter much. But if you're building something complicated, writing code, or designing mechanical parts or whatever, you're going to get addicted to it.
What I'm really curious about is the margin at $300 a month. Are they actually chewing up all that money on compute for you, or do they have significant margin at that price point? One thing I've been predicting for a long time, and it's inevitably going to happen soon, is that there will be use cases where you need that extra intelligence. When you're building a software product and you're prompting it, you absolutely need that extra level of intelligence. It makes you dramatically more efficient in moving forward. If you look at the cost of a software engineer's time, you can afford to go up another factor of 10 or even more in price point for this and still be glad that you paid it.
I think the escalation of pricing is going to come soon. The counterargument is that the competing models will then commoditize it. But I think people will pay a lot for marginally better improvement because of the effective product you get out the other side. It really accelerates your time to development, the quality of the design, or whether the solution to the math problem is right rather than wrong. It makes a big difference.
My guess is they're losing money.
You think so?
That's what OpenAI said for their Pro level, whereas the level below they make money. So I think the way I view this is as a loss leader, because if someone's paying $300, you enterprise-upsell them.
Mhm.
And then you do team things to get everyone doing it, because basically, right now, what we have is a UI problem. The reasoner is there. The way to hook it up and make it usable for as many people on your team isn't there. This is what Andrej Karpathy calls context engineering. What are the new UIs that will enable us to use this most efficiently and get our data in there? If you can crack that, then $300 a month for a high-level knowledge worker is nothing.
Yeah, you know, zero, right? Just like we used to pay $1,000 or $2,000 a month for Bloomberg when I was a hedge fund manager, mostly for instant messaging. Again, it's just not quite there, but it's about to flip there.
Yeah. Well, a lawyer will cost you that much per hour, or 3 to 5 times that per hour. Will this do the job of your legal document better? I can't wait. That's the 1 profession I would love to replace: lawyers. All right. You mentioned enterprise-level, Emad. Let's go there right now.
What else can Grok 4 do? We're actually releasing Grok 4, if you want to try it right now, to evaluate—run the same benchmark as us. It's on the API and has a 256K context length. We already see some of the early adopters trying Grok 4 through the API.
Our portfolio neighbor, the Arc Institute, which is a leading biomedical research center, is already using it to see how they can automate their research flows with Grok. It turns out it is able to help the scientists sift through millions of experiment logs and then pick the best hypothesis within a split second. We see this being used for CRISPR research, and Grok 4 independently evaluated as the best model for examining chest X-rays. Who would know?
In the financial sector, we also see Grok 4, with access to all the tools and real-time information, as one of the most popular AIs out there. Grok is also going to be available on the hyperscalers, so the xAI enterprise sector has only started 2 months ago, and we're open for business.
Open for business. So, Emad, you've been working on medical-related AI. The block here isn't the tech; it's going to be the regulations. When will an AI be able to fully replace a radiologist or fully replace any profession in the medical world? How do you think about that?
Well, I think it's augmentation first: reduce errors, increase outcomes, and then eventually it's replacement, because Google had its AI medical expert study, which compared a doctor, a doctor plus Google Search, a doctor plus AI, and then AI by itself.
Yeah. I want to touch on that because it was a really important article that came out. If you look again, the physician by themselves was getting something like 80% of the cases correct. The centaur—the physician plus the AI—was getting around 87%; the numbers are approximate. Then the AI, without the human biasing the output, was outperforming all of them at around the low 90s. Extraordinary.
Well, again, it's what you said: It's better than any postgrad at the moment. But right now, I think it's about empowering and accelerating in terms of integration, and we're way off the liability profile for replacement. I don't think you need replacement right now. What we need is fewer errors in something like medicine, right?
I think the doctor number by itself, Peter, was 70%, because I remember Daniel Kraft saying when you go to the doctor, you get the wrong diagnosis about 30% of the time, right? That's a staggering number of errors, by the way.
That means out of 4 of us, 1.5 got the wrong diagnosis the last time we went to the doctor. We need to figure out who that was. That's really ridiculous. You need an AI to take over that whole field.
Well, human bias—and getting human bias out of that—is also even more important.
The number of types of scans and sensors you can do is way, way outstripping any human ability to look at all the data that comes out of it. A lot of it isn't trying to beat a doctor. It's trying to assimilate data that never could have gotten into the diagnosis before.
That's a great point. That's a great point. All right, let's go on to the next one: available for an API. We've covered these areas already. Let's move on.
All right, I love this. You know, Elon is a gamer, and so it's not unreasonable for him to be talking about using Grok to make games. Take a listen.
The other thing we talked a lot about was having Grok make video games. Denny is actually a video game designer on X. We mentioned, “Hey, who wants to try out some Grok 4 preview APIs to make games?” And then he answered the call. This was actually just a first-person-shooter game made in a span of 4 hours.
One of the most underappreciated hard problems of making video games is not necessarily encoding the core logic of the game, but actually sourcing all the assets, all the textures, the files, and creating a visually appealing game.
I think one of the challenges is what we do with all of our time in the future, and we may be playing a lot of video games. This could actually light up the entire metaverse, because building the metaverse world and building those environments was the big limiting factor, and now you can do it at a very rich level. This could be really interesting to see what comes from this.
Yeah.
When did you guys first hear that Grok 4 was going to come out last night?
Well, he said a few days ago, didn't he?
A week ago. He was saying it was going to be this weekend, and then it got pushed to yesterday.
Yeah, because I feel like we had about 48-hour notice, plus or minus a day or two.
But it was amazing. If you look at the presentation from last night and compare it to Google I/O, Google I/O was scripted and staged with multiple presenters and clearly planned way in advance.
Last night was like, “Is it done yet, guys? Is it done? Does it work? Okay, if it works, we're launching tonight. Let's go. Get on stage. Let's go.” I think that's the way it's going to be in the future because it seems like getting to market 1 day or 2 days sooner actually matters a lot in this horse race. This is kind of the dynamic we should expect going forward.
By the way, that narrator is the AI voice of a geek who is living and breathing it.
That's what you want. All right, let's take a listen to Elon discussing video games and movie production, for example.
For video games, you'd want to use Unreal Engine or Unity, or one of the main graphics engines, and then generate the art, apply it to a 3D model, and create an executable that someone can run on a PC, a console, or a phone. We expect that to happen probably this year. If not this year, certainly next year.
It's going to be wild. I would expect the first really good AI video game to be next year, probably the first half-hour of watchable TV this year, and probably the first watchable AI movie next year.
The fragmentation of those industries is going to be incredible because normally we think of a video game coming out in a release, and all of your friends get the exact same release. It's a release that's maybe good for a year or more, and you're all on FIFA 23 now, or whatever, FIFA 25.
But here, because it's only 4 hours to create the next iteration, you can say, “No, I want a customized version.” There's going to be all this fragmentation. The version of the movie that I saw isn't the same as the one that Salim saw, so now we're debating how the movie ends because we saw a different AI-generated version.
It's going to be great. It's going to be really, really cool because we're going to have a lot to do with our time. Emad, listen, you spent so much time as CEO of Stability AI in this market arena of entertainment and video production and such. When I asked you earlier whether Hollywood is going to be disrupted, you said no. Can you explain that, please?
I think the thing that won't grow is people's attention. If you look at Netflix, their biggest competitor is video games, which is why they're going into video games. You only have so many hours in a day, and you're a consumer.
The video game sector right now, I think, is $450 billion. The movie sector is $70 billion. That's how fast it's grown. Education around the world is about 10 times larger, so it's 10% of education in terms of size.
If you think about that, then for Hollywood studios, this is great because the costs have come down, and it's been a dramatic shift. To give you an idea, the first video models—Stable Video Diffusion, I think, was pretty much the first—we released that in 2023. Now, with Veo 3 from Google and others, you're pretty much at a Hollywood level, close to it, but you need one more generation to get there.
The average Hollywood shot length is 2.5 seconds. It used to be 12 seconds. We can generate 8 seconds, and soon we'll be able to generate more.
You're getting to this point where you can make that. But again, people like having common stories to talk about—Barbie, Oppenheimer, and things like that. These marquee things can get the license of Cary Grant from back in the day and make him a star again.
Don't you think there's going to be so much supply that if I have a chance to watch a new episode of classic Star Trek, but I'm the character playing Captain Kirk, you're playing Spock, and my friends are taking the other roles, why would I not buy that entertainment from a source outside of Hollywood?
You'll buy that too, but I think one of the things we've seen in the AI world is that it's about distribution, distribution, distribution. You'll buy your interactive games and put yourself in the game, but you'll still have your marquee things. The cost of those will reduce dramatically, the distribution cost will decrease dramatically, and the impact will increase.
For companies, this is all great. For the individuals working in the industry, this is terrible.
I think this is the key thing. For individual creators, this is great because you can finally tell the stories. We'll see richer stories, but you've still got to distribute them. One of the examples I had to give is Taylor Swift—bless her heart, it's not the best music in the world, but she still causes earthquakes.
Yeah.
Yeah, no. Your point that the video game industry bypassed all other media combined—I think I read that—
And it's on a much faster growth trajectory as well.
But I think video games are far more compelling with AI components, AI players, AI voices, voices that are talking directly to you. That interactive media is going to get even more accelerated by this trend. Whether you call it movies, video games, or something else, the media is going to change. It always does.
It may not fit exactly in those swim lanes, but clearly the interactive, “talk to me” part is going to grow much, much faster than the passive-watching part.
I think it's the quality part, and it's the feedback for you to find flow. The movie industry has grown from about $50 billion to $60 billion in the last 10 years. The average IMDb score is 6.3.
The video game industry has quadrupled. It was $170 billion, and now it's about $500 billion. The average score has gone from 69% on Metacritic to 74%. Games are good now—
You need to be good to compete. Again, I think what we can see from this technology is that I, as a creator, can create the best things better because I can control every pixel. This is what Jensen Huang has said.
Every pixel will be generated exactly as it’s in your mind. Maybe you have to use a keyboard, but it just comes straight from your mind and can be on that screen. You can tell the stories you want, and on the other side, you’ve got the fast food. The general content farms get even better, so you’ve got your gourmet and you’ve got your fast food, and the quality of both of those will increase.
Of course, Grok for coding. Let’s take a quick listen.
Right. So, if you think about the applications out there that can really benefit from all those very intelligent, fast, and smart models, coding is actually one of them.
Yeah. So, the team is currently working very heavily on coding models. I think right now the main focus is that we recently trained a specialized coding model, which is going to be both fast and smart. I believe we can share that model with all of you in a few weeks.
I still remember, Emad, when you were on stage with me, like, 3 years ago at the Abundance Summit, and you said, “No more coders in 5 years,” and it was front-page news throughout India.
I got hate mail about that.
Oh my God. You scared the daylights out of everyone, and it’s true. I mean, it’s a big issue. It’s a big issue. Why would you be able to talk to a computer better than a computer can talk to a computer?
Well, hold on. Let me drill into that just for a second. Don’t you think we’ll end up with really good coders just creating 100 times more code?
No. Because what you’ll have is really good context engineers directing them to build things. Code is an intermediate step of language because the computers and the compilers couldn’t handle the complexity of what we wanted to talk about.
Now you can talk to the AI all day long about anything, and it understands, to a reasonable degree, what you actually want. Once we get the feedback loops really going, as we’ve seen with Cursor and other things like that, there’s a reason it got to $500 million in revenue in a year. There’s a reason Anthropic got to $4 billion. Probably two-thirds of that is code.
Mhm. Yeah, crazy. All right. It’s disappointing that we won’t have this for a couple of weeks. We’ll have to get back on the pod and check it out when it’s out. Somebody told me you can get to it through Cursor right now. I’m looking at Cursor as we speak, and I don’t see it popping up as an option.
Cursor is very much linked toward Anthropic, so it probably lobotomizes it. But Grok 3 and Grok 4 are already pretty good at code. They write clean code, and the coding model, I think, will be even better. But again, how much better are you going to get when you can output a 3D video game like that, or just about anything?
And I think this comes to the question: if you’re trying to create content, the AI is good enough already for just about anything.
If you’re trying to create something creative, this is the final part that requires planning and coordination and multi-agent systems. The UI/UX isn’t there yet for the feedback loops, et cetera.
Yeah. I can use all the horsepower they can give me, though, because when you’re writing a little code module, it’s all pretty much perfect already. But right now I can go to the best Claude model and say, “Build me a dashboard for this function,” and just give it that prompt, and most of the time it comes back great. It even thinks of things that I wouldn’t have thought of for that dashboard, and I can use another step up of capability in that area. So I’ll use it up as quickly as it comes out, believe me. All the tokens, Dave.
Okay, let’s hear from Elon about his video model training. What’s coming on input/output?
We expect to be training our video model with over 100,000 GB200s and to begin that training within the next 3 or 4 weeks. So, we’re confident it’s going to be pretty spectacular in video generation and video understanding.
So, 100,000 GB200s—more than anybody’s thrown at this. What is that? How does that hit you?
So, when we trained the first state-of-the-art video model 2 years ago—
That’s right.
We used 700 H100s.
700.
700 H100s. So, let’s say they’re 3 times slower. That’s the equivalent of 200 of the chips that he’s about to use, because these are the integrated GB chips from NVIDIA.
The top-level models right now, if you look at the Lumas of the world, the ByteDance models of the world, and the Veo 3s, use 2,000 to 4,000.
Wow.
He’s about to use 100,000 of those. The thing about video is, when you train a video model, it actually learns a representation of the world through computation. So, once we made a video model, we extended it to a 3D model that could generate any 3D asset. It understands physics and more.
So, actually, video models are world models that can be used to do all sorts of things, like improve self-driving cars by creating whole worlds, and other things like that as well. I think that’s the reason why, given they’ve got 300,000 chips, they’re putting 100,000 of these into their video model.
Well, and they’re planning a million GPUs by the end of this year. It’s like no small dreams here. Emad, when you pioneered this just a couple of years ago, like you said, the video model was trained completely separate from the large language model because it was just too much. You couldn’t put everything into one mega-model. Is he going to do a monster retraining of this model with video data, or is it a separate set of parameters and a separate model entirely?
This will be a separate model. So, we took the image model, then we created the video model from that, and then we created the 3D model from that. Now they’re doing training from scratch because the technology we developed for Stable Diffusion 3—the diffusion transformer architecture—is able to do that all at once. This is similar to what Veo 3 and others use. With optimizations, you can just pop that all straight in.
Now, the architecture that they use, like the Grok model for images, is actually the same architecture as for the language. They may do the same thing. I’m not sure how they’re going to train this model because, again, they’re super smart. But it’s a different model entirely.
They may all end up being the same model because, if you want a model that understands physics and the wonders of the universe and what’s the question to get to the answer 42, you probably want to train on everything that a human sees and more. It’ll train on everything a million humans can see and understand and read, and all sorts of stuff.
I mean, I’m excited about the idea that there are so many of my favorite science-fiction books that have never been made into movies or TV series, right? I mean, the ability to just say, “Hey, one of my favorite books is the Bobiverse series by Dennis E. Taylor. I love it. It’s a 4-book series. It’s extraordinary. Make it into a movie for me. Make it into a 20-part TV series for me. Here’s $100.”
100 bucks.
It would be really fun, actually, if you took the best books that have ever been turned into movies already and used that as training data. So, like, this book turned into this killer movie. Make the changes necessary to get from point A to point B. Okay, now here’s a book that never got made into a movie. From what you learned about those patterns, make the movie that’s most compelling.
The thing is, you won’t even have to do that. With the pace of chip improvements as we go through the generations, in 2 years you will have live 4K TV. You’ve already seen some people do live, low-resolution, interactive stuff. When Jensen says every pixel will be generated, he literally means it.
With the next-generation chips and a bit more improvement in the algorithms and optimization of the models, you can have live-streaming 3D or video where every single pixel is generated on your screen within a few years. So you can just say, “Stop, try this, adjust this,” and that’ll be the feedback loop.
It’d be fun to take some old movies and make them way better. Like, take the old Conan the Barbarian movie and make it really a proper movie.
Oh my God. You know what hits me? We’re sitting here having this conversation in 4 different cities around the world, where we’ve taken so much for granted in this video channel. Ten years ago, what did we have? We had just barely had Skype. And now it’s crazy.
We humans adapt so rapidly to awesomeness, and we normalize it very fast. It’s like your second Waymo ride, right?
Yeah.
Your first one’s like, “Wow.” And your second one was like, “Okay.”
Oh, for sure. So, any closing thoughts on Grok?
I have a question for Emad. You’ve been in the space for a while now. We have Grok 4, right? What are the types of things that Grok 5 will be able to do?
So, Grok 5 will be a multi-agentic system, but rather than having 4 boosters, it’ll have 60 or 600 or 6,000, depending on what you want.
It’ll probably have a world model plugged in, and it’ll have interconnectivity—and this is something that Elon mentioned yesterday—to every major type of system. So it knows how to use Maya, it knows how to use advanced physics simulators, and it will write its own Lean code and optimize it for mathematics. It’s just going to be an incredibly versatile worker.
And just like he’s going to unleash millions of Optimus robots, he’s going to unleash billions, if not trillions, of these things—GPU demand notwithstanding—into the economy, and that’s going to be a bit crazy. I think the way that you’ll interact with Grok 6, probably Grok 5, is that you’ll have a Zoom call with it, just like you have now.
Mm-hmm.
Gemini 3 and GPT-5. Let’s talk for one second about what you expect there. Are these going to just leapfrog Grok 4? Are they going to be sort of diverging in different directions? Emad, your thoughts?
I think they’ll probably all be kind of on the same plateau. Now it’s really about the UI/UX, and then how you wrap these into agents and then multi-agent systems.
And then how you make it so easy for anyone to use, like this.
So, you know, Google, in the work that they’ve done with their AR glasses, is enabling you to have a conversation with your AI and being able to have it see what you see. That’s a great step forward. OpenAI, with their voice mode, has been fantastic. Are there any versions of a user interface that we haven’t seen yet? I mean, BCI will be one of them for sure.
I mean, I personally think, again, the interface is just the interface that you have with a remote worker, and all the technology is almost in place for that. Like, get on a call, hit him on Slack, pretty much, and you just don’t know. That’s my AGI. My AGI is actually—I should call it—actually useful intelligence, right? This is, I think, probably what Salim would like. I don’t know if it’s an AI or not; it just gets the job done, and it doesn’t sleep.
The final part of it as well is that the task length of these AIs has gone to 7 hours now. I think I’ve seen from various entities that they’re getting that up to almost arbitrary length. So you can set teams away, and they have organizing AIs and others. They get the job done, and they check in whenever they’re unsure about something. This is that next step up for all these technologies.
But I think the 10^27 models will, as you said, all be pretty much similar, because they’re already above PhD level in everything. Now it’s about making them super useful and getting them out there. The demand for that is in the billions of agents.
Dave, you know what I find interesting? Elon’s got basically a limitless capital supply.
Yeah. Every time he’s gone to raise money, I’ve asked, “Well, how much can I get in the next round?” And it’s like, “Well, we’re oversubscribed already.”
Yeah, yeah. No, the constraint isn’t going to be the money. It’s going to be the GPUs. I have a question for you, Emad, about that, actually, because if you say GPT-5 will be out soon—hopefully in a couple of weeks—it’ll probably be on the same plane, probably leapfrogging, but in the same genre. Then Gemini 3 will come out and it’ll be somewhere similar, maybe a little better.
But the chip supply—you know, Google has huge amounts of GPUs and a massive cloud computing platform, plus they make their own TPUs. Then you’ve got a million chips going to Elon; we just talked about that. Sam at OpenAI has had a little bit of trouble with Microsoft recently. There’s definitely some kind of falling out there.
The way OpenAI got ahead of everyone in the first place was by getting access to compute from Microsoft. Is he going to have a problem catching up to a million concurrent GPUs training a single massive model? I mean, I think Stargate is in that order of magnitude when you look at the kind of gigawatts. Amazon has just announced support for Anthropic using Trainium for something that’s even bigger than Stargate with their latest kind of chip supply.
Google’s the leader in this. They have 3 million or so. But the thing that I come back to is that OpenAI basically slowed down when everyone was making Ghibli memes. If you think about the order of compute for Ghibli memes compared to the order of compute for useful work, it’s that versus that, right? Google is okay because Google is actually landing millions of its own TPUs, and they have the full stack. It has better interconnect for large context length. It’s actually really good 7th-generation hardware.
Elon will get the supply because he’s a beast. I think, again, OpenAI has the capital, but they’re moving more and more toward consumer with the Jony Ive acquisition and things like that.
The dark horse here is probably Meta, to be honest, because Zuck is going to drop $100 billion.
Yeah. On this, he dropped $30 billion on the glasses and the metaverse.
He thinks AGI is coming, and Meta is a $1.7 trillion stock. It will easily drop $100 billion.
Yeah, he’s got $70 billion of free cash right now to use and can pump it up.
Well, I did an interview with Yann LeCun at MIT not super long ago, and they had committed to and already bought 1 million GPUs for internal use at Meta. He had those on order already then. I’m sure they’re in-house now, so he has the compute in-house.
So basically, all the top guys can get 1 million. The next step is 10 million.
Well, there are only 20 million in the world. This is where it runs into a bottleneck.
You can’t even keep a straight face, can you?
Well, again, think about every pixel being generated, and think about the economic activity of actually having a single useful teammate or accountant. We’re talking about accountants and lawyers and other things like that on the other side of the screen. We’re not even talking about super-genius PhDs.
Is NVIDIA just going to keep going and going? Is anybody going to displace its production at all? All of the top chip manufacturers are good enough to run these models. The only question is who has enough gating supply. The reason for the Hopper thing was actually the packaging of the chips—the CoWoS.
So you have different supply-chain constraints, just like with robots. In 2 years, robots will be good enough to do 90% or 95% of human labor. The only reason the entire global economy isn’t going to flip over from $2 or $1 robots is supply chains.
What we’ve got is a complete replacement of the capital stock of the economy, from GPUs for virtual workers to robots, and it’s just supply-constrained. So, NVIDIA number one—you don’t go wrong, you don’t get fired getting NVIDIA—but you’ll get chips from wherever you can get them, because those chips are orders of magnitude cheaper than your team members.
I just asked Gemini in the background here what it costs at today’s market rate to train a ronnaFLOP. So, one of these models, just for compute, costs $312 million. Like you said, Emad, that’s like one signing bonus over at OpenAI these days. The cost isn’t the issue. It’s who has access to the compute.
What’s amazing to me in this entire conversation is that we haven’t said the word Apple once.
Yeah, and Apple controls about a third of the manufacturing capacity at TSMC for its M3 and M2 line chips. So it could easily become a player in the “get a big data center up and running” game. It would have an incredible asset, having that manufacturing toehold with TSMC. It’s just incredible that it hasn’t done that.
Well, I think this comes down to the thing: These models have economies of scope, in that once you train a model that’s good enough, do you really need another one? Then it becomes like electricity; it becomes a utility. Your genius models become utilities, and then what matters is the model that runs on the M3 or whatever.
Liquid AI is just releasing edge models. Those things become even more important, because the M3s and M4s have capacity.
Yeah, that’s a really big deal. By the way, Liquid AI—I didn’t appreciate how big a deal it was until recently—but people are going to want to use this stuff immediately. It’s so addictive, and inference-time compute is severely constrained. Liquid runs fine on the edge on these M3s. It runs really, really fast. It runs on the chips in cars, and they say it’s about 100 times more efficient than trying to run a brute-force transformer. That could be a huge unlock for people having access to AI, or at least more access to keep up with the demand.
Exactly. Because you’ll have your gated stuff, and then they might increase prices because they have to, because there’ll be so much competition for chips, even as you get them cheaper. Then you’ve just got this AI with you, but that AI will be smart enough to do your day-to-day.
And so you'll have a whole curve of intelligence, just like sometimes you need to have steady workers and sometimes you need your geniuses. I forgot you were actually the first guy to see Liquid AI when it was just a research project.
Yeah, I gave them all the compute to get going.
Yeah, that's right. That was—
—amazing. Now they're at a $2 billion valuation.
So listen, when you come back and join us next week—I think we have it scheduled—I want to hear all about the intelligent internet. I'd love you to break the news on what you've been working on in secret for the last year or so. I've seen pieces of it. It's awesome, but hopefully you'll spill the whole master plan for us. Dave, Salim, my Moonshot mates, thank you guys. Grok 4 special edition.
See you at Grok 5.
Yeah, we got Gemini 3.
In like 3 weeks.
We'll be back online soon.
All right, see you all. Thank you for joining us.
Take care, folks. Bye, guys.
Bye.