Steeve Morin:为什么 Google 将赢得 AI 军备竞赛,而 OpenAI 不会 | E1262
- Morin 的核心判断是:H100 是一个迟早会爆的泡沫——算力过剩、行业陷入困境的情形「可能就在今年」。 市场沿用 A100 的金融模型——第0代负责训练,上一代负责推理——但 H100 价格达到 5 倍,推理性能只有约 2 倍,还通过 GPU 本身作抵押、以6至7年摊销期融资。「总得有什么东西让步」——他已经收到从未听说过的服务商发来的冷邮件折扣,并猜测 GPU 可能会以「票面价值的30美分」被买走。
- 5年后的价值分配将是:「95% 推理,5% 训练」。 Agent 和 reasoning 将算力经济从吞吐量约束转为延迟约束——一张 GPU 可以在100路并发上服务每秒10,000个 token,却无法把每秒10,000个 token 全给一个用户——而且「这正是 Nvidia 可以被攻击的地方」。但如果今天被迫只能买一只股票,他仍然会选 Nvidia,「因为供给在那儿」,希望以后能回来告诉大家尽可能多地买 AMD。
- Google 是沉睡的巨人。 他认为胜负取决于产品、数据和算力组成的三角形——「谁三者都有?Google……它们可以把能力铺到各处」。肤浅的说法是 OpenAI 威胁搜索业务;现实是「OpenAI 很惊艳,但算力不属于它」——算力属于 Microsoft,而「如果你不拥有自己的算力,你一开始脚上就拴着东西」。
- 对于所有基于 AI 构建业务的人来说,真正可交易的洞见在于利润分层:TSMC 以60%毛利率出售,Nvidia 达到90%,Amazon 再拿走约30%。 「你只是一个非常大的蛋糕上薄薄的一层脆皮。」他建议创业公司尽可能不要转售算力;每花1美元,约98%都进了别人的利润池,根本无法与每月20美元的订阅价格相匹配。
- 尽管 AMD 的推理效率约高4倍,切换经济学仍解释了它为何没有追平差距。 「好7倍还不够让人切换——人们宁可什么都不要,也不要一个东西。」解决办法是自上而下:Microsoft 买下 AMD 的供给,并让 ChatGPT 跑在 AMD 上,这让 OpenAI 的推理业务「转为盈利」。ZML 的主张是把迁移成本降到零,这样「如果你比我好30%,我就切换到你」。
- 在新型芯片方面,Gro 和 Cerebras 证明了 SRAM 能带来单流速度,但这种经济性无法扩展。 Gro 每颗芯片只有230MB,而 BF16 下的70B模型需要140GB——「自己算算就知道」。他认为 Etched,以及很可能还有 VSORA,正在朝着更低价格的方向发展——「如果你能以接近 GPU 的价格交付,那你就赢了」——再往后的前沿则是存内计算,很可能包括 Rain AI 和 Fractile。
- 至于 Stargate,我不买账——它是一辆「AI 美国车」:体量大、耗油多,但归根结底不是一辆好车。 真正的限制只有「人才和能源」。与此同时,Blackwell 订单正因散热和芯片弯曲问题被取消,而 Jensen 真正的挑战是「如何走过下坡路——高点非常高,但不会永远持续」。
1. H100 泡沫可能爆裂——GPU 可能跌到票面价值的30美分
- Morin 的核心警告是:这个行业建立在 A100 金融模型之上——第0代负责训练,上一代负责推理——「而且运行得非常漂亮」。随后 H100 以5倍价格上市,推理性能却只有约2倍(上市时推理速度与 A100 相同)。「这里存在一个必须填补的资金缺口……总得有什么东西让步。」
- 最让他担忧的是:摊销期达到6年、7年,而且 GPU 本身就是抵押品。他设想的情景是:美国某处数据中心有1,000张 GPU,「人们可能会以票面价值的30美分买走」。时间点是:「可能就在今年。」
- 迹象已经出现:「我开始收到一些从未听说过的服务发来的冷邮件折扣」——大约从10月至11月开始。建设训练集群和建设云服务商是两种不同的生意,有些人手里留下了不知道该如何处理的资本开支。
- 同样的下坡逻辑也适用于 Jensen:Blackwell 订单正因散热和芯片弯曲问题被取消;H100 的供给在全年被摊平,但一些买家至今还没拿到去年的订单,B200 却已经出现了——「他们非常愤怒」。「高点非常高,但不会永远持续。」
2. Agent 和 reasoning 将算力从吞吐量约束推向延迟约束
- 他整套判断背后的机制是:聊天场景奖励流式吞吐量,但使用 Agent 和 reasoning 时,「你得等请求结束,才能拿到自己想要的东西」——真正重要的只有从发出请求到答案完成的时间。GPU 可以在100路并发上服务每秒10,000个 token,「但无法只给你一个人每秒10,000个 token」。他的结论是:「这正是我认为 Nvidia 可以被攻击的地方。」
- Nvidia 当前采用的是纵向堆叠路线——「更多的更多」——这正在逼近物理极限。Blackwell 把两颗芯片拼在一起,形成了大到会「开始弯曲」的表面,导致芯片与散热器失去接触;功耗上限达到1,000瓦,还需要液冷。
- 下一步是潜空间推理——不把思考翻译成 token,就像「你在心里思考时,不会把自己的想法大声说出来」。再加上可能属于 Yann LeCun 路线的能量模型,这些任务「从根本上说 GPU 无法大规模交付」,因为外部存储访问才是墙:HBM 已经是我们拥有的最好方案,但与 SRAM 相比仍然非常慢。
3. GPU 从来不是 AI 芯片——Nvidia 赢在训练,靠的是 Mellanox
- 他希望人们记住的历史是:GPGPU 大约20年前只是一个「很酷的技巧」——让图形渲染器去做并行数学运算。「GPU 是用于 AI 的好技巧,但它们不是为 AI 打造的。」TPU 才是最早的专用架构;对 LLM 来说,GPU 这套技巧在内存传输压力下「开始出现裂缝」。
- 他的修正性观点是:Nvidia 赢得训练市场,「靠的是 Mellanox,而不是计算能力」——矩阵乘法过程中,一个矩阵会被读取「数百次」,所以 InfiniBand 互联才是训练的「核心」。而推理「没那么重要。你不在乎」。模型大小之所以被这样设计,实际上就是为了让人们无需连接多台机器也能完成服务。
- 锁定效应为何仍在持续:PyTorch–CUDA 组合「非常、非常难以打破」,而一个自我强化的循环——人们买 Nvidia 再转售,因为所有人都在运行 Nvidia——让市场持续被效率最低的选项占据;在他看来,这甚至不是最好的软件平台。
4. 「你只是一个非常大的蛋糕上薄薄的一层脆皮」
- 他拆开利润链条后的秘密是:TSMC 的毛利率为60%,Nvidia 达到90%,然后 Amazon 再拿走约30%。因此,他对 AI 创业公司的长期建议是:「能不转售算力,就不要转售」——每花1美元,「可能有98%流向了别人的利润」,这很难被压进每月20美元的订阅里。
- 至于 CUDA 的护城河:「他们花了大量精力让你在意一些本不该在意的东西……谁在乎 CUDA?」一旦软件把各种底层差异抽象掉,「服务商就会围绕规格竞争,而不是围绕伪护城河竞争」——不过他也承认其中的政治现实:「如果你惹恼了 Jensen,可能就得去吻戒指,才能重新排队。」
5. 训练是研究,推理是生产——推理将占95%
- 两者服从「根本不同的构造性力量」:训练需要一切都更多,以及更快迭代——「就像给一辆正在行驶的汽车换轮子」;推理需要更少:少一些麻烦,能不做互联就不做。他的比喻是:画一幅画和画一百万幅画不同,前者看重落笔速度,后者看重可靠流程。
- 贯穿本期的数字是:5年后,「95% 推理,5% 训练」。
- 今天的推理「充满了胶带」,但它并不只是运行训练代码的前向传播。最大的未兑现机会是自动扩缩容:即使生产环境根本没人使用,用户仍然每天为1,000张 GPU 付费,「这太疯狂了」;按需配置的价值可以达到「5倍,有时10倍」。由于按需算力价格宰人,用户过度采购预留算力,进一步制造了叠加在真实稀缺之上的「虚假稀缺」。
6. 好7倍还不够——迁移成本必须降到零
- 每一家芯片厂商——AMD、Google、Tenstorrent——都面临同一个 GTM 问题:切换意味着放弃一份6年期 Nvidia 摊销计划,还要维护两套技术栈。他得出的经验指标是:「无论你选择什么指标,好7倍都不够让人切换。人们宁可什么都不要,也不要一个东西。我有实际案例。」TPU 就是证明:「它们有货,也很好,但没人关心」——因为迁移成本是6个月的工程投入。
- 反例是自上而下的打法:Microsoft 买下 AMD 的供给,并让 ChatGPT 跑在 AMD 上,这「让 OpenAI 的推理业务转为盈利」。计算方式是:8张 H100 只能容纳2个70B模型,而 GPU 数量翻倍,推理吞吐量只增加约10%——「这是没人谈的秘密」;AMD 的内存可以让每张 GPU 放下一个模型,芯片效率约高4倍,价格还低30%。「这根本不需要思考」——但前提是从产品出发、向下渗透;如果从底部把它们租出去,「人们不会租,事实就是如此」。
- ZML 的主张是通用解法:把迁移成本降到零——「你只要说,嘿,现在它是 AMD——砰,它就能运行。」这样一来,「如果你比我好30%,我就切换到你」,增量优势终于可以转化为销售。
- 如果被迫选一只股票,他说:「我仍然会选 Nvidia……因为供给。」但如果牌打得好,希望有一天能回来告诉大家「尽可能多地买 AMD」——或者,如果 Tenstorrent 上市,也可以买 Tenstorrent。
7. SRAM 经济学:Gro 和 Cerebras 惊艳,但负担不起
- 单流速度没有魔法,靠的就是 SRAM,也就是芯片上的内存。但 SRAM 会吞噬芯片面积、降低良率:Gro 每颗芯片只有230MB,而 BF16 下的70B模型需要140GB——「自己算算就知道」。Cerebras 的晶圆级引擎容纳44GB,用水冷却,并用铜针接触芯片:「简直疯狂……不过技术确实非常令人印象深刻,只是贵得离谱。」他怀疑 Mistral 和 Perplexity 的速度展示「是在亏本做」。
- 他认为正在朝着打破价格门槛方向发展的公司包括 Etched,以及很可能还有 VSORA——「如果你能以接近 GPU 的价格交付,那你就赢了」。
- 对 Nvidia 在 HBM 上的地位,他用了「Nutella 情形」来比喻:Nutella 拥有80%的榛子市场——「是的,你可以做一个竞争者,但你要从谁那里买榛子?」需要注意的是,纯 SRAM「在扩展性上是一条死路」;未来需要 HBM,同时也需要在更专用的架构中增加 SRAM。
- 再往后的前沿是存内计算:让处理器靠近数据,而不是来回搬运数据。他点名了两家公司:Rain AI,很可能 Sam Altman 是投资人之一,「并不意外」;以及英国的 Fractile。「可能不是今年,但它会到来。」
8. 产品、数据、算力——「这就是沉睡的巨人」
- 他判断 AI 胜负的模型是产品、数据和算力组成的三角形。「谁三者都有?Google……Android、Google Docs,等等——它们什么都有,可以把能力铺到各处。这是我心中的沉睡巨人。」当然,他也干巴巴地补充了一句:「前提是它们没有忙着重组。」
- 他反对的肤浅说法是:OpenAI 会摧毁搜索这只下金蛋的鹅。「OpenAI 很惊艳,但算力不属于它」——算力属于 Microsoft,而 Microsoft 是以「离谱的利润率」买来的。「如果你不拥有自己的算力,你一开始脚上就拴着东西。」那些购买数万张 GPU 的数据中心建设者告诉他,他们得到的唯一让步就是供给,没有折扣。
- 至于谁真正威胁 OpenAI:「我不确定目前谁是 OpenAI 的威胁。」他的母亲知道 ChatGPT;但如果问她 Gemini,「那是什么?」最终,借用 Fluidstack 的 Gary 的比喻,模型提供商会变成「汽车制造商」——不会赢家通吃,「每个人都会有自己的品牌」。
9. 暴力堆算力遭遇反击:DeepSeek,以及「一辆 AI 美国车」Stargate
- 今天的扩展路线是「暴力堆算力……一种非常美国式的路线——越来越多、越来越多、越来越多」,但物理规律正在反击:xAI 的集群不是100,000张 GPU,而是「4个25,000张」,因为互联成了约束——「某个时刻你会开始与物理规律搏斗」,就像逼近光速。他认可的两种反击路径分别来自 DeepSeek——「存在大量浪费……因为效率更高,实际上等于增加了算力容量」——以及可能属于 Yann LeCun 的路线:「这不是扩展,我们需要正视问题本身。」
- 至于 DeepSeek 本身:「约束是创新之母……如果你可以买更多,为什么要在乎?」中国在 ASIC 方面「有点晚」,大约还处在 A100 水平,但它正在水中练习:「它们注定会做得更好……用短期思维看待这件事是不对的。」同样的视角也影响了他为 Mistral 辩护:「散播 FUD 很容易……我知道的是,我希望它们不要有太多钱。」
- 主持人提到 Meta 的超大规模资本开支为600亿–650亿美元,Microsoft 为800亿美元;Morin 认为,这些支出追逐的是训练,而 Nvidia 的采购具有「增量性」——1,000张 Nvidia GPU 加上1,000张 Nvidia GPU,得到2,000张;但1,000张 Nvidia 加上1,000张 AMD,得到的是「两倍的一千」。
- 对 Stargate 的5,000亿美元,他的第一反应是:「我不买账……美国式做法——先提出一个主张,再想办法实现。」即便它是真的,也只是纵向扩展:「一辆 AI 美国车——体量大、耗油多,但归根结底不是一辆好车。」限制因素只有「人才和能源」;而在90%毛利率的背景下,「如果我们对这个利润率动手,就能把那个数字压下来」。
10. 模型正在消失——后端、世界模型与更好的数据
- 他对模型走向的开场判断是:闭源「模型」已经是后端的「星座」——当你要求生成一张图时,系统会在后台把请求路由给扩散模型。「模型」如果被理解为拿到权重并运行它们的东西,最终会消失,取而代之的是完整后端。
- 他对与可能属于 Yann LeCun 的世界模型路线、可能是 JEPA 的方案「非常看好」:「他向我解释它是怎么工作的,我当时完全被震撼了。」语言只是「通往世界的一扇非常窄的窗口……我们对世界的描述,严格来说有一定损失」;能量最小化式的思考更接近人类拿起 AirPods 盒子的方式:不会把整个动作过程讲出来。与此同时,非 Transformer 架构「可能会彻底淘汰 Transformer」,扩散式 LLM 则像人类一样以「分块」方式思考。
- 对合成数据,他依然保持原有的分寸:把模型输出重新注入训练数据「感觉不对……会发生退化」;但 AlphaGo 的跃迁来自合成棋局,而代码就是一个例子,因为「你只要运行代码」,就能获得真实答案——这也是 Poolside 的主张。蒸馏则「完全合理」:一个模型被提示生成一帧 Star Wars 画面,结果产出了真实截图,所以「你从一开始就把它拿走了……我拿起我的杯子,每天都非常享受那部电影」。最令人震惊的是,蒸馏后的小模型有时会击败教师模型——在你测量的那项任务上表现得更好。
- 发展方向是更小的模型:「很多时候,运行大模型其实是在浪费」,而 DeepMind 的人「甚至已经不再微调」——他们在运行时把数据直接倒进巨大的上下文窗口。Morin 说 ZML 正在研究作为 RAG 前沿的「注意力级搜索」;Harry 称这是独家消息。它延伸了 RAG 的那个「聪明技巧」:把语义上接近的片段预先拼接到上下文中。
The thing with NVIDIA is that they spend a lot of energy making you care about stuff you shouldn't care about, and they were very successful. Who gives a shit about CUDA?
OpenAI is amazing, but it's not their compute. Ultimately, if you don't own your compute, you're starting with something at your ankle. In 5 years, I would say 95% inference and 5% training. You have the products, the data, and the compute. Who has all 3? Google. That's Android, Google Docs, whatever—they have everything they can sprinkle everywhere. This is the sleeping giant in my mind.
Steeve, dude, I am so grateful to you for joining me today. I've wanted to make this one happen for a while, but when we were discussing who would be best for this topic, I was like, "We've got to have Steeve on." Thank you for joining me today.
Well, thank you. I feel humbled. I appreciate it. Thank you.
1. How Will Inference Change and Evolve Over the Next 5 Years
I want to start: can you just give a quick overview of ZML, and specifically your role in the infrastructure strategy today and where you sit at the very bottom of things?
ZML is an AI/ML framework that runs any model on any hardware, and it does so without compromise. We sit ultimately at the infrastructure layer. We enable anybody to run their model better, faster, and more reliably, but on any compute whatsoever. It doesn't really matter; it could be NVIDIA, AMD, a TPU, or whatever.
We do all that without compromise. That's the key point, because if there's a compromise, then it's not really agnostic, right?
Can I ask you, then, if we think about sitting between any model and any provider, whether AMD or NVIDIA, do you think we will exist in a world where people are using multiple models simultaneously and running them concurrently?
Yes. You can actually see it; it's been happening for a while. Models now are not the right abstractions. At least if you look at closed-source models, they're not really models—they're more like backends.
There are a lot of tricks that make you feel like you're talking into 1 model, but ultimately you're talking to a constellation, an assembly of backends, that produces an AI response. Probably the number 1 obvious thing would be that if you ask a model to generate an image, then it will switch to a diffusion model, not an LLM.
There are many, many more tricks. The Turbo models at OpenAI do a lot of tricks. Definitely, models in the sense of getting weights and running them is something that is ultimately going away in favor of full-blown backends.
You feel like you're talking to a model, but ultimately you're talking to an API. The thing is, that API will be running locally—or locally meaning in your own cloud instances and so on.
Okay, so we will have a world where we're switching between models, and there's this kind of trickery around it. Perfect. So we've got that at the top, then we've got ZML in the middle, and then you said on any hardware. Will we be using multiple hardware providers at the same time, or will we be more rigid in our hardware usage?
2. The Importance of a Top-Down Strategy for Microsoft and Google
No, absolutely. You can get probably an order of magnitude more efficiency depending on the hardware you run on. That is substantial. Not a lot of people have that problem at the moment because things are getting built as we speak, but a simple example is that if you switch from NVIDIA to AMD on a 70B model, you can get 4 times better efficiency in terms of spend. That is substantial—very much substantial.
If there's such a cost efficiency—4 times—why does everyone not do that?
3. Challenges and Innovations in AI Hardware
There are a few reasons. Probably the most important one is the PyTorch-CUDA duo, and that's very, very hard to break. These 2 are very much intertwined.
Can you explain to us what PyTorch is?
Yes, absolutely. PyTorch is the ML framework that people use to actually build and train models. You can do inference with it, but by far the most successful framework for training is PyTorch.
PyTorch was very much built on top of CUDA, which is NVIDIA's software. The way PyTorch works makes it ultimately very, very bound to CUDA. Of course, it runs on AMD, it runs on Apple, and so on, but there are always tens of little details that don't run exactly as you would expect. There's work involved.
Then there's also supply. Probably that's the number 1 thing. The second thing is that there are a lot of GPUs on the market, and pretty much all of them are NVIDIA.
The reason is that if you think in layers and say, "All right, I'm going to buy GPUs and sell them to folks to maybe not even do training, but just do inference," then most likely, if you look at it that way, you'll end up buying NVIDIA because everybody will want to run on NVIDIA. Nobody really knows how to do otherwise, and they've trained on NVIDIA, so they're thinking, "I can just reuse my code," and so on.
There's this self-perpetuating circle of people buying NVIDIA because they want to resell it, and people using NVIDIA because it's there. But it's by far not the most efficient platform, and arguably, even in terms of software, it's not the best software platform. Those are probably 2 of the most important reasons.
4. Nvidia's Market Position and Competitors
Before we move on, we were chatting about NVIDIA and AMD when DeepSeek happened and the stock crash happened. Why did NVIDIA rebound, do you think, in a way that AMD didn't?
There are a lot of things, but in my opinion, there's always going to be a need for inference. It's very hard to say whether it will be worth everybody's money to do it on an H100. That is a bubble that I think will blow sometime. I'm kind of afraid of that, to be honest.
Why do you think that's a bubble that will blow sometime? Why is that not legitimate?
Because it was built on the A100 financial model, which was at generation 0: we do training, but when it's the last generation, we do inference. It worked beautifully for the A100.
Then the H100 comes along, and inference is worth 5 times the price, but it maybe runs twice as fast in terms of inference. On training, it's a lot better, but on inference, it's maybe twice as fast. When it actually came out, it ran at the same speed as the A100.
There's a money gap that's going to have to be bridged sometime. The part that worries me is that I see amortization plans of 6 or 7 years with the GPUs as collateral, and I'm not sure how it's going to work. When they came out, they were worth 5 times the price and were only 2 times faster, so something has got to give.
Is the speed of development trumping chip-development speeds? Is it now becoming a real problem where models are far outpacing the speed of chip deployment?
Not much, ultimately. The 2 things that could really shake the chip industry, in my opinion, are agents and reasoning.
Why does that change things?
For agents and reasoning, you need to wait until the end of the request to get whatever it is you came for. You don't really care about the speed at which the text outputs, which is what you want in a chat. You only care about how much time it takes between the beginning of your request and the end.
That fundamentally changes the incentives from throughput-bound to latency-bound. If you're running GPUs at, let's say, 10,000 tokens per second, you very much like to do it 100 times in parallel. They can do that, but they cannot give you 10,000 tokens per second only to you, per stream, as we say.
In terms of agents or reasoning, this is exactly what you want, because you don't want to wait 50 seconds for whatever thinking is happening. Agents are the same. Those 2 things are the shot that might at least make NVIDIA change its course with respect to chips. They're not idiots, right?
How should agents change NVIDIA's strategy?
They're bound by the latency between the beginning of the request and the end of the response. You don't want your agents to take 30 seconds to generate whatever their response will be.
How should they change, then? How should they change their response?
Oh, you mean NVIDIA? It's hard to say because NVIDIA has a very, very vertical approach. They do more of more.
If you look at Blackwell, it's actually crazy what they did. They assembled 2 chips, but the surface was so big that the chip started to bend, which further perpetuated the problem because it then didn't make contact with the heatsink and so on.
They are very much pushing the pedal in terms of GPU scaling. The power envelope is pushed to 1,000 watts, it requires liquid cooling, and so on. They are very much operating at the limit in terms of GPU scaling.
The thing is, GPUs are a good trick for AI, but they're not built for AI. It's not a specialized chip; it's a specialization of a GPU, but it is not an AI chip.
Forgive me for continuously asking stupid questions, but why are GPUs not built for AI? If they're not, what is better?
The way it worked is that a screen can be thought of as a matrix. If you have to render pixels on a screen, there are a lot of pixels and everything has to happen in parallel, so you don't waste time.
It turns out that matrices are a very important thing in AI. There was this trick in which we essentially tricked the GPU into believing it was doing graphics rendering, while we were actually making it do parallel work. It was called GPGPU at the time, probably 20 years ago.
It was always a cool trick—very cool and very successful, mind you—but it was not dedicated to this. The pioneers were probably, of course, Google with the TPU, which is much more advanced on the architectural level.
5. Challenges of Incremental Gains in the Market
Essentially, the way GPUs work kind of works for AI, but for LLMs that starts to crack because they're so big and there's a lot of memory transfer and so on. That's why Groq, Cerebras, and all these companies achieve very high single-stream performance: the data is right on the chip. They don't have to get it from memory, which is slow, as a GPU has to do.
There are a lot of things that ultimately make it a good trick, but not a dedicated solution per se. That said, the reason NVIDIA probably won, at least in the training space, is Mellanox—not because of the compute.
You need to run a lot of these GPUs in parallel, so the interconnect between them is ultimately what matters. How fast can they exchange data? When you do a matrix multiplication, the matrix is read hundreds of times during the multiplication, so there are a lot of transfers going on.
So far, Mellanox with InfiniBand has the best technology. That's why a lot of people use it. When you do training, by the way, it is the name of the game: the interconnect. When you do inference, not so much. You don't care as much.
6. The Economics of AI Compute
Before we move to inference, I do want us to stay on chips and ask: we have TPUs, NVIDIA, and AMD. In terms of the distribution of gains, is this a winner-take-all market? Is it the cloud, where you have several providers who are dominant? What does the distribution of gains look like in the chip market?
I would divide it into 3 categories: the GPUs you can buy or rent, the TPUs you can rent, and the TPUs you can buy. This is how the market is structured today.
Right now, if you want to go dedicated, there are at least 2 options: TPUs and Trainium—TPUs on Google and Trainium on Amazon. These are available chips; you can rent them today if you want.
If you want to buy or rent GPUs, there are GPUs everywhere—we know that all the time. There's this new way of computing, which is dedicated chips that you can actually buy: Tenstorrent, Etched, and likely VSORA.
I think it will be a mix of whatever you can get. For instance, if you're in Google Cloud, of course you don't want to use NVIDIA because you get ripped off. Here's the dirty secret: NVIDIA, like TSMC, sells to you at a 60% margin. NVIDIA sells to you at around a 90% margin, and on top of that there's Amazon, which takes, let's say, a 30% margin. So you are a very thin crust on a very big cake.
That's why, to me, it's a bit of a losing game if you go all-in on 1 provider. You want optionality, with increasing competitiveness within each of those layers.
Do we not see margin reduction?
Absolutely, yes. But here's the problem. Let's say you're on Google Cloud and you're on TPUs. Suddenly, you remove that 90% chunk from the spend.
The problem is that, for multiple software reasons—which we are solving at ZML—they're not really a commercial success. They're very successful inside Google, but not much outside of Google. Amazon is pushing very, very hard for its Trainium chips as well.
7. Training vs. Inference: Infrastructure Needs
I would say the future I see is that you use whatever your provider has, because you don't want to pay an outrageous 90% margin and try to make a profit out of that.
I totally get you. When we move to inference and training, I think everyone's focused so much on training. I'd love to understand the fundamental differences in infrastructure needs when we think about training versus inference.
These 2 obey fundamentally different tectonic forces, if you will. In training, more is better. You want more of everything, essentially, and the recipe for success is speed of iteration. You change things, see how they work, and do it again. Hopefully, it converges. It's like changing the wheel of a moving car, so to speak.
On inference, this is the complete reverse: less is better. You want fewer headaches. You don't want to be waking up at night because inference is production. You could say that training is research and inference is production, and it's fundamentally different in terms of infrastructure.
The number 1 difference between the 2 is the need for interconnect. If you're doing production and you can avoid having an interconnect between, let's say, a cluster of GPUs, of course you will avoid that.
This is why models have the sizes they have: so people can run them without needing to connect multiple machines together. It's very constraining in terms of the environment. That is probably the fundamental difference—the need for interconnect.
Number 2 is: do you really care what your model is running on, as long as it's outputting whatever you want it to output?
Can you help me understand why training is more is more, while in inference less is more? Why do we have that difference?
Think of it like doing 1 painting versus doing 1 million paintings. The tools you use and the process you follow will be different. If you do 1 painting, what you favor is the speed at which you can make a stroke and iterate.
If you do 1 million paintings, what you want is a process—a process that is reliable and can deliver 1 million paintings efficiently. That's the same for training versus inference.
How do people then put inference into production? We've seen with training that NVIDIA has dominated so heavily. How do people put inference into production?
There's a lot of duct tape. One of the problems is that training, from first principles, is actually 2 passes: forward and backward. It's called the forward pass and the backward pass. Inference is running only the forward pass. That's how things are today, mostly.
There are people who are trying to specialize a bit because, at some point, duct tape doesn't really work out. When you're at big scale, that creates a problem. It's a problem that's growing because a lot of people are coming to the market with needs for inference that weren't there a year or a year and a half ago.
OpenAI had this problem, and maybe Anthropic had this problem, but it wasn't a universal problem yet. Now it's becoming one.
Can you articulate what problem OpenAI and Anthropic had with regard to inference?
For instance, probably the number 1 thing, depending on how you deploy, is what's called autoscaling. As your systems get more and more loaded, you want to provision compute because these things are tremendously expensive.
You don't want to say, "I have 1,000 GPUs 24 hours a day, and even if nobody is in production, I will pay for them," which, mind you, is what people are doing today. This is crazy.
What you want to do is provision compute as your needs grow. You want to scale up and scale down. That's probably the number 1 thing that gives you a lot of efficiency in terms of spend. We're talking about multiples—sometimes a 5x or even 10x improvement.
In regular backend engineering, this is a problem everybody knows. Everybody is doing it because the savings are so huge. But in AI, nobody really had that problem. Now they're coming up against it.
So the problem is that they're not doing provisioning. They're paying a ton more because they're fully in production all the time, versus provisioning as needed?
That's 1 example. Another one is choosing the right compute. It's a vicious circle because provisioning compute is very hard. If you lose compute, it's very bad, so you're essentially incentivized to overbuy.
In the case of Amazon or Google, that would mean buying reserved compute, which you're not going to use, because if you buy it on demand, you will get tremendously ripped off. That creates this scarcity of compute because people buy preemptively, pay a ton of money, and don't use it.
When you buy compute preemptively, does it not become outdated by the time you use it?
It might well be. Judging by the pace of development, it might well be. We are being spared a bit because Blackwell is late and H100s are getting canceled, so the H series is still active.
But yes, absolutely. What choice do you have?
We have a moment in time where there's this massive overhang, or oversupply, of compute that we've proactively bought ahead of time. But then the hyperscalers say, "We'd rather just burn it, buy fresh, and we have the money to do that."
I think it already started. I'm getting cold emails for discounts from services I've never heard about. I started getting these emails probably around October or November, so some people are left with a lot of CapEx that they don't know what to do with.
It's very hard to build a cluster and run a training run. That's a different thing from literally building a cloud provider, hyperscaler, or whatever you want to call it.
There are a lot of people who do their training runs on the regular providers, but then move to a regular hyperscaler when they go into production. I very much worry there will be an oversupply of these chips.
8. The Future of AI Chips and Market Dynamics
The problem is that the chips are the collateral. Somewhere in the United States or wherever, there could be a data center with 1,000 GPUs that people may buy for 30 cents on the dollar. I don't know, but this is what might happen.
What's the time frame for that?
Probably this year.
Jensen has made it very clear that inference opens up more revenue opportunity for NVIDIA. He said that 40% of its revenue today comes from inference. To what extent is that correct? Or, as Jonathan Ross at Groq said on the show, is NVIDIA not meant for inference and that market won't be won by NVIDIA?
Technically speaking, he's right, but realistically speaking, I'm not sure I agree. These chips are on the market. They're here. I can open a tab in Chrome and get one. Availability is something I don't take lightly.
I think NVIDIA is here to stay, at least if the H100 bubble doesn't burst. These chips are going to be on the market, and people will buy them and do inference with them. What remains to be seen is the OpEx, electricity, and so on, but that is a complicated question.
As far as I know, the only chips that are really frontier in that sense are probably TPUs and the upcoming chips. They're great chips, but they're not on the market, or they're available at outrageous prices—millions of dollars to run a model.
What chips are great, and why aren't they on the market?
If you look at, for instance, Cerebras, it's incredible technology and incredibly expensive. How will the market value the premium of having single-stream, very high tokens per second? There is value in that, as we saw with Mistral and Perplexity, but I'm not sure that was done profitably. I don't know the details, but I think Cerebras put it out at a loss.
Today, there are 3 actors on the market that can deliver this. I think this will be the pushing force for change in the inference landscape: agents and reasoning. That means very high tokens per second only for you, not for an aggregate of people.
What is forcing the price of a Cerebras to be so high? You heard Jonathan at Groq on the show say that they're 80% cheaper than NVIDIA.
There's a trick, because there's no magic. This little trick is called SRAM. SRAM is memory directly on the chip, so it's very, very fast memory.
Here's the problem with SRAM: it consumes surface area on the chip, which makes it a bigger chip and makes yield very difficult because the chances of problems are higher.
SRAM is very, very fast memory, which gives you a lot of advantage when you do very-high-speed inference, but it's terribly expensive. If you look at Gro, on this generation they have 230 MB of SRAM per chip. A 70B model in BF16 is 140 GB, so you can do the math.
Cerebras has 44 GB of SRAM in what they call their Wafer-Scale Engine, which is a chip the size of a wafer. Most likely it's interconnected, but it's huge, and it has to be water-cooled. They have copper needles that touch the chip to cool it. It's crazy stuff—very impressive technology, mind you, but very, very expensive.
My bet is that there will be chips on the market that do that at a much lower price. Two companies I see going in that direction are Etched and likely VSORA. If you can deliver this at a price comparable to GPUs, you've won.
Is minimizing SRAM the only way to reduce the unit cost on these chips?
It's hard to say. You need some SRAM, but if you can have a smaller process node and hook yourself up with external memory, then yes, you can do that a lot better.
But if you go full-blown SRAM, there's no magic: you will have to pay the price.
I'm so enjoying this. I'm also learning. My notes here are just expanding by the day. How do you think the inference market evolves over the next 3 to 5 years, pushed by reasoning?
Reasoning—not in the sense that you see in DeepSeek and whatever, but reasoning and what's called latent-space reasoning. Latent-space reasoning and agents will push the market toward different types of compute.
Can I just ask, what is latent-space reasoning?
The way models reason today is in tokens. It's as if, when you think to yourself, you say out loud what you're thinking.
It works, but it's a bit inefficient, and you lose information by doing this. Latent-space reasoning is reasoning without going into English or whatever language you're using—staying in what's called the latent space, which is where all the information of an LLM or an LM lives.
This is very much how we work as humans. We are moving toward what likely Yann LeCun calls energy-based models, in which we have different types of longer or shorter thinking times, if you will.
That fundamentally cannot be delivered by GPUs at scale.
Why can't GPUs deliver it?
Because access to external memory prevents it. HBM is all the rage, but compared with SRAM, HBM is absolutely dirt slow. That's the problem you get. HBM is the best we can do, but it's still slow versus SRAM.
When I had Jonathan on, he said that NVIDIA has such a stronghold because it's one of the only buyers of HBM, which gives it this unique position. Is being the sole buyer of HBM irrelevant if the world needs SRAM instead?
No, you want HBM, to be clear. SRAM will not deliver this. It's a dead end in terms of scaling. SRAM means consuming the surface area, which creates yield problems and makes everything explode.
You need some SRAM, so we will have bigger amounts of SRAM in chips and, of course, bigger amounts of external memory in chips. The issue with HBM is that it's still slow. Maybe NVIDIA has a stronghold and can prevent you from getting some, so I call it the Nutella situation: Nutella owns 80% of the hazelnut market. You can build a competitor, but who will you buy the nuts from?
This is a bit above my league, but there will be a need for HBM and a need for SRAM. Better, more dedicated architectures will be able to deliver these things. Then there's the next frontier after that, which is called compute-in-memory.
There are 2 companies that I know of, at least, that are in that market. One is called Rain AI.
Likely Sam Altman is one of the investors, so there’s no surprise. The other one is called Fractile; I think it’s in the UK, actually. This is the next frontier. The idea is that instead of transferring the data between external memory and the CPU and doing the compute there, you bring the CPU to the memory and do everything. It’s crazy stuff, but it’s coming—maybe not this year.
How does that change the situation?
It makes it much more efficient. What does that actually mean in reality? It means you get maybe not SRAM-level performance, but a lot faster performance in terms of compute. If you translate that to LLMs, let’s say you get much, much higher tokens per second in a single stream, which is exactly what you want when you go into reasoning.
You want your model to maybe think for half a second and then, boom, produce the answer. You don’t want to wait 50 seconds and context-switch to some other thing, which is the problem everybody has today, mind you. I think inference will be major. At least, this is my thesis: the compute landscape will be pushed to change because of these 2 constraints.
If you were to ascribe value between training and inference out of a pool of 100, is it 80 inference and 20 training? What does that look like in 5 years?
I would say 95% inference and 5% training.
Wow. Do you think NVIDIA owns both of those markets in 5 years’ time?
It depends on the supply. I think there’s a shot that they don’t, because here’s the thing: even if we take the same amount of compute, let’s imagine we have a new chip from Amazon that’s the same amount—wait, we do. It’s called Trainium. Why would I pay a 90% margin to NVIDIA if I can freely change to Trainium?
My old production runs on NVIDIA GPUs anyway. If you’re in production on dedicated chips, of course, there’s the issue of commoditization. If I’m on AWS, I can just click and, boom, it runs on AWS’s chips. Who cares? I just run my model like I did 2 minutes ago.
With that realization, do you think we’ll see NVIDIA move up the stack and also move into the cloud and models?
They are. They have a product called NIM that sort of does that. The thing with NVIDIA is that they spend a lot of energy making you care about stuff you shouldn’t care about, and they were very successful.
Who gives a shit about CUDA? I’m sorry, but I don’t want to care about that. I want to do my stuff. NVIDIA got me into saying, “Hey, you should care about this,” because there was nothing else on the market. That’s not true, but ultimately, this is the GPU I have in my machine, so off I go.
If tomorrow that changes, why would I pay a 90% margin on my compute? That’s insane. This is why I believe it ultimately goes through the software, because the software is my entry point to the ecosystem. If the software abstracts away those idiosyncrasies, as it does on CPUs, then the providers will compete on specs and not on fake moats or circumstantial moats.
This is where I think the market is going. Of course, there’s the availability problem. If you piss off Jensen, you might need to kiss the ring to get back in line. Ultimately, I don’t see this as being sustainable.
Can I ask, when we chatted before, you said something about AMD? I said, “Hey, I bought NVIDIA and AMD. NVIDIA—thanks, Jensen—I’ve made a ton of money. With AMD, I think I’m up 1%, versus the 20% gain I’ve had on NVIDIA.” You said that AMD had basically sold everything to Microsoft and Meta and had a go-to-market problem. Can you unpack that for me?
All chipmakers have a go-to-market problem. It’s all of them, whether it’s Google, AMD, or Tenstorrent. There are probably 2 fundamental problems.
The number 1 problem is that maintaining multiple stacks today is very, very hard. Let’s say I buy AMD. That means I’m going to abandon NVIDIA. Then I think, “Oh, crap. I have a 6-year amortization plan on that. What do I do?” Do I need to support both stacks? Maybe, until AMD tells me, “Hey, you have, I don’t know, 1,000 NVIDIA GPUs, and you’re about to buy 100,000 AMD GPUs.” I mean, come on. Then I’m like, “Okay, that makes it worth my while.”
Ultimately, the fundamental problem is that the stakes are very high. I need to have a lot of incentives to buy into that ecosystem, so I need to buy a lot of them. If you’re AMD, that is already a problem. Then Microsoft comes along and buys it all, which, by the way, puts OpenAI—or at least the inference side of OpenAI—in the green because of the efficiency gains.
Can I just try to understand? Are you saying the switching costs are really high from one provider to another, or are you saying that to get into one of these buying processes, you have to buy so much that it prohibits you?
It’s actually both. The buy-in is very high, so to make it worth it, you have to buy a lot. If you buy a lot, this is what everybody asks. We talk to all of them, and they always have the same questions. It’s completely understandable. They say, “This is great, great, but who’s the customer?”
Take Amazon, for instance, with Trainium. Anthropic just came and said, “Hey, we’re going to buy 100,000 of them.” You want to buy 10,000 and feel like the big shot? But go back in the queue, because Anthropic is before you. They have to have very high commitments to make it worthwhile.
You cannot be incrementally better. It’s very hard. You have to convince a lot of people. I can give you 1 metric, if you want. I know for a fact that being 7 times better in whatever metric you want—whether it’s spend or whatever—is not enough to get people to switch. People will choose nothing over something. I have stories. This is a very hard market to enter because you cannot compete on incremental gains.
9. The Zero Buy-In Strategy
Maybe you can go the Middle East route, where they sprinkle everything around and evaluate everything, but that’s not a very sustainable strategy in the long term, or at least in the medium term.
What is the right sustainable strategy, then? You don’t want to go so heavy that you can never get out and have those switching costs, but you also don’t want to sprinkle it around. What’s the right approach?
The right approach, to me, is making the buy-in zero. If the buy-in is zero, you don’t worry about this. You just buy whatever is best today.
How do you do that?
10. Switching Between Compute Providers
By renting. This is our overall promise—at least, ZML’s thesis: if the buy-in is zero, then you completely unlock that value. When you say the buy-in is zero, what does that actually mean? It means that you can freely switch compute to compute. You just say, “Now it’s AMD,” and boom, it runs. You say, “Now it’s Tenstorrent,” and boom, it runs.
How do you do that? Do you have agreements with all the different providers?
Yeah, but not agreements. We work with them to support their chips. As a user myself of AI technology, my view is that if it’s free for me to switch or choose whichever compute provider I want—AMD, NVIDIA, whatever—then I can take whatever is best today and whatever is best tomorrow. I can run both. I can run 3 different platforms at the same time. I don’t care. I only run what’s good at the moment.
That unlocks a very cool thing for me, which is incremental improvement. If you’re 30% better, I’ll switch to you.
Are you taking the risk on that hardware, then? If you’re the one providing on-and-off, on-demand provisioning, you name it, who takes the risk?
This is a great question. I think that if you’re doing it bottom-up, from infrastructure to applications, you’ll lose, because nobody will care, as they don’t today. If you look at TPUs, they’re available and they’re great, but nobody cares.
Why does nobody care about TPUs?
Because the cost of buying in is always the same. You have to spend 6 months of engineering to switch to TPUs. Mind you, TPUs do training. They’re the only ones with training now, but AMD can do training. In terms of maturity, by far the most mature software and compute is TPUs, and then it’s NVIDIA.
The buy-in is so high that people say, “Well, we’ll see. I’m not on Google Cloud. I have to just sign up.” These are tremendous chips and tremendous assets, but people don’t want to make that commitment.
11. Microsoft's Strategy with AMD
In terms of the risk, I think if you want to do it, you have to do it top-down. You have to start with whatever it is you’re going to build and then permeate downwards into the infrastructure.
Take Microsoft with OpenAI, for example. They just bought all of AMD’s supply and run ChatGPT on it. That’s it. That puts them in the green. That’s actually what makes them profitable on inference, or at least means they’re not losing money.
I’m sorry, how does Microsoft buying all of AMD’s supply make them not lose money on inference? Help me understand that.
I can give you actual numbers. If you run 8 H100s, you can put 2 70B models on them because of the RAM. That’s number 1. Number 2, if you go from 1 GPU to 2, you don’t get twice the performance. Maybe you get 10% better performance.
That’s the dirty secret nobody talks about on inference. You go from, let’s say, 100 to 110 by doubling the number of GPUs. That is insane. Would you rather have 2 machines with 1 GPU each than 1 machine with 2 GPUs?
With 1 machine of 8 H100s, you can run 2 70B models if you do 4 GPUs and 4 GPUs. If you run on AMD, there’s enough memory inside the GPU to run 1 model per GPU, so you get 8 GPUs’ worth of throughput. On the other hand, you get AMD GPUs at 2 to 2.5 times the throughput. That’s a 4× right there, just by virtue of this.
That’s the compute part. If you look at all of these things, there are tremendous amounts of memory. We talk to companies whose chips are coming with almost 300 GB of memory on them. That means 1 chip per model, which is the best thing you want if you’re running 70B models.
It’s not state of the art, but it’s the regular stuff people will use for serving. If you look top-down at what you’re going to build with them, it’s a lot better to go after the efficiency gains, because 4× is a big deal. These chips are 30% cheaper than NVIDIA’s. It’s a no-brainer.
But if you go bottom-up and say, “I’m going to rent them out,” nobody will rent them. That’s why I think it’s a good way to attack it from the software, because ultimately, do you really care whether your MacBook is an M2 or an M3? You just say, “Oh, it’s the better one,” and that’s it. Imagine if you had to care about these things. That would be insane.
When I listen to you now, I’m thinking, “I should sell my NVIDIA and buy more AMD.” If you were forced to buy 1—I’m not saying sell the other, I’m saying buy 1—which would you buy, and why? I mean the stock.
The thing is, I’m long. I used to think the market was efficient, so this is not investment advice. Probably, today, I would still go with NVIDIA because of the supply. If we play our cards right and ship our stuff, hopefully I’ll come back and tell you to buy as much AMD as you can—or Tenstorrent, if they go public—or whoever else. These chips are amazing, by the way.
What does everyone think they know about inference that they actually don’t?
Probably not a lot of people are accustomed to what it entails to run production. Inference is production, and production means somebody has to wake up at night. I used to be that guy. I don’t want to do it again.
Production is hard. Thankfully, we have a lot of software nowadays to do that a lot better, but there’s not a lot of reuse because the AI field, at least, isn’t really accustomed to that yet. It’s changing, but the discussions I had 1 year ago and the discussions I have today are not the same. They’re going in the right direction, but they’re not there exactly yet.
That would probably be the number 1 thing. Inference isn’t just training code running with only the forward pass. That’s not what it is.
12. Data Center Investments and Training
Can I ask how you evaluate the data center investment we’re seeing? When you look at Meta doing $60 billion to $65 billion, Microsoft doing $80 billion, and some of the intense capital expenditure you’re seeing, how do you think about that on the data center side?
They’re still going after training, so there’s still this frontier. It’s probably also why NVIDIA is the better buy right now. On the NVIDIA side, if you do training, it’s incremental. If you’ve bought 1,000 NVIDIA GPUs and buy 1,000 new NVIDIA GPUs, that gives you 2,000 GPUs. But if you buy 1,000 NVIDIA and 1,000 AMD, that gives you twice 1,000. It’s a bit different.
They’re still going after training, definitely, and they’re very pragmatic in doing so. They have the capital expenditure to spend. They’re not making their money out of it.
Google is probably the only one, by the way, that owns its compute. There’s this triangle of winning—that’s my mental model. You have the product, the data, and the compute. Who has all 3?
Google. Amazon doesn’t have products. They have Amazon and AWS, but they don’t have actual products. Google has Android, Google Docs, and whatever else. They have everything, and they can sprinkle it everywhere. This is the sleeping giant in my mind. If they’re not busy doing a reorganization, they might—
It’s fascinating, because if you’re a shallow thinker, you think OpenAI challenges the golden goose, which is search, and that Google is threatened more than ever now. OpenAI is amazing, but it’s not their compute. It’s Microsoft’s compute. If you own your compute, you own your margin, is essentially what you’re saying.
Yes. Even Microsoft, when they were running on NVIDIA, bought NVIDIA at outrageous margins. I talk to a lot of people who build data centers. These people buy tens of thousands of GPUs, and I ask them, “Do you at least get a discount or something?” They say, “No. The only thing we get is access to supply.”
13. How to Succeed in AI: The Triangle of Products, Data, and Compute
Ultimately, if you don’t own your compute, you’re starting with something around your ankle. This is why I like to think in this triptych, or at least this triangle: product, data, and compute. You can see where everybody sits and their weaknesses and strengths.
14. Scaling Laws and Model Efficiency
Can I ask you, if we move a little bit—you said it’s totally rational that everyone’s focusing on training still. When we think about that, it’s rational if you think that efficiency and scaling laws continue to place such emphasis on it. How do you think about model scaling and scaling methods coming into play?
There’s a brute-force approach to this. It’s a very American approach: more and more and more. But you look at, for instance, the xAI cluster, and it’s not 100,000 GPUs. It’s 4 times 25,000.
You’re starting to see that because of InfiniBand and, in the case of likely RoCE, the technology they use to bridge their GPUs together, you have a network bound. At some point, you’re fighting physics. You can push, but it’s like trying to get to the speed of light. As you approach it, the amount of energy you need gets higher and higher, and it grows and grows.
There are 2 approaches to that—or, sorry, 2 counters to that. The first is that we still scale, but there’s a lot of waste and excess spending on the engineering side. That’s the DeepSeek approach, which was very successful. They said, “If we do this and this differently, we get multiples sometimes.” Virtually, you increase your compute capacity because you’re more efficient.
15. Future of AI Models and Architectures
The other approach is likely Yann LeCun’s approach, which is that this is not scaling, and at some point we need to look the problem in the face and do something better. Of course, we push and push and push because there’s still capital, but between these 2 approaches, I think you can do more with less.
At what point do we stop and say, “There’s a lot of wastage, and we could do this better”? How far away are we from that?
Until somebody does it. DeepSeek was a good wake-up call. Suddenly, efficiency is in. That’s number 1.
Number 2 is until there’s a new architecture that comes out and changes the game. In the case of LLMs, for instance, you have what are called non-transformer models, which fundamentally change the compute requirements. That might be a frontier that completely obsoletes the transformers.
The transformers are the building blocks by which current models work. The way they work is that for each token, or syllable, if you will, the model looks at everything behind it. You can see that as you add more text, you have more work to do.
There are new architectures that don’t require this. That might change these things and probably shift the amount of compute needed to do training or inference. Then there’s the new thing, which is likely Yann LeCun’s thesis: the world model. LLMs are, at the end of the day, what we need is something that fundamentally understands the world.
What is it called?
It’s likely JEPA. I’m very bullish on this, but it’s very frontier.
Why are you bullish on it, and why is it so frontier?
It’s hard to explain. He explained to me how it worked, and I was blown away. It’s as simple as this, but it makes a lot of sense.
We’re creeped out because the machine talks back to us. That’s it. It’s not a new thing. When it exploded, it was new technology, but suddenly it was talking back, and that freaked us out, so we got crazy about it. Language is 1 form of communication, but it’s ultimately a very narrow window into the world.
We use language to describe the world, arguably with some loss. The JEPA approach, long story short, is that you have essentially 2 things you want to do, and you try to minimize the energy to do them. From this, understanding emerges. Physics emerges, and so on, because you’re trying to minimize the amount of energy to go from 1 state to another.
That actually makes sense. If you try to pick up this AirPods case, you’re not going to take a round trip around the block to get it. You just get it. In my brain, I’m wired to do the thing. If I go and talk to myself out loud—put my hand down, move to the left, and whatever—that feels very inefficient.
This will probably be something that changes. In the case of LLMs, there’s also good work on what are called diffusion-based LLMs. Instead of thinking autoregressively—that means you get a new token, reinject it, redo it, and so on—they think more like we do, which is in patches. Imagine a paragraph of text, and words appear until it’s done.
Wasn’t that distillation what DeepSeek did on OpenAI models, basically copying them?
Oh, distillation. They did distillation on OpenAI models.
If we’re all progressively moving toward a better future for humanity and more efficient models, is distillation not effectively open source in another wrapper?
I think it’s fair game, to be honest. There were some people who tried to ask—I don’t remember if it was an OpenAI model—but they asked it to generate an image from a Star Wars movie at a particular timestamp, and it came out with a screenshot from the Star Wars movie. Obviously, it was trained with it.
I think it’s fair game because there’s no free lunch. It was trained with data. You had a good ride, somebody was sneaky and took it, but you took it from the beginning, too. Let’s just accept that it’s fair game.
You also learn from their advancements.
Absolutely. I take my cup and enjoy it very much, every single day.
You mentioned training there. Obviously, data quality dictates a lot of training ability. When you think about the future of the data that feeds into training, how do you think about the balance between synthetic data and real data?
I’m a bit split on this. There’s a part of me that says if you reinject data into the system, the system deteriorates. That feels intuitive. But if you look at AlphaGo, for instance, the moment it ramped up its skills was when they started generating synthetic games.
I’m a bit split, but there are some verticals that benefit very much from this. Code LLMs, for instance—we can run code. That’s the Poolside thesis.
Why does it work for coding and not for other things?
You don’t use the AI model to generate output. You use the machine. You run the code, see what it makes, run all this code, and create data out of it.
Whereas if you run an LLM and say, “Generate me 2 trillion tokens of text,” it will do it with its own output. You may reinject that data and stuff, so there are a lot of tricks. Ultimately, my gut tells me it feels wrong because you reinject data that was already there, so it will deteriorate. There’s loss.
I’m a bit bullish, but I’m not sure exactly on which verticals. Code is one. We’ll see. Distillation is, in some sense, a bit like that. You distill and create synthetic data from a bigger model into a smaller one.
The most mind-blowing thing about distillation is that sometimes the smaller models become better than the bigger models through distillation.
I love that smaller models become better than bigger models purely because of the quality of the data that’s input into them.
One theory is that the smarter model is better at generating output that you would want it to generate. It’s not better in the general sense. It’s better at the task at which you were measuring it, because that’s what it learned to imitate.
How do you think about the future in terms of large, monolithic models versus more dynamic architectures?
Sometimes it’s wasteful to run big models. A lot of the time, it’s actually wasteful to run big models. I think there are going to be a lot of smaller models for efficiency reasons. Less is better.
There’s a but, though. You talk to people at DeepMind, and they don’t even fine-tune anymore because they have such a big context window, which is the data you inject into the model at runtime. Nowadays, they just dump data into it and say, “Do whatever that data tells you to do,” instead of fine-tuning as we used to do.
If the efficiency gains aren’t there yet, then we’re not there yet. But if the efficiency gains pass that threshold, we’ll just do it at runtime. We’ll have 1 great model that specializes itself for each request. That’s not for tomorrow, I think.
What is retrieval-augmented generation?
First, it’s a very clever trick. What you do is represent knowledge in what’s called vector space, or latent space. Imagine you have a 3D space that represents all knowledge, all of everything. A cat sits here, and a dog sits close because it’s an animal, but it’s far from some other property, and so on.
You run the user’s request through this same system. It’s called an embedding, and that gives you a vector. You take whatever is closer to you—that’s what’s called semantically close. Then you insert those pieces of text before the request.
It’s as if you said, “Knowing the following”—and you give the data, let’s say it’s law or whatever—“please answer my request.” That’s it. It’s a clever trick.
It’s a bit dirty because you’re limited by the amount of data you can input. There’s a problem around how you chunk the data that you input.
In a lot of the things we do, when we say, “Here’s a link; summarize it to the key points,” is that not retrieval-augmented generation? We’re inputting the data, and then it’s doing that.
Sometimes it is, yes. It depends on how it works. Think of it as a preamble to your question: “Knowing the following,” and the following is a tiny window into the content, “please answer my question.” Of course, as you talk more and more, it will forget, because that window is fixed.
How does that shift the movement from large, generalized models to smaller, more advanced models?
What pushes smaller models is efficiency—roughly, speed. Less is better. If we can do the same thing with less, then less it is. It’s as simple as that.
In terms of RAG, the key frontier is what we call attention-level search. This is something we’re working on. You have the exclusive now; I’m putting it out there. It doesn’t push model sizes. What really pushes model sizes is efficiency rather than specialization.
So, if you can get the same performance with a smaller model that’s fine-tuned with RAG or whatever, you’ll use the smaller one because, again, less is better.
Before we move into a quick-fire round, I do want to ask you: when we had DeepSeek, to what extent were you surprised that such innovation, I would argue—and I think many would agree with me—came from a Chinese competitor and not from a Western competitor?
I love it. Constraint is the mother of innovation. We can troll a bit about the Singapore gray market and all of these things, but ultimately, they had no choice.
Here’s the thing: if you can buy more, why would you give a shit? You can just buy more. If you’re pushed toward efficiency, then you will deliver efficiency. These are very, very skilled people.
The coolest thing to me about AI is that geography doesn’t matter anymore. You can just do things, appear out of nowhere, and, boom, you’re on the map. I’m very glad that they did it. I found the reaction very entertaining, to be honest. Constraint is a very good driver of efficiency.
16. Why OpenAI’s Position is Not as Strong as People Think
Do you think it’s a meaningful threat to OpenAI and ChatGPT? Bluntly, they still have the consumer loyalty and the consumer brand. To what extent is it actually a long-term threat?
I’m not sure who’s a threat to OpenAI at the moment. Here’s why: you look at the numbers, but we live in a bubble. We follow every new episode, every new model, who said what, and so on.
I go to my mother and ask her, “Do you know ChatGPT?” She says yes. Then I ask, “Do you know some other model?” I don’t want to dunk on anybody, but she says, “What is it?” Even Gemini—Google, right? They have a strong brand and a strong product.
Arguably, there’s a balance between the product and the models. Gary from Fluidstack told me that his mental model for model providers is that they’ll be like carmakers. There’s no winner-take-all. Everybody will have their own because, ultimately, human knowledge is—everybody has everything. We’re converging, but I like that analogy.
DeepSeek made very good waves, but the waves were amplified by the media, the narrative, and the drama.
Do you think export regulations inhibit China’s ability to compete in any way today, or tomorrow?
Maybe tomorrow. I’m not sure. They’re a bit late in terms of ASICs. They’re at the A100 level, but one of their unfair advantages is that they’re exercising under constraint. It’s like when you exercise in water. That’s their state.
They’re bound to do better. They just can’t buy their way into better compute. I think it hinders their success, but I think it’s short-term to think that way.
Are you fearful that Europe is going to regulate itself into constraints in a world of AI?
No, I don’t care. I have ZML. This makes me wonder sometimes. I understand the narrative and so on, but I’m absolutely not fearful. Let’s be successful first, and then we’ll talk about the politics.
I’m not Mistral. I’m not building gigawatt data centers and so on. If you build gigawatt data centers, you run into these problems. But if you’re successful, everything flows from there.
Steeve, I’m being direct here, but I’m asking you for the pros. Everyone says Mistral just doesn’t have enough money to compete. That’s the word on the street. To what extent is that fair?
They’re very competent. I don’t know. It’s easy to spread FUD. There’s a lot of FUD going around, especially about regulation and everything.
Here’s the thing: I look around me, and I don’t see what I read. I’m hardly convinced. Everybody was saying they were dead, and then, boom, they came out with their release, and it was insane.
I don’t know. What I know is that I hope they don’t have too much money, that’s for sure. You want to be clear-eyed, right?
Final one before we do a quick-fire round. Stargate was a $500 billion announcement. How did you evaluate that?
My first impression was that I didn’t buy it. It’s American style: you start with the claim and figure it out later. I don’t buy it. Ultimately, I’m not sure I care that much about it.
Let’s imagine it’s true. I don’t know if it is or not, to be clear, but let’s imagine it’s true. Congratulations, amazing. But it’s more of the same. It’s vertical scaling.
As you know, my days are spent on efficiency, so I look at these things and think, “All right, this is a bigger American car of AI. It’s big and consumes a lot of gas, but ultimately it’s not a good car,” or at least not to my liking.
There has to be sufficient capital, but at some point, I’m not sure it’s really a differentiator. That was my thesis before DeepSeek came, and it’s still my thesis. You need money, infrastructure, and so on, but the 2 limiting factors today are probably talent and energy. That’s it.
The rest—you can buy $500 billion worth of GPUs, by the way, at a 90% margin. If we work on that margin, we can shrink that number. That’s my view of the world. I’m probably wrong, but I’m not easily entertained by these numbers. I’ve seen how the sausage is made way too many times.
I want to do a quick-fire with you. I’ll say a short statement, and you give me your immediate thoughts. If you had to bet on 1 major shift in AI infrastructure over the next 5 years, what would it be?
Latency in reasoning, definitely this year. The shift from throughput to latency—how quickly my answer starts to appear and how long it takes for my answer to complete—is probably 1 of the fundamental shifts this year.
Longer term, I’m really rooting for non-transformer models, which will also change the compute landscape, and, of course, world models and energy-based models.
What’s 1 piece of advice you’d give to AI startups navigating the changing landscape of training, inference, and hardware?
The number 1 thing I’d say is: do not resell compute if you can avoid it. A lot of AI startups building on top of AI are trying to make a margin on top of a very big cake, and ultimately what they sell is compute.
If you look at a dollar of spend, maybe 98% of it goes to somebody else’s margin. If you do AI, try to verticalize on the product as much as you can, but not on the compute. If your business model implies buying a lot of tokens, it’s a very hard circle to square to put that into $20 a month. Look at it from that angle, and if you can, try to avoid it.
17. Challenges in AI Hardware Supply
What’s the biggest challenge that Jensen Huang faces today?
The highs are very high, but they don’t last forever. We’re seeing it, actually. It’s probably how to navigate the downslope. Blackwell is probably something that keeps him awake at night.
Why would that keep him awake at night? Would that not reenergize him—more orders, new enthusiasm, a new product?
Because orders are getting canceled. They have a lot of problems with these chips, so a lot of people are canceling their orders. These chips are on the frontier of scaling, and they were supposed to come out last summer, but they have heat-dissipation and motherboard-bending problems.
The people who are very pre-silicon told me, “This is what we call a pretty big problem,” end quote. Probably navigating the downslope is the challenge.
You may not know this, but the supply of H100s was actually smoothed out over the year so that they didn’t have a big spike in deliveries followed by a quarter with fewer deliveries, which pissed a lot of people off—especially people who bought a lot of them. Some still haven’t received their orders from last year, and they’re already seeing the new chip, the B200, and then the one after that. They’re super pissed.
I think navigating the downslope will be important at some point. The question is when and how. If there’s an H100 bubble, of course it will impact NVIDIA. I’m probably going to get a lot of flak for this, but I’ve seen some very worrying numbers about Blackwell and varying testimony from people who operate these things. I don’t know. Maybe that train will stop, or at least slow down.
Steeve, I’m not sure I’ve ever learned quite as much in 1 episode. I love what I do because I’m able to ask anything to the smartest people in their businesses, and I so appreciate you unpacking so much of it for me today. I’m thrilled to say that I finally get what you do after speaking with you. You’ve been a star. Thank you.
Thank you. I appreciate it.