DeepSeek 周落幕:AI 的 Moneyball、算力需求的未来、地缘政治现实检验及更多
与其把 DeepSeek 看作算力不再重要的证明,不如把它理解为 AI 领域的 Moneyball。 600万美元只覆盖了一次最终的 V3 训练,并不包括全部实验、数据采集或研发;而公开披露的优化方案是“送给美国的一份巨大礼物”。Thompson 的类比是:奥克兰运动家找到了战术,但美国实验室可以凭借大得多的预算,变成红袜队。
DeepSeek 实现了真正的效率突破,但没有超越美国的模型前沿。 蒸馏几乎不可能被阻止——限流只会把它变成类似盗版的猫鼠游戏——而 OpenAI 自己也靠蒸馏让 ChatGPT 具备经济可行性。几个月后达到 GPT-4o、Sonnet 或 o1 级别,意味着“按定义你并没有领先”:这是“效率突破”,不是能力突破。
算力充裕让美国软件文化忽视了优化能力,而限制迫使 DeepSeek 锻炼了这块肌肉。 它在混合专家训练、专家均衡方法和手写通信代码上的工作证明,“削弱限制确实奏效了”。Thompson 承认自己过去对优化存在偏见,并主张双轨制:一支团队不受约束地追求能力,另一支团队专门思考效率。
转化性 AI 产品迟迟未现,更像正常的采用滞后,而不是技术已经失败的证明。 ChatGPT 直到2022年11月才上线;互联网用了大约13年,Facebook 2006年的信息流才找到原生商业化形态。聊天客户端已经带来“大量尚未被挖掘的效率提升”,但采用仍受用户想象力限制;Sharp 原本预计 Apple Intelligence 能让 AI“变得傻瓜式易用”,结果却是“一场彻底的灾难”。
更便宜的模型可能增加而不是摧毁整体算力需求,因为推理模型能够制造训练数据。 推理模型生成思维链,这些输出训练更便宜的基础 LLM,改进后的基础模型再生成更强的推理模型,循环往复:“是 AI 在让 AI 变得更好。”Thompson 对方向的判断是绝对的,但没有判断时间点:这个时代可能需要“基本无限的算力”,包括专门用于训练的推理算力。
Microsoft 现在拥有颇具吸引力的 AI 期权,而 OpenAI 和 SoftBank 正在押注一场风险投资规模的前沿赌局。 Microsoft 保留 OpenAI API 的访问权和算力优先购买权,同时减少对前沿模型的融资义务——这些模型可能在几个月内商品化。Thompson 认为另一种结果的概率或许只有5%,但上行空间大到天文数字;这对股权合理,对债务不合理,而目前证据正倾向于商品化。
DeepSeek 暴露了美国芯片管制背后的地缘政治赌注:要维持持久的 AI 领先,美国必须弥补自身的工业弱点。 Thompson 仍然“陷在一堆泥里”,因为 EUV 是一个可执行的关键瓶颈,但切断中国获取芯片可能加速其国产替代、削弱 TSMC 的威慑价值,并鼓励美国守护过去的创新,而不是继续向前推进。这套政策最终与 OpenAI 共享同一个起飞前提:如果模型扩散并商品化,那么看似不断复利的技术优势可能无法持久。
1. DeepSeek 的 Moneyball 战术是送给富有竞争对手的礼物
Thompson 认可奥克兰运动家的类比:DeepSeek 找到了更富有的美国实验室可以用“巨额薪资预算”复制的战术,就像红袜队吸收 Moneyball 方法一样。更多算力仍然有用;认为算力突然变得一文不值,是“荒谬的”。
他从公开披露本身看出了地缘政治信号。如果 DeepSeek 的唯一目标是削弱美国技术,“聪明的做法”应该是把优化方案藏起来;公开发布既支撑了论文的整体可信度,也让整个行业都能复制这项工作。
Sharp 认为,叙事之所以被扭曲,是因为两组数字形成了极具诱惑力的反差:Sam Altman 一周前宣布 Stargate 需要5000亿美元,随后 DeepSeek 带着600万美元的数字出现。但这个数字描述的是一次最终训练,并不是背后的全部实验、数据采集或研发成本。
对投资者而言,关键区别在于时间。算力最终可能“填满所有可用空间”,但效率提升可以推迟需求消化、改变推理经济性,并让支付算力容量的公司受益,即便供应商仍将面对巨大的长期需求。
2. 蒸馏压缩的是成本,不是前沿差距
当被问到实验室能否阻止蒸馏时,Thompson 的回答是“基本不能”。供应商可以限制 API、设置速率上限,但数字输入和输出让执行看起来像打击盗版;意志坚定的运营者甚至可以编写脚本,让普通聊天客户端自动采集答案。
蒸馏也是前沿实验室以经济方式服务用户的手段。OpenAI 最初面对压倒性的 ChatGPT 需求,却要提供成本高昂的完整模型,于是现实压力推动了让其以更低成本服务数亿用户的突破:“这不是好事,也不是坏事。”
Thompson 保留了能力层面的区分:在发布几个月后达到 GPT-4o、Sonnet 或 o1 的水平,并不意味着 DeepSeek 领先。“效率突破是真实的,但它们是效率突破,不是能力突破。”
Sharp 拒绝美国式否认,也拒绝衰落主义恐慌。他的中间判断是:DeepSeek 获得了充足资金,使用了大量 NVIDIA 硬件,很可能为 V3 和 R1 蒸馏了 OpenAI 的输出,并贡献了令人印象深刻的创新——但这并不意味着一个600万美元的“支线任务”就终结了美国的领先地位。
3. 芯片限制迫使 DeepSeek 真正“贴近硬件编程”
一位听众把 DeepSeek 与主机游戏开发者类比:后者能从老旧的 Xbox 和 PlayStation 硬件中榨取出异常出色的画面。Thompson 接受这一机制的核心——静态限制会奖励深度优化——但指出,一旦资产制作成本超过引擎开发成本,现代游戏经济就会推动引擎转向跨平台抽象。
PS3 提供了反例:困难的硬件环境迫使开发者先为 Xbox 开发,再进行移植。Sony 的应对是收购工作室、锁定独占内容,这说明优化选择最终会重塑平台战略,而不会只停留在纯技术问题上。
DeepSeek 在混合专家方面的工作货真价实。它找到了让每个专家训练更高效的方法,把均衡逻辑移入专家内部的偏置因子,而不是集中计算所有内容——这挑战了混合专家收益主要来自推理阶段的假设。
工程师还手动编程计算单元,以便在受限带宽下改善通信。Thompson 的结论保留了两面性:DeepSeek 仍落后前沿模型数月,但“削弱限制确实奏效了”,而且工程工作“真的很惊人”。
4. 优化债务让 Google 的基础设施重新变得重要
Thompson 接受了自己过去低估优化的批评:“在这周之前,我从没这样说过”,此前他经常主张相反的观点。理性地遵循比较优势,可能会放弃一条学习曲线,而这项损失要到多年后才显现——这与他认为美国最终失去制造能力的过程相似。
Google 最初的基础设施战略提供了范本:用廉价的 x86 硬件替代定制 Sun 服务器,假设组件会故障,再围绕故障构建有韧性的软件。结果是更便宜、更具扩展性的基础设施,并成为 Google 成功的基础。
Thompson 现在支持双轨制:一支团队不受优化约束,另一支团队只思考优化。Gemini 1.5 的百万 token 上下文窗口,可能正体现了 Google 在基础设施、芯片和软件上的整合能力;“说到服务全世界,Google 仍然是最好的。”
但这种优势也有颠覆性风险。如果模型商品化,Google 可以以更低成本提供服务并压低竞争对手价格,但最容易受到低成本 AI 产品冲击的实体,可能正是 Google Search 本身。
5. AI 的产品空窗期只有两年
Thompson 的答案不令人满意,却有历史依据:基础性技术“总是比你想象的更久才发挥作用”。Transformer 于2017年出现,GPT-3 长期以一种未被充分使用的方式摆在人们眼前,而 ChatGPT 直到2022年11月才真正敲响唤醒警钟。
早期产品通常会把新技术硬接到旧有形态上。在网页文本旁边放上报纸式展示广告,并没有释放互联网经济的潜力;Facebook 的信息流于2006年出现,大约在早期互联网诞生13年后,创造了真正原生的商业结构。
现有聊天客户端已经非常强大,但其有用程度“取决于你愿意并且有能力想出可以让它做的事情”。Thompson 举的日常例子是让 ChatGPT 把一篇文章里的每个 H4 Markdown 标题都改成粗体,从而省去大约20分钟的烦躁操作。
Sharp 的反驳指向分发,而不是能力:他原本预计 Apple Intelligence 能让 LLM 变得毫不费力、无处不在,但称其为“一场彻底的灾难”。最终突破可能仍会遵循经典的创业者故事:有人解决自己的问题,把它产品化,并消除用户主动提出需求的必要。
6. OpenAI 的安全话术与其封闭策略发生冲突
一位听众纠正了 Thompson 对历史的描述:相较于 Google 的 DeepMind,OpenAI 确实更开放;模型卡片和 CBRN 风险分类也为负责任发布创造了有用语言。Thompson 立即承认:“说得好”,并基本接受了对方关于商业逻辑的全部论点。
对领先者而言,封闭权重是理性的;对 Facebook 和 DeepSeek 而言,开放权重则是一种战略筹码。DeepSeek 还直接受益于把技术分发到面临芯片约束的中国生态中,接收方可以直接使用权重,而不必蒸馏输出。
Thompson 剩下的异议在于话语范围。当“安全”从灾难性风险扩展到错误信息或偏见时,批评后两者的人可能会被指责为无视人类灭绝风险——这是一种“高地城堡与低地城堡”的话术,在他的经验中,它会压制正当的分歧。
治理结构进一步强化了这一担忧:如果一家实验室相信 AI 将控制世界,那么其非营利董事会结构实际上是在说,“将由我们控制世界”。Thompson 更偏好“更多 AI,而不是更少 AI”;既然已经跨过卢比孔河,扩散就不可避免,那么一家以安全为先的非营利机构不应把封闭模型与“自以为是的姿态”绑定在一起。
7. Microsoft 持有期权,OpenAI 买入登月赌注
在 Sam Altman 与 Satya Nadella 和解自拍之后,Thompson 认为 Microsoft 处于令人羡慕的位置。它保留完整的 OpenAI API 访问权和算力优先购买权,同时降低了为可能在数月内商品化的前沿模型提供融资的义务。
Sharp 质疑,在蒸馏没有持久防御能力的情况下,OpenAI 的战略是否仍然合理。Thompson 用概率回答:出现一个主导模型的概率或许只有5%,但其上行空间大到足以支撑一场风险投资式的股权押注。
这种收益结构很难用债务融资解释;股权投资本质上是一场风险投资式赌局,其上行空间足以覆盖风险。Microsoft 可以拒绝承担同样的风险,同时不切断对 OpenAI 模型的访问。
尽管如此,Thompson 认为证据正在向模型商品化方向移动。他对人物的判断则提供了剩余的治理风险:Altman 有一种“伊卡洛斯式的癖好”,可能总是飞得离太阳太近。
8. 合成数据把效率转化为更多算力需求
DeepSeek 使用2048块 GPU 进行的 V3 训练,并不代表其全部运作规模。Thompson 引用了拥有50000块 Hopper GPU 的报道——H800 也属于 Hopper——并认为受限的内存带宽可能限制了同步集群,而不是限制了公司对更多硬件的需求。
最终训练也不是主要消耗来源:研究人员需要进行大量实验,而 DeepSeek 令人震惊的低价推理已经在中国“彻底摧毁了定价模型”。大型数据中心越来越多是为了服务模型使用,而不只是进行一次性训练。
R1 更深层的突破在于反馈循环:推理模型生成思维链,实际上制造了近乎无限的合成数据;这些数据改进廉价的普通 LLM;更强的基础模型随后又支持更强的推理模型。“我们已经进入了这个良性循环。”
Thompson 称这正是“苦涩的教训”在发挥作用——通过蛮力生成并获取所有可能的答案。效率提升可能推迟近期部分推理采购,但训练、推理以及为训练服务的推理合在一起,会创造“比以往更大的”算力需求。
9. 算力增长会同时拉动电力与网络需求
更强大的集群仍然需要 NVIDIA 的系统优势。Thompson 指出,单颗 AMD 芯片可能更快,但 NVIDIA 在通过 CUDA 编程以及连接加速器方面仍然“好上几个数量级”,不会让通信开销摧毁整个集群。
因此,即便单次训练成本下降,网络、电力分配和液冷需求仍会保留。推理工作负载需要持续生成内容,而合成输出又会成为下一轮训练的输入。
Thompson 表示,Trump 政府正准备一项行政命令,推动建设与数据中心一一对应的专用离网电力,同时大幅放松监管。可预测的负载可以绕开电网瓶颈,再通过光纤把结果传输到其他地方。
核电是自然选择,因为数据中心需要持续供电;太阳能也可以部署在偏远地点,但必须配备大量电池以覆盖夜间。贯穿始终的结论是:DeepSeek 没有关闭基础设施机会,反而可能“真正解锁了”更多需求。
10. 出口管制押注美国的工业未来取决于 AI 起飞
Thompson 认为 NVIDIA“并不像一家美国公司那样运作”:它把稀缺芯片分配给 CoreWeave 等初创企业,投资部分客户,限制超大规模云厂商的议价能力,同时向中国销售芯片,而 Amazon 或 Microsoft 可能本来会买得更多。“NVIDIA 的手并不干净。”
但他对出口管制的立场仍然刻意保持未决。EUV 是一个清晰且可执行的关键瓶颈,其商业化用了大约10—13年,并多次依赖 ASML 的救援;更广泛的 DUV 和芯片限制反而可能损害美国供应商、加速中国替代,并耗尽“一张只能打一次的牌”。
对 TSMC 的依赖同样具有威慑冲突的作用:如果中国依赖台湾制造芯片,入侵就会承担额外代价;切断这种获取渠道会改变战略计算。因此 Thompson 担心,完全有效的管制可能比存在漏洞的管制更不稳定——漏洞或许起到了“压力阀”的作用。
他坦率的结论是:“我陷在一堆泥里。”未解决的交换关系是:究竟要在今天减缓中国的能力,还是让美国学会守护昨天的领先地位,而不是重建自身的创新与制造能力。
11. DeepSeek 刺破了美国的地缘政治自信
Sharp 提出的 Mistral 反事实揭示了情绪上的扭曲:如果完成这项工作的实验室来自法国,西方观察者可能会庆祝;因为它来自中国,同样的结果却同时引发否认与恐慌,也许还带来一剂有益的“谦逊注射”。
Thompson 的军事现实校验是物理层面的:中国能够造船,并主导无人机、机器人、电机、执行器和电池等零部件,而美国甚至难以扩大火炮产能。“打仗不是聚合用户、捕获需求。”
台湾暴露了其中的矛盾。台湾的先进制程晶圆厂支撑着美国安全,但留在岛内本身也是台湾的筹码;Thompson 认为,无论中国夺取台湾,还是战争摧毁 TSMC,美国都仍需要本土产能。他更倾向于直接为 Intel 提供采购保障,而不是单纯补贴供给;关税则是效率更低的需求工具。
芯片禁令最终与 Altman 和 Dario Amodei 共享同一个前提:AI 可能起飞、自我改进,并维持足以克服制造业弱点的复利式美国领先。DeepSeek 之所以重要,是因为如果模型最终扩散并商品化,那么许多战略决策都建立在一个“可能并不成立”的假设之上。
Hello, and welcome back to another episode of Sharp Tech. I'm Andrew Sharp, and on the other line is Ben Thompson. Ben, how are you doing?
I'm doing well, Andrew. How are you?
I'm doing all right. Happy Year of the Snake. It's New Year's over there. Are you a snake guy?
Not a big snake guy.
Okay.
Not a big fan. I am very happy I was born in the Year of the Monkey. Great animal.
Mm.
You find your blessings in all sorts of little ways, and God bless 1980.
There you have it. That was a test, and you passed. Not a snake guy, so we can continue the podcast for another 200 episodes here. Looking forward to that. And I'm a monkey guy also. I had a stuffed monkey growing up, so I always had an affinity for monkeys.
In any event, as for the show, we recorded our first episode this week on Sunday night in America, about 11 hours before a truly wild Monday morning on Wall Street. In the intervening days, you've written about DeepSeek on multiple occasions, and we've gotten a flood of great DeepSeek emails. So the plan for this episode is to bounce around and hit as many as we can. Does that sound good to you?
Yes. I'm going to try to keep my comments short and brief—
Mm.
—so we can get through all of them. It's something I'm very good at.
It's good to have goals. Billy says, “Is there anything to the analogy that DeepSeek is the Moneyball Oakland A's, having created new, efficient tactics that the rest of the industry will now take as standard? If that's the case, what happens when the American tech companies utilize these tactics with infinitely more resources, like the Red Sox did for the next decade after that Moneyball A's team? Do we yet know if the richer companies can add layers of AI-backed RL that would require more resources but ultimately produce better models?” What do you think, Ben? Grade Billy's analogy here.
I like the analogy. It's a good one. The part of the reaction that was definitely overwrought is this: Would DeepSeek like to have more computing? Yes. Do they have more computing in some respects? Yes. This idea that computing is no longer valuable is ridiculous.
I think there are fair questions. You mentioned the Wall Street angle. Everyone who's saying, “Why are these stocks down? AI is going to take over everything. Computing is going to fill all available space,” is correct. The challenge with stocks is that there can be a timing issue. At what point does all that computing get absorbed? Does that extend the timeline? Does that change what matters for inference, things along those lines? It's always difficult to say what's happening on Wall Street.
Mm-hmm.
But there are rational reasons to simultaneously believe this analogy is a good one, that more computing is useful, and yet also understand why some prices might be down and why some might be up.
Yeah.
Particularly the companies that are paying for all of this, right? The idea that the Boston Red Sox adopted the A's philosophy and paired it with a huge payroll is a great analogy. That's why, at the end of the day, I was emphasizing again and again that this was a tremendous gift to the US—
Yeah.
—and to the industry generally, to uncover these optimizations and these ways to be more efficient, because ultimately it's going to benefit everyone.
Frankly, I'm still pretty skeptical of the “this is all a lie, this is all a thing to just take down big tech” argument. The “smart thing” to do—quote-unquote, smart thing—if your number-one concern is geopolitical concerns would be not to disclose this at all.
Mm-hmm.
Just keep this knowledge internal. And the fact that it was shared, I think, actually buttresses the idea that the details are broadly correct. Again, when we're talking about the money involved, it was for 1 training run, not for all the experiments, not for all the data collection, and not for all the R&D. The paper was very clear about what the money was for.
Right.
It's gone through this filter, this game of telephone online—
It's gotten pretty warped.
It's been interesting to watch.
I think it's because people want to believe that story. It's a compelling counterpoint to Sam Altman announcing, a week earlier, “We need $500 billion for Stargate.” So it doesn't surprise me that certain people just took the $6 million number and ran with it, and it got out of control and twisted into something that it wasn't.
Right. No, for sure. Dario Amodei, the CEO of Anthropic—I know we're going to get to it—but he did mention, “Look, this is broadly in line with the cost reductions we've seen. Its capability is similar to what we did a year ago.” Of course, it's in his self-interest to say that and argue that, but I also think it's reasonable.
The whole thing about “They stole our data,” which is referring to distillation, which we've talked about in the last couple of podcasts—that's, yes, that happened.
Can I ask: Is there any way to prevent distillation going forward? Do any of these companies have any sort of defense against that behavior?
Not really.
Okay.
Because it just—there's work to stop API access at scale. You can put rate limits on, but at the end of the day, you're trying to block it. It's like trying to block piracy back in the day. You're dealing with digital bits, inputs, and outputs.
I've heard rumors of some pretty exotic things that folks in China or elsewhere have done to try to distill these models, even if they're rate-limited or their API access is blocked, including, as I mentioned in my article, pretty exotic things with chat clients, scripting them, and trying to get all the answers out.
Okay.
And it works. As I said, it works to their benefit internally. OpenAI is not serving up the original first-run model for users to access. That wouldn't be economically viable. A huge part of the post-ChatGPT launch was that they were overwhelmed with demand and drowning in costs because they were just serving up the whole model.
Right.
That was actually where a ton of the distillation breakthroughs happened, precisely out of desperation and a need to figure out how to serve hundreds of millions of people economically. Distillation is just a fact of life. It's not a good thing or a bad thing.
My point here is that people say, “Oh, well, they're claiming distillation. That's cope.” No, it's true. I think I've been pretty consistent: The US labs are still ahead. To do a GPT-4o-level model, a Sonnet-level model, or an o1 model months after those models were publicly available, by definition, you're not ahead.
The efficiency breakthroughs are real, but they're efficiency breakthroughs. They're not capability breakthroughs. So there's a bit here where everyone on all sides is arguing their interests. But that doesn't mean they're wrong.
Yeah. Multiple things can be true at once. And in terms of the meta-conversation, I think the fact that DeepSeek is a Chinese company kind of broke people's brains and drove the reactions here to an insane place.
I was joking about it with Bill Bishop earlier, because we recorded Sharp China a day late this week. Anytime there's a big Chinese breakthrough in technology, it goes to this place where you have 1 contingent of people who spend days explaining why the progress doesn't matter: It was subsidies, it was IP theft, or in this case, it was distillation. It's not a big deal whatsoever.
And then on the other side of the spectrum, you have these people who treat the breakthrough as evidence of imminent US decline and changes to the world order. So the truth is usually in the middle, and I think the truth was in the middle on this one as well.
The reality is that DeepSeek is very well-funded, not some $6 million side quest. They were using a ton of American hardware with NVIDIA chips. They distilled OpenAI's model to develop DeepSeek-V3 and DeepSeek-R1, and they introduced some really cool innovations that will force competition to be better and smarter going forward. It's not an end-of-days situation.
I wonder if Mistral had done this, would it have been seen as a really interesting twist and subplot and not one of the biggest stories in the world for the first couple of days of this week?
I think you're exactly right. I don't want to speak too much for the China side. I'm not Chinese. I think this is probably going to be very empowering for the Chinese ecosystem, would be my guess.
But from the American side, I do think there's a real tapping into a real underlying angst where—
Yeah, like fear and paranoia.
China is crushing us in all these areas. The one thing we have is technology, and, oh my God, we're losing that as well. As an American, I'm okay with that.
A little bit of urgency might be healthy.
That's right. That's right.
Yeah.
Exactly.
Indeed. All right, well, one more analogy. Mike says:
“The incredible efficiencies that DeepSeek displayed remind me of another era of impressive outputs on inferior hardware: the gaming scene of the late 1990s to the early 2010s. Video game consoles like the Xbox, PS2, PS3, et cetera, were typically launched with hardware roughly similar to a mid-tier gaming PC, but they would stay in the market for 5 to 6 years, meaning that midway to the end of their lives, the hardware would be very outdated relative to a new gaming PC.
“That said, games like Uncharted, Gears of War, and Gran Turismo would be released exclusively to consoles with graphics better than anything coming out on much more powerful PCs at the time. That's because console developers were ‘coding to the metal’ and eking out every bit of performance from the hardware, including often NVIDIA GPUs, maximizing against their static constraints, just like the DeepSeek team.
“PC developers, on the other hand, had to code to a moving target of a host of different PC configurations: different CPUs, RAM capacity, and different GPUs depending on the PC. It meant they couldn't come close to optimizing to the level of console developers. As they say, history doesn't repeat itself, but it rhymes.
“Thanks again for everything. Go, Commanders. The future is bright. Stick with it, Andrew.”
Thank you, Mike. I'm not a gamer or a video game historian, so I'll defer to you. Grade the analogy, part 2.
I think it's a solid analogy. It's not perfect, for lots of reasons that we could pick apart. But it was the case that you were getting this incredible output on old hardware with consoles, and PC gamers would be sort of annoyed at it.
One of the things that's—
That was just possible, yeah.
Yeah. One of the things that's shifted over the years is that consoles have become even more PC-like, and there's much more of a standard approach using game engines and less writing to the metal for a specific console.
Mm-hmm.
Part of it is just an economic thing. When the cost of asset creation started to outweigh the cost of engine development, you used to spend all your time building the engine, and when you're building the engine, that is close to the metal. Then they realized, actually, no, if we're going to make money on this game, it has to run everywhere so we can sell it to the most people, and everything sort of got abstracted away.
I think that, by and large, it's been the case that you're going to get the best graphical experience on a PC for basically every game, and that's been the case for a while. But that was part of the PS4, Xbox 3s—or whatever, Xbox One—generation. That was the generation where it really shifted.
We've talked about this in the past: Sony was ahead of this. They saw this coming, in part because developers had previously been reticent to write to the metal for the PS3, which was actually hard to develop for. They started just writing for Xbox and then porting to the PS3.
Ah.
People would write specifically for the PS2, or the PlayStation before that. Now they're doing this new approach. If we want to have differentiation, we need—
Let's buy some studios.
—to buy up all these studios—
Yep.
—get exclusives, and then people buy our whole thing. So it's actually tied to this discussion we've had before. Again, it's not a perfect analogy for different reasons, but the idea of constraints driving tremendous innovation in terms of efficiency and things on those lines is absolutely the case.
By the way, if you do the calculations of the amount of compute they had, their model efficiency was so high with their mixture-of-experts approach, particularly in training. One of the big innovations was that it was thought that mixture of experts made training much more complex, difficult, and less efficient, but you got the efficiencies in inference.
They came up with a new way, particularly of making sure every expert got trained fully and housing that sort of counting within the experts in a bias term instead of trying to calculate it all at the top end. That actually got them more efficiency in training.
They still had excess compute—relative to what they needed, and that included hand-programming some of their compute units to manage communication more efficiently.
That's crazy.
The nerfing worked. It really did, and it's a testament to the whole point. This is part of the whole discussion, right? We can talk about distillation. We can talk about the fact that they are still months behind. That should not at all diminish this incredible engineering work.
Yeah.
It's really amazing what they did. And, again, it's a blessing to everyone that they released it and made it broadly available.
Indeed. Okay. So, related to engineering work, Adam says:
“I've really enjoyed the discussion of the DeepSeek kerfuffle this week, and when discussing it with a buddy in IT, I was reminded that what's happened fits perfectly into established practice and philosophy in the big tech world.
“Ben has repeatedly mentioned that, because of constant advances in compute, the hard work of optimization has almost never been worth it. Why spend valuable time making something run better when next year's chips will run it good enough, and the opportunity cost of optimization is more features?
“My buddy drew a parallel to the bloated operating systems we have now and inefficient applications. We have utterly god-tier hardware, but often the subjective user experience is no more responsive than it was for early PCs. I'm starting to wonder if the efficiency chickens are coming home to roost. Have decades of skyrocketing compute capacity resulted in a software development culture where optimization is not given sufficient attention?”
Yeah, this is a really good point. I saw a couple of other people mentioning this and putting it in my face that I've made this optimization point. I think it was made most concretely in an interview I did with Kevin Scott of Microsoft.
Hmm.
He was making that point, but I was in broad agreement with him. I think there's two points. Number one, I think there's a bit where it's sort of like the comparative advantage question: you should always choose your comparative advantage. But the problem is, there's a time element to that, and there's a learning curve that, once abandoned, only comes back over time.
We're seeing this with the US and China. The US can't build stuff, right?
Yeah.
That's downstream from rational decisions made at the time that accumulate and lead to a wide divergence that ends up being a problem.
In this case specifically, I think what's coming home to roost is not that it's illogical to spend $100 million or whatever it is to build the future models if you really believe that AI is going to have this astronomical economic impact.
Mm-hmm.
There is a question, though: at what dollar figure does it feel like we've missed the boat, right?
Yeah.
I mentioned it in an update this week: how Google built up their infrastructure simultaneously with developing their product and completely revolutionized the whole server area, which used to be these bespoke machines dominated by Sun.
What they did was basically say, “No, we're going to run on x86, consumer-grade hardware.” Now, over time, Intel responded to that opportunity and made server-grade x86, but it was still drastically cheaper than what came before.
“We're going to assume that it's going to break. We're going to build systems that are resilient and handle stuff breaking down all the time.” It's actually much cheaper and much more scalable, and that was foundational to their success.
I think there's probably a bit here where we have gone too far in the “don't worry about optimization” approach. There probably should have been more dual tracks, where one team isn't worried about optimization and another team has that as all they think about.
Yeah.
That has probably been a mistake, but I totally own the “I never said this before this week” sort of critique. If anything, I've said the opposite. It's a very good point.
Well, as far as the challenge going forward, are you saying that's a muscle group that needs to be developed over time among today's engineers? Can you just flip a switch?
This is the reason why, again, I'm just all over the place on Google. I can never decide, and I usually go the wrong direction, so always take my Google analysis with a grain of salt.
But this is something I wrote a piece about last year regarding Gemini 1.5 and Google's nature, and I was talking about the million-token context window.
Mm-hmm.
There's almost certainly an aspect of how that works that is downstream from Google's infrastructure, the way they optimize and organize things, and how they design their systems from the ground up to work with their software. They do it in a scalable way, and Google is still the best at this.
They have the best data centers. They have the best infrastructure. They have their own chips. If anything, this is the part of Google that's still the most performant.
It's still working.
Yeah. So, if you want to be a Google optimist, I would have this very high on the list: at the end of the day, when it comes to serving the world, Google is still the best. They can do it the most economically. They can undercut anyone on price, particularly if models are commodities.
Again, my concern with Google is, if they're really good at it, the chief entity being hurt is Google Search.
Mm-hmm.
But setting that aside, the sort of classic disruption question, from the infrastructure perspective, they're still the best. To the extent that does matter, that's good for them.
Okay. Well, related to all of this, culture-wise, Dan asks, “Andrew and Ben, at one point in Monday's podcast, Ben discussed what DeepSeek could mean for the LLM product overhang—in other words, all the products we should eventually see if LLM progress stopped today. It reminded me of a discussion he had with Nat Friedman and Daniel Gross in June 2024, particularly this tidbit from Daniel, who said, ‘Obviously, hopefully there will be better breakthroughs and the models will become easier to use, but you could also take today's models and make awesome stuff out of them.’”
“And why that hasn't been happening, by the way, despite the best attempts of this podcast, I think is a very deep and interesting question. What is going on with Silicon Valley in general, and is that innovation pipeline from research breakthrough to iPhone waning a little bit? I don't know. So I guess my question is,” Dan says, “the same as Daniel's. What is going on with Silicon Valley's innovation pipeline? Why aren't we seeing awesome stuff? Is it simply that LLMs are such a massive innovation that it'll take time to productize, as Ben has theorized before, or is there more to it—culture, big tech dominance, inertia among potential customers? It seems like the theme has come up several times now, and at some point, we need to ask what's wrong, either with the technology or the people supposedly leveraging it to build stuff.””
Ben, do you have thoughts?
I think it's a little bit of all the above, but all of the above goes into the fact that historically, this stuff takes time, right?
Mm-hmm.
The internet—the World Wide Web—was in the early '90s, and you had the dot-com era, a lot of companies in which didn't work, and then suddenly it did. The same thing happened with the PC. It takes longer than you think.
Again, why? Is it because of culture? Is it because of customer inertia? Is it because of all these bits and pieces? Yes. It's because of all those things.
Yeah.
But also—
It hasn't been that long, just to be clear.
No, exactly. It's only been two years. Transformers were invented in 2017, so that's been a while, but even that speaks to it. Transformers were invented in 2017, and it took a while for folks to even start utilizing them.
Mm-hmm.
You had GPT-2 and GPT-3. When we started that podcast series with Nat and Daniel, GPT-3 was out. It was before ChatGPT, and we were like, “No one's building products around this. This is actually amazing.”
Yeah.
Then ChatGPT comes out, and it's like, “Oh, my.” That was the wake-up call for everyone. But to your point, that was only slightly over 2 years ago, in November 2022.
The history of technology is that you can have tremendous breakthroughs sitting in plain sight. They can not just be sitting in plain sight—the transformer was—but, like ChatGPT, they can have everyone awake and scrambling, and yet it still takes time. If anything, the larger the change and the larger the transformation, the more time it actually ends up taking.
This is an unsatisfactory sort of excuse, if you want to call it that. But a lot of things about human nature and humanity are unsatisfactory excuses. What we can do is just look at history: it takes time to figure this out.
A lot of the initial things are just taking the old paradigm, slapping on AI, and seeing if that works. Again, my favorite analogy is that you could put advertisements next to text on the internet, like a newspaper. That didn't make any money. You had to come up with a feed. That's what actually made money.
Facebook invented the feed in 2006, 13 years after the web was invented. This stuff just takes time. Now, are there questions about Silicon Valley's ability to innovate and all these bits and pieces? Sure.
Yeah.
There always are. But I think that is actually part of the broader critique.
Yeah. I mean, it's unsatisfactory because it would be so much more satisfying to have a take on the screwed-up culture in Silicon Valley and the complacency that pervades all these companies. I think it dovetails with some of what we were talking about earlier: when you have unlimited resources to throw at a problem, maybe you lack the forcing function to optimize and work on efficiency and really innovate in useful ways.
At the same time, though, as alluring as that story is, it's literally been 2 years, so I have to take a beat before I hold anyone's feet to the fire on that one.
Right. And the fact of the matter is, the chat clients are awesome. This is another human nature point: the usability of these chat clients is basically gated by your willingness and capability to think of things to do with them.
Yeah.
I was writing the DeepSeek FAQ article. I initially started putting every question as an H4 header in Markdown, which is 4 hash marks. Then I'm like, “No, this is not going to work out well. I should just make them bold.”
Do I go through and figure out a way to do it? No. I drop it in ChatGPT and say, “Change all these H4 headings to bold,” and it does the whole thing for me, right?
Right. It's not a huge thing, but that saves you probably 20 or 25 minutes or something.
Yeah. Just irritation, in general, or getting someone else to do it for me. I guess someone just lost their job—my text formatter.
Even internally, each of us has massively untapped efficiency gains that are downstream from our lack of imagination about how we can use this stuff. That applies at the micro level, and it applies at the macro level.
Yeah. I think that's true. I would also say that the people I know in tech love using LLMs on a daily basis, and I actually have more and more friends in mainstream society who are using them pretty frequently.
But I am waiting for somebody to productize them in a way that makes them more idiot-proof. That's what I thought Apple was going to do with Apple Intelligence, frankly: make this so easy that it becomes shorthand for everybody. Apple Intelligence has just been an absolute disaster.
At some point, I think there will be even more adoption than we've seen thus far.
Right. Maybe we're going to end up with more of a classic founder story, where someone solves a problem for themselves, makes a product out of it, and it ends up being a huge company. There are more myths to be made.
Yeah. Right now, it depends on your volition, and soon enough it will be second nature for all of us.
Right. But that was the thing with computers, too, and with the web, frankly. It depended on your volition, right?
Yeah.
The first people using PCs were eager to figure out what to use them for, and the first people on the internet were out there finding obscure communities. We talk about BBSs back in the day.
Carolina basketball message boards. Absolutely.
That's right. That's right.
So we'll get there. This is a long rant in response to a rant you had on Monday's episode.
Okay. All right. I'm going to go pour myself some coffee here.
That's right. Get comfortable on your end.
Bring it on.
William says, “At the end of the latest episode, Ben dropped some absolutely abysmal takes after correctly explaining and discussing R1. The best take on Twitter that I saw was essentially everyone sees R1 and then doubles down on their preexisting views.”
Ben was no exception. When discussing ‘Why are they called OpenAI? Har-har,’ you have to remember the context in which they were founded. There was not a new AI lab being founded every week at that point. The primary player was Google’s DeepMind. Have you read the DeepMind papers? There’s barely enough technical substance to do anything with.
They didn’t even allow external researchers to use the models at scale until AlphaFold. OpenAI was open relative to Google, and now the window for what it means to be open has moved.”
I just want to jump in.
Okay.
Great point. Very, very well put. Anyway, I’ll continue.
Okay, plus one to William. We’ll see how he does in the second half.
“If you want to say, ‘Fuck Anthropic or OpenAI for slowing down progress, what the hell? Where is the ire for Google?’ The whole reason we have any of this is because OpenAI was founded. They are the reason this stuff exists. Moreover, the take about holding back GPT-2 is also terrible. They did that to figure out: What is the minimum we should do to make sure this can’t be misused?
“They now have a process that they and Anthropic follow when rolling out new models. They produce model cards that rate CBRN risks. That’s good even if you don’t think it’s possible for these models to have meaningful CBRN risk. And if DeepSeek drops the weights of a model that does have high risk in one of these areas, it’s useful for us to have language to discuss it.
“From a business analyst perspective, the industry leader doesn’t benefit from open source or open weights. This is part of why Google didn’t open-source its infrastructure the same way Facebook did. It was a strategic credit for Facebook. The same is true today. It’s a strategic credit for Facebook to release its Llama models.
“It seems like DeepSeek is interested in open source, not from a business strategy perspective, though, so that’s pretty interesting. It may be that as long as Liang Wenfeng is in control, they will release the weights to all (most) of their models.”
So, Ben—
All right. You don’t need to reframe it. First off, William, fantastic email. The rant deserved a rant, and you delivered, so very good. I already gave you credit up front for the OpenAI versus Google. Great point.
To jump to the end, the business analyst perspective—yes, correct. From the distillation bit, what’s the easiest way to avoid distillation? Just get the weights, right?
Mm-hmm.
I would also add that, from a DeepSeek perspective, it makes perfect business sense to do open weights and open models for the same reason it made sense for Facebook to do open infrastructure. Even the counterargument that they should have kept this secret so the U.S. wouldn’t know about the efficiency gains—the benefit is that they spread this knowledge to everyone in China who is dealing with severe chip constraints.
Hmm.
All this stuff sort of makes sense. So I actually grant almost everything in this email. I will just explain the source of my rant.
Okay.
That was not a business analyst rant. Did it confirm some of my takes? Yes. My take is that there’s this tension between using the language of safety and the world-ending possibilities of AI.
To the extent that this is a real concern for you—and to be clear, I’m not flippant about this; I completely grant the concerns, and that’s part of what drives my irritation—it is detrimental to your cause to expand the definition of safety to mean a large number of things which we may find objectionable, but which are not safety-oriented.
Mm-hmm.
That’s why I went to the GPT-2 notes, talking about misinformation or biases or whatever it might be. Again, those are legitimate things to have a discussion about, but the way it manifests in an actual discussion is that you push back against that and you’re accused of not caring about the world ending.
Right.
That’s the motte-and-bailey sort of aspect of this argument. This is a fallacy where you go out to the bailey and make a point. Someone objects to that, and you retreat to the motte and say, “What? You’re talking about the world ending? You don’t care about this?”
You don’t care about safety? Yeah.
This might be motivated by my own personal position, which is—I promise you, every time I write about this, I get a lot of pushback that I don’t care about the world ending. So there is absolutely personal frustration about this sort of conflation that happens.
Number two is the overall structure of OpenAI. I’ve looked for this for ages. It might have been some conversation. I don’t remember. But there is a real element to even this board structure and this idea: if you really believe that AI is going to control the world, implicit in that with this structure is, “We will control the world.”
Right.
Right? Because the board’s going to make the decision. And who—
And we will decide what to release and when.
Who voted for them?
Yeah.
Right. Exactly. So if you take these arguments seriously, then I have real concerns.
Mm-hmm.
To the extent that this first point William made—doubling down on their existing views—yes, absolutely happened. One of my existing views all along is that, number one, if this does become dominant, I’m very worried about Sam Altman deciding the course of the world. This isn’t a shot at Sam. It’s a shot at any one person—
Any individual.
Having that degree of power.
Yeah.
Right?
I feel like that’s a reasonable stance.
Yeah. Number two, I don’t think that’s going to happen. I think there are going to be multiple models. This is very cliché, I admit, but I believe it.
The best way to fight against this, to defend ourselves against this, is to have more AI, not less AI. If we could go back and put this all back into Pandora’s box, sure, that’s fine. My belief is that we crossed that Rubicon a long time ago.
Mm-hmm.
It’s going to be out there. It’s going to be open, and R1 emphasizes that point. This is the whole basis of the chip discussion, right? China’s going to get this. They’re going to get enough chips, and they’re going to figure out how to build stuff over time. You’re fighting a losing battle that, by the way, is just the entire wrong mindset of playing defense instead of seeking to innovate.
Right.
Which, again, is totally reasonable to disagree with. I loved our discussion on that last time, in part because I feel like I could flip around and argue your point just as well.
Yeah.
It’s a discussion that needs to be had. But this is a prior of mine: AIs are going to be broadly available. If your number-one concern is truly safety, and you’re saying, “We’re a nonprofit. We’re doing this for the good of the world,” then you should be open-sourcing it.
Mm-hmm.
Now, to his point, from a business analyst perspective, of course they should not be open-sourcing it. But that means they’re being big hypocrites. And if they’re big hypocrites, I’m going to yell at them.
So, again, I think William makes all reasonable points. That segment—again, it’s a podcast. We get a little spicier on the podcast than I necessarily do on Stratechery. But that’s a reaction to what I perceive as the hypocrisy in this position, that if you take a very sharp analytical edge to it, looks like trying to capture power for a very small coterie of people.
Coterie. I can read that word. I can’t pronounce it either. Somebody in the emails—
All right. Well—
Let us know. But my question is: Would open-sourcing OpenAI’s models—and opening the weights, that is—help because it would disperse the innovation and allow other Americans to innovate on it? Is that right?
It’s like nuclear weapons and mutually assured destruction.
Mm-hmm.
Which, again, would we be better in a world with zero nuclear weapons? Well, actually, it’s an interesting question.
It is an interesting question.
I mean, on one hand—
I was pretty annoyed by the end of Oppenheimer because they made it sound like we unleashed some horrible era on the world, when in fact the nuclear bomb was invented and there’s been more relative peace than at any time in human history.
No, exactly. Exactly. This is the whole thing about the U.S.-China economies being so intertwined. Like, yeah, this is really a bad thing, and they’re like, “Well, maybe it’s a good thing.”
It’s the only reason there’s not a war right now.
Maybe it’s going to actually work out, right? The U.S.-China relationship is mutually assured economic destruction.
In any conventional world with conventional weapons and non-integrated economies at the level we are now, I think almost without question there would have already been a war over Taiwan.
Mm-hmm.
Right? So now the problem, and the very reasonable pushback, is, okay, if you want to take my arguments and put them back in my face, you might be right in the short to medium term. But if you think about the long term, all we're doing is setting ourselves up for a truly world-ending, destructive scenario.
Yeah.
And my response to that is, “You might be right,” and it's the same thing with AI.
It's all knock on wood right here.
We might be right, but the problem is AI is out there. It's a real thing. It is going to be diffused, in my opinion, and given that, the answer is not to try to put it back in the box.
Mm-hmm.
The answer is to run in the exact opposite direction. And again, from a business analyst perspective, William's totally right. OpenAI's right to close down, to not be open. Just save me all the—
The holier-than-thou—
—the self-righteous sort of posturing.
Yeah.
Yeah, exactly.
We're doing this for us.
That's what I was responding to. That's exactly.
Okay.
That's what I was responding to, just to be clear.
All right. Well, one more bit of Sam news. Ollie says, “Did you guys see the tweet celebrating Sam and Satya's rapprochement? Guess they're open to a soft landing now.” We got lots of emails about this specific tweet. Do you have any—
Just the one where they take the selfie together?
The selfie, yeah. Do you have an official—
Oh, a lot of good jokes.
—comment on the Sam and Satya selfie?
I'm not sure how many. There's a bunch of funny comments. I think my favorite one was something to the effect of, “When your parents come to tell you it's time for dinner after they've been arguing in another room for 6 hours,” or something like that.
Feels right.
Great stuff all around. No, look, Microsoft is looking amazing in many respects right now. They still have full access to all OpenAI's APIs. They have the right of first refusal to the compute, and they're also sort of letting themselves off the hook for pursuing the funding of these leading-edge models that get commoditized within months.
Mm-hmm.
And so that seems like a reasonable place to be. Now again, the OpenAI bet, which again is reasonable for OpenAI, and the SoftBank bet, which Masayoshi Son is making, is not reasonable. So I guess it makes sense for him and whoever ends up funding him. Their position is, no, actually, we're going to get there first, and it's going to dominate the world, and the economic returns to one model are actually going to be a thing.
Now again, I think the evidence is tilting toward models being a commodity—
Yeah.
—but it's still reasonable to bet in the other direction.
Is it, though? Is it reasonable if you—
No, it's reasonable if you have the appetite for the risk that entails, right? Maybe there's only a 5% chance, but the upside of that 5% is astronomical. It's a venture-capital type of bet where it's probably not going to be the case, but if it is, the returns are so large that it's worth putting money toward it.
Again, this is why the debt part doesn't really make sense to me. This feels like an equity-funding sort of situation.
Mm-hmm.
But if you're OpenAI, I get why you would go in that direction. And if you're Microsoft, I get why you wouldn't.
Yeah.
And again, it's not like they're cut off from OpenAI's models. They still have leverage that they can exercise to make sure they still have access.
Well, they're literally putting on a good face for the world this week, and I enjoyed the selfie. I can't wait for future updates on the Sam and Satya—
Just glad—
—relationship.
Yeah, just glad the parents aren't arguing anymore or—
Oh, man.
—you know, it's time for dinner.
I'll tell you what. I am more open to OpenAI than I have been in the past, in part because I use ChatGPT on a daily basis at this point. I will say, if there's one person in tech who is most likely to be the subject of 8 different documentaries that are sold to rival streamers at some point in the next 5 to 10 years, Sam Altman is the runaway favorite to be the subject of all those documentaries about—
No question.
—corporate abuse and corporate chaos.
Nope.
It's gonna be great.
Well, yeah. It's very reasonable to assume—and arguably make the case, based on episodes in the past—that Sam Altman has an Icarus sort of fetish.
Yeah.
He will always fly too close to the sun. So we'll see if it happens in this case.
The adventures will continue.
Peter says, “DeepSeek said their final training run of V3 was accomplished with a cluster of 2,048 GPUs, a far cry from xAI's apparent 200,000-GPU cluster. Assuming that compute needs are unchanged but are met using smaller clusters than previously assumed, what are the impacts on things like networking, power distribution, liquid cooling, and other areas that have been gating factors to creating large clusters?” Ben, do you have thoughts there? I think that could probably be a 45-minute conversation, but what comes to mind?
Well, I would say with DeepSeek, there's this report, and I think Dylan Patel originally said back in November that they had 50,000 Hopper GPUs. So number 1, that was misconstrued to be H100s. H800s are also Hopper GPUs. The difference is the constraint in memory bandwidth.
So first off, the 2,048 was almost certainly the biggest they could go, given the memory bandwidth constraints, because you have to communicate all these GPUs and keep them in sync. So would they have preferred to do a much larger run? Yes. They couldn't because of the nerfed memory bandwidth.
So is bigger going to be better? Almost certainly. I think GPT-4 was trained on somewhere around 24,000 GPUs or something. Maybe I can't remember what the exact number is. Maybe that's even high.
The constraint isn't that they didn't have more GPUs. It's that the communication overhead becomes overwhelming and the whole thing falls apart the larger you get. And so a huge part of this expansion is increasing those capabilities.
That's a big part of NVIDIA's moat. They're the best at building these systems. It's not just that the core AMD chip is faster than the core NVIDIA chip.
Mm.
NVIDIA's just a gazillion times better at, number 1, programming for them because of CUDA, but then, number 2, they're so much better at tying all these chips together—
Connecting them, yeah.
That's right. And so DeepSeek would have been better on more hardware.
Number 2, again, this was just the final model run. You get to a final model run that is so cheap and runs so well because you've done a gazillion other experiments and runs along the way. So their researchers are using a bunch of GPUs all the time to even get to this point to do the final run.
Number 3, they've been serving inference at shockingly low prices, which is arguably the biggest indicator. They completely destroyed the pricing model for inference in China about a year ago.
Mm.
So that's where most of their GPUs are probably going: people accessing it and needing to use it. A lot of these huge data centers are for people to actually run inference—
Use these products, yeah.
—not just training.
The other thing about this, and I should have made a bigger point about this, I think, in my FAQ, is one of the big implications of, number 1, the power of distillation, and, number 2, these reasoning models and their ability to generate chains of thought that are useful, is that the capacity for AI to train AI is exploding.
The big barrier is that, around GPT-4, we got to the point where we've surfed the whole internet, right? How do we get more data? The synthetic data question has always been the opportunity and always a question: Is it going to work? It looks like it's going to work.
So you think about it, if you want to get to the question we talked about on the first podcast of the year: Is aggregation theory dead? Is there a marginal cost to using AI that's going to mean that the way we've thought about internet economics is going to change?
Well, the counter has always been that regular LLMs are getting way cheaper, and they're really good at a lot of stuff. You're not going to use a reasoning LLM for everything. But the other thing you can use a reasoning LLM for is generating that many more answers to that many more questions—basically an infinite number.
Because the LLM, especially with reinforcement learning, gets into it, and you can…
And so you're generating more and more and more synthetic data, which is not a direct distillation, and you use that to train regular LLMs, which just poop out the answer. They don't think about it.
Mm-hmm.
And suddenly those are way more capable. And so you think about it from that perspective: the amount of potential knowledge in the world is infinite. We're talking about the bitter lesson, brute-forcing the creation and acquisition of all possible answers in the world, right?
So there will be infinite inference demand as we go forward here, right?
Well, is this training or inference? This is inference for the purpose of training.
Hmm.
So the self-contained need for inference—and this is where the $100 million OpenAI thing makes sense—if they see this as the path, we have a reasoning model that generates chains of thought that can then be used to train regular vanilla LLMs, their potential appetite for compute is actually larger than ever.
Yeah.
Because they want to train the reasoning model, then they want to use that reasoning model to generate all this synthetic data to train the base model. By the way, once the base model is better, you get a better reasoning model, so you can generate more synthetic data to make the base model better. We have entered this virtuous cycle which, to the credit of the people worried about real, capital-S safety—like AI taking over everything—is what they predicted.
Mm-hmm.
AI is making AI better, and we are in that era, and that era is going to demand a basically infinite amount of compute.
Okay.
So the reaction is overwrought. Again, I think there are legitimate questions about timing, right? Is there a bit where all this demand, particularly for inference, if we get way more efficient, might not be consumed? But the potential for that much more compute needed for training and inference—and inference that goes into training—is a huge thing.
And so there's going to be a massive demand for power distribution and liquid cooling.
Yeah, right.
Yeah.
Yeah.
It's all part of it.
It's all possible, right? Yeah, and by the way, there's been such a flood of stuff coming out of the Trump administration, it's hard to keep track. But one thing that slipped under the radar is that there's going to be an executive order for creating power that's off the grid, basically tied directly to these data centers. And a dramatic loosening of regulations. This is basically what we called for on this podcast—
Yeah.
A few months ago. There needs to be the capability for dedicated areas where you can cut through all the red tape and create power, and the power can be tied one-to-one to the data center.
Yeah.
The data center is very predictable in how much power it needs, and it's completely off the grid. You could build these anywhere. You could build one in the middle of the desert with a huge solar farm, even if you wanted to. The problem with solar is that it's intermittent, and at night it goes away.
Mm-hmm.
And so you need a lot of batteries because you need continual power. This is why nuclear is a natural fit for these data centers. But you could build that and then just have a fiber-optic line that sends the data somewhere else, and you don't even need all the grid stuff; you can completely bypass it. So the appetite here is larger than ever. There's actually an unlock of appetite.
Great.
I know I'm monologuing a little bit, but that's why my initial write-up on R1 was about the reinforcement learning, the synthetic data, and the distillation aspects.
Yeah.
Because that's the unlock. That's the real takeaway from R1. All this other stuff is kind of a distraction.
No, exactly. I came to R1 having read your piece, and then seeing the hysteria over the last week or so, it was all a little bit befuddling to me. I do think you nailed it with the comparison to the Huawei 7-nanometer chip, where a lot of people who didn't follow AI that closely were digesting this news and freaking out about it. But the synthetic data and what that suggests for the future is a massive deal.
So strap in. It's going to get pretty weird here as we all continue to learn the bitter lesson. Speaking of the Trump administration, though, a couple of follow-ups on the export-control conversation. Two questions from Thomas. The first is, “How would no chip controls on China have impacted the supply of GPUs for American companies? It was always said that the restriction on compute was how many GPUs TSMC and NVIDIA could make. If China had also been in the market, how would that change the amount of GPUs that could be delivered?”
I have no idea. I'm just curious whether you have any take on how that might have shaken out.
The question is, did it shake out at all? China seems to have been getting plenty of GPUs, number one. You had the H800—
Singapore.
The A800 and the H800.
A lot of GPUs going to Singapore.
Singapore is buying more GPUs than it has electrical capability to power, right?
Yeah.
We know where those are going, and I think this is all a very fair question. I think one of the more powerful critiques of NVIDIA that you can leverage is that NVIDIA does not operate like an American company.
Mm-hmm.
Which, again, is kind of the standard in tech. You're just global, and you sell to whoever you want. Was there a situation where NVIDIA could basically decide who gets its chips? Yes. Have they leveraged that to maybe not give the big tech hyperscalers as many as they want because NVIDIA realizes that's a danger to them in the long run? They don't want to give too much power to them and end up in a sort of monopsony situation.
Mm-hmm.
Yes. Has a lot of that gone to startups like CoreWeave and things along those lines that NVIDIA has also invested in? Yes. Has that entailed selling a lot to China when Amazon or Microsoft would have been happy to buy those chips? Also yes.
Mm.
NVIDIA's hands are not clean in this affair.
Yeah. Well, I'm not necessarily worried about NVIDIA, but it's pretty funny that NVIDIA is the company that has spent the last several years fighting like crazy to be able to sell its products into China, and in the past 2 months, it became the subject of an antitrust investigation in China, and then had DeepSeek use nerfed NVIDIA chips to make a model that somehow wiped out 20% of the stock's value.
Could be a chickens-coming-home-to-roost situation, but again, I think NVIDIA is going to be fine. All part of the adventure with that company.
Yep.
Thomas, part 2: “I remember hearing about how SMIC had stolen IP and people from TSMC and ultimately had to pay out a settlement to TSMC because of this in 2009. So the Chinese people and government have had a focus on semiconductor manufacturing capabilities for a long time. Related, and as noted on Monday's episode, a huge blocker on the leading edge has been the lack of EUV machines. So I'm wondering, does Ben think this is also a bad idea, and that all current machines and chemicals should be sold into China?”
What do you think? I was curious, too, whether that falls into your take.
I don't know what my take is, to be totally honest. You go back to ZTE and Trump banning them from chips, and I was immediately very uneasy. I was uneasy for these reasons: you're setting up a long-term problem for U.S. leadership in this area. The game-theoretical consequences for what this means for TSMC and Taiwan are very problematic. I'm not saying it's going to lead to war, but it increases the potential for it.
If China is dependent on TSMC, that's a very good reason not to invade Taiwan.
Right.
If they're cut off from TSMC, their cost-benefit analysis is fundamentally different.
Mm-hmm.
And so, again, go back and read what I wrote 8 years ago. I've been very uneasy on this point all along. Simultaneously, the EUV bit—the reason why I have so many hats on this, including a deeply self-interested hat, is that I'm trying to be honest about this—is that part of passing these laws is asking whether you can actually enforce them.
Yep.
And that was a very clear line and enforceable point of a break in the chain: you could stop EUV machines from going to China. That was a fundamentally different technology. It was hellishly difficult to develop. The idea of EUV goes back to the mid- or early 2000s, maybe even the late 1990s. It took 10, 12, 13 years and multiple rescues of ASML to even bring it to market.
Right.
And can China reinvent it? Yes, they can eventually. If you know something's possible, you can do it, but this is actually a place where you can draw the line. I have been much more uncomfortable with the, “Oh, they made us—they used DUV, which we let them buy for years.” What are you actually trying to accomplish here, other than handicapping US companies in the long run, increasing the prospects of competition, and losing your crown jewel in the very, very long run?
This is very valuable leverage. Is this the right place to trade it away? And from a China perspective, I said that Huawei's 7-nanometer was bad for China because the longer they stay on the “let's get on the leading edge with what we have” path, the worse it is. The best thing they can do is go back to basics and build up from ground zero.
Right.
All the semiconductor equipment.
Your take all along has been that you have to start with the trailing edge and build up node by node in order to develop—
The pushback is that China is doing that, and I just made the case before that companies should dual-track.
Mm-hmm.
I fully admit I'm arguing against myself on a lot of this stuff. I still don't know if the chip ban's a good idea. The thing that this week has brought back to me, when I was arguing against it with you—
Yep.
—and not wholeheartedly, because I'm still not totally sure, is that I am worried from a cultural perspective about the US and our—
What signal it sends to ourselves.
Yeah, that we're going to fight through blocking our past innovation instead of feeling the incentive and motivation that we have to run forward even more quickly. Where do I land? I land in a pile of mud. I freely admit that, and I think that's sort of the ultimate takeaway here.
Well, did you read the post from Dario Amodei, the Anthropic CEO? He wrote pretty articulately about export controls and argued for enhanced export-control enforcement. What did you think of his arguments there?
I mean, it's a very valid argument, and I think you've been making the argument, and you made it very well. If people agree with you and agree with him, I get it, because I might agree with you as well.
Mm-hmm.
I'm again down here in the mud. Have some mercy on me.
Yeah.
—
Greased pig. I know. I appreciate it. Well, and I—
Yeah.
It's interesting to me because, in a political context, I have no idea how the Trump administration might handle this, because obviously it's an unbelievably sensitive issue—
Well, let's get to the Trump angle in a second.
Okay.
The issue and concern I have with Dario articulating this—he's the one that wrote the GPT-2 post, right? There is a part of me that's instinctually a little suspicious of these folks who are simultaneously saying AI is the biggest danger, that it's a danger to humanity, and also devoting their lives to building it.
Yeah.
And again, if you want to throw it back in my face and say, “You're just arguing your preexisting priors,” yeah, I am. But that is my preexisting prior. It makes me concerned and suspicious, and it is deeply in their interest and OpenAI's interest to block this sort of thing. But, yeah, that's where I'm at.
Okay. Yeah. Well, and what I was going to say about Trump is that he's managing a much bigger relationship with Xi Jinping and the Chinese. I think there are factions of his security team that would want to enhance and expand the export controls on chips and chip-making equipment in particular, and get some of the allies on board, whether it's ASML, South Korea, or Japan.
I don't know whether Trump—Mr. Executive, we've talked about the broadening power—is willing to ruffle Xi's feathers over the next year or two here. Maybe he shouldn't, because I think it is a bulwark against continued tension heightening. So we'll see.
The thing that is worth keeping in mind—and this sort of takes us full circle to why it mattered that DeepSeek was Chinese—is actually the Chinese factor. I think your analogy to Mistral was a great one. If Mistral did this, everyone would be falling all over themselves with delight and glee—
Yeah.
—saying this is amazing.
“Wow, this is cool.”
Yeah. It's an excellent, excellent point. There's a bit where I mentioned that maybe this is a good wake-up call. I'm not sure that people in the US have fully internalized how precarious our position relative to China is, particularly in a war-fighting scenario.
Mm-hmm.
This is about their industrial capacity, the number of ships they can build. We can't build stuff. We can't build ships. We can't do XYZ. Could we spin that back up? Eventually. We didn't spin up artillery ammunition very well in Ukraine—we're still producing way less than Russia is.
Yeah.
I'm not sure we've fully internalized how bad that is, and also how much worse it's going to get in a new type of warfare defined by things like drones and robotics, all of whose components are completely housed in China. All the little actuators, motors, batteries, and all these sorts of things are dominated by China. If this is a world we end up in where that's what defines a war, we have a big problem.
There is probably an infusion of humility that's necessary, and if it comes in the form of maybe overstating what DeepSeek did but puncturing the attitude of presumed technological superiority, that's significant. When it comes to fighting a war, fighting a war is not aggregating users and capturing demand. Fighting a war is actually expending physical goods and actual people in a war of attrition. Again—
Mm-hmm.
My aggregation theory is not worth very much in that scenario.
Yeah.
Let's be very clear: In the US, everything about US tech and our economy as a whole is all about consumption. At a very fundamental level, we rule the world economically by leveraging our willingness to buy lots of crap. I don't want to give Trump too much credit. He likes tariffs, and he likes throwing out XYZ.
What I recognize, though, is that there's a real arrogance in the way we deal with this sort of stuff, and an assumption that the world as it was 30 years ago is the world as it is today. It's just not true.
Mm-hmm.
Is it actually wise to keep operating with that level of arrogance, assuming that we can dictate to the world what it is or what it isn't, if we increasingly don't have the goods to back up our talk?
When you think about something like Taiwan, people are flipping out over the potential application of tariffs to Taiwanese chips. I already wrote about this in Stratechery. I'll get to it again, but I kind of made my point that Trump's critique of the CHIPS Act—that it just subsidized supply—was fundamentally wrongheaded. You needed to subsidize demand. Tariffs are a way to do that, but I think a less effective way than just straight-up making buying guarantees for Intel to buy chips.
Yeah.
It's not insane when you think about what is necessary in terms of driving demand. There's also a broader point that our total dependence on Taiwan is a big problem. The Taiwanese government, quite rationally, doesn't want to allow leading-edge TSMC capabilities to go off-island—
Mm-hmm.
—because it is their trump card, no pun intended, to make sure the US comes to their defense. But how long is that sustainable if there are very real questions about our capability to fight a war?
Yeah. No, exactly.
And again, the analogy for Trump has always been a bull in a china shop, and there are real questions about his ability to put stuff back together.
Mm-hmm.
But sometimes there is value in stuff being broken, and one of those taboos that might be worth breaking is the total self-assuredness among most of the US that we can win a war.
Right.
In that world, is this sort of screwing Taiwan? Maybe. Does TSMC benefit from a manipulated currency, or a currency that doesn't seem to quite reach its right level, that gives them a fundamental cost advantage in making chips in Taiwan?
Maybe.
Maybe. The hands are not clean on all sides here. Taiwan is an ally, but they're not 100% in line with US interests.
Yeah.
Again, I'm not saying I agree or disagree yet. I'm still really thinking about it, but I do think there's a real paucity of humility and awareness of the situation that we're in right now.
Yeah.
And there might be an argument, by the way, that we need to lean into it, that we should come to some sort of agreement with China that locks them in as they're going to make stuff and we're going to buy stuff. And what does that mean for Taiwan? Well, we better have our own chip-making possibility, because the worry about Taiwan being a part of China is then China has the ability to cut off our chips.
Right.
Well, then we should figure out how to make our own chips domestically, because that's the exact same worry in a war scenario, which is that China blows up TSMC. Then we still need to make our chips domestically. It's very rational. Again, I love Taiwan. I've lived here for 21 years. Taiwan's democracy is amazing.
Mm-hmm.
I would love to stay in this gray area where it's functionally an independent country. I do feel truly free here. Well, COVID was a little iffy.
Yeah.
If you're concerned about national security, if you're concerned about these long-term things, you have to think about these issues.
Yeah.
That's why I wrote “The Geopolitics of Chips” years ago. This is a real problem.
Right. Well, and I don't say it trivially when I reference Trump's executive discretion. I sort of made it sound like a joke about 5 minutes ago, but I mean it. If we don't go the direction that I imagine a significant portion of his administration wants to in terms of expanding export controls, there may be a reason, and that actually may be a rational decision, because it will absolutely infuriate the PRC side. The relationship will devolve further, and it's not lost on me that the PRC has been in the midst of a manufacturing build-out that is not paralleled in the last 100 years beyond pre-World War II Germany and the early 20th-century United States. Like, they're—
Right. Well—
—in a much better position to fight a war than the US is, at least from a manufacturing standpoint, shipbuilding, and all of it.
Right. From the US perspective, the optimistic take is, well, our AI is going to become so good that we can technologically win.
Mm-hmm.
And that's the argument for the chip controls: We will overcome the manufacturing deficiency by virtue of having superior technology and AI. That's why, yes, we're playing our one-time-play Trump card now. Again, no pun intended. But this is the time to play it.
Well, and maintaining the advantage now will then compound in the years to come.
That's right, because the AI—
That was part of the logic.
—because the AI is going to make the AI better. Right.
Yeah.
And so this viewpoint—the funny thing is, the chip ban, if you dig down to the fundamental assumptions undergirding it, is actually the same as the Sam Altman view or the Dario Amodei view, which is: We're not going to have commoditized models. We're actually going to have takeoff, and we're going to have a sustainable, superior advantage that's going to maintain over time. That's why the DeepSeek thing is a big deal. If that's not the case, how many decisions were made with that assumption that might not be true?
Yeah. Well, and would DeepSeek exist if not for nerfed Nvidia chips that had been sold into China? That's an open question, too.
Well, to be clear, ByteDance and Alibaba all have pretty decent models.
Yeah.
But are they working with Nvidia chips, or are they working with what Huawei has developed? And Huawei has been buying up chip-making equipment for years, much to the dismay of people in the security community who think that you should be much tougher on Huawei. I don't know.
I think most of the training is probably happening on Nvidia chips. I think that's where, frankly, a lot of the smuggling is going.
Yeah.
Yeah, I think for inference it's more plausible. Inference needs less of that general bandwidth. You're more constrained by memory anyway, but I don't know. It's hard to say for sure.
Mm-hmm.
But the other thing is, you do have this constraint of power and data centers, which entails the ability to build stuff, and China's better at that than we are. And so you can just sort of brute-force scale, even if it operates inefficiently, relatively speaking, and it's okay because, particularly when you think about it, at the end of the day, the military is going to get all the best chips.
Mm-hmm.
So even if you constrain private industry, it's not like—
They'll figure it out.
I can make the argument for the chip ban as well.
Yeah. Can I read one more email on the chip ban? Do you have time here?
Yep.
All right. Michael says, part 1 of Michael's email is, “First, Andrew could not be more wrong with his comment that—
Oh, finally. Someone who takes my side. Geez.
—that somehow the chip export control—
We have to raise questions about you choosing these emails. I was wondering how that—
Oh, my God.
Yeah, fine. Continue.
Listen, I just wanted to give you a chance to respond to a fiery rant. Michael says, “Andrew—
No, that was a great rant. That was actually one of the best emails we've got in a long time. I will give credit where credit is due, but Michael's turn. Continue.
Terrific energy. “Andrew could not be more wrong with his comment that somehow the chip export control regulations were watered down by industry, and that's why they don't work. You spent a lot of time talking about cope on the last episode, and that's some hardcore cope. The regulations—
By the way, “cope” is such a great word. It's been a great addition to the discourse. I'm glad we're helping mainstream it here.
Oh, my God. My only request, honestly—the only reason I read this email is because I want to make a formal request to all our listeners, to Tech Twitter, to you. Can we all just take a 2-week break from saying the word “cope”? It's everything—
No.
Every other tweet—
We cannot. I love it.
I've read—
—is about cope. It doesn't have to be a permanent break. Just give me 2 weeks. Take a breath, everyone. All right. Michael says, “That is some hardcore cope. The regulations were completely dumped on industry without any engagement or opportunity for feedback. That's especially true of the recent AI diffusion and foundry rules. I don't know if they'd be better, worse, or whatever if industry were more involved, but industry was absolutely not involved in any meaningful way.”
This email made me smile because it reminded me how different the audiences are across the various podcasts we host in the bundle. The Sharp China audience, many of whom are in DC and in government, would have a very different take on what happened to the chip controls, but I'm sure there are lots of people in the Sharp Tech audience who feel like the government was dumping all this on them without having any idea what they're trying to regulate.
I'll just say what I recall is that there were efforts to get Greg Allen removed from his post at CSIS because he was too competent at explaining the areas in which the chip bans were failing to achieve their intended purposes. And in general, the lobbying behind the scenes has been pretty well chronicled. SemiAnalysis actually did a great job in October pointing out all the different ways the chip controls had failed and why they were strategically important.
But to Michael's point, I think the resulting incompetence is ultimately the responsibility of the Biden administration, and I also think he's referring to the AI diffusion and foundry rules, and I know less about that process. That probably was dumped on industry without any meaningful involvement or engagement, and I don't think those rules were the right way to approach any of this. So on that point, I can agree with Michael. Do you have any thoughts?
Yeah, I think you go back even also to the Biden executive order on AI.
Mm-hmm.
I just want to double down on that bit I made where they became true believers in this AI takeoff scenario, and it's like we have to do everything possible to stop this.
Yeah.
And that's why this is an important point, and the question of model diffusion is a really important point to talk through. It's meaningful—obviously super meaningful—from a business analyst perspective, right? Where does value accrue in the value chain?
But it actually undergirds these very deep, fundamental questions. Are we going to sacrifice our semiconductor industry in the very long run because this is the one time that we have to get it right?
Yeah.
Right? And so that's one point.
The other thing: it's not that you and Greg Allen and SemiAnalysis are wrong about these chip controls.
Yeah.
It is worth considering the extent to which the chip controls being leaky is a pressure valve, right? That prevents us from actually playing these scenarios through to the end, and we do need to think deeply about what happens if we had perfect chip controls.
Right.
If China actually did not get any chips, what is the rational response of China in that scenario?
Yeah.
It's not a pleasant one.
I wonder whether Trump is considering that as he weighs all this, as things move forward.
Who knows?
I'm not going to try to jump into that mind. But Michael, part 2: “Ben is correct on ICE engines. I had a fun opportunity earlier in my career to get the red carpet rolled out for me at Daimler and BMW. If you go into the Mercedes-Benz Museum, you go on this long elevator up to the top floor, and when the door opens, you're in a dark room with a spotlight shining on the first ICE engine the company produced. The key technology that they had was the engine. The same was true of BMW. The engine is the key technology. Everything else flowed from that. It's also true of Honda, based on my understanding of the company's history, and I also think it's true of Ford. They developed an engine first, and none of them have adapted particularly well to the EV environment. That their key differentiating technology was fundamentally disrupted isn't surprising.”
Ben, I include that part only to say that I would love to go on a tour of Daimler and BMW and get the red carpet rolled out for us.
I know.
Rolled out for us.
I feel like we need a Sharp Tech field trip here. I'm very jealous.
That's right. Let's go to Frankfurt sometime in the next year or 2. But thank you for the note, Michael, and thank you to everybody who wrote in. We got an avalanche of emails, so this is a bit of a longer episode, but we got through as many as we could.
You were efficient, taking your cues from DeepSeek on this episode, and here we are. We will be back next week, and it's not going to be wall-to-wall DeepSeek next week, but, Ben—
Yeah, don't promise things you can't guarantee. We'll see y'all.
That's true. Enjoy the weekend, and I will talk to you soon.
Talk to you later.