[BidClub_]
Dwarkesh Podcast · · 103 分钟

Jensen Huang——Nvidia 的护城河能否延续?

Jensen HuangDwarkesh Patel

YouTube
TL;DR
  • Jensen 认为,ASIC 威胁集中在单一客户身上。“Anthropic 是一个独特案例,并非趋势。没有 Anthropic,TPU 增长从何而来?100% 都是 Anthropic。”Trainium 也是同样的逻辑。切换的经济性本来就很薄:ASIC 毛利率约65%,Nvidia 约70%(“你到底省下了什么?”);他还向 TPU 和 Trainium 发起挑战,要求它们登上 Dylan 的 InferenceMAX——“没人愿意来。”
  • 护城河在供应链,而硬件瓶颈不会持续太久。已披露的采购承诺约1000亿美元(SemiAnalysis 报道为2500亿美元),此外还有 Jensen 亲自向供应商 CEO 做宣讲、由此释放的隐含上游资本开支。针对“AI 已经占到 N3 产能60%”的质疑,他给出绝对判断:EUV 和晶圆厂扩产“在两三年内很容易做到。你只需要一个需求信号”;真正更持久的约束是能源,以及“水管工和电工”。
  • 针对 CUDA 护城河正在被侵蚀的论点,Jensen 给出的答案是 TCO,而不是锁定效应。超大规模云厂商自己写 kernel,OpenAI 的 Triton 栈也在取代 cuBLAS/NCCL;但 GPU“就像 F1 赛车”,需要 Nvidia 配备的工程师,他们经常带来2倍、3倍甚至50%的模型提速,并且“直接转化为收入”——“全球无出其右的最佳性能/TCO。”
  • 他承认,自己的战略失误是没有充分意识到 Anthropic 确实没有其他选择。“我没有深入理解这一点:一家 VC 根本不会投50亿至100亿美元……这是我的失误。”这让 Anthropic 落入 Google/AWS 手中。如今,市场传闻 Nvidia 将向 OpenAI 投入最多300亿美元、向 Anthropic 投入100亿美元,Jensen 表示:“我不会再犯同样的错误。”
  • 在中国问题上,Jensen 把话说到了最大音量。出口管制等于放弃全球第二大市场和全球50%的 AI 开发者,只为一家公司的利益。中国生产全球60%的主流芯片,可能更多;Huawei 刚刚在创纪录的一年里出货了“数百万”颗芯片,能源充裕足以弥补7nm 的差距(“7nm 芯片本质上就是 Hopper”);“如果有一天 DeepSeek 率先在 Huawei 上发布,那将是对我们国家极其糟糕的结果。”
  • Dwarkesh 的反驳被完整保留在记录中。Anthropic 的 Mythos Preview 发现了数千个高严重性漏洞,其中包括一个存在27年的 OpenBSD 零日漏洞——算力是网络攻防能力的输入。Jensen 认为浓缩铀类比“很糟糕”,并将这种框架称为“失败者心态”;双方的分歧没有解决。
  • 可交易的新信息包括 Nvidia “加入 Groq”,并将 Groq 纳入 CUDA 生态。目标是服务高溢价、低延迟、高 ASP 的 token 市场——“拓展帕累托前沿”;路线图为 Vera Rubin → Vera Rubin Ultra → Feynman,同时要求其他 ASIC 团队证明自己也能让 token 成本“每年下降一个数量级”;产能分配从不竞价,而是按采购订单 FIFO 排队,需求暴涨时也不会提价。
  • 软件商品化逻辑被倒转了。代理和工具调用者不断增加,因此工具实例“很可能会暴增”——包括 Synopsys Design Compiler、布图规划器等;“工具调用会让软件公司腾飞”。之所以还没发生,只是因为“代理还不够擅长使用工具”。
摘要 · 为研究而整理的核心内容

1. 电子流入,token 流出——中间环节不会商品化

  • Dwarkesh 开场提出看空逻辑:Nvidia 向 TSMC 发送 GDS2 文件,SK Hynix、Micron、Samsung 提供 HBM,台湾 ODM 负责组装机架——“Nvidia 本质上是在做软件,制造由别人完成”;如果 AI 让软件商品化,Nvidia 是否也会被商品化?Jensen 没有回避这个框架,反而顺势概括:“输入是电子,输出是 token,中间是 Nvidia”——而这一转换“远未被充分理解,旅程也远未结束。我怀疑它会发生。”
  • 其运营哲学可以浓缩为一句话:“做到必要的最多,做到可能的最少”——凡是 Nvidia 不必亲自做的事情,就交给覆盖 AI “五层蛋糕”的“可能是全球最大合作伙伴生态”。Nvidia 必须亲自完成的部分,“事实证明,难得离谱”。
  • 对于软件估值崩塌,Jensen 看到的是“与人们所看到的相反”:代理和工具使用者呈指数增长,因此“工具实例很可能会暴增”——Synopsys Design Compiler、布图规划器、设计规则检查器。“工具调用会让软件公司腾飞”;之所以还没有发生,只是因为“代理还不够擅长使用工具”。

2. 供应链护城河靠 CEO 布道,而不只是合同

  • 已披露的约1000亿美元采购承诺(SemiAnalysis 报道未来规模为2500亿美元)只是显性部分。隐性护城河在于 Jensen 亲自撬动的上游资本开支:“让我告诉你这个行业会有多大……让我和你一起把逻辑推演一遍。”供应商选择为 Nvidia 而不是其他公司投资,是因为“他们知道我有能力买下他们的供给,再通过下游卖出去”;而且“如果一个架构的业务流失率很低,没人会为它建设供应链。”
  • GTC 的本质是基础设施教育:主题演讲刻意做得“有一点折磨人……像是在上课”,因为整条产业链都必须“像我推演问题一样,系统性地推演它”。
  • 最能说明这种布道奏效的案例是 Sanjay 和 Micron:5年前他们就相信这一判断,并在 LPDDR 和 HBM 上“真正加倍下注”——“这显然给公司带来了巨大的好处。”

3. 每个瓶颈都是2-3年问题——能源和水管工除外

  • Dwarkesh 用数据发起反击:Nvidia 是 TSMC 最大的 N3 客户,按 SemiAnalysis 数据,AI 今年占 N3 产能60%,明年占86%——“既然你已经是大多数,怎么还能翻倍?”Jensen 的回答是绝对的:EUV 和逻辑制程扩产“在两三年内很容易做到。你只需要一个需求信号。你能造1个,就能造10个;能造10个,就能造100万个。”只要 TSMC 被说服,ASML 就会跟上。
  • CoWoS 就是证明:过去两年他们“把所有人都调来猛攻它”,CoWoS 和 HBM 从专用技术变成了“主流计算技术”。如今 Nvidia 会提前数年预判瓶颈:通过 Lumentum 和 Coherent 布局硅光子,与 TSMC 共同开发 COUPE,并将专利授权出去以保持供应链开放,还引入双面探针等新型测试设备。
  • 真正让他担心的是下游:能源政策——“没有能源,就不可能创造一个产业”——以及最难的瓶颈:“水管工。水管工和电工。”他的反末日寓言是:10年前人们说不要去当放射科医生——“猜猜我们现在缺什么?放射科医生。”
  • Nvidia 是否会在2030年前把产品回移到 N7、释放先进制程产能?他说“没必要”——把现代数值计算和封装技术移植到老制程上,需要“没有人负担得起的研发投入”。“我们负担得起向前押注。”不过,假设产能此后永远不再增长:“当然,我会在一瞬间这么做。”

4. TPU 是另一种动物——可编程性才是复利优势

  • 面对排名前三的模型中有两个(Claude、Gemini)在 TPU 上训练,Jensen 表示:“我们做的是完全不同的东西。”加速计算覆盖分子动力学到粒子物理,设计目标是让其他人运营,并且存在于每一朵云中——“我们的市场触达远大于任何 TPU 或 ASIC 可能达到的范围。”
  • Dwarkesh 转述 AI 研究员朋友的观点:AI 无非是“可预测的矩阵乘法一遍遍重复”——TPU 的脉动阵列恰好适配收入正在增长的工作负载;既然如此,为什么要为 warp 调度器牺牲 die 面积?Jensen 的反驳是,新型 attention 机制、混合 SSM、扩散-自回归融合都需要通用可编程性,而“发明新算法的能力,才是真正推动 AI 快速进步的东西”。
  • 协同设计的证明点是:摩尔定律每年带来约25%的提升,但 Blackwell 相对 Hopper 公布的能效提升达到35倍——“没人相信。后来 Dylan 写文章说我有所保留,实际上是50倍。”这只有在处理器、互联结构(NVLink)、网络(Spectrum-X)、库和算法同时改变的情况下才可能实现:“没有 CUDA,我甚至不知道该从哪里开始。”
  • 为什么不做 wafer-scale 并行或 Dojo 式项目?“我们可以做……只是没有更好。我们在模拟器里把所有方案都跑过了,结果可以证明更差。”但工作负载变化确实会触发调整:“最近我们加入了 Groq,并将把 Groq 纳入 CUDA 生态。”如今 token 价值足以支撑一个高溢价、低延迟市场,因此 Nvidia 会通过更快响应、更低吞吐的推理产品“拓展帕累托前沿”。

5. 自己写 kernel 的客户,能让70%毛利率持续吗?

  • Dwarkesh 提出最尖锐的结构性问题:60%的收入来自5家超大规模云厂商,而它们有能力、也确实在绕开 CUDA——OpenAI 的 Triton 栈“下沉到 CUDA C++,而不是使用 cuBLAS 和 NCCL”,并能编译到其他加速器。如果最大的客户都负担得起绕过护城河,竞争是否会退化为“每美元规格”竞争?70%以上的毛利率还能维持吗?
  • Jensen 的类比是:CPU 是凯迪拉克,Nvidia 的 GPU“就像 F1 赛车”——任何人都能开到100英里时速,但要逼近极限则需要专业能力。Nvidia 配备的实验室工程师经常带来2倍、3倍或50%的模型提速,而在已部署的 Hopper/Blackwell 集群中,这些提升“会直接转化为收入”。
  • 他公开发起 TCO 挑战:“Nvidia 的计算栈是全球无出其右的最佳性能/TCO……没有任何一家公司”能击败它。InferenceMAX 对所有人开放,MLPerf 也是另一个基准——“TPU 不会来,Trainium 不会来……没人愿意来”;他还“欢迎 Trainium 展示它一直宣称的40%优势”。
  • 他纠正了客户集中度的前提:超大规模云厂商采购的 Nvidia 产能,大部分服务的是外部客户,而非内部使用(OCI 的全部、AWS 的大部分,以及 Azure 的全部)。飞轮效应来自“数亿块 GPU”装机量、可编程性、丰富的生态,以及最高的每瓦 token 产出——当数据中心功率为1GW 时,这决定了 token 收入上限。

6. 遗憾:Anthropic 因 Nvidia 开不出支票而投向他人

  • 如果 Nvidia 在纸面 TCO 上胜出,Anthropic 为什么刚刚还签下数GW 的 Broadcom/Google TPU 合同?Jensen 的回答是:“Anthropic 是一个独特案例,并非趋势。没有 Anthropic,TPU 增长从何而来?100% 都是 Anthropic。”Trainium 也完全一样。他还指出了被取消的 ASIC 墓地:“做出比 Nvidia 更好的东西,没那么容易。”
  • 他保留了原话中带有保留、未经修饰的自我检讨:“我没有深入理解这一点:他们确实没有其他选择;一家 VC 根本不会向一家 AI 实验室投入50亿至100亿美元……这是我的失误。”Google 和 AWS 投资了 Anthropic,而 Nvidia 从未在公司外进行过这种规模的投资。如今,市场传闻 Nvidia 将向 OpenAI 投入最多300亿美元、向 Anthropic 投入100亿美元,他表示:“我不会再犯同样的错误。”
  • 为什么要资助每一家基础模型实验室,而不是押注赢家?Nvidia 的60家早期3D 图形公司竞争者如今只剩 Nvidia 一家,尽管它最初的架构“完全错了……所有人都会认为我们已经出局。但我们走到了今天。”他从这份谦逊中提炼出的规则是:“不要挑赢家。要么让他们各自照顾好自己,要么把所有人都照顾好。”

7. Nvidia 为什么不会成为超大规模云厂商——以及新云 tension

  • 同一套哲学回答了“为什么不自己做云”的问题:如果 Nvidia 没有打造 CUDA(20年,其中大部分时间都在亏钱)、CUDA-X 或 cuLitho,“没人会做……这一点我完全确定”。云业务则相反:“如果我不做,就会有人出现来做。”
  • Dwarkesh 指出其中的矛盾:Jensen 否认在扶持新云,却又表示 CoreWeave(据报道有63亿美元的后盾融资、已投资20亿美元)、Nscale 和 Nebius 如果没有 Nvidia 就“不可能存在”。Jensen 的解释是:这些公司必须自己想要存在,并带着计划和专业能力前来;Nvidia 可以提供投资帮助它们起步,但“我们更愿意与融资行业的所有人合作,而不是自己成为融资方”。

8. 不竞价、不提价,以及值得押上全部筹码的路线图

  • 面对稀缺 GPU 的分配,Nvidia 从不把货给最高出价者——“这是糟糕的商业做法……如果我给你报了一个价格,我们就按这个价格来。就这样。需求冲上天,那就冲上天。”实际系统是先做预测,再按采购订单 FIFO 排队;如果客户的数据中心还没准备好,则进行调整。Larry/Elon 晚餐会上苦求 GPU 的传说也被直接否认:“他们从来没有乞求 GPU。他们只是下了一张订单。”
  • Nvidia 把可靠性本身当作产品:与 TSMC 合作约30年,而且“Nvidia 和 TSMC 没有法律合同。总会有某种粗略的公平。”路线图节奏是今年 Vera Rubin、明年 Vera Rubin Ultra,之后是 Feynman,再之后还有一款尚未命名的产品;他要求其他 ASIC 团队证明自己也能让 token 成本“每年下降一个数量级”。从一块显卡走到“1000亿美元的 AI 工厂,完全没问题”。“如今世界上只有我们这一家公司可以这么说。”

9. 中国已经拥有算力——Jensen 反对放弃中国

  • Dwarkesh 为出口管制提出了最强版本的论据:Anthropic 的 Mythos Preview 在每个主流操作系统中都发现了“数千个高严重性漏洞”,其中包括一个存在27年的 OpenBSD 零日漏洞——如果中国实验室拥有训练这类模型、并运行“数百万个实例”所需的芯片,这难道不是接近武器能力的技术?
  • Jensen 的反驳从现实出发:Mythos 使用的是“相当普通”的算力,而中国“到处都有”。中国生产全球60%的主流芯片,可能更多,拥有全球50%的 AI 研究人员,还有充裕能源,包括“幽灵数据中心……完全通电”。Huawei 刚刚迎来历史上最好的一年,出货量“非常多。数百万。数百万远多于 Anthropic 的数量。”他的结论是:“中国无法拥有 AI 芯片的说法完全是胡说。”
  • 替代逻辑是:在五层蛋糕中,底层能源可以与芯片形成替代——“如果你的瓦数完全充足……你还在乎每瓦性能干什么?”针对 EUV 才能制造先进 HBM 的说法,Jensen 表示,可以用硅光子把 HBM2 组装在一起,“就像我们用 NVL72 把它们组装在一起一样”。用 Jensen 的话说,Dwarkesh 关于 EUV 的前提“就是错的”。
  • 更深层的判断是:摩尔定律每年带来约25%的提升,但算法可以带来10倍提升——“AI 的大部分进步都来自算法进步”,因此中国庞大的研究者群体才是根本资产。最糟糕的场景是:“如果有一天 DeepSeek 率先在 Huawei 上发布,那将是对我们国家极其糟糕的结果”——开源模型在外国技术栈上完成优化,再扩散到全球南方。

10. “我无法理解那种失败者心态”——一场没有结论的交锋

  • Dwarkesh 用浓缩铀类比算力,Jensen 直接驳回:“这是一个糟糕的类比。一个不合逻辑的类比。”随后他指责对方从极端出发:“你的论证始于极端……这些极端情况很幼稚。”Dwarkesh 坚持认为,任何边际算力都能帮助训练具备攻防能力的模型;Jensen 承认其潜在代价,但重新定义真正的成本:放弃芯片层。
  • 市场数学很简单:中国约占全球科技产业的40%,拥有全球一半的 AI 开发者;放弃中国“是对我们国家的伤害……只为一家公司的利益”。生态系统不是汽车:“x86 交易之所以存在是有原因的,ARM 如此稳固也有原因。”他反复提到的先例是:类似政策曾导致“美国电信产业基本上被排除在全球市场之外”。
  • 当 Dwarkesh 以 Tesla 和 iPhone 为例,认为 Nvidia 无论如何都会失去中国份额时,Jensen 回应:“你面对的不是一个醒来后发现自己是失败者的人。那种失败者心态、那种失败者前提,我无法理解。”光刻也不是宿命——Blackwell 在3年内晶体管仅提升约75%,相对 Hopper 的性能却达到50倍:“架构很重要。计算机科学很重要。”
  • 他的收尾预言被设定为一个标记:几年后,当美国希望把自己的技术栈出口到印度、中东和非洲时,“我希望我们能再次进行同样的对话。我会把今天的对话原原本本告诉你:你的政策……毫无正当理由地让美国放弃了全球第二大市场。”如果失去它,那就失去它。但“我们为什么要主动放弃?”他的坚持是:“我们始终应该拥有这里最好的技术……并且最先拥有。”
Dwarkesh Patel

We've seen the valuations of a bunch of software companies crash because people are expecting AI to commoditize software. There's a potentially naive way of thinking about things, which is: Nvidia sends a GDS2 file to TSMC. TSMC builds the logic dies and the switches, then packages them with the HBM that SK hynix, Micron, and Samsung make. Then it sends them to an ODM in Taiwan, where they assemble the racks.

Nvidia is fundamentally making software that other people are manufacturing. If software gets commoditized, does Nvidia get commoditized?

Jensen Huang

In the end, something has to transform electrons to tokens. The transformation of electrons to tokens, and making those tokens more valuable over time, is hard to completely commoditize. The transformation from electrons to tokens is such an incredible journey.

Making that token is like making one molecule more valuable than another molecule, making one token more valuable than another. The amount of artistry, engineering, science, and invention that goes into making that token valuable—obviously, we're watching it happen in real time. The transformation, the manufacturing, all of the science that goes in there is far from deeply understood, and the journey is far from over.

I doubt that it will happen. We're going to make it more efficient, of course. The way that you framed the question is my mental model of our company. The input is electrons, the output is tokens. In the middle is Nvidia.

Our job is to do as much as necessary and as little as possible to enable that transformation to be done at incredible capabilities. What I mean by “as little as possible” is, whatever I don't need to do, I partner with somebody and make it part of my ecosystem.

If you look at Nvidia today, we probably have the largest ecosystem of partners, both in the supply chain upstream and downstream, all of the computer companies, application developers, and model makers. AI is a five-layer cake, if you will. We have ecosystems across the entire five layers.

We try to do as little as possible, but the part that we have to do, as it turns out, is insanely hard. I don't think that gets commoditized. In fact, I also don't think the enterprise software companies, the toolmakers, do.

Most software companies today are toolmakers. Some of them are not. Some of them are workflow-codification systems. But for a lot of companies, they're toolmakers.

For example, Excel is a tool, PowerPoint is a tool, Cadence makes tools, and Synopsys makes tools. I actually see the opposite of what people see. I think the number of agents is going to grow exponentially, and the number of tool users is going to grow exponentially.

It's very likely that the number of instances of all these tools is going to skyrocket. It's very likely that the number of instances of Synopsys Design Compiler is going to skyrocket, along with the number of agents using the floor planners, our layout tools, and our design-rule checkers.

Today we're limited by the number of engineers. Tomorrow, those engineers are going to be supported by a bunch of agents. We're going to be exploring the design space like you've never seen before, and we're going to use the tools that we use today.

I think tool use is going to cause the software companies to skyrocket. The reason why it hasn't happened yet is because the agents aren't good enough at using their tools yet. Either these companies are going to build the agents themselves, or agents are going to get good enough to be able to use those tools. I think it's going to be a combination of both.

Dwarkesh Patel

I think in your latest filings, you had almost $100 billion in purchase commitments with foundries, memory, and packaging. SemiAnalysis has reported that you will have $250 billion of these kinds of purchase commitments. Is Nvidia's moat really just the fact that you've locked up many years of these scarce components? Somebody else might have an accelerator, but can they actually get the memory to build it? Can they actually get the logic to build it? Is this really Nvidia's big moat for the next few years?

Jensen Huang

It's one of the things that we can do that is hard for someone else to do. We've made enormous commitments upstream. Some of it is explicit—these commitments that you mentioned. Some of it is implicit.

For example, a lot of the investments that are upstream are made by our supply chain because I said to the CEOs, “Let me tell you how big this industry is going to be. Let me explain to you why. Let me reason through it with you, and let me show you what I see.”

As a result of that process of informing, inspiring, and aligning with CEOs of all different industries upstream, they're willing to make the investments. Why are they willing to make the investments for me and not someone else?

The reason for that is because they know that I have the capacity to buy their supply and sell it through my downstream. The fact is that Nvidia's downstream supply chain and our downstream demand are so large, they're willing to make the investment upstream.

If you look at GTC, people are marveled by the scale of it and the people that go. It's a full 360 degrees—the entire universe of AI, all in one place. They're all in one place because they need to see each other.

I bring them together so that the downstream can see the upstream, the upstream can see the downstream, and all of them can see the advances in AI. Very importantly, they can all meet the AI natives—all the AI startups being built and all the amazing things happening—so they can see firsthand all the things that I tell them.

I spend a lot of my time informing, directly or indirectly, our supply chain, partners, and ecosystem about the opportunity in front of us. Some people always say, “Jensen, in most keynotes, it's one announcement after another.”

With our keynotes, there's always a part of it that's a little torturous, in the sense that it almost comes across like education. In fact, that's exactly on my mind. I need to make sure the entire supply chain, upstream and downstream, and the ecosystem understand what is coming at us, why it's coming, when it's coming, how big it's going to be, and are able to reason about it systematically, just like I reason about it.

Regarding the moat as you describe it, we're able to build for a future. If our next several years are a trillion dollars in scale, we have the supply chain to do it.

Without our reach, the velocity of our business—just as there's cash flow, there's supply-chain flow, there's churn. Nobody is going to build a supply chain for an architecture if the business churn is low. Our ability to sustain the scale is only because our downstream demand is so great.

And they see it, they hear about it, they see it all coming. That allows us to do the things we're able to do at the scale we do them.

Dwarkesh Patel

I do want to understand more concretely whether the upstream can keep up. For many years now, you guys have been 2X-ing revenue year over year. You've been more than tripling the amount of FLOPs you're providing to the world year over year, and 2X-ing at this scale now is really incredible.

Jensen Huang

Exactly.

Dwarkesh Patel

But then you look at logic. You're the biggest customer on TSMC's N3 node, and you're one of the biggest on N2. AI as a whole this year is going to be 60% of N3. It's going to be 86% next year, according to SemiAnalysis.

How do you double if you're the majority? And how do you do that year over year? Are we in a regime now where the growth rate in AI compute has to slow because of upstream? Do you see a way to get around this? How do we build 2X more fabs year over year, ultimately?

Jensen Huang

At some level, the instantaneous demand is greater than the supply upstream and downstream in the world. At any instant, we could be limited by the number of plumbers, which actually happens. The plumbers are invited to next year's GTC.

By the way, great idea. But that's a good condition. You want an industry where the instantaneous demand is greater than the total supply of the industry. The opposite is obviously less good.

If we're too far apart, if one particular component is too far away, the industry swarms it. For example, notice people aren't talking very much about CoWoS anymore. The reason for that is because for 2 years we swarmed the living daylights out of it.

We doubled, doubled, doubled on several doubles. Now I think we're in fairly good shape. TSMC now knows that CoWoS supply has to keep up with the rest of the logic demand and the memory demand.

They're scaling CoWoS and future packaging technologies at the same level as they scale logic. This is terrific, because for a long time, CoWoS and HBM memory were rather specialty. But they're not specialties anymore.

People now realize they're mainstream computing technology. Of course, we're now much more able to influence a larger scope of our supply chain. At the beginning of the AI revolution, all the things that I say now, I was saying 5 years ago. Some people believed in it and invested in it—for example, Sanjay and the Micron team.

I still remember the meeting really well where I was clear about exactly what was going to happen, why it was going to happen, and the predictions of today. They really doubled down on it. We partnered with them across LPDDR and HBM memories, and they really invested in it. It obviously has been tremendous for the company.

Some people came a little bit later, but now they're all here. Each one of these bottlenecks gets a great deal of attention. Now we're prefetching the bottlenecks years in advance. For example, the investments that we've done with Lumentum, Coherent, and the silicon photonics ecosystem over the last several years really reshaped the supply chain.

We built up an entire supply chain around TSMC. We partnered with them on COUPE, invented a whole bunch of technology, and licensed those patents to the supply chain to keep it nice and open. We're preparing the supply chain through the invention of new technologies, new workflows, new testing equipment like double-sided probing, investing in companies, and helping them scale up their capacity.

You can see that we're trying to shape the ecosystem so that the supply chain is ready to support the scale.

Dwarkesh Patel

It seems like some bottlenecks are easier than others. Scaling up CoWoS versus scaling up—

Jensen Huang

I went to the hardest one, by the way.

Dwarkesh Patel

Which is?

Jensen Huang

Plumbers. Plumbers and electricians. This is one of the concerns that I have about the doomers describing the end of work and killing of jobs. If we discourage people from being software engineers, we're going to run out of software engineers.

The same prediction happened 10 years ago. Some of the doomers were telling people, “Whatever you do, don't be a radiologist.” You might hear some of those videos still on the web saying radiology is going to be the first career to go and the world is not going to need any more radiologists.

Guess what we're short of? Radiologists.

Dwarkesh Patel

Going back to this point about how some things you can scale, and other things—how do you actually manufacture 2x the amount of logic a year?

Jensen Huang

Ultimately, memory and logic are bottlenecked by EUV. How do you get to 2x as many EUV machines year over year? None of that is impossible to scale quickly. All of that is easy to do within 2 or 3 years. You just need a demand signal.

Once you can build 1, you can build 10, and once you can build 10, you can build 1 million. These things are not hard to replicate.

Dwarkesh Patel

How far down the supply chain do you go? Do you go to ASML and say, “Hey, if I look out 3 years from now, for Nvidia to be generating $2 trillion a year in revenue, we need way more EUV machines”?

Jensen Huang

Some of them I have to directly, some of them indirectly, and some of them—if I can convince TSMC, ASML will be convinced. We have to think about the critical pinch points. But if TSMC is convinced, you'll have plenty of EUV machines in a few years.

My point is that none of the bottlenecks last longer than a couple of years—2 or 3 years, none of them. Meanwhile, we're improving computing efficiency by 10x, 20x, and in the case of Hopper to Blackwell, 30x to 50x.

We're coming up with new algorithms because CUDA is so flexible. We're developing all kinds of new techniques so that we drive efficiency in addition to increasing capacity. None of those things worry me. It's the stuff that's downstream from us: energy policies.

You can't create an industry without energy. You can't create a whole new manufacturing industry without energy. We want to reindustrialize the United States. We want to bring back chip manufacturing, computer manufacturing, and packaging.

We want to build new things like EVs and robots. We want to build AI factories. You can't build any of these things without energy, and those things take a long time. More chip capacity—that's a 2–3-year problem. More CoWoS capacity, 2–3-year problem.

Dwarkesh Patel

Interesting. I feel like I have guests tell me the exact opposite thing sometimes. In this case, I just don't have the technical knowledge to adjudicate.

Jensen Huang

The beautiful thing is you're talking to the expert.

1. Will TPUs break Nvidia’s hold on AI compute?

Dwarkesh Patel

True. I want to ask about your competitors. If you look at the TPU, arguably 2 out of the top 3 models in the world, Claude and Gemini, were trained on TPU. What does that mean for Nvidia going forward?

Jensen Huang

We build a very different thing. What Nvidia built is accelerated computing, not a tensor processing unit. Accelerated computing is used for all kinds of things: molecular dynamics, quantum chromodynamics, data processing, data frames, structured data, and unstructured data. It's also used for fluid dynamics and particle physics. In addition, we use it for AI.

Accelerated computing is much more diverse. Although AI is the conversation today and is obviously very important and impactful, computing is much broader than that. Nvidia has reinvented the way computing is done, moving from general-purpose computing to accelerated computing.

Our market reach is far greater than any TPU or ASIC can possibly have. If you look at our position, we're the only company that accelerates applications of all kinds. We have a gigantic ecosystem. All kinds of frameworks and algorithms run on Nvidia.

Because our computers are designed to be operated by other people, anyone who's an operator can buy our systems. With most of these home-built systems, you have to be your own operator because they were never designed to be flexible enough for others to operate.

Because anybody can operate our systems, we're in every cloud, including Google, Amazon, Azure, and OCI. If you want to operate it to rent, you better have a large ecosystem of customers in many industries to be the offtakers. If you want to operate it for yourself, we obviously have the ability to help you operate it yourself, like we did for Elon with xAI.

And because we can enable operators in any company and any industry, you could use it to build a supercomputer for scientific research and drug discovery at Lilly. We can help them operate their own supercomputer and use it for the entire diversity of drug discovery and biological sciences that we accelerate.

There are just a whole bunch of applications that we can address that you can't do with TPUs. Nvidia built CUDA to be a fantastic tensor processing unit as well, but it also handles every life cycle of data processing, computing, AI, and so on.

Our market opportunity is just a lot larger, and our reach is a lot greater. Because we support every application in the world now, you can build Nvidia systems anywhere and know that there will be customers for it. It's a very different thing.

Dwarkesh Patel

This is going to be a long question. You have spectacular revenue, and you're not making $60 billion a quarter from pharma and quantum. You're making it because AI is an unprecedented technology that is growing unprecedentedly fast. The question then is what's best for AI specifically.

I'm not in the details, but I talk to my AI researcher friends and they say, “Look, when I use a TPU, it's this big systolic array that's perfect for doing matrix multiplies, whereas a GPU is very flexible. It's great when you have lots of branching or irregular memory access.”

But what is AI? It's just these very predictable matrix multiplies again and again and again. You don't have to give up any die area for warp schedulers or switches between threads and memory banks. And the TPU is really optimized for the bulk of this growth in revenue and use case for compute that is coming online right now. I wonder how you react to that.

Jensen Huang

Matrix multiplies are an important part of AI, but they're not the only part. If you want to come up with a new attention mechanism, disaggregate in a different way, or invent a whole new type of architecture altogether—like a hybrid SSM—you want an architecture that's generally programmable.

If you want to create a model that fuses diffusion and autoregressive techniques, you want an architecture that's just generally programmable. We run everything you can imagine. That's the advantage. It allows for the invention of new algorithms a lot more easily because it's a programmable system.

The ability to invent new algorithms is really what makes AI advance so quickly. TPUs, like anything else, are impacted by Moore's Law, which we know is increasing by about 25% per year.

The only way to really get 10x or 100x leaps is to fundamentally change the algorithm and how it’s computed every single year. That’s NVIDIA’s fundamental advantage. The only reason we were able to make Blackwell 50x more energy efficient than Hopper—when I first announced Blackwell was going to be 35x more energy efficient than Hopper, nobody believed it. Then Dylan wrote an article saying I sandbagged, and it’s actually 50 times.

You can’t reasonably do that with just Moore’s Law. The way we solve that problem is with new models, like MoEs, that are parallelized, disaggregated, and distributed across a computing system. Without the ability to really get down and come up with new kernels with CUDA, it’s really hard to do. It’s the combination of the programmability of our architecture and the fact that NVIDIA is an extreme co-design company.

We can even offload some of the computation into the fabric itself, like NVLink, or into the network with Spectrum-X. We could effect change across the processors, the system, the fabric, the libraries, and the algorithm simultaneously. Without CUDA to do that, I wouldn’t even know where to start.

Dwarkesh Patel

This gets at an interesting question about NVIDIA’s clientele. 60% of your revenue is coming from these big 5 hyperscalers. In a different era with different customers—let’s say professors running experiments—they need CUDA. They can’t use another accelerator. They just needed to run PyTorch with CUDA and have everything optimized.

But these hyperscalers have the resources to write their own kernels. In fact, they have to in order to get that last 5% of performance they need for their specific architecture. Anthropic and Google are mostly running their own accelerators or running TPUs and Trainium. But even OpenAI, using GPUs, has Triton because they need their own kernels.

Down to CUDA C++, instead of using cuBLAS and NCCL, they’ve got their own stack, which compiles to other accelerators as well. If most of your customers can and do make replacements for CUDA, to what extent is CUDA really the thing that is going to make frontier AI happen on NVIDIA?

Jensen Huang

CUDA is a rich ecosystem. If you want to build on any computer, building on CUDA first is incredibly smart. Because the ecosystem is so rich, we support every framework. If you want to create custom kernels—for example, we contribute enormously to Triton—the back end of Triton has huge amounts of NVIDIA technology. We’re delighted to help every framework become as great as it can be.

There are lots and lots of frameworks. There’s Triton, vLLM, SGLang, and more. Now there’s a whole bunch of new reinforcement learning frameworks coming out, like verl and NeMo RL. With post-training and reinforcement learning, that entire area is just exploding.

If you want to build on an architecture, building on CUDA makes the most sense because you know the ecosystem is great. You know that if something happens, it’s more likely in your code and not in the mountain of code underneath. Don’t forget the amount of code you’re dealing with when building these systems. When something doesn’t work, was it you or was it the computer?

You would like it to always be you and to be able to trust the computer. Obviously, we still have lots of bugs ourselves, but our system is so well wrung out that you can at least build on top of the foundation. That’s number 1: the richness, programmability, and capability of the ecosystem.

The second thing is, if you’re a developer building anything at all, the single most important thing you want is an install base. You want the software you write to run on a whole bunch of other computers. You’re not building software just for yourself. You’re building it for your fleet or everybody else’s fleet because you’re a framework builder.

NVIDIA’s CUDA ecosystem is ultimately its great treasure. We have several hundred million GPUs out there now. Every cloud has them. It goes back to the A10, A100, H100, H200, the L series, the P series. There’s a whole bunch of them. They’re in all kinds of sizes and shapes.

If you’re a robotics company, you want that CUDA stack to actually run in the robot itself. We’re literally everywhere. The install base means that once you develop the software or the model, it’s going to be useful everywhere. That is just incredibly valuable.

Lastly, the fact that we’re in every single cloud makes us genuinely unique. If you’re an AI company or developer, you’re not exactly sure which cloud service provider you’re going to partner with or where you’d like to run it. We run everywhere, including on-premises for you if you like. The combination of the richness of the ecosystem, the expansiveness of the install base, and the versatility of where we are makes CUDA invaluable.

Dwarkesh Patel

That makes a lot of sense. The thing I’m curious about is whether those advantages matter a lot to your main customers. There are many people for whom they might matter. The kind of person who can actually build their own software stack makes up most of your revenue.

Especially if you go to a world where AI is getting especially good at the things which have tight verification loops, where you can do RL on them, this question of how you write a kernel that does attention or MLP the most efficiently across a scale-up is a very verifiable sort of feedback loop. Can all the hyperscalers write these custom kernels for themselves?

NVIDIA still has great price performance, so they might still prefer to use NVIDIA. But then the question is, does it just become a question of who is offering the best specs—the best FLOPs and memory bandwidth for a given dollar? Whereas historically NVIDIA has just had, and still has, the best margins in all of AI across hardware and software—over 70%—because of this CUDA moat. And the question is, can you sustain those margins if, for most of your customers, they can actually afford to build instead of relying on the CUDA moat?

Jensen Huang

The number of engineers we have assigned to these AI labs is insane, working with them and optimizing their stack. The reason for that is because nobody knows our architecture better than we do. These architectures are not as general-purpose as a CPU.

A CPU is kind of like a Cadillac. It’s a nice cruiser. It never goes too fast. Everybody drives it pretty well. It’s got cruise control, and everything’s easy. But in a lot of ways, NVIDIA’s GPUs and accelerators are like F1 racers. I could imagine everybody’s able to drive it at 100 miles an hour, but it takes quite a bit of expertise to be able to push it to the limit.

We use a ton of AI to create the kernels that we have. I’m pretty sure we’re going to still be needed for quite some time. Our expertise helps our AI lab partners to get another 2x out of their stack easily, oftentimes. It’s not unusual that by the time we’re done optimizing their stack or optimizing a particular kernel, their model sped up by 3x, 2x, or 50%.

That’s a huge number, especially when you’re talking about the install base of the fleet that they have, of all the Hoppers and Blackwells that they have. When you increase it by a factor of 2, that doubles the revenues. That directly translates to revenues.

NVIDIA’s computing stack is the best performance per TCO in the world, bar none. Nobody can demonstrate to me that any single platform in the world today has a better performance-TCO ratio. Not one company. In fact, the benchmarks that are out there—Dylan’s InferenceMAX is sitting out there for everybody to use—and not one… TPU won’t come, Trainium won’t come. I encourage them to use InferenceMAX and demonstrate their incredible inference cost.

It’s really hard. Nobody wants to show up on MLPerf. I would welcome Trainium to demonstrate the 40% that they claim all the time. I would love to hear them demonstrate the cost advantage of TPUs. It makes no sense in my mind. It makes absolutely zero sense. On first principles, it makes no sense.

So I think the reason why we’re so successful is simply because our TCO is so great. Secondly, you say 60% of our customers are the top 5, but most of that business is external. For example, most of NVIDIA in AWS is for external customers, not internal use. Most of our customers at Azure—obviously, all of our customers are external. All of our customers at OCI are external, not internal use.

The reason why they favor us is because our reach is so great. We can bring them all of the great customers in the world. They’re all built on NVIDIA. And the reason why all these companies are built on NVIDIA is because our reach and our versatility are so great.

So I think the flywheel is really the installed base, the programmability of our architecture, the richness of our ecosystem, and the fact that there are so many AI companies in the world. There are tens of thousands of them now. If you were one of those AI startups, what architecture would you choose? You would choose an architecture that’s most abundant. We’re the most abundant in the world.

You’d choose the one that has the largest installed base. We’re the largest installed base. And you’d choose the one that has a rich ecosystem. So that’s the flywheel. That’s the reason why, between the combination of, one, our performance per dollar is so great that they have the lowest-cost tokens, and, second, our performance per watt is the highest in the world.

So if one of these companies—if our partners—built a 1-gigawatt data center, that 1-gigawatt data center better deliver the maximum amount of revenues and number of tokens, which directly translates to revenues. You want it to generate as many tokens as possible and maximize the revenues for that data center. We are the highest-tokens-per-watt architecture in the world.

Lastly, if your goal is to rent the infrastructure, we have the most customers in the world. So that’s the reason why the flywheel works.

Dwarkesh Patel

Interesting. I guess the question comes down to: What is the actual market structure here? Because even if there are other companies, there could have been a world where there were tens of thousands of AI companies that had roughly equal shares of compute.

But even through these 5 hyperscalers, really, the people on Amazon using the compute are Anthropic, OpenAI, and these big foundation labs who can themselves afford and have the ability to make different accelerators work.

Jensen Huang

No, I think your premise is wrong.

Dwarkesh Patel

Maybe. But let me ask you a slightly different question.

Jensen Huang

Come back and make me correct your premise.

Dwarkesh Patel

Okay. Let me just ask you a different question. But still make sure to make me come back and fix it, because it’s just too important to AI. It’s too important to the future of science. It’s too important to the future of the industry. That premise—look, let me just finish the question, and then we can address it together.

Jensen Huang

Yeah.

Dwarkesh Patel

If all these things about price, performance, and performance per watt are true, why do you think it is the case that, say, Anthropic, for example, just announced a couple of days ago that they have a multigigawatt deal with Broadcom and Google for TPUs and the majority of their compute? Obviously, for Google, TPU is the majority of its compute.

So if I look at these big AI companies, it seems like a lot of their compute—there was some point where it was all NVIDIA, and now it’s not. So I’m curious how to square this: If these things are true on paper, why are they going with other accelerators?

Jensen Huang

Anthropic is a unique instance, not a trend. Without Anthropic, why would there be any TPU growth at all? It’s 100% Anthropic. Without Anthropic, why would there be Trainium growth at all? It’s 100% Anthropic. I think that’s fairly well known and well understood.

It’s not that there’s an abundance of ASIC opportunities. There’s only 1 Anthropic.

Dwarkesh Patel

But OpenAI’s deals with AMD—they’re building their own Titan accelerator.

Jensen Huang

Yeah, but I think we could all acknowledge they’re vastly NVIDIA. We’re going to still do a lot of work together. I’m not offended by other people using something else and trying things. If they don’t try these other things, how would they know how good ours is? Sometimes you’ve got to be reminded of it.

We have to continuously earn the position that we’re in. There are always big claims. Look at the number of ASICs that have been canceled. Just because you’re going to build an ASIC, you still have to build something better than NVIDIA. It’s not that easy building something better than NVIDIA. It’s not sensible, actually.

NVIDIA’s got to be missing something, seriously. Because of our scale and our velocity, we’re the only company in the world that’s cranking it out every single year—big leaps, every single year.

Dwarkesh Patel

I guess their logic is, “Hey, it doesn’t need to be better. It just needs to be not more than 70% worse,” because they’re paying you 70% margins.

Jensen Huang

No, don’t forget: Even in ASICs, margins are really quite high. NVIDIA’s margin is 70%, let’s say, but ASIC margins are 65%. What are you really saving?

Dwarkesh Patel

Oh, you mean from Broadcom or something like that?

Jensen Huang

Yeah, sure. You’ve got to pay somebody. I think the ASIC margins are incredibly good, from what I can tell. They believe it, too. They’re quite proud of their incredible ASIC margins.

So, you asked the question why. A long time ago, we just didn’t have the ability to do it. At the time, I didn’t deeply internalize how difficult it would be to build a foundation AI lab like OpenAI and Anthropic, and the fact that they needed huge investments from the supplier themselves.

We just weren’t in a position to make the multibillion-dollar investment into Anthropic so that they could use our compute. But Google and AWS were. They put in huge investments in the beginning so that Anthropic, in return, used their compute. We just weren’t in a position to do that at the time.

I would say my mistake is that I didn’t deeply internalize that they really had no other options, that a VC would never put in $5–10 billion of investment into an AI lab with the hopes of it turning out to be Anthropic. So that was my miss. But even if I had understood it, I don’t think we would’ve been in a position to do that at the time.

But I’m not going to make that same mistake again. I’m delighted to invest in OpenAI, and I’m delighted to help them scale, and I believe it’s essential to do so. And then, when I was able to—when Anthropic came to us—I was delighted to be an investor and delighted to help them scale. We just weren’t, at the time, able to do it.

If I could rewind everything—and NVIDIA could have been as big back then as we are now—I would’ve been more than happy to do it.

2. Why doesn’t Nvidia become a hyperscaler?

Dwarkesh Patel

This is actually quite interesting. For many years, NVIDIA has been the company in AI making money, making lots of money. Now you’re investing it. It’s been reported that you’ve done up to $30 billion in OpenAI and $10 billion in Anthropic.

But now their valuations have increased, and I’m sure they’ll continue to increase. So if, over these many years, you were giving them the compute, you saw where it was headed, and they were worth one-tenth what they’re worth now—a couple of years ago, or even a year ago in some cases—and you had all this cash, there’s a world where either NVIDIA itself becomes a foundation lab, does a huge investment to make that possible, or makes the deals you’ve made now at current valuations much earlier on. And you had the cash to do it. So I’m curious, actually: Why not have done it earlier?

Jensen Huang

We did it as soon as we could have, and if I could have, I would’ve done it even earlier. At the time that Anthropic needed us to do it, we just weren’t in a position to do it. It wasn’t in our sensibility to do so.

Dwarkesh Patel

How so? Was it a cash thing?

Jensen Huang

Yeah, the level of investment. We had never invested outside the company at the time, and not that much. We didn’t realize we needed to.

I always thought that they could just go raise from VCs, for God’s sake, like all companies do. But what they were trying to do couldn’t have been done through VCs. What OpenAI wanted to do couldn’t have been done through VCs. I recognize that now. I didn’t know it then.

But that’s their genius. That’s why they’re smart. They realized then that they had to do something like that. And I’m delighted that they did.

Even though we caused Anthropic to have to go to somebody else, I’m still happy that it happened. Anthropic’s existence is great for the world.

I'm delighted for it.

Dwarkesh Patel

I guess you still are making a ton of money, and you're making way more money quarter after quarter. It's still okay to have regrets. So the question still arises: Okay, now that we're here and you have all this money that you keep making, what should NVIDIA be doing with it?

There's one answer, which is that there's this whole middleman ecosystem that has popped up for converting CapEx into OpEx for these labs so that they can rent compute. Because the chips are really expensive, they make a lot of money over their lifetime because the AI models are getting better. So the value that they generate—their tokens—is increasing, but they're expensive to set up.

NVIDIA has the money to do the CapEx. In fact, it's been reported you are backstopping CoreWeave up to $6.3 billion and have invested $2 billion. Why doesn't NVIDIA become a cloud themselves? Why doesn't it become a hyperscaler themselves and rent this compute out? You have all this cash to do it.

Jensen Huang

This is a philosophy of the company, and I think it's wise. We should do as much as needed, as little as possible. What that means is, the work that we do with building our computing platform—if we don't do it, I genuinely believe it doesn't get done.

If we didn't take the risk that we take—if we didn't build NVLink the way we built it, if we didn't build the whole stack, if we didn't create the ecosystem the way we did, if we didn't dedicate ourselves to 20 years of CUDA while losing money most of that time—if we didn't do it, nobody else would have done it.

If we didn't create all the CUDA-X libraries so that they're all domain-specific, a decade and a half ago, we pushed into domain-specific libraries because we realized that if we didn't create these domain-specific libraries—whether it's for ray tracing or image generation or even the early works of AI, these models, or for data processing, structured data processing, or vector data processing—nobody would.

I am completely certain of that. We created a library for computational lithography called cuLitho. If we didn't create it, nobody would have. So accelerated computing wouldn't advance the way it has if we didn't do what we did. We should dedicate our company, all of our might, wholeheartedly to go do that.

However, the world has lots of clouds. If I didn't do it, somebody would show up. So, following the recipe—the philosophy—of doing as much as needed but as little as possible, that philosophy exists in our company today. Everything I do, I do it with that lens.

In the case of clouds, if we didn't support CoreWeave to exist, these neoclouds, these AI clouds, wouldn't exist. If we didn't help CoreWeave exist, they would not exist. If we didn't support Nscale, they wouldn't be where they are today. If we didn't support Nebius, they wouldn't be what they are today. Now they're doing fantastically.

Dwarkesh Patel

Is that a business model?

Jensen Huang

We should do as much as needed, as little as possible. So we invest in our ecosystem because I want our ecosystem to thrive. I want the architecture and AI to be able to connect with as many industries as possible, as many countries as possible, and make it possible for the planet to be built on AI and to be built on the American tech stack. That vision is exactly what we're pursuing.

Now, one of the things that you mentioned: There are so many great, amazing foundation model companies, and we try to invest in all of them. This is another thing that we do. We don't pick winners. We need to support everyone. It's part of our joy of doing so. It's imperative to our business.

But we also go out of our way not to pick winners. So when I invest in one of them, I invest in all of them.

Dwarkesh Patel

Why do you go out of your way not to pick winners?

Jensen Huang

Because it's not our job to, number one. Number two, when NVIDIA first started, there were 60 3D graphics companies. We are the only one that survived. If you would have taken those 60 graphics companies and asked yourself which one was going to make it, NVIDIA would be at the top of that list not to make it.

This is long before you, but NVIDIA's graphics architecture was precisely wrong. It's not a little bit wrong. We created an architecture that was precisely wrong, and it was an impossible thing for developers to support. It was never going to make it.

We reasoned about it from good first principles, but we ended up with the wrong solution. Everybody would have counted us out. And here we are. So I have enough humility to recognize that. Don't pick winners. Either let them all take care of themselves, or take care of all of them.

Dwarkesh Patel

One thing I didn't understand is you said, “Look, we're not prioritizing these neoclouds just because they are neoclouds and we want to prop them up.” But you also listed a bunch of neoclouds and said they wouldn't exist if it wasn't for NVIDIA. How are those 2 things compatible?

Jensen Huang

First of all, they need to want to exist, and they come to ask us for help. When they want to exist and they have a business plan, expertise, and the passion for it, they obviously have to have some capabilities themselves. But if, at the end of the day, they need some investment in order to get it off the ground, we would be there for them.

Your question was, “Do we want to be in the financing business?” The answer is no. There are people in the financing business, and we'd rather work with all the people in the financing business than be a financier ourselves. Our goal is to focus on what we do, keep our business model as simple as possible, and support our ecosystem.

When someone like OpenAI needs an investment of a $30 billion scale because it's still before their IPO, and we deeply believe in them—and I deeply believe that they're an extraordinary company already today—they're going to be an incredible company. The world needs them to exist. The world wants them to exist. I want them to exist. They have the wind at their back. Let's support them and let them scale.

We'll do those investments because they need us to do it. But we're not trying to do as much as possible. We're trying to do as little as possible.

Dwarkesh Patel

This may be an obvious question, but we've lived many years in this situation where there's a shortage of GPUs, and it's grown now because models are getting better. We have a shortage of GPUs.

Jensen Huang

Yes.

Dwarkesh Patel

NVIDIA is known for divvying up the scarce allocation, not just based on the high bidder, but rather on, “Hey, we want to make sure that these neoclouds exist. Let's give some to CoreWeave, let's give some to Crusoe, let's give some to Lambda.” Why is it good for NVIDIA? First of all, would you agree with this characterization of fracturing the market?

Jensen Huang

No. No. Your premise is just wrong. We're sufficiently mindful about these things. We're very mindful about these things.

First of all, if you don't place a PO, all the talking in the world won't make a difference. Until we get a PO, what are we going to do? So the first thing is, we work really hard with everybody to get a forecast done, because these things take a long time to build, and the data centers take a long time to build.

We align ourselves with demand and supply and things like that through forecasting. Okay? That's job number 1. Number 2, we've tried to forecast with as many people as possible, but in the final analysis, you still have to place an order. Maybe, for whatever reason, you didn't place your order.

What can I do? At some point, first in, first out. But beyond that, if you're not ready because your data center's not ready, or certain components aren't ready to enable you to stand up a data center, we might decide to serve another customer first. That's just maximizing the throughput of our own factory. We might do some adjustments there.

Aside from that, the prioritization is first in, first out. You've got to place a PO. If you don't place a PO… Now, of course, there are stories about that. For example, all of this kind of started from an article about Larry and Elon Musk having dinner with me where they begged for GPUs. That never happened.

We absolutely had dinner, and it was a wonderful dinner. At no time did they beg for GPUs. They just had to place an order. Once they place an order, we do our best to get the capacity to them. We're not complicated.

Dwarkesh Patel

Okay. So it sounds like there's a queue, and then based on whether your data center is ready and when you place a purchase order, you get them at a certain time. But it still doesn't sound like the highest bidder just gets it. Is there a reason to do it…?

Jensen Huang

We never do that.

Dwarkesh Patel

Okay.

Jensen Huang

We never do.

Dwarkesh Patel

Why not just do high bidder?

Jensen Huang

Because it's a bad business practice. You set your price, and then people decide to buy it or not. I understand that others in the chip industry change their prices when demand is higher, but we just don't. That's just never been a practice of ours.

You can count on us. I prefer to be dependable, to be the foundation of the industry. You don't need to second-guess. If I quoted you a price, we quoted you a price. That's it. If demand goes through the roof, so be it.

On the other end, that's why you have a productive relationship with TSMC, right?

Dwarkesh Patel

Yeah. NVIDIA's been in business with them for, I guess, coming up on 30 years.

NVIDIA and TSMC don't have a legal contract. There's always some rough justice. Sometimes I'm right, sometimes I'm wrong. Sometimes I got a better deal, sometimes I got a worse deal. But overall, the relationship is incredible. I can completely trust them. I can completely depend on them.

One of the things you can count on with NVIDIA is that this year, Vera Rubin is going to be incredible. Next year, Vera Rubin Ultra will come. The year after that, Feynman will come. And the year after that, I haven't introduced the name yet.

Every single year, you can count on us. You're going to have to go find another ASIC team in the world—pick your ASIC team—where you can say, “I can bet the farm, I can bet my entire business that you will be here for me every single year. Your token cost will decrease by an order of magnitude every single year. I can count on it like I can count on the clock.”

I just said something about TSMC. For no other foundry in history can you possibly say that. You can say that about NVIDIA today. You can count on us every single year.

If you would like to buy $1 billion worth of AI factory compute, no problem. If you'd like to buy $100 million, no problem. You'd like to buy $10 million, or just 1 rack, not a problem. Or just 1 graphics card, okay, no problem.

If you would like to place an order for $100 billion of AI factory, no problem. We're the only company in the world where you can say that today. I can say that about TSMC as well. I want to buy 1, buy $1 billion, no problem.

We just have to go through the process of planning for it, and all the things that mature people do. So I think this ability for NVIDIA to be the foundation of the world's AI industry, this is a position that has taken us a couple of decades to arrive at. Enormous commitment, enormous dedication. The stability of our company, the consistency of our company, is really important.

3. Should we be selling AI chips to China?

Okay. I want to ask about China. I actually don't know what I think about whether it's good to sell chips to China or not, but I like to play devil's advocate against my guests. So when Dario was on, who supports export controls, I asked him, why can't America and China both have a country of geniuses in the data center?

But since you're on the opposite side, I'll ask you in the opposite way. One way to think about it is, Anthropic actually announced a couple days ago Mythos Preview. This model, Mythos, they're not even releasing publicly because they say it has such cyber-offensive capabilities that we don't think the world is ready until we make sure these zero-days are patched up.

But they say it found thousands of high-severity vulnerabilities across every major operating system, every browser. It found one in OpenBSD, which is this operating system that's been specifically designed not to have zero-days. It found one that's existed for 27 years.

So if Chinese companies and Chinese labs and the Chinese government had access to the AI chips to train a model like Claude Mythos with these cyber-offensive capabilities and run millions of instances of it with more compute, the question is, is that a threat to American companies, to American national security?

Jensen Huang

First of all, Mythos was trained on fairly mundane capacity, and a fairly mundane amount of it, by an extraordinary company. The amount of capacity and the type of compute it was trained on is abundantly available in China.

So you just have to first realize that chips exist in China. They manufacture 60% of the world's mainstream chips, maybe more. It's a very large industry for them. They have some of the world's greatest computer scientists.

As you know, most of the AI researchers in all of these AI labs are Chinese. They have 50% of the world's AI researchers. So the question is, considering all the assets they already have—they have an abundance of energy, they have plenty of chips, they've got most of the AI researchers—if you're worried about them, what is the best way to create a safe world?

Victimizing them, turning them into an enemy, likely isn't the best answer. They are an adversary. We want the United States to win. But I think having a dialogue and having research dialogue is probably the safest thing to do.

This is an area that is glaringly missing because of our current attitude about China as an adversary. It is essential that our AI researchers and their AI researchers are actually talking. It is essential that we try to both agree on what not to use the AI for.

With respect to finding bugs in software, of course, that's what AI is supposed to do. Is it going to find bugs in a lot of software? Of course. There are lots and lots of bugs. There are lots of bugs in the AI software. That's what AI is supposed to do, and I'm delighted that AI has reached a level where it could help us be so much more productive.

One of the things that is underemphasized is the richness of the ecosystem around cybersecurity, AI cybersecurity, AI security, AI privacy, and AI safety. There's a whole ecosystem of AI startups that are trying to create this future for us, where you have 1 AI agent that's incredible, surrounded by thousands of AI agents, keeping it safe, keeping it secure.

That future surely is going to happen. The idea that you're going to have an AI agent running around with nobody watching after it is kind of insane. We know very well that this ecosystem needs to thrive.

It turns out this ecosystem needs open source. This ecosystem needs open models. They need open stacks so that all of these AI researchers and all these great computer scientists can go build AI systems that are as formidable and can keep AI safe.

So one of the things that we need to make sure that we do is we keep the open-source ecosystem vibrant. That can't be ignored. A lot of that is coming out of China. We ought to not suffocate that.

With respect to China, of course we want the United States to have as much computing as possible. We're limited by energy, but we've got a lot of people working on that. We've got to not make energy a bottleneck for our country.

But what we also want is to make sure that all the AI developers in the world are developing on the American tech stack, and making the contributions, the advancements of AI—especially when it's open source—available to the American ecosystem.

It would be extremely foolish to create 2 ecosystems: the open-source ecosystem, and it only runs on a foreign tech stack, and a closed ecosystem that runs on the American tech stack. I think that would be a horrible outcome for the United States.

Dwarkesh Patel

Since there are a lot of things, let me just triage the response. I think the concern, going back to the FLOP difference in the hacking, is yes, they have compute, but there are some estimates that because they're at 7 nm—they don't have EUVs because of chipmaking export controls—the amount of FLOPs they're able to actually produce, they have 1/10 the amount of FLOPs that the US has.

So with that, could they eventually train a model like Mythos? Yes.

But the question is, because we have more FLOPs, American labs are able to get to these levels of capabilities first. Because Anthropic got to it first, they say, “Okay, we’re going to hold onto it for a month while all these American companies—we’ll give them access to it. They’re going to patch up all their vulnerabilities, and now we release it.”

Furthermore, even if they train a model like this, the ability to deploy it at scale matters. If you had a cyber hacker, it’s much more dangerous if they have 1,000,000 of them versus 1,000 of them. So that inference compute really matters a lot.

In fact, the fact that they have so many AI researchers who are so good is the thing that makes it so scary, because what is it that makes those engineer-researchers more productive? It’s compute. If you talk to any AI lab in America, they say the thing that’s bottlenecking them is compute. There are quotes from the DeepSeek founder, or Qwen leadership, or whatever. They say the thing they’re bottlenecked on is compute.

So then the question is, isn’t it better that we get American companies, because they have more compute, to get to the Mythos-level capabilities first, and prepare our society for it before China can get to it because they have less compute? We should always be first, and we should always have more.

Jensen Huang

But in order for that outcome you described to be true, you have to take it to the extremes. They have to have no compute. If they have some compute, the question is how much is needed?

The amount of compute they have in China is enormous. You’re talking about the country that is the 2nd-largest computing market in the world. If they want to aggregate their compute, they’ve got plenty of compute to aggregate.

Dwarkesh Patel

But is that true? People do these estimates, and they’re like, “SMIC is actually behind on the process nodes.”

Jensen Huang

I’m about to tell you. The amount of energy they have is incredible. Isn’t that right? AI is a parallel computing problem, isn’t it? Why can’t they just put 4x, 10x as many chips together because energy’s free? They have so much energy.

They have datacenters that are sitting completely empty, fully powered. You know they have ghost cities; they have ghost datacenters too. They have so much infrastructure capacity.

If they wanted to, they could just gang up more chips, even if they’re 7nm. Their capacity for building chips is one of the largest in the world. The semiconductor industry knows that they monopolize mainstream chips. They have overcapacity; they have too much capacity.

So the idea that China won’t be able to have AI chips is complete nonsense. Now, of course, if you ask me, would the United States be further ahead if the entire world had no compute at all? But that’s just not an outcome. That’s not a scenario that’s true. They have plenty of compute already.

The threshold they need for the concern you’re worried about—they’ve already reached that threshold and beyond. So I think you misunderstand that AI is a 5-layer cake, and at the lowest layer is energy.

When you have an abundance of energy, it makes up for chips. If you have an abundance of chips, it makes up for energy. For example, the United States is scarce on energy, which is the reason why NVIDIA has to keep advancing our architecture and do this extreme co-design so that, with the few chips we ship—because the amount of energy is so limited—our throughput per watt is off the charts.

But if your amount of watts is completely abundant, it’s free, what do you care about performance per watt for? You get plenty. You can use old chips.

So 7nm chips are essentially Hopper. I’ve got to tell you, today’s models are largely trained on the Hopper generation. So 7nm chips are plenty good. The abundance of energy is their advantage.

But then there’s a question of whether they can actually manufacture enough chips. But they do. What’s the evidence? Huawei just had the largest single year in the history of their company. How many chips did they ship? A ton. Millions. Millions is way more than Anthropic has.

Dwarkesh Patel

There’s a question of how much logic SMIC can ship, and there’s a question of how much memory.

Jensen Huang

I’m telling you what it is. They have plenty of logic, and they have plenty of HBM2 memory.

Dwarkesh Patel

Right. But as you know, the bottleneck often in training and doing inference on these models is the amount of bandwidth. So if you have HBM2—I don’t know the numbers offhand, but versus the newest thing you have, there could be almost an order-of-magnitude difference in memory bandwidth, which is huge.

Jensen Huang

Huawei is a networking company.

Dwarkesh Patel

But that doesn’t change the fact that you need EUV for the most advanced HBM.

Jensen Huang

Not true. Not at all true. You could gang them together, just like we gang them together with NVL72. They’ve already demonstrated silicon photonics, connecting all of this compute together into 1 giant supercomputer. Your premise is just wrong.

The fact of the matter is, their AI development is going just fine. The best AI researchers in the world, because they’re limited in compute, also come up with extremely smart algorithms. Remember, I just said that Moore’s law is advancing about 25% per year. However, through great computer science, we could still improve algorithm performance by 10x.

What I’m saying is that great computer science is where the lever is. There is no question, MoE is a great invention. There’s no question, all the incredible attention mechanisms reduce the amount of compute. We have got to acknowledge that most of the advances in AI came out of algorithm advances, not just the raw hardware.

Now, if most advances came from algorithms and computer science and programming, tell me that their army of AI researchers is not their fundamental advantage. We see it. DeepSeek is not an inconsequential advance.

The day that DeepSeek comes out on Huawei first, that is a horrible outcome for our nation. Why is that? Because currently you can have a model like DeepSeek that can run on any accelerator if it’s open source. Why would that stop being the case in the future?

Suppose it doesn’t. Suppose it’s optimized for Huawei. Suppose it’s optimized for their architecture. It would put ours at a disadvantage.

Dwarkesh Patel

You described a situation that I perceive to be good news. A company developed software, developed an AI model, and it runs best on the American tech stack. I saw that as good news. You set it up as a premise that it was bad news.

Jensen Huang

I’m going to give you the bad news: AI models around the world are developed and they run best on non-American hardware. That is bad news for us.

Dwarkesh Patel

I guess I just don’t see the evidence that there are these huge disparities that would prevent you from switching accelerators. American labs are running their models across all the clouds, across all the different accelerators.

Jensen Huang

I am the evidence. You take a model that’s optimized for NVIDIA and you try to run it on something else.

Dwarkesh Patel

But American labs do that.

Jensen Huang

And they don’t run better. NVIDIA’s success is perfect evidence. The fact that AI models are created on our stack and run best on our stack—how is that illogical to understand?

Anthropic’s models are run on GPUs, they’re run on Trainium, they’re run on TPUs. A lot of work has to go into it to change.

Dwarkesh Patel

But go to the Global South. Go to the Middle East. Coming out of the box, if all of the AI models run best on somebody else’s tech stack, you’ve got to be arguing some ridiculous claim right now that that’s a good thing for the United States. But I guess I just don’t understand the argument.

Say Chinese companies get to the next Mythos first. They find all the security vulnerabilities in American software first, but they can do it on NVIDIA hardware and they ship it to the Global South.

Jensen Huang

They do it on NVIDIA hardware. How is that good?

Dwarkesh Patel

Okay, it runs on NVIDIA hardware—

Jensen Huang

It’s not good. It’s not good.

Dwarkesh Patel

Right. It’s not good.

Jensen Huang

So let’s not let it happen.

Dwarkesh Patel

Why do you think it’s perfectly fungible, that if you didn’t ship them compute it would exactly be replaced by Huawei? They are behind, right? They have worse chips than you.

Jensen Huang

There’s evidence right now. Their chip industry is gigantic.

Dwarkesh Patel

You can just look at the FLOPs, bandwidth, or memory comparisons between the H200 and the Huawei 910C. It’s like half to a third.

Jensen Huang

They use more of it. They use twice as many.

Dwarkesh Patel

It seems like your argument is they have all this energy that’s ready to go, right? And they need to fill it with chips. They’re good at manufacturing. I’m sure eventually they would be able to just out-manufacture everybody. But there are these few critical years.

What is the critical year you’re talking about?

Jensen Huang

These next few years.

Dwarkesh Patel

We've got these models that are going to be able to do all the cyberattacks. In that case, if the next years are critical, then we have to make sure that all of the world's AI models are built on the American tech stack in these critical years. If they're built on the American tech stack, how would that prevent them, if they have more advanced capabilities, from launching Mythos-equivalent cyberattacks?

Jensen Huang

There's no guarantee either way. But if you have it early, we can prepare for it.

Listen, why are you causing one layer of the AI industry to lose an entire market so that you could benefit another layer of the AI industry? There are 5 layers, and every single layer has to succeed. The layer that has to succeed most is actually the AI applications. Why are you so fixated on that AI model? That one company? For what reason?

Because those models make possible these incredibly offensive capabilities, and you need compute to run them. The energy, the chips, and the ecosystem of AI researchers make it possible.

Dwarkesh Patel

Okay, stepping back, it has to be the case that China is able to build enough 7nm capacity. And remember, they're still stuck on 7nm, while you'll move on to 3nm and then 2nm or 1.6nm with Feynman. So while you're on 1.6nm, they're still going to be on 7nm, and they have to produce enough of it to make up for the shortfall.

They have so much energy that the more chips you give them, the more compute they'd have. So it comes out as a question of, ultimately, they are getting more compute. Compute is an input to training and inference.

Jensen Huang

Listen, I just think you speak in absolutes. I think the United States ought to be ahead. The amount of compute in the United States is 100 times more than anywhere else in the world. The United States ought to be ahead.

Okay. The United States is ahead. NVIDIA builds the most advanced technologies. We make sure that the U.S. labs are the first to hear about it and have the first chance to buy it. And if they don't have enough money, we even invest in them. The United States ought to be ahead.

We want to do everything we can to make sure the United States is ahead. Number one point, do you agree? We're doing everything we can to do that.

Dwarkesh Patel

But how is shipping chips to China keeping the U.S. ahead if they're bottlenecked on compute?

Jensen Huang

No, no. We've got Vera Rubin for the United States. We have Vera Rubin for the United States. Now, am I in the United States? Do you consider me part of the United States?

Dwarkesh Patel

Yes.

Jensen Huang

NVIDIA. You consider NVIDIA a United States company? Okay. Number one, why is it that we don't come up with a regulation that's more balanced so that NVIDIA can win around the world instead of giving up the world? Why would you want the United States to give up the world?

The chip industry is part of the American ecosystem. It's part of American technology leadership. It's part of the AI ecosystem. It's part of AI leadership. Why is it that your policy, your philosophy, leads to the United States giving up a vast part of the world's market?

Dwarkesh Patel

I guess the claim here is—Dario Amodei had this quote where he said that it's like Boeing bragging that we're selling North Korea nukes, but the missile casings are made by Boeing. And that's somehow enabling the U.S. technology stack. Fundamentally, you're giving them this capability.

Jensen Huang

Comparing AI to anything that you just mentioned is lunacy.

Dwarkesh Patel

But AI is similar to enriched uranium, right? It can have positive uses; it can have negative uses. We still don't want to send enriched uranium to other countries. The analogy is that enriched uranium is like compute.

Jensen Huang

It's a lousy analogy. It's an illogical analogy. But if that compute can run a model that can do zero-day exploits against all American software, how is that not a weapon?

First of all, the way to solve that problem is to have dialogues with the researchers, and dialogues with China, and dialogues with all the countries to make sure that people don't use technology in that way. That's a dialogue that has to happen. Okay? Number one.

Number two, we also need to make sure that the United States is ahead, that Vera Rubin and Blackwell are available in the United States in abundance—mountains of it. Obviously, our results would show it. Abundance, tons of it. The amount of computing we have is great. We have amazing AI researchers here. It's great. We ought to stay ahead.

However, we also have to recognize that AI is not just a model. AI is a 5-layer cake. The AI industry matters across every single layer, and we want the United States to win at every single layer, including the chip layer.

Conceding the entire market is not going to allow the United States to win the technology race long-term in the chip layer, in the computing stack. That is just a fact.

Dwarkesh Patel

I guess then the crux comes down to: how does selling them chips now help us win in the long term?

Tesla sold extremely good electric vehicles to China for a long time. iPhones are sold in China—extremely good. They didn't cause them to lock in. China will still make their version of EVs, and they're dominating. Their smartphones are dominating.

When we started the conversation today, you acknowledged that NVIDIA's position is very different. You used words like “moat.”

Jensen Huang

The single most important thing to our company is the richness of our ecosystem, which is about developers. 50% of the AI developers are in China. The United States should not give that up.

Dwarkesh Patel

But we have a lot of NVIDIA developers in the U.S., and that doesn't prevent American labs from also being able to use other accelerators in the future. In fact, right now they're using other accelerators as well, which is fine and great. I don't see why that wouldn't be the case in China as well if you sell them NVIDIA chips, just the same way that Google can use TPUs and NVIDIA—

Jensen Huang

We have to keep innovating and, as you probably know, our share is growing, not decreasing. The premise that even if we competed in China, we're going to lose that market anyway—you’re not talking to somebody who woke up a loser.

That loser attitude, that loser premise, makes no sense to me. We're not a car. We are not a car. The fact that I can buy this car brand one day and use another car brand another day—easy. Computing is not like that.

There's a reason why the x86 deal exists. There's a reason why ARM is so sticky. These ecosystems are hard to replace. It costs an enormous amount of time and energy, and most people don't want to do it.

So it's our job to continue to nurture that ecosystem, to keep advancing the technology so that we can compete in the marketplace. Conceding a marketplace based on the premise you described, I simply can't acknowledge that. It makes no sense.

Because I don't think the United States is a loser. Our industry is not a loser. That losing proposition, that losing mindset, makes no sense to me.

Dwarkesh Patel

Okay. I'll move on. I just want to make sure that—

Jensen Huang

You don't have to move on. I'm enjoying it.

Dwarkesh Patel

Okay, great. Then I won't. I appreciate that.

But I think maybe the crux—and thanks for walking around in circles with me, because I think it helps bring out what the crux here is—

Jensen Huang

The crux is you're going to extremes. Your argument starts from extremes: that if we give them any compute at all in this narrow moment, we will lose everything.

Dwarkesh Patel

No, I think what my argument is—

Jensen Huang

Those extremes, they're childish. Let me just make my argument for myself. The idea is not that there is some key threshold of compute. It's that any marginal compute is helpful.

Dwarkesh Patel

So if you have more compute, you can train a better model. And I just want you to acknowledge that any marginal sales for the American technology industry are beneficial.

Jensen Huang

I actually don’t. If the AI models that run on those chips are capable of cyber-offensive capabilities, or the chips are training models with cyber capabilities and running more instances of those models, it is not a nuclear weapon, but it enables a weapon of a kind.

Dwarkesh Patel

The logic that you use, you might as well say it to microprocessors and DRAMs. You might as well say it to electricity. But in fact, we do have export controls on the technology that is relevant to making the most advanced DRAM. We have all kinds of export controls on China for all kinds of chip-making stuff.

We sell a lot of DRAM and CPUs into China, and I think it’s right. I guess this goes back to the fundamental question of, is AI different? If you have the kind of technology where they can find these zero-days in software, is that something where we want to minimize China’s ability to get there first, to deploy it widely? We want the United States to be ahead. We can control that. How do we control that if the chips are already there and they’re using them to train that model?

Jensen Huang

We have tons of compute. We have tons of AI researchers. We’re racing as fast as we can.

Dwarkesh Patel

Again, we have more nuclear weapons than anybody else, but we don’t want to send enriched uranium anywhere.

Jensen Huang

We’re not enriched uranium. It’s a chip, and it’s a chip that they can make themselves.

Dwarkesh Patel

But there’s a reason they’re buying it from you.

Jensen Huang

We have quotes from the founders of Chinese companies that say that they’re bottlenecked on compute. Because our chips are better. On balance, our chips are better. There’s just no question about it.

Dwarkesh Patel

Can you acknowledge that Huawei had a record year? Can you acknowledge that a whole bunch of chip companies have gone public? Can you acknowledge that?

Jensen Huang

Yes.

Dwarkesh Patel

Can you also acknowledge that we used to have a very large share in that market, and we no longer have a large share in that market?

Jensen Huang

We can also acknowledge that China is about 40% of the world’s technology industry. To concede that market for the United States technology industry is a disservice to our country. It is a disservice to our national security. It is a disservice to our technology leadership, all for the benefit of one company. It makes no sense to me.

Dwarkesh Patel

I guess I’m confused. It feels like you’re making 2 different statements. One is that we’re going to win this competition with Huawei because our chips are going to be way better if we’re allowed to compete. Another is that they would be doing the same exact thing without us anyway. How can both of those things be true at the same time?

Jensen Huang

It’s obviously true. In the absence of a better choice, you’ll take the only choice you have. How is that illogical? It’s so logical. The reason they want NVIDIA chips is that they’re better.

Dwarkesh Patel

Yeah. Better is more compute. More compute means you can train a better model.

Jensen Huang

No, it’s just better. It’s better because it’s easier to program. We have a better ecosystem. But whatever the better is, whatever the better is, of course we’re going to send them compute. So what? The fact of the matter is that we get to benefit.

Don’t forget, we get the benefit of American technology leadership. We get the benefit of developers working on the American tech stack. We get the benefit, as those AI models diffuse out into the rest of the world, that the American tech stack is therefore the best for it. We can continue to advance and diffuse American technology. That, I believe, is a positive. It’s a very important part of American technology leadership.

Now, the policies that you’re advocating resulted in the American telecommunications industry being policied out of basically the world, to the point where we don’t control our own telecommunications anymore. I don’t see that as smart. It’s a little narrow-minded, and it led to unintended consequences that I’m describing to you right now that you seem to have a very hard time understanding.

Dwarkesh Patel

Okay, let’s just step back. It seems like the crux here is there’s a potential benefit and there’s a potential cost. What we’re trying to figure out is, is the benefit worth the cost? I guess I’m trying to get you to acknowledge the potential cost.

Jensen Huang

Compute is an input to training powerful models. Powerful models do have powerful offensive capabilities, like cyber attacks. It is a good thing that American companies got to Mythos-level capabilities first, and then now they’re going to hold off on those capabilities so that the American companies and American government can make their software more protected before that level of capability was announced.

If China had had more compute or more cloud compute, if they could have made a Mythos-level model earlier and deployed it widely, that would have been very bad. One of the reasons that hasn’t happened is that we have more compute thanks to companies like NVIDIA in America. That is a cost of sending it to China. So let’s leave the benefit aside for a second. Do you acknowledge that this is a potential cost?

Dwarkesh Patel

I’ll also tell you the potential cost is we allow one of the most important layers of the AI stack, the chip layer, to concede an entire market—the second-largest market in the world—so that they could develop scale, so that they could develop their own ecosystem, so that future AI models are optimized in a very different way than the American tech stack.

As AI diffuses out into the rest of the world, their standards, their tech stack, will become superior to ours, because their models are open. I guess I just believe enough in NVIDIA’s kernel engineers and CUDA engineers to think that they could optimize—

Jensen Huang

AI is more than kernel optimization, as you know.

Dwarkesh Patel

Of course, but there are so many things you can do, from distilling to a model that’s well-fit for your chips.

Jensen Huang

We’re going to do our best.

Dwarkesh Patel

You have all the software. It’s just hard to imagine that there’s a long-term lock-in to the Chinese ecosystem, even if they have a slightly better open-source model for a while.

Jensen Huang

China is the largest contributor to open-source software in the world. Fact. China’s the largest contributor to open models in the world. Fact. Today, it’s built on the American tech stack—NVIDIA’s. Fact.

All 5 layers of the tech stack for AI are important. The United States ought to go win all 5 of them. They’re all important. The one that is the most important, of course, is the AI application layer. The layer that diffuses into society—the one that uses it most will benefit from this industrial revolution most.

But my point is that every layer has to succeed. If we scare this country into thinking that AI is somehow a nuclear bomb, so that everybody hates AI and everybody’s afraid of AI, I don’t know how you’re helping the United States. You’re doing it a disservice.

If we scare everybody out of doing software engineering jobs because it’s going to kill every software engineering job—and we don’t have any software engineers as a result of that—we’re doing a disservice to the United States. If we scare everybody out of radiology so nobody wants to be a radiologist because computer vision is completely free and no AI is going to do a worse job than a radiologist, we misunderstand the difference between a job and a task.

The job of a radiologist is patient care. The task is to read a scan. If we misunderstand that so profoundly and we scare everybody out of going to radiology school, we’re not going to have enough radiologists and good enough healthcare.

So I’m making the case that when you make a premise that is so extreme, everything goes from zero or infinity, we end up scaring people in a way that’s just not true. Life is not like that. Do we want the United States to be first? Of course we do. Do we need to be a leader in every layer of that stack? Of course we do. Of course we do.

Today, you’re talking about Mythos because Mythos is important.

Dwarkesh Patel

Sure. That’s fantastic.

Jensen Huang

But in a few years’ time, I’m making you the prediction that when we want the American tech stack, when we want American technology to be diffused around the world—out to India, out to the Middle East, out to Africa, out to Southeast Asia—when our country would like to export, because we would like to export our technology, we would like to export our standards, on that day, I want you and I to have that same conversation again.

I will tell you exactly about today’s conversation, about how your policy and what you imagined literally caused the United States to concede the second-largest market in the world for no good reason at all.

We shouldn't concede it. If we lose it, we lose it. But why do we concede it? Now, nobody is advocating all or nothing, meaning we ship everything to China at all times. Nobody's advocating that.

We should always have the best technology here. We should always have the most technology here, and the first. But we should also try to compete and win around the world. Both of those things can simultaneously happen. It requires some amount of nuance, some amount of maturity instead of absolutes. The world is just not absolutes.

Dwarkesh Patel

Okay. The argument hinges on this: they've built models that are specified for the best chips that they make in a few years. Those chips get exported around the world. That sets the standard. Because of EUV export controls, as we said, you're going to move on to 1.6nm. They're still going to be on 7nm, even a few years from now.

It may make sense that domestically they would prefer, "Hey, we've got so much energy, we can manufacture at scale. We'll still keep using 7nm." But on the exporting thing, their 7nm chips have to be competitive against your 1.6nm chips. Their models have to be so far optimized for the 7nm that it's better to run their models on 7nm than to run their models on your 1.6nm.

Can we just look at the facts, then? Is Blackwell 50 times more advanced in lithography than Hopper? Is it 50 times?

Jensen Huang

Not even close. I just kept saying it over and over again. Moore's Law is dead. Between Hopper and Blackwell, from the transistors themselves, call it 75%. They were 3 years apart—75%. Blackwell is 50 times Hopper. My point is, architecture matters.

Computer science matters. Semiconductor physics matters as well, but computer science matters. The impact of AI largely comes from the computing stack, which is the reason why CUDA is so effective, which is the reason why CUDA is so beloved. It's an ecosystem, a computing architecture that allows for so much flexibility that if you wanted to change an architecture completely—create something like MoE, create something like diffusion, create something that's disaggregated—you could do so. It's easy to do.

So the fact of the matter is, AI is about the stack above as much as it is about the architecture below. To the extent that we have architectures and software stacks that are optimized for our stack, for our ecosystem, it is obviously good, because we started the conversation today about how NVIDIA's ecosystem is so rich.

Dwarkesh Patel

Why do people always love programming CUDA first?

Jensen Huang

They do. They do. So do the researchers in China.

But if we are forced to leave China, first of all, it's a policy mistake. Obviously, it has backlash. It has turned out badly for the United States. It enabled, it accelerated their chip industry. It forced all of their AI ecosystem to focus on their internal architectures. It's not too late, but nonetheless, it has already happened.

You're going to see in the future, they're not stuck at 7nm, obviously. They're good at manufacturing. They will continue to advance from 7nm and beyond. Now, is there a 10x difference between 5nm and 7nm? The answer is no. Architecture matters. Networking matters. That's why NVIDIA bought Mellanox. Networking matters. Energy matters. So all of that stuff matters. It's not simplistic, like the way you're trying to distill it.

4. Why doesn’t Nvidia make multiple different chip architectures?

Dwarkesh Patel

We can move on from China, but that actually raises an interesting question. We were discussing earlier these bottlenecks at TSMC and in memory and so forth. So if we're in this world where you're already the majority of N3—and at some point you'll be N2 and you'll be a majority of that—do you see that you could go back to N7, the spare capacity at an older process node, and say, "Hey, the demand for AI is so great and our capacity to expand the leading edge is not meeting it, so we're going to make a Hopper or Ampere, but with everything we know about numerics today and all the other improvements you described"?

Do you see that world happening before 2030?

Jensen Huang

It's not necessary to. The reason for that is because with every generation, the architecture is more than just the transistor scale. You're doing so much engineering and packaging and stacking, and the numerics and the system architecture.

When you run out of capacity, to go back easily to another node—that's a level of R&D that no one could afford. We could afford to lean forward. I don't think we could afford to go back.

Now, if the world simply says—if, on that day, we go, "Listen, we're just never going to have more capacity ever again"—would I go back and use 7nm? In a heartbeat. Of course I would.

Dwarkesh Patel

One question somebody I was talking to had was, why doesn't NVIDIA run multiple different chip projects at the same time with totally different architectures? So you could do something like a Cerebras-style wafer-scale. You could do a Dojo-style huge package. You could do one without CUDA. You have the resources and the engineering talent to do all of these in parallel.

So why put all the eggs in one basket, given who knows where AI might go and where architectures might go?

Jensen Huang

Oh, we could. It's just that we don't have a better idea. We could do all of those things. It's just not better. We simulate it all in our simulator—provably worse—so we wouldn't do it. We're working on exactly the projects that we want to work on.

If the workload were to change dramatically—and I don't mean the algorithms, I actually mean the workload, and that depends on the shape of the market—we may decide to add other accelerators. For example, recently we added Groq, and we're going to fold Groq into our CUDA ecosystem. We're doing that now because the value of tokens has gone up so high that you could have different pricing of tokens.

Back in the old days, just a couple of years ago, tokens were either free or barely expensive. But now you can have different customers, and those customers want different answers. Because the customers make so much money—for example, our software engineers—if I can give them much more responsive tokens so that they're even more productive than they are today, I would pay for it. But that market has only recently emerged.

So I think we now have the ability to have the same model, based on the response time, have different segments. That's the reason why we decided to expand the Pareto frontier and create a segment of inference that has a faster response time, even though it's lower throughput. Until now, higher throughput is always better.

We think there could be a world where there could be very high-ASP tokens, and even though the throughput is lower in the factory, the ASPs make up for it. That's the reason why we did it.

But otherwise, from an architecture perspective, if I had more money, I would put more behind NVIDIA's architecture.

Dwarkesh Patel

I think this idea of extremely premium tokens and just the disaggregation of the inference market is very interesting. The segmentation of it.

Jensen Huang

Yeah.

Dwarkesh Patel

All right, final question. Suppose the deep learning revolution didn't happen. What would NVIDIA be doing? Obviously, games.

Jensen Huang

Accelerated computing, the same thing we've been doing all along. The premise of our company is that general-purpose computing is good for a lot of things, but for a lot of computation, it's not ideal. So we combined an architecture called a GPU, CUDA, to a CPU, so that we can accelerate the workload of the CPU.

Different kernels of code or algorithms could be offloaded onto our GPU. As a result, you speed up an application by 100x, 200x. Where can you use that? Obviously, engineering and science and physics, data processing, computer graphics, image generation, all kinds of things.

Even if AI doesn't exist today, NVIDIA would be very, very large. The reason for that is fairly fundamental, which is that the ability for general-purpose computing to continue to scale has largely run its course. And the way to do that is through domain-specific acceleration.

One of the domains that we started with was computer graphics, but there are many other domains. There are all kinds: particle physics and fluids, structured data processing, all kinds of different types of algorithms that benefit from CUDA.

Our mission was really to bring accelerated computing to the world and advance the type of applications that general-purpose computing can't do, and scale to the level of capability that helps break through certain fields of science. Some of the early applications were molecular dynamics, seismic processing for energy exploration, image processing, of course—all of those kinds of fields where general-purpose computing is simply too inefficient to do so.

If there were no AI, I would be very sad. But because of the advances that we made in computing, we democratized deep learning. We made it possible for any researcher, any scientist, anywhere, and any student to be able to access a PC or a GeForce add-in card and do amazing science.

That fundamental promise hasn’t changed, not even a little bit. If you watch GTC, there’s the whole beginning part of it. None of it is AI. That whole part of it with computational lithography, our quantum chemistry work, data processing work—all of that stuff is unrelated to AI, and it’s still very important.

I know that AI is very interesting and quite exciting, but there are a lot of people doing a lot of very important work that’s not AI-related, and tensors are not the only way that you compute it. We want to help everybody.

Dwarkesh Patel

Jensen, thank you so much.

Jensen Huang

You’re welcome. I enjoyed it.

Dwarkesh Patel

Me too.