[BidClub_]
The a16z Show · · 99 分钟

Dylan Patel 谈 AI 芯片竞赛——NVIDIA、Intel 与美国政府对决中国

Erik TorenbergDylan PatelSarah WangGuido Appenzeller

YouTube
TL;DR
  • Nvidia–Intel 的联手是一场战略背书,既可能降低 Intel 的资本成本,也可能重绘 PC 与数据中心版图。 在 SoftBank 投入20亿美元、美国政府投入100亿美元后,Nvidia 又承诺投入50亿美元,但 Dylan Patel 仍认为 Intel 需要约500亿美元;Jensen Huang 先带来半导体行业的“巴菲特效应”,再为更大规模融资铺路。Guido Appenzeller 认为 x86 与 Nvidia 的整合产品极具吸引力,并警告称,当“两大宿敌突然联手”时,AMD 和 Arm 面对的是最坏的消息。

  • Huawei 在技术上具备可信度,但制造规模——尤其是高带宽内存——仍是决定成败的承重墙。 Huawei 在2020年将7 nm Ascend AI 芯片推向市场,后来通过中间商获得约290万颗 TSMC 制造的芯片,如今又提出将 prefill 与 decode 分开的产品路线,并配套定制 HBM。中国或许能生产大量7 nm 逻辑芯片,甚至利用现有设备做到5 nm,但 Patel 强调,设计不等于生产:HBM3 良率、刻蚀产能,以及从库存储备转向国产规模化生产,仍未解决。

  • 中国拒绝 Nvidia 芯片,同时是产业政策、一次高风险产能押注,也可能成为谈判筹码。 ByteDance 和其他模型厂商仍偏好 Nvidia,因为它“好得多”,但北京可以强制采用国产芯片,同时让 Huawei 发布雄心勃勃的路线图;Patel 称这一招是“1万 IQ”,华盛顿“下跳棋,而他们在下国际象棋”。如果国产供应无法足够快地爬坡,中国可能暂时回撤,届时必须在主权与以接近美国的规模部署“超级强大的 AI”之间做选择。

  • Nvidia 的短期看多逻辑,建立在 AI 资本开支显著高于华尔街模型之上,而不是进一步提升市场份额。 银行共识预计6家超大规模云厂商明年资本开支约3600亿美元;Patel 对数据中心和供应链的研究则指向4500亿至5000亿美元,其中 Nvidia 大体随市场增长、同时守住份额。仅 OpenAI 就与 Oracle 签下超过3000亿美元的合同,行业看多逻辑最终变成“每年在 AI 基础设施上投入数万亿美元”——但 Patel 拒绝预测5年以后,因为“第5年基本就是 YOLO”。

  • Nvidia 的护城河,在于它反复愿意押注库存、临近量产时重做设计,并在竞争对手完成修改前先交付可用芯片。 据称 Jensen 在 Microsoft 正式授予 Xbox 业务前就下了量产订单,押注不可取消的产能时甚至超过客户自身计划,并在投片前仅数月为 Volta 加入 tensor cores。Nvidia 往往在 A0 阶段就出货,而 Intel 有一款数据中心处理器走到了 E2——约第15次修改;这体现了“我讨厌电子表格,我就是知道”和按季度谨慎决策之间的文化差异。

  • Amazon 和 Oracle 即使没有最好的加速器,也能获得 AI 云份额,因为带电产能与愿意动用资产负债表的交易对手都很稀缺。 Patel 预计,随着 Anthropic、Trainium 和 GPU 部署填满 Amazon 的闲置产能,AWS 收入增速将触底并重新超过20%;Trainium 仍然“非常难用”,但服务少数高流量模型的实验室可以对其进行手工优化。Oracle 的硬件中立工程能力,以及为 OpenAI 需求提供承保的意愿,使其成为更大胆的交易对手,不过 OpenAI 在2028–29年能否每年支付超过800亿美元,仍是核心风险。

  • GB200 的经济性取决于工作负载,而基础设施能力不足的团队可能因可靠性问题损失纸面性能。 相比 H100 约1.6倍的总拥有成本,GB200 在部分训练场景中可能只有约2倍性能,但在 DeepSeek 推理中,每颗 GPU 的性能可能超过6–7倍;问题在于,如今一次故障会发生在包含72颗 GPU 的 NVLink 域内。因此,一些云厂商对64颗 GPU 承诺约99%的可用性,对完整72颗 GPU 却只有95%;“故障的爆炸半径”与基准测试速度同样重要。

  • GPU 市场再次趋紧,而 Nvidia 下一个战略难题变成:每年潜在的2500亿美元自由现金流该怎么用。 Patel 在同一句话中还使用了一个含义不明的“2亿美元”数字。随着推理模型推高需求,几家大型 neocloud 的 Hopper 产能已经售罄;Blackwell 的部署耗时超过 Hopper,价格数月前已触底并开始缓慢上行——少量分配仍容易获得,但大规模即时集群并不容易。Patel 偏好的资金出口是数据中心和电力,即 GPU 增长的瓶颈,但即使如此,也未必能消化这些现金而不把 Nvidia 变成一家文化完全不同的公司。

摘要 · 为研究而整理的核心内容

1. Nvidia 对 Intel 的投资,将竞争转化为战略依赖

  • Patel 开场的反应带有金融戏剧色彩:Nvidia 宣布投资50亿美元于 Intel,Intel 股价上涨约30%,按他的说法,这笔持股已经产生了“10亿美元利润”。更重要的是,一位潜在的大客户正在投入资本并承诺产品路线图,让 Intel 在重返公开债务或股权市场前就获得了验证。

  • 产品逻辑异常强。Intel 会把自己的一个 chiplet 与 Nvidia 的一个 chiplet 并排封装,改写 Intel 因芯片组业务遭遇反垄断诉讼、并向 Nvidia 支付和解金的历史。Patel 称这一转折“很有诗意”:Intel“某种程度上正爬向 Nvidia”,但一台完全整合 Nvidia 图形能力的 x86 笔记本,可能是市场上最好的设备。

  • 已宣布的支票相对于 Intel 的需求仍然很小:Nvidia 投入50亿美元,SoftBank 投入20亿美元,美国政府投入100亿美元,而 Patel 此前估计 Intel 需要立即获得约500亿美元。更多战略投资者——他提出 Apple 作为一种可能——可能在 Intel 扩充资产负债表之前,先制造“沃伦·巴菲特买入一只股票”般的效应。

  • Appenzeller 从客户角度给出的评价既兴奋又残酷。Intel 可以重置其缺乏竞争力的内部图形和 AI 项目,包括讨论中提到的“Gaudi F4”;AMD 如今要面对两大历史宿敌联手,Arm 则失去了“任何想避开 Intel 的客户都会自然选择它作为伙伴”的卖点。“牌局被重新洗牌了。”

2. Huawei 在制裁来临时已是一家接近前沿的芯片公司

  • Patel 要听众从2020年开始回顾,而不是从当前路线图开始。Huawei 将 Ascend 加速器提交给独立的公开基准测试,成为第一家将7 nm AI 芯片推向市场的公司;在禁令切断完整的海外供应链之前,它与 Nvidia 的技术差距并不大。Huawei 的 TSMC 订单量还曾超过 Apple。

  • Trump 政府的限制,恰好在 Huawei 已具备挑战市场的要素时切断了这条供应链。Nvidia 加速扩张,Huawei 则围绕 SMIC 重建供应链、寻求韩国内存,同时通过空壳公司下单购买 TSMC 芯片。Huawei 已经用有限的原始 Ascend 库存训练了大量模型。

  • Patel 称,到2024年末,这条中间商渠道已在约5亿美元订单的基础上,产出约290万颗 TSMC 制造的芯片,随后被当局叫停。他提到有报道称 TSMC 可能面临10亿美元的美国罚款,但明确表示不确定罚款是否已经开出;关键是,公开可见的 Ascend 部署当时还没有消耗完这批库存。

  • 2025年的 H20 禁令又移除了 Nvidia 在中国超过200亿美元的收入,SemiAnalysis 曾对此作出估算。Nvidia 注销库存,后来获准转售这些芯片,也面临更艰难的选择:重启一条可能再次被切断的供应链,还是任由 Huawei 和 Cambricon 吞下一个当前仍使用海外晶圆和内存的“国产”市场。

3. 中国设计前沿架构的速度,快于制造它的速度

  • Patel 对设备规则的解读,比其公开标签更宽松。尽管华盛顿称限制范围延伸至14 nm,他认为真正被禁止的设备主要用于7 nm 以下工艺。因此,中国应能生产大量7 nm AI 芯片,也可能利用现有设备把工艺拉到5 nm,只是经济性会更差。

  • Huawei 的路线图将推理拆成专用产品:一颗芯片用于推荐系统和 prefill,另一颗用于 decode。Nvidia 和多家初创公司也在进行同样的架构拆分,因此 Huawei 真正令人意外的并非拆分本身,而是为 decode 定制 HBM——Nvidia 和 AMD 当时也只是在准备于次年采用类似方案。

  • 进口轨迹仍然说明瓶颈所在。此前中国将30–40%的设备进口用于囤积光刻设备,历史上晶圆厂设备组合中光刻设备约占17–18%,EUV 时代约占25%;如今刻蚀设备进口正在激增。HBM 需要先在每一层上刻蚀硅通孔,再将12层或16层芯片堆叠起来。

  • Huawei 曾经采样 HBM2,但据 Patel 所述,到当时还没有开始量产数年前就已推出的 HBM3。设备可得性和良率爬坡都很重要:中国追赶的速度可以快于发明这项技术所需的时间,因为工艺已经存在,但几个月的进口无法复制韩国多年积累的产能。Torenberg 将最终结果概括为“只是时间问题,不是会不会的问题”;Patel 自己强调的则是,制造产能和良率仍是瓶颈。

4. 北京禁用 Nvidia,制造出危险的过渡缺口

  • 中国可以先拒绝 Nvidia,因为它正在把2024年的库存转化为加速器。真正困难的阶段,会出现在这些芯片耗尽,而国产逻辑芯片和 HBM 尚未达到量产时。Patel 预计中国最终会爬坡,但“会再花一点时间”,这可能造成一段北京回撤并放松政策的窗口期。

  • 商业激励与产业政策背道而驰。ByteDance 被形容为“苦苦哀求 Nvidia 芯片”:它使用部分 Huawei 和 Cambricon 硬件,但希望用 Nvidia 训练最好的模型,并高效部署推理。政府可以强制采购国产芯片,但这并不意味着 Nvidia 已经失去竞争力。

  • 走私和转出口仍以低到中低规模持续,但无法填补全国部署缺口。中国最终必须决定,在多大程度上优先建设内部供应链、又要如何保持强大 AI 的发展速度;否则,它部署的“AI 芯片数量会少得多”,落后于美国。

5. Huawei 越是想要 Nvidia,越可能夸大自身实力

  • Patel 的谈判解读极具对抗性:如果中国想要更好的美国芯片,就应该宣传国产自给自足,公布尽可能“疯狂”的多年期 Huawei 路线图,并宣布禁用 Nvidia。美国供应商随后会警告官员,一个不可替代的市场正在消失,并游说放宽出口限制。

  • 他将这套打法称为“1万 IQ”——“我们在下跳棋,而他们在下国际象棋”——同时承认 Huawei 的许多架构是真实存在的。夸大之处主要在于对制造确定性的描述:路线图把希望中的产能和良率,当成了已经存在的现实。

  • 因此,出口政策不能简化为全部出售或全部禁售。Patel 建议比较中国在各性能层级上能够实现的产量,再决定哪些接近或略高于该水平的美国产品可以出售。AI 的潜在终端市场远大于半导体设备,因此只追求当前芯片收入是不够的。

6. Jensen 真正担心的中国以外竞争者,是 Huawei

  • Patel 认真对待 Jensen 对 Huawei“强大”的描述。在制裁之前,Huawei 的 TSMC 订单量和多个市场的手机份额都超过 Apple,随后又在没有西方供应链的情况下开始恢复。基于这段历史,对 Huawei 的担心超过 AMD,并不只是游说话术。

  • Jensen 最有力的论点,是让 Huawei 雄心勃勃的路线图变成地缘政治现实:说服政策制定者相信制造产能不是约束,相信 Huawei 将拿下中国、中东、东南亚和南亚、欧洲以及拉丁美洲。Patel 的反驳很窄,却很关键——产能和良率是真实约束,即便只是暂时的。

  • 另一种策略是“日本化”中国:将其隔离到硬件和软件都针对国内市场高度优化,像 Patel 提到的那些不同寻常的日本 PC 一样,把全球平台留给美国公司。但隔离也可能迫使中国走上一条更优路径,而西方的硬件—软件协同设计则固化在局部最优。“我不知道这是否准确,但这是一个有意思的说法。”

7. Nvidia 可量化的看多逻辑,是超大规模云厂商4500亿至5000亿美元资本开支

  • 银行共识预计 Microsoft、CoreWeave、Amazon、Google、Oracle 和 Meta 明年资本开支约3600亿美元。Patel 基于逐个数据中心、组件和供应链建立的模型,得出的数字约为4500亿至5000亿美元。Nvidia 已经占据主导地位,无法有意义地继续提升份额;盈利杠杆在于守住份额,同时让基础设施总规模以快于共识的速度扩张。

  • Oracle 的合同说明了需求规模。OpenAI 承诺在数年内支付超过3000亿美元,年支出升至超过800亿至900亿美元,尽管它目前没有现金完成支付。OpenAI 的 ARR 已达到约200亿美元;对下一年退出时 ARR 的估计介于350亿美元和450亿美元之间,而在实现盈利前,预计年度现金消耗约150亿至250亿美元,盈利目标定在2029年。

  • 将 OpenAI、Anthropic 和其他实验室的类似收入与融资叠加起来,5000亿美元的超大规模云资本开支就变得可信。Nvidia 的最大化论点是:AI 基础设施每年将达到“数万亿美元”,GPU 将成为商业代理、编程和消费级陪伴服务的中介。

  • 当被追问 Nvidia 的最终天花板时,Patel 拒绝给出虚假的精确度。如果 AI 反复改进 AI,价值创造可能达到数百万亿美元,但供应链只能看到未来3到4年:“第5年基本就是 YOLO。”白领生产率是否翻倍、被取代,或依赖持续不断的 token 流,超出了可投资的5年预测范围。

8. Nvidia 通过反复押注公司生死,构筑了护城河

  • Nvidia 早期多次失败,也反复作出攸关存亡的承诺。一位行业老兵告诉 Patel,Jensen 在 Microsoft 正式授予合同前就下令生产 Xbox——这个故事可能比传说本身更复杂——但订单确实先下了。Nvidia 第一颗成功芯片同样必须在公司唯一负担得起的掩模组上一次成功,否则公司就会耗尽现金。

  • 加密货币繁荣期间,Nvidia 说服供应商相信需求来自持久的游戏、可视化和数据中心增长,促使它们建设产能。加密货币崩盘后,Nvidia 吸收库存减记,供应商则留下空置产线。AMD 的挖矿芯片效率更高,却没有激进扩大产量;这是合理的风险政策,但也放弃了上行空间。

  • 这种行为延续到了超大规模云厂商。Nvidia 预订了高于 Microsoft 内部计划的不可取消、不可退回产能,随后 Microsoft 将自己的计划上调至接近 Nvidia 的数字。Patel 用 Jensen 对 CFO 说过的话概括这位创始人的决策体系:“我讨厌电子表格。我不看它们。我就是知道。”

  • Jensen 的弹球比喻解释了这种时间跨度:“你之所以赢,是为了能够再玩一局。”胜利为下一代产品提供资金,而不是服务于一份固定的15年计划。创始人的记忆在这里很重要——他记得 Nvidia 曾几乎破产,却仍然认为公司必须继续承担类似风险,而不是围绕可预测的华尔街季度优化。

9. 首版硅片与临门改设计,将愿景变成市场份额

  • Nvidia 长期稳定的工程组织里,既有有远见的架构师,也有愿意说“我们现在就得把这颗硅片做出来”的执行者。Patel 形容其中一位私下里的工程负责人几乎像传说人物,另一位同事则以砍掉技术人员喜爱的功能闻名,把这些功能留到下一颗芯片,而不是拖延出货。

  • 可量化的优势在于 stepping。Nvidia 经常交付 A0 硅片,偶尔交付 A1;Intel 有一款数据中心处理器走到了 E2,约为第15次修改。每次新的 stepping 可能耗费约1个季度,因此验证质量会从工程卫生问题变成上市武器。Patel 回忆说,Intel 曾公开羡慕 Nvidia“始终能在第一次修改时交付”。

  • Nvidia 还会先开始晶体管层生产,在必要时于最终金属布线前暂停,等验证完成后再进入量产。仿真、验证与有计划地保留在制品相结合,使 Nvidia 能在竞争对手完成掩模修改前作出响应。

  • Volta 是这类押注的典型:Nvidia 观察 P100 Pascal 上的 AI 工作负载后,仅在投片前数月加入 tensor cores。如果没有这次临时改动,AI 加速器市场可能会被别人拿走。同样重要的是,其软件组织足够快地交付驱动和基础设施,使首版硬件一出厂就能立即发挥作用。

10. Nvidia 的资产负债表正在变成战略问题

  • Patel 提到 Nvidia 每年约2500亿美元的自由现金流;同一句话中还出现了一个含义不明的“2亿美元”数字。监管机构不允许 Nvidia 收购 Arm,即使是50亿美元的 Intel 投资也需要审查,而没有明显的收购标的能够在不引发反垄断或整合问题的情况下吸收数千亿美元。

  • 对 CoreWeave、模型实验室和其他 neocloud 的小额投资有助于分散买家,但仍只是“小菜一碟”。Nvidia 可以为 Anthropic、xAI 或 OpenAI 的整轮融资提供资金,但挑选赢家会让所有未被选中的客户感到不安,并强化它们采用 AMD、TPU、初创公司或自研芯片的动机。

  • 更干净的用法,是扩张互补环节:数据中心和电力。Nvidia 可以为数据中心和能源基础设施提供资金,却不亲自运营云层,从而移除 GPU 增长的物理瓶颈,同时让日益强大的云竞争者负责出租算力。

  • 未解决的风险在于文化。浇筑混凝土和建设电力基础设施,需要的是与设计加速计算不同的人才;无休止的回购则会让人把 Nvidia 与 Tim Cook 治下的 Apple 相比——供应链执行出色,但在 Patel 看来,近10年几乎没有转型性投资。“没有什么东西需要3000亿美元资本”,却又能提供明显回报。

11. Amazon 的闲置电力可能扭转其 AI 云放缓

  • Patel 在2023年第一季度提出的“Amazon 的云危机”认为,AWS 针对的是上一个时代:弹性横向扩展网络、定制 CPU 和降本,而不是在绝对成本上升时仍追求每美元性能最大化的紧耦合 AI 系统。Neocloud 会把这一优势商品化,AWS 随后成为表现最弱的超大规模云厂商。

  • 他最新的判断是转向,而不是撤回此前的结构性批评。随着 Anthropic、Trainium 和 Nvidia 产能开始贡献收入,AWS 同比增速应在当前季度触底,并重新超过20%。在当下的短缺环境中,有带电空间可以填满,可能比拥有行业最干净的软件或加速器更重要。

  • Amazon 历来运营密度异常高的设施——当同行使用12 kW 机柜时,它使用约40 kW——其对效率的痴迷让数据大厅感觉“像一片沼泽”,又热又潮。改造网络和冷却不如专门建设基础设施优雅,但相对于 GPU 成本很低;Patel 提到,有些2 GW 场址已经配备电力、变压器、湿式冷却器和干式冷却器。

  • 这次改造创造了组件赢家。SemiAnalysis 曾基于 Amazon 的连接订单等因素,将 Astera Labs 的目标价设在约90美元;Patel 指出,该股次月已升至约250美元。他的意思并不是 Amazon 拥有更优架构,而是只要新增网络和冷却硬件能让 GPU 机柜出租出去,这些成本就无关紧要。

12. Anthropic 能容忍 Trainium,因为它服务的是狭窄工作负载

  • Trainium 仍然“非常难用”,一家主持人投资组合公司的说法是,2023年夏天它几乎不可能使用。Patel 没有声称开发体验已经变好。他更窄的论点是,即使在 Nvidia 上,生产级推理也需要手写 kernel、底层优化,有时还要做类似汇编的工作。

  • TPU 和 Trainium 使用更大、更简单的核心,通用功能更少;据称一些 Anthropic 工程师在深入到底层后反而更喜欢这种设计。普通客户仍然会遇到困难,但 Anthropic 只需要优化少数几个模型,不必支持全世界的架构。

  • Patel 有意给出的粗略例子,是主要服务 Sonnet——“Sonnet 3.5,或者等等,4.5,不管现在是哪一个”——同时把其他模型留在 GPU 或 TPU 上。如果 Anthropic 的 ARR 达到数百亿美元,而一个模型承载大部分流量,那么即使 Trainium 难用,为约150亿美元的 Trainium 产能投入大量优化成本也具备经济合理性。

13. Oracle 通过承保 Microsoft 不愿承保的需求赢得订单

  • Oracle 将庞大资产负债表与几乎不带硬件教条的态度结合起来。它可以部署 Arista Ethernet、白盒网络、Nvidia InfiniBand 或 Spectrum-X,并由强大的网络和软件团队提供支持。这种灵活性使它成为承接 OpenAI 极端需求的天然交易对手。

  • Microsoft 的独家权变成了优先购买权:OpenAI 可以提出每年800亿美元、或多年3000亿美元的需求,而 Microsoft 可以拒绝。Oracle 接受这场押注,是因为 OpenAI 没有资产负债表,而 Oracle 有;Microsoft 的谨慎可以辩护,但也因此留下了机会。

  • SemiAnalysis 通过许可证、卫星图像、电力设备、冷却器、变压器和发电机,绘制 Oracle 已签约和潜在场址。按照 GB200 时代的简化假设——一颗 GPU 约1000瓦,整套系统约2000瓦,每颗 GPU 的全包加速器资本开支约5万美元——它估算每兆瓦每年约1200万美元的租赁收入,再将每个场址每季度通电的进度纳入 Oracle 预测。

  • 这套方法与 Oracle 公布的2025–27年收入路径以及2028年的大部分表现高度吻合;未被观察到的2028–29年产能造成了偏差。OpenAI 能否每年支付超过800亿美元仍不确定,但 Oracle 只会在出租前1到2个季度购买 GPU。它更早作出的承诺主要是数据中心产能,因此搁浅资产风险有限,后续购买 GPU 则可以通过债务融资。

14. xAI 的优势,是把监管当成另一项工程约束

  • AI 基础设施已经从百分比增长进入数量级增长。一个超过100 MW、拥有10万颗 GPU 的集群曾经非同寻常;Patel 的团队如今追踪的这类集群约有10个,看到另一个200 MW场址时甚至只会发一个打哈欠的表情。“只有做到 GW 级别才值得兴奋。”

  • Elon Musk 在 Memphis 的第一个项目仍然突出:xAI 在2024年2月前后买下一座工厂,6个月内便用约10万颗 GPU 训练模型。它部署了大规模液冷、移动变电站、发电机和 CAT 涡轮机,还接入附近的天然气管道——在传统运营商还在寻找另一个场址时,xAI 已经把资源集中到位。

  • Colossus 2 在接近 GW 的规模上重复了这一做法。Memphis 的扩张受到政治和环境阻力后,xAI 又买下约10英里外的另一座设施,将其建在 Mississippi 州边界附近,并在 Mississippi 收购了一座电厂,因为那里的监管不同。Patel 的第一性原理总结是:别人说那里建不了电力设施,Musk 说,“那就跨过边界。”

15. GB200 奖励正确的工作负载,也惩罚薄弱的运营能力

  • SemiAnalysis 估算 GB200 的总拥有成本约为 H100 的1.6倍。如果工作负载只能获得2倍性能,升级值得做,但优势有限;对于 DeepSeek 推理,Patel 提到每颗 GPU 的性能超过6–7倍,并且仍在持续优化,这会把60%的成本溢价转化为约3–4倍的每美元性能。

  • B200 在运营上更简单:传统服务器内放置8颗 GPU,上行空间较小,但可靠性更熟悉。GB200 NVL72 建立了一个72颗 GPU 的统一域,推理收益大得多,但热量、新颖性和系统复杂度让它更难伺候。一颗 GPU 故障的“爆炸半径”远大于8 GPU 服务器中的一次故障。

  • 具备复杂基础设施能力的实验室,会将高优先级任务放在64颗 GPU 上,将低优先级任务放在8颗 GPU 上;一颗 GPU 故障时,可以借用健康 GPU,并推迟实体维修。云厂商无法随意把这些备用 GPU 租给其他客户,因为 NVLink 域必须保持一致。

  • SLA 如今体现了这种折中:Patel 的示意案例是,64颗 GPU 的正常运行时间约为99%,完整72颗 GPU 则约为95%,具体取决于供应商。大型实验室可以围绕这一点排程;小公司则可能因 GPU 闲置、停机,以及无法混合高优先级和低优先级工作负载,损失理论上的性能收益。

16. Rubin CPX 将 prefill 变成独立的芯片市场

  • 现代训练越来越多地由类似推理的强化学习生成组成,而推理本身又分为 prefill 和 decode。Prefill 处理提示词并构建 KV cache;decode 则自回归地生成 token。两者对硬件的压力差异足够大,领先实验室已经将它们放在不同的 GPU 池中。

  • Chunked prefill 可以填补未使用的批处理容量,但会拖慢并发 decode。Disaggregation 让运营商能够分别对长输入和长输出流量自动扩容,同时保证首 token 时间——这是用户最能感知的延迟。一位主持人说,“反正我也读不了那么快”,但用户仍会放弃那些在开始输出前迟疑的产品。

  • Decode 会反复移动模型参数和用户专属的 KV cache,因此内存带宽是核心。一个64,000 token 的 prefill 请求则包含足够多的计算,可以单独占满一颗加速器;相比快速加载参数,原始 FLOPS 更有价值。

  • Nvidia 的 Rubin CPX 专门针对计算密集型 prefill 阶段,并去掉昂贵的 HBM;Patel 称 HBM 占 GPU 成本的一半以上。如果 Nvidia 将相近的毛利率传导给客户,CPX 就能显著降低长上下文推理的成本,而不必迫使数据中心运营商重新设计周边设施。

17. 大型 GPU 买家重新陷入趋紧的现货市场

  • Patel 的 GPU 采购原则刻意令人难忘:“买 GPU 的方式,就像买可卡因。”买家会联系或私信多个供应商——“喂,你们有多少?什么价格?”——而不是进行一场干净利落的企业 RFP。他的团队与约30家 neocloud 保持 Slack 联系,并实时传递客户需求。

  • 几家大型 neocloud 的 Hopper 已经售罄,而 Blackwell 产能仍要等数月。推理模型和收入增长速度超过了供给;与此同时,Blackwell 的可靠性和部署学习曲线,使可用产能相对于 Hopper 延后,后者在1到2个月内就能安装并产生收入。

  • Hopper 价格大约在3到6个月前触底,随后开始缓慢上涨。Patel 没有断言市场回到了2023–24年的短缺状态:少量 GPU 容易获得,但能够立即交付的大型集群很难找到。产能再次先于正式价目表成为关键。

Dylan

How you buy GPUs is like buying cocaine. You call up a couple of people, text a couple of people, and ask, “Yo, how much do you have? What’s the price?”

Guido Appenzeller

If your 2 arch nemeses suddenly team up, that’s the worst possible news you can have. I did not see this coming. I think it’s an amazing development.

Dylan

Like Warren Buffett coming into a stock, Jensen is like the Buffett effect for the semiconductor world. It’s kind of poetic that everything’s gone full circle and Intel is sort of crawling to NVIDIA.

Erik Torenberg

Dylan, welcome back to the podcast.

Dylan

Thanks for having me. It just so happens that there’s some big news as we’re having you on: NVIDIA announced a $5 billion investment in Intel, and they’re teaming up to jointly develop custom data centers and PC products. What do you think about the collaboration?

I think it’s hilarious that NVIDIA can invest, have it get announced, and their investment is already up 30%. A $5 billion investment, a billion-dollar profit already, right? I think it’s fun because they need their customers to really have buy-in. So when their potential customers buy in and commit to certain types of products, it makes a lot of sense.

It’s kind of funny because, in the past, there was this whole thing around how Intel was sued for being anticompetitive with its chipsets, and NVIDIA actually got a settlement from Intel way back when, when the graphics were separate from the GPU and were really put on the chipset, which had all this other I/O, like USB and all this stuff.

So it’s kind of a funny turn of events that now Intel is going to make a chiplet and package it alongside a chiplet from NVIDIA, and then that’s a PC product, right? It’s kind of poetic that everything’s gone full circle and Intel is sort of crawling to NVIDIA. But actually, it might just be the best device, right?

I don’t want an ARM laptop because it can’t do a lot of things, so an x86 laptop with NVIDIA graphics fully integrated would probably be the best product on the market. Are you optimistic? How do you think this will go?

I mean, sure. I hope so. I’m a perpetual optimist on Intel because I have to be. I was thinking that the structure of the deal that a lot of the government folks and Intel were trying to pursue was that big customers and the biggest suppliers would directly give capital to Intel.

But this is sort of the other way around, where they’re buying some of the stock and having some ownership, but they’re not really diluting the other shareholders. The other shareholders will get diluted—everyone will get diluted—when Intel finally does raise capital from the capital markets.

But because they’ve announced these deals, and they’re pretty small—$5 billion from NVIDIA, $2 billion from SoftBank, and $10 billion from the US government—these are still relatively small.

Dylan

Pretty small.

Yeah, in the grand scheme of things. I mean, last time I think I said Intel needs something like $50 billion right now. When they go to the capital markets, it’s better, and hopefully they get another couple of these announcements.

There’s all sorts of speculation that Trump is involved in getting these companies to invest in Intel. Now the government is involved as well, of course. Is Apple going to come invest and also do something with Intel? Who else will come in?

That will really boost investor confidence, and they can dilute or go get debt, like Warren Buffett coming into a stock. Jensen is like the Buffett effect for the semiconductor world.

Guido, you were the CTO of Intel’s Data Center and AI Group. What are your thoughts?

Guido Appenzeller

I think it’s really good for customers and consumers in the short term. Specifically for the laptop market, having Intel and NVIDIA collaborate is amazing.

I wonder what’s going to happen with any of the integrated graphics or AI products at Intel. They might just push a reset and give up on that for now. They currently don’t have anything competitive. There was the Gaudi F4, which is more or less done, and there were the Intel graphics chips, which never really competed at the high end.

From that perspective, it makes a lot of sense for both sides. For Intel, they needed a breath of fresh air. They were sort of desperate, so I think it’s a very good thing.

I think AMD is screwed, right? If your 2 arch nemeses suddenly team up, that’s the worst possible news you can have. They were already struggling. Their cards are good; their software stack is not. They were getting very limited traction, and now they have a bigger problem on that side.

I think Arm is a little bit screwed as well, because its biggest selling point was, “Look, we can partner with everybody that doesn’t want to partner with Intel.” In a sense, NVIDIA is probably the most dangerous of the future CPU competitors. Now NVIDIA suddenly has access to Intel technologies and might go in that direction.

It reshuffles the cards. I did not see this coming. I think it’s an amazing development.

Erik Torenberg

Yeah, it will be very interesting to see this play out. To my point, it’s been a packed news week.

The other thing we wanted to pick your brain on, since we have you here, Dylan, is the other news about Huawei unveiling its AI roadmap. Obviously, they’re hyping up the capabilities. I think you guys have been ahead of the curve in trying to gauge what the Atlas 950 SuperCluster can actually do.

I’d love your thoughts on everything that’s going on on the China front. This is coupled with DeepSeek saying its next models are going to be on domestically produced Chinese chips, and the Chinese government banning companies from buying the NVIDIA chips produced specifically for China.

There are just a lot of dominoes falling right now in the semiconductor market in China. I’d love your take overall, and I’d like you to drill into some detail.

Dylan

When you zoom out, let’s walk from 2020, because I think it’s really important to recognize how cracked Huawei is, even historically. They’ve always been really good. Sure, initially they stole Cisco source code and firmware and all this stuff, but then they rapidly passed Cisco up, as well as every other telecom company.

In 2020, they released an Ascend chip and submitted it to impartial public benchmarks. They were the first to bring 7-nanometer AI chips to market. They were the first to do that. You could still say NVIDIA was ahead, but the gap was almost nothing. The market was so nascent then that Huawei could have really taken it over.

Huawei got banned by the first Trump administration from accessing the full foreign supply chain, and that went into effect in 2020. They were only able to make a small volume of these chips, but they had trained significant models on the chips they made.

Over the next couple of years, NVIDIA continued to accelerate. Because Huawei was banned from TSMC, it had to figure out how to manufacture at SMIC, the domestic TSMC. In parallel, Huawei was trying to use shell companies to manufacture at TSMC and acquire memory from Korea, and so on and so forth.

By the end of 2024, this had gotten into full swing, and it was caught. They finally shut it down, but Huawei was able to acquire 3 million chips—2.9 million chips—from TSMC through these other entities, roughly $500 million worth of orders.

That ends up being a $1 billion fine that the US government gave TSMC, if I recall correctly, or at least there was a Reuters article about it. I don’t know if they actually issued it, which is important and interesting to gauge, because the number of Ascends floating out there has not consumed this entire capacity yet.

Now we get to 2025. The H20 got banned at the beginning of the year. NVIDIA had to write off huge amounts of money. Our revenue estimate for NVIDIA in China for just the H20 was north of $20 billion, because that’s what they were booking in capacity and then had to write off.

They cut the supply chain. They just said, “No, we’re not doing this anymore.” Their inventory got reapproved, and they resold the inventory, but now NVIDIA’s question is, “Do we even restart production?”

China is now saying, “We don’t need NVIDIA. We have domestic alternatives,” whether it’s Huawei or Cambricon. These companies have capacity, but most of this capacity is still foreign-produced, whether it’s wafers from TSMC or memory from Korea—Samsung and SK hynix.

So the question is, how much can they do domestically? There are 2 fronts. There’s logic, meaning replacing TSMC, and there’s memory, meaning replacing SK hynix, Samsung, and Micron.

Dylan Patel

On the logic side, they are behind, but they’re really ramping there. I think they can get to the production-capacity estimates needed, and the US is still allowing them to import pretty much all the equipment necessary. The bans are really for beyond the current generation of technology—beyond 7 nanometers. Even though the government says they’re for 14 nanometers, the actual equipment that’s banned is only for below 7 nanometers.

They’ll be able to make a lot of 7-nanometer AI chips and maybe even get to 5 nanometers using existing equipment for 5 nanometers, rather than using the new techniques. There’s the logic side and then there’s the memory side. The aspect of Huawei’s announcement that was surprising was that they’re doing custom memory.

Speaker 2

Right? Yeah.

Speaker 1

That’s the part that’s really exciting. They announced 2 different types of chips for next year: 1 focused on recommendation systems and prefill, and 1 focused on decode.

Speaker 2

There’s the trend these days.

Speaker 1

Yeah. At NVIDIA, it’s the same thing. They just announced a prefill-specific chip recently. Numerous AI hardware startups are really focusing on prefill versus decode, so this split of inference into 2 workloads is becoming common. Huawei is doing the same thing for its chip next year.

What’s interesting is that the decode one has custom HBM. What does that mean? What’s the manufacturing supply chain? That’s the tricky one, right? How much of that custom HBM can they manufacture? NVIDIA and others are also adopting custom HBM starting next year, so it’s not like the manufacturing capacity isn’t there. Maybe it’s going to consume a bit more power, and maybe it’s going to have slightly lower bandwidth, but the fact that they’re able to do some of the same things that NVIDIA plans to do and AMD plans to do in their memory is evidence that they’re catching up.

The main question that remains is production capacity. As far as NVIDIA being banned in China—China saying, “Don’t buy NVIDIA chips”—I think that’s fine for a period of time, from China’s perspective. If I’m China, that’s fine because you have all this capacity that you shipped in 2024 that you haven’t turned into AI chips. Now you’re turning them into AI chips and running that stockpile down.

What about the transition from running that stockpile down to ramping your new stuff? That transition is the one that’s really tricky. China is either shooting itself in the foot by not purchasing NVIDIA chips during that period, or it’s able to ramp.

Dylan Patel

I think they’ll be able to ramp. I think it’ll take a little bit longer, and there’ll be a gap in between where China probably backtracks and says it’s fine. ByteDance will be begging for NVIDIA chips. They don’t want to use Cambricon; they use some Cambricon and some Huawei, but they really want to use NVIDIA because it’s way better.

They don’t care about the domestic supply chain. They want to make the best models and deploy their AI as efficiently as possible. The government can mandate them not to do it. So it’s not that NVIDIA isn’t competitive; it’s that the government is trying to instigate it.

The last thing is that there’s always the argument of, “Hey, if banning NVIDIA chips to China is so good for China, why didn’t China do it for itself?” They’re finally doing it for themselves. So again, it’ll be interesting to see.

Smuggling is still happening. Reexportation of chips from other countries to China is still happening at some volume—low volume, lower-medium volume. Direct shipments of NVIDIA chips that are legally allowed to China aren’t necessarily happening today, but they may have to restart at some point because China won’t have the production capacity. It would just have so many fewer AI chips being deployed domestically versus the US, and at some point you have to pick: “Am I all about the internal supply chain, or am I all about chasing super-powerful AI?”

Erik Torenberg

Is there a negotiation angle here as well? There are still discussions ongoing about exactly what the boundaries are and what can be exported to China. These are sort of well-timed announcements if you want to make a point that the US should allow more exports. Do you think that’s a factor or not?

Dylan Patel

Yeah. In the report we did a few weeks ago about Huawei’s production capacity and the supply chain, there was a bit in there about how, honestly, if you’re China and you want NVIDIA, you do want NVIDIA chips. How do you play this? It’s by hyping up your domestic supply chain.

Erik Torenberg

And it’s like, yes, we can do everything. Huawei announced the craziest shit possible—7 years of shit, or 3 years of roadmaps that are so—did they read your report? Basic question.

Dylan Patel

I think they do. I mean, say we’re banning NVIDIA, right? The government official is going to think alongside the lobbying from domestic players, “Of course we want to ship them better AI chips. We’re losing this market. We can’t lose this market.” It’s sort of 10,000 IQ. We’re here playing checkers while they’re playing chess.

Erik Torenberg

Negotiation aside, in that report you talked about HBM, or high-bandwidth memory, being a bottleneck for Huawei. To your point about one of the surprising aspects of the announcement, do you think it’s credible that it’s no longer a bottleneck based on what they’re saying, or is it just hype?

Dylan Patel

Production-capacity-wise, it is still absolutely a bottleneck. Certain types of equipment required for making HBM need to be imported. They’re working on domestic solutions, but as far as we know, they have not imported enough equipment for this.

Although if you look at Chinese import data for different types of equipment, fabs spend—depending on the process technology—roughly different amounts of money on lithography, etch, deposition, and metrology, these different steps. Historically, lithography has hovered around 17% or 18%; with EUV, it grew to 25%. But China wanted to stockpile lithography and was worried about the coming ban, so it was importing lithography at a much higher rate. Thirty to 40% of its equipment imports were lithography, and it was just stockpiling lithography equipment.

That has sort of reversed now. If you look at the monthly import-export data, both into provinces in China and out of countries, you can see that etch—specifically—is skyrocketing. The main thing about stacking HBM is that with each wafer, you have to etch and create something called a through-silicon via so it can connect from the top to the bottom. Then you stack them on top of each other—12 high or 16 high for HBM. That’s how you make super-high-bandwidth memory, and their imports of etch equipment are skyrocketing now.

Erik Torenberg

They don’t have the production capacity yet. How fast can they ramp it as a function of how much equipment they can get? And B, the yields, right?

Dylan Patel

Improving yields is really hard in manufacturing. Intel and Samsung are really good, and TSMC is just amazing. Not that those companies suck; I think that’s a better way to put it. Those are the 2 things: yield and production capacity.

Yield-wise, they haven’t even started production of HBM3. They’ve only done some sampling of HBM2. HBM3 came out a few years ago, so there’s still quite a ways to go up the learning curve. I expect them to catch up faster than it took for the technology to be developed because it exists. We know how to do it in the world; it’s just a matter of actually doing it versus inventing it.

The other issue is production capacity. A couple months of import-export data isn’t enough to set up years’ worth of supply-chain buildup, which is what we have today in Korea for the Korean companies. SK hynix is also investing in the US, in Illinois, and Micron is primarily in Japan. The American memory companies are primarily in Japan and Taiwan, but they’re also expanding in Singapore and the US now.

There’s so much capital that’s been invested. It would take some time for China to build up that production capacity to actually match the West. When I say the West, I mean East Asia—in non-China East Asia—in production capacity. It’ll take some time to get there. I don’t think it’s a question of whether they can design this; it’s always a question of whether they can manufacture it. The thing Jensen would say is that you’re betting on China not being able to manufacture.

Erik Torenberg

That’s a matter of when, not if.

Dylan Patel

And that’s the whole calculus that I think the US government has to be aware of when it’s like, “Hey, what level of AI chips do we sell? Do we sell everything?” Probably not, because AI is far more powerful, and the end market for AI is going to be way larger than the end market for semiconductors and equipment.

Do we sell, you know, what level do we sell at? How much can China make at each specific performance tier? Then analyze that: What’s the volume? Figure out what is okay, which is maybe a little bit above or around the same level.

Erik Torenberg

Yeah. So, to your point on playing chess versus checkers, if you’re Jensen, what would your next move be, given the situation at hand?

Dylan Patel

It’s partially true that he’s more afraid of Huawei than he is of AMD.

Erik Torenberg

Right? He called them formidable.

Dylan Patel

Yeah. I mean, Huawei has beaten Apple, right? They passed Apple in TSMC orders, and they passed Apple in phone market share—not in the United States, but in many parts of the world—before the bans came down. Even now, they’re growing back again in market share without Western supply chains.

They’ve done this to numerous other industries. I would say Huawei is a formidable competitor, right? They’ve beaten a lot of industries, and so it’s reasonable that he’s afraid of them. It’s sort of—you know, he’s not afraid of AMD.

I think the best thing is to try and see whether what Huawei announced is reality rather than their aspirational target.

Erik Torenberg

Yeah.

Dylan Patel

We shouldn’t wave away all doubt about manufacturing capacity, which I think is not fair. Manufacturing capacity is a real bottleneck for them. Yield learning is a real bottleneck, although perhaps temporarily. We’ll see how long that lasts, how fast the rest of NVIDIA’s technology advances past what Huawei is capable of, and how fast Huawei is able to close the gap.

I think his main pitch would be that Huawei is real. They’re a formidable competitor, and they’re going to take over not just the Chinese market, but also foreign markets—whether it’s the Middle East, Southeast Asia, South Asia, Europe, or Latin America. Everywhere besides America.

I think Noah Smith has this analogy. The whole idea is that you should let China go, right? Make them have their own domestic industry that is so different from the rest of the world—kind of what happened with Japan in the 1970s, 1980s, and 1990s. Their PCs were so specific and hyper-optimized to the Japanese market, with the weird scroll wheel on these Japanese PCs. You literally go like this and it scrolls, and the touchpad is a circle with that around it. Things like that are so weird.

Erik Torenberg

Totally.

Dylan Patel

The rest of the world doesn’t care, but the Japanese market likes it, right? His whole idea is: Let them go—keep their technology within China—and then that’s deadweight loss, and they never expand outside of China, versus serving the whole world.

But the whole risk is that the opposite can also happen, right? Our technology is hyper-optimized to run language models at this scale and reinforcement learning. Hardware-software co-design can take you down a branch of the tree that is a dead end. Then China, because they’re not allowed to access this tree, says, “Oh, okay,” and ends up in the optimal spot, right? We hit a local minimum; they had a global maximum. That sort of technological goal-posting is what Noah Smith’s analogy is. I like it a lot. I don’t know if it’s accurate, but it’s an interesting one.

Erik Torenberg

Yeah. I love that. Well, actually, maybe just taking a step back from current events—even though there’s so much to talk about right now—last time you appeared with us, Nvidia came up, obviously, and you talked about a couple of the potential paths forward for Nvidia. Give us maybe the bull case, the bear case.

Dylan Patel

Fair enough. There’s a lot embedded in their numbers now. What’s interesting is that the consensus from the banks across the hyperscalers—Microsoft, CoreWeave, Amazon, Google, Oracle, and Meta—is $360 billion of spend next year across all of them. Those are the 6 hyperscalers I would consider hyperscalers.

My number is closer to $450–$500 billion. That’s based on all the research we do on data centers, tracking each individual data center, the supply chains, and so on.

Erik Torenberg

So this is just Nvidia spend?

Dylan Patel

This is capex for the hyperscalers. Capex gets split up across different companies, but the vast, vast majority still goes to Nvidia.

Nvidia is in a position where they can’t take share. They grow with the market and defend share.

Erik Torenberg

Yeah.

Dylan Patel

The question is: How fast is the growth rate of capex for hyperscalers and other users? The reason I included Oracle and CoreWeave as hyperscalers, even though they’re traditionally not called hyperscalers, is because they are OpenAI’s hyperscalers.

Erik Torenberg

Right. Right.

Dylan Patel

When you look at the Oracle announcement, I don’t understand why people don’t think this is crazier. They did the most unprecedented thing in the history of stocks, public companies, and companies ever: They gave 4-year guidance. It made Larry the richest man in the world, among all these other things.

The question is: How fast does revenue grow? Do you think OpenAI, which signed a $300 billion-plus deal with Oracle, will actually be able to pay $300 billion through raising capital and revenue? It gets to a rate of over $80 billion or $90 billion a year in just a handful of years.

Do you believe the market will grow that fast? It’s very possible. For OpenAI, what is their revenue going to be exiting next year? Some people think $35 billion, some people think $40 billion, and some people think $45 billion in ARR by the end of next year. This year they hit $20 billion.

If that growth rate is maintained, then all of that cost goes to compute, plus all the capital they continue to raise. The financials they gave investors for their last round were, “We’re going to burn $15 billion next year.” It’s probably more likely to be closer to $20 billion. They’re not generating cash flow, and they’re not going to be profitable until 2029.

They’re going to continue to burn $15–$25 billion of cash each year, plus revenue growth. That’s their compute spend. You do this for Anthropic, you do this for OpenAI, and you do this for all the labs. It’s very possible that the pie gets to more than $500 billion—not $360 billion—next year in total capex, and the pie continues to grow for hyperscalers.

Nvidia says it’s actually going to be multiple trillions of dollars a year in AI infrastructure, and Jensen is going to capture a huge portion of it. That’s his bull case: AI is actually so transformative that the world gets covered in data centers, and the majority of your interactions are with AI—whether it’s business productivity and telling an agent to write code, or you’re just talking to your AI girlfriend, Annie. All of this is running on Nvidia, for the most part. The bear case is, you know, even if it does grow a lot.

Erik Torenberg

Yeah, so the bull case for a second: I think fundamentally the value creation is there. I mean, trillions of dollars of value with AI—I can totally see this happening. So assume it’s true: Where will Nvidia top out?

Dylan Patel

I guess it depends on how much you believe in takeoffs. If there is a takeoff scenario where powerful AI builds more powerful AI, which builds more powerful AI, and each level of intelligence enables more for the economy, how many monkeys can you employ in your business versus how many humans? You know, what is the value creation of a human versus a dog? It’s sort of the same with AI.

Erik Torenberg

I mean, in this case, the value creation could be hundreds of trillions, if not more than that.

Do you even need this? If you take every white-collar worker and make them twice as productive with AI, that’s in the hundreds of trillions, isn’t it?

Dylan Patel

Yeah, but what is “twice”? If you talk to people at the labs, what does “twice as productive” even mean? It’s replaced them, right? It’s 10 times better than that. I don’t know how soon, but if white-collar work is essentially useless without a constant stream of LLM tokens that make people productive, at that point you basically can tax every single knowledge worker in the world, which is most workers in the world long term.

Erik Torenberg

Yeah. So I don’t know. What’s your guess? Give us a number. What’s the cap?

Dylan Patel

Cap? I mean, why aren’t we making a Matrioshka brain? I don’t know. At some point, the machine says humans don’t need to live, and we need even more compute.

Speaker 1

One step before that, maybe—

Speaker 2

Are we colonizing Mars yet?

Dylan Patel

TBD. I don't know, man. I find it completely impossible to predict anything beyond 5 years, given how much stuff is changing. Linear time is a large number; I'll leave that to economists, right? Honestly, supply-chain stuff is 3 or 4 years out, and that's it.

Speaker 1

And then the fifth year is sort of YOLO, right?

Dylan Patel

I just try to ground myself in the supply-chain stuff. What is the adoption of AI? What is the value creation? What's the usage like? You can see that in a short horizon. Beyond that, I don't know. Are we all going to be connected to computers through BCIs and stuff? I don't know, dude.

You saw Elon's thing, right? He's like, "Yeah, humanoid robots are why Tesla's worth more than $10 trillion." Okay, great. What is all that being trained on? Great, Nvidia. Okay, awesome. So that's worth also $10 trillion, right? I don't know. It's too out there for me. I don't like the out-there discussions.

Speaker 1

Very fair.

Speaker 2

Read some sci-fi books.

Speaker 1

So, just pulling out the thread where you talked about how market share can't really grow just because it's such a dominant market share. We talked about—or you guys talked about—the moat of Nvidia last time, and obviously this moat is tied to maintaining the very high market share they currently have. I love the historic journey you took us through with Huawei earlier. Can you walk through what Nvidia did throughout history to build its moat?

Dylan Patel

It's super awesome because Nvidia failed multiple times in the beginning, and they bet the whole company multiple times. Jensen is just crazy enough to bet the whole company, right? Whether it was certain chips, ordering volume before he knew they even worked, when it was all the money he had left, or ordering volumes for projects he had not won yet.

I heard a rumor—or, not a rumor, but a story—from someone who's a graybeard in the industry and who I think would know. He said, "No, Nvidia ordered the volume for the Xbox before Microsoft gave them the order." They were literally just like, "Fuck it. YOLO." I don't know how true that is. I'm sure there's more nuance there, like a verbal indication or whatever, but the order was placed before he got the order, is what he said.

With the crypto bubbles, there were a couple of them. Nvidia did its damn best to convince everyone in the supply chain that it wasn't crypto, that it was gaming, and that it was durable, real demand. It was gaming and data center and professional visualization, and therefore everyone should ramp production.

They all ramped production and spent all this capex on increasing production and building out new lines for Nvidia. They paid per item, bought them, sold them, and made shitloads of money. Then, when it all fell apart, they just had to write down a quarter's worth of inventory, whatever.

Speaker 1

Everyone else was like, "Well, crap, I have all these empty production lines," right?

Speaker 2

But what did AMD do then? Its chips were actually better for crypto mining, right? In terms of the amount of silicon cost versus how much you hash. But AMD just didn't. AMD was like, "We're not really going to raise production," right? As a reasonable thing. It wasn't a matter of striking while the iron was hot.

Dylan Patel

Nvidia has done the same thing in recent times. They've ordered capacity that no one believes, multiple times. They see the demand, obviously, but in many cases their number for Microsoft was higher than Microsoft's internal planning. Microsoft's internal planning went up, but Nvidia's number for Microsoft was way higher. It was like, "We just don't think Microsoft is going to need this much, even though they tell us this." Who the heck is like, "No, no, no, customer, you're going to buy more"?

When the orders come through the supply chain, it's like, "I have to put, 'Pay NCNR'—non-cancelable, non-returnable."

Speaker 1

I asked a question in Taiwan once. Colette, the CFO, and Jensen, the CEO, were both there. It was a room full of mostly finance bros, and they were asking stupid finance questions 3 days before earnings, so obviously they couldn't answer anything because of SEC regulations.

My question to them was, "Look, Jensen, you're so vibes-driven, gut-feel-driven, and visionary. Colette's a CFO. She's amazing in her own right, but those personalities clash. How do you work together?"

He's like, "I hate spreadsheets. I don't look at them. I just know." That was his response. Of course, the best innovators in the world have really good gut instinct, right? So the gut instinct to order—

Speaker 2

With non-cancelable orders, when you don't know, they've had to write down over their history multiple times—many, many billions of dollars in cumulative orders, whether it be the H20, which is more regulatory, or other cases where they've ordered and had to cancel.

Speaker 1

Is that many billions?

Dylan Patel

It's many billions.

Speaker 2

Peanuts.

Dylan Patel

Well, it depends. The crypto write-down was multiple billion when their stock was less than $100 billion. It's peanuts compared to the upside, right?

Speaker 1

I think everything Nvidia did was right, and everything AMD did was wrong in that scenario. But it is crazy, especially in a cyclical industry like semiconductors, where companies go bankrupt all the time, which is why we have all this consolidation.

Dylan Patel

These bets were totally worth taking. Yes.

Speaker 2

If you look at it from a risk-return perspective, these bets were totally worth taking. If you look at it from the perspective of, "I'm a CEO and I want to have predictable quarters for Wall Street," it's a very different story. I think that's where part of the tension is now.

Speaker 1

We made one of Jensen recently and put it on social media—Instagram, TikTok, XHS, Redbook, and Twitter, of course. I really liked it because he's like, "The goal of playing is to win, and the reason you win is so you can play again." He compared it to pinball, where you just play all day and keep getting more rounds. His whole thing is, "I want to win so I can play the next game."

Dylan Patel

It's only about the next generation. It's only about now and the next generation. It's not about 15 years from now, because it's a whole new playing field every time—or 5 years from now. I think the risk-reward is correct.

Speaker 2

But few people take these kinds of risks. Nvidia is the only semiconductor company worth north of $10 billion that was founded as late as it was.

Dylan Patel

MediaTek was founded in the early 1990s, and Nvidia and everyone else are mostly from the 1970s.

Speaker 1

The big ones.

Speaker 2

I think you raised this great point about betting the farm, and he's actually been wrong a couple of times, to your point.

Speaker 1

Mobile, right? What the hell happened with mobile?

Dylan Patel

Exactly. He still takes those risks. Mark actually had this great conversation with Eric where he talked about being founder-run, where you have this memory of the risks you took to get to where you are today.

In a lot of cases, if you're a CEO brought on later, you're sort of like, "Okay, continue to steer the ship as is." But in this case, Jensen remembers all the times they almost went belly-up, and he's like, "I've got to keep making bets like that."

How do you think he's changed? He's been one of the longest-running CEOs—over 30 years. He's kind of right up there with Larry Ellison now.

Speaker 1

I mean, obviously, I'm 29. I don't freaking know what he was like.

Speaker 2

I've watched a lot of old interviews.

Dylan Patel

I won't say he wasn't—

Speaker 1

Longer than you've been alive.

Dylan Patel

Yeah, exactly. Nvidia was founded before I was born. I'm '96, right? Maybe anything over the last couple of years is more relevant.

Speaker 1

No, no, probably better. I think even watching old interviews, right? I watched a lot of old interviews and a lot of old presentations he's given.

Dylan Patel

One thing is that he's just sauced up and dripped up. The charisma he's gotten has only gotten stronger.

Erik Torenberg

Right?

Dylan Patel

Which is an interesting point. I don't know if it's quite relevant.

Erik Torenberg

Totally agree with that.

Dylan Patel

But the man has learned to be more of a rock star, even though he was always charismatic. He's a complete rock star now, and he was a rock star a decade ago, too. It's just that people maybe didn't recognize it.

I think the first live presentation of his that I watched was at—what's the conference? CES, like 2014 or 2015 or whatever. He was talking only about AI. He was telling all these gamers about AlexNet and self-driving cars. It's like, know your audience, first of all, but also, it had nothing to do with consumer electronics. It was gaming, you know. At the time, I was a teenager moderating gaming and gaming-hardware subreddits.

At the time, I was also half like, “Holy crap, this is amazing.” But I was also half like, “I want you to announce a new gaming GPU, right?” On the forums, everyone was quickly like, “Screw this. I want to hear about the gaming GPUs.” Nvidia's price gouging.

Of course, Nvidia has always had the attitude that it prices for the value, plus a little bit, because it's just smart enough to know. I'm guessing Jensen just has the gut feel for how to price things.

Erik Torenberg

He'll change the price, at least on gaming launches, right up until right before the presentation. Wow.

Dylan Patel

So it really is a gut-feel thing, probably. Anyway, he had that charisma to know what was right, but I think a lot of people were like, “Oh, no, whatever. Jensen's wrong. He doesn't know what he's talking about.”

But now, when he talks, people are like, “Oh, very, very...” It might just be that he's been right enough.

Erik Torenberg

Yeah. There's a post on X recently that said he had moved up into god mode with a select group of CEOs.

Dylan Patel

Who's god? Who are the other gods?

Erik Torenberg

It was Zuck. Who are the other gods?

Dylan Patel

Elon.

Erik Torenberg

Elon, Zuck, and Jensen.

Dylan Patel

Nice. Nice. Okay.

Erik Torenberg

A cool crew to be in.

Dylan Patel

So we pray to Silicon Valley. It's sort of a cult now, is it?

Erik Torenberg

Exactly. Just on one last thing about people: you mentioned Colette Kress, his CFO. There's a famously loyal crew at Nvidia, even though all of the OGs could retire at this point. Is there anyone akin to Gwynne Shotwell at SpaceX, or previously Tim Cook to Steve Jobs at Apple, who is at Nvidia today?

Dylan Patel

I mean, he had 2 cofounders, right? Let's not overlook that. 1 of them isn't involved and hasn't been for a long time, but the other one was involved until just a few years ago. So it's not just Jensen running the show, right?

Erik Torenberg

Totally.

Dylan Patel

Although he was running the show, there are quite a few people on the hardware side. There's someone at Nvidia who's mythical to me. When you talk to the engineering teams, he leads a lot of the engineering teams. He's a private person, so I don't want to say his name, actually.

Erik Torenberg

Fair enough.

Dylan Patel

But he's effectively the chief engineering officer. That's his role, and people within his organization know who he is. I think there are people like that. He's intensely loyal, and there are a number of these types of people.

There's another fellow who's associated with all these innovative ideas at Nvidia, but he's the guy who literally says, “We need to get this silicon out now. We're cutting features.” That's what he's famously known for, and all the technologists at Nvidia hate him. This is a second guy, also intensely loyal to Nvidia, who's been around for a long time.

When you have such a visionary, forward-looking company, 1 problem is that you get lost in the sauce. You think, “I want to make this. It's got to be perfect and amazing.” You have to have people who say, “Screw it. Cut it. We'll put it in the next 1. Ship now. Ship faster.” That's really hard to do in a space like silicon.

The thing about Nvidia that's always been super impressive, going back to the beginning, is their first successful chip. They were going to run out of money, and Jensen had to get money from other people just to finish the development. Even then, he had barely enough money because Nvidia had already had a failed chip before that.

When the chip came back, it had to work. Otherwise, Nvidia wouldn't make it. They could only pay for what is called a mask set. Basically, you put these stencils into the lithography tool, and that tells it where the patterns are. You put the stencil in, deposit material, etch material away, and repeat the process.

You stack dozens of layers on top of each other to make a chip. These stencils are custom to each chip, and today they cost tens and tens of billions of dollars. Even back then, it was still a lot of money. They could only pay for 1 set.

The typical thing with semiconductor manufacturing is that, no matter how well you simulate or verify everything, you'll send a design in and have to change it. There's always going to be something. It's so hard to simulate everything perfectly. The thing about Nvidia is that they tend to get it right the first time.

Erik Torenberg

Yeah.

Dylan Patel

Even great-executing companies like AMD or Broadcom often have to ship revisions. They're denoted as A followed by a number or B followed by a number. Nvidia almost always ships A0. They sometimes ship A1.

The letter is basically the transistor layer, and the number is the wiring that connects all the transistors together. Nvidia will start production of the A layer and ramp it really high, then hold it right before transitioning to the metal layers, just in case it needs to change them.

The moment Nvidia confirms that it works, it can blast through a lot of production. Everyone else is like, “Let's get the chip back. Okay, A0 doesn't work. We have to make this tweak, make that tweak, and get the chip back.” It's called a stepping.

At Intel, we were very jealous of Nvidia at that time. Nvidia consistently delivered on the 1st revision. We did not. In the data center CPU group, there was 1 product where Intel got to E2.

Erik Torenberg

E2 is like a 15th revision. This is—

Dylan Patel

This was around the peak of AMD's market share, when AMD skyrocketed in market share versus Intel. Intel was at E2—15 steppings.

Erik Torenberg

Because it's quarters of delay, right? I mean, it's catastrophic for a go-to-market.

Dylan Patel

Yeah. Each 1 is a quarter of delay or something like that. It's absurd. I think that's the other thing about Nvidia: “Screw it, let's ship it. Let's get the volume as soon as possible.”

Nvidia has some of the best simulation and verification, which lets it go from design—or from idea—to shipment as fast as possible. It cuts out unnecessary features that could delay the product and makes sure it doesn't have to do revisions, so it can respond to the market as quickly as possible.

There's a story about Volta, which was the first Nvidia chip with tensor cores. Nvidia saw all the AI work on the prior-generation P100 Pascal and decided to go all in on AI. It added the tensor cores to Volta only a handful of months before sending it to the fab. It said, “Screw it. Let's change it.” It's crazy.

If Nvidia hadn't done that, who knows? Maybe someone else would have taken the AI chip market. There are all these times when Nvidia makes major changes, but there are often minor things that have to be tweaked, too, like number formats or some architectural detail.

Nvidia is just so fast. The other crazy thing is that it has a software division that can keep up with that. If you come out with a chip and basically require no stepping, it's immediately in the market. Being ready with the drivers and all the infrastructure on top of that is just super impressive.

Guido Appenzeller

Yeah, I love that point because you think of Nvidia benefiting from tailwind after tailwind, but I think both of you are saying you've got to move fast enough and execute well enough to take advantage of those tailwinds.

Erik Torenberg

And if you think about it—and by the way, I loved your CES story. I'm just envisioning him more than 10 years ago talking about self-driving cars. But if you think about nailing the video game tailwind, VR, Bitcoin mining, and obviously AI now, one of the things that Jensen talks about today is robotics and AI factories.

Maybe my last question on Nvidia: What do you think about the next 10 to 15 years? I know calling beyond 5 is hard, but what does Nvidia's business look like? It's really a question of—and this is, I think, every time I've talked to some executives at Nvidia, I've asked this question because I really want to know, and they won't answer it, obviously—but what are you going to do with your balance sheet? You are the highest-free-cash-flow company; you have so much cash flow.

Sarah Wang

Now the hyperscalers are all taking their cash flow way down, right, because they're spending on GPUs. What are you going to do with all this cash flow? Even before this whole takeoff, he wasn't allowed to buy ARM, right? So what can he do?

Dylan Patel

With all this capital and all this cash, right? Even this $5 billion investment in Intel has regulatory scrutiny there. It's in the announcement: “Yeah, this is subject to review,” right?

Erik Torenberg

Yeah.

Dylan Patel

You know, I imagine that'll get passed, but he can't buy anything big. He's going to have hundreds of billions of dollars of cash on his balance sheet. What do you do? Is it starting to build AI infrastructure and data centers? Maybe. But why would you do that if you can just get other people to do it and take the cash?

Sarah Wang

Well, he's investing those, right?

Dylan Patel

Investing peanuts, right?

Guido Appenzeller

Right. You know, he recently gave CoreWeave a backstop because today it's really hard to find a large number of GPUs for burst capacity. Like, “Hey, I want to train a model for 3 months. I have my base capacity for all my experiments, but I want to train a big model for 3 months, and then I'm done.”

Sarah Wang

We know from our portfolio, yeah.

Dylan Patel

Yeah, so Nvidia sees this issue. They think it's a real problem with startups; it's why the labs have such an advantage. But what if I could—right now, most companies in the Valley spend, what, 75% of their round on GPUs, right? Or at least—yeah, with [inaudible].

Guido Appenzeller

What if you could do 75% in 3 months on 1 model run? You know, really scale and have some sort of competitive product, and then you have the model. Then you raise more capital or start deploying. What do you do with it? Is it—

Erik Torenberg

Start buying a crapload of humanoid robots and deploying them? But they don't really make good software. They don't make really that amazing software for them. In terms of the models, they make the layer below, which is great. Where they deploy their capital is the question.

Sarah Wang

He has been investing up and down the supply chain a little bit, though, right? Investing in the neoclouds, investing in some of the model-training companies.

Dylan Patel

Yeah, but again, it's small fries. He could have done the entire Anthropic round if he wanted to. Of course, he didn't. He could have done the entire OpenAI round, or the entire xAI round.

Erik Torenberg

Do you think these are things he should be doing?

Sarah Wang

Yeah, good question.

Dylan Patel

I don't know, right? I think picking winners is obviously really tough for him because he has customers all across this ecosystem. And if he starts picking winners, then his customers will be even more anxious to leave and give even more effort to whether it's AMD or some startup or their internal efforts, et cetera, et cetera, right? Buying TPUs, whatever it is. He can't just invest in these—

Erik Torenberg

We'll quote you for the next round that we're raising, but anyways—

Dylan Patel

He could make venture a dead industry, take all of the best rounds.

Sarah Wang

Put a lot of business in it. Yeah.

Erik Torenberg

You know, you could do the seeds and then have Jensen mark you up.

Guido Appenzeller

I don't like it.

Erik Torenberg

Is he also reshaping his market? I mean, look, a couple of years ago there were 4 big purchasers of these cards. You just listed 6. To what extent is that a strategy?

Sarah Wang

Him and Nebius—there's a long list there, of course. Yeah.

Erik Torenberg

Is there a strategy?

Dylan Patel

It is. I think it absolutely is. But he didn't have to put much capital down to do this.

No, but if you look at the grand amount of capital that he spent investing in the neoclouds, it's a few billion, but he has a lot of other levers if he wants to.

Erik Torenberg

Right, right. Allocations, as you mentioned. What's nice is that historically you gave volume discounts to hyperscalers, but because he can use the argument of antitrust, he's like, “Everyone gets the same price.” So what should he do with the cash, or what should guide his decisions?

Dylan Patel

I mean, I think there is an argument that he should invest in data centers—and only the data-center layer, not what goes in the data centers—so that more people build data centers. If market demand continues to grow, data centers and power are not the issue. Invest in data centers and power.

I've said that to them: They should invest in data centers and power, not in the cloud layer, because the cloud layer is not commoditized, but it's quite a commoditized complement, right? That's the whole phrase. And I won't say being a cloud is commoditized, but you certainly have a lot of competitors who are decent now. You've educated commercial real estate and other infrastructure investment firms into going into AI infrastructure as well. So I don't think it's the cloud layer that you invest in.

Do you invest in data centers and energy? Yes. Do you invest in it because that's the bottleneck for your growth, really? A, how much people want to spend and can spend; and B, the ability to actually put them in data centers.

And then robotics. I think there are areas he could invest in, but nothing requires $300 billion of capital. So what do you do with the capital? I really don't know, and I feel like Jensen has to have some idea. There's some visionary plan here because that's what shapes the company.

I mentioned $200 million of free cash flow, $250 billion of free cash flow a year. What do they do with it? Do they just buy back stock forever? Do they go the Apple route? The reason Apple hasn't done anything interesting in nearly a decade is they've got a non-visionary at the head. Tim Cook is great at supply chain, and they're just plowing the money into buybacks. They're not really doing anything in automotive; the self-driving-car thing failed. We'll see what happens with AR/VR. We'll see what happens with wearables, but Meta and OpenAI might be even better than them. We'll see, and others, right?

So what does he invest in? I have no clue. But nothing that requires so much capital and actually gets a return is the tough question.

Sarah Wang

Because the easy thing is my cost of equity, right? I just buy back stock—

Guido Appenzeller

And it doesn't completely change the company culture. I think that's another thing, right? There are probably areas you could invest in, but you suddenly end up with the company doing 2 completely different things, which are very difficult to keep aligned.

Dylan Patel

But they do 10 completely different things, right? I mean, one way to look at it is: We build AI infrastructure, and humanoids around the world are AI infrastructure. Data centers and energy are AI infrastructure.

Sarah Wang

So the humanoids would totally work, right? If you suddenly start pouring concrete and building power plants, that's a completely different culture, a completely different set of people, and it gets much, much harder.

Erik Torenberg

There are all these different areas where he could use capital to allow something to happen, right? Not necessarily owning it himself.

Guido Appenzeller

And look, bear in mind, at Intel, one of the biggest problems we had was that our customer base sucked, right? Most of the chips we sold went into the large hyperscalers. They were way too concentrated, and they build their own chips, so they can push down your prices. So, honestly, spending it on diversifying the customer base…

Erik Torenberg

In 2014, you guys should have just charged so much that your margins were 80%. What would the world have done?

Guido Appenzeller

Nothing. The margins were pretty good back then. That wasn't the problem. That wasn't the primary problem. They were 60, 65.

Erik Torenberg

They were 80 still.

Guido Appenzeller

Yeah.

Erik Torenberg

Oh boy. Jensen's PTSD is kicking in here.

Well, wait. I think Guido's comment is actually a really good segue into something else we wanted to talk to you about, which is the hyperscalers. One of the reasons that I love reading SemiAnalysis is that you guys make these out-of-consensus calls that you're often right about.

Dylan Patel

Only often. But you have a Jensen hit rate that's very high.

Erik Torenberg

Where's my billion-dollar, PV-positive bet?

The one that caught my eye was Amazon's AI resurgence. I wanted to talk to you a little bit about that because I think we found it pretty interesting being on the ground and helping our portfolio companies pick who their partners are. We have some microdata on this, so can you walk us through why they are behind?

Dylan Patel

Yeah. So, in Q1 2023, I wrote an article called “Amazon's Cloud Crisis.” It was about how all these neoclouds were going to commoditize Amazon. It was about how Amazon's entire infrastructure was really good for the last era of computing: what they do with their Elastic Network Adapter, ENA, and EFA, their NICs, and the whole protocol and everything behind them; what they do for custom CPUs, et cetera. It was really good for the last era of scale-out computing, and not this era of scale-up AI infrastructure.

It was also about how neoclouds were going to commoditize them, and how their silicon teams were focused on cost optimization, whereas the name of the game today is maximum performance per cost. Often that means you drive performance up like crazy, even if the cost doubles or triples, because then the cost per performance still falls. That's the name of the game today with NVIDIA's hardware.

It ended up being a really good call. Everyone was calling us out, saying, “No, you're wrong,” because Amazon was the best stock at the time, Microsoft really hadn't started taking off yet, and neither had Oracle or all these other companies. Since then, Amazon has been the worst-performing hyperscaler.

The call here is that they still have structural issues. They still use Elastic Fabric Adapter, although that's getting better. They're still behind NVIDIA's networking and behind Broadcom and Arista-type networking NICs. Their internal AI chip is okay, but the main thing is that they're now waking up and are able to actually capture business.

The main call here is that, since that report, AWS revenue growth has been decelerating consistently, and our big call is that it's actually going to start reaccelerating. That's because of Anthropic and because of all the work we do on data centers—tracking every single data center, when it goes online, and what's in there.

If you know how much the chips, networking, and power cost, and you know generally what the margins are for these things, then you can start estimating revenue. When we build all that up, it's very clear to us that AWS revenue growth troughs this quarter. This is the lowest AWS revenue growth will be on a year-over-year basis for at least the next year, and it's reaccelerating to north of 20% again.

That's because of all these massive data centers they have online with Trainium and GPUs. It depends on which one, depending on the customer. The experience isn't as good as, say, CoreWeave or whatever, but the name of the game is capacity today. CoreWeave can only deploy so much. They can only get so much data center capacity, and they're really fast at building.

The company with the most data center capacity in the world—and still today, although they may get passed up in the next 2 years—is actually Amazon. Based on what we see, they will get passed up, but incrementally, Amazon still has the most spare data center capacity that's going to ramp into AI revenue over the next year.

Erik Torenberg

Is that the right type of data center capacity? For the high-density AI buildouts today, you need massively more cooling. You need enough water close by, and you need enough power close by. Is it in the right place, or is it the wrong type?

Dylan Patel

Data center capacity, in this sense, means everything from power being secured, to substations being built, to transformers, to being able to provide power whips to the racks.

Historically, Amazon has had the highest-density data centers in the world. They went to 40-kilowatt racks when everyone else was still at 12. If you've ever stepped foot inside most data centers, they're pretty cool and dry-ish. If you step inside an Amazon data center, they feel like a swamp. It feels like where I grew up. It's humid and hot.

Because they're optimizing every percentage point, your point here is that Amazon's data centers aren't equipped for the new type of infrastructure. But when you compare that to the cost of the GPU, having a complex cooling arrangement is fine.

We made a call on Astera Labs a couple of months ago, when it was at 90, and it went to 250 the month after because of the orders Amazon is placing with them. There are certain things with Amazon's infrastructure—I won't get too much into it—but its rack infrastructure requires it to use a lot more Astera Labs connectivity products. The same applies to cooling. On the networking and cooling side, they just have to use a lot more of this stuff. But again, this stuff is inconsequential in cost compared to the GPU.

Erik Torenberg

You can build, right? My question was more like, look, I may need a major river close by for cooling at this point. In many areas, I just can't get enough water. It's probably power in the same region, too.

Dylan Patel

There are 2 gigawatt-scale sites where they have all the power secured. Wet chillers and dry chillers—all secured. Everything's fine. It's just not as efficient, but that's fine, right? They're going to ramp the revenue. They're going to add the revenue.

Not that I necessarily think Amazon's internal models are going to be great, or that their internal chip is better than NVIDIA's or competitive with TPUs, or that their hardware architecture is the best. I don't necessarily think that's the case. But they could build a lot of data centers and fill them up with stuff that will be rented out. It's a pretty simple thesis.

Erik Torenberg

How important has Anthropic been to the co-design for Trainium? I remember we had a portfolio company—this was summer 2023—and they invited them to AWS. They spent, I think, 8 hours with them over the course of a week trying to figure out Trainium back then. It was just impossible to work through. Obviously, that portfolio company hasn't gone back and tried it now, but how different is it based on what you're hearing?

Dylan Patel

Oh, it's still bad.

Erik Torenberg

Okay. Got it.

Dylan Patel

It's tough to use. This is sort of the argument that every inference company offers, including the AI hardware startups:

Erik Torenberg

Because I'm only running 3 different models at most, I can hand-optimize everything, write kernels for everything, and even go down to an assembly level, right? How hard can it be?

Dylan Patel

Yeah, it is pretty hard. But you tend to do this for production inference anyway. You aren't using cuDNN, which is NVIDIA's library that's super easy to use. You're still not using these ease-of-use libraries. When you're running inference, you're either using CUTLASS, stamping out your own PTX, or, in some cases, people are even going down to the SASS level.

When you look at, say, OpenAI or Anthropic, when they run inference on GPUs, they're doing this. The ecosystem isn't that amazing once you get all the way down to that level. It's not like using NVIDIA GPUs is easy. You have an intuitive understanding of the hardware architecture because you work on it so much, and everyone's worked on it and you can talk to other people, but at the end of the day, it's not easy.

Whereas, on Anthropic's Trainium or on TPUs, the hardware architecture is a little bit simpler than a GPU.

They have larger, simpler cores rather than all this functionality. They’re less general, so they’re a little easier to code on. There are tweets from Anthropic people saying that, when they’re doing that low-level work, they actually prefer working on Trainium and TPUs because of the simplicity.

Erik Torenberg

Interesting.

Dylan Patel

To be clear, TPUs—and Trainium especially—are very hard to use, not for the faint of heart. It’s very difficult, but you can do it if you’re just running—if I’m Anthropic and I must only run Claude 4.1 Opus or Sonnet, and screw it, I won’t even run Haiku. I’ll just run Haiku on GPUs or whatever, right? I’m just going to run 2 models.

Actually, screw it, I’m just going to run Opus on GPUs too and Sonnet on TPUs. Sonnet is the majority of my traffic anyway. I could spend the time. How often am I changing that architecture—every 4 or 6 months?

Erik Torenberg

Right. How much?

Dylan Patel

It’s not even changing that much, honestly, right?

Erik Torenberg

I think from 3 to 4 definitely did change, right?

Dylan Patel

Yeah. I mean, define architectural change. At a high level, the primitives are more or less the same across the last couple of generations.

Erik Torenberg

I don’t know enough about Anthropic’s model architecture, to be honest. But I think, from what I’ve seen at other places, there have been enough changes that it takes time to program this. The main thing is, if I’m Anthropic and I have, what, $7 billion ARR now, or whatever—north of $10 billion by the end of next year, north of $20 billion, right? ARR is maybe even $30 billion—and my margins are 50%, 70%, that’s $15 billion of Trainium that I need, right, that I can run Sonnet on.

Most of that is going to be Sonnet 3.5—or, sorry, 4.5, whatever it is, right? It’s going to be 1 model serving most of the use cases. So I could spend the time, and it’ll work on this hardware.

Yeah, totally. Maybe on the topic of non-consensus calls you’ve made, I’ll move to another cloud. In June, you guys said that Oracle is winning the AI compute market. In this pod, we’ve already referenced the big jump that Oracle had. I think it was the single largest gain that a company with over $500 billion in market cap has ever had.

Was it Nvidia in Q1 2023? Wasn’t that bigger? It might have been smaller. Okay. I think it was maybe close. We’ll fact-check ourselves. That’s amazing. But, obviously, this is the massive commitment that was announced. Can you walk us through why you made that call, and why Oracle is poised to do so well in such a competitive space?

Dylan Patel

Yeah, so Oracle has the largest balance sheet in the industry that isn’t dogmatic to any type of hardware. They’re not dogmatic to any type of networking. They’ll deploy Ethernet with Arista, Ethernet through their own white boxes, and NVIDIA networking—InfiniBand or Spectrum-X. They have really good network engineers. They have really great software across the board.

Again, ClusterMAX—they were ClusterMAX Gold because their software is great. There are a couple of things they needed to add that would take them higher, and they’re adding those to Platinum, which was where CoreWeave was.

When you couple 2 things, OpenAI has insane compute demand. Microsoft is quite pansy. They’re not willing to invest because they don’t believe OpenAI can actually pay the amount of money. I mentioned earlier the $300 billion deal: OpenAI doesn’t have $300 billion, and Oracle’s willing to take the bet.

Of course, there’s a bit more security in the bet in that Oracle really only needs to secure the data-center capacity. That’s how we came across the bet. We’ve been telling our institutional clients, especially in a super-detailed way—whether they be hyperscalers, AI labs, semiconductor companies, or investors—in our data-center model, because we’re tracking every single data center in the world.

Oracle doesn’t build their own data centers either, by the way. They get them from other companies and co-engineer them, but they don’t physically build them themselves. They’re quite nimble in terms of being able to assess and engineer new data centers. So we saw all these different data centers Oracle is snatching up, in deep discussions, signing, et cetera.

We have a gigawatt here, a gigawatt there, a gigawatt there. Abilene, 2 gigawatts. You have all these different sites that they’re signing up for and discussing, and we’re noting them. We have the timeline because we’re tracking the entire supply chain. We’re tracking all the permits and regulatory filings through language models, using satellite photos constantly, and then the supply chain for chillers, transformer equipment, generators, et cetera.

We’re able to make a pretty strong estimate, quarter by quarter, in our data-center model, of how much power there is for each of these sites. Some of these sites that we know of aren’t even ramping until 2027, but we know that Oracle signed them, and we have the ramp path.

Then it’s a question of, let’s say you have 1 megawatt, for simplicity’s sake, which is a ton of power, but now it doesn’t feel like much—we’re in the gigawatt era. If you’re talking about 1 megawatt, you fill it up with GPUs. How much do the GPUs for 1 megawatt cost?

Actually, it’s even simpler to do the math. If I’m talking about a GB200, each individual GPU is 1,000 watts, but when you talk about the whole system, it’s roughly 2,000 watts. All-in, for simplicity’s sake, it’s $50,000 per GPU. The GPU itself doesn’t cost them that; there are all the peripherals. So that’s $50,000 in capex for 2,000 watts, or $25,000 for 1,000 watts.

Then what’s the rental price for a GPU? If you’re on a really long-term deal, volume is $2.70—$2.60 in that range. Then you end up with, oh, it costs like $12 million per megawatt to rent a megawatt.

Erik Torenberg

Yeah.

Dylan Patel

Each chip is different, so we track each chip—what the capex is and what the networking is. We know what each chip is, so you can predict what chips they’re putting in which data centers, when those data centers go online, and how many megawatts by quarter.

Then you end up with, well, Stargate goes online in this time period. They’re going to start renting at this time. It’s this many chips at each Stargate site. Therefore, this is how much OpenAI would have to spend to rent it.

Then you price that out, and we were able to predict Oracle’s revenue with pretty high certainty. We matched pretty dead-on what they announced for 2025, 2026, and 2027, and we were pretty close on 2028. The surprise for us was that they announced some 2028 and 2029 data centers that we haven’t found yet, but we’ll find them, of course.

This methodology lets you see what data centers you’re getting, how much power, what they’re signing, and how much incremental revenue that is when it comes online. That’s the basis of our Oracle bet.

Obviously, in the newsletter we included a lot less detail, but it was that thesis: they have all this capacity, and they’re going to sign these deals. In our newsletter, we talked about 2 main things: the OpenAI business and the ByteDance business.

Presumably, tomorrow—on Friday—there’s going to be an announcement about TikTok and all this. But the ByteDance business involves huge amounts of data-center capacity that Oracle is also going to lease out to ByteDance. We did the same methodology there.

With ByteDance, it’s pretty certain they’ll pay because they’re a profitable company. With OpenAI, it’s not. There have to be some error bars as you go further out in terms of whether OpenAI will exist in 2028, 2029, or 2030, and whether they’ll be able to pay the $80-plus billion a year they’ve signed up to Oracle for. That’s the only risk here.

If that happens, Oracle’s downside is also somewhat protected because they only sign the data center, which is a minority of the cost. The GPUs are everything, and they purchase those 1–2 quarters before they start renting them. The downside risk is pretty low for them. If they don’t get the deal, they don’t get the revenue, but it’s not like they’re stuck with a bunch of assets they bought that are worthless.

Erik Torenberg

Yeah. Yeah. Is there another angle here? I mean, OpenAI and Microsoft were BFFs, and now they’ve filed divorce papers. They just want to diversify, and that’s pushing them toward other providers.

Dylan Patel

Yeah. Microsoft was the exclusive compute provider. It got reorged to a right of first refusal.

Erik Torenberg

Is it not your last choice or something like that?

Dylan Patel

No, it’s still a right of first refusal. Microsoft and Oracle—those 2 are not mutually exclusive.

Erik Torenberg

Well, if OpenAI is like, “We’re going to sign an $80 billion contract or a $300 billion contract for the next 5 years. Do you guys want it?” and they’re like—

Dylan Patel

“No, what? Okay, cool.” Right? OpenAI needs someone with a balance sheet to actually be able to pay for it. They’ll make tons of money off OpenAI on the margins on the compute and the infrastructure and all these things, but someone’s got to have a balance sheet, and OpenAI doesn’t have one. Oracle does.

Although, given the scale of what they signed, we also had another source of information: they were talking to the debt markets, right? Oracle actually just needs to raise debt to pay for this many GPUs over time. They won’t do it immediately; they can pay for everything this year and next year from their own cash, but in 2027, 2028, and 2029, they’ll start to have to use debt to pay for these GPUs. That’s what CoreWeave has done, and most of the neoclouds are debt-financed.

Even Meta went and got debt for its Louisiana megadatacenter. It’s literally better on a financial basis to do buybacks with your cash and get debt because the debt is cheaper than the return on your stock. It’s a financial engineering thing. Who’s out there, right? It could be Amazon, Google, Microsoft—a very short list.

Erik Torenberg

Or it could be Oracle or Meta, right? Meta’s obviously not. Microsoft has chickened out. Amazon, Google, and Oracle—that’s all that’s left.

Dylan Patel

Google would be an awkward fit. So—

Erik Torenberg

Yeah, Google would be an awkward fit. Amazon would be a fine fit, but—

Dylan Patel

Exactly. Right. It’s like—

Erik Torenberg

Yeah. Well, I guess maybe on the topic of these giant data center buildouts, you guys just released a piece on xAI and Colossus 2. Are you getting less impressed by these feats of building something this massive in 6 months, or is it still very impressive to you guys?

You know, this is the thing I’ve said about AI researchers: they’re the first class of humans to think about things on an order-of-magnitude scale, whereas people have always thought about things in terms of percentage growth. Ever since industrialization—and before that—it was just absolute numbers. Humanity is evolving in how we think because things are changing faster. Everything is an upscale.

Dylan Patel

And so it was really impressive when GPT-2 was trained on so many chips, and then GPT-4 was trained on 20,000 H100s. It was like, “Holy crap.” Then it was the era of 100,000-GPU clusters. We did some reports around 100,000-GPU clusters, but now there are 10 100,000-GPU clusters in the world. I was like, “Okay, this is kind of boring.”

But 100,000 GPUs is over 100 megawatts. Now, literally, in our Slack and some of these channels, it’s like, “Oh, we found another 200-megawatt data center.” There’s someone who puts the yawning emoji every time, and I’m like, “Dude, what?” Now it’s only exciting if you do gigawatt scale.

Erik Torenberg

Gigawatt era. Yeah.

Dylan Patel

Yeah. And I’m sure maybe we’ll start yawning at that, too. But the log scale of this is like—

Erik Torenberg

The capital numbers are crazy, right? It was crazy enough that OpenAI did a $100 million training run. Then they did a $1 billion training run; now we’re talking about $10 billion training runs. It’s crazy that we think in log scale, but yes, things are only impressive—

Dylan Patel

Yeah, when they do it like what Elon’s doing. What Elon’s doing in Tennessee, in Memphis, the first time was crazy, right? 100,000 GPUs in 6 months. He bought a factory in February 2024 and had models training within 6 months.

He did liquid cooling—the first large-scale data center at this scale for AI using liquid cooling—all these crazy firsts. He put generators outside, CAT turbines, and all these things to get the power; mobile substations; all these different crazy things. He tapped the natural gas line running alongside the factory. All of these—

Erik Torenberg

You know, 200, 300 megawatts, right? Now he’s doing it at a gigawatt scale, and he’s doing it just as fast. You would think this is obviously way more impressive that he did it again.

Dylan Patel

Yeah.

Erik Torenberg

But—

Dylan Patel

Like—

Erik Torenberg

Maybe I’m desensitized, but it’s like you’ve given the child too much candy, right?

Dylan Patel

Exactly.

Erik Torenberg

And now the child has no—it’s like he doesn’t like apples, right? I don’t know.

Dylan Patel

So, yeah, a gigawatt data center. There were all these protests around his Memphis facility. People were saying, “Oh, you’re destroying the air.” And it’s like, “Have you looked around that area of Memphis?” There’s a gigawatt gas-turbine plant that’s just powering that area generally. There’s a sewage plant servicing the entire city of Memphis. There are open-air pits—there’s open-air mining. There’s all sorts of disgusting shit around there, which is needed. We need that stuff for a country to run, to be clear. People were complaining about a couple hundred megawatts of power—

Erik Torenberg

Of generation. So he got protests from all sorts of people. He got super into the political side of things, and the NAACP even protested him. He really got some local municipalities to be like, “Oh, I don’t like this.”

Dylan Patel

And so he couldn’t do as much as he wanted to in Memphis. But he still needed the data center to be close because he wanted to connect these data centers with super-high bandwidth, super close. He already had a lot of infrastructure set up there, so he bought another distribution center at this time. It’s still in Memphis, but the cool thing about Memphis is that it’s right across the border from Mississippi. Right. So now—

Erik Torenberg

You know, it’s 10 miles away from his original one, but his facility is a mile away from Mississippi, and he bought a power plant in Mississippi. He’s putting turbines there because the regulation is completely different. If the question is really to galvanize resources and build it really fast, maybe Elon is ahead of everyone. He hasn’t made the best model yet, or he doesn’t have the best model, at least today. You could argue Grok 4 was the best for a little period of time, but it’s truly amazing how fast he’s able to build these things.

From first principles, most people are like, “Shit, we can’t build the power. We can’t do power here anymore. I guess we have to find a new site.” It’s like, “No, just go across the border.”

Dylan Patel

Go to Mississippi.

Erik Torenberg

My favorite thing is that Arkansas is right there, so Mississippi gets mad. I don’t know—the regulation—all future data centers built in places where multiple states meet, is that the—

Dylan Patel

Four Corners, yeah.

Erik Torenberg

The optimal regulation, I think. There’s one. There we go. Is there a point in the U.S. with 5? I know there’s a point with 4—4 states intersect. Yeah. Maybe that’s going to be a data center kind of concern. All right.

Dylan Patel

I’m going to buy real estate in that area, right?

Erik Torenberg

Well, I guess on the topic of just maybe new hardware, you had this piece analyzing TCO for GB200s. I’m kind of going to ask this question on behalf of our portfolio companies, which it sounds like you’re helping them already. One of the findings that I thought was really interesting was that TCO was sort of 1.66x H100s for GB200s. Obviously, there’s this point where that’s the benchmark for the performance boost that you’re going to need to at least make the performance-cost ratio benefit from switching over.

Maybe just talk about what you’ve seen from a performance standpoint, and what do you recommend to portfolio companies, maybe on a smaller scale than xAI, who are thinking about new hardware? Try to get it—there are capacity constraints, obviously.

Dylan Patel

Yeah, that’s a challenge, right? With each generation of GPU, it gets so much faster that you want the new one. In some metrics, you could say GB200 is 3 times faster—or 2 times faster—than the prior generation. In other metrics, you can say it’s way more than that. If you’re doing pretraining versus inference, right—

Erik Torenberg

You can run everything for a bit, right?

Dylan Patel

Yeah, if you can run it for a bit, or just inference, and take advantage of the huge NVLink—NVL72—there are ways you can squint and say GB200 is only 2 times faster than H100. In which case, with 1.66x TCO, it’s—

Erik Torenberg

You know, it’s worthwhile, right? It’s worth going to the next generation—

Dylan Patel

But more marginal.

Speaker 1

It’s more marginal. It’s not a big deal. Then there are other cases where, if you’re running DeepSeek inference, the performance difference per GPU is north of 6–7x, and it continues to optimize for DeepSeek inference. Then it’s like, “I’m only paying 60% more for 6x,” so it’s a 4x or 3x performance-per-dollar gain. Absolutely, right? If you’re running inference of DeepSeek, that can also include RL, right?

Then there’s the question of the GPU being new. There’s also B200, GB200, and B2000. B200 is much simpler from a hardware perspective; it’s just 8 GPUs in a box. So it’s not as much of a performance gain, especially in inference, but you have all this stability. It’s an 8-GPU box; it’s not going to be unreliable. The GB200s are still having reliability challenges, though those are being worked through and getting better by the day. When you have an H100 or H200 8-GPU box and one fails, you take the entire server offline and fix it. If it’s GB200 and one GPU fails, what do you do with 72 GPUs? Do you break the whole thing and get a new 72? The blast radius of a failure is huge. GPU failure rates are at best the same and likely worse generation on generation because everything is getting hotter and faster. A lot of people run a high-priority workload on 64 and use the other 8 for low-priority workloads. When a high-priority workload has a failure, instead of taking the whole rack offline, you take some GPUs from the low-priority workload and put them in the high-priority one, then let the dead GPU sit there until you service the rack later. There are all these complicated infrastructure challenges.

Dylan Patel

That 3x or 2x performance increase in pre-training is lower because the downtime is higher, I’m not using all the GPUs all the time, and I’m not able—or I don’t have the infrastructure—to manage low- and high-priority workloads. It’s not impossible; the labs are doing it. It’s just—

Speaker 3

I mean, if I’m running a cloud, it’s actually really hard, because I probably have to rent the spares out as spot instances or something.

Speaker 2

No, no, no, no. Because it’s a coherent domain; it’s NVLink. You don’t want anyone touching that. So the end customer has to leave them as empty spares. That’s even worse.

Speaker 3

The end customer usually will just be like, “I want them, and I’ll use them.” The SLAs and pricing account for that, right?

Speaker 2

Generally, when you have a cloud, you have an SLA. It says uptime is going to be 99%, blah, blah, blah, for this period. With GB200, it’s 99% for 64 GPUs, not 72, and then it’s 95% for all 72. It differs across every cloud; every cloud has a different SLA. But they’ve adjusted for this because they’re like, “Look, this hardware is just finicky. Do you still want it? We will credit you, in that 64 of them will always work.”

Speaker 3

Right, not 72.

Speaker 2

So the end customer has to be capable of dealing with the unreliability.

Speaker 3

The whole reason you want this 72-GPU domain is so you can have some of these gains, right?

Speaker 2

But you have to be smart enough to be able to do it, and that’s challenging for small companies.

Speaker 3

Totally. So NVIDIA just announced the Rubin prefill cards, like CPX—

Speaker 2

CPX. CPX.

Speaker 3

There we go. What’s your take on that? Does it cannibalize?

Speaker 2

Dude, by the way, I don’t know if this is brain rot or what, but I can’t remember what I had for lunch yesterday, yet I know the model number of every fucking chip.

Speaker 3

In your dreams.

Speaker 1

We’re broken. We’re broken. Living the dream.

Speaker 2

No, no, no. Why do you pre-announce a product that’s 5x faster for certain use cases?

Speaker 3

Is that that much?

Speaker 2

I think, historically, AI chips were AI chips. Then we started getting a lot of people saying, “This is a training chip; this is an inference chip.” Actually, training and inference are switching so fast in terms of what they require that now it’s still like one chip.

Speaker 3

Actually, there are still workload-level dynamics that differ.

Speaker 2

The main workload is inference, even in training, because of RL. Most of that is generating stuff in an environment and trying to achieve a reward, so it’s inference still. Training is now becoming mostly dominated by inference as well.

Inference has 2 main operations. There’s calculating the KV cache for prefill: here are all these documents; do the attention between all of them, between all the tokens, however—whatever type of attention you use. Then there’s decode, which is to autoregressively generate each token.

These are very different workloads. Initially, the infrastructure techniques—the ML systems techniques—were, “Okay, I’ll just make the batch size for every single forward pass this big. I’ll make it, let’s call it, 1,000 big, and maybe I’ll run 32 users concurrently. That way, I still have 900-something left—960 left.”

That 960 is actually doing the prefill. If a request comes in, it chunks it. It’s called chunked prefill. You prefill chunks of it, and now you get really good utilization on GPUs. But that ends up impacting the decode workers. The people who are autoregressively generating each token end up having slower TPS, and tokens per second is really important for user experience and all these other things.

Speaker 3

Everyone, everyone, everyone—

Speaker 2

Together AI, Fireworks, all these guys do prefill-decode disaggregated. They run prefill on a set of GPUs and decode on a certain set of GPUs.

Why is this beneficial? Because you can autoscale them. All of a sudden—or not all of a sudden, but over time—if my traffic mix is not long input and short output, but short input and long output, I can have more decode workers. This way, I can guarantee that my prefill time is at a certain level.

What’s really important in search is how fast you get the page to start loading, not when the response is finished. What do people do in games? The loading screen often has some sort of interactive environment, or it blends in over time, or it has tips and tricks—ways to distract you. The same thing is true here. There are studies and papers out there showing that users prefer a faster time to first token, with the first token streamed to them sooner, even if the total time to get all their tokens is a little bit longer.

Guido Appenzeller

I can’t read that fast anyway, right?

Dylan Patel

I mean, I like to give—I like to give—

Guido Appenzeller

Yeah. Most models return above speed-reading speed.

Dylan Patel

But you need that, right? I think the idea is that you want to guarantee time to first token at a certain level for user-experience reasons. Otherwise, people say, “Screw this, I’m not using AI.” Decode speed matters a lot, too, but not as much as time to first token.

By having separate prefill and decode, you can do this. But now you’ve already done this—and this is all in the same infrastructure. What’s the next logical step? These workloads are so different.

For decode, you have to load all the parameters and the KV cache to generate a single token. You batch a couple of users together, but very quickly you run out of memory capacity or memory bandwidth because everyone’s KV cache is different.

Guido Appenzeller

Yeah. The attention of all the tokens, right? Whereas on prefill, I could even just serve 1 or 2 users at a time. If they send me a 64,000-context request, that is a lot of FLOPs, right?

I’ll use Llama 70B because it’s simple to do the math on—70 billion parameters. That’s 140 gigaflops per token. Seventy times 64,000—that’s many, many petaflops. You can use the entire GPU for about a second, potentially, depending on the GPU, just to do the prefill. And that’s just 1 forward pass.

Dylan Patel

So I don't necessarily care about loading all the tokens or all the parameters into the KV cache quickly. All I care about is the FLOPs. That leads us to CPX, but I had to give this long-winded explanation because it's hard for people to understand what CPX is. I've had a lot of—even my own clients—we sent multiple notes explaining it, and they're like, “I still don't understand.” I'm like, “[expletive], okay.” Send them the Attention Is All You Need paper, and—

You can't expect—I mean, think about a networking person. They're like, “I don't know. I don't need to know about this. You know, Attention Is All You Need, right?” Or think about an investor. There are all these data center operators, and they're like, “Oh, there are 2 chips. Why should I build my data center differently?” It's like, “I have to explain everything.” Or just like, “No, you don't have to build differently.”

At Stanford, at least 25% of all students—not CS students, all students—read the paper Attention Is All You Need. That's low. The majors? And you don't like the philosophy? I find this amazing, anyway. Sorry.

Erik Torenberg

The Middle East—I can't remember what country it is—has AI education starting at around 8, and in high school they have to read Attention Is All You Need.

Dylan Patel

Wow.

Erik Torenberg

Someone told me that their son had to read Attention Is All You Need.

Dylan Patel

Which is—I don't know. Look, top-down mandates for education: maybe they work, maybe they don't. Maybe people like homeschooling their kids. I don't know. I went to public school, but—

Erik Torenberg

Back to your question.

Dylan Patel

Yeah. Just on the topic of hardware cycles, I wanted to maybe—I actually explained what CPX is. CPX is a very compute-optimized chip for prefill, whereas decode, statistically speaking, is like the rest: the normal chips with HBM. HBM is more than half the cost of the GPU.

If you strip that out, you end up with a much cheaper chip that's passed on to the customer. Or, if NVIDIA takes the same margin, the cost of this prefill chip is much, much lower. Now the whole process is way cheaper and more efficient, and long context can be adopted.

Erik Torenberg

All right.

Sarah Wang

Yeah. I love that we're actually going into all this detail because I had a more 10,000-foot-view question for you. I haven't been following the semiconductor market as closely as you have. I probably started with the A100. I remember helping Noam at Character.AI in June 2023 chase down GPUs, and the only thing that mattered at that time was the delivery date because there was a huge capacity crunch.

Then to see that evolve over the last 2 years—let's say 6 to 12 months ago, people were doing these RFPs to 20 neoclouds, right? To some degree, the only thing that mattered was price.

Dylan Patel

People actually do RFPs for GPUs?

Sarah Wang

Yes.

Dylan Patel

So, just to be clear, my opinion on how you buy GPUs is that it's like buying cocaine—or any other drug. This was described to me, not by me. I don't buy cocaine.

Erik Torenberg

Someone tells me this. Someone tells me this. I'm like, “Holy [expletive], it's right.” You call up a couple of people, you text a couple of people, and you ask, “Yo, how much do you have? What's the price?”

Dylan Patel

It's like—

Erik Torenberg

Exactly. This is [expletive] like buying drugs. Oh, sorry. Sorry.

Dylan Patel

No, I mean, to this day, it's the same way. You just send—we have Slack Connects with around 30 neoclouds, as well as some of the major ones, and we just send them a message: “Hey, the customer wants this much. This is what they're looking for.” Then they send quotes.

Erik Torenberg

I know this guy.

Dylan Patel

I know a guy.

Sarah Wang

Well, I think that's actually a very accurate description, and I've sent countless people your original ClusterMAX post because I thought it did a really good job breaking them down. Maybe one question to end on for me is: What era are we in now, with Blackwells coming online? Are we sort of back to the summer of 2023 era, where GPUs were tight, or what is your view on where we are?

Dylan Patel

Very good question. For one of your portfolio companies, after their difficulties with Amazon, we tried to actually get you GPUs. The original deals we got you were gone, but here were some other deals. It turned out that multiple major neoclouds had sold out of Hopper capacity, and their Blackwell capacity comes online in a few months.

So it's a bit of a challenge, right?

Sarah Wang

Due to inference?

Dylan Patel

Inference demand has been skyrocketing this year.

Sarah Wang

Reasoning models, yeah.

Dylan Patel

These reasoning models—the revenue has been skyrocketing this year. Also, Blackwell comes online, but it's hard to deploy, so there's a learning curve to deploying it. You could buy Hopper, install it in the data center, and have it running within a month or 2. For Blackwell, it's a longer time frame because of reliability challenges. It's a new GPU; it's just learning curve and growing pains.

There was this gap in how many GPUs were coming onto the market as revenue started to inflect, so a lot of capacity got sucked up. Actually, prices for Hopper bottomed 3 or 4 months ago—or 5 or 6 months ago.

Sarah Wang

Yeah.

Dylan Patel

They've actually crept up a little bit now. They're still not too bad. I don't think we're quite back to the 2023–2024 era of GPUs being tight, but certainly, if you want just a few GPUs, it's easy. If you want a lot, it's hard.

Sarah Wang

You can't get capacity instantly.

Dylan Patel

Yeah.

Erik Torenberg

Wow. What a time. Shall we wrap on that? Dylan, this was another instant classic. Thank you so much for coming to the podcast.

Dylan Patel

It was like 2 hours, bro. What? I missed—thank you. We couldn't stop. Thanks so much. This is great. Thank you so much for having me.

Dylan Patel 谈 AI 芯片竞赛——NVIDIA、Intel 与美国政府对决中国 — 文字稿与摘要 | BidClub