Dylan Patel 谈 GPT-5 的路由器时刻、GPU 与 TPU、商业化
Dylan Patel × Erin Price-Wright × Guido Appenzeller × Erik Torenberg
GPT-5 与其说是前沿算力的跃升,不如说是围绕路由器打造的“经济释放”。 Dylan Patel 认为,重度用户失去了 GPT-4.5 和 o3 的使用权限——o3 平均思考约30秒,而 GPT-5 只需5-10秒——但免费用户有时获得了此前从未享受过的推理能力。路由器让 OpenAI 可以在普通、mini 和 thinking 模型之间选择,并在服务受限时“优雅降级”,以最高能力换取显著更大的 token 容量。
路由器更大的价值,在于让推理支出匹配每次查询可变现的价值。 “为什么天空是蓝色的?”可以交给 mini;而寻找最好的酒驾律师、航班或商品,则可以获得“大得离谱的算力”,因为 OpenAI 可能直接完成交易并抽成。考虑到 ChatGPT 已占 Etsy 流量约10%,Dylan 认为,代理式商业才是免费用户的变现路径,而不是让结果变差的广告。
统一费率的 AI 订阅,正撞上消费者使用量相差20倍的现实。 Anthropic 和编程产品在重度用户利用“负毛利”套餐后收紧了速率限制;据称一名开发者把睡眠切成水手式的短暂小睡,Reddit 使用排行榜上还有人的订阅消耗价值约每月3万美元。企业或许能通过承诺用量或平均费率来消化差异,但消费者越来越指向按使用量收费——即便产品仍会用订阅制和更好的审查界面来制造用户黏性。
AI 可能已经创造了超过基础设施成本的价值,但实验室只能捕获其中一小部分。 Dylan Patel 的编程思想实验是:3000万名开发者、生产率翻倍、每人创造10万美元价值,仅一个用例就对应3万亿美元的潜在 GDP 价值;而 Dylan 认为,OpenAI 捕获的 ChatGPT 已创造价值“甚至不到10%”。这种错配未必会阻止资本开支继续增长:超大规模云厂商的支出还可能再增20-30%,CoreWeave、Oracle、基础设施基金和主权资本则会补充更多短期经济意义较弱的产能。
只有当 AI 需求继续集中在少数巨型买家手中时,定制芯片才是 Nvidia 最大的威胁。 Google 正在生产数百万颗高利用率 TPU,Amazon 生产数百万颗 Trainium,Meta 也在大幅增加内部芯片订单;Dylan 认为 Google 应该直接销售 TPU,而不只是出租算力。但如果开源模型和廉价部署让需求分散,Nvidia 的通用生态反而会更强——独立挑战者必须做到“好上5倍”,供应链、软件和毛利率劣势才不会抹平领先优势。
美国 AI 部署受限的主要不是电价,而是已通电场址、电网设备和建设速度。 Dylan 估算,一座 Blackwell 数据中心约80%的成本来自 GPU、网络、建筑和电力转换资本,土地、电力、制冷、备用电源及相关项目只占20%。因此,为提前3个月上线而支付溢价是理性的;与此同时,中国当前受制于资本和芯片质量,而不是电力,尽管其发电扩张速度更快。
Intel 需要立即进行运营手术并获得资本支持,几家平台型巨头也需要重新建立产品紧迫感。 Dylan 表示,Intel 长达5-6年的设计周期以及最多14次芯片修订,必须压缩至2-3年和1-3次;如果没有大规模现金注入或大幅削减成本,公司可能在正式分拆完成前“真的”破产。他还提出几项更广泛的建议:Nvidia 应将预计超过1000亿美元的现金储备用于基础设施再投资,Google 应开放 TPU,Apple 或许应投入500亿美元建设 AI 基础设施;Erik 则认为,尽管 Microsoft 起点极其优越,仍必须“把公司彻底摇醒”。
1. GPT-5 真正的突破,是掌控推理经济学
Dylan 的失望取决于用户层级:付费重度用户失去了 GPT-4.5,而他仍认为 GPT-4.5 在部分工作上是更好的预训练模型;他们也失去了 o3,后者平均会思考约30秒。GPT-5 thinking 通常只运行5-10秒,因此分配给他普通查询的算力更少。
这并不意味着没有进步。GPT-5 的模型规模大致相同,却显著优于上一代 vanilla 模型,同时避免了 o3 那种病理式推理——比如花48秒判断猪肉究竟是红肉还是白肉。Anthropic 早已证明,同等甚至更好的答案不必耗费那么长的思考时间。
路由器现在可以在普通模型、触发速率限制后切换的 mini,以及 thinking 之间做选择,并控制推理持续多久。免费用户偶尔会得到强得多的答案;当容量收紧时,OpenAI 也能“对他们进行优雅降级”。Guido 抛出一个 meme,并明确说不是真的:OpenAI 是否把 o3 和更小模型放在路由器后面,以更低的混合价格提供服务;Dylan 的回答是,“确实有一点像。”
2. 查询价值可以决定 OpenAI 要投入多少智能
Dylan 的商业判断是:“路由器指向了 OpenAI 的未来。”传统广告与有帮助的助手存在冲突——植入推广会恶化答案,横幅广告也不适配;因此,免费用户需要一条原生于交易的变现路径。
他的对比把分配逻辑说得很具体。“为什么天空是蓝色的?”值得交给 mini;而寻找最好的酒驾律师,可能值得搜索法院文件、联系本地律所、比较结果,并投入“大得离谱的算力”,因为 OpenAI 可以从一笔高价值交易中抽成。
购物和旅行是最明显的切入口:Dylan 提到,来自 ChatGPT 的流量已经占 Etsy 流量的10%,但 OpenAI 一分钱也赚不到,部分原因是 Amazon 屏蔽了 ChatGPT。他的建议是接入信用卡、日历,以及靠过道还是靠窗等偏好,然后让代理完成预订,并收取约定比例的交易佣金。
3. AI 定价暴露了订阅制隐藏的使用极端值
OpenAI 表示,它为大量用户将速率限制翻倍,并大幅增加了服务 token 数量,让 GPT-5 成为一次“经济释放”。这意味着混合推理成本更低,而不只是 GPT-5 在 MMLU 或其他智能基准上成为新的赢家。
重度编码用户暴露了统一费率套餐的脆弱性。Anthropic 同时施加按小时和按周的限制后,据称有用户像独自航海的水手一样把睡眠切碎;Reddit 排行榜上则出现了一个通过订阅每月消耗约3万美元的用户。“人们正在利用负毛利。”
消费者的使用量可能相差约20倍,推动模型厂商转向按使用量收费;企业可以对全职开发者较为可预测的使用行为取平均,也可能支付大额承诺费用,以避免账单无上限。一项产品黏性的来源是验证闭环:在可视化变更文件、后果和图表,以及提供快速反馈和复杂反馈方面,做得比竞争对手更好。
4. Nvidia 的增长跑道,建立在需求加速和充足的投机资本之上
Dylan 将当前芯片需求大致分成三等份。OpenAI 和 Anthropic 单独就占约30%:Anthropic 的算力来自 Google 和 Amazon,OpenAI 的算力来自 Microsoft、CoreWeave 和 Oracle。Meta 和 ByteDance 等广告买家占另一等份;剩下那些经济价值不那么清晰的服务商,可能难以持续融资越来越大的轮次。
主持人仅从编程场景挑战需求上限:即使是传统的 GitHub Copilot 推广,也可能带来15%的生产率提升,而 Dylan 坚持认为更好的产品可以做到远不止这些。假设3000万名开发者、生产率翻倍、每人创造10万美元价值,理论规模就达到3万亿美元,还没有计入其他 AI 应用。
Dylan 对“6000亿美元问题”的修正是:现在购买的基础设施,对应的不是一年的平坦收入,而可能是未来5年不断增长的收入。他认为 AI 已经创造了超过支出的价值,但“价值捕获出了问题”:他的4人团队用廉价的 Gemini API 分析许可证、监管文件、卫星图像、发电机、冷却塔、变电站和施工进度,随后捕获的下游价值远高于模型提供商。
经济上合理的资本开支存在上限,但可用资金并不存在同样的限制。超大规模云厂商明年的资本开支还可能增长20-30%;CoreWeave 和 Oracle 可以通过资本市场筹集更多资金;Brookfield、Blackstone、G42、GIC 以及其他主权或基础设施资金池“几乎还没有真正开始接触 AI”。
5. 市场集中度决定 TPU 还是 Nvidia 胜出
定制芯片是 Nvidia 面临的核心威胁。Guido 表示,Google 和 Amazon 正在生产数百万颗芯片,Meta 正在大幅增加订单;Dylan 说 Google 的 TPU 利用率基本达到100%,而 Amazon 的 Trainium 利用率仍然落后。Guido 形容 Microsoft 的定制芯片“有点烂”;Dylan 后来称,Microsoft 的内部芯片项目是所有超大规模云厂商中最差的。
决定性变量是需求集中度。少数巨型 AI 买家可以摊薄自研芯片成本并压缩 Nvidia 的利润;由中国开源模型和廉价推理库驱动的分散需求,则有利于 Nvidia 获得广泛支持的平台。因此,Dylan 认为 Nvidia 可能会长期保持全球市值最高公司。
Dylan 问 Google 是否应该开始向所有人销售芯片,Guido 认同 Google 应对外销售 TPU,而不只是出租 TPU 算力。这样做需要 Cloud、TPU、JAX 和 XLA 之间进行文化重组。Dylan 认为,更大的 TPU 业务可以支撑更高的 Google 估值;但 Guido 指出,Google 管理层仍会认为 Gemini 最终的价值要高得多。
6. 初创公司必须领先 Nvidia 5倍,现实才会把差距压回去
在公开芯片问世之前,资本已经在为 Etched、Rivos、MatX 等公司提供资金,同时还有 Groq、Cerebras、SambaNova、Tenstorrent 和 SoftBank 旗下 Graphcore 等成熟挑战者。超大规模云厂商拥有它们都没有的优势:一个愿意把专用芯片当作压缩利润手段来接受的 captive customer。
独立厂商必须同时设计芯片和软件,组装 IP,管理芯片、机架、网络、内存和客户,最后还要赚取利润。AMD 说明了其中的难度:尽管工程实力很强,但在相近性能下需要更多芯片面积和内存,毛利率接近50%,而 Nvidia 约为75%。
当研究方向发生变化时,硬件与模型的协同设计会变成陷阱。Cerebras、Groq 和 SambaNova 曾针对当时领先的工作负载,强调更多片上 SRAM、减少 DRAM;但更大模型和视觉 Transformer 改变了经济性。新一代初创公司为密集 Transformer 优化巨型脉动阵列,却遇上了类似 DeepSeek 的工作负载:形状更小、矩阵乘法更多。
Nvidia 拥有更好的网络、HBM、制程获取能力和爬坡速度,并在 TSMC、SK hynix、机架供应商和线缆厂商面前拥有更强的议价能力。挑战者的“5倍”架构优势,可能在供应链惩罚后变成2.5倍,再在 Nvidia 压缩利润、部署软件防御后变成约50%。与此同时,“你必须沿着科技树继续升级”,却不知道模型最终会向哪个方向分叉。
7. 中国有电力,但资本效率和芯片获取仍是约束
Dylan 提到,中国各省有规定称 H20 的效率不足,尽管他认为 H20 是中国现有最好的 AI 芯片,并表示 Huawei 仍然落后。在电力受限的美国,即便 H20 免费也可能缺乏吸引力:用更弱的芯片消耗稀缺兆瓦,意味着更低的算力容量。
Nvidia 推动出口的理由是控制生态。中国开发者贡献了有价值的 Nvidia 兼容软件,包括 Triton 扩展,因此销售 GPU 可以阻止 Huawei 建立竞争性软件栈。Dylan 认为,模型为社会创造的经济价值高于硬件;供应 H20 和削弱版 Blackwell,因而可能转移更多经济能力,而不是只捕获芯片销售收入。
中国当前的直接瓶颈“始终是资本”,而不是电力。中国 AI 资本开支的百分比增速高于美国,但绝对金额更低,单位美元产出也更差;中国每年已经为半导体补贴约1500亿-2000亿美元,如果愿意,完全可以资助一个 Meta 或 Google 规模的国家级项目。
8. 已通电的土地,而不是廉价制冷,才是稀缺的 AI 资产
中国企业通过在海外租用更先进的 GPU,或借助与新加坡有关联的实体建设,绕开国内限制。ByteDance 是 Google Cloud 最大客户之一,也从 Oracle 和 Microsoft 租用算力,因为海外 Blackwell 产能在单位产出成本上可能优于自建国内基础设施。
美国的物理瓶颈导致已购买的芯片闲置。Google 有等待已通电设施的 TPU,Meta 等公司也有同样状态的 GPU;芯片本身约占集群成本的60-80%。电网互联、输电、变电站、电气承包商和外地电工,都已成为决定工期的关键投入。
CoreWeave 的价值很大一部分在于它愿意快速行动:改造加密货币场址、维护裸金属集群并购买已通电资产。Google 为获取电力持有 TeraWulf 8%的股份;随着速度压倒此前承诺,超大规模云厂商实际上已经对自身的可持续发展承诺说了“管它呢”。
Guido 表示,制冷是次要问题:如今苜蓿的用水量约为 AI 数据中心的100倍,而水下设施节省的成本可能只有5-10%,却会变得无法维修。对于 Blackwell 设施,约80%的成本是资本开支,20%用于土地、电力、制冷、备用系统和发电机;Elon 购买昂贵的临时发电机和冷却机是理性的,因为这让数据中心提前3个月上线。
9. Intel 能否存活,比是否优雅分拆更重要
Guido 表示,世界需要 Intel,因为 Samsung 在2纳米级制程开发上似乎落后更多,而 TSMC 实际上垄断了先进制程。尽管 TSMC 拥有定价权,明年也只提价3-10%;如果台湾发生意外,Intel 可能拥有全球最先进的制程技术,只是经济性很差。
Intel 的晶圆厂和设计业务最终应当分拆,但执行分拆可能耗尽公司本就没有的管理时间。眼下最紧迫的问题在运营层面:从设计到上市需要5-6年,有时要进行14次 tape-out 修订;强劲竞争对手则大约需要3年,修订1-3次。
Dylan 对 CEO Lip-Bu Tan 的建议是:让两种文化分别运营,削减层级、清除能力不足的管理者,留下引领制程技术20年的工程师,并将设计到上市的周期压缩至2-3年。即使没有 AI 领先式增长,x86 和 PC 业务仍可以保持高利润,员工数量甚至可能只需要现在的三分之一或一半。
晶圆厂要推进后续几代制程,需要更多资本。如果没有大规模注资或激进削减成本,Dylan 警告 Intel 可能在重组完成前“真的”破产;Erik 所期待的兜底方案是,每家主要超大规模云厂商先贡献约50亿美元,而 TSMC 的毛利率未来可能接近75%。
10. Nvidia、Google 和 Meta 应把算力所有权转化为分发能力
Dylan 会建议 Jensen Huang 把 Nvidia 的现金储备用于基础设施层。按照新的 Trump 税法,GPU 集群第一年的折旧对客户有重大税务影响;Nvidia 自己到年底可能持有超过1000亿美元现金,若用于回购和分红,就会浪费加速已通电产能、控制更多产业栈的机会。
Dylan 认为 Google 应进一步开放 TPU 和 XLA 周围的生态,销售硬件,加快数据中心建设,并夺回曾经的算力领先地位。Sergey Brin 正与 DeepMind 密切合作,但物理基础设施和产品交付仍然过慢,采购代理商正在威胁要绕过 Google,截走可变现的搜索查询。
Mark Zuckerberg 已经意识到紧迫性:搭建临时“帐篷”,在收购 Thinking Machines Lab 或 SSI、出价300亿美元未果后积极招聘,并将超级智能与可穿戴设备及助手连接起来。Dylan 的批评在于 Meta 围墙之外的执行力:产品经常“有点平庸”,因此应该更快推出明确对标 ChatGPT 和 Claude Code 的竞争产品。
11. Apple 和 Microsoft 可能失去界面,Elon 可能失去焦点
Apple 的硬件和形态设计仍然很强,但 Dylan 认为,如果不投入或许500亿美元建设 AI 基础设施,公司可能“错失这艘船”。随着代理接入日历、消息、偏好和交易,AI 将成为计算界面,削弱 Apple 通过触控、键盘和封闭生态控制用户体验的能力。
Microsoft 在2023年和2024年行动激进,随后却收缩数据中心投入,OpenAI 也开始从其掌控中滑走。Dylan 称其内部芯片项目是超大规模云厂商中最差的,MAI 已经失败,Azure 则容易受到 Oracle、CoreWeave 和 Google 的冲击。Erik 表示,尽管 Microsoft 起步时拥有 GitHub、最好的源代码仓库、企业分发能力、先发优势,以及与一家模型公司的牢固关系,GitHub Copilot 和 Microsoft Copilot 仍然很弱。
Dylan 对 Elon Musk 的判断仍然保留条件:色情模型可能加速 xAI 收入,robotaxi“开始看起来不错”,而 Musk 仍然是吸引顶尖建设者的磁石。但人才流失、项目被砍和仓促决策,如今正与其历史性的上行空间一起造成损害;最后的建议很简单:“重新专注于产品。”
Nvidia is going to have better networking than you. They’re going to have better HBM, a better process node, and they’re going to come to market faster. They’re going to be able to ramp faster and have better negotiations with whether it’s TSMC or SK hynix on the memory and silicon side, or with all the rack people, copper cables—everything.
They’re going to have better cost efficiency. You can’t just do the same thing as Nvidia. You have to really leap forward in some other way. You have to be 5× better.
Dylan, welcome to the podcast.
Thank you for having me.
We’ve been trying to get you for a while. You’re a busy man, but it worked out. Guido, do you want to introduce why we’re so excited to have Dylan on the podcast and what we’re excited to discuss?
Look, I think Dylan, you’ve done an exceptional job covering what’s happening in the AI hardware space, the AI semiconductor space, and now more and more in the data center space as well. Currently, the most valuable company on the planet is an AI semiconductor company, right? The biggest IPO so far in AI was an AI cloud company.
This is currently where it’s happening. In any gold rush, in the early days, it’s the picks and shovels that make money, and I think this is the stage that we’re in. So, we’re super excited to have you here today.
Awesome. Thank you. Happy to talk about my favorite topics.
Amazing. Well, maybe let’s start with GPT-5. We just had some of the researchers from OpenAI—Christina and Isabella—on here last week. You said it was disappointing. Why don’t you share your reactions, what capabilities you were hoping to see, or your overall reaction?
I think it depends on what tier of user you are, right? If you’re just using GPT-5 and before you were a $20 or $200-a-month subscriber, you no longer have access to GPT-4.5, which, in my opinion, is still a better pretrained model for certain things. Or you no longer have access to o3, which would think for 30 seconds on average, maybe, right?
Whereas GPT-5, even when you’re using thinking, only thinks for 5 to 10 seconds on average, which is an interesting phenomenon. Basically, GPT-5 is not spending more compute, per se. The model did get a little bit better on a vanilla basis. GPT-4.5 is actually quite a bit better.
But when you think about what this curve of intelligence looks like, the more compute you spend, the better the model gets. That’s whether it’s a bigger model—which GPT-5 isn’t; you can see it’s not a bigger model, it’s roughly the same size—or whether you think more. OpenAI’s first thinking models, the first few generations of o1 and o3, would think for a long time and waste a lot of tokens, if you will.
When you look at, for example, Anthropic’s thinking models, they actually think a lot less when you put them in thinking mode, while getting to the same results or better results than OpenAI was. OpenAI optimized a lot of that. The silliest example I had was when I asked o3, “Is pork red meat or white meat?” It thought for 48 seconds. I was like, “What are you doing? This should just tell me the answer.”
The nice thing is that GPT-5 will think a lot less, even if you select thinking manually. More importantly, they have the auto functionality—the router—which lets them decide whether to route you to the regular model, perhaps to a mini model if you’re out of rate limits, or to the thinking model, and how much to think.
In general, the thinking model will think less. So, there’s less compute going into a power user’s average query than before.
But isn’t it even more interesting? OpenAI can now control how much compute it wants to allocate to you, right? If we’re in a high-load situation, maybe it tunes the router a little bit so it uses less. I have no idea what they’re doing behind the curtain, but there’s this meme out there at the moment that all they did—which is a meme, right? It’s not true—is take o3, plus a couple of smaller models, put a router in front, and offer that at a lower blended price, essentially.
I think there’s a little bit of that. Cost suddenly matters, and they figured out a way to steer that. They talked about how they’ve been able to dramatically increase their infrastructure capacity. I myself was regularly using o3 or GPT-4.5, and now I’m forced to use auto, which sometimes gives me the o3-equivalent thinking model but sometimes gives me just the regular base model, which sucks.
For the free user, though, it’s actually quite interesting. The free user was not getting thinking models pretty much ever—not using them, or in many cases just opening the website and asking a query. Now, sometimes the query gets routed there, so sometimes they get a way better model. But now OpenAI can gracefully degrade them if they need to.
I think the router points to the future of OpenAI as a business. You can look at the model companies: Anthropic is fully focused on B2B—APIs, coding, Claude Code, whatever it is. OpenAI has that business, Codex, and its API business, but the majority of its revenue is really from consumers. It’s consumer subscriptions, but it has no way to upsell or make money off all the free users.
In any other consumer app, the free user still pays via ads. But that’s not compatible with AI. It’s a helpful assistant; you can’t just make the user’s result worse by injecting ads. Banner ads don’t really work in AI, either. So how do you monetize them now?
I think with the router, they’re getting really close to figuring out how to monetize that user. If you saw the new CEO of applications’ product that she launched at Shopify—I think it was Shopify—it was an agent for shopping. Now this immediately clicks: if the user asks a low-value query like, “Why is the sky blue?” just route them to mini. The model can answer perfectly fine, and that’s a large chunk of queries.
But if they ask, “What’s the best DUI lawyer near me?” all of a sudden, you’re in jail and you have one shot. You’re like, “Screw it. Let me ask ChatGPT what the best DUI lawyer is.” Soon enough, the model will be able to contact all the lawyers in the area, figure out what their results are, search their court filings, and book the best lawyer—or an airplane.
Negotiate a cut as part of that, you know.
Yeah, of course they’re going to take a cut. But this is a much better way of monetizing the free user. Etsy gets 10% of its traffic now from ChatGPT, and OpenAI makes nothing off of that, but they really will soon. Partially, that’s because Amazon blocks ChatGPT, but there’s a way to make money from shopping decisions, whether it’s booking flights or looking for items.
You can say, “Free user, I don’t care. I’m going to send you to my best model. I’m going to send you to agents. I’m going to spend ungodly amounts of compute on you because I can make money off of this.” But if it’s a query like “Help me with my homework,” I’ll send you to a decent model. I don’t need to spend money on you.
That’s how I think OpenAI can finally make money off the free user. I think that’s the biggest thing about the router.
This is super interesting. I think this is the first time we’ve seen a launch of a new model where, to some degree, cost is the headline item. So far, it was always, “Who has the smartest model? Who has the highest MMLU score?” Now we suddenly have people who use models for coding for 8 hours a day and are surprised that, if you take a large context window and the best model, it creates thousands of dollars of cost a month.
Cost matters. To some degree, the shift on the Pareto frontier between cost and performance is the new benchmark for model competitiveness. Is that what we’re seeing here, or is cost alone…?
I mean, definitely. OpenAI said it doubled its rate limits for large numbers of users. It has dramatically increased the number of tokens it’s serving from this launch, which effectively says this is an economic release.
And it probably also means the tokens are all cheaper, right? Otherwise—
Yeah, for sure. I think the funniest thing about this whole cost issue is that we’ve seen it in the coding space. Cursor had to pull away the unlimited cloud code. Initially, it had a super-expensive plan with unlimited usage, and then it had only a weekly rate limit.
Now they have hour-based rate limits. I saw the craziest thread on Twitter where this guy said he changed his sleep schedule, modeled after how sailors operate when they're solo sailing. You can't sleep uninterrupted, so they'll take power naps when they get to the right spots so they can still be safe.
They do that in the morning, when it's not very windy. But they can't sleep uninterrupted, right? Because Anthropic had to put rate limits in place that are based not just on weeks, but on a number of hours, he basically sleeps multiple times a day in small chunks just so he can maximize his usage. There's also a leaderboard on Reddit where people are competing to see how many tokens they're using through their subscription. There's a guy spending $30,000 a month.
So I'm going to find a developer in India to do pair programming with. I can get the day cycle, he can get the night cycle, and we can both maximize the account's quota. Is that the future, then?
It's clear people are taking advantage of the negative gross margin in the subscriptions that are offered. I think Anthropic probably makes a positive gross margin off my subscription. I don't code enough, but plenty of people are definitely losing money. As you said, it's an economic issue.
It'll push more and more, I think, toward usage-based pricing. If you have an underlying commodity that you're reselling to some degree, and that commodity is that large a part of your cost of goods, you need to go to usage-based pricing.
How much do you think the customer capture and stickiness for these code products is? I'm curious what you think. Once you use an IDE, once you integrate one of the CLI products, how sticky is it, or is it?
A billion-dollar question. That's a very conservative estimate. Look, Andrej Karpathy has this great slide where he basically says that if you're building an agentic system today, fundamentally what it is is this loop. Half of the loop is the model thinking and trying to do something. The other half is the user verifying what the agent did, deciding whether it's the right thing, providing feedback, and trying to steer it in the right direction, because we can't run forever. Eventually, you need to steer it back.
One half of that is the model provider. They're trying to build the best models. The other half is really about designing the best possible UI to enable the user to give feedback. I think there's value in that, so I think there's a certain amount of stickiness there. Take code editing: What are all the different tools for visualizing the code changes? How can I most easily visualize what they impact and which files? How can I get very quick feedback for small changes, versus complex feedback for complex ones? There are some tools that actually draw diagrams for you of what they do.
I think this will be the battle. I think there's stickiness in that. How much exactly? This was a great question.
So in that sense, people should be doing subscriptions to get people locked in, instead of moving to usage-based pricing?
I think it's the customers that don't want to do usage-based pricing, because it's so hard to guarantee. It's hard for usage to get away from them. You actually want guarantees, and you're willing to commit to pretty high spend in order to avoid usage-based pricing. I think it's the model companies that want usage-based pricing.
I think with consumers, it's frankly very hard not to have usage-based pricing, just because the variability is so massive. Us coding versus somebody who does this as their full-time job—you just have a factor of 20 or so difference in usage. If that costs a lot of money, I think for enterprises we could see more flat-fee pricing, because you can average it out more.
You have a developer who's using it all day. You kind of know, in a general sense, how many hours a day they're programming and what that sort of looks like.
The vibe quotas are harder.
Yeah.
Before we leave OpenAI, I want to ask a broad question. If someone was sitting here and saying, “Hey, Dylan, I'll listen to anything you tell me to do. Any advice you have, as long as it makes OpenAI more valuable,” what would you tell them?
I would say immediately launch a method for you to input your credit card into ChatGPT and agree that, for anything it agentically does for you, it'll take a cut. Then launch that product, because shopping is where we know Anthropic, OpenAI, and all the other labs are buying RL environments from Amazon, Shopify, Etsy, and all the different ways to shop on the internet, as well as airline websites.
Just integrate my calendar. I want to fly there on Thursday. Make sure I don't miss a meeting. Great—book it. Do that integration extremely well. Know my preferences, whether I like an aisle or window seat, and just take a take rate. I think this will make them so much money the moment they launch it.
I think they're working on it already, but I'd like to hear how Sam Altman thinks about it, because he's shifted his tone massively on ads over the last 6 months. He used to be like, “No way,” and now he's like, “Maybe there's a way to do it without harming the user.” I think this is how you monetize the free user. That's probably what I'd tell him, or ask him about—a whole line of questions around this.
I want to shift to Nvidia. Nvidia is having a monster year. They're up almost 70%. What are the possible paths from here? How do you see it playing out?
It depends on how bullish you are on the continued growth, but I think you guys have a good vantage point. We have a good vantage point of how fast revenue is growing for a lot of these companies, especially the code companies, but even many other applications. I think we can clearly see the demand side is accelerating.
On the training side, I think the race is on. Meta is upping hugely. Google is upping hugely. If you just look at OpenAI and Anthropic and the compute that they have and are getting this year—from Google and Amazon for Anthropic, and from Microsoft, CoreWeave, and Oracle for OpenAI—30% of the chips are going to those 2 companies alone.
Well, 1/3 of it is ads, whether it be ByteDance, Meta, or many of the other people who are doing ads. So then it's still, okay, where are the rest of these 1/3 of the chips coming from? They're mostly uneconomic providers, and I don't think it's an obvious bet that they're going to keep growing and raising bigger and bigger rounds. So what happens there?
We talked about coding earlier. Qwen3-Coder is actually super cheap if you're running it on-prem or in the cloud with all these inference libraries, and there's stuff like that as well. I think the question is how much does it keep growing? Clearly, the first 1/3—OpenAI and Anthropic lab spend—is definitely skyrocketing. The second 1/3, ads, is going to grow. It's not going to grow like crazy, but I think there's definitely an inflection point that could be hit with generative AI ads.
I know Meta has been experimenting with it a lot, but I could totally be convinced that there's going to be a huge inflection in the take rate there, where you start showing me personalized ads. Every ad looks like me, and I'll be like, “Okay, yes.” Except it's slightly better, so I feel better, and I want to buy it.
It's interesting. I have no idea how this is going to scale, but if you ask how much it could scale, how much value we're creating here, can we create enough value to actually keep growing for a long time? If you just take AI software development as an example, right?
We know we can easily get about 15% more productivity out of—
I don't think that's right. I think it's way higher.
No, no. With a straight—I've talked to a lot of enterprises—a classical enterprise, straight-up GitHub Copilot deployment gives you about 15%. We can do much more than that.
But, bro, you know how bad GitHub Copilot is? Look at the revenue ARR chart. It's so funny: in 3 months, Claude Code has surpassed them, and Cursor easily surpassed them. Even companies like Replit and Windsurf/Cognition are going to pass them. It's like you're preaching to the choir. So look, let's assume we can get this to 100%.
So we can double the productivity of a developer, right? About 30 million developers worldwide, give or take. Let’s say $100,000 in value added per developer. This might be a little high worldwide; the U.S. is low, but worldwide is high. So it’s $3 trillion.
Yeah, yeah.
Right. So we’re probably building technology here that adds $3 trillion of GDP value.
Just from a coding model?
Just from this one use case. So at least in theory, the value generation is here to keep growing, right?
Now, how that translates to the industry is much more complicated. There’s the whole famous $300 billion problem—or $200 billion problem. Now it’s the $600 billion problem. I’m sure someone is going to put out the $1 trillion problem soon enough. But there is some reality in that, of course. It ignores that infrastructure spend today is accounting for 5 years of revenue, not 1, and the revenue looks like this, not a flat line.
But I think the main thing is that AI is already generating more value than the spend. The value capture is broken, right? I legitimately believe OpenAI is not even capturing 10% of the value they’ve created in the world already, just through usage of ChatGPT. I think the same applies to Anthropic, Cursor, and whoever else you’re looking at. I think the value capture is really broken.
Even internally, I think what we’ve been able to do with 4 developers in terms of automation is remarkable. Our spend on the Gemini API is absurdly low, and yet we go through every single permit and regulatory filing around every single data center with AI. We take satellite photos of every data center, and we’re able to label our data set and recognize the generators people are using, the cooling towers, the construction progress, and the substations.
All this stuff is automated, and it’s only possible because of generative AI. We do it with very few developers, and the value capture that I’m able to generate by selling this data and consulting with it is so high. But the companies making it get nothing out of it, right? I think there’s a value-capture challenge here that far exceeds the creation. As you get models like GPT-5 or open-source models continuing to drive costs down, the value capture is just harder and harder and harder for these companies because they’re making 50% gross margin on inference, if they’re making that, or less in many cases.
In so many words, you’re saying we’re getting commoditized, and therefore you can’t capture the value. Thus, you should temper your expectations of how much you can spend on GPUs.
Well, no. I think there are still ways to inflect hugely on value capture, right? I mentioned that ads are a huge value capture.
That needs to happen before we see a massive increase.
No, I think the other thing is that there’s a lot of capital that hasn’t been spent, right? The hyperscalers can still grow CapEx 20% to 30% next year from what they’re doing this year. In addition, companies like CoreWeave and Oracle, because they’re tapping capital markets, can raise way more than 20% to 30% in CapEx.
And you go down the list further, and it’s like, “Oh, the largest infrastructure funds in the world, like Brookfield and Blackstone.” Actually, they’re turning all of their attention to investing even more in AI infrastructure. Then you have the sovereign wealth funds of the world, like G42, Norway’s sovereign wealth fund, or GIC in Singapore. These people have barely started touching AI.
So I think there’s a whole lot more CapEx that can come without it necessarily being economically motivated on day 1. I’m more so saying economically motivated CapEx can only grow so much, but there’s so much other spending where it’s not clear from a spreadsheet, if you’re basing it on a real business, that you should actually spend this much. But people will because they believe—I believe, I think you believe, and I think people in infrastructure believe—that you’ll get profit out of it. But there’s no 100% certain way to argue it.
How threatened, if at all, is Nvidia by custom silicon? I think that’s the biggest thing, right? When we look at orders from Google and Amazon, especially, and Meta—not Microsoft—their custom silicon kind of sucks. But the other 3 are really upping their orders massively over the last year.
Amazon is making millions of Trainium chips, and Google is making millions of TPUs. TPUs are clearly 100% utilized, right? Trainium’s not there, but I think Amazon will figure out how to do that. And Anthropic will.
I think that’s the biggest threat to Nvidia: people figuring out how to use custom silicon more broadly. If AI is concentrated, then custom silicon will do better. And that’s not even talking about OpenAI’s silicon team and stuff, right?
If AI is really concentrated, then custom silicon will do better. But if it gets dispersed broadly because there are all these open-source models from China, along with open-source software libraries from Nvidia and China, and that makes the deployment costs rock-bottom, then potentially—hear me out here—if Google’s TPU is able to compete with Nvidia, in theory, Google could do it on the open market. The TPU business is worth more than Google these days. Shouldn’t Google start selling its chips to everyone? In theory, it should be able to achieve a higher market cap.
I absolutely think so. I think Google’s even discussing it internally. It would require a big reorganization of the culture, how Google Cloud works, how the TPU team works, and how the JAX and XLA software teams work.
I totally think they could do it. It would just take them shaking themselves pretty hard to be able to do it. I totally think Google should sell TPUs externally—not just renting them, but selling them physically.
It’s kind of funny if a side hobby, in theory, has higher company-value potential than your—
—entire business, especially as you think about the degradation of the search business.
Yeah. But I think if you were to ask Sergey, “Do you think selling chips and racks is more valuable than cloud or Gemini?” he’d be like, “No, no, no, no, no. Gemini is going to be worth way, way, way more. It’s just not there yet today, right?”
So I think today you say Nvidia is the most valuable company in the world. Again, it’s a whole concentration thing, right? If the world is super concentrated in terms of customers, then Nvidia will not be the most valuable company in the world. But if it gets dispersed more and more—which arguably we’re starting to see, with a lot of these open-source models getting better and better and better, and with the ease of deploying them improving—then you could argue Nvidia will remain the most valuable company in the world for a long period of time.
Historically, no pun intended, software has eaten the world in most markets. If you look at the early networking days, Cisco was the most valuable company on the planet for a while. It’s no longer. The companies that built services on top, like Google, Amazon, or Meta, eventually eclipsed them.
Well, which is why Nvidia is making all these software libraries, right? That’s why they’re trying to commoditize inference. I don’t think you guys even have an inference API provider investment, do you?
I think I talked to one of the team members—maybe Rajko or someone—about why you guys didn’t invest in Together or Fireworks. The argument was, “We think just serving models alone, without making them, will be commoditized,” right?
We have all kinds of model providers.
Model providers, but I’m talking about a pure API provider investment. I think that’s right, isn’t it? I think I talked to one of the team members about why you guys didn’t invest in Together or Fireworks. The argument was that just serving models alone, without making them, would be commoditized, right?
We have some in the Stable Diffusion ecosystem, like fal. It’s a little bit different dynamically there, I think.
They tend to make much more compound models than the LLM folks. I think that’s a little bit different. But you guys don’t have one of these, like Baseten, or any of these API investments, because you think—this is from someone on the infrastructure team—that it’ll get commoditized.
The software Nvidia is making, along with vLLM and SGLang, which are open-source software projects coming out of Berkeley and now sort of have their own environments and are supported by many, means that API providers aren’t necessarily worth a ton, right? That’s sort of your argument, maybe. I think that’s relevant to this whole thing, which is: Why would you do this?
Shifting gears, what about the silicon startups? What’s your take on those? There’s a ton of capital flowing into that, right? We’ve seen—I don’t have the numbers—but probably billions being invested in chip startups.
Yeah, for sure.
Whether you're looking at companies like Etched and Rivos, and a number of other companies like MatX and others, I think it's pretty impressive that they've gotten the amount of funding they've had without even launching a chip. In the past, silicon companies would make money or raise money, but they would at least launch a chip before they got a big round. But Etched and Rivos have raised a lot of money without ever launching a chip publicly, which I think speaks to—
Yes, silicon is super capital-intensive if you're building a chip, especially an accelerator, which has so many moving pieces. There are 10 different AI accelerator companies out there that are newish in the last few years.
I think there are a lot more.
Yeah. Yeah. Yeah. That's fair. And then there's the old guard, which continues to raise money, right? Groq, Cerebras, SambaNova, Tenstorrent, and so on and so forth. Or Graphcore getting bought out by SoftBank, with SoftBank dumping money into this effort as well, right? There's a lot of capital being invested to displace Nvidia's top position.
They're a captive customer, which is themselves, right? It's huge—
And they can just win on supply chain, right, by using cheaper providers.
It's a margin-compression exercise, essentially.
Yeah. And maybe for certain workloads, like Meta's recommendation systems, they'll have a better—you know, they can specialize more. But for the most part, it's like, no, we're targeting the same workloads. We can just simplify the supply chain or bring a lot of it in-house, compress margin, and it'll be fine.
But in the case of these other companies, it's like, well, they don't have a captive customer, so now you have to contend with the fact that you're using the same ecosystem. Either I can use some custom silicon provider who's going to take a margin anyway on top, and that's going to compress what I can sell for, or I can try to bring everything in-house. But then it's like, this is really hard, right? I'm going to do all the software design, all the silicon design, build all this different IP, and manage the supply-chain pain on chips, on racks, on everything. It ends up being a huge effort in terms of team size.
At the end of it all, it's like, hey, I make a 75% gross margin as Nvidia. AMD sells its GPUs for a 50% gross margin, and they have a hard time out-engineering Nvidia—and they're great at engineering, right? But they still take more silicon area and more memory to achieve the same performance, and they have to sell for less, so their margin gets compressed.
That makes sense. But look, historically, if you look at it, typically, if new entrants in markets didn't win by marginally improving on something existing—that happens sometimes—but more likely, they jumped up some kind of disruptive technology leap, where it's like, we have a different approach, we have different technology. Is that possible here?
I mean, to some degree? Maybe this is simplifying it a little bit, but I think part of the reason why the transformer model won was because it runs so incredibly well on GPUs. A recurrent neural network is similarly performant, but it runs terribly on a GPU. So, did we sort of pick the model for an architecture? And now it's hard to come up with an architecture that—
Well, it's hardware-software co-design, right? There's all this hype about neuromorphic computing, right? Theoretically, it's amazing and super-efficient. It's like, okay, great, but there's no ecosystem of hardware and there's no ecosystem of software. It would take tens of thousands of people who are the best at AI today focusing on that to even prove out whether it's worthwhile or not, right? On the hardware side, on the software side, and on the model side.
And so you look at Groq, Cerebras, and SambaNova. They all sort of over-indexed on the models that were leading at the time when they designed their chips, and so they made certain trade-offs, right? They put a lot more memory on-chip, and Nvidia was like, well, we're not going to do that.
A lot faster, at least, right?
Well, more like—if you compare the amount of SRAM on NVIDIA's chips, it's much, much lower.
Yes, correct. They went SRAM instead of DRAM, but then they usually have less DRAM. So there's a trade-off there as well.
Right? There's less DRAM, there's more SRAM, and because there's more SRAM on the chip, you have to have less compute on the chip. And so they ended up losing, right? Because the model sizes got too big and all this, right?
And so you have this super-weird dynamic where they bet on something that was actually better, right? I have no doubt that Cerebras would run certain types of models better than NVIDIA or Groq or, hey, Dojo, right? Dojo runs certain types of models way better than NVIDIA's chips because they're optimized for that.
But then it's like, oh, well, actually, even in vision tasks, people use vision transformers now. So it's like, okay, cool. Model sizes grew and all these things. It ends up being a catch-22 in that you optimize for something, and so now today you have this new age of AI accelerator companies that are like, okay, we're going to optimize for transformers.
But by the time they started designing, they're like, okay, transformers are dense models that are this big. What's the best—you know, the hidden dimension is 8K, and your batch sizes are this big and your sequence lengths are this big, so let's just make a super-large systolic array so you can create the maximum efficiency.
And then it turns out—oh, look at DeepSeek, or go look at what the labs are doing. Actually, their shapes are much smaller. You need to do a bunch of small matrix multiplies, not massive, massive, massive single matrix multiplies per layer. And then it ends up being, oh, well, that chip you're designing for that is actually not super-effective for that.
There's still the software angle, right? NVIDIA has fantastic—
Yeah, and then there's software as well, right? But NVIDIA's going to have better networking than you. They're going to have better HBM, they're going to have a better process node, they're going to come to market faster, and they're going to be able to ramp faster. They're going to have better negotiations with whether it's TSMC or SK hynix on the memory and silicon side, or all the rack people, or copper cables—everything. They're going to have better cost efficiency, so you have to be 5× better.
But to be fair, if somebody had a viable competitor that was even marginally cost-competitive, my guess is many of the big consumers of GPUs would immediately shift some revenue there just to have a number 2, right? Just to—
There's still pretty limited traction, though, right?
Sure. Meta continues to buy from them, and Microsoft did buy a bunch, and then they stopped because it's like, well, yes, AMD is giving you all these advantages, but it still ends up not being better on a performance-per-watt basis. They have a way bigger software team, and they're somewhat competitive on all these dynamics that I mentioned, right?
So you can't just do the same thing as NVIDIA and do it better, or try to execute better. Like AMD, you have to really leap forward in some other way. But the design cycle takes so long that models will shift, right? Because they're like, okay, what does the next generation of TPU and GPU look like? Okay, let's optimize for that. And the research path is, great, yes, neuromorphic computing could be the most optimal thing for us to do.
But no one's working on that because you have to advance in the tech tree. You've chosen it, right? If you restart the tech tree, you're going to be like, "Well, this sucks." And so, if it branches this way and you're over here, you're screwed, because you have to be 5x. There's a moat, because the supply chain stuff means that 5x actually turns into a 2.5x, and then NVIDIA can compress its margin a little bit if you're actually competitive. Then that 2.5x becomes like 50% better, and so it ends up being way too difficult.
And defense supply chain, for sure.
Yeah, defense supply chain. And then they get that right. Lutnick himself said, "We had to do this for rare earth minerals," and it's interesting. There are provinces in China with rules saying that the H20 is not efficient enough to be deployed, which is super bizarre because it's clearly the best AI chip China has. Huawei is still a little bit behind.
Well, what's interesting is that efficiency is so much less of an issue in China than here because they have the power infrastructure to support it. So even if they're running less powerful chips, you would imagine that it doesn't really matter, because China has such an infinite supply of power that they'd sort of be okay with it. So it's interesting.
Which is a big challenge in America, right? There have been companies that said—Jensen keeps saying he couldn't give away H20 in America for free, but I've literally heard companies say, "No, I mean, I wouldn't, because I only have this much power. How am I going to power data centers ready to go over the next year if I bought an H20? I'd literally have less compute capacity, and then I'd lose." Even if it was free, it doesn't make sense. Whereas China doesn't care. They can build these data centers; they have the muscle.
I'm curious how this all shakes out. China's posturing really hard. They even put out something that was like, "We're investigating to see if there are back doors in the H20." It's like, "There's no back door in the H20. Chill." GPUs are usually firewalled from the public internet anyway. You step through stuff before you get to the GPU clusters, so a backdoor wouldn't even matter. I don't know. I think it'll be interesting to see, because China can definitely deploy way, way, way more power to AI the moment they decide to, but there are competing interests, right?
Because they want Huawei to be better than NVIDIA.
Yeah. And this is how NVIDIA argued to the administration: if we don't do this, I think it's a very powerful argument. For example, within Triton, which is a common ML library, ByteDance has open-sourced some stuff that plugs into it that's super awesome. There are all these other libraries—not just models—that Chinese companies open-source for NVIDIA.
In a sense, NVIDIA's argument was that by selling GPUs, they were able to stop Huawei from building up a software ecosystem, and the Western ecosystem is better. But on the flip side, if you believe the models deliver more economic value to society than the hardware—which I actually think they do; it's just that there's a value-capture problem today—then you're giving China way more by giving them H20s and soon a cut-down version of Blackwell, as Trump said, versus selling them the chips. The economic value derived from selling them the chips is not as large as being able to somehow sell them AI services.
So, is China gatekeeping power for AI?
I don't think so. Again, what we see is that even with H20s being sold into China, as well as future versions of the chip—the H20E and other chips—we still see Chinese companies like Alibaba renting GPUs outside of China because the GPUs they can get outside of China are so much better on a dollar-spend performance basis. They're renting them or even going through a Singaporean company that's effectively a Chinese company, building data centers, and putting chips in them.
I don't think China is limiting the power per se. It's that Chinese companies are growing their capex way more than US companies on a percentage basis next year. The absolute dollar number is obviously still higher for US companies spending on AI. But on a percentage basis, Chinese companies are growing more next year, and you still have the problem that dollar spend to AI output—in tokens or whatever—is going to be lower because these chips are worse.
Power is not the gating factor. It's always capital, at least today. China can spend a lot more capital if it wanted to. It's subsidizing the semiconductor industry to the tune of $150 billion to $200 billion a year through SOEs and through capex that's not generating revenue, et cetera. So it's not like China couldn't do this to the AI ecosystem, right?
Given that Meta's capex is like $60 billion and Google's capex is like $80 billion, China could totally spend way more than that on a single effort. They just haven't decided to. And I think for the US, our build-outs are constrained by power, right? Google has a ton of TPUs sitting and waiting for data centers to be powered and ready, as does Meta with GPUs. We posted about how Meta is now building these effectively as tents.
Isn't this to some degree also coupled to their unwillingness to sell them to a broader ecosystem? I mean, if they want to be confined to their own data centers and they didn't rent data-center build-out for their own hyperscale use cases quickly enough, then yes, that constrains them, right? If they were on the open market, would we still be constrained?
Yeah, yeah, for sure. Because companies like CoreWeave—why is CoreWeave valuable? It's really because they build infrastructure really fast. Their software is nice, I think, but a lot of their customers are bare metal: just replace the GPUs whenever they're broken and network aggressively, and I think they'll go anywhere.
Jensen likes them as well.
Yeah, they'll go anywhere. Yeah, that's very important as well. But they'll go because it deconcentrates the ecosystem, which is better for NVIDIA. Having worked at Intel, I know exactly what's going through his mind.
Yeah. So, I think what's really important is that CoreWeave doesn't care, right? They're like, "Oh, crypto data center. I will convert it to an AI data center." They bought a company for like $10 billion that was doing crypto mining, which was worth like $2 billion a couple of years ago. And it's not because their Bitcoin-mining business is growing. It's because they have powered data centers, right?
Anywhere and everywhere, people are trying to build powered data centers. Companies like CoreWeave and Oracle are moving to that. Actually, today Google just bought 8% of a crypto-mining company called TeraWulf, right?
Not because they're getting into crypto mining.
No, because they need the data. They want the power.
They need the power, right? And all the hyperscalers have said, "Screw it," to their sustainability pledges because they need power as fast as possible. They're doing things that take a little bit longer to move the ship, but even if you didn't do it in your own self-built data centers, there are still a lot of challenges in the open market.
There's a deficit, right? And that's constraining American AI build-outs heavily. Others could maybe do it a little bit faster, like CoreWeave or others. Oracle has an open mind as well, but it's still constraining US build-outs heavily. Even though the capital's been spent, the chips are 60% to 80% of the cost of the cluster, depending on what chips you're getting. They've already bought the chips; they just can't put them anywhere because the data centers aren't ready. That applies to Google, Microsoft, Meta, and a lot of folks.
I mean, it's really hard to build infrastructure in the US. Power-grid interconnections, transmission, substations—all of this stuff, including electrical contractors and electricians. In Texas, if you're willing to be a traveling electrician, it's oil pay, right? It used to be that if you're physically adept, you could go make $100,000 in West Texas, but who the fuck wants to do that?
Now it's like, well, you could go 200 miles away from Dallas, to what's still a reasonable town, and build a data center and work on the wiring within the data center and all this other stuff—the transmission stuff—and your pay is up 2x now versus what it was just a few years ago.
This labor problem is a challenge too. I think in China they don’t have any of these problems, but they just haven’t spent the capital yet. Capital is an issue as well, because of the scale of what’s being spent, right? NVIDIA’s revenue this year is going to be over $200 billion, and next year it expects over $300 billion. Plus, Google is going to spend around $50 billion on TPU data centers, and Amazon is going to spend tons and tons on Trainium data centers.
The scale of dollars is quickly growing to nation-state-level stuff. What’s more important is being able to decide to spend the dollars and what’s cost-effective. To some extent, China is still constrained by that, but it can smuggle chips in. It can build data centers outside China and rent data centers outside China, and have the most cost-effective Blackwell chips or whatever, right?
ByteDance is either the biggest or the second-biggest customer of Google Cloud for a reason, right? They’re getting many, many Blackwell chips from them, right? The same is true with Oracle and Microsoft. All these other companies are renting tons of chips to China anyway, because it’s more cost-effective to do that than build it yourself.
It’s not like China has this mentality where it only has to—well, the government does, but the infrastructure companies don’t. Alibaba, Tencent, ByteDance, and so on don’t.
So what’s the end game for data centers? We need more power; we need more cooling. Will, at the end, all data need to be next to a nuclear reactor or lots of solar, next to deep-sea water that we use for cooling or something like that? What’s the end game?
I think that cooling—the physical cooling of a data center—is not as significant as people think. There’s this whole narrative that AI uses so much power, and it’s not really true. Farming alfalfa uses 100 times the water of AI data centers; even by the end of the decade, it’ll be the same. Alfalfa is worth very little.
People have experimented with undersea data centers to reduce cooling costs, but that doesn’t make sense. It’s like 5% to 10% savings, but if you want to get the water out of the ocean, then put the data center into the ocean. If you want to service it, you’re screwed, right?
The same goes for power. We talk a lot about power, but it’s not actually that expensive. It’s just hard to build, right? It’s about getting to the right place, getting to the right space, and converting it down to the voltages and all the stuff that chips need.
So it’s less the magnitude of power and more where it is and how it moves.
The magnitude too, right? In terms of total worldwide energy consumption, AI data centers are still a fraction of a percent. Even by the end of the decade, the US will have around 10% of its electricity going to data centers, which is still a small fraction of total energy consumption. In terms of energy, that’s an even smaller fraction, right?
Oh, yeah, because shifting to electric vehicles could probably make a bigger swing than all the AI data centers we can build.
But outside the US—in Europe, for example—that number is not moving up that fast. In all these other countries, I think we need to build a lot more power, but it’s not some crazy amount. Doing it properly is the hard thing.
Again, the cost of power—if you look at the deals people are signing, even though the price has skyrocketed from a few cents a kilowatt-hour for these massive, massive purchases to around 10 cents, it’s still, when you think about the full TCO of the cluster, the GPU cost, the networking, and all of that stuff, far outstrips the power. The same goes for cooling.
What percentage is power? If I do a 4-year amortized GPU data center, what percentage would be power?
About 80% of the cost of a GPU data center, if you’re building Blackwell, is capital, right? It’s the GPU purchases, the networking, the physical data center, the power-conversion equipment. All of this stuff is 80% of the cost.
Then 20% is your land, power, cooling, cooling towers, backup power, generators, and all this stuff. It’s like nothing, which is why it doesn’t matter if you spend 10% or 50% more on that.
At the end of the day, the expensive thing is the infrastructure itself. This is why what Elon did would seem silly: They spent a lot more money on generators outside the data center and mobile chillers to cool the water for their liquid cooling instead of the more cost-effective option, because it got the data center up 3 months faster.
That 3 months of additional training time is worth way, way more on a TCO basis, right? The performance you got out of the chips, the time to market, and all of this is way, way faster. Therefore, it was the right decision, even though this part of the data center ballooned in cost.
Everything else is still there, and you’re still paying for the chips. If they were sitting idle, it’s not worth it, right?
Just by bypassing the grid, bypassing anything to do with interconnect, anything to do with public utilities.
Exactly. Exactly.
What’s your take on Intel? Where’s Intel going?
I think the world—well, the US needs Intel. The world needs Intel because Samsung is doing worse than Intel on leading-edge process development, in my opinion. Based on various customers in the industry having done test chips at Intel versus Samsung, I think the industry generally agrees that Intel is further along in 2-nanometer-class process technology than Samsung, but both are way behind TSMC. TSMC is a monopoly, to some extent.
The number-one question people ask is, why is TSMC not making more money? Why are they only raising prices 3% to 10% next year, depending on what it is? TSMC is a monopoly. They could raise prices a lot more, but they’re good Taiwanese people rather than dirty American capitalists.
If TSMC were owned or managed by Americans—I think most of the ownership is actually American, in terms of the stock market, on the New York Stock Exchange and all this—they would have raised prices a lot more. There’s this difficult thing to be done: There’s 1 island that controls all leading-edge semiconductors, and not just all leading-edge—the majority of trailing-edge production as well. Something needs to be done. Intel is behind, but not absurdly so. If something were to happen to Taiwan, Intel would have the most advanced technology in the world. It's just not economic.
Can you keep Intel as 1 company if you want it to be competitive?
I think the process of splitting it would take so much executive time and effort that Intel would have gone bankrupt by then, right? That’s the big challenge. I think Intel should be separate, right? But properly splitting the company, with all the management time that’s needed, is absurd.
Instead, what you need is Lip-Bu Tan, the CEO of Intel. There’s a lot of drama going around about him because he’s one of the greatest semiconductor investors ever, right? He’s invested in so many different companies. First, he was on the board of SMIC, which is effectively China’s TSMC, and that’s a big source of drama. He was the first investor in some of the biggest tool companies in China because it was a multipolar world there and he was making good investments.
Now people are getting mad about that, but he recognizes the companies. He understands the supply chain. He needs not to spend his time splitting the company, because then he never actually fixes the company, right?
Intel’s problem is that it takes 5 to 6 years to go from design to shipping the product. In some cases, more. When they tape out a chip, you send the design to the fab, the fab brings back the chip, and they go through 14 revisions in some cases. The rest of the industry goes through 1 to 3 revisions if they’re good: Send the design in, get the chip back, test it, send the design in again. For a public launch, they’ll launch a chip in 3 years.
But if you look at Intel today, they still don’t have a competitive entry on the AI side, and they will, right? So what does that mean for their offering? They’re still doing great on CPUs. They don’t have a good AI chip product. Is that long-term sustainable positioning as a standalone chip company?
IBM still makes more money off every mainframe launch. So it’s not like x86 is dead. You don’t get the growth rates, but you could totally run this as a very profitable enterprise.
The same goes for PCs. There’s some turmoil, some ARM entry, and some AMD competition, but I think it can be a very profitable business if it had 1/3 the people or half the people working on it.
And so, Lip-Bu Tan, to fix Intel, needs to go into the design side of the company and lay off a shitload of people, but keep all the good people. He needs to make sure they’re designing fast and launching fast—that from design conception to launch is 2 to 3 years, not 5 to 6. That’s on the design side, and he needs to make that profitable.
On the fabs, you have to do the same thing. There are all these people. One of the heads of fab automation at Intel—I explicitly told Lip-Bu Tan, because we have a couple of ex-Intel people who are actually good in the company who worked on the fab side. We were like, “Who’s the worst?” and they were like, “Oh, this guy sucks.”
I explicitly told Lip-Bu Tan. He had never talked to the guy because he was 4 layers down. The company has absurd amounts of hierarchy. He goes and talks to the guy, and the guy’s out, right? He figures out who’s bad and who’s good. The vast majority of the team at Intel is the team that led the world in production and process technology for 20 years, right? But there’s a lot of built-up crap.
So, he has to go figure this out, right? He can’t waste his time on, “Oh, all this structuring to split.” I think it would be better if the company split; I just don’t think he can spend the time to do that.
If the design side of the company isn’t really going to get into AI, you have to make some money there. But the fabs, I think, could truly become a competitor, but they’re going to go bankrupt by the time anything can happen. So, they have to figure out how to get capital.
I think the goals are completely correct. I think the big challenge, just reflecting back on my time there, is that right now, if you look at Intel, they essentially have software, chip design, and then the core manufacturing part. They have 3 very different cultures, and it’s very hard to get everything under 1 umbrella. I think that is the big challenge.
I think you could—you should—run the company separately, but you can’t physically separate them entity-wise because it’s going to take so long to sever all these things, and he doesn’t have time. Intel is literally going to go bankrupt if they don’t have a big cash infusion or lay off half the company. Some could argue you need to lay off 30% of the company anyway, but a lot of bad things happen if that happens.
They need to spend a lot more on building the next-generation fab. Even if they fix the fab, they don’t have money for that. There are a lot more important problems than physically separating the company, even though I think long term, yes, the fab has to be separate from the chip design and software part of the company. That’s going to make each company much more accountable and able to service its customers better. It’s just that’s going to take too long, and they’re going to go bankrupt by then.
Awesome. I hope someone does something right. I hope, I pray someone does something right: You get a big capital infusion. No, no, no—the big hyperscalers have muscled in like, “Oh, okay, wait. If TSMC eventually grows its margin to 75% because of the monopoly, plus it takes in all this stuff like co-packaged optics and power delivery and all this, all of a sudden the cost is going to spike. So, we should actually just throw $5 billion at Intel each, right? Screw it.”
That could actually give Intel enough of a lifeline to potentially get to something and maybe be competitive. That’s the hope. Can we finish by finishing this game that we started when we gave Sam Altman advice? If Jensen was here, what advice would you have for him?
I think he has a massive, massive balance sheet. Jensen does. His free cash flow is ridiculous. The new Trump tax bill institutes something really incredible, which is that you can depreciate all of the GPU cluster costs in year 1. We put out a note about how the tax implications for Meta are $10 billion a year, and across each of the major hyperscalers, it’s massive.
NVIDIA is going to spend tons and tons of cash. They’re going to spend, like, tens of billions of dollars on taxes. Why don’t they get into the infrastructure game somehow? Now, this is obviously crazy because they’re buying GPUs—their own GPUs—and putting them in data centers, and they’re competing with their own customers. But they’re already doing that anyway because their customers are trying to make chips.
They should accelerate the data center ecosystem with investments, because we think we can have a very high degree of accuracy on what they’re going to do next year in terms of revenue. It’s just the number of data center watts that are being built. This is a harder thing to shift up and down. Right now, there’s a little bit of share difference between how much is TPU versus GPU, but you have to accelerate the infrastructure, and you need to spend all of this capital that you’re building, right?
Do you want to go the route of doing buybacks and dividends? Great—you’re a loser if you do that, right? You can make more money by reinvesting in building a bigger company that’s not just chips into the ecosystem or servers into the ecosystem, but actually controls the infrastructure end to end somehow. I think there’s something he could do there with this massive war chest and this massive balance sheet.
There’s a reason NVIDIA has done some buybacks and increased dividends, but the cash on its balance sheet keeps growing. They’re going to have north of $100 billion of cash on their balance sheet by the end of this year, I think. What are you going to do with that? I think there’s something in moving much more into the infrastructure layer that they could do if he really wants to be the king of the world, which I think he does.
Sergey and Sundar.
I think they should open up the ecosystem around TPUs, right? Start selling them. Open up the software—open-source a lot more of the XLA software, because there’s OpenXLA and there’s XLA, but the vast majority is closed source. Really, really open up the ecosystem around that, and be a lot more aggressive.
They’re still not very aggressive on data centers. They’re not very aggressive on a lot of elements of the company. The TPU team’s next-generation designs are not very aggressive, partially because a lot of the TPU team has left to go to OpenAI—the best people that I knew. It was really annoying. I knew 4 or 5 people, and they all went to OpenAI, so now I don’t get as much in. I met some other people, right? But I think they could be a lot more aggressive in many ways across the company.
They don’t have to be, but they could. Because AI—take ChatGPT, for example—the shift of search queries, especially the monetizable ones, from Google to purchasing agents is going to really screw Google long term if they don’t get their act together. I think they’ve gotten their act together on DeepMind. There are still some inefficiencies, but Sergey works within DeepMind a lot, and they’re driving hard.
They’re still a little bit behind, but I think physical infrastructure, TPUs, how much money they could make, and how much they could take the wind out of everyone else’s sails if they start selling TPUs externally all matter. They should reorganize around building data centers much faster so that they once again have the most compute in the world, because they did. There are certain companies that are going to surpass them over the next few years if they don’t really get their act together. So, that’s what I would say for them. Also, learn how to ship products better, right?
Zuck.
I think Zuck—it remains to be seen what goes on with superintelligence, but they’re trying to move super fast with the data centers. Screw it, they’ll build tents instead of physical data centers because they only need these for 5 years anyway. The superintelligence moves, you could say whatever you want, but trying to buy Thinking Machines Lab for $30 billion or SSI for $30 billion didn’t work out. So then they spent not even that much on hiring—not $30 billion on hiring all these people.
I think that he recognizes the urgency with the models and the infrastructure. I really think he needs to move faster. If you read his website post about AI, I think he sees the vision: There are the wearables, there’s integrating AI into those, and there’s being your AI assistant to do all this purchasing and stuff. I think he sees the vision, but I think he also needs to focus on actually releasing that faster.
The products that they do outside of their core IP—every time they launch something, it’s kind of mid, right? Meta’s Reality Labs is doing well, but I think they should go more explicit...
Have a ChatGPT competitor, have a Claude Code competitor—just start releasing way more products. They're really focused on their individual gardens rather than branching outside of them.
Do you think Apple should have that same sense of urgency? If Tim Cook were here, what would you tell him?
The funny thing is, some of Apple's best AI people are now at Meta Superintelligence Labs. They're building an AI accelerator. They have AI models, but they're just way slower. They did mention on the last earnings call that they're going to allocate more capital to this, but it's like, guys, Apple, you're going to lose the boat if you don't spend like $50 billion on infrastructure.
You don't think the current Siri will cut it?
I think more and more you'll see people say, “Great, Apple has this walled garden,” but they can only do so much to protect it, right? IDFA—they shut down ads, or data sharing, to Meta, but Meta made better models, and now they have way more data and way more power over the user than they ever did before. It was good that Meta kicked the crutch off of them—or Apple did.
The same applies to AI. Yes, they have access to the text and they have access to this, but I think other people are going to be able to integrate user data. Agents will be able to integrate all this user data, and they'll start to lose control of what the user experience is as more and more gets disintermediated by AI being the interface, rather than touch, rather than a touchpad and keyboard.
I don't think they've truly realized what happens when the interface to computing is AI. They market it, but that's going to shift computing really heavily. They have great hardware, and their hardware teams are working on awesome stuff and form factors, but I just don't know if they get what is actually going to happen to the world in the next 5 years truly well enough, and they're not building fast enough for it.
What about Microsoft, to that end?
Microsoft has the same problem, I think. They were super aggressive in 2023 and 2024, and then they pulled back heavily, right? Now OpenAI is slipping through their grasp. There's that whole thing there. They cut back on data center investments heavily. They were going to be the largest infrastructure company in the world by a factor of 2x, which—you could argue maybe that was too much, and maybe it wouldn't have been economical—but they're losing their grasp on OpenAI.
Their internal model efforts are failing spectacularly. They're on LLM Arena right now, and they're pretty decent there, but that's just a synthetic model. It's a code name, but whatever. MAI is failing. Azure is losing a lot of share to Oracle, CoreWeave, and Google, and so on and so forth. Their internal chip effort is by far the worst of any hyperscaler. They're just mis-executed.
GitHub—how is GitHub not the highest-ARR software code model? They had the best IDE, the best source-code repository, the best enterprise sales force, the best relationship with a model company, and they were first to market. They had everything going for them.
And there's just nothing, right? GitHub Copilot is failing, and Microsoft Copilot is still crap, right?
Yeah, it's unusable.
What is going on? You need to shake the crap out of the company. I think they win a lot because they have the best business-to-business relationship with so many enterprises.
Best sales force on the planet.
Yeah, but they end up not having the actual product to sell them, which is really scary. They need to really work on product.
Satya has done great on sales and stuff, but—yeah.
If Elon were here, what advice would you give him?
A lot of people at xAI are mad about the porn models, the porn stuff. It's fine. You're going to make a ton of money off this. This is how you accelerate the revenue of that company.
But he's losing a lot of talent and axing a lot of good projects. Elon is a magnet for amazing talent and building stuff, so I won't bet against him. Since he left the administration and focused on stuff again, I don't know. I think he's focused on a lot of things.
I think the robotaxi is starting to look good again, actually. I haven't ridden one yet, but I have some friends who have ridden one. It looks pretty decent. He could not make these snap decisions, which often are the reason why he's amazing, but some of these snap decisions are hurting him. I'm not sure I can give Elon that much great advice.
I think maybe it's just: focus on the products again, more. But he is working on that stuff a lot. Yeah.
I think that might be a good place to wrap. Awesome. It was a great discussion, Dylan. Thanks so much for joining us.
Thank you for having me.
Thank you.