[BidClub_]
The Next Big Thing · · 67 分钟

Dylan Patel 谈驱动 AI 革命的基础设施|下一件大事

Dylan PatelChristopher Gannatti

YouTube
TL;DR
  • 记忆是本期节目的主线:SemiAnalysis 从2023年初认为“记忆是 AI 最大输家”,转而在 o1 促使 Patel 预测 KV cache 将爆发后,认为记忆会成为最大赢家;如今 Patel 表示,“这场短缺会持续数年”——未来3年产能年增速仅20%-30%,而需求将翻倍。价格已经上涨约4倍,接下来还会再涨2倍、3倍;缺乏价格弹性的买家将被迫退出,已有数据显示中国中低端智能手机出货量下降40%;明年 iPhone 和 MacBook“必须涨价”,不是100美元,而是“几百美元”,直到“AI 得到满足”。毛利率将升至85%-90%,随后腰斩回到70%区间或更低。
  • CPU:需求出现真正拐点,但只是小周期,不是超级周期。强化学习环境和智能体工具调用确实推动需求拐点出现;SemiAnalysis 在11月的机构研究中率先指出这一点,ARM、Intel 和 AMD 股价也已大涨,但卖方“根本不懂技术,只是在编故事”。计算很简单:单颗 CPU 约5,000美元,而单颗 Blackwell 是“5万多美元”;3,000亿-5,000亿美元的 Blackwell 对应的 CPU 销售额只有300亿-500亿美元。如今的狂热,只是为过去3年出货的约1,000万颗 GPU 或其他 AI 芯片补配 CPU,而这些芯片此前几乎都没有配 CPU。“这个市场此前定价过低,现在更合理了。”
  • CPO 被推迟,转而看好铜。“市场对 CPO 有点过于兴奋”:规模化共封装光学不是2027年,而是“2028年末,但真正放量要到2029年”。Reuben 全部采用铜连接,GPU 上的 Fineman 也仍然采用铜,因此中期更看好 Amphenol 等公司和非 CPO 光学方案,等待 CPO 的制造能力和良率追上来。
  • 电力是这轮数据中心建设的硬约束,建设规模将从今年的20 GW升至30 GW、再到50 GW,解决方案也越来越多转向现场供电。未来几年内,新增数据中心电力的一半将在表后发电——从 GE Vernova 联合循环燃气轮机,到改装成燃气发电机、由汽车修理工维护的卡车发动机(“麻烦得要命,但能奏效”)。约2年后,光伏加电池的成本将低于燃气;Nvidia 取消原本可能用于 Reuben Ultra Kyber 的800V方案,也会推迟这条转换供应链。
  • ROI 之争,数据已经给出答案:Anthropic 4月和5月均实现自由现金流为正且盈利,6月也呈现同样趋势,营收已超过500亿美元 ARR,毛利率超过70%——“Anthropic 在印钞”(“Anthropic is printing”)。SemiAnalysis 自身的 AI 年化支出也从11月不到10万美元,升至如今约1,100万美元,峰值周折算年化达到1,400万美元;对于一家90人的公司,转折点来自 Claude Code。那些压缩 AI 支出的公司“会被远远甩在身后”。
  • 成本优化意味着使用最新模型,而不是最便宜的模型:一个用 Claude 4.6 Opus、分多轮消耗100,000个 token 才能完成的任务,用 4.8 Opus 一次性消耗25,000个 token 就能完成。“Anthropic 击败 OpenAI 的主要原因”是 token 效率,而不是基准测试领先;OpenAI 模型能解决前沿科学、数学和代码中的极端案例,但会消耗3-4倍 token,人类反馈循环也更慢。对于嵌入固定流程的 AI,逻辑正好相反:冻结质量,跟随每年约60倍的成本下降曲线——DeepSeek 约2年后比可能的 GPT-4 便宜600倍,而连续两次60倍降价本应意味着3,600倍。
  • 在 SemiAnalysis 的开源 InferenceX 基准测试中,Blackwell 在 DeepSeek V3 上沿性能区间某处测得比 Hopper 快30倍。这一结果甚至超过 Jensen 发布时被嘲讽的25倍说法——“Jensen,我错了。你当时是在故意压低数字”(“Jensen, I was wrong. You were sandbagging”)。Jensen 在 GTC 用5分钟展示 Patel 的图表和“Inference King”腰带,证明自己并未压低数字。
  • 应对每一个“下一场短缺”(MLCC、PCB 钻头、铜箔)的通用框架是:资金流入比例 × 需求弹性 × 市场结构。关键在于每1美元 AI 支出有多少美分流向该产品,定价究竟是记忆那样的现货商品模式,还是 TSMC 那样相对稳定的合作伙伴模式(只涨价“5%、10%”),以及市场是寡头垄断,还是由数百家公司组成、分散在台湾、日本和韩国交易。
摘要 · 为研究而整理的核心内容

1. 从少年 shitposter 到90人研究公司

  • Patel 讲起 SemiAnalysis 的起点:“SemiAnalysis 的起源其实就是 shitposting”——在自己拥有智能手机之前,就开始发布智能手机 SoC 和显示屏规格相关内容;12岁时已经在主持 Android、Intel 和 Nvidia 论坛,随后做了2年量化交易员:“没错,你能赚钱,但实际没看起来那么神奇。”2020年,他辞职,用真实姓名开设 WordPress 博客。这个选择也受到他在佐治亚州乡村、从小住在父母汽车旅馆里的经历影响——“我就是有点懂生意。”
  • 第一篇文章确立了写作模板:Huawei 失去 TSMC 供应后,美国市场认为 Qualcomm 会赢;Patel 却判断 MediaTek 是最大赢家,因为“从地缘政治角度看,中国更愿意从台湾公司采购,而不是美国公司”。技术、供应链、金融和地缘政治由此融为一体;他每年参加40场上下游会议,其中有些只有300人规模、甚至只用日语:“我当时根本没有固定住处。”
  • 向机构研究公司的转型来自第3号员工 Myin。Myin 有对冲基金背景,回应了2023年初那篇付费文章中埋在付费区的一则招聘信息;那篇文章当时认为记忆是 AI 最大输家。随后团队加入模型和数据服务,“雪球开始滚下山坡”:员工数从2人增至7人(2023-24年),再从7人增至20人、从20人增至60人,如今达到90人,今年新增30人。
  • Patel 认为,公司真正的护城河是别人没有的人才密度:上游有来自 ASML、Applied Materials 和 Lam Research 的工程师;下游有来自 Intel、TSMC、Nvidia、OpenAI、Tesla FSD 的人员,以及一位可能来自 Cohere 的员工;“我公司里还有人曾在哈萨克斯坦建过电厂。”另一半则来自对冲基金,或“互联网上那些极其投入的普通人”。

2. Jensen 登台:“Dylan 说我在压低数字,但我没有”

  • InferenceX 是 SemiAnalysis 的开源基准测试套件,每天晚上都会运行,因为 CUDA、PyTorch、驱动或推理引擎的版本随时可能更新。测试覆盖8种 GPU,以及 Google TPU 和 Amazon Trainium,使用的捐赠硬件价值超过5,000万美元,捐赠方包括 OpenAI、Microsoft、Amazon、Google、CoreWeave、Nebius、Crusoe 和 Oracle。
  • 它给出的结论是:Jensen 在发布时声称 Blackwell 比 Hopper 快25倍,很多人却说“不、不、不,只有3倍左右”;SemiAnalysis 的模拟器测得15-20倍。随后 InferenceX 在 DeepSeek V3 上测得 Blackwell 沿性能区间某处比 Hopper 快30倍。Patel 发邮件说:“Jensen,我错了。你当时是在压低数字。”
  • GTC 体育场内有20,000人,Jensen 把 Patel 的图表和 WWE 风格的“Inference King”腰带带上台展示了5分钟,以证明自己并未压低数字:“整场发布会里,他谈论我们的时间比任何其他内容都长。”获得类似展示时长的只有 OpenClaw。

3. Anthropic 在印钞,SemiAnalysis 自身 AI 账单增长超过100倍

  • Patel 给 ROI 怀疑者的第一个答案是:Anthropic 4月和5月都实现了自由现金流为正且盈利,6月也呈现同样趋势;经常性收入“已经飙过500亿美元 ARR”,毛利率超过70%。“Anthropic 在印钞。”随着 Codex 采用率上升,OpenAI 的收入也正在出现拐点。
  • 从自己的账本看,Patel 称之为 ARS,即“年度经常性支出”(annual reoccurring spend):11月不到10万美元,当时为每名员工购买200美元的 ChatGPT 席位;Claude Code 搭载 Claude Opus 4.5 和4.6进入拐点后,1月底达到400万美元;如今约1,100万美元,峰值周折算年化达到1,400万美元。AI 已经占员工支出的三分之一以上,“取决于 Methos 和其他模型最终表现,年底可能达到一半”。
  • 对于年薪约30万美元的优秀开发者,AI 支出“正开始接近一比一”;而在 SemiAnalysis,“我们很多最大的支出者其实不会写代码”,他们只是“不断迭代、不断迭代、不断迭代”。
  • 企业正面临分岔路口:很多公司在第2季度就花完了全年 AI 预算。有的削减传统 SaaS,有的裁员而不是削减 AI,有的则直接收紧 AI 支出——“但这些公司在生产力提升方面会被远远甩在身后。”

4. 成本优化意味着最新模型,而不是最便宜的模型

  • Patel 将 AI 工作负载分成两类。流程嵌入式 AI,例如检查每份流入文件是否包含 XYZ,需要先达到质量门槛、冻结模型,然后沿着成本曲线下降——在质量不变的情况下,模型每年每项任务的成本下降约60倍。“DeepSeek 比可能的 GPT-4 便宜600倍”,这让很多人感到震惊。按每年60倍的曲线计算,约2年后本应便宜3,600倍,但实际结果是600倍。
  • AI 助手则完全相反:“成本优化往往意味着选择最新模型。”一个需要 Claude 4.6 Opus 消耗100,000个 token、经过多轮交互才能完成的任务,用4.8 Opus 消耗25,000个 token 就能一次完成——token 更少,人类耗时更短,真实成本更低。
  • token 效率是“Anthropic 击败 OpenAI 的主要原因”。OpenAI 的模型可以解决领先科学、数学和代码领域的极端案例,这是 Anthropic 做不到的,“但它们耗时3倍、token 数量多4倍”,人类参与反馈的循环也会恶化。因此 SemiAnalysis 仍然是“以 Anthropic 为主的公司”,只有过夜任务才交给 OpenAI Codex。
  • 4.6→4.7 Opus 和4.7→4.8 Opus 两次发布后,他的成本都曾连续约一周下降,随后又突破前高,因为员工适应了新模型:“好,我之前做的工作完成了,那我再多做一些。”

5. 记忆:以年计算的短缺,最终由你的下一部 iPhone 买单

  • 这条思想演进很关键:2023年初,SemiAnalysis 认为记忆是 AI 最大输家,因为 AI 服务器的内存容量远低于普通服务器;随后在 o1 于2024年12月发布时转向。推理能力让 Patel 预测 KV cache 的使用量会爆发:无论上下文是1,000个 token 还是100,000个 token,权重读取方式都相同,但 KV cache 的内存读取量会随上下文增长,而计算量几乎不变。结论因此变成:记忆将是最大赢家。
  • 2026年1月,当有人问记忆价格上涨50%是否已经见顶时,他再次强调:“不、不、不——我不认为你们真正理解了。”未来3年,产能每年增长20%-30%,需求却会翻倍。“这不是短期短缺,而是一场会持续数年的短缺。”
  • 市场出清机制非常残酷:价格不断上涨,直到缺乏价格弹性的买家退出。小米等中国中低端手机厂商报告称出货量下降40%;高端市场目前尚未受到影响,因此“明年 iPhone 价格必须上涨,明年 MacBook 价格必须上涨”。而100美元不足以改变市场,所以“涨幅必须达到几百美元”,直到“AI 得到满足”。
  • 他明确表示,这仍然是周期性行业:价格已经上涨约4倍,未来还会再涨2-3倍;记忆毛利率将升至85%-90%——“记忆不应该享有85%的毛利率”——随后再腰斩回到70%区间或更低。周期依然存在,只是底部不断抬高。

6. 通用框架:资金流入比例 × 需求弹性 × 市场结构

  • 这是本期节目反复用于分析各行业的工具:每1美元 AI 支出中,有多少美分流向某个产品——“这个产品可能只有1美分,另一个产品可能有5美分”?终端市场是增长50%、翻倍,还是增长4倍?接下来是谁有能力重新定价:TSMC 缺乏弹性——“我们对客户相当公平……只会上调5%、10%的价格”;记忆则能通过现货和合约市场完成出清,因此价格上涨4倍。ASML 的波动几乎可以忽略。
  • 将框架应用到外围领域,就是 MLCC、PCB 钻头和铜箔——“你上网就会看到,这是下一场短缺,那是下一场短缺。”局部涨价确实存在,但规模“很小且数量很多”,相关公司分布在台湾、日本和韩国交易,“投资者并不容易接触到这些标的”。

7. CPU:智能体和强化学习让这块被遗忘的芯片重新变得紧俏

  • SemiAnalysis 在11月的机构研究中率先指出这一点:OpenAI 和 Anthropic 正在签订协议,租用 Amazon、Google 和 Microsoft 服务器集群中几乎全部的 CPU。主持人也注意到,过去3年 AI 讨论中几乎听不到“CPU”,如今却到处都是。
  • 机制在于:预训练几乎不需要 CPU,但强化学习会用环境检查每一代结果——单元测试、编译器、沙盒网站和购物流程;智能体推理则会持续调用现实世界中的工具,包括搜索、数据库和 Python 解释器。两者都以聊天模型从未有过的方式消耗 CPU。
  • 第三条需求支柱来自部署后的输出。GitHub 提交量较去年增加“数倍”——“很多代码很粗糙,但确实有大量代码正在部署”;网页抓取器和业务流程自动化则运行在标准、具备成本效益的 CPU 核心上。
  • 赢家分布也很清楚:Intel 和 AMD 都提高了价格;ARM 进入市场,股价“已经疯涨”;Amazon 租用 Graviton 而不是出售芯片,获得了“惊人的利润”;Nvidia 的独立 CPU Vera 给出了200亿美元收入指引。

8. 但 CPU 是重新定价,不是下一个超级周期

  • Patel 也在给自己的判断降温:“卖方根本不懂技术,只是在编故事”——当前一些比例已经暗示 CPU 规模会超过 AI 算力,这是“错误的”。一颗完整 Blackwell 的价格是“5万多美元”,而一颗 CPU 约5,000美元;即便 CPU 与 AI 算力按1:1配置,3,000亿-5,000亿美元的 Blackwell 也只对应300亿-500亿美元 CPU 销售额。
  • 真正发生的是积压需求追赶:过去3年出货的“约1,000万颗 GPU 或其他 AI 芯片,基本都没有配 CPU”。等这批存量设备补齐 CPU 后,需求就会降至增量配套水平——“我们现在处在一个 CPU 小周期里。”他的结论是:“这个市场此前定价过低,现在更合理了。”
  • 根据核心数量定律,设计空间取决于工作负载:核心尺寸翻倍,单核性能只提升约50%。Nvidia 的 Vera 针对的是 AI 算力因等待 CPU 响应而停顿的工作负载,因此选择少于100个高速核心;AMD 的256核心方案和 Graviton,则适用于高度批处理、无需等待任何单一核心的工作负载。“有些工作负载需要 Vera,有些则需要 Graviton 或 AMD CPU。”

9. CPO 推迟到2029年,过渡期看好铜

  • 网络连接相关支出增长速度快于其他任何类别:从 AI 芯片相关支出的不足10%,升至10%以上;CPO 到来后将达到20%-30%。电信光学公司——可能包括 Ciena——股价也已经大涨。
  • 但就共封装光学本身而言:“市场有点过于兴奋。”“按我的判断,2027年不会到来——真正要到2028年末,但规模化共封装光学真正放量要到2029年。”制造量、良率和芯片设计都还不到位;交换机级 CPO 会早于 GPU 级 CPO 到来。
  • 周一的机构研究报告判断:中期看好铜和非 CPO 光学,对 CPO“某种程度上偏空”。Reuben 完全采用铜连接,GPU 上的 Fineman 也仍然是铜,而 Reuben 可能才刚刚开始出货。因此,背板制造商——可能包括 Amphenol——“未来几年实际表现会远好于此前预期”。
  • 铜在局部场景持续胜出的原因是:“归根结底,集成光学的成本远高于电信号传输。”除非距离足够远,迫使系统使用中继器或光学方案。假设5年后回头看,光学的规模会“大得多”;其中一部分已经反映在价格中,另一部分还没有。

10. 电力:从卡车发动机到太空,最终会向资金让路的约束

  • 规模正在迅速扩大:今年部署20 GW数据中心,明年30 GW,后年50 GW;制约因素依次是能源、政治和建设。电力供应有3个环节,其中输电“最难看多”——公用事业垄断和成本摊销规则限制了弹性;真正有机会的是发电和转换。
  • 表后发电是释放瓶颈的出口:SemiAnalysis 预测,未来几年内,新增数据中心电力的一半将在现场生成。供应链上游是 GE Vernova、Mitsubishi 和 Siemens 的联合循环燃气轮机,下游则包括改造成发电设备的火车、船舶和卡车发动机;卡车发动机可以改用燃气,反向驱动电动机,再由电池提供缓冲,并交给“很多汽车修理店出来的人”维护。已有超过10 GW数据中心计划采用这类技术。“麻烦得要命,但能奏效。”这条路径的两端,一端是“彻底用脏能源”,另一端是“完全走向太空”,因为太空中的太阳能板甚至不需要电池。
  • 下一次成本交叉点将在“约2年后”:受益于中国的制造规模,“光伏加电池将比燃气更便宜”。但可靠性要求仍是限制条件:只需准备一晚电量的电池很便宜,要应对连续3个雨天则完全不同。
  • 电力转换本身也是一条供应链,包括 IGBT、碳化硅、GaN MOSFET、12V→54V→800V DC 的转换、固态变压器和超级电容;而这条链条刚刚遭遇打击:Reuben Ultra Kyber 很可能不再采用800V方案,导致相关供应链被推迟。SemiAnalysis 最大的研究垂直领域甚至不是半导体,而是“DEI 团队”——数据中心、能源和工业(内部是个双关语),因为团队要追踪每一个数据中心和电厂;“Google 关心 Meta 能部署什么。”

验证说明

  • 原始字幕存在内部矛盾:“memory isn't a shortage”之后紧接着说“这是一场会持续数年的短缺”;本摘要采用后者。
Christopher Gannatti

I'm joined by my colleague, Klay Hyman, and Dylan Patel, our new partner and the founder of SemiAnalysis, the research group. It's an exciting day, because we're going to go through the current state of the AI infrastructure landscape with Dylan.

You've probably heard Dylan on many different podcasts. I follow his content all the time, as well as the newsletter on the SemiAnalysis site. Most recently, one of the pieces I saw was about data centers in space. If anyone is interested in that topic, they have a great piece with a lot of detail on it.

Dylan, I would love to hear a bit about where the idea for SemiAnalysis came from. I know that in the Substack community these days, people are talking a lot about the firm's revenue and the success it has achieved, but I feel like a lot of the time you see the current level of success and forget about the journey, the origins, and how much work probably went into all of it.

1. SemiAnalysis Grows From Scratch

Dylan Patel

Yeah. I think the origin of SemiAnalysis really comes from shitposting—posting online in a less serious manner. When I go back to my first posts about semiconductors, they were from when I was a tween. I was just posting on the internet about chips, smartphones, smartphone displays, and smartphone SoCs, before I ever even had a smartphone.

It was the same with gaming hardware: PC hardware and console hardware. I was posting about this stuff all the time on the internet, on various forums. By the time I was 12, I was moderating and creating a lot of forums related to Android, Apple, Google, Intel, NVIDIA, and AMD—basically, all the hardware topics of the world. There were various forums on Reddit related to these things.

That's sort of where the origin of it all comes from. I've always been a poster. I've always posted my opinion, replied, fought, and taken comments. One of the things the people on my team say now—we're a 90-person organization, so I actually have people in marketing—is, "Dylan, stop replying to random bozos on the internet. You're making us look bad."

I just have that burning desire to respond to anyone on the internet when they try to criticize me. Maybe it's a bad thing, but throughout my teenage years I was moderating these forums. I started investing as soon as I started making money in my later teenage years. I was a quant for 2 years, and then I started my firm, but the whole time I was posting, posting, posting.

I had anonymous blogs and anonymous posts. Then, in 2020, I got fed up with my job. The disillusionment of being a quant is that it's not as amazing as it seems. Yes, you make money, but it's not as amazing as it seems.

I quit my job and started my company. I wasn't exactly sure how it was going to go, but I was posting on a WordPress website that I made. I was posting under my real name and writing about a mix of technology, business, finance, and supply chain—the things I was most interested in.

I grew up in a small business. I grew up in a motel in rural Georgia. My parents owned the motel, and we lived in it. We later had gas stations, too, so I sort of lived in business and grew up in it. I always loved business.

Supply chain was always interesting from the perspective of investing and from the perspective of how things are made. I think that's always just been a knack of mine: How are things made? The technology aspect, of course, is super exciting, and the finance aspect is super exciting.

Taking the combination of those things, my first posts were about things like China banning—or the U.S. banning—Huawei from access to TSMC. My first post was actually about how MediaTek was the biggest winner. Huawei had the number-one market share in China for smartphone chips and smartphones in general, and obviously that was going to tank because Huawei no longer had access to TSMC.

The U.S. market thought Qualcomm was going to win. But MediaTek, the Taiwanese firm, was actually going to win a lot more share because, geopolitically, China would rather buy from a Taiwanese firm than from a U.S. firm, given that we had just banned Huawei. Both firms benefited, but MediaTek benefited a lot more.

It was that mix of technology, supply chain, finance, and geopolitics, all melded together. Over the coming years, I converted my WordPress into a Substack. At one point, I started charging for it, and I was writing about all these topics across the entire semiconductor and AI supply chains.

I followed AI a lot. I did a little bit of AI when I was a quant, in addition to following semiconductors out of a passion. The Substack just grew and grew and grew.

For 4 years, I traveled all around the world and went to every conference I could. I was going to 40 conferences a year. I went to AI conferences like NeurIPS, ICML, and ICLR. Those are mostly researcher conferences, although some companies go there and present their research.

I also went all the way downstream to random conferences for, let's say, chemicals that are inputs into the semiconductor supply chain. I went up and down the stack, whether it was servers, networking, fabrication, or AI. I went to 40 conferences a year.

Some were super niche, with 300 people there, and they only spoke Japanese except for 5 people. I'd think, "Well, whatever. That is what it is." Others had 10,000 or 20,000 people and were huge. It was the whole spectrum and continuum.

I was able to go across the whole ecosystem. When you go to a conference 3 times, you actually know the language, you know people there, and you build all these contacts so you can ask them questions. I developed this whole ecosystem of knowledge, and I was covering the inflections at each of these points.

I was very curious technologically, but once something stood out—whether it was technology or supply chain—I learned from a conference where it would lead on a supply-chain or finance basis. I was writing about the whole mix.

Sometimes reports would be centered around technology and no one in finance would care. Other times, people would say, "Wow, this is the bottleneck," or, "This is the inflection that's happening," or, "This company is going to gain a ton of market share because they have next-generation technology." I would call it before anyone else on the Street, before any hedge fund, before anyone.

That was the start. As the Substack grew, in 2022 I started hiring people. My first 2 hires were people I had known from Discord for years. After that, my third hire was Mihir. He had worked at a hedge fund before and was moving to Japan to live with his wife, so he was sort of a free agent.

I had put up a post about how, at the time, it was interesting that in early 2023 memory was the biggest loser from AI. The reason was that the amount of memory AI chips and AI servers used, versus regular servers, was a lot less.

In regular servers, about half the bill of materials was memory; in AI servers, it was much less. Part of that was that NVIDIA's margins were much higher, and part of it was that there were a couple of different factors. Of course, NVIDIA's next-generation chips have increased the memory content dramatically, and now it's way more. But at the time, the argument was that memory was the biggest loser.

In the paid section, I said, "Hey, I'm hiring." Myin reached out, and he was the first person from a hedge-fund background. The other 2 people had technical backgrounds. As soon as he joined the firm, we started making all these models, and we really converted the business from what it had been.

We still do the newsletter, and we still post a lot of amazing content there—more than ever before. But we converted the business into one centered around selling information services: selling these reports and selling these data sets.

As that started to happen, the ball started tumbling down the hill. From 2023 to 2024, I went from 2 people to 7. Then, from the end of 2024 to the beginning of 2025, I went from 7 to 20. From 2025 to 2026, I went from 20 to 60. Now we're at 90, so we've added 30 people this year.

It's just been a ball tumbling down the hill, and we've added sectors. I've always been interested in everything, but now I've been able to add people who are experts. I think the most exciting thing about SemiAnalysis is that I don't know another firm with the level and concentration of expertise that we have.

I have people who have worked at ASML, Applied Materials, and Lam Research—the equipment companies that build wafers—all the way upstream to people who have worked at Intel, TSMC, NVIDIA, Microsoft, and Amazon. We also have people who have worked at OpenAI on models, someone who has worked at Tesla on FSD, and someone who has worked at Cohere.

We've got people who have worked on the model layer, and then, on another vertical, we've had people who have worked on data centers.

There’s someone at my company who built a power plant in Kazakhstan. We just have this insane talent density, which is awesome. Half the firm is people who have done engineering across the industry, and the other half is either ex-hedge-fund people or just random people from the internet who are super passionate. I found them on Twitter or Discord, and I’m like, “You’re smart. Come work for me.” And it works.

This is what’s built up, and now SemiAnalysis has many different lines of business: obviously data services, consulting, and information services. We have the newsletter. We’re doing all this media. We’re having a conference soon that’s going to be big, and all sorts of different stuff that we do. It’s just a hell of a ride.

Christopher Gannatti

Speaking of the ride, I had a moment, Dylan, where WisdomTree and SemiAnalysis have been on this journey, working together for a number of months. NVIDIA GTC happens in March, as it does every year, and I’m sitting here watching the livestream in Charlotte, North Carolina. I think there were 55,000 other people watching the livestream, and suddenly—

Dylan Patel

There are 20,000 people in a stadium. There are 20,000 people in the stadium, dude.

Christopher Gannatti

And he’s referring to you directly. You said he was sandbagging a certain number, and your charts were right up on the stage. Admittedly, I had a moment, probably on your behalf, watching the CEO of the biggest company in the world refer to your research and how you were criticizing his take on some of the numbers. I’d love for you to tell us about it. I guess you were in the stadium, from what it sounds like.

Dylan Patel

Yeah. That moment was quite surreal. Basically, one of the things that SemiAnalysis does is we have a number of engineers, and we do open-source benchmarking of all the AI models that are open source, as well as all the hardware. It’s a pretty awesome effort. There are a number of engineers on my side, but we collaborate with the industry heavily.

We get hardware, so we have over $50 million of hardware donated to us from companies like OpenAI, Microsoft, Amazon, Google, CoreWeave, Nebius, and Crusoe—all of the major clouds you can think of have donated to us. Oracle has donated hardware to us that we run these benchmarks on. We have 8 different kinds of GPUs: H100s, H200s, Blackwell, and AMD’s various GPUs. In addition, we have TPUs from Google and Trainium from Amazon.

What we do is run benchmarks on the latest version of software every single day. The reason is that every night a CUDA version could be released, a PyTorch version could be released, a driver update could be released, or an inference-engine version could be released. We run these benchmarks every single night on the entire curve—from how fast you want the tokens to how cost-effective you want them in the optimal scenarios. We run all of this every single day. It’s an automated benchmarking suite that runs.

When Jensen originally launched Blackwell, he claimed it would be a 25x improvement. At the time, no one believed him. It’s Jensen, right? He’s marketing. Even we at the time were like, “Oh, okay.” We were more bullish than ever. We thought it could be a 15x to 20x improvement based on some of the simulations we were running, because we have a simulator for performance.

As we built out this inference benchmarking called InferenceX, we got to the point where we realized, “Oh, wow. In DeepSeek V3, Blackwell is 30x faster than Hopper somewhere on the continuum.” I emailed him as soon as we had the results, and they were automatically published to the open-source GitHub. It’s an open-source collaboration; NVIDIA people helped, and they knew. But I highlighted it to him. I was like, “Hey, Jensen. Back in 2024—or back when you launched Blackwell—you said 25x, and everyone gave you crap. Here are all the people that gave you crap. Even I gave you crap. I was like, ‘There’s no way it’s 25x. It’s maybe 15x to 20x.’”

A lot of people were like, “No, no, no, it’s like 3x.” We were quite bullish, but, Jensen, I was wrong. You were sandbagging it. It was 30x. He took that, and I didn’t know that he was doing anything with it. I’d heard from a couple of customers. Someone at Meta told me there was a meeting they had, and Jensen was using that as proof that he doesn’t sandbag numbers. He was talking about the next-generation chip, and anyway, this all happened. I didn’t expect it to happen onstage.

In addition, in InferenceX, we created this belt. It looks like a WWE belt, and it says “Inference King.” We sent it to all of our collaborators. We sent it to NVIDIA, AMD, and some of the other folks—SGLang, vLLM—all these different people who helped us with the benchmark, as well as people who donated hardware.

It’s an open-source effort where I’m spending a couple million dollars a year on engineer salaries, and other people are spending millions on hardware and donating it, or millions on engineer salaries and donating them toward this open-source effort that we run. I sent Jensen this belt, and he had it on the slide. He held it up, and then there were our charts. He was on the slide for 5 minutes talking about how, “Dylan said I was sandbagging, but I wasn’t. Our performance is the best.” It was such a surreal moment.

Christopher Gannatti

He talked about us longer than anyone else in the entire presentation. The only other thing he talked about as much was OpenClaw, which is obviously taking the world by storm. It was an incredible moment.

Dylan, you mentioned a couple of things there. I think it’s quite interesting that you mentioned open source, and now that we’re transitioning a bit toward some of the recent developments and the markets, there have been discussions around the actual inference efficiency of some of the open-source models versus the closed-source models. There’s also, even to this day, a lot of investors who are questioning the ROI on all of this.

In just the last week or 2, we had a Bloomberg economist talk through how a lot of these AI initiatives may be potentially failing at some of the firms out there. I know you’ve highlighted how your firm is using AI extensively and really leaning into giving your employees lots of access to tokens. You just highlighted that you’re hiring. I’m curious: What’s your take on this end demand, and on the fact that this end demand really does drive the big buildout we’re seeing, which is full of all these constraints that have, at least over the last month—besides the most recent couple of days in the markets—been driving up a lot of the different stocks tied to some of these themes?

2. AI Spending Pays Off

Dylan Patel

Yes. I would say a few things. When you look at the overarching question here—ROI, are companies making enough money from AI, is this going to continue, and are the companies using AI actually getting value out of it?—it’s an overarching question that a lot of people have. When I look at it and think about it, there are a few ways to dissect it.

First and foremost, Anthropic is free-cash-flow positive, and they are profitable in Q2. Even in April, when they closed April’s books, they were profitable. In May, they were free-cash-flow positive and profitable. June looks like it’s going to be the same way. It’s not fully closed yet, but at least for 2 of the 3 months, they’ve been free-cash-flow positive and profitable.

Their recurring revenue has soared past $50 billion ARR, and they’re doing fantastic. That’s one side of the coin: Anthropic is printing. Obviously, there are a lot of companies that aren’t printing, but they’re getting there. OpenAI’s revenue has started to inflect as Codex’s adoption has grown, and others as well. These companies are all getting much more profitable. Anthropic’s gross margins are really, really high; they’re above 70%.

Ultimately, that’s one side of the coin. The other side of the coin, which you were alluding to, is: What about the companies’ spending on AI? At least at SemiAnalysis, we went from our annual recurring spend—I like to call it not ARR, annual recurring revenue; it’s ARS, annual recurring spend. Our annual recurring spend in November, before Claude Code really started to take off for us in December last year, was less than $100K.

What we had was a subscription to every model, or we had a subscription to the $200 tier for ChatGPT for every user, and that’s about it.

So we were spending less than $100,000 on this. If people wanted xAI or Claude, we’d give it to them as well, but our standard was giving everyone the $200 OpenAI subscription. That was the state in November, and I think we were on the bleeding edge even then.

But then Claude Code really started to hit its inflection point with Claude Opus 4.5 and 4.6, and so on and so forth. By the end of January, our ARS—our annual recurring spend—had hit $4 million. That’s because people were using Claude Code.

Today, it’s about $11 million. The highest, if we take a week of spend and multiply it by 52, we’ve been at $11 million. The highest we’ve ever had was $14 million. We’ve oscillated a lot based on what work people are doing, but right now the average looks to be about $1 million of spend a year for a 90-person firm.

That’s freaking insane, right? I just want to be clear. We’re spending more than a third of employee spend on AI, and we’ll probably get to half by the end of the year, depending on how Methos and other models start coming out and getting better and better.

That’s a huge amount of spend. Now the question is, what’s the ROI? I think there’s been huge ROI because we’ve been able to build products, sell more, and increase the efficiency of everyone in the company.

A lot of companies are questioning, “Hey, if I’m spending hundreds of thousands of dollars—if I take a really good developer and they make, call it, $300,000 a year or more, right? There are a lot of devs who make a lot more, of course—their spend on AI is starting to approach one-to-one for good developers.”

For non-developers, the spend ranges and can be lower. But even at SemiAnalysis, a lot of our biggest spenders are people who don’t know how to code. They just tell the model what they want, iterate, iterate, iterate, and get what they want.

You see this soaring spend per employee, and a lot of companies are now rightfully asking, “Hey, we blew through our entire AI budget for the year in Q1 or Q2. We blew through it already. Now what do we do?” The question is, do we cut spend, or do we cut elsewhere?

A lot of companies are saying, “Oh, maybe we need to slow down on AI spend.” But a lot of companies I’ve seen are starting to cut elsewhere. They’re cutting other SaaS products that they’ve used historically. They’re saying, “Hey, we can grow faster, so we’ll just do it.” They’re saying, “Hey, it’s okay to spend on AI. We’ll take the hit temporarily. AI keeps getting cheaper.”

As adoption soars, what I used to do 6 months ago is much cheaper with AI today. Of course, what I’m doing with AI today is much more extensive than what I did 6 months ago. There are a variety of different approaches people are taking.

Some people are even cutting employees instead of cutting AI. Some people are clamping down on AI, but those companies are going to get left in the dust in terms of productivity gains and what they’re able to build.

Christopher Gannatti

Gotcha. One of the ways to, I’ll say, mitigate some of the incremental cost is choosing cheaper, maybe sometimes less intelligent models—maybe not always being at the leading edge there. I’ve said it’s very early, I think, in terms of some of the rumblings there, but I’m curious: is there a point where firms like yours decide there are some use cases that are more optimal for using a DeepSeek V4-type model for a certain type of work, and then obviously you might need to rely on things that cost a bit more for things that require a lot more intelligence? Is that part of the calculus here?

3. Models Differ By Workload

Dylan Patel

I think that’s absolutely part of the calculus for some folks. You have to break out AI workloads into 2 types. One is, “Hey, this is AI integrated into a process that I have.” In those cases, it’s like, “Oh, when a customer sends me a document, I check it for XYZ. I put it into the model, the model checks it, and it’s done.”

There, I just need to hit some level of quality, and then from there I can stop improving the model and start decreasing cost by waiting for newer models, cheaper models, or cost efficiency. We’ve seen AI models improve at a rate of about 60x per year in cost. You take a quality level, and a year later it’s 60x cheaper.

People freaked out about DeepSeek because it was 600 times cheaper than GPT-4. That was actually about 2 years after GPT-4, so 3,600x—60x times 60x, 3,600x—and it actually ended up being 600x cheaper. Somewhere on the curve is how much cheaper it’s getting each year.

DeepSeek V3 versus GPT-4 was 600x cheaper in 2 years. As you step forward, it’s somewhere in that range. If you have a workflow and integrate AI into that workflow, then you get to a quality level, and then you go cheaper.

The other range of work is an AI assistant. That’s where I think there’s actually a bit of a misnomer. If I’m doing my day-to-day work and asking the model to help me with this, help me find this, or help me figure out that, cost optimization isn’t actually going to a cheaper model.

Cost optimization is oftentimes taking the newest model, because the newest model can be much more efficient. Claude Opus 4.6 would take 100,000 tokens to do a task, and it might take a couple of turns—me talking to it back and forth—so it might take 100,000 tokens and 10 minutes of my time. Claude Opus 4.8 can do it in a quarter of the tokens, 25,000 tokens, and it might only take 1 back-and-forth.

The cost is actually less because the number of tokens being generated is less, and the amount of time I’m using is less. When I look at a developer or someone doing intelligence work, how do I reduce the cost? It’s actually not by using a cheaper model. It’s by taking an existing task that could sometimes be done with the model after fighting with it, going to newer and newer models, and now it’s able to do it in just 1 iteration or 1-shot the entire workflow. It’s able to do it in fewer tokens.

What we saw from Claude Opus 4.6 to when Claude Opus 4.7 came out was that my cost actually fell for a week before it soared back up, because people were using it more and more. Why did it soar back up? People had to adjust to the new workflow: “Okay, the work I was doing is done. Let me do more.”

Likewise, when Claude Opus 4.8 came out after 4.7, the cost fell for about a week or a week and a half, and then it soared back up because people were like, “Oh, yeah, now I can do more work.” You have to measure productivity alongside cost.

When it’s an AI assistant, token efficiency is really important. This is why Anthropic has been beating OpenAI: its models are more token-efficient than OpenAI’s. Actually, OpenAI’s models, on the edge cases—in terms of leading science, leading math, and leading code—can oftentimes do a task that Anthropic’s models cannot, but they take 3 times as long and 4 times as many tokens.

Therefore, it costs a lot more, and the feedback loop between human and AI is not as rapid. It ends up being worse on a customer-perception basis. It’s one thing to say, “Hey, model, do this task,” and then come back and check whether the task is done. It’s another thing to say, “Hey, I have 4 hours to do this task,” and whether it’s 1 call to the model and it does work for 4 hours, or 4 calls to the model and it goes back and forth—which one does it better?

It turns out Anthropic, when you have this human-in-the-loop feedback loop, is actually way faster and better because it’s more token-efficient. That’s the main reason why we still remain a majority-Anthropic shop.

For some tasks, people do use OpenAI. Oftentimes, the tasks they let run overnight are the ones they give to OpenAI Codex. But most tasks they keep with Claude Code.

This is one of the interesting factors of what’s going on with the models and token efficiency: cost is a bit hard to parse out. For some tasks, you freeze the model quality and wait for the models to get cheaper, and for others, you actually just want the smartest model because it is cheaper.

Christopher Gannatti

Dylan, I was curious about your thoughts, shifting it a bit to the hardware side. I know earlier this year, in one of the newsletters—I’ve been a big newsletter fan for multiple years, if the audience hasn’t realized it—there was an article talking about memory.

Memory has usually been a cycle. Meaning, maybe it’s 18 to 24 months: you go up, and then 18 to 24 months, you go down. We know that it feels like almost everything is in shortage. If you’re involved in a component that goes into a data center, it feels like it’s not a question of whether you can even get the component.

It's more, okay, how long are you going to have to wait? Because it feels like in the world today, you can barely get any component. With your experience having looked across the hardware side, what do you expect is going to change with something like memory, which used to always be this commoditized product? You ride the upswing, you go through the downswing, and it just repeats, going back the last 40 years.

4. Memory Supply Tightens

Dylan Patel

Yeah. I'm not saying there aren't going to be cycles anymore. I think cycles will happen. Obviously, we're in a supercycle where the upswing is crazy, and there will be some downswing, and it'll be brutal as well. But the downswing—trough to trough—there's still a lot of growth, right?

So I think what's relevant now about memory and other components is the shifting phases of what's happening. Historically, upcycles would be up 50% for the end market, and therefore for commodity markets like memory, where pricing is more elastic, you'd end up with those stocks 2–3x. What we've gotten today is that, instead of being up 50%, spend has already doubled over just the last few years, and it's going to double again.

When you look at the elasticity of different end markets, memory pricing has gone up like 4x, and it's going to go up another 2–3x again, in addition to capacity growth. So you've got the stocks just ripping like crazy before going back down. What's really exciting about memory is that it's not just an end-market thing. It's not just that the market is ripping and it's a very elastic good, and memory is a commodity, and therefore its pricing is very elastic with end-market demand.

What's actually interesting—and this is something we wrote in 2024 when o1 came out—was that OpenAI released o1, the first reasoning model. It created a new boom of reasoning models that OpenAI, Anthropic, DeepSeek, and many others have been exploiting to get models to go after long-horizon agent tasks.

When we look at that, what's interesting is that when o1 came out, the immediate thing we noticed was that the workload changed dramatically. When we were doing chat, when you're talking to ChatGPT, you may send a prompt that might be 50 words or 500 words, but you're going to send a prompt and it's going to give you a response back. That ratio—the context length—is a few thousand. You might have a context length of, let's call it, 2,000.

When you're running inference, every time you generate a token, you read all the weights into the chip, you read all the context into the chip, you process a token, and then you iterate again. You read all the tokens, the context, and the weights. The context is called the KV cache, right? It creates this relationship between all these tokens.

What's interesting is that when you're running model inference on the weights side, whether the context length is 1,000 or 100,000, you still have to read all the weights. So memory intensity on inference is the same on the weights side. But on the side of the KV cache—the context—the memory intensity when you have 1,000 tokens that you're reading in versus 100,000 tokens is a humongous difference, even though the compute amount is roughly the same.

The compute amount is roughly the same because of KV-cache caching in memory and things like that, so you can sort of get away with it: your compute costs don't soar, but your memory costs soar. What we highlighted in our o1 note was that, in December 2024, we talked about the scaling laws and how pretraining scaling laws were giving way to reasoning scaling laws. o1 was a big step-function change. We talked about how the KV cache was going to explode because of reasoning and, therefore, memory was going to be the biggest winner.

We did that in December 2024, and multiple times in 2025, we were really excited about memory. But in January 2026, I think, is when we wrote the note saying that, at the time, people were like, "Okay, memory has gone up 50%. Is it the top of the cycle? Do we need to keep going?" And we wrote a note that was basically like, "No, no, no. I don't think you guys get it." Memory capacity is only growing 20–30% a year for the next 3 years, and yet demand is doubling.

What's going to end up happening is that memory prices are going to keep soaring. Users of memory who are less elastic, or less capable of adapting to the elasticity of pricing, will drop out of the market. Smartphones and laptops, because their costs are going to soar so much, will drop out of the market, and that's going to all give way to AI. The price is just going to have to soar until that happens because capacity is not going up enough.

Ultimately, our point there was that memory isn't a shortage, and this is not a short-term shortage. It's a shortage that's going to last years. What we've seen so far over the rest of Q1 and now Q2 is that memory has just been gangbusters. It's been soaring. There have been days where it's gone down 7–8% for some random reason, but ultimately the chart has been up and to the right.

That's not investment advice, but we see it continuing to soar because pricing continues to go up. We still haven't seen the high-end market get impacted yet, although we've had some Chinese smartphone makers in the mid-range and low-end, like Xiaomi, say their shipments are down 40%. Next year, iPhone prices have to go up. Next year, MacBook prices have to go up.

Right now, if MacBook prices or iPhone prices go up $100, that market's not going to adjust too much. But memory is going to keep getting more and more expensive until AI gets its fill. That means smartphone prices aren't just going to go up $100; they're going to have to go up a few hundred bucks.

At some point, there's going to be an equilibrium where AI gets the demand it needs and mobile and consumer hardware gets pushed down enough. Obviously, at some point people still need new phones and new laptops, so they'll still buy. We're going to have to reach a new equilibrium because memory capacity for this end market doesn't grow fast enough.

As we extend across the ecosystem, what really matters is that a lot of different components are in shortage. Who is taking elasticity? Who has an elastic price and who doesn't? An example is TSMC, which is not elastic on pricing. They're a pretty good company, pretty fair with their customers, and they partner long-term. They're like, "We'll take up prices 5–10%."

Memory companies are in a commodity market. They let the spot market and contract market, with supply and demand balancing, really adjust pricing. So you see pricing increase 2–3x, and someday you'll see pricing halve, because memory doesn't necessarily deserve an 85% margin, which is where it's headed, though. We're still not at 85–90% gross margins for memory, but we'll get there. Then at some point from there, it'll also halve back down to the 70s or maybe even lower.

We'll see this oscillation in memory. In TSMC, you don't see so much oscillation. In other areas, like ASML, we don't see much oscillation in pricing. They make equipment, but different parts of the ecosystem will oscillate differently based on, first, how much of the end AI demand flows through to them.

Different parts of the supply chain are going to have different exposure. For every dollar spent on AI, it might be $0.01 on this product, but it might be $0.05 on this product. So, obviously, there's a difference in this end market—memory versus something else. Ultimately, different end markets in the infrastructure supply chain will benefit differently in terms of demand.

In addition, what are the market dynamics there? Is it one where there's a monopoly or an oligopoly? Is it one where there's a very competitive, large market? Is it one where pricing is pretty stable and there are a lot of long-term agreements? Or is it quite a commodity market where pricing is based on supply and demand?

All of these factors determine what happens in a specific end market, whether it be memory or the shortages people are now talking about—MLCCs, PCB drill bits, PCB foil, copper foil, and all these random components. You'll go online and see, "This is the next shortage. This is the next shortage." What matters is how much flow-through of demand there actually is. Is this end market doubling? Is it going up 50%? Is it quadrupling? How much is pricing going to go up based on the market structure? These are what really determine what happens in the infrastructure supply chain.

Christopher Gannatti

And if you take that framework, it feels like year by year the market wakes up to exactly what you said: a new, quote-unquote, potential shortage. Earlier this year, we had the OpenClaw virality on various sites, which awakened people to the world of AI agents and all the possibilities. Taking the framework you just described, I'd be curious to hear your take on the CPU market, which, for the first 3 years of AI, I don't think I heard the word CPU, and this year I'm hearing CPU everywhere.

5. CPU Demand Inflects

Dylan Patel

Yeah. On the side of CPUs, what's interesting is that in some of our institutional research for our clients, in November last year, we started talking a lot about it. That's because OpenAI and Anthropic had started striking deals with Amazon, Google, Microsoft, and others to buy all the CPUs they had in their fleets and rent them out. Over the course of late last year and now this year, CPU demand has just been inflecting.

Let's talk about the reason first. Initially, when AI was training and doing inference—and inference was mostly a short-context thing—it was mostly just predicated on compute and networking, right? But as pre-training shifted to reinforcement learning, and as chat-style inference turned into agentic inference, we had this big inflection where CPUs became more in demand.

Why is that the case? In pre-training, you're training the entire web dataset into your model. Whereas in reinforcement learning, the model generates some synthetic data or a reasoning trace, and then it checks it against an environment. That environment may involve running unit tests on code. It may be a sandbox that looks like a website, or a sandbox that looks like an engineering system or some other platform that you would use, whether it be a website, a shopping site, or what have you.

It might involve compiling the code. Those environments require a lot of CPUs, whereas before, in pre-training, the actual processing of tokens didn't require much CPU; it was all the environment checking. I've generated these tokens—are they valid? What do they look like inside an environment, whether it be Python or a C compiler, or within a website if I'm trying to buy something through e-commerce? Whatever it is, as an agentic workflow, I'm testing these things constantly, and that requires a lot of CPU.

The flip side is live inference. When you're doing chat, it's, "Okay, I tell it something, it gives me an answer back, and we're done." I might ask it a few more questions, but that's it. But now, when I talk about agentic workflows where the model is making tool calls, it's, "Okay, I'm going to go search for this. I'm going to look this up in a database. I'm going to ask the Python interpreter, and I'm going to write a little bit of code to check my work. I'm going to write some code, compile it, and deploy it."

These agentic flows end up requiring more and more CPU because they have to actually interact with the regular world, right? It was one thing when the human was interacting with the model: I'm telling the model something, the model gives me a response, I read it, and I'm like, "Okay, copy and paste it into whatever it is." It's a different thing when the model is interacting with the internet, right?

There ends up being a lot more compute in the loop, a lot more AI in the loop—or, sorry, a lot more CPUs in the loop—that are bouncing the answers back and forth. Both reinforcement learning and agentic workflows need a lot of CPU.

Now what's ended up happening is that, okay, we need a lot of CPU, but let's evaluate the prior things in our framework. What is the market structure? There are a few people in the market. There's Intel and AMD. Arm is now releasing a CPU, and Arm's stock has gone gangbusters because of that, because they're a new entrant into the market that looks pretty competitive.

Then you've got Amazon, which is the leader in this, as well as Microsoft and Google, releasing their own CPUs that they've developed internally. You've also got NVIDIA releasing its own CPU. So you've got a lot of different competitors in the market, but up until 2 years ago, all of the market was Intel and AMD. Now Amazon has gotten a good amount of share, and NVIDIA and Arm are starting to get more share.

Ultimately, what happens in the end market is that Intel is actually able to increase its price. AMD is also able to increase its price, so they've both increased their pricing. They've obviously gotten demand to go up a lot. Amazon is able to extract incredible margins out of CPUs because they don't make them and sell them; they make them and rent them. Their Graviton CPUs are renting like crazy, and they've increased their orders massively.

NVIDIA, which was previously only selling CPUs attached to its GPUs, is now selling CPUs standalone with Vera. They've given guidance of $20 billion of CPU revenue. For NVIDIA, that doesn't really scratch the surface. It's like, okay, that's a few percentage points of growth. No, I'm just kidding. But when you look at other companies like Intel, AMD, Arm, and Amazon, which gets the revenue instead of just the sales revenue, there are huge things happening there.

Christopher Gannatti

Dylan, maybe on the back of that, with CPUs now, some of the discussion that I've heard has been that CPUs for agents are different than historical CPUs in some regards. The cores are more optimized for agentic activity, is what I remember hearing Jensen saying or implying around the Vera CPU.

There's also a lot of discussion around this GPU-to-CPU ratio, which obviously highlights maybe the direction of the demand and need for CPUs. Can you give us a little bit more color on each of those topics? The concept makes a lot of sense to people at a high level, but there are some technical things that are probably pushed under the rug, if there are any. I'm not sure if it's just marketing or if there's a reality to this.

Dylan Patel

When it comes to agentic workflows, the use of CPUs varies a lot. You have some agentic workflows where the model is running, and then I send a response—all the tokens—to some CPU workflow. I'm waiting on the CPU to do something, and then I send it back to the model and the model works some more.

The question is: Did the compute that the model is running on stall while you're waiting for the CPU? In some cases it does; in some cases it doesn't. In the cases where it does stall, the compute that's running the model just stalls while waiting for the CPU's response. Then the CPU needs to be architected very differently.

The basic concept is: Do I want more cores, or do I want faster cores? There's sort of a law within CPU architecture, which is basically that if you make the CPU core twice as big—which means I have half as many CPU cores on the chip—my performance doesn't go up 2 times per CPU core. My per-CPU-core performance may only go up 50%. Obviously, there's a lot of engineering involved, and the trade-off isn't that simple, but to simplify it, that's a simple way to think about it.

If I look at an NVIDIA Vera CPU, it has fewer than 100 cores, but those cores are faster than the AMD cores. AMD's leading CPU has 256 cores. So you've got this big delta in the number of CPU cores, but the NVIDIA core is faster. It's not twice as fast as an AMD CPU core. There's this trade-off that people are making in the design space.

For some workloads where the AI compute has to stall while waiting for the CPU, then you need to ask: Who cares if I have half as many cores and they're only 50% faster? In total, the performance of the CPUs is lower, but the per-core performance is higher. Therefore, I'm not waiting on the CPU cores as often. I don't need a super-parallel workload. What I really need is this one workload done now.

In that case, where the AI compute is stalling, I want to have the fastest core possible, and I'm willing to sacrifice multicore performance. That's true for some types of agentic workflows.

Other types of agentic workflows are different. If I talk about how I use Claude day to day, or how the team uses Claude—how we spend $11 million a year on Claude on an ARS basis—what is that? I'm calling Claude, and Claude is processing a bunch of tokens, but it's not just using me. It's batching hundreds of thousands of users together across all of its compute.

If I get the response back and now it's waiting on me to implement it, whether it's waiting on me or a CPU core to implement it somewhere, that's okay because the computer is still running, just not for me. It's running for other people. So if the CPU is slower but I get way more of them, it's a different sort of task.

Another question is: Is it the active use of AI, or is it what's AI-generated and then taking what AI generated and deploying it? The beauty is that, if we look at GitHub commits globally, they're up multiple times versus last year. It's not just that total GitHub commits are up 10% or 50%. They're up multiple times.

What that means is that all this code is being generated for the world, and people are deploying a lot of the code. A lot of the code is sloppy, but a lot of code is being deployed.

And when it gets deployed, it’s being put on CPUs. It might just be a web scraper, an analytical engine, or some business-process automation. That doesn’t necessarily need to be on a super-fast CPU core; it can be on a cost-effective CPU core.

When you look at the continuum, NVIDIA has built the highest-performance CPU core, but it’s not necessarily giving you the maximum number of CPU cores times the performance per core if you have a CPU chip. They’re actually not so great at that. Whereas if I look at AMD and Amazon, they have a lot more cores—hundreds—but they have less per-core performance.

You end up with this trade-off. ARM is on that end, too. Where in that continuum do you want to go? For some workloads, you do want Vera, and for some workloads, you want the Graviton or the AMD CPU. I wouldn’t say it’s as simple as that.

As far as the other question you mentioned, which is the ratio, it is indisputable that CPU demand is going up. We were the first to call it out late last year in our institutional research and in January of this year in our newsletter. Since we published that, some of these CPU stocks have ripped: ARM has gone up multiple times, Intel has gone up multiple times, NVIDIA has gone up multiple times, and AMD has ripped. These stocks have ripped.

But now the sell-side, which doesn’t really understand technology at all, is just making things up. It’s getting to the point where the ratio of CPUs to GPUs, or the ratio of CPUs to AI compute, is getting lopsided to the point where it’s more in favor of CPUs than AI compute. That’s false.

Just to reiterate, if you look at a full, all-out Blackwell, it’s like $50,000-something per chip. If you had a 1:1 ratio, CPUs cost something like $5,000. For $300 billion of Blackwell to sell, or $500 billion of Blackwell to sell, you would only get $30 billion or $50 billion of CPU sales. That’s another thing that people are starting to miss.

Yes, this end market is ripping. Ultimately, the majority of the dollars are still going to AI compute and memory. This market was underpriced, and it’s more fairly priced now. I think that’s something that people need to recognize: it’s not like CPUs are going to keep growing and growing in demand beyond that of AI ASICs.

It’s a bit of a rightsizing. In 2023 and 2024, there were years of selling millions of AI chips and very few CPUs. Now, all of a sudden, CPU demand has inflected, and the ratio shouldn’t be here; it should be here. People are in catch-up mode, so now they need to buy a bunch of CPUs to catch up with all the compute they’ve historically bought, in addition to the compute they’re currently buying.

Once I catch up with that backlog of all these AI chips that I had bought previously, and I catch up on CPUs, that demand isn’t there anymore. I’ve already caught it up, and now it’s only the incremental demand. If there’s a ratio of, let’s say, 1 CPU to 2 GPUs, and each of those GPUs costs $50,000 while each of those CPUs costs $5,000, then for every $100,000 I’m spending on GPUs, I’m spending $5,000 on CPUs.

That’s still a great market dynamic in terms of CPU growth. It’s way better than it used to be. But if you think about having 10 million GPUs and AI ASICs that I shipped over the last 3 years without any CPU really attached, then that $5,000 has a huge catch-up component. That’s what we’re experiencing right now: a huge catch-up, as well as a shift in the ratio. You’re seeing demand be ridiculous, but it will eventually subside and reach a steady state. We’re in sort of a mini CPU cycle.

Christopher Gannatti

No, that’s excellent context. Really, really helpful. Then maybe just moving to networking to move into another area of the stack. I think this is one that has come to a lot of investors’ attention, particularly as they dive down into the optics supply chain and some of the constraints there.

We’re seeing some estimates that co-packaged optics is something that’s talked about a lot but is really probably going to be deployed around 2028—2027 or 2028. As you think about it, there’s obviously this concept of “using copper when you can, optics when you must,” and this kind of transition from optics to copper. We also had Jensen talking a lot about it at Computex, along with a lot of other discussions popping up, or at least bringing a lot of attention to firms like Marvell.

Is there any additional thought you have around optics and how you see the architecture of the data center within the networking domain evolving over the next 2 years?

6. CPO Ramp Gets Delayed

Dylan Patel

As models get bigger, how do we run them across nodes? How do we train models? There are a lot of different domains within the optical stack. There are telecom optics—companies like Ciena have been ripping, along with many of the constituent supply chains around them.

Then you have datacom: chip-to-chip communications. That has a copper domain and an optics domain today, and those are all ripping because the growth of networking content is faster than the growth of any other content, in percentage terms. Networking is going from below 10% to above 10% of the spend associated with AI chips. When we get to CPO, networking grows even further—it’s like 20% to 30%.

We’ve got this huge uplift in networking content. On the flip side, CPO is such a huge step-function change in the industry, and everyone recognizes CPO now. Currently, I think people are a little too excited about CPO. It’s not coming in 2027, in my view. It’s really coming in the tail end of 2028, but 2029 is the real ramp for scale-up co-packaged optics.

There have been a lot of problems. It’s a manufacturing thing: if we could deploy it today at a good cost, everyone would do it. But it’s really hard. The manufacturing volumes aren’t there, the yields aren’t there, and the chips aren’t really designed for it yet. It’s a very complex, difficult thing to ramp.

People are going to stay in copper as long as they can. That means Reuben is all copper. Fineman on the GPU is still copper, which is the next-generation NVIDIA GPU. After Reuben, there’s Reuben Ultra, then Fineman, and we’re not even at Reuben shipping yet. Reuben is just starting to ship, so we’ve got a few generations of chips before we get to co-packaged optics on the GPU.

There’s co-packaged optics on switches, which is coming earlier than on the GPU or the AI ASICs. Ultimately, even without that, as the cluster size gets bigger, you need more optics per GPU, or more active electrical cables and things like that.

We’ve seen this big dynamic shift. On Monday, we released a note at SemiAnalysis for our institutional research subscribers, which was on a localized timeline—not saying anything about the end market, right? Obviously, our view is that CPO is going to happen; we’ve been pushing that for a long time. Our view is that copper will be subsumed over time, but on a medium-term basis, we’re actually very bullish on copper, very bullish on optics that aren’t CPO, and actually kind of bearish on CPO because of certain delays on chips that we see downstream.

Fineman isn’t going to be full CPO, along with other things that we’ve seen. Copper names like Amphenol, which makes all the backplane connectors and cables, are actually going to do way better over the next few years than previously expected because we previously thought CPO would ramp sooner, but now it’s delayed.

These things happen in the supply chain. Ultimately, optics is an area that, if you close your eyes today and open them 5 years from now, is going to be way bigger. A lot of that is priced into stocks, a lot of it isn’t, and I’d say there are some localized dislocations.

That’s part of the research that we do, and the work that we’ve been doing with you folks as well: how do we weight that? How do we weight what amount is CPO-favored optics versus non-CPO optics, typical optical transceivers versus copper? Copper has actually got a long way to go. There are a lot of things happening in the copper industry that are innovating and pushing back CPO.

Why would I do CPO? At the end of the day, integrating optics is so much more expensive than sending something electrically. Except if I have to send something electrically, I can’t go that far unless I add repeaters or optics. There’s this trade-off and continuum, and CPO will happen, but it looks like it’s getting pushed out a little bit.

Christopher Gannatti

And Dylan, as we go into what is probably our last overall topic, we’ve done models, GPUs, CPUs, memory, and networking.

We'd probably be remiss not to mention the elephant in the room at any data center: How are you getting the electricity, and how are you getting that electricity into the right form factor? I know you've written about certain things in the newsletter, at least, about direct current versus alternating current.

When you see the hyperscalers spending all this money building these data centers, potentially even putting the power plants on-site, behind the meter, how are we to think about the electrical demand—the grid versus non-grid? I know it's a big topic, but I feel like we'd be remiss not to at least mention it here.

7. Data Centers Need Power

Dylan Patel

Yeah. I would say data center growth is massive. This year, we're deploying 20 gigawatts of data centers. Next year, that number goes up 50% to 30 gigawatts, and then it'll be 50 gigawatts the year after that. The growth in data center capacity is massive.

There are a lot of local dislocations that people are having to deal with, and energy is one of the biggest ones. The other one is political, and the third is construction. Building data centers and getting permits and filings is politically difficult. People are trying to stop it, but the primary factor gating it is really the energy at the end of the day.

What's happening there is that energy can be broken down into a few things. There's generation: Where do I generate the electrons from? There's transmission: How do I transmit the electrons from where they were generated to the data center? And then there's conversion, because the power that gets transmitted is in a form factor that the chips cannot consume. The chips need to consume it in a different form factor. What does that conversion pipeline look like?

I think in all 3 of these areas, there are very bullish aspects. The transmission side is the hardest to be bullish on because of the regulatory and political difficulties with building more transmission capacity, the way local utility monopolies work, and how, if they build a utility line, they have to amortize it across all users, not just the individual user. There are all these various weird dislocations with transmitting power, and so building more grid capacity is difficult on a transmission basis.

But generation-wise and conversion-wise, there are 2 interesting things. Generation-wise, obviously, there's more generation happening on the grid. There's also this big shift toward generating power for the data center. We predict that in a couple of years, half of the power for data centers—the incremental new power for data centers—will be generated on-site, not off-site. Behind-the-meter generation is soaring.

We see this with the behind-the-meter tracker we have in our data center and energy models. I mentioned that someone on our team built a power plant in Kazakhstan. She's leading our energy model—Ellie. She has been tracking this, and we've been building a model of the entire grid: every generation asset, every transmission asset, all the load assets, as well as all the behind-the-meter work.

What's interesting is that we've seen this huge boom in behind-the-meter generation. There's been a lot of fighting on the permitting and regulatory side, whether it be people not wanting to allow air permits or people not wanting to allow the gas pipeline to be built to the site, or things like that. We've seen that with an Oracle data center. There are a lot of different aspects of this that are happening, but ultimately, the end state is that behind-the-meter generation is soaring.

A lot of it is gas. A lot of it was combined-cycle gas turbines from GE Vernova, Mitsubishi, or Siemens. But beyond that, there have also been a lot of different types of energy sources. There's reciprocating engines, industrial gas turbines, and various types of diesel engines. People have taken train engines, boat engines, and truck engines and converted them into power generation for data centers.

We see a sea of innovation happening there. It's not like we don't have the industrial capacity. The U.S. can make millions of reciprocating engines a year. These are just engines that burn fuel and spin, and it's pretty trivial to retool those to run on gas rather than diesel. But even if it's diesel, that's fine.

Then you stick an electric motor on it and basically back-drive it, and that generates electricity. You can do this in large volumes to generate power. We see 10-plus gigawatts of data centers that are going to be built with technologies like this—taking diesel truck engines, converting them to gas, which can be done at the time of production very simply, putting an electric motor on them, back-driving the motor, and then sticking them on-site at a data center.

You have hundreds of these powering a data center, and then you hire a bunch of people from car mechanic shops. These things need to be serviced, so they just run around servicing these diesel engines all day. You have some buffer, so that when they go down, you can service them and keep them going and have maximum power. Obviously, you need some batteries in between because you don't want the ups and downs of the data center to mess with or blow up the engines that you've got.

You've got this entire supply chain of behind-the-meter generation, which is exciting. In addition, in about 2 years, solar plus battery will be cheaper than gas. The supply chains for solar plus battery are difficult, and it depends on what level of reliability you want.

If you have just enough battery to get through the night, it's cheaper. But what if you need enough batteries to get through the night—3 days, right? Because it might rain for 2 days. How many nines of reliability do you want? Solar plus battery is getting cheaper and cheaper at an incredible pace because of China's manufacturing excellence and some of the subsidies, too.

It's going to get cheaper at some point to do solar plus battery. Then you've got space data centers, where you don't even need a battery. You just stick it in space, you've got a solar panel, and that's it. You've got this whole continuum of ways to generate power, whether it be taking diesel engines and making them gas engines, using combined-cycle engines, or, all the way downstream, saying, "Let's just ship the chips into space instead."

There's the whole continuum, and there's a lot of money to be made there. There's a lot of interesting, dynamic things to do there. That's why, actually, the largest data set and research vertical for SemiAnalysis—which you think is semiconductors—is actually data centers, energy, and industrials. We call it the DEI team: the data center, energy, and industrials team. It's a pun internally. The tag is @DEI team.

Jeremy leads that team. He came up with the name. Data centers, energy, and industrials is actually our biggest research vertical because we're tracking every data center and every power plant. When we identify a delay, or that something is happening, or that there's going to be a build, or that a company is going to have a certain number of data centers go online in a particular quarter, it's something that no one else in the industry can do. That's why it's one of our biggest verticals.

Everyone's interested in that. Google is interested in what Meta is able to deploy. Meta is interested in what OpenAI is able to deploy. All of these companies are also looking at what the supply chain is able to do and who has capacity, and all the investors are looking as well. That's our biggest data set.

But it's a market where it's very decentralized. In the case of memory, there are 3 names, so it's pretty simple. In the case of accelerators, there are just a few names. In the case of semiconductor wafer fabrication equipment, there are just a few names.

In this case, there are hundreds of names in the supply chain making all these random little widgets. There are dozens of companies building data centers, and there are dozens of companies trying to do different things, whether you're an independent power producer, doing it behind the meter, offering some sort of battery service, or doing all these other things. It's a very complex supply chain, but one that has a lot of dynamism. Ultimately, I think there's a lot of innovation happening.

While data centers will continue to be a constraint of sorts, they will also not be a constraint because it depends on how crazy you're willing to go. As I said, you can just take truck engines, convert them, hire a bunch of mechanics, and run a site like that.

It's not going to be the best. A lot of people say, "That's disgusting. How reliable is that going to be?" or, "That's going to be really annoying to do." But people are doing it, and it will work. It's a pain in the ass, but it will work.

You know, all the way to, “I’m going to shoot it into space.” It’s going to be a pain in the ass. It’s going to be really hard to make it work, but it will work. And so you’ve got solutions to the data center problem, whether it be going full dirty or going fully into space, whereas other parts of the supply chain, you don’t. I think that’s what makes this market so dynamic: you’re going to see people go up and down a lot.

Then I guess the other part—that’s on the generation and transmission side—but on the conversion side, the other thing is: How do you get the power from where it is generated or transmitted to what the chips want? There’s an entire supply chain of stuff going on there, whether it be IGBTs, silicon carbide, various types of MOSFETs, GaN, gallium nitride MOSFETs—all the names there. What happens when we go from 12-volt to 54-volt to 800-volt DC in the conversion supply chain? What happens with solid-state transformers as those get innovated?

All these things are happening in the space. What happens with UPSs—uninterruptible power supplies—battery backups, supercapacitors, and all these other different ways to smooth out the power, make it from the dirty, variable power that gets created on the left side to the super-clean power, but also the variable usage of it on the right side? How do you match that? That entire conversion pipeline is super, super exciting.

We had a blog on that and 800-volt just last week. We’ve talked more recently to our institutional subscribers about some delays that are happening there on the NVIDIA side as they delay it out of Kyber. Reuben Ultra—Kyber doesn’t have 800 volt anymore. So what does that mean for the supply chain? Well, it gets pushed out a little bit.

Christopher Gannatti

So, Dylan, I want to thank you profusely on our side. This is the first episode, if we think in terms of chapters. This will be the first time we’ve had Dylan on the podcast, but certainly not the last, because there’s a lot more information. As he said numerous times, everything’s changing all the time across the entire stack. It’s a bear to keep track.

Dylan Patel

The other thing I would say is this supply chain is so freaking crazy. A lot of times we talk about the big ones—memory, CPUs, data centers—but actually, when you drill down to the supply chain, the local bumps are very small and many. For a couple of months, we were talking about PCB drill bits—the drill bits that drill into PCBs for the holes for copper foil that goes on PCBs.

There are all these random small things in the supply chain that also have these dislocations. Also, the companies that exist in them are all over the world. They could be trading in Taiwan. They could be trading in Japan. They could be trading in Korea. They could be trading in all parts of the world. It’s not just easily accessible to investors.

I think that’s what’s really exciting about our partnership and the way we’re working together: We get to influence what’s going on. We get to talk a lot about these supply chain disruptions, but also what’s really interesting in the framework that I laid out earlier and the entire landscape that we’re trying to cover. I’m looking forward to coming back on the show more and to our other collaborations.