[BidClub_]
Sequoia Capital · · 70 分钟

为什么软硬件协同设计才是 AI 真正的100倍:SemiAnalysis 的 Dylan Patel

Dylan PatelShaun MaguireSonya Huang

YouTube
TL;DR
  • 核心判断:协同设计就是100倍。 分别优化硬件、系统软件和模型,3个2倍相乘只有8倍;跨越所有层级协同优化,结果“实际上是100倍”。后果是模型与芯片绑定:DeepSeek V3的专家形状为 Hopper 打造,V4则面向 Blackwell 和 Huawei 的芯片,因此“TPU运行 DeepSeek 很差”;按 OpenAI 模型的发展方向,使用 TPU 将是“一个糟糕的决定”,而 Anthropic 和 Google 用 GPU 训练也同理。
  • CUDA护城河已经不再是CUDA本身。 如今模型会写 kernel,“所有软件都会商品化”,护城河已经转移到生态:DeepSeek、Kimi、Alibaba 等中国开源模型都围绕 Nvidia GPU 协同设计,进而把下游所有推理 API 和 RL 公司拖到 Nvidia 阵营。如果 Google 开源真正优秀的模型,同样的引力会把生态拉向 TPU,这正是 Gemma 的意义。
  • 这轮算力紧张由需求驱动,而且具有持续性。 今年新增20 GW,明年即使计入延期也会超过30 GW,但 Fable 5/Mythos 5的 TAM“绝不只是 Opus 的2倍”;而在 Opus 4.5发布后的7个多月里,全球算力并没有翻倍。Anthropic 在剔除股票薪酬后于Q2实现净利润盈利,Opus 4.8 API 的 token 利润率“超过80%”,因此即使高于市场价格租 GPU 仍能持续赚钱:“我想付什么价格,就能付什么价格。”
  • GW之间并不可互换。 Trainium 的每GW租金低于100亿美元,GPU为120亿–130亿美元,SpaceX–Google 的交易约为每GW250亿美元;托管价格已从每千瓦每月60美元涨到120–160美元,最高200美元。Google 通过调配电力,在1 GW机房中塞入约1.5 GW硬件;“给 Anthropic 的1 GW,客观上比给 OpenAI 的1 GW能带来更多收入。”
  • 所有人都在做ASIC,但通用算力仍会存活。 各家实验室“根本不知道1年后会采用什么架构”,专用芯片可能在局部最优点上竞速,最终被淘汰。Google 同时运行3种不同架构的 TPU 项目,却仍以每GPU小时11美元向 xAI 租 GPU。Cerebras 的风险在于,资金会集中到最好的模型;如果前沿模型达到10T+参数、上下文达到百万 token,基于 SRAM 的芯片就装不下。
  • Jensen正在构建一个多极世界。 如果超大规模云厂商掌握全部算力,或3家封闭实验室掌握全部模型,“那将是一个他会很惨的世界”,所以他支持 neocloud 和 neolab,也乐见中国实验室发展——“把一堆诱饵撒进水里,最好的鱼会自行找到方向并活下来。”
  • 太空算力是真趋势,但还没到现在。 到2030年,太空算力占比仍低于1%;但到2040年,它可能贡献“超过一半的新增算力”。仅 OpenAI 和 Anthropic 到2030年合计就将超过100 GW,2040年达到太瓦级,推理支出将占 GDP 的“多个百分点”。
  • 他真正认可的熊市情景,与让他恼火的那一个并不相同。 Sonya 对杠杆的担忧确实成立:Crusoe 的一名客户刚要求暂停建设;如果模型产生的有用工作不再以快于算力增长的速度扩张,“潮水就会退去”。但“AI没有 ROI”让 Dylan 恼火,Shaun 否认模型进步则可以用一句话概括:“在能力维度上,这条线一直在向右上方走。”
摘要 · 为研究而整理的核心内容

1. 从汽车旅馆脚凳到 Starcraft 宗师——一个偏执狂是如何炼成的

  • 按他自己的说法,父母经营一家汽车旅馆和一家加油站,而他个子太矮,够不到香烟,于是学会根据进店的人提前把脚凳放到合适位置。“我训练的第一个神经网络”就是依据人们何时进入加油站、挑选哪种香烟,对其进行种族和外貌画像。小生意的经济学烙印一直没有离开他。
  • 他的 Xbox 360 出现“红环死亡”,他拆机后把温度传感器短接,结果“打开了潘多拉魔盒”。12岁时,他已经在 Reddit 硬件论坛担任版主;即使还是青少年,他也因利润率问题为 Nvidia 力压众人偏爱的 AMD:“它用更小的芯片,在更低功耗下获得更好的性能。”
  • 贯穿始终的是偏执式投入:他一度成为北美 Starcraft 2 天梯宗师。“痴迷是好事。”至于成绩?“足够应付亚洲父母。”

2. 2020年那场人生崩盘,催生了 SemiAnalysis

  • 创业前是一连串打击:他让量化公司创造了远超1000万美元的无风险收入,却“被克扣了奖金”;祖母患上痴呆症,并在摔倒后去世;随后 COVID 爆发。一次网络争论又导致他被人肉曝光,他停更3周,问自己“我到底在做什么?为什么在乎这些?”,并在24岁生日当天启动了 SemiAnalysis 博客。
  • 从2020年年中到2024年,他实际上一直无家可归:先是开着卡车、带着帐篷游历美国国家公园,同时阅读半导体教材;之后去了拉美;再后来每年参加供应链各环节的40多场会议。SPIE 光刻技术第一次让他彻底受挫——第一次只理解约10%,第二次一半,第三次75%:“供应链有些环节极其晦涩、极其深奥……即使到现在,我也没有全部理解。”
  • 一个日本人用蹩脚英语讲出的往事成为行业传说:1980年代,世界上生产某种化学品的唯一工厂被烧毁,存储器价格随之翻了2倍或3倍。“和今天没有太大不同。”
  • 如今 SemiAnalysis 约有90人,其中很大一部分是供应链技术专家和工程师,另一大部分来自对冲基金;内部争论——“这不重要”“但成本很重要”“这项技术太酷了”——都是自然发生的。至于1亿美元收入的传闻:“信息有多准确,这个说法就有多准确。”

3. InferenceX 跟踪模型成本每年下降60倍

  • 静态时点的 benchmark 一发布就会过时:每周都有新模型,vLLM、SGLang、PyTorch 每周更新2次,“一个接一个、毫不停歇的突破”推动同等质量模型的成本每年下降约60倍。因此,InferenceX 每天自动运行,覆盖中美最好的开源模型和约15种芯片,使用来自 CoreWeave、Crusoe、Nebius、Oracle、Microsoft、Amazon、Google 和 OpenAI 捐赠的超过5000万美元硬件;等 TPU 和 Trainium 到位后,规模将超过1亿美元。
  • 核心是找到 Pareto 最优前沿:供应商往往拿自己的最优点去对比竞争对手的非最优点——“如果我开一辆 Porsche 和赛车手开一辆赛车比赛,显然我会开得更慢”。InferenceX 会在曲线上的每个点,为最优配置开源容器。
  • 他的总框架是“吞吐量—交互性曲线”:“硬件、基础设施、模型、应用层的大多数事情——一切都由这条曲线决定。”一刀切的推理时代正在结束:批处理工作负载拿到4倍成本节省,低延迟用户则为此支付4倍价格;Claude Code 的 fast mode 和 OpenAI 的 priority queue 都是按价格切分用户的例子。

4. 太空算力与每瓦智能账本

  • 被问到多少推理会在太空完成时,他的判断是:2030年低于1%;但“20年后,我认为绝大多数算力都会进入太空”,到2040年,太空可能贡献“超过一半的新增算力”。如果能买股票,他会买 SpaceX IPO——“不是投资建议”,这句话他重复了2次。
  • 背后的规模是:到2030年,“仅 OpenAI 和 Anthropic 合计就会超过100 GW”;到2040年达到太瓦级,推理市场规模超过石油,“占 GDP 的多个百分点”。
  • 每瓦智能的提升约为40倍,相比之下成本方面是60倍。与人脑相比,我们仍相差多个数量级——“幸好这其实不重要”:给计算机供电远比给人供能容易,后者还要面对“疾病、食物偏好……睡眠”等问题。

5. 核心论点:协同设计把8倍变成100倍

  • Shaun 的框架是3个层次:硬件、kernel 级系统和模型,而且近期收益主要来自硬件;Dylan 直接反驳:“Shaun,我完全不同意。”在最优部署下,Hopper 到 Blackwell 让 DeepSeek 获得约30倍提升,但模型层带来的提升更大:3年前前沿模型还是 GPT-4,如今一个小型 Qwen 就强得多——总计27B参数,活跃参数约2B。
  • 机制在于:DeepSeek V3 的专家形状针对 Hopper 优化,V4则针对 Blackwell 和 Huawei 的芯片优化。结果是,“TPU 客观上是一款非常出色的芯片……但 TPU 运行 DeepSeek 很差”。形状、网络 I/O、集合通信和注意力算术强度都在全栈协同优化,“很难把这些收益彼此拆开”。Google 也是如此:每一代 Gemini 都为对应一代 TPU 打造,迁移到旧硬件上“真的不怎么样”。
  • 对于中国是否更擅长这套方法,他的答案是否定的:“西方不会告诉你他们在做什么。”GPT-4o 与 DeepSeek V3 的规模大致相同,前者略小,而且更早发布。
  • 逐层优化是2倍×2倍×2倍,合计8倍;跨3层协同设计,“不再只是相乘得到8倍,而是实际上达到100倍”。一家可能与 Naveen Rao 有关的公司——也是 Sequoia 的一笔投资——代表了更长期的路线:同时设计芯片、软件抽象和模型,甚至尝试将模拟计算与能量模型结合。“大概率不会成功,但这很令人兴奋……肯定不会很快成功。”

6. CUDA护城河已经不再是CUDA本身

  • Shaun 观察到,模型公司已经不再害怕其他芯片:“Claude 和 Codex 已经很擅长完成大量优化工作”,而且市场上只有几十家模型公司,不是 CUDA 护城河理论假设的几千家。Dylan 也承认:“CUDA 护城河和软件护城河至少部分已经脱钩……模型非常擅长写代码,所有软件都会商品化。”
  • 剩下的是生态引力:DeepSeek、Kimi、可能还有 Zhipu、Alibaba、可能还有 Tencent,以及 Xiaomi,都围绕 GPU 协同设计开源模型,于是下游所有参与者——推理 API 提供商、RL 定制公司——最后都会得出结论:“既然生态都在用 Nvidia,那我大概也得用 Nvidia。”如果 Google 开源真正优秀的模型,引力就会反向转向 TPU——Gemma 正是为此存在。
  • 大型实验室根本不需要开源软件栈:OpenAI 很早以前就 fork 了 PyTorch。前沿竞争的逻辑是:“我会选择最好的硬件,并把模型和基础设施软件从头到尾协同设计;然后让 AI 帮我写出所有这些软件。”

7. Nvidia 与 TPU 双赢——真正的风险是局部最优

  • 他拒绝二选一:2年后,Google 每年将生产超过1000万颗 TPU,对应每年1000亿美元以上的 TPU 规模;Nvidia 则会出货数千万颗 GPU,规模超过5000亿美元——“这不是具体估算”。他可以一本正经地论证 GPU 远胜 TPU,也可以论证 TPU 远胜 GPU;最终答案仍是协同设计:按 OpenAI 模型的发展方向,“使用 TPU 将是一个糟糕的决定”,而 Anthropic 和 Google 用 GPU 训练也可能同样糟糕。
  • 具体差异包括:OpenAI 的模型稀疏得多,Anthropic 的模型则是“更稀疏、但更稠密”;矩阵乘法单元的尺寸不同;NVLink 交换机连接72颗 GPU,而 Google 无交换机的 ICI 以高带宽连接8000颗芯片,数据会经过其他芯片。Google 同时运行3个独立的 TPU 设计项目,分别由 Broadcom、MediaTek 和一家未披露的公司负责,架构确实各不相同。
  • 最终,每家公司都会部署数十亿到数百亿颗自研 ASIC,Google 每年的规模将达到数千亿美元;但通用算力仍会存在,因为实验室“根本不知道1年后会采用什么架构”。专用芯片可能在局部最优点上竞速,却因一次注意力机制突破而被淘汰。即便 Google 也会以每GPU小时11美元向 xAI 租用 GPU;Google 的部分非 Gemini 项目——药物发现或 Waymo,“我不说是哪一个”——主要运行在 GPU 上。
  • 对 Cerebras,他认为这家公司“非常创新”,SemiAnalysis 几乎“只用” fast mode;但风险在于,fast mode 对最好的模型最重要,收入也会集中到这些模型上——Fable 5 和 Mythos 5 发布当天,尽管价格更高,仍有大量用户切换过去。若未来出现10T+参数、百万 token 上下文的模型,Cerebras 和 Groq 这样的 SRAM 芯片就无法容纳。至于 token 数量而非收入:“谁在乎 token 的数量?关键是美元。”这就是他用 F-150 对比 Camry 的观点。

8. 物理瓶颈:古老的存储单元、1W/mm²墙与柴油发动机

  • 从技术而非供应链角度看,NAND 单元已有约25年历史,DRAM 约40年;HBM 的进步主要只是堆叠更多层、提高速度。未来几年会出现的新路径是把存储器直接堆叠到计算裸片上,“让带宽爆炸式增长”。
  • 功率密度过去20年一直被锁在约1瓦/mm²:今天芯片功耗达到1400W,Rubin 达到2000W,Rubin Ultra 约4000W,全部只是继续增加硅面积。当前正在研发的突破方式是打破这面墙,让单颗芯片使用更少硅,但代价是极其困难的散热和电气干扰工程。共封装光学不是“做不做”的问题,而是“什么时候做”的问题,争论范围横跨“27年、28年、29年、2030年”。
  • 能源问题也有不性感但有效的解法:把美国能够生产的数百万台柴油卡车发动机改成燃气发动机,让电动机反向运行充当发电机,再从汽车修理店调机械师来维护这些站点。Shaun 提出的、更深层的西方问题得到 Dylan 认可:“既然你可以给广告做广告赚钱,为什么还要去硬件行业工作?”

9. 算力紧张是真实存在的——而 Anthropic 能支付高于市场的价格

  • 供给正在复合增长:今年20 GW,明年超过30 GW,且已计入延期;但需求增长更快。Mythos 5/Fable 5 的 TAM“绝不只是 Opus 的2倍”,而在 Opus 4.5发布后的7到8个月里,全球算力并没有翻倍。
  • 其下方的经济学是:Anthropic 在剔除股票薪酬后于Q2实现净利润盈利,可能到Q3连同股票薪酬也能盈利;Opus 4.8 API 的 token 利润率“超过80%”。在75%毛利率下,即使算力成本翻倍仍有50%毛利,因此“我租来的每一颗 GPU,都能立刻拿去卖 token,并保持正利润……我想付什么价格,就能付什么价格。”Anthropic 甚至以高于市场的价格从 SpaceX 买过 GPU,但仍低于 Google 后来支付的价格,因为 Google 签约更早。
  • Sonya 的反驳值得保留:所有人的杠杆都押在“必须建出来”这件事上,而 Crusoe 已公开表示,一名客户要求暂停某项建设;“高杠杆、高增长作为投资者让我非常紧张。”Dylan 先调侃她——少量股权就有巨大上行空间,“你又不是债权投资人……去上私募股权学院”——但最终承认了核心问题:如果模型产生的经济价值不再以快于算力供给增长的速度扩张,“潮水就会退去”。
  • 他的基准情景是潮水不会退:模型进步速度比6个月前更快,形成一个“伪递归自我改进循环”——模型编写基础设施,让下一代模型更早上线。资本才是约束条件:即便 SpaceX 约1000亿美元股票将在9个月后解禁,Google 仍然融资;Larry Page 当年以100亿美元估值投入10亿美元,“这是有史以来最伟大的投资之一。干得好,Larry。”Meta 宣布融资后股价却下跌。至于那个“让我非常恼火”的观点——“AI没有 ROI”、模型进步已经停滞——他的回应是:“兄弟,能力这条线一直在向右上方走……看看新 benchmark,它们正在飙升。”

10. GW之间并不可互换——Jensen想要一个多极世界

  • 价格已经在区分不同算力:Trainium 每GW租金低于100亿美元,GPU历史价格为每GW120亿–130亿美元;SpaceX–Google 的交易“好像是每GW25、甚至是疯狂得难以置信的数十亿美元”——具体约为250亿美元/GW,也就是每兆瓦每年2500万美元。托管价格已从每千瓦每月60美元升至120–160美元,最高200美元——适用于信用较弱但设施优质的租户;印度最低约80美元。在算力层面,“给 Anthropic 的1 GW,客观上比给 OpenAI 的1 GW能带来更多收入”;两家公司都能卖掉手里的全部算力,尤其是在 Codex 5.5 发布之后。
  • 运营能力会进一步放大差异:Google 通过根据工作负载知识调配电力,能把约1.5 GW硬件塞进1 GW数据中心;它还会签下2 GW的电力合同,只允许每年3天例外。Shaun 补充说,SpaceX 的 Starlink 网络能力和 Tesla 的电力管理经验,“可能是很多人分析时严重忽略的部分”;SpaceX 卖的是“现在就能运行的算力,买就行”,而 Google 卖的是提前6个月、停留在纸面上的算力,并用它为自己的采购订单融资。
  • Neocloud 之所以存在,源于他2023年的《Amazon Cloud Crisis》报告:超大规模云厂商的优势——Nitro NIC、租户隔离、定制 SSD、定制网络——原本是为按时间切片的 CPU 云设计的;当 AI 客户改为通过长期合同租用整机架后,这些优势反而变得中性甚至有害。Microsoft 的数据中心团队在需求预测翻倍时“摔得很惨”。激励机制也不同:超大规模云厂商里没人会因更快建设而变富,而 Crusoe 的“高杠杆股权所有者”会。
  • 按 Dylan 的说法,Jensen 的棋局是:“Jensen 绝对痛恨一个所有超大规模云厂商都掌握全部权力的世界。”如果世界只剩 OpenAI、Anthropic 和 Google 的模型,或只剩超大规模云厂商掌握算力,“那将是一个他会很惨的世界”。因此,他为 neocloud 提供支持、投资 neolab,并支持中国实验室——因为5年后 Crusoe 和 CoreWeave 仍然存在,会削弱 TPU 和 Trainium。“把一堆诱饵撒进水里,最好的鱼会自行找到方向并活下来。”早期回报已经出现:Thinking Machines 的 Tinker 上线约6个月不到,ARR已达到几亿美元。
Dylan Patel

I think it's really fun inside of SemiAnalysis because we have 90 people, and a big chunk of them are technologists and engineers across the whole supply chain. Another big chunk is people who are formerly at hedge funds. You see these arguments happen organically: people say, “Oh, well, that doesn’t matter,” and then someone says, “Well, but cost,” and one of the engineers says, “No, no, no, but this technology is the coolest.” You see them organically fight it out.

We’re pretty informal, and given the fact that I was a former moderator, you can imagine why they’re enjoying it. You don’t wrestle with a pig because a pig enjoys it, right?

Shaun Maguire

Exactly. We’re here in the SemiAnalysis office with Dylan Patel. I’m Shaun from Sequoia, and my partner is Sonya Huang. It’s pretty insane what you’ve done. Semis 5 years ago were not very sexy in the West. They were sexy in the East, but people here in the West had kind of forgotten about them.

You did not forget about them, though. You went very long, and you created probably the premier research company in the space. You’ve been educating the world on the state of the art, from very technical details to the supply chain to the bigger picture. There are rumors that SemiAnalysis recently passed $100 million in revenue. I don’t know how accurate those are. Whatever the numbers are, you guys are crushing it.

Dylan Patel

It’s as accurate as the information is.

Shaun Maguire

Yeah, cool. You never know. There are also rumors that you might start a venture fund. I hear all the time in the ecosystem about people wanting affiliation with SemiAnalysis. You’ve built this trusted brand, and whatever you do, it’s working. This is clearly just the beginning of the journey for you. Congratulations on all of that.

But how did this happen? My first question is: What’s your background? How did you get to where you are now?

1. Motel Kid Origins

Dylan Patel

When I was a young boy, coming out of the womb—no. Okay, I grew up in a small business. My parents had a motel; we lived in the motel and ran a gas station. I joke that the first neural network I trained was racially and visually profiling people based on when they entered the gas station and which cigarette they would pick.

The cigarettes were all displayed across the top, and I was too short to reach them. Technically, it wasn’t legal to sell cigarettes at that age, but whatever. I had to move the step stool over to the right area.

Shaun Maguire

I started working my first job before it was legal, too. It’s good experience.

Dylan Patel

Well, I didn’t get paid. It’s a family business.

Shaun Maguire

Same.

Dylan Patel

We had our motel, and across the street was our gas station. Sometimes someone would walk in, and if an old white lady with curly hair walked in, I’d move the ladder or step stool over to where the Camels were. If someone of a different age, demographic, profession, race, or whatever walked in, I’d move the step stool over accordingly.

2. Xbox Repair Spark

I joke that this was the first neural network I trained because if I waited for them to tell me what they wanted, I’d have to move it over and then step up, versus just being ready. Menthols versus 100s, slims, and all these things—I joke that was the first neural network I trained.

It all really goes back to when I was 8 years old. My birthday’s in May, and it was April when the Xbox 360 was announced. For my birthday, I didn’t ask for the Xbox—or I didn’t ask for a birthday gift. My parents asked what I wanted, and I asked for it for Christmas.

We celebrated Christmas, but there was no way—at least at the time, I thought there was no way—they would give me the Xbox 360 for Christmas. Anyway, Christmas comes around, and I get it. Fast-forward a couple of months, and my cousin, who lived in Alabama—we also lived in a motel—was going to come over for spring break.

He’s between me and my older brother in age. My older brother was a bit more of a jock, so he didn’t really care too much about the Xbox. He played sometimes, but he didn’t really care. I wanted my cousin to think I was cool, so I bragged many times on the phone: “Yeah, I got an Xbox.”

Then the Xbox broke. There was a hardware defect called the Red Ring of Death. Long story short, I had to open it up and short the temperature sensor, and it fixed it. There were many other tricks I tried first, and none of them worked. That’s how I got into hardware. It was like opening Pandora’s box.

3. Internet Forums to Semis

By the time I was 12, I was on these forums a lot, reading and posting. This was around the time Reddit ate all the other forums, and I became a moderator of Android, Apple, and Google, as well as hardware. I was watching Intel, Nvidia, and AMD and all these other forums. I was building PCs, watching, reading, and posting a lot, and moderating some of them.

I watched smartphones develop from very simple to much faster and eventually become technologically more advanced than PCs in many ways, architecturally. The same was true of all the new GPUs—just tracking and watching that, and reading every comment.

I always had an economic tinge because I grew up in a small business. I was always looking at the economics. There was a time when all the neckbeards on the internet loved AMD GPUs, and I personally had bought an AMD GPU, too, because of price-performance. But when it came down to what was technically better, I’d always be like, “No, no, no, Nvidia is better because they use a smaller chip to get better performance at better power efficiency, and their margins are better.”

I would always talk about how Nvidia’s margins were better than AMD’s in the GPU landscape, and it was very fun.

Shaun Maguire

And you were 12 at the time.

Dylan Patel

I started moderating when I was 12, but this was all through my teenage, tween, and high school years.

Shaun Maguire

Do you have any other weird hobbies, or was it just semis?

Dylan Patel

I played a ton of StarCraft. At one point, I was Grandmaster on the North American ladder in StarCraft II.

Shaun Maguire

Very serious. You’ve gotten obsessively good at multiple things.

Dylan Patel

Obsession is good.

Shaun Maguire

How were your grades?

Dylan Patel

They were decent. I had mostly A’s, but there were classes that I thought were really boring or that I just didn’t enjoy. Spanish—I got not the greatest grades. I speak fluent Spanish, by the way, so it’s really dumb, but it’s just sort of—

Shaun Maguire

Maybe that’s why you didn’t get a good grade.

Dylan Patel

I didn’t learn Spanish until later, to be fair. But, yeah, my grades were fine—fine enough for Asian parents. I was better than most of school, but it wasn’t like try-hard maxing for all A’s.

4. From Quant to Founder

Shaun Maguire

Okay. So you’re very much a student of the internet. This is how you developed this expertise. At what point did you decide to start SemiAnalysis, and what’s been the biggest surprise since starting the company?

Dylan Patel

I went to school and got a few degrees in stuff that wasn’t related to semiconductors. I was a quant for 2 years at a small quant risk firm. Basically, there was a culmination of events that happened.

One was that I got screwed out of a bonus. I had made my company many millions in risk-free revenue because I exploited a risk-related thing in the market—well over $10 million, I think—and then someone else took credit for my work and all this sort of stuff. Eventually, I did get rightsized, but I lost the social contract with the company I was working with.

My grandparents lived in the motel with us, so I was very close with them. My grandmother got dementia and forgot who I was. She fell down some stairs, had a tragic accident, and passed away. All of that happened in early 2020. Additionally, there were some girl things, so there were a few things that happened that made me very sad.

All of those things culminated, and then COVID happened. My brother was like, “Dude, just come stay with me.” He lived in Nashville, so I went and stayed with him in Nashville. We thought lockdowns would last a few weeks: I could stay with him while they happened, then go back home. Famous last words.

Lockdowns lasted much longer. Living with my brother for a few months was like, okay, I didn’t know what I was doing. I was now at my brother’s home, and everything was his rules. He and his fiancée at the time, now wife, were there, so I basically had to tiptoe around. But I didn’t care about my job, so I was posting even more than normal.

I’d always been posting a lot on the internet, and I’d always been trading stocks a lot. I made a lot of money shorting COVID and going long on COVID, and all this stuff.

Semiconductor shortages happened around then, too. Anyway, I was very much obsessed with posting and things like that. Eventually, around that time, I got into an argument with someone on the internet, and they doxed me. They publicly revealed my identity for my anonymous account. At the time, I was like, “Oh no, I’m scared.” I stopped posting for 3 weeks, and I was like, “What am I doing? Why do I care?”

So then I started posting under my real name. I had blogs and stuff as well. I made a real blog, SemiAnalysis, and on my 24th birthday, I posted 2 blogs. From there, it took off. It wasn’t a newsletter, but I got so much traction because now, instead of posting under an anonymous name, it was a real name, and I put a lot more effort into those 2 posts than I usually did. Instead of just posting on the internet, I put real effort into the blog.

5. Homeless Research Roadtrip

You can actually go back and read those if you want. They’re not that great, but they were good for the time. They were the best stuff you could find on the internet about semis. I just kept posting, posting, posting, and I started getting a lot of consulting business.

6. Cerebras Speed and Limits

In 2020, I was again crashing out. I didn’t know what I wanted to do, so I packed everything up. I took my truck, bought a tent that fit on the back of the truck, bought an air mattress, and drove around all these national parks all around America. For 2, 3, or 4 days of the week, I’d stay in a random motel where I negotiated the price down to around $30 a night for a room, and I would work on other stuff. On the weekends, I’d read books and often textbooks while in some random national park or hiking, and listen to audiobooks about semiconductors, AI, and all the things I cared about.

I got way more educated over those 6 months of going to every national park. The whole time, I was alone and posting blogs. Everyone was like, “Dylan, what the fuck are you doing?”

Shaun Maguire

Pre-Starlink or the very early days of Starlink?

Dylan Patel

Pre-Starlink. Yeah, it was very much like, “What are you doing?” I traveled around LatAm again for about a year, initially with my friend and then with my ex. Then, from the end of 2021 through 2024, I was completely—I’m still completely homeless since mid-2020—but I was traveling around to every conference in the world.

I go to 40-plus conferences a year, no matter where in the supply chain they are. I’m like, “Oh, that looks interesting. I guess I’ll go to that.” I went to 1 conference and thought, “Wow, this is amazing. You get to talk to the experts, and they’ll talk to you because you’re so excited.” In the case of semiconductors, everyone’s a boomer, so it’s great to talk to them. They don’t see young people who are excited about it, so they’re really happy to tell you stuff, and you just have to ask.

Shaun Maguire

Was there a part of the supply chain, or one of these conferences, that particularly changed your view of the semiconductor world, or that you felt then—or feel now—is particularly underrated?

Dylan Patel

I think trade shows and conferences range really widely. Obviously, some of the ones I have the most fun at include NeurIPS.

Shaun Maguire

Why is that?

Dylan Patel

Because it’s 20,000 AI researchers, and they’re generally in my age range. It’s a lot of fun, but they’re also leading AI researchers, so you learn a lot. There are also a lot of parties.

Then it ranges all the way to a random chemical conference in Japan, where it’s 300 Japanese dudes. It’s 20 guys from ASML, 20 guys from TSMC, 20 guys from Intel, and those are the only people who speak English. Everyone else speaks only Japanese, and you’re like, “Huh, I guess they’re still pretty interesting and fun.”

I think one skill set I have is that I’m able to bond with anyone, regardless of their background or who they are. I’m able to talk to them and find something interesting to talk about. Oftentimes, it’s the tech stuff, but it’s not always. I think the most interesting conferences are oftentimes the really big ones, because that’s where the biggest stuff is happening.

But I think the niches that are really exciting are SPIE. There’s IEEE, which is the Institute of Electrical and Electronics Engineers, and there’s SPIE, which is another ecosystem. SPIE conferences are super, super deep in detail. Every single one I went to—especially SPIE Advanced Lithography and SPIE Photomask—I didn’t even understand 90% of what I heard the first time.

Then I read, read, read, and built some context, of course, and the next time I went, I understood half of what I heard. The third time I went, I understood 75% of what I heard. Even now, I go and think, “I still don’t understand everything that’s going on.”

Whereas, when you go to NeurIPS a couple of times, you can understand, “Okay, what’s neurosymbolic reasoning? What’s this? What’s that?” You can kind of get a map of what everything is pretty quickly. But some parts of the supply chain are so arcane, so deep, and so technical that it takes many visits for you to even understand what’s happening in everything.

You go to a conference for a few reasons. You understand the research being published, but what you really care about is understanding how that research intersects with technology. Also, how does that research differ from what’s there today? None of these research papers tell you what’s happening today, but then you just ask people, build contacts, and learn.

Then you learn about the supply chain: this company supplies that company, even though it’s not publicly stated anywhere. You learn that this chemical costs about this much, a tool uses about this much, and you hear the horror stories of how a chemical shortage totally threw off this part of the supply chain. Then it turns out there are only 3 companies in the world that make that chemical.

My favorite one is that I learned from a Japanese guy at that specific Japanese conference I went to, where almost no one spoke English. In very broken English, he told me about how his father worked in this industry in the 1980s, that the only factory in the world that built this chemical burned down, and that caused memory prices to double or triple.

Shaun Maguire

Not too different from today.

Dylan Patel

Not at all.

7. InferenceX and Benchmarking

Shaun Maguire

Crazy. Is inference going to be the biggest market on Earth—the biggest market beyond oil? Agree or disagree?

Dylan Patel

Obviously, the use of tokens is going to be the biggest market, and the value that’s created from tokens is going to be the biggest market. But I think tokenomics—the use of tokens and adoption of AI—is the most important thing that’s happening. Inference, whether it’s open models or closed models, will be one of the biggest markets in the world, much bigger than oil. I think inference of AI will be many percentage points of GDP.

Shaun Maguire

Yeah, right. What you’ve done with InferenceX, I think, is industry standard. Maybe say a word on why you started it, what it does, and what people misunderstand about performance benchmarking for inference.

Dylan Patel

Yeah. To zoom back, SemiAnalysis does a lot of stuff. A lot of it is research for institutional clients and our subscription products, but a lot of it is also, “Hey, this would just be cool to figure out. Let’s figure out how to figure it out and just post it publicly.” That gets more and more scale.

We’ve done this with a lot of GPU benchmarking and testing, training performance, and inference performance. Ultimately, we saw that inference benchmarking was point-in-time. You test it, take some time, release it, and it’s slow, arcane, and outdated because models change all the time. Every week, I feel like there’s a new model, whether it’s a Chinese model or, you know, Mythos 5, Fable [?] dropped today.

On the software layer, PyTorch, vLLM, SGLang, new drivers—something new drops all the time. In fact, the update cycle for most of these libraries is twice a week. You basically have the software updating all the time and, therefore, performance changing. New inference optimizations are coming out and getting updated.

I feel like it’s a relentless breakthrough after breakthrough after breakthrough that keeps driving efficiency and cost down, which is why we’ve seen model costs drop for equivalent quality by 60× a year. It’s incredible. But to stay on top of that, you can’t have point-in-time benchmarking. You need benchmarks to be living and breathing—constantly running on the latest hardware and the latest models.

We embarked on a project and got a lot of buy-in from the ecosystem. This was only possible because we had enough aura with some of the ecosystem to get CoreWeave, Crusoe, Nebius, Oracle, Microsoft, Amazon, Google, and OpenAI to contribute compute to us. We were able to work with SGLang and vLLM, and now Red Hat and Inferact, which are the private companies leading those open-source efforts, to collaborate with us.

We're able to get NVIDIA, AMD, Google, and Amazon involved now because we're adding TPUs and Trainium to collaborate. Now we've got all these people collaborating. We've got over $50 million of hardware donated to us. Once we launch TPUs and Trainium, it should actually be over $100 million of hardware.

Maybe about 15 different chip types are all running these benchmarks every single day on all the latest models, right? The best model from Moonshot, the best model from Alibaba—there are about 5 different Chinese models—the best open-source models from the best Chinese labs. We run benchmarks on their models every day, and also on the best U.S. open-source models: GPT-OSS, Nemotron, et cetera. So we're running these benchmarks every day in an automated fashion, and they run on servers dedicated to us for inference benchmarking. We sweep across so many different configurations and optimization types.

What it creates—and all the results are public, and all the configurations are public—is the Pareto-optimal curve, because a lot of times when people are comparing inference performance, they're taking a suboptimal curve or point for someone else and comparing it to their optimal one. It's like, well, yeah, I can make—I can stick—if I drove a Porsche versus some race car driver, obviously I'd drive it slower. The same thing applies to inference benchmarking.

What we did is create open-source containers for the optimal points across every point on the interactivity—that is, how fast it's responding to me—versus batch size—that is, how many users I'm simultaneously serving—curve. Now anyone who wants the optimal point can just go to InferenceX, download it, and run that as the optimal point. They can check every day if they want, or they can even auto-download the most optimal point for that model, and their inference performance will be near peak.

Shaun Maguire

Is that curve the most important curve, in your opinion?

Dylan Patel

The throughput-interactivity curve is the most important one. Yeah, I think most things in hardware infrastructure, the model application layer—everything is downstream of that curve, right? Is it something that needs to be super, super fast, super-low latency, and I don't really care about the cost, so I make batch size very low and use techniques like speculative decoding or multi-token prediction heavily? There are so many possible techniques there.

Or is it something where I'm actually batch-processing a ton of documents and I don't really care about all these things? I don't use these techniques that are actually worse for cost efficiency but help with speed for an individual user, because I just want to pack a bunch of users together. I don't care if the document takes all night to process, right?

Right now, the way we treat AI infrastructure, it's like one-size-fits-all. But over time, we're going to get to the point where there are workloads where you have batch processing, or you need an instant response, and the whole curve is going to matter for users. We see this with Anthropic, right? Claude Code's fast mode costs way more than regular mode, and the same is true with OpenAI's priority processing.

Shaun Maguire

Sorry, dumb question. How does cost factor into the chart? So, in an imaginary example, I have a batch size of 100—

Dylan Patel

Okay?

Shaun Maguire

—and I can do 10 tokens per second per user. In total, I'm doing 1,000 tokens per second off that one piece of compute. That's one side of the curve: super slow, 10 tokens per second. The other side is that I have 500 tokens per second, but I only have 1 user. So maybe 250 tokens per second, 1 user. Then there are points in the middle that are Pareto-optimal, right? The average person actually wants 50 or 100 tokens a second, and maybe the number of users I can batch together.

So the curve is, okay, 1,000 tokens total per second or 250 tokens total per second, depending on how many users I batch, and there's a curve in the middle. If you had to guess, and you choose the time frame—10 years or 15 years—what percentage of inference compute do you think will happen in space?

Dylan Patel

Ultimately, some workloads will actually want the 4× cost decrease because the same unit of hardware can do 1,000 versus 250, and some users will pay 4× more because they don't care about the price; they care about time because the person using the tokens is expensive, or the feedback loop I have here is expensive.

It can be 0%, 50%—Shaun, 99%? This is a tough one. You choose the time frame, like 10—whatever time frame—and you're—

So I think the nonconsensus, or at least the view against the SpaceX thing—I love SpaceX, by the way, and I would totally buy the IPO if I could buy stocks.

Shaun Maguire

Not investment advice.

Dylan Patel

Not investment advice. Thank you. Thank you. Not investment advice from either. I don't think space data centers will really matter in the next 3 to 5 years. With that said, I think in 20 years the vast majority of compute will be going into space.

The real factor there is the cost: the time frame, the cost of building power on terrestrial land, and how much power you're going to be able to have on terrestrial land. Obviously, my view of where inference is going—how many gigawatts or terawatts are devoted to inference—is a crazy curve for me personally.

Shaun Maguire

What's your forecast? How many gigawatts?

Dylan Patel

By 2030, just OpenAI and Anthropic will have over 100 gigawatts combined. Then you'll add Meta and Google, and so on and so forth. It's a humongous amount of compute that will be dedicated to inference. By around 2040, it'll be terawatts—the curve of productivity that we're going to get.

Inference deployments are going to be huge. If you look at 2040, I think probably more than half of the incremental compute will be going into space. But if you look at 2030, I think it's sub-1%.

Shaun Maguire

Do you think intelligence per watt has been increasing? It seems like there's still a giant gap between where we are in intelligence per watt versus human biology. Do you think we're going to close that gap? And if so, where is that gain going to come from?

Dylan Patel

Yeah, I think it often depends on what you're doing, too, right? A TI-84 is way more intelligent per watt in terms of doing math than us, and it's 30 years old, right? Obviously, this is a dumb, dumb—

Shaun Maguire

General intelligence.

Dylan Patel

Yeah. But general-intelligence-wise, one of the things InferenceX does is measure the power and cost of all this hardware. So we offer not just throughput versus interactivity; we offer cost versus interactivity and power versus interactivity.

As far as whether intelligence per watt has been increasing, I mentioned that it's been a 60× cost decrease for the same benchmark level. We've also seen the same thing with intelligence per watt. It hasn't been exactly 60×; it's been closer to 40×. Some of the efficiencies are non-power-related, but there's been a humongous improvement in intelligence per watt on an annual basis—at least so far this year, last year, the year before, and the year before that. I expect that to continue.

As far as where we are compared with the human brain, we're many orders of magnitude away. Thankfully, it doesn't really matter. We can devote a lot of power to computers. It's much easier to power computers than human brains. We have sickness, disease, food preferences, and sleep.

Shaun Maguire

Yeah, exactly.

Let me ask one more question on the general theme. In terms of intelligence per watt or intelligence per dollar—any of these metrics—I think there are 3 levels of input. You can get hardware improvements, where the hardware is more efficient. You can get low-level systems optimizations, like kernel-level improvements, matrix multiplication libraries, things like that. Or you can get high-level, model-level algorithmic improvements.

It seems to me that in the last 3 years, most of the gains have come from the hardware level, with some from the model level. Do you agree with that? Do you think that's what it will look like in the future? Do you think there's a bunch of juice to squeeze in, say, the kernel level, like—

Dylan Patel

Yeah, Shaun, I completely disagree with you, by the way.

Shaun Maguire

Great. That's why I'm asking this question.

Dylan Patel

Okay, so I think one way is to look at it as these 3 different layers. In that sense, from Hopper to Blackwell, which is all we've had over the last 3 years, there's roughly a 30× improvement on DeepSeek on the most optimized deployment, which you can see on InferenceX. There's about a 30× improvement.

But over the last 3 years, we've had way more improvement in intelligence per watt, a lot of that coming from the model layer, right? If you look back 3 years, it's GPT-4. Now it's maybe Qwen, one of the smaller Qwen models, that's about 27B parameters total and 2B active, and is way better.

You've got this huge improvement on the model layer, and you've got this pretty sizable improvement on hardware, but it's that co-design layer. I think that's what's important, right? If you look at the architecture of any of these models, DeepSeek is the most famous one, at least among the public models that people have seen.

Shaun Maguire

Yeah, DeepSeek got huge efficiency gains from co-optimization, or kernel-level optimization of memory.

Dylan Patel

Yes, I think it’s kernels, of course, but it’s actually that you build the hardware architecture for the chip. If you look at the shapes of all the experts in DeepSeek V3, they were all optimized for Hopper. If you look at V4, they’re optimized for Blackwell and Huawei’s chip.

What’s interesting is that despite the fact that TPUs are objectively an amazing chip—they run all of DeepMind, and they do all the training for Anthropic as well, at least on the pretraining side—TPUs are terrible at running DeepSeek. But they’re really, really great at running other kinds of models that don’t run well on NVIDIA.

There’s some level of such deep optimization that’s been done, whether it’s the shapes, network I/O patterns, how you do the collectives, or things around the arithmetic intensity of the attention mechanism. All these different things are co-optimized between the model, the hardware, and the infrastructure software in between. It’s hard to say you can disentangle the gains.

Shaun Maguire

Do you think that my understanding is that China has done this a lot better than the West over the last few years? DeepSeek was one of the first models to really do this.

Dylan Patel

I don’t necessarily think so. I think it’s more that the West doesn’t tell people what they do, right? OpenAI didn’t tell people how sparse GPT-4o was, what the shape size was, or any of these things. But GPT-4o is roughly the same size—slightly smaller than DeepSeek V3—and 4o came out a little bit earlier, if I recall correctly.

Shaun Maguire

So is your view that all 3 of these things have been happening simultaneously at roughly the same rate, and the biggest gains come when you just co-optimize?

Dylan Patel

I would say there have been more gains on the model layer than on the co-optimization layer, the software infrastructure layer, and the hardware layer. But there have been innovations on every layer, and really, the biggest gain—and the beauty of the best labs—is when they co-optimize all 3.

When Anthropic has used many different kinds of hardware, they don’t really do much inference on TPUs. They mostly train on TPUs, and they run a lot of inference on Trainium and GPUs. GPUs are more of a jack of all trades, but they’ve optimized their hardware, they’ve optimized their model, and they’ve optimized everything so they can do that.

OpenAI’s prior models were optimized for Hopper; now they’re more optimized for Blackwell. You step forward through time with these labs, and it’s the same with Google. They’ve optimized Gemini 2 for the TPU v6e, Gemini 3 was, and then the next Gemini that’s coming out is really optimized for TPU v7.

A lot of these things are being co-optimized, and when you pull that model and run it on the old hardware, it’s really not that great. I think a lot of this co-optimization is the most important thing. It’s called software-hardware co-design, and that’s what’s really exciting about what I think my day-to-day is like: You get to look at one layer, but there are all these innovations happening here, and there are all these innovations happening on every layer.

The real breakthrough innovation is when you leapfrog a few layers, co-optimize and co-design them, and all of a sudden you’ve taken what could have been a 2x gain here, a 2x gain here, and a 2x gain here. Instead of being multiplicative to 8x, it’s actually 100x because you’ve optimized across all 3 layers.

That’s what’s really exciting about what you see at the labs, and what you see at a company like NVIDIA, which isn’t co-optimizing on the model layer per se, but is co-optimizing a little bit from the model layer all the way downstream to silicon. Or you look at a company like TSMC: They’re co-optimizing not just fabrication, but everything from the components, consumables, and tools all the way upstream to what the designs of their customers’ chips are telling them.

This is co-optimization across many layers of the abstraction stack.

Shaun Maguire

There will always be bottlenecks somewhere in that optimization, though, that are lagging behind and then need to get pulled forward. If you had to predict at any level of the stack—it can be literally anywhere—what are some of the bottlenecks you’re tracking most acutely over the next year?

Not necessarily in the supply chain, not in scale, but in terms of the actual... It can be in the supply chain too, but is it memory improvements? Is it just scaling?

Dylan Patel

Memory is an easy one that everyone’s talked about, but I’m not going to talk about it from a supply-chain angle. I’m talking about it from a technology angle. Memory capacity and bandwidth have been improving very slowly.

The NAND cell was invented about 25 years ago. The DRAM cell was invented about 40 years ago, and there’s been no major breakthrough in how a NAND cell or a DRAM cell works. Obviously, NAND is a very simple gate. There is stuff that could come down the pipeline that could be hugely innovative.

But even over the last 5 years, all we’ve really done is add more stacks to HBM and make it faster. Actually, there are new innovations coming in the next few years where, instead of stacking the HBM separately from the chip, you stack the memory directly on the chip, and that makes your bandwidth explode. There are interesting companies in that space and interesting proofs of concept that companies are trying to do there.

I think memory bandwidth is one of the biggest bottlenecks. Another one is that, for the history of silicon—basically, for the last 2 decades at least—how many watts a chip is can be easily predicted just by looking at it. For a data-center or desktop chip, it peaks out at 1 watt per square millimeter. So if a chip is 100 square millimeters, generally the power consumption is around 100 watts or a little bit less.

If you look at the newest NVIDIA silicon and the newest TPU silicon, they’re still in that range of 1 watt per square millimeter. Chips are now getting to 1,400 watts. The next generation is 2,000 watts for NVIDIA, with Rubin and such. As you move forward to Rubin Ultra, it’s going to be 4,000 watts or something like that.

But really, that’s just increasing the amount of silicon. What’s exciting is that we’re now finally doing things—and it’s in development right now—where you can actually pump the amount of power into the silicon to be way more than 1 watt per square millimeter. All of a sudden, that means you need less silicon. Obviously, it’s running at higher power and is less efficient in some cases, but you reduce the amount of silicon and you’re able to—

Shaun Maguire

Like, over thermal issues?

Dylan Patel

Thermal issues. There are electrical interference issues. There are all sorts of different issues that crop up, and that’s why it’s a hard engineering problem. That’s why we’ve stuck at about 1 watt. But what’s exciting is that the world is trying to change these things.

In a different part of the supply chain, people will talk about how energy is hard and how we have energy bottlenecks. But there are actually very simple solutions one could think of. Take the millions of diesel engines for trucks that the U.S. has the capacity to make. You can very trivially convert them to use natural gas on the assembly line and then connect them to an electric motor, back-driving it so the electric motor generates electricity rather than causing the rotation of the wheel, for example—doing it in the opposite direction.

Now you’ve generated electricity by pumping gas into something that the U.S. can make millions of. That sounds like a pain to service, because now you have to have hundreds of these on a data-center site. But you can just pull people out of car mechanic shops and have them run around and repair truck engines.

Actually, it’s pretty trivial to do. I don’t want to say it’s trivial; I couldn’t do it.

Shaun Maguire

I think you’re making a really good point, which is that because the West wasn’t really thinking about semiconductors, or even hardware more broadly, over the last 20 or 30 years, we didn’t have much innovation. We’d have the best minds thinking about how to improve these—

Dylan Patel

Why would you want to go work in hardware when you can make ads to ads?

Shaun Maguire

Yeah, exactly. Okay, I’m dying to ask: NVIDIA versus TPU. What are your thoughts?

Dylan Patel

I think everyone wants to pick one or the other for this, but it’s really a function of the fact that, if you look 2 years from now, Google is going to make 10+ million TPUs through their supply chain, and NVIDIA is going to make many more—tens of millions of GPUs.

Both are going to be $100+ billion. Well, Google is going to be $100+ billion of TPUs created a year, and NVIDIA will be $500+ billion, or whatever. I’m not making a specific estimate.

Shaun Maguire

This is not a revenue forecast.

This is just a thought experiment.

Dylan Patel

Yeah. Or research.

Shaun Maguire

You've been media trained.

Dylan Patel

Absolutely. You know, getting ready for the SpaceX IPO.

Are you guys big in SpaceX?

Shaun Maguire

Okay, so that makes sense. We're very lucky to be very large investors.

Dylan Patel

Awesome. So I would say the case of Google TPUs versus NVIDIA GPUs is that they both have points that are really in their favor, right? NVIDIA will be like, “Oh, well, we have switches and we're general purpose.” And TPUs will be like, “Well, we're more optimized, actually more energy efficient, and our network is more optimized for certain types of network architectures.”

8. Sparse vs Dense Models

So you have these counterpoints that both would really get into. I could, with a straight face, argue with you that GPUs are way better than TPUs or that TPUs are way better than GPUs, but it comes down to hardware-software co-design. Actually, given the way OpenAI's models are headed, it would potentially be a terrible decision for them to use TPUs. And given the way Anthropic's and Google's models are headed, it could potentially be a terrible decision for them to train with GPUs.

Shaun Maguire

What's the fundamental difference there?

Dylan Patel

There are various things, right? The size of the matrix-multiply unit is different, as a very simple example. Therefore, the shape of the matrix multiply you do, the attention mechanism you use, the way that attention mechanism is structured, and the way the experts are structured are all different.

Shaun Maguire

So you think OpenAI and Anthropic are converging on very different model architectures?

Dylan Patel

I think they have quite different model architectures. In fact, OpenAI's models are much more sparse, and that has benefits. Anthropic's models are still sparse, but more dense in general, and that has different benefits.

9. Interconnect Shapes Architecture

There are many other things, right? The network topology is another one. NVIDIA's chips are all connected to NVLink switches. Google has no switch. What they've done is enable their interconnect to connect 8,000 chips at super-high bandwidth, but you have to pass through other chips to get there because there's no switch.

There are trade-offs there—positives and negatives—and that influences the model architecture. It's not necessarily that you should claim one is better than the other because, at the end of the day, how do you say that this is better than that when you can't measure them in isolation? It also extends up to the model layer, right?

10. CUDA Moat Is Shifting

Shaun Maguire

I remember, for a long time, thinking that the programmability of NVIDIA and CUDA itself was such a big moat. It seems to me that narrative has changed, at least in my mind, over the last 3–6 months. Model companies no longer care whether they have to write custom kernels for another chip; so be it. They'll work with 4 or 5 chips if they have to.

Claude and Codex are actually quite good at doing a lot of that optimization work. It also isn't as though there are 10,000 model companies that each need programmability. There are on the order of tens of model companies, maybe. So it seems to me that the fundamental premise of tens of thousands of big customers that need CUDA compatibility—that thesis seems to be changing over the last 3–6 months.

Dylan Patel

Certainly, the CUDA moat and the software moat are at least partially disentangled because models are just great at coding, and all software gets commoditized in that case.

11. Ecosystems and Co-Design

I do think there is some level of open source, and what people call the CUDA moat is not actually anything to do with CUDA. It's the fact that DeepSeek, Kimi, Zhipu, Alibaba, Tencent, and Xiaomi—all these companies—have models that are co-designed for GPUs. Therefore, if I want to run them on TPUs, in some cases they don't run really well on TPUs.

Google just has to create its own open-source model ecosystem or open-source models themselves. They have the Gemma models. So you end up with, well, that's not really CUDA as a moat; it's that the downstream product is more optimized for NVIDIA. In these cases, these companies are just open-sourcing them. Nemotron is just open-sourcing it.

The users of those models—for example, the open inference API providers, the RL companies that are trying to take open models and customize them for companies' business use cases—all these different companies are downstream of the fact that, “Okay, I guess I need to use NVIDIA because the ecosystem uses NVIDIA,” even though I don't particularly care about writing CUDA kernels because the models are great at that.

It's the shape of the model: this expert, the d_model is this, and the hidden dimension is this. Therefore, it's better to run on NVIDIA GPUs than it is on TPUs, and vice versa. If Google were to open-source really good models, people would take those models and say, “Oh, wow, these don't run that well on NVIDIA GPUs. I should actually just rent or buy TPUs and do it on there.”

For small teams, you're going to want to use all the open-source software, like vLLM, SGLang, PyTorch, and all that stuff. But the big labs don't necessarily need to use all of that. OpenAI forked PyTorch long ago, and Anthropic and all these other companies don't necessarily rely heavily on the open-source implementations of these things. They forked things or built them on their own already, so they don't need to rely on the open source.

Therefore, now it's more like, “I'll choose the best hardware, and I'll co-design my model and infrastructure software through and through for the hardware that is the best and most cost-efficient.”

Shaun Maguire

And I'll have AI help me write all that software.

What do you think of Cerebras?

Dylan Patel

I think Cerebras is a really innovative company. In some parts of the market, they're really, really good. They have very fast inference, and I think that's a big market. We use fast mode almost exclusively at SemiAnalysis.

Shaun Maguire

By the way, I love how disciplined you've been about accounting for the dollars spent and the ROI on each task. I don't know if that was one exhibit you did or if you do it consistently, but it's awesome analysis.

Dylan Patel

Yeah, we do it pretty diligently, so thank you. That was the Dark GDP article that we wrote. We also track everyone's token spend by day, and if someone's spend has spiked, I'm like, “What did you do?” It's like, “Okay, thank you for telling me. That seems worth it.” Then I continue with my day.

I think fast mode is obviously worth a lot for high-end tasks, right? I could see so many different use cases where super-fast tokens are worth it. I can also see the flip side, where there are a lot of use cases where super-fast tokens aren't needed, and therefore the market won't pay for them. They'll use GPUs and TPUs instead.

I think the big risk for Cerebras is that I mostly think the best models are the ones that you want to use fast mode on. With small models, you might not necessarily use fast mode. I could see that being wrong with financial markets, maybe—like Jane Street high-frequency trading or something like that, or medium-frequency trading.

Ultimately, running really large models at really long context lengths is very difficult on SRAM-based chips like Cerebras and Groq. So now it all of a sudden becomes, what happens if the models get too big? If OpenAI's model is not on the order of hundreds of billions of parameters or low trillions of parameters, but is actually 10+ trillion parameters, I don't think that will fit on Cerebras.

And if that doesn't fit with a long context length—if you have a 1-million-token context length—that makes it really difficult to justify. So far, we've seen the bulk of revenue and usage at the labs be on their best model, even when the model price has gone up.

There's some data that shows that even though Fable just released today, they've had incredible numbers of people switch to Fable and Mythos, sort of that next-tier model, even though it's way more expensive.

Shaun Maguire

And that's volume by dollars, totally? But was that volume by tokens?

Dylan Patel

Well, I guess who cares about volume by tokens? It's about the dollars.

Shaun Maguire

Fair enough.

Dylan Patel

Right. If I don't care that there are 200,000 Mini Coopers or Toyota Camrys sold if Ford F-150s have 5× the ASP and sell only half as many, then the most lucrative market is pickup trucks in America. I'm mostly being facetious, but—

Shaun Maguire

I do think this is one of the things you've done so well, and it differentiates you from almost everyone else: You care so much about the economics in addition to the technology. Very few people bridge those 2 things well.

Dylan Patel

I think it's really fun inside SemiAnalysis because we have 90 people, and a big chunk of them are technologists and engineers across the whole supply chain. Then another big chunk are people who were formerly at hedge funds.

You see these arguments. People are like, “Oh, well, that doesn't matter.” Then someone's like, “Well, but cost.” And then the engineers are like, “No, no, but this technology is the coolest.”

You see this organically fight it out. We’re pretty informal, and given the fact that I was a forum moderator, you can imagine how much we enjoy it.

Shaun Maguire

You don’t wrestle with a pig because a pig enjoys it.

Dylan Patel

Exactly.

12. ROI Debates and Hot Takes

Shaun Maguire

Just on this topic, before going to the next question: are there trigger topics in semis for you? Like, if someone says, “Memory is the bottleneck,” which is such a meme, does that trigger you?

Dylan Patel

I mean, it’s true, but the one that really gets me is when people say, “AI has no ROI.”

Shaun Maguire

It infuriates me, right? There’s the “What’s the ROI?” argument, or people denying model progress. There are these people who say, “Models aren’t getting better. They’re not reasoning. They can’t think. They’re going to dead-end and plateau.” It’s like, bro, the line has been up and to the right in terms of capabilities this entire time. They’re like, “Look, this benchmark didn’t improve because it was at 90 percent. Look at the new benchmark—you saturated it, and now the scores are skyrocketing.” Right?

Dylan Patel

I think that’s more so the issue and challenge. Semis are really complex, and I don’t fault people for lacking an understanding of them. I learn stuff every day about the semiconductor supply chain from people, and I’ve been studying it for arguably 18 years, since I started moderating the forums when I was 12. I’ve been studying it for that long, and it’s all I care about, but there are so many layers of the abstraction stack.

I learned about a new chemical that does $100 million in sales yesterday, and I was like, “Whoa, I didn’t know this one existed or what process it was used in.” You learn about things all the time. It’s like, okay, $100 million in sales in a couple-hundred-billion-dollar industry is whatever, but—

Shaun Maguire

But it’s essential.

Dylan Patel

It’s essential, and actually every chip requires it. It’s like, wow, I guess there are 1,000 process steps. You like semiconductors? Name every process step. Come on.

What I think is the funniest is when people have all the facts in front of them and then get the conclusion completely wrong.

Shaun Maguire

That happens in our job all the time, too.

Dylan Patel

Yeah.

Shaun Maguire

Yeah. I mean, I can’t—I get it. I think my attitude is not to be mad that you do that; it’s to do it as fast as possible.

13. Ten Year Tech Bets

I think the industry is so important. AI is the most important thing in the world right now, and there are so many near-term bottlenecks. We talk a lot about the near term, but are there longer-term things that you’re really excited about, say, on a 10-year time frame? We talked about orbital data centers, but what about silicon photonics? Do you think it’s underrated or overrated on a 10-year time frame? Are there other things that excite you on that time frame?

Dylan Patel

On space, I think space is super crazy and awesome on a 10-year time frame. I’m super excited about space data centers, mining asteroids, and all these things. I’m super excited about the vision of SpaceX. Again, not investment advice before you hop in.

On the semiconductor side, tremendous market movements and tremendous things can happen just when things happen 1 year later or sooner. That’s all technology that, in terms of co-packaged optics, everyone knows is going to happen by the end of the decade. The debate is whether it happens in 2027, 2028, 2029, or 2030, but at some point along there, it’s going to happen.

I think the more interesting thing is that there are companies like Naveen Rao’s company. Did you guys invest in Naveen Rao’s company?

Shaun Maguire

We did.

Dylan Patel

Okay, yeah. He’s trying to innovate on the silicon layer, the software abstraction layer, and the model layer simultaneously, and he fully understands that it’s not a “we’re going to do this in a few years” kind of thing.

Shaun Maguire

It’s not a 2-year time frame.

Dylan Patel

Yeah, it’s not a few-year time frame. It’s a long-term bet. Stuff like that is exciting. It’s like, okay, we’re going to bring analog compute, energy-based models, and all this crazy stuff together at once. That’s exciting. It probably won’t work, but that’s exciting, and I really look forward to—

Shaun Maguire

It definitely won’t work quickly.

Dylan Patel

Yeah, definitely won’t work quickly is what I should say. I believe in Naveen. I met him very early, and I think he’s one of the first people I met in the industry, funnily enough, in 2020 or 2021. Actually, 2020. That says something about him.

In my experience, he’s always trying to help the younger generation and identify talent.

Shaun Maguire

I baited him on the internet. I baited him on the internet.

Dylan Patel

He’s always trying to help the younger generation. He’s trying to identify talent.

Shaun Maguire

He was also so ahead of his time with Mosaic. I remember getting pitched.

Dylan Patel

No, it was 2019. I was still anonymous then, actually. I baited him on the internet, and he started replying. Then I took it to DMs, then took it to a call, and that was the first person who was really important that I talked to in the entire semiconductor industry. Funny.

Shaun Maguire

But yeah, sorry to interrupt. That’s funny. What do you think is the end state of the ecosystem? Do you think every lab and every hyperscaler just has its own chips? Trainium seems like it’s working now, right? Do we end up with every lab and every hyperscaler having its own chips, at least for inference, and then maybe going to NVIDIA or whoever for training? What do you think is the end state?

Dylan Patel

I think everyone will try and then stop trying. Ultimately, supply chains matter, and what technology you can bring in matters. More and more, as the industry gets bigger, supply-chain diversification happens.

Right now, everyone’s chip more or less looks the same. It’s a big logic-compute die in the center, with some HBM on the right and left and on the top and bottom. The top side is networking, and the bottom side is PCIe and other I/O. That’s the exact same structure for Trainium, TPU, and NVIDIA chips. Most of the startups—not Groq and Furiosa, which are doing something weird, but that’s cool—are doing the same thing.

As you step forward, we’re going to get more bifurcation of hardware architecture and model architecture, and therefore people are going to co-optimize them. Some of them will end up in local minima. If this is like gradient descent, people are trying to get to the most optimized solution. Some people will race to a local minimum, and then the question is: how do you leap, how do you scoot back over to the absolute minimum?

To some extent, a general-purpose NVIDIA chip will always be more general-purpose than anyone else’s chip, at least on a parallel AI-compute basis, because NVIDIA has so many customers who care about different things and who will always give them feedback in the design. The minimum will always be better for them, but is that minimum a local minimum? Is the TPU, Trainium, Groq, Cerebras, or whoever’s design optimized awesomely for here, but in the end state, you actually have to go over here, and so they’re the wrong choice?

Shaun Maguire

Maybe they make a great chip. They’re great for a little bit of time, but then they end up being wrong. That’s the real question.

Dylan Patel

I think there will be a big market for general-purpose AI compute. You talk to people at labs, and they don’t even know what architecture they’re going to be doing in a year. They literally don’t know what architecture they’re going to be doing in a year. They have bets, many research bets, and that’s the exciting thing, but they don’t know where it’s going.

Generally, they know what hardware they have, and they’re trying to co-optimize. Ultimately, if a new breakthrough happens in model architecture, it’s like, just replace the attention mechanism with something else. Who knows? All of a sudden, if something happens, the best hardware will change.

Therefore, are people going to make 5-year investments in hardware solely on an ASIC that is more specialized? They’re going to have some bucket of more general-purpose compute. You see this with Google paying $11 an hour per GPU to xAI for H100 GPUs, right? That’s insane. Obviously, compute is limited and so on and so forth, but it’s a very high amount. At the same time, despite the fact that Google has TPUs, there are some questions there: Why do they do that?

Google actually has 3 different design programs for TPUs. They’re making a TPU with Broadcom. That’s a different architecture from the TPU with MediaTek, which is a different architecture from the third one. I won’t disclose the details of that one, which is being developed by Google Research. They’re making different architectures. It’s not just, “Oh, they’re making TPUs with a couple of vendors using the same architecture.” They’re making different architectures, and the third one is very different from the first 2.

I think people recognize that local minima can happen. Therefore, I think everyone will have their own ASIC program. Everyone will deploy billions of dollars of their own ASICs—tens of billions of dollars, and in Google’s case, hundreds of billions of dollars a year on their own ASICs. Ultimately, they’re also going to have workloads that don’t use TPUs. Some of Google’s bets that are not Gemini or DeepMind actually primarily use GPUs.

They don't use TPUs. Some of them also primarily use TPUs, right? It's a broad thing, but maybe for drug discovery or for Waymo, you might not want to use TPUs. I won't say which one it is, but there are different architecture bets and different paths for AI.

AI for science may have different algorithmic patterns than general-intelligence AGI models. I think we'll see diversity continue to proliferate.

Sonya Huang

Yeah, and because the market has gotten so big, niches will be carved out. That makes it possible for companies to have their niche and actually make money, even if the majority of the pie goes to NVIDIA, TPU, and Trainium.

14. Compute Crunch and NeoClouds

Okay, love that. Can we talk about the data center buildout? First, it seems like, by all accounts, if you look at the charts—dollars per compute hour—we're in the middle of a crazy compute crunch. It seems like it's both a demand- and supply-side crunch, right? Demand for long-running agents is skyrocketing, while supply— all these data center buildouts—are delayed. Do you think we're in a compute crunch for the foreseeable future, or do you think it alleviates at some point?

Dylan Patel

Yes. Every quarter, we're deploying vastly more compute than the prior quarter, and there are more data centers built than in the prior quarter. This year, there's going to be 20 gigawatts, even accounting for the delays, and next year there's going to be more than 30 gigawatts, accounting for the delays. Of course, delays happen on everything, right? Anything involving hardware can have a delay. That's just the reality of life.

Are we going to have a compute crunch for the rest of our lives? It depends on what happens with models. But the TAM for Mythos 5 and Fable 5 is not just 2x that of Opus, right? The model is so much better, and it can do so many more tasks, that the TAM for it is way larger than that.

And yet compute in the world did not double in the last 6 months, right? From Opus—or maybe 7 or 8 months since Opus 4.5 launched—to now, 4.6, 4.7, and 4.8 were improvements, but Fable and Mythos were a huge step-function improvement. The world's compute did not double, quadruple, or whatever in that same time frame. But the demand for useful tasks that can be done by AI—the number of useful tasks and the value of them that can be done by AI—has.

So now the question is, what happens? Obviously, Anthropic in Q2 is net-income profitable, excluding stock-based compensation. I think by Q3 they may even be profitable including stock-based compensation. That's how profitable they're getting. Their margins on an Opus token—at least an Opus 4.8 token—are north of 80% at the API price.

They've got a lot of deals where their total corporate gross margin gets clawed down a little bit because of how they do Bedrock deals, Vertex deals, and things like that. But ultimately, their per-token margin is so high. If you have the capability to pay, you can ultimately buy every GPU above the market rate. They also bought GPUs at above-market rates from SpaceX, which is below the rate of Google, but that's because they signed earlier.

That's something that other companies—maybe a venture-backed company or a company that doesn't really have positive margins—can't necessarily do, right? What is the cost-benefit ratio? Every GPU I rent because I'm out of compute capacity, I can immediately turn around and sell tokens on it. Every TPU or Trainium I rent, I can immediately sell tokens on it at a positive margin.

If I'm running a 75% gross margin and I double the cost of the compute, it's fine. I'm still running a 50% gross margin, and spinning up more compute nodes isn't necessarily a human-requiring task for them if they're renting them. Ultimately, it's like, well, my NOI still goes up, right? I'm going to rent GPUs at whatever price. At some level, whatever price I want to pay, I can pay.

Sonya Huang

I have almost the reverse question. At some point, does this compute buildout go bump in the night? Earlier today, I think there was a tweet that Crusoe publicly said one of its customers had asked to halt construction on one of its data center buildouts. It seems like everybody in the ecosystem is so leveraged right now to, “We have to build. We have to go build. We have to build.”

High leverage, high growth makes me very, very nervous as an investor.

Dylan Patel

Wait, hold on. High leverage, high growth means a small amount of equity has huge upside. You're not a debt investor.

Sonya Huang

You're a credit investor—you're an equity investor, right?

Dylan Patel

Let's go.

Look, you've got to go to the school of private equity. Levered buyouts only.

Sonya Huang

I actually come from the school of private equity.

Dylan Patel

Oh, awesome.

Shaun Maguire

She forgot the school. She's been a VC for too long.

Sonya Huang

Yeah. No, I just do revenue multiples. But are you seeing any signs of that? Are you worried about that?

Dylan Patel

I see what you mean, right? That sort of goes back to the model point. Obviously, if the models expand the total economic value of the work—the “Dark GDP” report that we did, which you mentioned earlier—if the work that these models can do does not expand faster than the compute capacity, then that tide turns, right?

Over the last 6 months, that tide has been very much levered in this direction: The models can do more work, or are expanding the TAM of work they can do, faster than compute is increasing. So prices go up.

It's very possible that all of a sudden model progress stops. You talk to anyone at Anthropic or OpenAI—maybe they're drinking the Kool-Aid—but you talk to basically all of them and they're like, “No, no, no, no. Model progress still goes up.”

Ultimately, current methods could stall somewhere. I'm not sure where that would be. It seems like we have a line of sight to model improvement—rapid model improvement. In fact, models are improving faster than they were 6 months ago or a year ago because there's—I wouldn't call it recursive self-improvement, but basically, the models are helping the engineers write all the infrastructure and launch the next model sooner and sooner and sooner.

You've got this pseudo-recursive-self-improvement loop going, and so the models are getting better and better and better faster.

Ultimately, capital is a big problem, which is why Google raised capital. They've got an ungodly amount of cash, right? They own about 5% of the company.

Sonya Huang

I think a little more, but yeah. Maybe. I think at one point they had around 10%.

Dylan Patel

Larry Page invested $1 billion at a $10 billion valuation and got 10% of the company. It got diluted, and all of that. But that was one of the greatest investments of all time. Good job, Larry—the guy.

So they know they have something like $100 billion in the bank that they can sell in 9 months or whatever, once the lockup expires. They have all the gross profit they do, and yet they still modeled that and were like, “We need to raise capital.” So they did an offering, and it's like, that's insane. That tells you how much they think they need to spend.

Capital is a real problem. Meta did announce that they're going to do a raise. The stock tanked. People don't like it, but all these companies are going to raise capital, whether it be debt or equity. At some point, money spigots will have to slow down.

But right now, every GPU that Amazon adds is making higher revenue, or every TPU or Trainium that anyone adds is making gross profit. I do a little bit of a tee-up on this to turn it into a question for you, but—

Sonya Huang

As we talk about this, the thing going through my head is almost an alternative hypothesis for the Crusoe example. I'll use an analogy in oil. In oil, Saudi Arabia has a way lower cost per barrel to produce oil than a lot of other countries. There's also the purity of the oil. Saudi Arabia generally has very low contaminants in its oil, which makes refining easier, and all of this.

The question for me is, when you look at every gigawatt that's being put in the ground—call it the 20 gigawatts coming online today—how much homogeneity do you see in those gigawatts? I don't know what metric you think is right, but are Google's gigawatts 2x more valuable than, say, most neoclouds because they have optical switches, they've been doing it for a long time, and they know how to do power smoothing?

I think this could be the alternative hypothesis: Some of the people who are good at building data centers should just do it to the max because there's so much demand and they're so much better at it. But maybe we're starting to see the early signs of the people who aren't as good at it getting hit a little. I don't know the reality here. I'm just curious how you think about this.

Dylan Patel

So far, there are metrics for this, right? Trainium sells at a sub-$10 billion-per-gigawatt rental rate to Anthropic and OpenAI. GPUs, at least before the craziness of the last 6 months, usually went around $12 billion to $13 billion per gigawatt.

So the rental rate—and this is from a neocloud versus Amazon, even—and now, when Amazon sells GPUs, they'd also be $13 billion or so—

Sonya Huang

My understanding of that is also that Amazon subsidized those numbers a little bit. I actually think the numbers were even lower—I think the disparity was even more.

Dylan Patel

It's less than 10.

It's less than 10, but there's some weirdness, basically.

Shaun Maguire

And look, my understanding, obviously, is that Anthropic played a big role in making Trainium useful in terms of writing all the libraries, et cetera.

Dylan Patel

Everything I hear is that Trainium's really freaking good hardware, and it's getting way better. Obviously, Anthropic is now using it a lot, so hopefully we would see that price go up. Per the deal they did, there was actually a floor mechanism in it: if it didn't do well, it would be cheaper, to the point where it was cancellable, and if it did really well, the price would be higher. But effectively, less than 10 is where Trainium shakes out, whereas the xAI GPU deal, again, was $25 billion or something crazy per gigawatt, or $25 million per megawatt as a yearly rental rate with Google. I was like, that's a crazy divergence.

Now, obviously, if Amazon was selling Trainium today, it'd probably be more expensive than 10 because of the compute shortages. But you do see this already in the sense that, with data centers, oftentimes a rental price of a data center—if you're doing colocation, not compute in there, but just power—gets priced generally in dollars per kilowatt per month.

They used to be $60 per kilowatt per month, and now you see things transacting anywhere from $120 to $160. But with different-quality data centers, I've seen data centers go as high as $200 when the customer doesn't have such a great credit rating and the data center is a pretty good one. I've seen stuff go as low as $100 still, or in India as low as $80, because the grid isn't reliable, the internet connection isn't great, and it's a pretty mid-level data center—but at least it's a data center.

In the case of data center construction, usually the pitfalls are that they just fail. There are a lot of people who claim they're going to build a data center—there are like 4 guys, and they're like, “Yeah, we bought some turbines. We put the money down for them. We're going to build a data center.” Then they get delayed, delayed, delayed, and fail.

You have to probability-weight and time-weight the time lag, and account for the teams that suck versus the teams that don't. Our data center model does that. We track every data center and try to do this for every single one based on the equipment they're using and all these things.

One of the things you mentioned about Google is that, in a gigawatt data center, they'll actually put 1.5 gigawatts of hardware. Because they understand the workload so well, they're able to slosh the power around. Instead of constantly having a gigawatt of compute—which typically runs at 60% or 70% utilization in terms of power consumption, not utilization of the hardware, because someone's always renting it—they're now running it so that 60% to 70% means it's at a gigawatt and they're using the full gigawatt.

You see people doing deals, including Google with utilities, where they're like, “Oh, well, I know this grid can sustainably take a gigawatt, but except for 3 days of the year, you can actually do 2 gigawatts. So give me 2 gigawatts and then just tell me to turn off.” And so they'll do that.

These sorts of tricks require supreme management of workload, backup power, and all these things, plus generators on-site, to figure out how to actually keep 2 gigawatts sustainably. When people do this, they're able to charge more. Whether it be, “I'm actually selling 2 gigawatts despite only having 1 gigawatt because those 3 days I'm able to deal with via battery, gas, et cetera,” or, “I figured out how to build power on-site. Now I have a gigawatt where no one else does, and so I'm able to do it quickly.”

It's not necessarily transacting for a higher price; it's that I'm selling more gigawatts. Sometimes there are levers where selling more gigawatts means that each gigawatt is selling at a different price. I think, on the data center and energy layer, it's more about just having it versus not, and then whether that is delayed or not. It's more binary.

But on the compute side, I do think there's a lot more interesting work there, right? A gigawatt given to Anthropic is objectively worth more revenue than a gigawatt given to OpenAI. It seems that both of them could sell every gigawatt that they have right now, given rate-limit problems, token max limits, and all these sorts of things at OpenAI and Anthropic. Especially since Codex 5.5 came out, it's much better.

Shaun Maguire

My guess—my suspicion—is that they probably make better use of the hardware than most people. I think people underestimate how much networking experience they have from Starlink in particular, and also how much power-management experience they have via Tesla.

Dylan Patel

Yeah, people like Brett Mayo are incredible.

Shaun Maguire

Pretty good.

Dylan Patel

Yeah.

Shaun Maguire

And so I think that, for me, that's probably the thing that might be missing from a lot of the analysis people are doing. I don't actually know the answer, but I think that might be what's missing.

Dylan Patel

I think it's also the fact that when CoreWeave builds a gigawatt, even though its GPU compute is objectively better than Amazon's, Google's, or Microsoft's in terms of performance—we've tested the performance and reliability—the problem is that CoreWeave sells it 6 months before they have it up. They need to turn around and take the paper they signed to get debt, with that credit backing, so they can actually pay for the purchase order they've already issued—for the order they've already issued.

Whereas xAI was like, “No, no, no. This is running now. Buy it.” It's a big discrepancy when you have a balance sheet to do that versus not, and that also helps your revenue per megawatt be much higher.

Shaun Maguire

Why does the neocloud opportunity even exist? If you had asked me 5 years ago, I would have said the hyperscalers were going to own this. You mentioned just now that CoreWeave has better performance than the hyperscalers. Why does this opportunity exist, maybe at the macro level and then at the execution level?

Dylan Patel

Yeah. So in 2023, I wrote a report that had Amazon really hating me. It was called “Amazon Cloud Crisis.” I talked about how Amazon was the best cloud because they had their Nitro NICs, which offered tenant isolation. The hypervisor ran on the NIC, and then you could sell all the cores.

They had custom SSDs that they made. They'd buy the raw NAND and build their own SSDs, so they had lower costs, and they had their custom Graviton CPUs that drove down the cost per core. They had all these things that enabled them to sell more cores, have better security, and have good networking and storage, but this was all for the traditional CPU and traditional cloud world.

In the AI cloud, a lot of this stuff hurt performance. These Nitro NICs were bad for performance and still have worse performance, although they've caught up a lot because they've had a couple of iterations to improve them. They're still worse for performance.

A lot of the security stuff doesn't matter because it's not like I'm time-slicing users or splitting a socket among many users. No one rents a single GPU in an 8-GPU server. No one rents a single GPU in a 72-GPU rack; they rent the whole rack, and in fact, they rent many of the racks.

There's no, “I rent for 6 hours and give it back.” Everyone has these long-term contracts. The mechanics of the GPU rental market meant that a lot of the expertise of the hyperscalers fell away, and a lot of the expertise that they did have was actually detrimental.

For Google and Amazon, they had custom networks that were better for traditional CPUs and for the stuff they were doing, but actually worked worse for AI. In other cases, Microsoft would save money by building its own data centers, but its data center teams actually weren't that great. When building was predictable, it was fine; when it came time to double the forecast for the year, they fell on their face and had to go get a bunch of neocloud capacity.

I think performance—I think time to market's another one. These massive organizations—no one's getting rich from building this data center faster. But you look at Crusoe, for example: Chase and all the other people on the team. I was going to name some people on the team, but I'd rather not. These people are getting rich if they deliver this compute faster. They're hyper-levered equity owners.

Shaun Maguire

Hey, look, they're also all coming from Bitcoin. And you know, you're not supposed to say that.

Dylan Patel

Their main data center guy came from Microsoft.

Shaun Maguire

I don't know. I'm just teasing. But you learn a lot when you're in a high-fluctuation market.

What do you think Jensen was playing for in this chess game?

Dylan Patel

Jensen absolutely hates a world where all the hyperscalers have all the power.

There's a reason he's blowing money on random AI labs that I don't even know if it makes sense to, but he's blowing money and pumping them up and going to everyone around the world and saying, “You should invest in this company,” because he wants to create a multipolar world. That's why he loves Chinese labs: He wants to create a multipolar world. A world where OpenAI, Anthropic, and Google models are the only models is one in which he's screwed.

Shaun Maguire

Yep.

Dylan Patel

Right. A world in which the hyperscalers are the only ones building compute is one he's screwed in.

Shaun Maguire

Yeah.

Dylan Patel

And so, of course, he needs to point the allocation gun at neoclouds, help backstop their clusters, and do anything and everything because, while today a GPU sold to Crusoe, a GPU sold to CoreWeave, and a GPU sold to Google and Amazon are all the same price for him, 5 years from now, Crusoe and CoreWeave existing means Google TPU will be weaker, and it means Amazon Trainium will be weaker. More inference being done with non-closed-model labs is better for the firm.

So I think the neocloud ecosystem is these people that are Wild West. These neoclouds as well—a lot of them have investments from NVIDIA. It's the Wild West. Some will fail, many will fail, but some will emerge as really great teams, whether it be, oddly, Crusoe, who's a bunch of crypto guys who then started building data centers and doing flared-gas stuff, or CoreWeave, who initially was a bunch of New York hedge fund guys and crypto guys, but then they built. There were a lot of people who didn't bubble up like them, started around the same time, and just failed. Right. And so I think—

Shaun Maguire

I gotta say, both those teams are phenomenal. They deserve a lot of credit, and that's your point.

Dylan Patel

Yeah, my point is you throw a bunch of bait into the water, and the best fish will figure it out and survive, right? In the same way with the neoclouds, he hopes the neolabs will do the same as well. We'll see if any of the neolabs really bubble up, but Thinking Machines has a few hundred million dollars of ARR, right? That's pretty impressive, even though in the media it's, “Oh, they've lost all this talent.” It's like, well, but Tinker is doing a few hundred million dollars of ARR. That's pretty impressive for a product that's less than 6 months old or whatever.

We hope the same happens to other neolabs, and so he wants a multipolar world.

Shaun Maguire

Truly, congratulations on the success. Thank you. Just the last thing I'll say is, I've seen a little bit of this. I think the public can probably tell from listening to you how hard you work, but it's clear you've just been working your ass off for more than a decade, and it led to the last few years of being in the right place at the right time. It's unbelievable what you've accomplished, and I know it's just the beginning.

Dylan Patel

Thank you so much.

Shaun Maguire

Thank you for doing this.

Awesome.