Dylan Patel——两家实验室很快将控制全球大部分劳动力
- Dylan Patel 的核心预测是:Anthropic 和 OpenAI 获得的年度新增算力占比,将从今年约30%升至明年的40%–50%;“到明年年底,新增算力的一半已经流向 Anthropic 和 OpenAI。” OpenAI 今年初拥有2 GW算力,Anthropic 不到2 GW;到年底两家都将超过5 GW。由于 GB300、TPU v7 和 Trainium 3 的每瓦性能提升3–5倍,如果趋势延续,到2028年底两家实验室可能控制全球大部分可用 FLOPs。
- 单位经济性已经从负毛利翻转为每兆瓦3–5倍的收入差额。 算力成本为1000万–1500万美元/兆瓦;Anthropic 的收入最高已达到5000万美元/兆瓦,并于Q2开始实现盈利;市场认为 OpenAI 依靠 Codex 和 GPT-5.6,最早可能在Q3某个时点实现盈利。其飞轮机制是:“如果我在推理容量上花10美元,实际上能创造50美元收入,然后我就能把这部分利润全部增量投入训练。”
- 算力本身也必须重新定价:任何人在1000万–1500万美元/兆瓦的价格上部署一台 GB300 机架,再通过 Kimi、vLLM 或 SGLang 接入 OpenRouter,都能赚钱,因此实验室可能需要支付2500万–5000万美元/兆瓦,才能拿下70%–80%的供给。 SpaceX 和 Meta 正把算力囤在资产负债表上,可以在内部使用与高价卖给实验室之间做选择。即便如此,Dylan 预计到明年年底,大部分算力交易价格仍将低于200亿美元/GW,因为算力通常在建成前就已签约并完成融资。
- Dylan 的非共识判断是:随着时间推移,实验室会把更少算力用于推理,把更多算力投入研究和训练,包括打造 AGI。 Anthropic 每月新增算力持续上升,但1月收入激增后,月度新增收入已经趋于平台期,这意味着边际产能中投入研发的比例在提高。当前预算大致为50%研究、10%开发、40%推理;Mythos 的预训练约使用不到200 MW,持续约2个月,而 RL 在任何单一地点或时点的用量更低,但按顺序累计的总算力可能更高。
- 监管和部署约束,而不只是模型能力本身,可能限制每兆瓦收入的上限。 Dylan 提到,OpenAI 表示将暂停训练2周;Anthropic 也没有发布其安全评估认为是 Model 2 的模型,外界普遍认为这就是下一代 Mythos。Dwarkesh 称 Mythos 已被削弱,且无法提供给“我们”做推理优化;Dylan 则说 Astra 尚未在内部大规模部署。如果没有这些约束,Dylan 认为实验室可以实现每兆瓦1亿美元以上的收入,并支付5000万美元/兆瓦。
- 出口管制迄今扩大了中美算力差距:目前中国仅占新部署数据中心 AI 功率的不到10%,到2028年可能拥有不超过30 GW算力,而且芯片性能更差。 Dwarkesh 推测,即使2029年中国拥有主要由国产芯片构成的50 GW算力,实际价值可能也只相当于约20 GW美国水平的算力。Dylan 认为50 GW是合理预期,同时强调出口管制、MATCH Act 和国产设备生产仍然重要;Dwarkesh 则指出,如果 AI 起飞速度更慢,中国就有机会大幅追赶。
- 融资算式指向一次利率冲击:Dylan 的模型显示,2024–2029年资本开支约为11万亿美元,其中约6万亿美元由现金支付,超过5万亿美元通过信贷融资。 Meta 最近以5%–6%的利率融资,并且愿意支付8%;如果增量债务的20%利率仍低于向 SpaceX 租算力的成本,Anthropic 也会接受。利差上升将重估整个经济,压垮非 AI 股票的估值倍数,迫使高债务国家走向违约,并可能最终把2030年代的利率推升至两位数。
- 双方都认为中心化可能是最终结局,但尚未找到解决方案。 Dylan 估计,前沿 AI 的有效劳动力规模每年增长约10倍,因此到本十年末,一个实验室拥有的等效劳动力可能超过地球人口。Dwarkesh 强调训练的规模经济和稀缺性加价;Dylan 则补充了部署数据和 RSI。Dylan 说:“我不信任政府,也不信任 Dario,也不信任 Sam。”他提出应思考一个去中心化的后 AGI 未来,并总结道:“有80,000个世界,只有其中1个世界里,Anthropic 不拥有整个世界。”(“There’s 80,000 worlds and in only one of them, Anthropic doesn’t own the whole world.”)
1. 实验室经济性从风险资本输血的亏损,翻转为每兆瓦5000万美元的机器
- Dylan 的基准判断是:今年上线的算力中,约1/3最终将服务于 OpenAI 和 Anthropic;今年 AI 资本开支“略高于1万亿美元”,2028年将超过2万亿美元。根据已经签署的合同,实验室年度支出将从数百亿美元一路扩大,甚至在本十年末逼近“每年数万亿美元”。
- 毛利翻转才是核心故事:基础算力成本为1000万–1500万美元/兆瓦。OpenAI 在 NVIDIA Hopper 上运行 GPT-4 时毛利为负,但 GPT-5.6、Opus 5 和 Fable 5 的变现能力已远超增量成本——Anthropic 的收入最高达到“5000万美元/兆瓦”。Anthropic 于Q2开始实现盈利;市场认为 OpenAI 依靠 Codex 和 GPT-5.6,最早可能在Q3某个时点实现盈利。
- 自我融资飞轮的原话是:“如果我在推理容量上花10美元,实际上能创造50美元收入,然后我就能把这部分利润全部增量投入训练。”新资本仍会流入,但增长越来越多由收入而非资本注入驱动。
2. 两家实验室将在明年年底拿走全球新增算力的一半
- OpenAI 今年初拥有2 GW算力,Anthropic 不到2 GW;到年底两家都将超过5 GW。Dylan 认为这意味着算力规模增长3–4倍,占今年新增算力约30%。已经签署的交易则将这一比例推升至明年的40%–50%:“真正到了明年年底,新增算力的一半已经流向 Anthropic 和 OpenAI。”
- 全球今年新增算力约30 GW,明年50 GW,2028年70 GW,2029年90–100 GW;但采用 GB300、TPU v7 和 Trainium 3 的新增功率,每瓦性能高出3–5倍。因此,如果趋势延续,到2028年底,两家实验室可能“单独控制全球大部分可用 FLOPs”。
- 算力建设者正在变化:SpaceX 将成为主要参与者,把算力租给出价最高的一方;OpenAI 正在设计自研芯片,Anthropic 则采购 Google TPU,并通过 Fluidstack 部署。会计口径上,Amazon 通过 Bedrock 提供 Anthropic 模型服务,在 Dylan 的框架中也算 Anthropic 算力,因为这部分最终计入 Anthropic 收入。
3. 100倍晶圆厂套利,以及镜片瓶颈为何缓慢解除
- Dwarkesh 基于 Dylan 的晶圆厂设备模型测算:1 GW Vera Rubins 需要约55,000片 N3 晶圆、6,000片 N5 晶圆和170,000片 DRAM 晶圆。模型估算设备投资约30亿–40亿美元;若加上晶圆厂洁净室、厂房主体及相关资本开支,总额约60亿美元,才能实现每年生产1 GW。
- 每1 GW每年可产生约1000亿美元收入。按5年周期计算,第1 GW获得5年收入,第2 GW获得4年,以此类推;约60亿美元晶圆厂资本开支最终可以带来超过1万亿美元的终端 AI 收入。即使 Dwarkesh 保守地扣除一半归其他参与者所有,仍然存在约100倍的差距。
- 对于“资本主义难道不会修复这一套利吗”的问题,Dylan 的回答是:“这是一根鞭子,鞭梢很久才能收到信号。”Carl Zeiss 直到最近才接受现实:到2030年,它需要为每年100台 EUV 设备准备足够的镜片——“建起来就是要花这么久”。他的套利逻辑是:任何能以4亿美元买到 ASML EUV 设备的人都应该买下,再以“超过10亿美元”的价格转售。
- 整条产业链不会立即同步扩张:明年实验室收入将达到数千亿美元,但全行业资本开支约为2万亿美元,实验室还无法靠现金流独自为整个扩张周期融资。Dylan 不认为今年、明年或后年会出现类似私募股权的全栈式扩张,因为“这个世界受资本约束”。
4. 算力重新定价:每兆瓦1000万–1500万美元人人都能赚钱,实验室可能支付2500万–5000万美元
- 算力价格底部已经被验证:“去买一台 GB300 机架,下载 Kimi 权重,再下载 vLLM 或 SGLang……把它放到 OpenRouter 上……你开始获得的收入就会超过算力成本。”这已经在推动1000万–1500万美元/兆瓦的价格上行;如果到2028年全球合计达到100 GW,实验室可能需要支付2500万、4000万,甚至5000万美元/兆瓦。
- SpaceX 利用了“先签合同”的融资模式,在获得通常所需的客户承诺之前,就提前准备好了可用算力。Dylan 认为,Elon 可以按2500万–4000万美元/兆瓦出售容量,并在1年内收回全部资本开支。Dwarkesh 则举了另一个例子:SpaceX 可能以400亿美元/GW的价格向 Google 出售“B300 或类似产品”。
- Meta 和 SpaceX 是“唯一有可能成为第3名”的候选者,因为它们把算力囤在真实的资产负债表上,既可以内部使用,也可以按高利润率出售给 Anthropic 或 OpenAI。
- 尽管如此,Dylan 对短期市场仍保持克制:即使到明年年底,大部分算力交易价格仍低于200亿美元/GW,因为多数算力在建成前就已签约;客户承诺为信贷市场融资提供了基础。供应链将以“甩鞭效应”重新平衡:存储和基板价格反应较快,而 TSMC 的调整更慢。
5. 价值最终流向哪里:用户,而非实验室
- 当前产业链中,终端用户拿走的价值最多。Jane Street 通过与 OpenAI 签署 GPT-5.6 Ultrafast mode 独家合约,从购买的 token 中获得的价值“远远远远”高于 Anthropic 获得的利润;Meta 此前一度被传占 Anthropic 业务的最高10%,但它从广告算法和用户参与度提升中获得的变现远超支付给 Anthropic 的费用。到目前为止,应用层创造的价值非常有限。
- 一年前,模型层还依靠 VC 资金维持负毛利,价值主要被硬件供应链拿走。2023年,存储公司虽然掌握 HBM 这一关键环节,却几乎没赚到钱;如今 Dylan 认为,TSMC 拿走的价值反而少于存储公司。“价值捕获已经大幅换了位置。”
- 一个用于理解收入上限的直觉是:1 GW算力如果能够维持100万名白领,每人的年薪为10万美元,对应的经济价值就是1000亿美元——“这其实低得出奇”。完整 AGI 可能意味着“每1 GW创造数千亿美元”的价值。
6. 监管与部署约束可能封顶每兆瓦收入
- Dylan 提到,OpenAI 表示将暂停训练模型2周;Anthropic 也没有发布其安全评估认为是 Model 2 的模型,外界普遍认为这就是下一版 Mythos。Dylan 更广泛的抱怨是:“世界上现存的最佳模型是在2月训练出来的。”
- Dwarkesh 称 Mythos 已被削弱,无法提供给“我们”用于推理优化及其他用途。另一方面,Dylan 说 Astra 甚至没有在内部大规模部署。这里的区别很重要:访谈记录并没有证明 Mythos 在内部不可用。
- Dylan 对2027年底收入的预测带有很强的条件性:每兆瓦5000万美元以上,混合口径可能达到7000万–8000万美元,“甚至更高”。Dwarkesh 则根据近期进展外推,追问如果到2027年底出现一次 GPT-4o 到 Mythos 2 级别的跃迁,收入会如何变化,同时指出 Mythos 和 Mythos 2 都尚未完整发布。
- 在不受约束的情景下,Dylan 的原话是:“在一个安全不重要的世界里……他们可以开始实现每兆瓦1亿美元甚至更高的收入,也可以支付5000万美元/兆瓦。其他人没有任何合乎逻辑的理由再拿算力做别的事,只能说:‘Dario,请把我的算力全部拿走。’”供给端限制也在拖慢扩张:纽约州禁止数据中心,得州实施暂停措施,俄亥俄州则提出要求数据中心为一定半径范围内的物业缴税。
7. 非共识判断:推理占比收缩——边际兆瓦用于打造 AGI
- Dylan 明确逆市场共识判断:“实验室将把越来越少的算力用于推理……这是非常非共识的观点。”当每兆瓦可变现6000万–7000万美元时,实验室可以选择分红和回购,也可以用这些容量打造 AGI。Dylan 认为,对 Anthropic 和 OpenAI 及其董事会而言,显而易见的选择都是后者,因为后者更赚钱。
- 这一趋势已经开始:Anthropic 每月新增算力,但收入增量在1月激增后便趋于平台期。“它们获得的边际兆瓦中,投入研发的比例高于投入推理的比例。”
- 当前预算结构约为50%研究、10%开发、40%推理。Mythos 预训练使用不到200 MW,持续约2个月;RL 在任何单一地点或时点的用量更低,但按时间顺序累计的总算力可能更高。多地点协同、共址部署和发布限制,使单次训练规模受到约束,而其余算力则被研究任务消耗。自动化研究员和持续学习可能进一步把算力结构推向训练。
8. 出口管制迄今扩大了中美算力差距
- 自2022年以来,格局已经逆转:2022年美国新增算力约占全球的45%–50%,如今占新部署功率的约70%。中国目前占新部署数据中心 AI 功率的不到10%,依靠走私芯片、TSMC 原本以为并非供给 Huawei 的客户最终转给 Huawei 的芯片,以及 Samsung 出口的 HBM。
- SMIC 和 CXMT 的晶圆厂将在2027–2028年开始上线,国产芯片年产量将达到数百万颗,单在2028年就新增约5–10 GW国产芯片。访谈中的预期是,这些芯片性能将弱于 NVIDIA 或 Google 在2028年的芯片,因此单看 GW 数量会高估中国的有效算力。
- 在 Jensen 访谈中为合作可能性做了最强论证后,Dwarkesh 改变了看法:“我之前没意识到算力状况已经像你说的这么糟……我认为这其实是一个值得注意的成功。”中国头部实验室合计最多只有100–200 MW,ByteDance Seed 是例外;相比之下,Anthropic 到年底将拥有超过5 GW。
- Dylan 认为,2029年中国拥有50 GW算力“完全合理”。Dwarkesh 推测,如果其中大部分来自国产芯片,50 GW可能只相当于约20 GW美国水平的算力。MATCH Act、进一步的设备管制以及中国建设国产设备的能力仍然关键;Dylan 说,中国半导体补贴超过世界其他地区的总和。Dwarkesh 补充,如果 AI 起飞速度更慢,中国将有机会大幅追赶。
9. 11万亿美元资本开支、超过5万亿美元信贷,以及第二次 Volcker 冲击
- Dylan 的模型显示,2024–2029年资本开支约为11万亿美元,其中约6万亿美元由现金支付,超过5万亿美元通过信贷融资。仅看关键 IT 支出会低估总额:按当前价格计算,100 GW每年就需要约5万亿美元;如果把预先建设的电厂、数据中心和下游产能算进去,总额可能达到7万亿–10万亿美元。电厂大致是30年期资产,数据中心则是15–20年期资产。
- 超大规模云厂商已经通过现金流、转向资本开支的回购资金和债务,为大部分建设提供融资。Google、Meta、Amazon,以及最终可能包括 Microsoft,都有能力举债;半导体公司、基础设施投资者和更广泛的信贷市场也会成为额外资金来源。回购可以成为融资来源,但并不是所有公司的回购计划都会停止。
- 利率传导机制是这样的:Meta 最近以5%–6%的利率融资,并且“愿意支付8%”;Anthropic 愿意为新增数十亿美元债务支付20%,因为这仍然比以500亿美元/GW的价格向 SpaceX 租容量便宜。利差上升250个基点会重估所有资产——银行的负债重定价速度快于资产,因而遭受损失;更高的贴现率则会击穿非 AI 股票估值:“我为什么要为 Johnson & Johnson 支付这么高的价格?”
- Dwarkesh 讨论主权债务时,Dylan 调侃他在讲“你上个月才学会的东西”:企业所得税占联邦收入不到10%,而自动化可能压缩工资税和个人所得税。在 Dwarkesh 的假设中,利率上升1个百分点,5年后偿债支出占税收收入的比例会从20%升至25%;上升5个百分点则会超过40%,如果政府继续每年借款2万亿美元,最终将超过60%。他认为,只要对美国数据中心征税,美国仍可能安然无恙,但 Pakistan 和 Nigeria 可能严重暴露。
- Basil Halperin 提出的“第二次 Volcker 冲击”类比,指向1980年代模式的重演:当年 Paul Volcker 将实际利率提高至约8%,约40个国家、主要是拉丁美洲国家发生违约。通过 Damon Binder 的投入产出研究延伸到奇点情景:如果 AI 让劳动力规模翻倍,经济最终可能每年翻倍,从而把2030年代利率推向“两位数”。
- Dylan 对市场的补充判断是:如果你真的彻底相信 AI 叙事,“经济中的一切都应该按2倍或3倍盈利交易”。存储股未必还会再涨10倍;市值约1.5万亿美元的 Meta 被他说成“荒谬”,因为它的现金流和囤积的算力可能值更多。
10. 一切力量都把系统推向中心化——信任问题无解
- Dylan 估计,前沿 FLOPs 每年增长4–5倍,而实现同一能力所需的算力每年下降3倍,这意味着有效 AI 人口每年增长约10倍。OpenAI 可能从今年的1000万名 AI 劳动力,增长到明年的1亿名、后年的10亿名;到本十年末,一个实验室拥有的等效劳动力超过地球人口并非不可能。
- Dwarkesh 指出规模经济所在:训练成本可以摊薄到数十亿次会话中,而算力不足的领先者可以收取更高的稀缺性加价。Dylan 则补充,更广泛的部署会提供更多学习数据,最强模型还可以通过 RSI 帮助生产下一代最强模型。
- Dylan 提到 Gavin Baker 与 Dario 的争论:如果相信 RSI,也相信实验室拥有更强的算力变现能力,那么中心化就会成为必然结果。AGI 的限制因素不是研究工程师能否“拧动齿轮”,而是“世界其他地方允许它做到什么程度”。
- 2030年暂停发布6个月的风险由 Dylan 而非 Dwarkesh 提出:在这6个月里,某个实验室可能在内部进行递归自我改进,而公众仍然落后数年。他同时表示,慢速起飞仍有可能——“至少,这是我的希望”——因为监管、融资和部署都存在约束。
- Dylan 提议开展一项思想工程,寻找一个去中心化、让更广泛人群拥有能力的后 AGI 未来,同时认真对待规模经济。他说自己不信任“政府”、Dario 或 Sam。他关于用户拿走大部分价值的判断,明确被称为自己的“自我安慰”:Jane Street 可能捕获每兆瓦3亿–5亿美元的价值,而 Anthropic 捕获1亿美元/兆瓦;但如果内部 AI 研究的价值更高,Anthropic 就有动力把算力留在内部。最后的问题仍未解决:“有80,000个世界,只有其中1个世界里,Anthropic 不拥有整个世界。”
Okay, I’m back with Dylan Patel, founder of SemiAnalysis. Our version of a family Thanksgiving dinner is a regular yearly podcast. But we’re not actually related. Don’t tell the people this; it will destroy the myth.
Basically, where the world economy is headed is more and more becoming a function of where lab economics are headed, where the compute market is headed, et cetera. I want to understand where the crazy future ends up within a few years. But let’s start with where we are today. Walk me through lab compute and lab revenue right now, and maybe project out a year or two.
When we go back to last year, even at the end of the year, most of GDP growth in America was just AI infrastructure. As we look toward this year, about a third of the compute coming online is for the labs—for OpenAI and Anthropic. It may be built by others and then rented to them, but the end customer is them.
As we go forward into the future, the numbers for compute are ballooning. We’re at a little bit over a trillion dollars of CapEx this year. As we go out into ’28, it’s going to be more than $2 trillion. The labs are also taking an increasing percentage of this.
So ultimately, you’ve got a very interesting situation where the labs are going from companies that spend tens of billions of dollars a year to hundreds of billions of dollars a year, to being forecast to spend trillions of dollars a year even toward the end of the decade. This is at least what some of the contracts they’ve begun signing with their partners imply. This requires a big reshaping of what happens with their economics.
Up until now, they have been companies that mostly lost money. Anthropic started turning a profit in Q2. It’s believed that at some point in Q3, OpenAI could start turning a profit, even, with the bigger rise of Codex and GPT-5.6 and all this. But if we go back a year ago, all the money they had was venture-funded losses.
If we go back to even the beginning of this year, it was venture-funded losses. They’ve now turned the corner and are actually starting to profit. That doesn’t mean they’re not taking in new capital. The new capital is still coming in to accelerate the growth further. But ultimately, more and more of their business is being funded off of their own revenue rather than capital injections into them.
Over the last year and a half, their margins have really skyrocketed. The base cost of compute tends to be around $10 or $13 or $15 million per megawatt. The most interesting aspect about what’s happening now is this: Before, if they served a model—GPT-4 being served on NVIDIA Hopper GPUs—it was generating negative gross margin for OpenAI.
But now, when OpenAI serves GPT-5.6 or Anthropic serves Opus 5 or Fable 5, their revenue generation has passed well beyond the incremental $10–15 million per megawatt. In the case of Anthropic, the revenue has gone as high as $50 million per megawatt.
What that now enables them to do is say, “Hey, if I spend 10 bucks on inference capacity, I actually generate 50 bucks of revenue, and then I can turn around and incrementally spend all of that profit on training.”
One thing I’m very interested in understanding is how you see the centralization of compute happening at the labs, or the relative ratio of compute that goes to the world versus the labs. If you say right now a third of marginal compute is going to the labs, by when is over half of the incremental compute in the world going to the labs? By what point do the labs have basically a vast majority of the world’s compute?
At the beginning of this year, OpenAI started at 2 gigawatts and Anthropic at less than 2. By the end of this year, they’re both above 5. So they’ve 3–4×ed their compute as a whole. When you look at the incremental compute added, that’s about 30% of the compute added this year.
As we step forward to next year, given what’s already been signed, penned, and inked, you’ve got something even more dramatic. Anthropic and OpenAI are taking as much as 40%–50% of compute next year. This centralization doesn’t look like it’s slowing down or stopping. In fact, it looks like it’s only accelerating.
Who’s building that compute for them will change. Next year, a big new entrant is, for example, SpaceX, which is building a ton of compute. They’re actively going to lease quite a bit of it to Anthropic and OpenAI, most likely, because they’re the ones who have the marginal capability to pay the highest price.
In addition, OpenAI and Anthropic are also starting to build their own compute—OpenAI with their own chips, Anthropic with TPUs that they’re purchasing from Google and deploying with Fluidstack. So you ask, “Hey, when does half of the world’s incremental new compute go to just OpenAI and Anthropic?” It’s really by the end of next year when half of the incremental compute is already going to Anthropic and OpenAI.
Because compute is growing so fast, incremental compute is going to be basically most of compute. So it’s very soon—you’re saying maybe within a year and a half or two years—that most of the world’s compute is owned by two labs, or at least is serving the demand from two labs.
There’s this trend where maybe world compute in gigawatts doubles every year, but the compute at the frontier labs triples every single year. If you keep the current trend going, it goes from 2 at the beginning of this year to close to 6 at the end of this year, just multiplying out by 3. It’s 18 by the end of 2027, 54 by the end of 2028.
Are you thinking, “Okay, at that point, they simply can’t continue tripling, given the amount of world compute”? How do you see the world compute situation over the next few years?
If the incremental compute this year adds 30 gigawatts, next year 50 gigawatts, and the year after that roughly 70, you end up with this really interesting phenomenon. A new watt deployed this year is significantly more efficient than the watts deployed two years ago.
A humongous percentage of the world’s compute was deployed this year. Even though it didn’t double the number of watts deployed, I’m deploying GB300s, TPU v7s, and Trainium 3s, which are way, way, way more efficient. They’re 3–5× more performance per watt than the prior-generation chips.
So ultimately, you’ve got a huge ladder here. If Anthropic and OpenAI take on 45% of compute next year, you’ve got them, by December ’27, having taken on half of the world’s incremental new compute. But that half of the world’s new incremental compute is actually at a higher performance than everything else before it. So you’ve got another multiplier on that.
By the time you’re toward the end of 2028, if this trend continues—and I see nothing that’s stopping it—you’ve got them just controlling most of the usable FLOPs in the world on their own.
1. 6 billion in fab capex enables $1t+ of end revenue
The thing I’m confused about is why you think we only add 80 gigawatts in 2028 if we enter a world in which the value of compute increases so much. That’s the upper bound, by the way—the “I’m so fucking bullish” case.
Okay, let’s do some chain of thought here. When I interviewed you a few months ago, you said that in order to make a gigawatt of, I think, Vera Rubins, you need 55,000 N3 wafers, 6,000 N5 wafers, and 170,000 DRAM wafers. I don’t know if those numbers have changed.
I’m going to troll you, but the way you said “wafers” was so fucking Indian. “Vafers.”
By the way, when we first moved to the US, I had the v/w thing pretty bad, and I was a vegetarian.
I remember you told me about this.
In North Dakota, I was in elementary school, and I’d be like, “Can I get a ‘wedgie’? Can I get some ‘wedgies’?”
So that’s for one gigawatt. I had an LLM run your wafer fab equipment model and figure out how much the tooling costs to produce a gigawatt of compute basically every single year. It said $3–4 billion.
Now suppose you add in cleanrooms and shell and everything else at the fab. So $6 billion of fab CapEx produces a gigawatt every single year. A gigawatt produces $100 billion of revenue right now.
But also, that $6 billion in CapEx is producing a gigawatt every single year, and that gigawatt is producing $100 billion every single year. So over the course of 5 years, the first gigawatt has generated 5 years of profits, the second gigawatt the fab has produced has generated 4 years of profits, and so on.
$6 billion of CapEx at the fab level will have generated over a trillion dollars of end-AI revenue.
Yeah. There’s a lot of OpEx along the way. There’s a lot of other CapEx, like the data center and the power. And you had to pay OpenAI for the R&D and installation. There are a lot of different people who need money here. Take away half of it for all these middlemen.
That still means there’s a 100x discrepancy between fab CapEx and the end revenue generated. More than that, actually, but we’re just being very conservative. As a result, this is capitalism. You have this huge discrepancy where you can turn $1 into $100. They’re not going to figure out a way to make more mirrors?
They are. It’s just that these mirrors take some time to make. But the urgency is so big that Anthropic and OpenAI are like, “We could make $1 trillion right now, but we’re just bottlenecked on the mirrors that go into the ASML machines.” How can we make more mirrors if we spend $100 billion on this? That’s the situation we’re going to be in pretty soon.
We’re not going to be able to solve that supply constraint?
That just seems quite hard to imagine. You’ve seen people do funny arbitrages here where they buy turbines and then try to resell them, because the value of a turbine is way more since it’s the thing bottlenecking your data center. I think if anyone had $400 million and the ability to convince ASML to sell them an EUV tool, they should totally just go buy one, wait, and then sell it for north of $1 billion.
But ultimately, yes, capitalism will cause these things to expand. But it’s a whip. It takes a long time for the whip signal to get to the tail end of that. The supply chain doesn’t react immediately. In fact, you go talk to someone at Carl Zeiss, and they’re like, “Yeah, yeah, yeah, we need to make 100 EUV tools by the end of the decade.”
When we had our episode earlier this year, they didn’t even think they needed to make enough mirrors to make 100 EUV tools a year. Now they’re like, “Okay, we need to do that.” But in reality, because of all the economics of what’s going on, it should be even more. It takes so long to build.
Suppose that every single company in the stack got private-equitied. Somebody came in who was super AGI-pilled and was like, “We’re going to maximize production.” What do you think the physical constraints on making more things would be?
The reason I ask is that we’re pretty soon going to be in a world where the lab revenue—or just AI cash flows, because obviously the accelerators also have these huge cash flows—will be so big that you can just fund extreme expansion of all this production from cash flows themselves.
I do agree generally. There are obviously some physical constraints. The way the supply chain is expanding currently, 100 is roughly still the right number.
For 2030?
100 ASML tools for 2030. But if you said, “Carl Zeiss, here’s $10 billion. Please fucking just expand production,” that would change things. You would have to do this with every company in the supply chain.
But you don’t think that’s going to happen next year?
I don’t think it’ll happen this year. I don’t think it’ll happen next year. I don’t think it’ll happen the year after, because the world is capital-constrained. But in a world where, say, the top labs are generating, even combined, $1 trillion in revenue next year, they’re not able to take $10 billion of that—I don’t think they’re going to do that, but—
Or hundreds of billions at least?
It just seems like they realize where the world is headed. I feel like they could just make—
The thing is, the labs are going to generate hundreds of billions of revenue next year. But ultimately, CapEx next year is like $2 trillion. So you’ve got this big mismatch. The wafer fabrication equipment supply chain will do something on the order of $200 billion. The data center market supply chain will do even more.
The accelerator supply chain will do even more. The energy supply chain will do a number. You sum all this up, and it’s going to be well north of $2 trillion of CapEx. So the labs have not yet gotten to the point where their cash flows can fund this stuff.
Obviously, they will never get to that point, because you want to keep your CapEx higher than your returns.
2. Compute prices will rise if the labs outbid everyone
Yeah, you reinvest. The key question I really want to understand is: If the current trend continues, it’d be north of 50 gigawatts per lab by the end of 2028. So between them, they’d have 100 gigawatts. Those gigawatts, as you’re saying, drive many-fold more throughput or performance by 2028 than they do now, because the hardware’s gotten better. Not only have flops per watt increased, but the hardware also gets better at working with AI workloads.
Okay, so 100 gigawatts for the labs by the end of 2028. How much is world compute?
I think that may be a little difficult, given that by 2028 they’ve taken 70–80% of incremental compute. And I’m not sure what happens to markets then. How much does the price of compute skyrocket for them to actually be able to buy 70–80% of compute? Is Google or Meta or Amazon willing to sell even that much?
Also, there’s one caveat when we’re talking about these gigawatt numbers. When Amazon is serving Bedrock Anthropic models, that counts as Anthropic compute in our worldview, because it is effectively, at the end of the day, counted as revenue for Anthropic, even though there’s a revenue share and credit back for all that.
But ultimately, in 2028, if they get to 100 gigawatts combined, they will have done really disruptive things to the market. Because anyone can make money off of $10–15 million per megawatt of compute today. I kid you not, it’s not that hard. Go get a GB300 rack, go download the Kimi weights, go download vLLM or SGLang, and set it up.
Codex and Fable can actually help you do this. It’s pretty simple. It’s not trivial, but it’s not rocket science. Go put it on OpenRouter. It’s very simple. You’ll start generating more revenue than you’re paying for the compute. This has already led to this compute pricing—$10–15 million per megawatt—starting to inflect up.
To get to that 100 gigawatts in 2028, you have to believe that the labs can outbid everyone for compute, because anyone can make money at $10–15 million. Does compute now get to $25 million a megawatt? Does it get to $40 million a megawatt?
As you’re saying, it’s already the case that the labs are generating way more revenue per megawatt than everybody else. If they stay as far ahead as they are currently, you would expect that to continue being the case.
If there’s some kind of recursive self-improvement where the AI labs are relatively uplifted—or they have models internally they’re not releasing externally that are helping them make their next model better—you’d expect that to be even more the case.
Aren’t you already seeing this, where xAI, or whoever is slightly further behind, will just sell compute to the highest bidder if they can’t internally monetize it as well as the labs? You’d expect them to keep bidding for larger and larger shares of the compute market.
I think that is my worldview. They will continue to gobble up more of the compute. But ultimately, they can’t do it at current pricing or anywhere close to it. They do have to start paying $25, $30, $50 million a megawatt to really gobble up 70% of the world’s compute in 2028, to get to 100 gigawatts by 2028, which is a very aggressive goal.
The other aspect of this that’s really challenging is that we’ve already seen a huge slowdown for the AI labs. This regulation that they advocate for is actually slowing down the labs a lot more than it slows down the open-source Chinese-language models.
OpenAI not releasing Astra. OpenAI stopping training for 2 weeks. Anthropic not releasing what their safety assessment says is Model 2, which is widely believed to be the next version of Mythos. They’re clearly not releasing their best models, in which case their revenue per megawatt stalls or can even start to decline again because other models are competitive again.
It’s not that they’re falling behind. It’s just that they’re not releasing their best stuff.
What if there is some regulatory impact that prevents them from releasing their best models?
Now their revenue per megawatt does not climb as fast. Their ability to buy that incremental compute for a higher price than everyone else starts to diminish, and then maybe they can’t get to that 100 gigawatts.
But in a world where safety doesn’t matter, I do believe that’s exactly what happens. They can start generating $100 million per megawatt or more, and they can pay $50 million a megawatt. No one else has any logical reason to do anything with their compute besides say, “Please, Dario, take everything off of my hands.”
But there are forces at play, which we cannot describe, that would potentially slow this down.
I think a good intuition pump is: What if the AI models were literally as good as a fully automated software engineer?
They’re not currently there yet. I think they’re far from being able to fully automate the job of a full white-collar worker. But white-collar workers earn 6 figures or north of that a year. If you have a gigawatt that can sustain a population of, say, roughly 1 million white-collar workers, then off the back of that—
That would be $100 billion. That’s actually surprisingly low.
3. Which layer will capture most of the surplus?
Yeah, $100K per person and a million-person population—I don’t know. But it would be many hundreds of billions of dollars per gigawatt if you get full AGI. The other aspect of this—and we’ve continued to see this—is that most of the value capture is not happening. Most of the value that these models generate does not get given to OpenAI and Anthropic. Thankfully, so far, it is mostly just being given to the users.
Take Jane Street, with its exclusive contract with OpenAI for GPT-5.6 Ultrafast mode, or as one of Anthropic’s biggest customers. It’s generating way, way, way more value out of the tokens it’s paying for than Anthropic is generating in terms of profit, because it gets to make money off the market.
Or take Meta, which at one point was rumored to be as much as 10% of Anthropic’s business. They’re generating way more efficiencies by optimizing their ad algorithms or what have you, getting engagement time 5% longer, all these things. They’re making way more money off using these models than Anthropic is.
That’s what’s required. Sure, if you had a million new software engineers, the cost for a software engineer would also fall. One thing I’m confused about is whether the market comes into equilibrium. If it comes into equilibrium, would you just expect the price of compute to equal whatever Anthropic and OpenAI can generate from it, or be very close to it, with a small amount of markup for Anthropic and OpenAI?
Right now, it’s really weird that there is a 4× or more difference between what compute sells for and how much money Anthropic can make from it. In a world where the revenue per gigawatt continues to increase, if Anthropic’s ability to monetize a gigawatt doubles or triples, it’d be weird if the gap continued to increase. Anthropic, just by having some weights, can take something that costs them $10 and turn it into $100.
This is always a fun question: Where does the value go in AI? AI is generating all this value. You’ve got the end user, which I think we all agree is generating more value than anyone else, hence they’re paying a lot for these models. Then you have the app layer. So far, the app layer has generated very little value.
Then you’ve got the model layer, which up until a year ago was generating negative gross margins and is now generating massive positive gross margins. It looks like it’s on the path to generating $100 million per megawatt—turning $10–15 into $100, as you said.
But if we go back a year ago, the hardware supply chain was generating all this gross margin while literally everyone else was losing money on it. OpenAI and Anthropic were just plowing VC money in, as were many other startups. Many of these hyperscalers were building infrastructure without knowing if there was going to be a payoff.
So ultimately, you had this negative value being created on the model layer, if you will, because they were selling the tokens for less than it cost them on the infrastructure side. All the value was being captured at the chip and fab layers.
Initially, in 2023, the memory companies were making no money off HBM or memory for AI, even though theoretically the value they were delivering was humongous. Now you’ve got—well, actually, TSMC captures way less value than the memory companies. So the value capture has shifted around a lot, which is very fun for people tracking or participating in the market, like Jane Street, as an example. This is not an ad. This is not an ad. This is not an ad. They’re a sponsor but you don’t have to plug them that hard.
So what happens going forward? Anthropic and OpenAI have slowly started to balloon in value capture. Do they balloon and take all the value capture?
Well, that was a thought, and then Elon showed, “Actually, no. I can sell my compute for $25 million a megawatt or $40 million a megawatt to Anthropic and Google. Even if it’s a short-term thing, I’ve sold it for this price, and I’ll recoup my entire CapEx in a year.”
What’s your prediction of how much the relevant tranche of compute—B300s or whatever SpaceX sold for $40B a gigawatt to Google—will sell for at the end of next year?
I think most compute will continue to transact at sub-$20 billion per gigawatt.
Even at the end of next year?
Because all of it has to be financed. If Meta, Microsoft, Amazon, or SpaceX can build compute without finding a customer—just saying, “Fuck it, I’m going to build this compute”—and then turn around and wait until it’s already built, they now control what’s going on.
Most compute is contracted well before it’s built. This is what Elon took advantage of in the market. He actually had all this compute. He was like, “Hey, Anthropic, I know you’re making $60-plus billion per gigawatt. Why don’t you just buy my stuff for a crazy amount of money?” Obviously, it’s not like Elon decided this or Anthropic decided this. The market figured itself out.
Other people—you go to a random cloud, and they’re like, “Okay, I’m going to build a gigawatt of compute, or 100 megawatts of compute. I’m going to spend the CapEx. I need to turn around and find a customer. If I want to find a customer, I need to find the capital. Who’s going to give me the capital and the customer? The customer has to sign a deal. Then I take the customer’s commitment to the credit markets and I raise the capital.”
So there’s this completely different power structure where Meta is effectively hoarding compute. Meta and SpaceX are the only plausible #3, because they’re hoarding all this compute. They’re using their balance sheets and capabilities to build compute without an end customer that’s monetizing at a huge degree.
They have an actual balance sheet, so they can go to the credit market. You build a gigawatt, you can make your margin—not a crazy margin, but a good margin. Now I have all this compute. Now Meta and xAI have this optionality of looking around and being like, “Is my internal use case going to make me more money, or should I go out there and sell it to Anthropic or OpenAI at crazy margins?”
So now we’ve entered a regime where SpaceX and Meta are saying, “Actually, I’m going to build the compute, and I can rent it out for not $13. I can sell it for $25, $50, and more.”
4. Will datacenter regulation slow down AI?
What do you think their revenue per gigawatt is by the end of 2027? For Anthropic or OpenAI, by the end of ’27?
I think it’s highly dependent on who has the best model and whether they’re allowed to keep releasing their best models. But I don’t see why it wouldn’t be $50-plus million a megawatt.
By the end of ’27?
Oh, by the end of ’27? That’s where it gets more challenging, but I think it could get higher than that—to like $70 or $80 million a megawatt, blended across the company, if not higher.
Seems low. So if that’s the case, then what happens to the price of compute?
Well, if I’m Anthropic, incremental compute is worth it. Maybe I spend $40 million a megawatt on SpaceX compute. If I’m SpaceX, I look to the supply chain and I’m like, “Well, I’ve struck this deal with Jensen, where he’s now all of a sudden using Twitter.”
And Elon’s saying they’re exclusive to NVIDIA, but why doesn’t Jensen raise his prices? Then SK hynix, Micron, and Samsung look at it and they’re like, “Well, why don’t we raise our prices?” So with the value capture, I think there’s a bullwhip effect here. Just because someone has raised prices doesn’t mean the entire supply chain rebalances immediately.
But over time, the supply chain will rebalance, and things will cost more and more. To get that incremental capacity, you have to. So TSMC is raising prices very slowly, but memory companies are raising prices very quickly. Substrate companies are raising prices very quickly. Elon wouldn’t have sold if it was $15, but he’s selling because it’s $25+. So obviously, he raised his prices really quickly.
I’m surprised you think that revenue per megawatt doesn’t increase way more than even $100 million per megawatt by the end of next year. When does RSI happen? When does takeoff happen? Or even if RSI doesn’t happen, just say the current rate of progress continues. Just look at how much progress we’ve made in, let’s say, the last year and a half. What was the model from a year and a half ago? Claude 3.5 or something?
My problem with this is that the best model that exists in the world was trained in February. So you’re saying maybe we just won’t be allowed to release the labs’ best models. OpenAI says they’re not training models for 2 weeks, man. What the hell?
There’s one thing: internally, are they getting enough use for it that they’ll bid up the price of compute? Another is, does AI progress as a whole slow down because of regulation?
Yeah, but they’re not even allowed to use this new model internally. Astra’s not even widely deployed internally.
But still, if you have a model that is… What was the model released at the beginning of last year? GPT-4o? Was that 4o?
Yeah. You’re talking about a GPT-4o-to-Mythos-2-sized leap by this point—again, by the end of 2027.
Yeah, but Mythos 2’s not out. Or even Mythos. That leap again. Even Mythos is not allowed to be out. They’ve neutered it. We can’t use it to optimize inference performance. We can’t use it to optimize all sorts of things.
Yeah, maybe there’s some slowdown in AI progress or the deployment of AI that means the revenue per gigawatt can be lower. But that’s the only way I could see it being only $100 million per megawatt by the end of next year. As long as the model gets better, the value generated out of it gets better. Obviously, who captures the value is still up for debate, but ultimately everyone’s going to raise their prices because they can, and it’s super inflationary.
Right now, so far, the method of regulation is just “don’t release the models.” But more and more, the method of regulation is New York banning data centers. Texas is holding moratoriums. Ohio’s saying, or at least trying to say, you have to pay everyone’s property tax in a certain radius. These sorts of things are going to decrease supply and increase cost. That’s going to get passed on as well.
You start to end up in a spot where progress does slow, at least in the external sense, even if the models internally keep getting better and better. In a takeoff scenario, why would Anthropic not have their best model 6 months ahead of what is externally available? Because of safety and regulation, but also the competitive advantage? That 6-month difference, if progress accelerates, is actually a bigger differential. So that’s the thing that would cap revenue-per-megawatt gains to much lower growth than we’ve seen in the first half of this year.
5. Labs are shifting compute from inference to R&D
Here’s something I’m very interested in. As these companies go public and they’re accountable to investors, let’s say by the end of next year they have close to 20 gigawatts. So 10% of compute is 2 gigawatts. Let’s say they want to go from 60% of compute devoted to training to 70% of compute devoted to training. And their investors are like, “Well, if you’re going to be able to generate $100 billion per gigawatt, you’re basically saying no to $200 billion of revenue in order to increase your training compute.”
So investors are like, “What the fuck? You’re already spending so much on training. Why are you spending even more on training?” As a public company, what do you think would happen if they’re just like, “No, we will keep increasing the share of compute we spend on training to offset the increase in revenue that each gigawatt of compute is giving us”?
This is what I personally believe: the labs are going to allocate less and less compute to inference over time. I think that’s very non-consensus. The standard belief of most people is, “Oh, most compute will go to inference.” Most of it will go to forward passes for training, not necessarily revenue-generating inference.
Ultimately, if they’re generating $30–40 million per megawatt today, you allocate 40% to inference. If you now get to generating $60–70 million per megawatt, do you still allocate 40% to inference and generate all this profit and then do dividends and share buybacks? Or do you go build AGI? I think the obvious answer from Anthropic and OpenAI, not just at the executive level but also their boards, is to go build AGI, because it’s way more profitable. So ultimately, you’re going to see them ratchet up their percentage of compute dedicated to training.
While each increment of compute is getting more and more profit-generating if they had dedicated it to inference.
Right. The whole point is, if I’m selling tokens, is OpenAI releasing ultrafast mode just for external use, or are they doing it internally too? It turns out, no. Actually, I’m going to allocate it to internal and external, because the internal value I’m generating from super-fast AI or the best AI model is way more than what someone external is.
So ultimately, sure, I could generate $100 million per megawatt, but if I turn that toward AI research, what is the incremental progress that I get? What does that do toward my future earnings potential—the discounted cash flows of whatever the hell I’ve done? They’re not going through that calculation, but ultimately it makes more sense to dedicate more and more compute internally. The only reason to have inference compute be so large is so you can grow your training fleet.
I think this is an interesting economics question that I feel we can have the models digest. What would have to be true about a world where they reduce the fraction of compute spent on inference?
I think they have over the last 3 months already. I think that at parts of this year, they were increasing the fraction of compute going to inference. Let’s just take it month by month. You would agree that every month, Anthropic has added more compute than the prior month. There might be some noise when they sign a SpaceX deal or whatever, but in general, the amount of compute is a curve up.
So in January, they added less compute than in December, and yet their revenue adds skyrocketed. Then they’ve sort of plateaued. They’re not adding $25 billion of ARR every month now. That means the marginal megawatt they’re getting is going as a higher percentage to R&D than it is to inference. So they are factually increasing their compute toward R&D today.
6. China gets less than 10% of new compute, but its labs need less
If I look at the numbers you said for how fast world compute grows, here are some things I want to understand. It seems like if I add up the numbers you just said, it would be over 200 gigawatts of world compute by the end of 2028, right?
Yeah, globally.
Okay. How fast can global AI compute continue growing after 2028?
30 this year, 50 next year, 70 in ’28. ’29 should be on the order of 90–100.
Then just 100 more every single year or something?
I think the slope can continue to go upward. It’s hard to predict anything more than 4 years out. Who knows whether we’re in an RSI regime? When will the world economy be growing at 10% a year? Because if you’re at 100+ gigawatts a year, you’re at absurd GDP growth.
If you think there’s 200 gigawatts globally in 2028, how much is in China by that point? How does Chinese compute continue increasing through this whole trend? Because if the RSI stuff kicks off in the West before China has a large amount of compute, maybe we’re living in a different world than when it doesn’t.
If we level-set back to 2022, the US was adding about 45–50% of the world’s compute. China was adding about 30–35%. The rest was being taken up by the rest of the world.
Since 2022, we’ve had big regulations against China and a dramatic increase in America. So today, 70% of watts are being deployed in America. China is really a very small number. Less than 10% of watts being deployed for data-center AI compute are in China.
As we step forward, they’re still at a very small number. Their domestic production is quite small. Their purchasing from NVIDIA is still quite small, and a lot of that ends up in other places as well—Malaysia or what have you. So ultimately, China domestically still continues to have less than 10% of incremental new compute.
In 2028, it might start to inflect up, I think.
But it’s pretty easy to say China will have 30 gigawatts of AI compute or less by 2028.
Yeah, in 2028.
Okay. And then how fast does their hockey stick go up?
I do think in 2028 they have a big uplift in the amount of compute they’re able to deploy. In 2026, they’re still mostly relying on a lot of smuggled chips, a lot of the chips that TSMC made for companies that they thought weren’t Huawei but ended up being Huawei, or a lot of HBM that Samsung is shipping.
In 2027, fabs start to go up. In 2028 especially, fabs start to go up from SMIC and CXMT and such, where domestic production is actually reaching many millions of units a year. Now they’re incrementally adding 5–10 gigawatts, just in 2028, of domestically produced chips.
Those chips are definitely worse than the chips that NVIDIA will have in 2028, or Google will have in 2028, or OpenAI will have in 2028. So even the gigawatt number overstates things—you’re saying. It’s 30 gigawatts, but it’s really much worse chips.
But if you think the world is going to add 100 gigawatts the following year—I know you said you can’t really say that far out—how much is China able to add the subsequent year? Basically, I want to know: do they just hockey-stick at the point at which they are able to start shipping large amounts of compute, or is it still going to be less than the US plus allies?
There’s a lot left to whether or not the US passes the MATCH Act, whether or not tools continue to get export-controlled, and how fast China can build the new equipment that they’re starting to be able to produce domestically. But ultimately, China is definitely going to hockey-stick. If there’s anything China’s really good at, it’s scaling manufacturing really, really quickly.
I imagine China will start to be able to extract more and more purchasing of even foreign chips into domestic China, or at least close the gap in what the US is allowing NVIDIA to sell them, or what have you.
But do you think China could be adding 50 incremental gigawatts in 2029?
I think that’s completely reasonable. Part of that could also be purchased from foreign sources. But, yeah, I think it’s completely reasonable that China in 2029 can do 50 gigs.
But if most of those are domestic chips, there is some factor there where that 50 gigawatts is really worth as much as 20 gigawatts from American chips.
Right. So you’re actually projecting a world where maybe the leading lab in 2028 has more compute than all of China will have in 2029 or even 2030, if you weighted gigawatts by their quality—implying that there’s nothing done to slow down the US labs.
That’s right.
But clearly the government and politicians are starting to do that. Whereas China’s not going to slow down AI. In fact, the only thing they’re going to do is accelerate it.
Honestly, when I interviewed Jensen and asked about export controls, I’m a libertarian person—I wasn’t genuinely sure what I thought about this issue. I was steelmanning the opposite view from what he has, because I think it’s important to hash out ideas. I thought, “Yeah, maybe there’s a world where, if we just cooperated with China, it would be better for us, especially since they control so much of the supply chain and the other things that will be needed for robotics.”
But I didn’t realize the compute situation was as fucked as you’re saying. Actually, the export controls do seem to have really made a difference. If they ship the amount that you’re saying, that’s a huge difference. By the time we have an automated coder and are getting into an automated researcher, China is way far behind on the compute stock. If that ends up being the case, that would have worked.
I think that’s actually a notable success. The only caveat there is that some of it is export controls, but some of it is also just financial systems. American financial systems are more willing to YOLO into startups than Chinese financial systems. But once Chinese financial systems choose an industry to focus on, they’ll subsidize it a hell of a lot more.
So the Chinese semiconductor industry has significantly more subsidies than the rest of the world’s semiconductor industries combined.
If takeoff is not as fast as you’re implying but actually takes longer, then ultimately China will catch up drastically on the semiconductor side, which then is compute at some point.
The other noteworthy aspect of this is that Chinese companies today are not that far behind in AI models, at least as perceived by the public, relative to the amount of compute they have. The leading Chinese labs have 100–200 megawatts total of compute at most, ByteDance Seed being the one outlier where they have significantly more than that.
But Kimi is not running a gigawatt or anywhere close to it, whereas Anthropic is more than 5 gigawatts by the end of the year. So the question is, does it matter?
I think right now this difference in compute doesn’t matter that much. When we break down the compute ratio, or budget, of a lab, so far it’s been 60% training and 40% inference. But that training gets broken down further. Actually, 50% of the compute is research, 10% of the compute is development, and then 40% is inference.
What I mean by research and development is that researchers are generating ideas, testing new architectures, testing new data mixes, testing new hyperparameters, new attention techniques, blah, blah, blah. But ultimately, when they do the training run—when Anthropic trains Mythos—it’s sub-200 megawatts.
The pre-train or the whole thing?
The pre-train. It’s sub-200 megawatts for, call it, 2 months. Then the RL is even less.
You think the RL was less compute than the pre-train?
At least in terms of a single site for pre-training, yeah.
But total compute was probably higher, right?
But it’s sequential. At most, the most they ever used at 1 point in time was maybe 200 megawatts. In reality, they had multiple gigawatts, so most of their compute was going to the research, not the development of a model.
There are reasons for this. It’s hard to coordinate all these clusters. It’s hard to co-locate all of them. It’s hard to do multi-site training. It’s hard to do RL. Generating even more rollouts during RL does not necessarily make it better. There are all sorts of reasons why you may not be able to leverage all 2 gigawatts that you have onto training. Actually, I can only leverage 200 megawatts.
As we get further and further down the path of automated coding and automated research, I actually expect the percentage of the compute budget that goes to research versus training to become a lot more fuzzy, or even higher for training. Also, things like continual learning—all of these things start to mean that more and more is actually going to training the model.
If you end up in a world where you’re doing 100 gigawatts a year, at current prices, that would be $5 trillion of CapEx every single year. Then stack on the fact that you have to build the power plants way before then. It’s also a 30-year asset. You stack on the fact that the data centers are a 15- to 20-year asset, and you have to build that then too.
So the $5 trillion, once you account for future years’ growth, is actually going to be more like $7 or $10 trillion of CapEx.
Wait, I didn’t understand. That doesn’t include the fact that there’s not the infrastructure for the power generation in the data center itself.
Right, exactly. When you talk about AI CapEx, people are saying $40 or $50 billion. But that’s really just the critical IT: the servers, the networking, the fiber, the transceivers, optical communications, and all this sort of stuff.
It doesn’t account for the data center itself or the power plants themselves, which are being built ahead of time. If I’m building 100 gigawatts this year and 150 gigawatts next year, then all of the buildings for that 150 gigawatts need to be built in CapEx this year. If I’m building 200 gigawatts the year after that, all those power plants need to be paid for too. You have to buy the turbines this year.
So actually, it’s much bigger than even $5 trillion if you’re building 100 gigawatts.
Right. Very plausibly, incremental CapEx every year is getting close to $10 trillion by the end of 2030, which is going to be close to a tenth of the world economy. If all of it’s going up in the US—
The US economy will have grown as well.
But still, at the current size of the US economy, it’ll be like a third to a quarter of the US economy just going toward data centers. As I say that out loud, I’m like, “Maybe you’re right and we just won’t allow it, and that’s the reason this doesn’t happen.”
Because for this exponential to continue, a quarter of America’s economy is just building data centers. I believe in capitalism and reallocation of resources toward the most profitable thing. But at the same time, politics exist, credit markets exist, and capital markets exist.
So to enable, let’s say, that 100 gigawatts by 2030—or let’s even pare it down to 2028, where it’s like $3 or $4 trillion of CapEx across all of these items: over $2.5 trillion toward IT CapEx, and then another $1–2 trillion on data centers and energy, and all the supply chain downstream, like semiconductors and all that stuff.
If you’re at $3 or $4 trillion of CapEx, where does all this cash come from? No one is generating that much cash from the business yet. Hyperscalers funded all of the growth up until now: Google, Microsoft, Amazon, and Meta. They funded a huge percentage of it.
They were more than half of compute, but they now don’t generate cash. They actually spend everything on CapEx. In addition, they raise debt and spend everything on CapEx. You’ve seen Meta do it, even Amazon, even Google. Microsoft will be there soon.
Everyone is raising debt to pay for their CapEx. Now, who is the incremental person to pay for this who was not doing it before? In the case of Google, it was pretty simple for them to stop doing buybacks, or for Meta to stop doing buybacks, and turn around and buy compute infrastructure. That doesn’t have a huge effect on the market, but it does have some effect.
But as you step forward to 2028, where the hyperscalers are now raising hundreds of billions of dollars of debt, and then all of their supply chain is raising hundreds of billions of dollars of debt, who pays for this? There are a few different ways.
There are semiconductor companies like Nvidia and Broadcom, and the memory companies, turning around and deciding to fund some of this CapEx. There are the traditional infrastructure investors who are gathering capital and investing in infrastructure. Instead of bridges, it’s data centers.
Then lastly, there’s everyone in the economy who’s realizing, “Maybe I shouldn’t buy a home, or maybe I shouldn’t invest in credit that’s helping people buy homes, or maybe I shouldn’t buy government debt. I should just buy hyperscaler debt, or I should buy this data center’s debt, or I should buy Anthropic’s debt.”
Because Anthropic’s willing to pay 20% rates for the incremental billion dollars to build their capacity. They know their revenue from it’s going to be huge, and they’re going to pay 20% because it’s still better than renting it from SpaceX for $50 billion a gigawatt.
So you’ve got all of this contention. But if you now do this, the whole world economy is really shifted around.
Antithesis is a deterministic software testing platform that enables perfect reproducibility. It also unlocks some pretty insane approaches to debugging. Like time travel. With Antithesis, you can jump to any point in a trajectory and start from there. So when there's a crash, you can rewind to the exact moment that something went wrong and freeze the entire system: the application, the database, even the environment itself. This lets you do something that would otherwise be impossible, which is to observe every part of a distributed system at the exact same instant. Time travel also allows you to add telemetry and logging to an event that has already happened. For example, you can rewind to five seconds before a crash and decide to capture all the network traffic. Most powerfully, Antithesis gives you a live terminal into your system that you can use to perturb anything you wish. Kill a node or disable a feature, then hit play and see what happens. Then go back and try something else. In production, you often only get one shot on goal with this sort of destructive analysis. If you restart a deadlocked service, for example, the exact deadlock you needed to study disappears. But with Antithesis, the original timeline is always reproducible, so you can test as many hypotheses as you need. And if you don't want to do all this time traveling yourself, you can just have your agents do it for you via the Antithesis API. Go to antithesis.com/dwarkesh to learn more.
7. Will AI cause a sovereign debt crisis?
You and I have been debating off-air for the last few days whether there will be a sovereign debt crisis as a result of AI. The logic is this: As we were mentioning, you have a situation where very little investment turns into a lot of money. So the rate of return—
What a fucking problem, dude. Oh my God. I can’t believe it.
No, it is a huge problem for everybody else who can’t turn a little money into a lot of money. So the rate of return is incredibly high.
Even at the data center level, if you build a data center and you’re trying to get it rented out to an Anthropic or an OpenAI for 10x what it costs you on a depreciated basis to build it, it’s fucking crazy. You turn $1 into $2 or $10 or something at the end of the year.
That pushes the interest rate higher. Now, if the interest rate goes higher, and if it does that for the entire economy, people are borrowing more and more money. They’re competing against the other lending that the government would’ve done, or that other companies would’ve done, or that you as a consumer or a mortgage buyer would’ve done. That’s making it more expensive for everybody else to borrow.
This has huge implications for tons and tons of people. Sorry, I’m going to go on a bit of a monologue here, but we’ve been thinking about this together.
I think the US will be fine at the end of the day. Because if the data centers are built in America, you can fundamentally just tax the data centers. But the way the current tax system is set up, corporate income is less than 10% of federal revenues. More than 80% is payroll taxes and income taxes, which, as more and more automation happens, will shrink.
At the same time, on the spending side, currently 20% of tax-revenue spending goes toward servicing the debt, paying interest on the debt. Now, a lot of the debt is short-duration, so it rolls over every 5 years. Why are you fucking laughing?
Because it’s things you’ve learned in the last month. Like it’s any different for you. Like you got a degree in fucking financial economics.
I didn’t. The internet thinks I’m a beekeeper. A few months, a few months. This is our business, Dylan.
I know, I know. Sorry, sorry. Now I’m self-conscious. Fuck.
No, it’s good. You’re doing good. I just think it’s funny.
A million people listen to this guy who just learned about debt this month.
Suppose the interest rates rise 1%. Over a 5-year basis, the fraction of tax revenue that goes toward servicing the debt goes from 20% to 25%. If it rises 5 percentage points, that would go north of 40%.
But if you take into account the fact that the government is borrowing $2 trillion every single year, then that goes from 40% to north of 60%. So 60% of tax revenue just goes toward paying interest on the debt.
Now, I think the US is going to be fine because the tax base will increase if we let data centers get built in America. Other countries are absolutely fucked, in my opinion.
I was just looking at which countries have a lot of debt, have very little tax revenue, and also have a lot of their debt serviced quite often. Those countries, like Pakistan or Nigeria, I think are just going to be very fucked in this new interest-rate regime.
This crowding-out effect is the reason it’s not YOLO 1 billion gigawatts. You’ve got all these industries and countries that use a lot of debt, all these impoverished countries that you mentioned earlier that are just going to default.
You’ve got consumer packaged goods, all of these companies that make things you see at Trader Joe’s or wherever. They use a lot of debt. All these telecom companies use a lot of debt. Banks use a lot of debt.
So if interest rates go up in the market—not necessarily the government-set interest rate, but the spread between what the government says their federal funds rate is versus what everyone else is charging, because Amazon wants to raise $100 billion of debt next year or whatever the hell the number is, probably less—you end up with this really challenging problem of where the cash comes from.
There is some level that is funded by cash flows, and the cash flows keep going up. But the logical thing to do is to invest way more than your cash flows because then the returns in future years will be amazing. So you have this delta.
Then what’s pushing down on the delta is all of these other things: regulations against data centers, consumers getting mad, politicians getting mad, regulations against AI, and the AI labs not releasing their latest models because of safety reasons. Interest rates going up are an influence on all of these things.
So all of these things bend the curve from what capitalism wants in terms of pure, simple economics to what the complex system that we have wants, and bend it lower and lower to where not as many gigawatts as should be built will be built.
Well, the interest rate is part of capitalism, right?
Yeah, but in the simple economic model versus the more complex model of what we have.
What is the rate at which you think Amazon or Anthropic or whatever will be issuing bonds for debt next year, if they do hundreds of billions of dollars of debt?
What is the average rate?
I don’t think Amazon will do hundreds of billions of dollars of debt.
In total. Let’s say the big tech guys—the hyperscalers in total, and all the clouds.
In the modeling that we do, we have about $11 trillion of CapEx from 2024 to 2029.
Total?
Total. If you fund a lot of this with cash flows, as much as you can, you still end up with north of $5 trillion of credit that needs to be issued for this $11 trillion-plus build out.
So you don’t think AI revenue continues even 3x-ing year over year?
AI revenue does go up. I don’t think it can go up forever without certain constraints being hit. Labs will have certain incentives. Labs are not the ones building all the compute in many cases, even though they’re increasingly trying to go that way.
But they’ll have all this cash flow. How much did you say the revenue will be? You think they’ll not have that much revenue?
No, I’m just saying till 2029 there’s something on the order of $11 trillion of CapEx. $6 trillion of that is funded with cash, and $5 trillion of that is funded with debt.
If that’s the case, $500 billion of debt being raised across the whole ecosystem does make interest rates go up. Then what prevents that?
There are a couple of things. One, do labs increase their revenue per megawatt and keep inference allocations large? In which case, they’re accumulating all the profit across the S&P 500 because everyone’s paying to reduce their costs. Of course, their profits will also go up, but cash has to come from somewhere.
So there’s an upper limit on how fast their revenue can grow versus the value they deliver into the world. And there’s a diffusion aspect of the technology. But ultimately, labs’ revenues keep going up. They can’t cash-flow fund everything.
The optimal scenario is you actually use credit as much as you can to fund it, because even if cash flows from the labs fund a lot of stuff, you want to build more than that. So there is some amount of credit that gets built. Our current modeling has $5 trillion of credit and $6 trillion of cash-funded infrastructure investments through ’29.
When you take that, this is not enough compute relative to what the demand growth is from the AI models. So you’ve got the obvious answer, which is revenue per megawatt keeps going up.
That makes sense. How much do you think interest rates will increase by 2029 as a result of all this?
Dude, this is vibing a number, but if you’re vibing a number out—
Growth in the world economy is going up a lot, so why wouldn’t interest rates for Amazon go up from where they are today?
This is going to be extremely vibed out, but recently Meta has raised at 5% to 6%. I don’t see why they wouldn’t pay 8%. They would happily pay 8% because the return from the compute that they’re going to build is humongous. The market won’t want them to, but they’ll want to pay 8%.
The flip side is that if they pay 8% versus the 5%, 5.5%, 6% they do today—a 250-basis-point increase—that makes everyone else in the economy also pay 250 basis points more, which then causes a lot of things.
Banks will scream, because if their credit spread goes up, their debt reprices faster than their assets reprice. They ultimately end up losing tons of money if their credit spread blows up.
The other consequence of this—this is a point you made—is that if interest rates rise, the discount rate increases, which means that the discounted cash flows of all equities crater.
Which means that even though the stock market as a whole might be doing fine—the S&P 500 will be fine—any individual stock will probably have just cratered in value, especially the Buffett, Berkshire-type, pay-good-cash-flows-for-30-years type stocks.
Yeah. It’s like, “Why would I pay this much for Johnson & Johnson?” They’re seen as a stable stock: good cash flows, they’ll return their cash flows over time. Or a railway company. Why the fuck would I invest that much if my discount rate isn’t 3% or 5%? It’s now 8% or 10%.
For developing countries, Basil Halperin, who’s a good friend and an economist, made this point that we’ll see a second Volcker shock. In the ’80s, to fight inflation, Fed Chair Paul Volcker raised interest rates more than 5%, to something like an 8% real interest rate. That caused some 40 different countries, mostly in Latin America, to default in that decade. I think that will probably happen again.
Okay, now we’re getting into singularity talk. We’ve been talking about what happens if interest rates rise—I think this all happens before singularity, by the way.
Yeah, that’s what I’m saying. We were talking about, before singularity, interest rates rising 2%–3%, et cetera. At some point, I think it’s very likely that the world economy will be doubling every single year.
This is not happening in 5 years.
But it’ll happen eventually. There’s this researcher, Damon Binder, who’s done great work on this. If you look at input-output tables in a fully automated economy, what would it take to double the entire stock of things in the economy every single year?
Yeah. If the economy grows at 3% a year, then, by the rule of 70, that’s 20-something years.
Right. But he was like, “Okay, right now we’re bottlenecked by the fact that there are people, and you can’t double people every single year.” But in a world where you can also double the labor force every single year, how fast can the economy grow?
I think it could double every single year. At the very least, it would be tens of percent every single year.
Okay. The rate of interest should be pretty close to the growth rate. It won’t be exactly that because of consumption, but it should be pretty similar. Then we’ll go into a world, I think in the 2030s, where the rate of interest is tens of percent. Part of my brain is like, “It might be hundreds of percent,” but let’s say it’s at least tens of percent.
I’m just like, okay. Every country that is not involved in the production of AI defaults. Every stock that is not an AI stock is worth basically zero because discounted cash flows are worth nothing. If the federal government can’t figure out a way to tax AI, servicing the debt is more than the current tax revenue. And you have all these other effects that I’m sure we’re not even pricing in: you can’t get a mortgage, et cetera, et cetera.
Fundamentally, what is happening in this world? This is all nerd speak, right? But let’s step back. What’s happening?
Just now it started, the nerd speak?
We’d be entering a totally different growth regime. The economy’s basically saying, “Hey, the opportunity cost of the government borrowing money to pay people’s pensions is extremely high now. Because that money could be spent building a robot factory that builds a robot factory that builds a robot factory.”
The opportunity cost of capital is going to increase a ton. That’s fundamentally the cause of all of these things we’re talking about.
As interest rates go up, equity markets get pummeled. Even AI companies. Some people who really believe in AI are like, “Why does Micron or SK Hynix or Kioxia trade at 2 or 3 times earnings?”
It’s like, “Well, if you’re really AI-pilled, everything in the economy should trade at 2 or 3 times earnings.” If you’re not AI-pilled, then sure, they’re over-earning.
It’s an argument for why—I think memory is going to do great—memory stocks shouldn’t 10x or whatever again. Because if we’re in a market where there’s that much demand for memory—which means AI’s caused this drastic change in the economy—then everything should trade at 2 or 3x multiples and the stock market should fucking crash.
In a sense, Meta trading at—I think they’re like a $1.5 trillion company—it’s like, what? Silly. They’re worth way more than that, at least in a logical sense. You just look at their cash flows, all the infrastructure they’re hoarding, and all the compute that they’re going to be able to sell for crazy amounts of dollars per watt, either as tokens because their lab works, or just to Anthropic and OpenAI.
It ultimately becomes a question of: you have to reallocate all the capital to AGI. You do that by pricing everyone else out. So the limiter on AGI is not how fast the research engineers, like our roommate Sholto, can crank the gears. It’s actually just how much the rest of the world lets that happen.
Because they’re going to regulate. They’re going to obviously increase interest rates. They’re going to say, “No data centers.” They’re going to say, “Stop building fabs.” They’re going to say, “Oh shit, every company’s equity value is tanking, so how can I pay for AI to increase my business?”
Well then, Anthropic and OpenAI have to start building their own stuff.
They’re building their own chips already, or at least designing their own chips, and it’ll expand out. They’re contracting their own data centers and building their own infrastructure in the next couple of years. There’s the question of how this reallocation of the economy happens. There’s a lot of downward pressure on it not being just a straight takeoff, even if the models were capable of it.
I think you and I believe we’re in a world where models are capable of that. But slow takeoff is possible—at least, that’s my hope—because of everything in the economy and regulatory world. The government is saying, “Don’t release your models.” The government is saying, “Actually, you can’t even use your models internally that much,” because that’s going to happen soon. They’re already saying you can’t release your models.
The thing I’m most worried about is a singularity, which external deployment is actually helping. So the fact that we’re preventing external deployment is stupid.
Does that prevent singularity?
Right now, it would lead to more revenue because the models are incapable of RSI. But I’m worried about a world where it’s 2030 and the government’s like, “We’re going to wait 6 months before you can release your newest model to the public.” Six months, 100x. Let’s go.
In those 6 months, they do recursive self-improvement internally. They just have all kinds of crazy shit happening in the company. Meanwhile, the rest of us are stuck with models that are, at the current pace, years behind.
Here’s my thought. Suppose that the whole world gets in on this conspiracy to try to slow down AI.
I don’t think it’s a conspiracy. It’s explicitly stated by every politician.
Suppose they slow down AI by a year. If compute is increasing 2–3x every single year, they prevent a whole year of AI deployment, such that you’re a year behind where you would otherwise have been. During RSI, you’re getting 3–6 years of AI progress in a single year.
But they don’t just limit compute. They also limit the lab’s ability to release the model internally. We saw that.
If they did that, that would be ideal. Anthropic had to stop giving Mythos to foreign employees for a bit. I didn’t know that was true internally as well?
That’s what they claimed.
I thought that was just a different checkpoint that was not Mythos, but it was basically Mythos.
But stuff like that is not going to be allowed either. The government is dumb, but they’re not that dumb, I would hope, at least. Governments—at least the US government, which has the cards here—are not going to want Anthropic to use Mythos 4 internally. They’re going to be like, “Hold the fuck on. Slow down,” because of all of these regulatory reasons.
Everyone who’s elected is going to hate AI. Even the people who are elected already hate AI. All the constituents. I bet you at some point your parents are going to call you and be like, “Dwarkesh beta, you’re doing a terrible job. You’re making AI progress happen faster.”
Because of my podcast, I’m accelerating AI progress?
Maybe. You educate people. Maybe if they’re smarter, they’re progressing AI faster. Anyway, you’re going to have real-world constraints on the progress, development, and deployment of AI. Even though it will happen eventually, we could tear ourselves apart before we get there.
Jane Street is hiring for two separate ML internships right now: one focused on ML engineering and the other focused on ML research. I sat down with Alok, who helps run the research track, to learn more about that program. I think this domain is fundamentally understudied. Often we have unanswered questions within our deep learning research team where we don't understand, say, some market participants' behaviors or certain dynamics of how trading happens. These unanswered questions make for really good intern projects because they are ultimately topics that we care about and just haven't gotten around to figuring out yet. So even as an intern, you'll be contributing to real research, not working on some sort of contrived exercise. The Jane Street team follows frontier LLM research closely. A relatively common intern project is adapting a recent paper to financial markets, which come with their own set of gnarly problems. Ultimately, we're trying to model thousands of interconnected irregular time series. The signal-to-noise ratios are extremely low because we have a lot of competitors trying to do the same. So we have this adversarial, non-stationary, extremely high-dimensional problem that we're trying to solve. To be clear, you don't need to know anything about finance in order to be a good fit. As long as you have a background in ML research, Jane Street can teach you the rest. Their 2027 internship applications are open now. Apply at janestreet.com/dwarkesh.
8. Will the world's future workforce belong to a few companies?
One thing I find crazy about these scenarios is just how much of the world’s future labor supply ends up in very few companies, and also how fast that labor supply grows year over year. If compute at the frontier in FLOP terms is growing 4–5x a year—and, further, the compute required to achieve a level of capabilities is decreasing 3x a year—basically, the effective AI population size at the frontier labs is increasing 10x year over year.
That doesn’t really matter that much right now because AIs are not good enough to do full jobs or be as autonomous as people in their capacity to do work or pull off schemes or whatever. But if the current trend continues, you have a world where OpenAI goes from having 10 million AI laborers this year to 100 million the next year, to 1 billion the year after that.
Pretty soon, even if compute scaling slows down, it doesn’t take many more years before each company individually has more labor equivalence than there are people on Earth. I think it’s very plausible by the end of this decade that there’s more AI labor—more effective population—within a single lab than there are people on Earth.
We talk often about centralization of power because of nationalization or whatever. But we don’t think enough about the fact that we’re actually moving very fast into a regime where most “people,” in terms of work output, are concentrated within 2 labs that are consuming more and more of the world’s compute.
If these AIs are misaligned, then most of the world is misaligned because most of the world’s minds are there. But even if they’re not, very few companies have a lot of influence or a lot of control.
There was the whole spat recently where I think Gavin Baker was like, “Dario believes that there’s only going to be 1 company in the world.” Then Sholto and Dario came out and were like, “No, no, no. We didn’t say that.”
But ultimately, if you believe in RSI, if you believe the labs are the most effective users of compute and can generate the most value from the compute, then the only thing that’s going to happen is centralization of compute. If you believe in AI researchers, RSI, and AGI, then all of this exists. All of this is the base.
This is even true if there’s no RSI. The effective population of the frontier is currently increasing 10x year over year for a given level of capabilities. So if you get to the level of capabilities of a very competent remote worker, a very competent software engineer, or a very competent researcher, the population of those is increasing 10x year over year at the current rate of capabilities growth.
I see, and without RSI.
Then once you have RSI, it’s even crazier. Then it’s maybe growing 100x a year or 1,000x a year. Or their intelligence is increasing but the population isn’t increasing. Or some mixture of the 2, right?
What world do you see, Dwarkesh, where everything is not centralized? Because it seems to me that every force is screeching toward centralization. And that’s scary as hell.
I would love for it not to be centralized completely. But maybe that’s the whole point of “Machines of Loving Grace,” right? It is everything, and it makes our lives great. It’s so hard to think about the future. But I agree with you.
I think the fundamental problem is that AI training has huge economies of scale because any effort you spend on training an AI for a specific skill or a specific set of knowledge gets amortized across billions of sessions or billions of users. So that’s 1 effect.
The other effect is that if you’re slightly ahead in the AI race and compute is in shortage, you can charge a much higher markup because you can better economize this scarce resource. So there are 2 effects that give more and more to the person who’s ahead in the AI race.
There may be more. If models are learning from deployment, and 1 model is deployed much more widely than another one, it’s getting much more real-world data.
Your point is taken that whether it’s user deployment and continual learning, whether it’s training and having these economies of scale, and whether it’s the incremental progress where the best AI model helps you make the next-best AI model—RSI—all of these things point to centralization. I think one of the big intellectual projects, honestly, that we should spend some time thinking about—or at least I’ll spend some time thinking about—is this: What is a vision of a decentralized, broadly empowered future after AGI that takes these economies of scale seriously?
The alternative vision is that the government controls it, and maybe you think that you can trust the government more because it’s not a private corporation. I don’t trust the government, and I don’t trust Dario, and I don’t trust Sam. That’s a problem, right? Obviously, it’s very easy to be wrong about the future. You don’t anticipate a key effect or something that changes everything.
But ex ante, it’s very hard to see how we avoid a scenario where we have to choose one source of centralization. It’s why capitalism worked, right? It’s decentralized decision-making and decentralized power. And it’s why super-centralized capitalistic economies actually grew slower than super-decentralized capitalist economies, to some extent. You have to have rule of law and all this.
But then AI flips all this on its head. And ultimately you’re like, “Actually, private ownership is probably not the most efficient economy, and therefore it grows slower than an AI economy, which is centralized.” Well, it’s still private ownership, but how many firms are really involved in this share of the economy? It’s, what, maybe 2% of the economy right now? $1 trillion divided by 30.
NVIDIA is a huge share of it, and Anthropic and OpenAI and these hyperscalers. Obviously, there are other firms involved, but a large share of the AI stuff is just happening from very few companies. So it could be private property, but very few companies are involved. I mean, this is what the structure of the market is doing. So what can prevent it?
I don’t know. Unless AI progress slows down, unless governments regulate the fuck out of it, this is all that happens. In which case, we’re headed for a world where either we have super concentration of resources and we pray that that one company gets everything right, or we have governments slow everything down and people slow everything down, and you have a slowdown of progress somehow, hopefully, and there is more of a balance of power. Even as we go towards AGI, ASI, RSI, everything along the way will still lead to someone capturing more resources. So it’s kind of hard to find a framework in which AI doesn’t lead to super concentration.
Now, the one positive thing here is that today Anthropic does not capture most of the value. We can talk all we want about how they went from $20 million per megawatt to $100 million per megawatt, but they’re still paying $13 million for a lot of the compute they’re buying. But at the end of the day, the reason they’ve gone to $100 million per megawatt is because Jane Street is capturing $300 million per megawatt or $500 million per megawatt.
Or Dwarkesh, from researching his podcast and learning about credit, is capturing how many dollars per megawatt? Now, how much can you use? Tough. But I think that’s the one saving grace: The rest of the economy maybe profits so much more from Anthropic—
No, but the whole logic you were laying out earlier—them reallocating inference to AI R&D—the whole logic of that is that the returns to labor inside AI labs are much higher than the returns outside.
Yes. This is my cope. I agree. In all scenarios of the world, there are 80,000 worlds, and in only 1 of them, Anthropic doesn’t own the whole world.
Again, power concentrates because I don’t want to send the tokens outside. They’re more valuable inside. So it’s the same thing. Why would I let Jane Street make all this money off of these degenerate options traders?
Hey, they’re a sponsor, come on. Jesus Christ.
No, I think it’s great. It’s a good value for the world to make it an efficient market. Jane Street making all this money off of getting the worldview correctly, making money off of degenerate options traders, whatever it is—why would Anthropic allocate compute to that? If the end monetization that Jane Street has per megawatt is $200 million, so they’re willing to pay Anthropic $100 million, well, what if Anthropic can just generate hundreds of millions of dollars per megawatt by using that compute internally? That’s what’s happening.
On that somber note, I guess we’ll meet again when the RSI is officially kicked off.
You’re not going to have me on your podcast again for 2 months? All right, cool. Thanks, dude.