AI抛售与数据不符|顶级AI投资人解读
- 7月是“一个月里的2022年”——AI标的直线下跌40%-60%。 但Baker花了一周在硅谷压力测试,试图找出“一个负面的量化指标”,最终没有找到明确答案:GPU供给、GPU租赁价格、现货DRAM价格和token增长都在加速。Nvidia目前的远期市盈率处于10年来最低水平;“市场100%认为它们赚得远超正常水平……也许确实如此。”
- 核心做多逻辑是算力重估。 ’24-’25年所有人的模型都假设GPU价格会下降,但老GPU价格“将在2026年垂直上行”——一家初创公司以每GPU小时2美元中段的价格租下相同的B200集群,7个月后却希望以略低于4美元的价格续租;一家推理云计划在合同续期时支付高出100%的价格。低于市场价的合同陆续到期后,超大规模云厂商的存量资产将被重新定价:“几乎所有超大规模云厂商都赚得不够。”
- 他认为,信用是唯一真正的看空催化剂。 超大规模云厂商CDS全面走阔、Meta债券定价不佳、实际收益率上升,但共识是按Aier一代的费率将即将到来的Blackwell/Reuben吉瓦级算力变现,对应超大规模云厂商1.3-1.4万亿美元的经营现金流;如果只按当前Blackwell价格打个折,规模就是“约2万亿美元,拿掉7000亿美元的信用需求”。兜底逻辑是:“如果信用不存在,只意味着现有的flops会变得更有价值。”
- 开源恐慌的方向搞反了:“token就是token”。 GLM 5.2和Kimmy K3在相同的flops、内存和功耗下,把token从毛利率约90%的前沿模型转移到毛利率约30%的开源模型,利润从模型层迁移到基础设施层。关键观察是:如果开源不利于Jensen的生意,他不会成为“全世界最大的开源支持者”。
- 内存LTA如今已成为决定厂商业务与产业地位的关键。 4家买方达到规模、市场份额由供给分配决定之后,“违反一份LTA,可能摧毁整个业务和产业地位”。Nvidia的答案——“如果GPU价格高于底价,就通过信用包装加收入分成”,再加上到处持股——市场并未理解;如果他经营Hynix,也会照搬这一套。
- 市场微观结构已经变了。 所有人都把每条新闻输入Claude,“Claude有点像股市的Walter Cronkite”——它很聪明,但并不总是正确,把一个持续3年的日本电容周期压缩成6周。他的锚点是:“投资中最重要的3个词不是安全边际,而是我不知道。”
- 最大的风险是监管,而不是基本面。 纽约的数据中心暂停令看起来可能只是第一例,整个行业还在打一场糟糕的公关战,其中一部分建立在作者承认的1万倍用水量错误之上——尽管数据中心是“我这一生中蓝领工资经历的最棒事情”。
- SpaceX是被低估的算力机器。 它按每吉瓦约500亿美元的规模变现,而市场共识对明年营收的预测只有730亿美元;一份公开报告称其将在18个月内新增8GW,Baker“几乎不敢相信”,但IPO以来基本面已因Grok 4.5和Cursor变得更好,“轨道算力每天都显得更真实”。
1. 7月是“一个月里的2022年”——但最终只出现一个有争议的负面数据点
- Baker给这个月的标签是“一个月里的2022年”,AI标的“一个月内直线下跌40%-60%”。他这一周的任务是“压力测试……告诉我一些负面的东西”,但在GPU供给、GPU租赁价格、DRAM现货价格和token增长等指标中,没有找到明确的负面数据:“事实上,每个指标都在加速”,而且依据的是数据,不是感觉。
- 他反复提到的估值事实是:Nvidia的远期市盈率处于过去10年来最低水平,只有DeepSeek和Liberation Day时期更低;而这两次都形成了V型反转底。“市场100%认为它们赚得远超正常水平……我们需要保持谦逊。也许确实如此。”
- 他在市场光谱中的位置是:“这里几乎所有人都比我更乐观。”他读到的一篇文章设想以约250,000美元租用一张H100一年,大约是当前现货价格的15倍;这“甚至没有进入我考虑过、但最终认为完全不可能的结果”。“我看着股市发生的一切,感觉自己像个愚蠢的乐观主义者”——但在硅谷,他却是那个空头。
2. 恐慌的构造:除了信用,所有催化剂都被误读
- Meta“出租算力”被市场解读为产能过剩和即将削减资本开支。Baker认为:“事情完全不是这样。”Meta看到SpaceX以“真正巨大的溢价”将面向训练优化的大型集群卖给客户,获得了展示IRR的样板交易,甚至可能是在股权融资前布局。资本开支的遥测数据从未改变,资本开支也没有削减;Meta长期以来最好的模型之一Muse 1.1虽然被Grok 4.5的光芒盖过,但说明油门仍踩到底。
- Silicon Data的token指数下跌,看起来像需求开始转弱,实际是结构迁移:GLM 5.2和Kimmy K3把token从前沿模型定价——推理毛利率可能是“80%、90%或95%”,具体仍有争议——转向开源模型。市场把它当成负面信号;他看到的却是利润迁移,而非需求萎缩。
- 中国DUV光刻机的消息重创半导体设备篮子。他的类比是:DUV相当于螺旋桨飞机,EUV则是喷气涡轮机——即便落后25年,这仍是一场真实的阶段跃迁,因为芯片制造依赖在实践中学习:“你不可能瞬移到未来。”他的结论是:这件事意义重大,但市场可能仍然反应过度。(对于无法核实的走私EUV传闻,他说:“这得是多么惊人的间谍成就——那些东西那么大。”)
3. Token就是token——开源拿走的是利润,不是算力需求
- 他反驳的核心是:“token就是token,生成一个token所需的算力完全相同”——相同的flops、相同的内存、相同的功耗。开源模型夺取份额,会剥离前沿模型层的利润,同时把利润和需求弹性推向基础设施层。关键观察是:如果开源损害了Jensen的生意,他会成为“全世界最大的开源支持者”吗?
- 路由器经济学解释了企业AI账单下降的悖论:一些公司在3个月内把AI支出放大20倍并烧光预算,随后安装路由器、削减支出,但消耗的GPU小时数可能反而上升——因为节省来自用毛利率约30%的开源token替代毛利率约90%的前沿token,“很多情况下效果略好,成本只有一半”。
- “开源在某种程度上是公开市场的暗物质”——难以测量,但正在将其变现的美国推理云,包括Fireworks、可能还有Baseten、Modal和Together,都显示GLM 5.2/Kimmy K3完成能力跃迁后,需求正在加速。唯一偏软的数据点是第三方数据所显示的Anthropic曲线出现弯折;这“很可能是真的”,但Anthropic股东强烈反驳,争议正酣。
- 他考虑过“Anthropic/OpenAI/Grok最大主义者”的观点:前沿模型一旦达到RSI,就会在每个智力水平上以更低成本完成蒸馏,开源模型因而没有生存空间。但他对此存疑:AI原生公司可以在专有数据上对模型进行RL(Fireworks的Nexus只需“三行代码”;Harvey在被收购前和Cursor都在积极采用),不再只是包装层;而便宜的“120 IQ”开源模型,反而可能让“160 IQ”的编排器更有价值。
4. 信用是唯一真正的担忧——而算力重估可以化解它
- 他无法轻描淡写带过的催化剂是:实际收益率上升、利差走阔、所有人的CDS“全面走阔”,以及Meta债券“没有按你想象中Meta债券应有的价格成交”。“如果我们需要债务来融资这轮建设,那会非常、非常可怕”——债务驱动的建设要求立即偿还,这正是互联网周期瓦解的原因。
- 他的模型是:共识实际上按Aier费率——也就是落后两代的费率——来计算即将到来的Blackwell/Reuben吉瓦级算力,对应1.3-1.4万亿美元的超大规模云厂商经营现金流。如果只按当前Blackwell价格打个折,规模就是约2万亿美元,“拿掉7000亿美元的信用需求”,并改善正令市场恐慌的信用指标。
- 这已经开始体现:Microsoft、Meta和Amazon的经营现金流本季度从28加速到32;剔除一笔异常的一次性项目、主要是欧盟罚款后,则是从28加速到35——在这个体量上,这是“实质性的加速”,而且还没算入Reuben溢价和合同重定价。兜底逻辑是:“如果信用不存在,只意味着现有的flops会变得更有价值。”
5. 现货与合同:存量算力的收益远远不够
- 这次行程中的故事是:一家热门初创公司租下数千张B200,价格为“每GPU小时2美元中段”;7个月后,它希望以略低于4美元的价格租用基本相同的集群——现货价格上涨了50%-60%,而2024-25年的每一个模型,无论看多还是看空,都假设GPU价格会下降。“我不认为24年或25年有人想到老GPU的价格会在2026年垂直上行。”
- 一家推理云在播客中表示,合同到期时计划为Blackwell支付高出100%的价格——“这只意味着几乎所有超大规模云厂商都赚得不够。”随着合同陆续到期,合同价格将向现货价格重估,即便现货价格本身下跌也一样。
- 需求背景是:全球可能只有250,000-500,000人在使用智能体AI,同时算力正处于严重短缺状态。“当我们从500,000人增长到100,000,000人,再到500,000,000人时,会发生什么?”他最喜欢的验证方式是:“你听过有人说GPU太多了吗?”一个也没有——“听起来像毒品市场。”
6. Claude是股市的Walter Cronkite
- 一位投资者(音频中姓名不清)提出的理论是,市场多样性崩溃会先于泡沫和崩盘;如今这一理论需要更新:所有公开市场股票投资者都会把每条新闻输入Claude或Claude代理,“它解读这些新闻的方式可能没有太大差异……Claude有点像股市的Walter Cronkite”——它非常聪明,“但并不总是正确”,而投资本质上是对未来进行概率化的贝叶斯解读。
- TBU的日本电容股图表就是例证:“我们在6周内经历了完整的电容周期”,股价先垂直上冲,随后快速坠落;一个本应持续3年的周期,在基本面甚至尚未兑现前就被压缩完成。
- 一位Fidelity朋友对这个时代的箴言是:“以最快速度做最愚蠢、最表面的事情,然后在这些事情之间循环。”这个月,市场依据各种叙事削减风险,而“除了信用之外,这些叙事都有点荒谬”。但股价仍在下跌,他也尊重技术派的警告:“按定义,真正击中你的子弹一定是你没看到的那颗”,同时坚持自己的锚点:“投资中最重要的3个词不是安全边际,而是我不知道。”
7. 内存LTA与Nvidia的信用包装:用上涨空间换取持久性
- 他承认自己判断错误的转变是:内存厂商用带有预付款、底价和上限的长期供货协议,换取短期价格暴涨的空间。内存是整个体系的轴心——每个flop配更多内存,就意味着单位算力能生成更多token,这是“你能做的最重要的一件事”,也解释了需求为何没有出现负弹性。
- 2027-28年打破LTA的博弈论很清楚:有4个买方真正重要——Amazon的Trainium、Google的TPU、AMD和Nvidia,“比其他所有人加起来都大得多”——市场份额由供给分配决定。打破其中一份协议,等议价权重新回到内存厂商手里时,“你就会倒闭……你可能摧毁整个业务和产业地位”。这与Apple时代不同,当时内存厂商总会重新接纳最大的买方。
- Nvidia的新模式是:“如果GPU价格高于底价,就通过信用包装加收入分成”——这不是供应商融资,因为贷款方是第三方——同时在各处持股(“几乎每次没有持股,最后都证明是个错误”)。它可能建立一项由版税驱动的巨型云业务,提高每吉瓦收入,并加固护城河:竞争芯片买方在TSMC和HBM DRAM上支付更高价格,而“没有什么比Nvidia GPU更容易融资。没有。”这让其10年低位估值倍数“有点难以理解”。
- 当被问及如果经营可能是Hynix的内存公司会怎么做时,他回答:“现在Nvidia做的事情,我会一模一样地做。”也就是投入现金、分享持续收入,这是LTA交易的自然延伸。“我相信我们的朋友Blackstone和Apollo已经在向内存公司建议某种类似方案。”
8. 没有人松开油门——以及技术层面的变量
- 每家实验室都吸取了同一个教训:如果Anthropic当时跟上OpenAI的算力攻势,“它本来会一骑绝尘”。OpenAI已经重回竞争,Grok位于Pareto前沿,SpaceX则通过Grok 4.5和Cursor进入战局。Dario面临的两难是:买太多算力会破产,买太少会输掉竞争;这带来一个问题:“尤其是在可以用经营现金流融资的情况下,近期有人会退缩吗?”
- 他最大的技术收获是:很多人认为自己“接近解决持续学习和样本高效学习”问题(SSI表示其模型将在8月推出)。如果一个模型只需在10万亿token上训练,然后在真实世界中学习,而不是使用300万亿token,那么训练占算力的比例“最终会趋近于一个并非接近0、但非常小的水平”——这“对世界来说很棒”,很难相信会对基础设施需求不利,但“再次强调,我在努力保持开放心态”。
- 硬件变量是推理解耦:在不带HBM DRAM的芯片上做prefill,在高功率HBM芯片上做attention,再用老制程制造的SRAM加速器运行前馈网络——“对于那个前馈网络,SRAM就是无可匹敌”。他认为这“对AI的ROI非常、非常积极”。
- 具备《权力的游戏》规模的黑马包括Fireworks的Lynn(“绝对的杀手”)、Cognition的Scott Wu,以及一个音频中听不清的名字。推理云本身可能才是沉睡的变量:它们的早期增长速度几乎和前沿实验室一样快,却消耗很少现金;按Rule of 40口径计算,数据“疯狂得惊人”。
9. 监管是最大风险,而行业正在输掉叙事
- “监管一定是最大的风险……你不能忽视纽约实施数据中心暂停令”——这看起来可能只是“第一例”,而“AI行业的公关做得非常糟糕”。华盛顿的叙事是:数据中心推高你的电费、拿走你的水、抢走你的工作。
- 他的反驳是:表后供电交易通常会降低当地电价;承诺期开发商如今正在建设医院、学校、警察局和消防站;就业也会通过持续维护和升级而延续。“从很多角度看,数据中心是我这一生中蓝领工资经历的最棒事情”——但民主党这个名义上的蓝领政党却在反对它们。
- 用水恐慌源自一位作者承认的1万倍高估,这个数字至今仍在传播——就像80年前把菠菜含铁量小数点弄错的“大力水手”错误。他半认真地提出的解决方案是:由基金会或政治行动委员会在四强赛和NFL比赛期间投放“数据中心究竟做什么”的广告,同时讲清楚AI治疗疾病的故事(今年ASCO大会上出现了单次会议中最多的科学突破,部分归功于AI)。
10. SpaceX:市场没有给这台算力机器定价
- Patrick问SpaceX是否是“最重要的新上市公司”;Baker认为,IPO以来基本面已经改善:Grok 4.5、收购Cursor(Cursor“显然已经实现了有意义的加速”),以及过去3年里以比任何人都更快、更低成本的方式上线算力。他们曾在现货价格高点一夜之间向市场投放海量算力,但“货运列车完全没有减速”——这本身就是更看多需求的数据点之一。
- 一份公开报告称SpaceX将在18个月内新增8GW——他“几乎不敢相信”,但SpaceX按每吉瓦约500亿美元的规模变现,而市场共识对明年营收的预测是730亿美元,因此只要接近这一水平,就足以淹没市场预期。只有超大规模云厂商、CoreWeave、Crusoe和SpaceX曾在一年内新增超过500MW。Elon的说法是:“我们的专长是让不可能的事情迟到。”纽约大型对冲基金的核心做空逻辑,需要现货算力价格下跌约90%。
- 在Starbase待过之后,他说:“轨道算力每天都显得更真实。”他的理智校验对象是Benchmark——与Elon生态没有关系——该机构正在为StarCloud提供资金,而SpaceX正与其进行某种合作,可能允许其使用Starlink激光技术:“也许我疯了,也许Elon疯了,也许Benchmark也疯了……但这件事不太可能这么巧。”
核验说明
- 原始字幕无法确认多样性理论中提到的姓名;本摘要保留匿名处理。
I want to be scared. I don't want to feel like a lunatic watching these stocks get cheaper, thinking the expected forward returns are going up. My main mission out here this week is to pressure-test.
Yeah.
Tell me something negative.
But I haven't been able to find one that's a quantitative metric. The underlying fundamentals in stocks are improving. NVIDIA is actually, as we record this, at its lowest forward P/E of the last 10 years. The market 100% thinks they're significantly overrated.
Gavin, it's only been 2 months. The model release cycles—the gap between our podcast episodes—are shortening.
We're basically—you and I are basically on a model-release cadence at this point.
Well, I was sensitive to criticism that somebody pointed out our podcasts were coincident with local market peaks, and nobody can say that after this. What's on your mind? It's been a crazy, crazy month.
1. AI Selloff vs. Fundamentals
I would describe July as 2022 in a month.
Yeah.
There are some fundamental negatives that we should talk about, but on the whole, the balance of fundamentals is improving significantly. Loads of AI names are down 50% to 60% from their highs—call it 40% to 60% in a month, in a straight line. I asked you before we started: you've been out here for the summer. Have you heard a single negative quantitative metric about AI?
Yeah.
A single instance of deceleration?
Nothing.
In fact, every metric is accelerating.
And to your point, it's not just blind optimism from people excited about AI. Here's some data they can show you from their different vantage points.
Absolutely. However you cut it—whether you cut GPU availability, GPU rental pricing, the spot price of DRAM this month, or token growth—everything has actually accelerated.
I do think a big part of the problem is, one, the market does not have visibility into Anthropic and OpenAI. Then I would say these open-source inference clouds that monetize inference here in America—Fireworks, Baseten, Modal, and Together. The picture looks very different when you see that, because open source has accelerated massively because of GLM-5.2 and Kimmy K3, and then Neatron continues to chug along.
We had a great, very small American open-source model release. OpenAI has accelerated, and Anthropic continues to grow really strongly and is almost certainly pumping out significant amounts of free cash flow. I just think there's this chart that everybody looks at, of semiconductor cash flow going like this and hyperscalers' free cash flow going like that, and you're missing these private companies.
But I also think that chart misses something very important. Everyone in 2024 and 2025 thought—even if you were really bullish—that the price to rent a GPU would decline slowly. If you were bearish, you thought it would decline precipitously. I don't think anyone in 2024 or 2025 thought that the prices of old GPUs would still be going vertical in 2026.
Yeah. Everybody thought, "Hey, we're going to be smart. We're going to sign these long-term contracts." To some degree, a lot of the neoclouds had to do that because they needed an offtake agreement to finance the GPUs.
Essentially, you have the contracted base of installed compute trading at a massive discount to the current spot market. As those contracts roll off and compute gets repriced higher, spot can decline and compute will still get repriced higher. I think you're going to see a lot of acceleration that's going to answer these ROI questions. You've started to see that this quarter.
If we look at operating cash flow—not free cash flow—operating cash flow from Microsoft, Meta, and Amazon has accelerated from 28 to 32. There are some actually pretty big unusual items now. These hyperscalers always seem to have billions of dollars of unusual legal expenses, mostly fines to the EU, but there was an unusual amount of one-time items this quarter. If you strip that out, we went from 28 to 35, and that's a material acceleration at this scale.
That's really before they start to light up the Reubins, which will come at a meaningful premium, and before these contracts reprice.
It's been a challenging month. Is it helpful to walk through the month and how we got here?
Yeah. First, Meta is going to rent out compute, and this is seen as very bearish. They have excess capacity. They're going to cut capex. This is a disaster. This is not at all what it was. They just reported that they didn't cut capex.
What it was is they saw SpaceX have a big installed base of compute and sell some big training-optimized clusters into the market at a truly massive premium to these contracted rates, and at least the analysts liked that. They saw an opportunity. There's a lot of speculation that they're going to raise capital.
Maybe what they were thinking is, "Hey, we will show on a small chunk of capacity that we can generate really strong IRRs, then we'll go raise equity capital, and we'll be off to the races and probably raise capex." That doesn't look like what they're doing. Nonetheless, the market sold off because it interpreted this very negatively, and I was really sure it wasn't negative.
There's a lot of telemetry into Meta's capex plans. None of that telemetry had shifted at all. If anything, it was continuing to get more aggressive. Then, shortly after that, they released their best model in a long time, Muse 1.1, which is actually a very good model. It was overshadowed by Grok 4.5, but it was a good model—way better than you think in 2 years. There's just no chance they're taking their foot off the gas.
Then Kimi came out, and there was this huge freakout about open source. At the same time, this Silicon Data token index dipped and flattened, and the two are connected. What the Silicon Data token index captures is mix, and they don't see all the tokens. Because of GLM-5.2, too, and then Kimi—which took a while to layer in—there's kind of a mix shift from more expensive frontier tokens, which probably have an inference margin—
We can debate whether it's 80%, 90%, or 95%.
—but super high—toward open-source tokens. For whatever reason, the market thought this was negative. But the reality is, a token is a token, and you need the exact same amount of compute to make a token. All else equal, it takes the same amount of FLOPs, the same amount of memory, and the same amount of watts.
Now, tokens are not equal, but broadly speaking, all open source taking share does is take margin dollars out of the frontier-model layer and, because there is elasticity, drive token demand. You need more demand for compute. Anthropic and open-source models all run on the same underlying cloud providers, which charge the same amount for compute, so you're literally just taking margin from frontier models and driving more margin dollars into the AI infrastructure layer. I think that was a catalyst.
Well, this combination of things—well, yeah. It's like Jensen is the world's largest supporter of open source. He's a super-idealistic guy. He's a patriotic American. I think he always does what's right. But does it really stand to reason that Jensen would be the world's biggest supporter of open source if it were bad for his business?
He'd still support it if it was the right thing for the world, but maybe it wouldn't be his signature issue.
Yeah. And by the way, I think open source is really important in a world where there's just 1 or 2 dominant frontier models that charge 90% margins. It's not good for humans. It might not be good for society. I think we want a lot of models, as we've discussed before.
2. Financing the AI Buildout
Then China has a DUV machine, and this causes a huge sell-off in semicap equipment. Then we get to what I think is, in a lot of ways, the real concern: real yields have gone up, which makes sense. We're investing a lot to fund this investment, and for sure credit is an increasing part of it, even if the majority—overwhelming majority—is still funded out of operating cash flows.
So real yields go up and spreads widen. Meta priced a bond last week, and it did not price where you would think a Meta bond would price. This just shows that the credit market—
And maybe a CDS was blowing out.
All of these CDSs for everybody are blowing out. Very smart private-capital people just said, "Hey, this is exactly what you'd expect. These are just banks hedging their commitments." But nonetheless, it doesn't look good.
These are undeniable facts: CDS is up, spreads widened, and real yields are up. That would be really, really scary if we needed debt to finance this buildout. That's where I think this differential between spot and contract pricing for the installed base of compute is so important.
It's so important to understand what the financing will be like for the next 6 months or something. The degree to which this buildout is going to require credit, right? That would be the classic capital cycle. We start to overextend ourselves with debt, and that's where things get difficult.
100%. Debt-fueled buildouts demand immediate repayment, so if supply and demand get a little bit out of whack, things can unwind very, very quickly. That's what happened in the internet.
If one believes, as I do—rightly or wrongly, and after this month I'm super open; I'm looking at this like I've been pressure-testing all of these, and I really went deep on credit because this is real—if we need credit to fund this buildout, this is a significant negative.
If you model it out, if you look at the amount of gigawatts that are supposed to come on and the consensus estimates for hyperscalers, they're effectively modeled—and these are gigawatts of Blackwell and Reuben—to monetize roughly at the rate of Aier, which is 2 generations behind. Not at Hopper, but Aier. There's $1.3 to $1.4 trillion in hyperscaler operating cash flow.
If you just assume that they're not going to monetize at the rate of Ampere—and I think it's very unlikely they will—we could go into why. Some of it comes from just seeing what is happening on the ground with demand here from real quantitative metrics. But let's just say they monetize at a discount to current Blackwells.
Then it's more like $2 trillion—
—of operating cash flow.
And that kind of takes $700 billion of credit demand out.
And, ironically, that improves all the credit ratios as these installed bases of compute reprice. We're going to continue accelerating; consensus is modeling a deceleration, which I think is unlikely. Then the credit metrics look better, and all of a sudden it gets easier to finance with credit.
Now, whether they choose to do that or not, we'll see. But this is all a little bit—I think we spoke 2 months ago.
No, but the time before that, we talked about the risks of a Blackwell air pocket, where you're spending hundreds of billions of dollars on Blackwells. They're mostly being used for trading initially. Trading does not generate a return, and this could be a risk.
We actually really saw that kind of in the first quarter. I think one reason I got comfortable with that risk when we did the podcast 2 months ago was that you were seeing such incredible things out of Anthropic. Then it's like, okay, well, the market's kind of going to look past this.
And it did look past it in April, May, and June. Then in July, because of this confluence of things, it stopped looking past it just as operating cash flow started to really accelerate. This is just a fact: It is accelerating at big scale.
So essentially, what this all comes down to is: Do you believe that the quantitative demand signals we're seeing on the ground here in Silicon Valley from private companies are going to continue, such that the installed base of compute reprices higher as contracts roll off?
And operating cash flows go up.
Operating cash flows go up, and you could fund most of this out of operating cash flows—maybe all of it. If it reprices at current rates, you could probably fund all of it for the next several years.
It's been a very unusual episode in the market. In some ways, the fact that—well, we should talk about what the fundamentals are that are getting better. Technicians would say it's like 2022: The market is worried about a recession, rates going up, and inflation. That's what the market was worried about in 2022. You knew exactly what it was. DeepSeek—you know what it's worried about. Liberation Day—you know what it's worried about. There's something very clear. In a weird way, that's comforting and reassuring.
And here, we talked about a lot of specific things, but it just feels like all those specific things, with the exception of credit, are just kind of ridiculous. The fact that it is still going down, a technician would say, "Hey, that's a little scary." It's definitionally the bullet you don't see that gets you.
3. GPU Prices Keep Rising
You know, I think we've talked before about how I think the 3 most important words in investing aren't "margin of safety," but "I don't know." You've been out here for 2 months. I've been out here, and I literally spoke to a company this morning that rented a cluster of several thousand Blackwells. This is one of the sexiest startups that people want to be in business with.
They had rented a cluster of several thousand Blackwells at, let's just call it, somewhere in the mid-$2 per GPU-hour. They're renting the exact same cluster—the exact same size cluster, essentially identical in every way, B200s, no differences—and they're hoping, 7 months later, to pay just under $4.
You hear this today.
That's pretty crazy because, again, you would expect a really gentle decline in prices to be bullish. Instead, we're up, depending on the starting point, 50% to 60% in 6 or 7 months.
There have been so many anecdotes like that. I think one of the inference clouds—I think it was based in—I’m not sure. They went on a podcast and essentially said, "We are planning to pay 100% more for Blackwells when our contract expires." That just means that essentially all the hyperscalers are under-earning.
I haven't found anything negative. My main mission out here this week was to pressure-test everything.
Like, pressure-test.
Yeah.
Find—tell me something negative. The question I asked you was: Is there one negative quantitative metric you've heard? That's what I've been asking everyone.
The main thing people are saying is that the third-party data on Anthropic suggests that the Anthropic curve started to go off its trajectory a little bit. That's the only thing that I—
I think that may very well be true.
But then you have OpenAI and open source massively accelerating the competition. If you look at the sum, it is net accelerating. I don't know that it looks the same; I think it may have accelerated.
Open source is a little bit like— they talk about dark matter in the universe. Open source is kind of dark matter to the public markets. It's hard for public markets to measure it, but if you just track what these inference clouds are saying, demand is clearly accelerating.
That makes sense because you had this huge capability leap with GLM-5.2 and Kimmy K3, which I think we're going to see continue. I think you're going to see NVIDIA bring Neatron steadily closer to the frontier.
It's been a very humbling, challenging month, but it's also like, wow, I've kind of pressure-tested every assumption. The underlying fundamentals are improving, and NVIDIA is actually, as we record this, at its lowest forward P/E of the last 10 years.
Crazy.
The only times NVIDIA has been cheaper were Liberation Day and DeepSeek, and those were kind of V-bottoms.
And that means to you just that the market thinks they're significantly over-earning?
Yeah, the market 100% thinks they're significantly over-earning, and we need to be humble.
Maybe they are.
Maybe they are. But my mission out here this week was to look for negative data points as hard as I could. Normally, you come to Silicon Valley and there's a mixture of, "Okay, here's something negative; here's something positive." On balance, it's positive. Technology creates value over time.
But I haven't been able to find one that's a quantitative metric, other than that Anthropic third-party data. I would say that seems to be hotly contested by the Anthropic shareholders, who are chomping at the bit to tell you what they know.
We're also very scared they're not going to get an IPO allocation [laughter] if it gets back to the company that they're the ones who said, “Actually, things are great.” You can just see Anthropic shareholders; they want to be like, “It’s not true.” [laughter] I mean, it’s hard for me to believe that open source and OpenAI have accelerated to the extent they did, but yeah, Anthropic is clearly kind of in the pole position.
And, oh, by the way, Grok and Cursor have also—you can see from third-party data that July was a pretty transformational month, with Grok 4.5 builds coming out. So it has been a tricky month, and I have a friend at Fidelity who just says the way to have navigated the last 3 years is just to do the dumbest, most superficial thing as quickly as possible and cycle between them.
What is that? What is that now?
Well, that’s just been to cut risk. Yeah.
All month, in response to these kinds of narratives that, factually, except for credit, are just not true, people have been cutting risk. The work we’ve done makes me think that credit just isn’t going to matter. Has this repriced? Let’s just say you do need credit to build the FLOPs we need. Well, if credit isn’t there, it just means the FLOPs that are there are going to be even more valuable.
4. Claude Moves Markets
There was an interesting essay that got sent to me. I think we’ve talked before about Mike Mikeson's theory that a breakdown in diversity is kind of what leads to bubbles and crashes. Essentially, everyone I know in the public-equity investment business, whether retail or institutional, feeds everything immediately—every piece of news—into Claude, Claude Code sometimes, or a Claude agent. Claude is probabilistic, but there’s probably not that much variation in the way it’s interpreting this news.
So it’s almost like we’re back to, in stock-market terms—there’s never really been this way in the stock market before, but people talk about the fragmentation of media and how it used to be like Walter Cronkite was the only voice of truth, and now we don’t have that anymore. It’s like Claude is kind of Walter Cronkite for the stock market, and everybody just believes whatever it says.
Funny.
Really? And, by the way, it’s really smart, but it’s not always right. Its interpretation isn’t always correct. With the stock market, you are fundamentally dealing with a probabilistic Bayesian interpretation of the future. So it just feels like, in the market, there’s this: here’s this piece of news; it gets fed through Claude; Claude interpreted it this way. So, a huge chunk of people trade on Claude’s view.
There’s this guy, TBU[?]. He’s part of the anonymous semiconductor mafia, but he posted this amazing chart of Japanese capacitor stocks. He said, “We’ve had a capacitor cycle—an entire capacitor cycle—in 6 weeks.” And it’s true. The stocks—whether they double, triple, or quadruple, I don’t know—went vertical and then whoosh. The actual fundamentals haven’t even hit, and yet you’ve already had what probably would have normally been a 3-year cycle in 6 weeks.
What’s your sense of being out here? It makes me especially curious about the innovation that is going on here to improve the efficiency in every aspect of serving inference, of training models, et cetera, and how that will affect public markets over time. Have you learned anything interesting about the long-lead-time innovation-type stuff that has you especially excited or curious?
Yeah, I am very curious. A lot of people seem to feel like they are very close to solving continual learning and sample-efficient learning, which we’ve talked about before. It is possible that, if those are solved, that could be a temporary discontinuity in demand.
If, instead of having to—I think somebody told me that it was trained on effectively 20 billion tokens, and then these models are trained on 300 trillion tokens—and if you can train something on 10 trillion tokens and then let it out into the world and learn sample-efficiently, that doesn’t sound good for training demand. But training as a percentage of semiconductor demand, compared to compute, is going to asymptote to something—not approaching zero, but very small.
I would say that is the most interesting thing. Who knows if it’s a long horizon or a short horizon? SSI says that they’re going to come out with their model in August, and there’s this whole generation of new labs that are focused on this.
And this would be good for the world. To be clear, this would be awesome for the world.
Yeah, we all want this.
Yeah, we want this. It would be amazing for the world, and it’s hard for me to believe that would actually be negative for AI infrastructure demand. But again, I’m trying to be really, really open-minded. I would say that was probably the biggest, whether we call it scientific or technical, takeaway.
But it’s just—you just don’t know.
5. What Could Break the Thesis
Well, yeah. And also, Nvidia is heavily involved with all of these startups. If you were forced to come up with the set of circumstances that would really switch you around and get you really scared, would it just be that this operating-cash-flow thing doesn’t play out, and therefore we need to debt-finance this? If operating cash flow does not continue to accelerate, that would be negative.
That, to some degree, is going to be a function of how Anthropic, OpenAI, Grok, Cursor, and open source do. If there was a pretty dramatic contraction in GPU prices that was kind of sustained, the market would react to that instantly. That would be worrisome. If it started to get really easy to get GPUs, I mean, have you heard anyone say they have too many GPUs?
Not a single person.
In fact, it’s the opposite. It sounds like a drug market or something.
Yeah, it really does. It’s just wild. But, yeah, I think there’s a long list of pretty obvious things. If Anthropic, OpenAI, Grock, Cursor, and open source—the sum of these labs—plateaus or starts to decline, that’s really negative, unless it’s just because open-source tokens are net growing the pie and taking share.
I do really think the future is multimodel. Particularly for the AI natives, they’re going to want to take an open-source model. All these inference clouds have gotten really good at supervised fine-tuning and reinforcement learning. So you can take your data, customize an open-source model, and then get something that you can put behind a router. The router routes it to, often first, your model, and then Claude—a frontier model, whatever—or Claude or Grok checks it. In a lot of cases, you can get slightly better outcomes at half the cost.
But again, that half the cost—I think a lot of people hear that and they’re like, “That’s bad for AI demand.” It’s actually not at all, because the cost the user pays is just a function of the margin on the tokens. You’re literally just shifting tokens from really expensive tokens with 90% gross margins to tokens with, maybe, let’s call it a 30% gross margin. That’s where the savings are coming from, but the tokens cost the same amount of compute to produce.
Also, all these things are happening at different cycle times. All these big public companies are like, “Oh my God, my AI spend is 20x’ed. I’ve burned my budget in 3 months.” So they set up a router, and that actually cuts their AI spend. But it doesn’t really impact—it may actually increase—the amount of tokens that they are generating just by shifting them to these cheaper open-source tokens. That’s just more compute.
A company getting smarter about which model to use for which task may lead to a stabilization in their spend, or even a decline, but it actually has nothing to do with the amount of GPU compute hours they are effectively consuming behind these model layers of the router. The GPU compute hours probably are going up as you shift to these cheaper tokens that you can use more of.
That’s happening at the cutting edge of public companies. Then you have this whole wave of AI natives, and they are leaning into this so hard. They’re not hiring humans; they’re just really putting it mostly into tokens, so they’re not slowing down. Then you have companies on the East Coast of America that have barely adopted AI, companies broadly speaking in other parts of the country that maybe are just getting started, and Europe, which is just trying to figure out how to regulate AI.
Before using it.
Yeah. So there are these differential waves of adoption all happening at the same time. But the thought I can’t get out of my mind is—I think I said it maybe last time—but, I don’t know, 500,000 people in the world, 250,000 maybe, are using agentic AI, and we’re in an acute compute shortage.
That’s—there are 7 or 8 billion people on the planet. What happens when we go from 500,000 to 1%, to 100 million, to 500 million?
And then I do think it’s interesting. A lot of people are just like, “Okay, well...” I do think it’s helpful to post on X to see the pushback, and a lot of people are saying, “Well, where fundamentally is the—okay, we accept your argument that hyperscalers are under-earning, and as compute reprices, their operating cash flow is going to accelerate, and maybe we could fund this, but where is that operating cash flow going to come from? Where is the customer?”
And, definitionally, it has to either come from faster economic growth through productivity.
Kind of Satya's comments are: either we're going to start growing 10%, or we're going to have labor substitution. For sure, in a lot of these AI-native companies, you're seeing labor substitution, but not because they're firing people. They're just not hiring nearly as many humans. The gross profit dollars per FTE at a16z, Iconiq, and a bunch of companies that have done this work are very high, particularly relative to past generations of startups.
And then it is interesting: are you doing any surveys of your companies and their token spend relative to labor spend?
Oh, yeah. I mean, it's always reported as a percentage of total comp spend or something like this. What ranges have you seen?
I mean, in the really AI-led companies, it gets really high—20%, 25%, something like that.
Well, our friend Dylan Patel at SemiAnalysis—[laughter]—he's an ASIC skeptic, but he's at 30%.
Yeah.
That's probably the highest one I've heard.
I've actually heard of 50%. There's $25 trillion in knowledge work. Let's take your 20% number: that's $5 trillion, and that either comes out of labor substitution or faster economic growth. We really, really, really want, as humans, for it to come from faster economic growth.
One interesting thing I heard this morning from one of the great, leading technology CEOs who's founded several companies is that, if you look at founder-controlled companies and adjust for some of the COVID-era overhiring, nobody's really laying people off. These are the people who would probably be quickest to adopt AI to become more efficient or whatever, but they're not really doing jack aside from huge-scale layoffs, which probably tells you something about where they think there will be lots of opportunity to still have people plus $100K in token spend.
100%. Well, in the bull case—
So growth, not labor—labor growth in the bull case. And you've seen charts from Cognition, Ramp, and Stripe—
—that the companies that are spending the most on AI are growing meaningfully faster.
Yeah, I love that Cognition Index.
Yeah, the Cognition Index is wild. Now, all the skeptics will point out, rightfully, that it's not really controlling for industry. But then, if you dig down into it, I think one of them gave an example of—I forget if it was a plumber or an HVAC contractor—but everybody who's a blue-collar worker is doing great because of AI.
6. The Memory Supply War
By the way, something that I think we should touch on, and we could do it now or later, is that everybody is citing these LTAs. We're in a situation where everything's in a shortage right now. If there's weakness, it's just because we can't energize the gigawatts fast enough.
The gigawatts are going to get energized. Regulatory policy is moving in a good way. The turbine manufacturers and the diesel generator manufacturers are ramping up. You're ripping turbines off old airplanes, reconditioning them, and repurposing them. There are crazy things happening. Capitalism is very, very good at this.
But I do think one of the most important questions in the market—and a transition in the market that I got wrong—is that we are shifting, particularly for memory more than anything else, from crushing numbers in the short term to what they call supply-chain agreements, or long-term agreements—LTAs—where they essentially agree. There are many flavors, but the customer prepays, and there's a floor and a ceiling.
This comes back to the point about labor because a lot of people, after kind of firing too many people during COVID, were really reluctant to lay people off. They talked about labor hoarding. If you remember a few years ago, you remember this? Let's just think about the game theory of breaking an LTA.
There are 4 companies that matter at scale. There's Amazon with its Trainiums, there's Google with its TPUs, there's AMD, and then there's Nvidia, which is much bigger than everybody else combined. Let's just say it's 2027.
Mhm.
It's very important to realize that memory is the—more memory you put with FLOPs for a given unit of compute, the more tokens you get out. It's the single most important thing you could do to increase token output per unit of compute. That obviously lowers costs, which is why demand hasn't responded negatively at all. There's been no elasticity, just because it's the only axis that's dominating all others.
At some level, this is a giant Game of Thrones or IP war between these companies. It's 2027 or 2028, and you're vaguely tempted to break one of these LTAs and try to get a lower price. But, to a large degree, market shares, I think, for the next several years are going to be determined by supply—by supply-chain allocations and by what you have pre-purchased.
If you break the LTA—and this is assuming we're not in a severe oversupply situation, although the logic, almost the game theory, even holds in a severe oversupply situation—and then, in the next 2 or 3 years, for any reason, leverage shifts back to the memory guys, you're out of business.
It's over.
It's over. Let's say Google breaks an LTA. There's an oversupply—I'm making this up—and in 2028 or 2029 they break their LTAs. If they're breaking their LTAs, it probably means there's oversupply, prices are coming down, and then capacity naturally contracts. What do you think is going to happen to Google's allocations? This is a cyclical industry, and oversupply is followed by undersupply. What do you think they think is going to happen to their allocations next time?
Given that this is the axis around which everything is revolving, you might blow up your entire business and your franchise by breaking an LTA. That was never the case before. Apple—who cares? They're buying, and they don't have a competitor. They're overwhelmingly the largest purchaser. Going back 3, 4, or 5 years, they knew they could do whatever they wanted with no consequences because their volume was so big. Even if they super-screw SK Hynix, Micron will of course take them.
This is just different. You have at least 4 players, and you have all the startups you're an investor in, et cetera. If you break an LTA and they just say, “Okay, fine. You broke the price agreement. We're going to break the volume agreement and screw you. We're going to give the volume to your competitor,” you just lost share.
7. Nvidia’s New Playbook
I think Nvidia's dominance—and the extent to which the current environment favors Nvidia—is a little hard for me to understand why it's trading at such a low multiple. In other words, if you need to be able to finance the chips—and you do—nothing's more financeable than an Nvidia GPU. Nothing. If you need to get land and power, they're doing a very good job of playing that chess game in matchmaking.
Then they've rolled out this really clever new business model, which I would describe as kind of like a credit wrapper.
With a revenue share if GPU prices are above a floor.
Yeah. This could lead to them having a really giant cloud business effectively through royalties, really quickly. It is another way of alleviating this cash-flow mismatch: “Hey, we're making all the cash.”
Yeah.
And this isn't really vendor financing because they're not loaning them the money. Somebody else is loaning the GPU buyer the money, so it's not quite vendor financing. They're still making equity investments, but it's not like you're just putting money into someone in return for them. Some of that money is used to buy chips, even though Nvidia said they write into all their equity investments that the money can't be used to buy Nvidia chips. Obviously, money is fungible.
Funny thing—
What's that?
Just like a funny little thing.
Yes. [laughter]
Makes sense.
Yeah, but I think at some level it probably makes everybody feel better.
What would you do if you were the CEO of likely SK Hynix?
I'd do the exact same thing Nvidia is doing right now. I would be going to the buyers of GPUs, Trainiums, and whatever, and saying, “I'll participate in the Nvidia credit wrapper.”
Now, their business is just inherently less stable and predictable, but maybe they just put up some cash up front so they're not on the hook. I'm just making this up, but do something like that because you have money now and credit markets are revolving. I'm sure our friends at Blackstone and Apollo are suggesting some variant of this to the memory companies: “Hey, we will put up some amount of money from our cash flow today, and then it's gone.”
It's equity that makes the person who's extending the debt feel better, but we want some sort of a cut of the ongoing revenues as well, right?
Like, that is 100% what I would do. And it's almost like a logical extension of the LTAs, where they're kind of trading upside for durability here.
You can effectively get a royalty on recurring revenues, and that is what NVIDIA is doing. I do think that is very misunderstood, and I think it would serve NVIDIA well to really explain this one. They're really bullish on AI. Essentially, every time they haven't taken an equity stake in something, it's been a mistake.
They've taken equity stakes in everything, essentially, except the memory companies and, for a long while, Anthropic. Then they took an equity stake in Anthropic. But why not? If you have cash flow and you're bullish on AI—and Jensen sees every lab and knows all the advances, all these continual-learning labs, and Safe Superintelligence is now working with them—he sees everything, and what he sees makes him bullish.
So, one, have some equity upside, and two, have a revenue share. You're generating hundreds of billions of dollars of free cash flow and helping to bridge what is clearly a gap, at least given everybody's gone free-cash-flow negative, until the operating cash flow accelerates enough that you can internally fund this. It's almost like—it's very opportunistic, and it's significant in a good way. It significantly increases their revenue per gigawatt, and it also strengthens their competitive position.
You and I both have startups, but okay, that's great—use that startup's chip. What prices are they paying at TSMC? Higher than NVIDIA and all these guys. What prices are they paying for HBM DRAM? Higher. Can you finance those chips easily at the same rate as NVIDIA? No. And so there's a real burden, particularly if you use HBM DRAM. You're in the crosshairs of this, unless maybe they made really different—
Architectural—
Architectural choices. Everything that's happening is actually pretty good for him.
Just going back to game theory, Anthropic, if they had been as aggressive on compute as OpenAI had been, they would have run away with it.
Yeah.
And so now OpenAI is back in the game. I think Grok is in the game. Those are the companies on the Pareto frontier.
And they have the compute.
And do you think, after watching that, anyone is going to let off the gas, especially if it could be funded out of operating cash flow?
Right, because it was, I think, 4 months ago that Dario was talking about how it's really, really hard. If you buy too much compute, you could go bankrupt at the scale of these things, but if you don't buy enough, you could lose.
Well, we saw what happened.
OpenAI just got back into the game, and now SpaceX is in the game in a big way with Grok 4.5 and Cursor.
After watching that, from a game theory perspective, is anybody going to back off anytime soon?
Have you met anyone in your travels out here that you would say is way more bullish than you? And if so, what do they believe that you don't?
Essentially everyone out here is more bullish than me, man.
[laughter]
I read this thing that [name unclear] wrote, and I was like—
The 3x compute-price thing or whatever.
Yeah. Well, he was—I forget what it was.
No, it was like 15x or something.
Yeah. But basically, renting an H100 for a year would cost $250,000.
And that's 15x the current spot or something.
Exactly. Like, wow. That was just like—that was in my book. That wasn't in my Bayesian probability space of expected outcomes. That wasn't even in my considered-but-dismissed-as-totally-unlikely outcomes.
And then that guy is very dorky. He's a very smart guy. He's very plugged in. And then he pointed out that something like—I think he just said margins on compute are going up, the amount of compute is going up, and inference margins are going up. If you multiply those 3, that's how you're getting this crazy acceleration in the sum of the labs plus open source. Although, obviously, the margins on open source are not really going up.
[laughter]
Yeah. I look at what's happening in the stock market, and I feel like a foolish optimist. Then when I talk to people—whether it's people at the labs or anyone in this ecosystem—I'm bearish relative to essentially everyone, which is just a strange state of affairs.
What do you make of the DUV news out of China? I've seen reactions really along a spectrum, from this being the equivalent of what ASML had in 2001 to this being the first bit of news in a new story for how we should think about the global supply of cutting-edge compute.
I think both could be true. Let's make an analogy. Let's say a DUV machine is like a propeller plane and an EUV machine is like a jet turbine—or whatever it's going to be. They didn't have it before, and now they allegedly do, and that is a phase transition. You've gone from liquid to solid.
Now, that solid—that jet engine, prop plane, whatever—is 25 years behind, but still, it's important, and I don't think it should be dismissed. But I also think it's kind of funny. You just see this in the stock market. The stock market massively overreacts, and if this ever hits ASML's orders, maybe it hits them in 5 years.
The market has forgotten about it, gotten worried about it, forgotten about it, gotten worried about it, forgotten about it multiple times along the way. So I do think that was probably an overreaction, but we shouldn't dismiss that either. If you're China, this is really important to you.
There are some reports that an EUV machine had been smuggled into China. And I mean, what a feat of espionage, because those things are giant—huge.
I don't know if that's true. There's some noise about it, but China is really, really good. They're really, really smart. They work brutally hard, and they see this as super important for them as a country.
But are they going to go from the year 2001 to 2026, or even 2030? It really is—
It's a learning by doing.
It's a learning-by-doing process, and you kind of have to—yeah. You can't—
You can't accelerate the doing. You can't teleport into the future. You actually have to go through those learning cycles. So, is it significant? Yes. Did the market overreact? Probably.
But I think it's very hard as an American to really understand what is happening in China and have total conviction and clarity. For better or worse, we are decoupling, and that is a process that has been set in motion. At this point, it almost feels like it's kind of self-reinforcing on each side, and that's unfortunate.
But we are where we are, and they're not going to stop. Neither are we.
Any commentary on every other company in America? I feel like right now it is 10 companies, plus a couple of private companies.
Well, not last month. I mean, everything but AI was vertical.
I do think open source getting closer to the frontier, and companies like Fireworks making it really easy to customize a model such that you can get, in some cases, better-than-frontier performance for meaningfully lower cost, is a godsend for the software industry. It’s also a godsend for all these AI natives.
Our friend Eric Vishria, I think he said 2 years ago, “I’ve never seen more companies go from being founded to $50 million a year in revenue and generating cash flow in 9 months.” It’s hard to know if any of them were durable because, back then, a lot of people would dismiss them as ChatGPT wrappers.
Well, now with open source, you’ve actually generated some data that’s unique to you and your use case, whatever the vertical you’re going after as a wrapper is. Fireworks did come out with a really cool product called Nexus. If you’re using Claude Code, OpenAI Codex, or Grok Build, it is literally 3 lines of code, 20 words, and Fireworks ingests your data. They can RL a model, and then there’s a router that sends the query, and they’ve had amazing results.
This is kind of the solution for every AI native. That’s why you saw Harvey, before it was acquired, and Cursor lean so heavily into this—Harvey, Legora, all of them—because if you can go from just using 1, 2, or 3 frontier models to using—
Whatever’s optimal.
Those frontier models for whatever it is, 30% to 60% of your token consumption, and then use your own RL model, all of a sudden you’re not a wrapper. You’re way more defensible.
I was so interested by that Cursor thing that came out. I think it was Cursor where AI is sort of speedrunning what we’ve learned amongst humans: you could use the frontier model to plan and then farm out tasks to the dumber models.
And it’s 15 times more efficient, or whatever the metric was. It may be that—and this is super ironic—it may be that lower-margin, open-source tokens that are just a little bit behind the frontier are more valuable.
We have friends who believe that once a frontier model hits ASI, it will actually have a dramatically lower cost to serve at every level of intelligence by distilling this, and then there’s no place for open source. I would say that’s an Anthropic, OpenAI, Grok maximalist view, and we shouldn’t dismiss anything. Anything is possible. We want to be very humble. I particularly want to be humble after the month I’ve had. But that doesn’t seem that likely to me.
Why?
One, because there are so many of these AI natives that have actually generated a decent amount of domain-specific proprietary data.
Yeah. Before open source had this moment, and these inference clouds and routers really developed, you kind of didn’t have a choice. Whatever the terms of service were, you accepted them. But if you can now get off that treadmill, that gives you a degree of independence, maybe durability and safety.
Going back to your point, it may be that these cheaper tokens just massively inflate the value of the most cutting-edge frontier tokens. If today you have—I’m going to make this up—120-IQ open-source models, and they’re really cheap to run, doesn’t that make a 160-IQ model that can orchestrate them more valuable?
And so, just—we talked last time about how I’ve been really surprised that so much of the economic returns accrued to the frontier. Now that is changing with what we’re seeing with these inference clouds—Together AI, Modal, and Baseten—in a very cache-efficient way.
What’s shocking about those business models is they’re growing almost as fast as the frontier labs did in the early days, but burning very little cash.
Right. It’s pretty extraordinary. To go back to silly SaaS metrics, from the Rule of 40 perspective, these are crazy numbers.
Do you think there’s a lot of instruction in just the distribution of pay inside of an organization? The CEO makes X times more than the median person at a company. Maybe that’s frontier tokens versus open source.
Absolutely. It may be that what we discussed last time happens: frontier tokens may lose a little. The pie is growing really, really fast, so they may continue to capture the overwhelming majority of economic value, but not all of it the way they have been.
Open-source tokens might be the majority of tokens processed. Again, going back, that’s great for infrastructure demand because a token is a token, and it takes the same amount of FLOPs, watts, space, and cooling to make.
What’s the worst thing that could happen in AI? Is it regulatory? Is it some sort of—
I think regulation has to be the biggest risk. It’s the most obvious risk.
That was one reason I was excited to be here this week: I want to be scared. I don’t want to feel like a lunatic watching these stocks get cheaper and cheaper, thinking the expected forward returns are going up while it feels like the on-the-ground fundamentals have materially improved in July relative to even June.
But I still come away thinking regulation has to be the biggest risk. You just can’t ignore New York making a data center moratorium, and we’re living in this weird post-factual, post-logical political world.
I think the AI industry has done a terrible job of PR, and I do think—
I think it at least realizes that now.
Maybe if it’s not fixed, it realizes it.
Yeah. The narrative in Washington—the political narrative amongst a lot of ordinary Americans—is: data centers are going to raise your electricity prices, they’re going to take all your water, and then they’re going to take your job.
The reality is that, given the deals being cut now, when a data center goes in, electricity prices actually generally go down for everyone around there because of behind-the-meter deals. This is that data center pledge that Trump asked people to sign.
Generally, the data center developer—it used to be that they just had to get the police department and the fire department new trucks and new cars, and new body armor or whatever. Now it’s, “We’re going to build you a hospital, a school, a new police station, and a fire station, and we’re going to lower your power bills. How does that sound?”
By the way, the jobs are ongoing because it turns out that you need these plumbers, electricians, and HVAC contractors. Data centers are, in a lot of ways, the best thing to happen for blue-collar wages in my lifetime.
And yet you have the Democrats, who ostensibly represent these blue-collar workers, taking those jobs away. It’s just kind of wild how—a lie could go around the world—
Faster than truth gets out of bed.
Yeah. An author made a mistake in a book and overestimated the amount of water usage in data centers by 10,000 times. Not a little bit—not 1 order of magnitude, not 2 orders of magnitude, not 3. She’s admitted that mistake many times: “I was completely wrong.” It’s been super debunked.
Did you ever hear the Popeye effect?
No.
Popeye ate spinach. The reason was the same deal: in an academic book, they misplaced the decimal 1 place. Spinach does not have more iron than everything else. It was just this 1 source, and then that propagated, and people still say it has more iron.
I literally thought it had more iron.
That’s wild. I literally thought spinach had more iron.
That’s amazing.
Crazy.
Yeah. You learn something new every day.
Same thing, though.
Yeah, it’s the same thing.
Somebody just needs to tell the truth. I feel like the industry—and I thought, maybe if nobody else is going to do it, I’ll do it. There needs to be some sort of foundation. Maybe it’s a PAC that runs ads during the Final Four, during NFL games, during college football games—
Here’s what a data center does: your power—
A data center that signed this pledge in your community—
Yeah, your power prices are going to go down. They’re almost certainly going to contribute to the community in a material way. You’re going to see a massive influx of super-high-paying blue-collar jobs that are going to persist.
And I think a lot of people thought that they were one-time, and they're just not. There's for sure a spike, and then that moves to the next data center. But there is an ongoing need for R&M and upgrades at these data centers, and technology is changing. So you're going to have more jobs, cheaper power, and a wealthier community. There's going to be no impact on water, no impact on the environment, and it's easy to build the data center 10 miles out of town.
That story needs to be told along with the stories about how AI is increasingly saving lives and curing rare diseases. I think—we talked about it last time—I can't remember if it was at ASCO this year, but the vibe was, “Hey, this is the most scientific breakthroughs we've ever seen at a single conference,” and for sure some of that is due to AI. We need to tell those stories. If you have a sick child, a sick parent, or a sick loved one, AI meaningfully increases the odds of them recovering.
We just need—everybody needs to tell this. I think people out here find all of this so blindingly obvious that they assume everyone else already knows.
They can't process that this is a true but wildly divergent view from most Americans. I think the industry really needs to tell its story better, because New York just feels like the first of many. Even in some of these deep-red states that are super pro-growth, they're just like, “Hey, you guys are not doing a good job telling your story. We can't tell your story. If you tell your story, though, we can retell it, but you're the experts.”
If you don't speak your own truth, who else will?
Yeah.
What have we missed, then?
I do think something that's missing from all of this conversation about compute is what is going to happen when you put these SRAM-based accelerators that are not constrained by HBM DRAM and are often made on older nodes that are not competing with the latest, greatest GPUs. When you disaggregate inference, people talk about prefill and decode, but decode is 2 parts: attention and feed-forward network. The ultimate holy grail is if you could do prefill on 1 chip that probably doesn't have HBM DRAM, do the attention on a super-high-powered chip with HBM DRAM, and then do the feed-forward network on one of these SRAM chips.
But the ROI on adding these SRAM accelerators to the existing installed base of compute and new compute is that you do better. You just can't beat SRAM, in particular for that feed-forward network, and you almost can't, no matter how much you try to get the ratio of compute to HBM DRAM to SRAM on the chip correct. The workloads are always changing, and there are different workloads. Being able to disaggregate into these 3 parts, I think this is going to be really, really positive for the ROI on AI.
For some reason, I just thought of a funny question. I love the framing of Game of Thrones versus all these people. Can you imagine a player who is not currently on everyone's mind becoming relevant at the major Game of Thrones scale? That could be Micron all of a sudden, as a sample answer to the question of someone who becomes as important as Anthropic, OpenAI, Microsoft, Amazon, and Nvidia. So, what dark-horse Game of Thrones names come to mind?
Lee Buu is probably a dark horse. I do think Lynn at Fireworks—she is just an absolute killer. I think our friend Scott Woo—Cognition is kind of like—
You're high on that one?
Yes. I think those are the most obvious names.
What about SpaceX? What's it been like watching that be digested by public markets, at least initially? Do you think the market understands it as a company—the most important new company to be public?
It doesn't really feel like it does, because everything to me is—the fundamentals have gotten better since it IPO'd. Grok 4.5, the Cursor acquisition—Cursor has clearly accelerated meaningfully—and then they have shown that they can bring on more compute faster than anyone at lower prices. Now we know that they can even, adjusting for the spot-versus-contract gap—their big advantage was they came into the market and just hit those spot highs.
In a strange way, one of the more bullish things for compute is that they put a vast amount of compute into the market overnight, and it wasn't even really a blip. The market just utterly absorbed it. The freight train didn't slow down at all.
But a Substack writer reportedly thinks that SpaceX is going to try to bring on 8 gigawatts of compute. I will never bet against Elon, but that would be a truly incredible feat. Their rates have gone up since they signed those loss contracts, not down, and they're monetizing at something like 50 billion a gig. Consensus estimates for next year are 73 billion.
Forget Starlink V3, forget likely Starlink Direct-to-Cell, Grok 4.5, and Cursor. The sum of that probably hits a $10 billion ARR pretty quickly. Forget all of that. Forget the core-base Starlink business. If they bring on anywhere near that, the consensus estimate is 73 billion, and that's 8 gigs at 50 billion a gig. Obviously, that would not all be lit up at the beginning of 2027.
And it seems very implausible to me. I almost don't believe the Fund AI report.
But to this day, the only companies that have brought on more than 500 megawatts of power in a year are the hyperscalers, CoreWeave, Crusoe, and SpaceX. SpaceX has brought it on both the fastest and at the lowest cost. People do actually really like their clusters. But again, the market is going to need to see that.
That would not be the market's interpretation of SpaceX today.
No. No. And it does feel like there's a big New York hedge fund short case on it. I think they think, “Oh, the spot price for compute is going to go down 90%, and you're going to bring on all this compute. It's not going to generate nearly as much revenue as you think.”
Maybe. But I also want to be really clear: I have seen Elon's companies do really impressive things over the years. Bringing up the public report, which says 8 gigawatts in 18 months—I'm just quoting that because it's public. It's available to everyone.
Yeah.
One of Elon's phrases is, “We specialize in making the impossible late.”
I never heard that. That's great.
Yeah. There's a lot of truth to that.
Yeah.
But I just think very little is built into that stock, from my perspective, for the amount of compute that they might be able to bring on. Again, I don't think it's anywhere near 8. It's going to be really hard. Energizing these GPUs is really hard, but they've been good at it. It doesn't feel like that's in estimates or really in people's thinking.
I'm thinking about that funny meme that says, “SpaceX, the data center company.”
Oh, no. 100%. Yes, absolutely. I did spend a lot of time at Starbase, and orbital compute feels more real every day. Pretty cool to see Starship landing the other day.
Pretty cool to see the Starship landing. It's funny: our friends at Benchmark funded Starcloud, which is an orbital-compute company that SpaceX is kind of partnering with. They're going to, I think, let them use the Starlink laser technology, which is really important for orbital compute.
I do think that's a good sanity check. Last time I checked, the Benchmark guys were pretty smart, and they're not coming from the Elon ecosystem at all. They chose to fund an orbital-compute company at a decent valuation without the internal launch cost that SpaceX gets. That's, to me, a good—hey, am I crazy?
Am I crazy? Well, maybe I'm crazy, maybe Elon's crazy, maybe Benchmark is also crazy, and maybe the SpaceX engineers are also crazy. But, man, that just doesn't seem that probable to me.
Should we say whose offices we're in? We're sitting in the Benchmark office. [laughter] Yes, this is their famous table for their famous dinners. We will see where all of these stocks are in a year.
And the great thing is, time will tell.
You know, people are going to be right or wrong. The future's probabilistic, but we are at an exciting moment.
Well, if we keep doing this on the model release cycle, I'll see you in a couple of weeks.
Yes.
Yeah. Maybe here again at Benchmark. That's always a blast to do with you.