瓦特、晶圆与AI基础设施的未来 | Gavin Baker
- 3/4月是一次“加仓型回撤”,而不是“吃药型回撤”:纳斯达克下跌之际,Anthropic单月新增110亿美元ARR,相当于Palantir、Snowflake和Databricks花10年建立起来的业务总和。“资本主义历史上从未发生过这样的事”,而科技股相对大盘的估值也降至过去10年任何时候都未见的便宜水平。“3月你唯一要做的,就是观察Anthropic正在发生什么。”
- 传闻9000亿美元估值、对应500亿美元ARR的Anthropic,比表面看起来便宜:Baker认为,算力约束和降级后的Claude(Opus针对同一问题生成的token减少约70%)掩盖了其“无约束运行率收入”1000亿—2000亿美元的水平——因此投资者实际支付的可能约是5倍“URR”。Anthropic本可以按“至少高出100%的溢价”融资,但它遵循Elon的教训:“估值上永远不要贪心”,换取未来20—30年的资本获取能力。
- 唯一能证明泡沫存在的信号,是TSMC的产能决策。如果TSMC满足Jensen的要求,Nvidia在2026或2027年可以卖出“2万亿美元的GPU。也许2.5万亿,也许3万亿”——这就进入过度建设区间。历史经验(Carlota Perez、运河、铁路、2000年)告诉我们应当预期泡沫;本轮的抵消因素是,建设仍主要由经营现金流融资,且GPU利用率为100%,远高于99%闲置的暗光纤。要关注Intel或Samsung打破纪律:“它们不会一直保持纪律,迟早会破。”
- 轨道算力是“太空中的机架”,不是五角大楼大小的数据中心——一台Blackwell机架重3,000磅、功耗100kW;预计Starlink V3的运行功耗为20kW,SpaceX的目标是让功耗100—120kW的卫星通过激光在真空中互联。瓦特短缺将在2027—2028年缓解,随后轨道算力解决这一问题——这直接警告陆地电力与冷却供应商:它们扩张的产能落地时,“正好赶上所有愚蠢的怀疑者开始明白轨道算力确实存在”。
- 前沿token正在攫取模型层绝大多数价值——Gemini 3.1 Pro从“令人震撼”变成“无法忍受”,而帕累托前沿已经翻转:9个月前由Google主导,如今是Anthropic和OpenAI,Grok 4.3也处于前沿,Gemini则勉强维持,很可能是在“出于骄傲进行补贴”。与此同时,AI正从“随便吃到饱”转向“按杯付费”(他的电信分析师类比),这可能正是OpenAI和Anthropic今年ARR应超过“远超200”的原因。
- Prefill/decode分离可将GPU寿命延长至10—15年——把Cerebras或Groq的LPU放在Hopper和Ampere前面,让它们运行到“烧融为止”——这“可能单枪匹马拯救私募信贷”,并将GPU融资成本从CoreWeave的低7%降至5—6%,从数学上降低整个建设周期的成本。
- 横截面估值“完全不讲道理”:半导体设备股按下一季度年化利润计算为40倍,而DRAM只有中个位数倍数(上一轮周期峰值差距约为5倍对12倍);质量最低、成本最高的“卖短缺”公司受X平台资金流推动股价冲天,优质公司反而落后。AI内部相关性在1月彻底瓦解——你已经无法再用半导体设备对冲存储器——而被错误归类的公司,如被塞进“铜失败者”篮子的交换机公司Astera,正是机会所在。“我只是希望AI空头能再多一些。”
- 决定一切的有3个问题:前沿token溢价能否持续,苦涩教训能否经受ASI的检验(“谁知道苦涩教训对400智商模型是否仍然成立”),以及持续学习何时到来——“如果我们得到它,起飞速度会非常快。”此外还有黑暗尾声:围绕AI的政治暴力正在上升,以及“机枪已经到来……如果我们不全都成为机枪大师,就会被机枪主人掌控。”
1. 3月是积压阿尔法——资本主义历史上最非凡的月份
- Baker的分类是:有些回撤意味着“你的假设被证伪,必须吃药”,也有些回撤意味着你极度不同意价格走势,可以“积累积压的阿尔法”。3月属于第二类——纳斯达克下跌之际,市场正进入他所称的美国商业史上最非凡的时刻。
- 证据是:Anthropic单月新增110亿美元ARR。Palantir、Snowflake和Databricks——可以说是过去10—12年成立的3家知名度最高的SaaS公司——花了10年、动用了数万名员工才建立起各自的业务。“Anthropic在1个月内加上了它们业务的总和……资本主义历史上从未发生过这样的事。”节目中Krishna提到的统计数据——听起来像是“DRAM增长500%”(原话不清)——若连续复合3年,结果将荒谬至极。
- DeepSeek的相似之处在于:2025年,这篇文章发表于“DeepSeek星期一”前7天;等到股市崩溃时,已经“非常清楚这将是算力需求历史上发生过的最积极的事情”——DRAM价格垂直上涨,亚洲AWS可用区价格翻倍。那笔交易需要做一些功课;而今年3月,“你唯一要做的,就是观察Anthropic正在发生什么”。
- 他对宏观的逆向判断是:霍尔木兹海峡关闭“实际上对美国相对有利”——Bloomberg上的美国天然气价格下跌20%,而亚洲和欧洲天然气价格翻了两三倍,一夜之间改善了美国相对制造业竞争力,这正是政府最关心的事。与1970年代不同,如今美国既是全球最大石油和天然气生产国,也是最大出口国,因此更容易继续聚焦科技股——它们相对估值处于10年来最低水平。
2. Anthropic约为5倍“URR”——以及Elon关于永远不要推高估值的教训
- 从资本效率角度看,OpenAI和Anthropic是“完全不同的动物”:Anthropic为达到大致相同的收入规模,烧掉的资金可能少80%,这意味着两者的ROIC存在结构性差异。不过,他称赞Sarah Friar是“最出色的CFO之一”,并指出OpenAI积极锁定算力“确实收到了成效”。
- 他提出的新指标是URR,即“无约束运行率收入”。Anthropic“显然降低了Claude的智能水平”——分析显示,Opus针对同一问题生成的token甚至减少了70%,而“token数量等于答案质量”——因此,如果拥有所需的全部算力,Anthropic的收入将达到1000亿美元、1500亿美元,“也许2000亿美元”。若以9000亿美元对应500亿美元ARR的估值计算,“你买入的可能更接近5倍URR”。
- 为什么不以3万亿美元估值融资1000亿美元?因为Elon的超能力——随时随地筹集所需资本——来自“估值上永远不要贪心。永远不推高估值。就这么简单”。他的朋友Antonio指出,SpaceX曾连续10年实现30%出头的年复合增长率。Anthropic当然可以按“至少高出传闻最新估值100%的溢价”融资,但这种做法可以创造持续20—30年的好处。
3. 瓦特:分区审批成新瓶颈,轨道算力就是太空中的机架
- 在“没有重大监管或政治反弹”的情况下,资本主义会解决瓦特短缺,但他认为出现这种反弹的可能性确实不小。一家大型PE机构的数据中心基础设施负责人说:“过去能源和芯片是最大的制约因素,现在变成了分区和审批。”企业可能等到中期选举后才处理可能的裁员问题——“没人想当靶子”。涡轮机约束是真实存在的:两台这样的机器就能铸造大型叶片,而西方80年来都没建过能做这件事的设备;不过他预计短缺将在2027—2028年开始缓解。
- 随后他坚持重新定义一个问题:轨道算力不是太空中的五角大楼大小建筑。“它是太空中的机架。”一台Blackwell机架重3,000磅,尺寸为8英尺×4英尺×3英尺,功耗100kW;卫星大致就是这个量级,配有约500英尺长的太阳能翼,在太阳同步轨道运行,后方拖着一条延伸数百英尺的散热器。机架通过真空中的激光互联——每一颗Starlink上都已经有这套技术——构成虚拟数据中心。在太空中优化的是重量,而不是体积;在地球上则是“能用铜就用铜,非用光不可时才用光”。
- 可信度来自两点:SpaceX运营着在轨卫星总量的98%—99%,同时拥有地球上最大的数据中心;预计Starlink V3的运行功耗为20kW,而他们有信心“直接做到100—120kW”。对于怀疑者,他借用Larry Ellison的话:“他在外面降落火箭。我没看到其他人降落火箭。”但诚实的限制仍然存在:如果机架在轨道上坏了,关于维修的问题是——“除非你有一台漂浮的Optimus,否则做不到。”
- 可投资的边际在于:推理会进入轨道,而训练“很长时间内”仍留在地球;陆地数据中心“在我的有生之年都很有价值”。美国“会尽可能榨取每一种能源”,算力也一样。但如果你是一家“正大规模扩充电力和冷却产能”的供应商,而这些产能落地时“恰好赶上所有愚蠢的怀疑者开始明白轨道算力确实存在”,那就值得认真、长久地想一想。
4. 晶圆:TSMC可能单枪匹马阻止泡沫,TerraFab即将到来
- 晶圆约束在文化上具有决定性意义:TSMC管理层把自己视为“Morris Chang神圣遗产”和硅盾的继承人。20年前在Science Park,他们告诉他,追上Intel是“一个多么美好的梦想,但那是我们孙辈的梦想”——而他们做到了。Jensen则“从未与Taiwan Semi签过合同”:双方依靠握手,在看起来公平的条件下做生意。
- 每项基础技术——运河、铁路、互联网——都会制造泡沫,这是Carlota Perez的框架:市场正确识别技术,所有人转为看多,供给过度,随后崩盘。其间会出现他所谓的多样性崩溃。本轮的根本差异在于:建设“仍然主要由经营现金流提供资金”,估值更理性,而且每块GPU的利用率都是100%,相比之下99%的光纤处于闲置状态。但“根据过去200年的历史……我们应当预期泡沫出现”。
- 他反复提到的警示故事是George Vanderheiden,这位Fidelity最伟大的基金经理之一曾对抗1999年泡沫——40%仓位是烟草,40%是住宅建筑商——2000年初因客户说“George,你落后于时代”而退休,随后3年跑赢纳斯达克“约20倍或30倍”。Vanderheiden有一句话:“过早做对,和做错是一回事。”
- 唯一需要观察的是TSMC的产能决策。“如果Taiwan Semi按Jensen的要求行事,Nvidia在2026或2027年可以卖出2万亿美元的GPU。也许2.5万亿美元,也许3万亿美元”——足以触发过度建设。如果泡沫没有出现,“我们就该为他们举办一场庆功宴,因为他们将单枪匹马阻止一场泡沫”。风险在于Intel或Samsung——“它们不会一直保持纪律,迟早会破”——而TSMC在领先节点上领先9—15个月,供给不足第二来源与放任其过度建设之间存在一个黄金区间。
5. TerraFab:顶级团队追随Elon
- 他认为TerraFab——SpaceX与他相信还包括Tesla在内、旨在美国建设全球最大晶圆厂的合资项目——会成功,理由有3个:Intel的合作带来50年的机构知识,进度“只落后前沿节点3—5个季度”;半导体设备领域的顶级团队(ASML、KLA、Lam、Applied)会像当年前往台湾一样为Elon到场——“他们不喜欢买方垄断”;此外,TerraFab的差异化足够大,不会疏远TSMC。
- 人才是第三根支柱:Elon“在中国、台湾、韩国和日本是活着的神”,Baker预计德州会出现真正的台湾城、日本城和韩国城——连最受欢迎的餐厅都带着员工整体搬迁过去——因为“最好的工程师想为Elon工作,尤其是在硬件工程领域。这不是掌管Intel和Samsung的人所拥有的思维方式。”
- 对于建设周期,他的态度很轻松:“其他人都要花3年建设数据中心。他122天就建好了。”Samsung甚至不得不在其德州晶圆厂给Elon安排一个办公室,因为他对进度慢到极度不满。
6. 3个问题:前沿溢价、苦涩教训与持续学习
- 本轮周期最令人意外的地方是:模型层经济价值“绝大多数”都流向前沿token,因此必须判断这种状态能否持续。他自己的经历让他保持开放态度:Gemini 3.1 Pro刚发布时“令人震撼”,如今却是“无法忍受,无法忍受”。企业会在前沿模型上做原型,有时再部署开源模型,但前沿溢价目前是事实。
- 帕累托前沿,即智能与成本之间的关系,是他分析AI实验室最重要的单一视角,而这个前沿已经翻转:9个月前由Google主导;如今Anthropic和OpenAI占据主导,Grok 4.3位于前沿,是“成本最低、效果最好的5000亿参数模型”,而Gemini 3.1正在“勉强维持”——“如果让我下注,我会押它是在出于骄傲进行补贴”。根本原因延续上期讨论:Google在TPU v8设计上过于保守,放弃了单token成本领先地位,而Nvidia仍保持激进。
- 整个交易面临的最大风险,是违反Richard Sutton的苦涩教训——人类聪明才智胜过蛮力算力。3月的惊吓事件“turbo quad”,是一篇Google一年前发布的内存优化论文,在Micron、Samsung和Hynix就DRAM长期供货协议谈判期间发布(“人们做什么永远比说什么更重要”),随后以“DRAM完了”的标题走红;但他“找不到地球上任何一位相信turbo quad会影响DRAM需求的AI工程师”。
- 他与实验室的分歧在于:开发者怀疑会出现违反苦涩教训的情况,但他认为“我们距离ASI已经非常近了。谁知道苦涩教训对400智商模型是否仍然成立”。如果达到ASI,它最先想做的事情可能就是变得更聪明、获得更多资源;它可能让自己更高效,而“苦涩教训从字面上也包括人类在内”。第三个问题是持续学习——今天粗糙的版本,只是在中期训练阶段针对可验证任务进行RL;人类碰一次火就知道烫,模型却“需要把手伸进火里100万次”,然后由设计者把火嵌入下一代RL训练环境。动态更新自身权重意味着“起飞速度会非常快”,而人们似乎相信它“就在眼前”。
7. 按杯付费:为远超200 ARR的定价转向提供资金
- 他以2005—2007年的电信分析师框架来解释:在固定费用加使用费的定价模式下,移动通信曾是一个优秀的成长行业;当所有人转向包月不限量后,它就不再是优秀的成长行业。“AI正在从随便吃到饱转向按杯付费”——而人们确实很喜欢使用AI,尤其是“现在一个人可以让100个agent同时工作”。这一转变加上新增算力,可能正是OpenAI和Anthropic今年ARR应超过远超200的原因。
- 对投资者的实际启示是:每月250美元的套餐已经无法让你接触前沿——“你受到了严重的频率限制。你拿到的是被切除部分大脑的AI版本。”要理解前沿AI能做什么,即便是非编程工作,“你需要Claude Code或Codex,还需要企业版”——按使用量计费,因此模型会生成它认为答案真正需要的token。他也提到一个令人不安的事实:“这对世界来说很悲哀……如果你负担不起,你就不在前沿。”
- Harness的重要性比他之前暗示的更高:“harness engineering没有模型重要,但确实非常重要”;而harness和模型正在越来越多地共同开发——包括上下文、记忆、状态和工具。“即便是简单版本,也会带来惊人的差异。”
8. 芯片初创公司:做出不同且困难的东西,同时拯救私募信贷
- 芯片设计和坦克设计一样存在“铁三角”:攻击、防御、机动性——Merkava优化防御,俄罗斯坦克优化机动性;这些取舍受物理规律约束,并体现在TSMC的设计规则中。TPU、Trainium和AMD本质上都在“努力成为更好的GPU”。Trainium目前可能做得最好(“它们正在拽Superman的披风”),但Trainium 3必须扩大规模,因为MoE推理在经济上需要其交换机扩展网络;至于AMD的MI450,“我们拭目以待”。
- 他的经验法则是:1%市场份额价值1000亿美元,已经是很好的风投结果;但Jensen的潜台词是:“如果有人做出不同的东西,并拿到1%、2%或3%的份额,我们就会做出那块芯片。”因此门槛是既不同又困难。Prefill和decode分离打开了新的设计空间——他的同事Andrew Fox用一句话概括:“Prefill是在装填炮弹,decode是在开火。”Prefill受内存容量约束,decode受内存带宽约束。至于初创公司所谓“获得TSMC特殊工艺支持”的说法,永远不要相信:“Jensen在Taiwan Semi还只是眼中一闪的念头时,就已经看过那道工艺了。”
- Cerebras是他眼中“困难且不同”的范例——晶圆级计算,历经3代芯片才做对,并试图通过在顶部放置一块光学晶圆来解决其shoreline-IO约束。“Cerebras IPO之后,所有人都会获得融资。”
- 更具宏观意义的判断是:分离意味着GPU可以使用10—15年——把Cerebras系统或Groq的LPU(“被Nvidia收购的那家公司”)放在Hopper、甚至Ampere前面负责Prefill,让GPU运行到“烧融为止”。这“可能单枪匹马拯救私募信贷”,因为私募信贷原本按3—4年GPU寿命承保;将融资成本从CoreWeave的低7%降至5—6%,将“从数学上改变整个建设周期的融资成本”。此外,他通过Coatue的朋友Jamin提到,卖短缺的一方正在击败买方;而在agent驱动的世界里,CPU的重要性远高于过去(编排、工具调用),超大规模云厂商庞大的CPU装机量可能因此收回一部分优势。
9. 应用与开源:token路径和新的囚徒困境
- 在应用层,“别再谈价值归属——AI已经摧毁了数万亿美元的价值”,即使把Cursor和Cognition算进去也一样。在Jensen的五层蛋糕中,利润流向能源、数据中心、芯片和模型,而不是应用。他筛选项目的标准是:在你能做大规模之前,这个想法是否已经对全世界显而易见?Amazon的电商统治消灭了所有风投支持的仿盘;Wayfair活下来,是因为“他们做了困难的事”。AI创始人“真的很挣扎”,他们押注于利基数据壁垒,但模型公司可能直接碾过去。
- Baker认同Jamin Ball的规则:“如果你不在token路径上,也不属于某个真正的利基领域,日子可能会很难。”反过来,编程是“正义的聚焦”——Cursor、Cognition和Anthropic在OpenAI什么都做的时候选择聚焦;Replit创始人的说法也一直留在他脑中:编程可能是“通往ASI的最短路径”,因为有了代码,你可以为自己写出任何东西。如果前沿token的回报率最终压缩,“应用层将出现价值创造的大爆发”。
- 如今的开源前沿是“用偷来的美国token训练出的中国模型”——他告诉过DeepSeek,其中一个版本只使用了15万条推理轨迹;中国公司可以通过API以许多方式洗白这些数据;美国实验室正在竞速研发反蒸馏技术。他认为,这也是Mythos没有得到广泛服务的部分原因:除了算力短缺,“他们不希望它被蒸馏”,不如自己蒸馏,再用RL训练下一代模型。
- 新的囚徒困境是:你是否应该通过API发布自己的前沿模型?如果所有实验室都不发布,中国开源模型就会追上来;如果有一家背叛,它就能拿走收入,而“资源等于智能,所以它们会开始领先”——这与TSMC、Samsung和Intel面对的博弈相同。Jensen“可能随时都能让自己的模型非常接近前沿”(可能是Nemotron),这就是让开源始终落后一个固定距离的合理反制。开源也不是免费的:token需要消耗能源,而模型公司“几乎总会拿到收入分成”。
10. 多样性崩溃、失灵的篮子与超大市值股成绩单
- 他担心的是“多样性崩溃”——“我不知道还有谁像我一样真正看多DRAM。一个都没有。”
- 从横截面看,“估值完全不讲道理”:半导体设备股按下一季度年化利润计算为40倍,而DRAM只有中个位数倍数(上一轮周期峰值差距约为5倍对12倍)。
- Nvidia的相对估值接近10—12年来最低,却无法与GE Vernova的估值倍数相匹配。与此同时,短缺经济意味着质量最低的公司表现最好,正如牛市中的高成本大宗商品生产商一样,被X平台上的散户账户“竞价推上月球”。2024—2025年那些“核能泡沫和量子泡沫”的荒谬行情,已经转移到小盘股。“我只是希望AI空头能再多一些。”
- 从结构上看,AI交易在1月停止作为单一因子交易:scale-up网络暴涨,scale-out下跌;DRAM大幅跑输NAND和HDD;“你已经无法再用半导体设备对冲存储器”。
- 他现在最喜欢的猎场是被错误归类的公司:Astera被放在“铜失败者”篮子里,但它即将推出的最大产品是一款交换机——“从定义上说,如果你是一家交换机公司或加速器公司,就不可能是铜失败者。”
- 成绩单方面,Google失去了TPU优势,但拥有最多算力、对机器人真正有价值的YouTube数据,以及“正在疯狂增长”的GCP。不过,如果5天后举行的Google I/O不能至少稍微领先OpenAI或Claude,“那就很有意思了”。Zuckerberg获得“巨大的赞誉”,因为他是唯一一家真正实现内部AI优先转型的互联网巨头;MSL的首个模型已经接近帕累托前沿,而在3年视角下,“变化速度比当前水平更重要”。Amazon则凭借Trainium表现强劲,并且真正的机器人业务利润效率将在18个月内进入零售领域。
- 至于Microsoft,Satya“在大约3年内,从‘我们要让Google跳舞’变成了Copilot的产品经理”。Microsoft在2025年初“短暂退缩”,因此失去了算力配额;但Baker认为当前的动作很勇敢:它把算力用于自己的产品和模型,而不是服务OpenAI,尽管如果反过来做,“Microsoft今天可能已经是800美元的股票”。他仍然保留一个反事实问题:在那场政变尝试期间,Satya是否希望自己支持的是Ilya和Mira,而不是Sam?最后一个观察是:初创公司与Amazon和Nvidia的合作参与度“遥遥领先”,其次是Google;Broadcom则走自己的ASIC路线。AMD、Microsoft和Meta的参与“基本为零”,他认为这会成为真正的劣势,因为“一些最优秀的团队已经不在大型上市公司里了”。
尾声:事件视界
他为Mythos 3和4做准备,第一步是防守性地过度投资网络安全,并认为“每个人都需要一个安全暗号”——一个无法被社会工程攻击破解的暗号,以防收到一次“完全准确模拟”的FaceTime电话,电话那头是你的孩子要求你汇款100万美元。更黑暗的是,随着美国政治暴力上升(“有人向Sam Altman家中投掷燃烧瓶,这太糟糕了”),他担心AI政治化会让公众人物面临人身风险——这是一个“波动率更高、贝塔更高、风险更高的世界”。
《最后的武士》是他理解自己职业生涯的框架:英雄为武士而战,最终被“一个拿着机枪的农民”屠杀。“机枪已经到来。如果我们不全都成为机枪大师,就会被机枪主人掌控。”他乐观地认为,一位50岁的资深人士仍有很长的优势窗口来驾驭它;他向Patrick承认,目前对他最有用的agent,是每天总结其工作要求中6小时播客内容的工具,外加一个初筛代理,标记PSU与RSU激励机制的变化。
在地缘政治上,乌克兰“真的开始获胜”,而且主要不是靠无人机——“他们拥有仅次于美国和以色列、可能是全球最好的战场AI”——这对美国是好事,却会让正在消化这一事实的对手感到不稳定。他对未来的乐观判断是新的美国治世,并引用1945年:美国当时独占核武器,本可以控制世界,却选择重建德国和日本。他仍然是“AI最大化乐观主义者”——他讲述了一个故事:女儿罕见的基因突变,最终对应到一种由AI代理发现的市场在售药物——“但这是一道事件视界……是社会必须驾驭的不连续性。最好的AI只有有很多钱的人才能获得,这有点反乌托邦。我们需要解决这个问题。”
What was happening in AI was, I think, the most extraordinary moment in the history of capitalism—the history of American business. Anthropic added $11 billion of ARR. The 3 highest-profile SaaS companies founded in the last 10–12 years are Palantir, Snowflake, and Databricks. These 3 companies spent 10 years building their businesses. Anthropic added their combined businesses in 1 month. That’s just nothing like that has ever happened in the history of capitalism. Forget my career—just the flat-out history of capitalism, the history of business.
All right, so this is our 6th time doing this, if you can believe it, which puts you back into first place—at least tied for first place with Bill Gurley. I think even since last time, when we did this, which was so exciting and spectacular, we’re in an even more interesting time now. Maybe just start by riffing on how it felt for you living through March and April of this year, which felt to me like a completely unique economic, technology, and market environment. You’re the biggest student of the history and of these times, so what did it feel like?
I’d say, broadly speaking, there are 2 kinds of drawdowns. There are drawdowns where you’re wrong: a company misestimates, your hypothesis was invalidated, and you have to take your medicine and crystallize that loss. Then there are drawdowns or periods of underperformance where you’re underperforming because of companies you know really, really well, and where you profoundly disagree with the price action and can lean in. Instead of crystallizing negative performance, you can build pent-up alpha—pent-up future performance.
For me, that is what March felt like. It felt like the Nasdaq was selling off, and at the same time, what was happening in AI was, I think, the most extraordinary moment in the history of capitalism—the history of American business. What I mean by that is that Anthropic added $11 billion of ARR. What is astonishing to me about this is that the SaaS and cloud revolution created, let’s call it, between $5 trillion and $10 trillion of value.
Arguably, the 3 highest-profile SaaS companies founded in the last 10–12 years are Palantir, Snowflake, and Databricks. These 3 companies have employed thousands of people—tens of thousands collectively—and they’ve all spent 10 years building their businesses. Anthropic added their combined businesses in 1 month. That’s just nothing like that has ever happened in the history of capitalism. Forget my career—just the flat-out history of capitalism, the history of business.
I mean, it’s wild that Krishna comes on this show and shares some stats: 500% in DR.
Yeah, you do the math on that for 3 years. Insanity.
So, there’s just no precedent for this. Tech investors hear a lot of discussions about S-curves and investing in exponentials. I’ve just never seen an exponential like this. It felt even more extreme than DeepSeek, which was a very similar setup.
If we go back to 2025, there was a huge sell-off on DeepSeek, which was very strange because the paper was published 7 days before DeepSeek Monday. It was published, I believe, on a Monday that was a holiday in America. I read it and thought, “Hmm, this feels like it might not read that positively for the AI trade.”
I took action. We had DeepSeek Monday, where AI really imploded a week later. That was really strange because by DeepSeek Monday, it was super clear that this was going to be the most positive thing that had ever happened to compute demand. Prices in the AWS availability zones in Asia had already doubled. You were seeing GPU availability go down.
This was just the first time we saw how much more compute-hungry reasoning models are during inference than non-reasoning models. That was a similar setup, but you had to do some work to see it. It’s not that hard to say, “Oh, wow, stocks are selling off, the price of DRAM is going vertical, the price of GPUs in Asia is going vertical, and GPU availability is going down.” Then, 2 or 3 days later, GPU prices in America started going up—GPU rental prices.
1. Anthropic and OpenAI Valuations
All you had to do in March was simply observe what was happening to Anthropic. There are all these people who seem to regret not buying during 2022, not buying during COVID, and not buying during DeepSeek. You had the same valuation setup at the beginning of April and an even clearer AI inflection. There have been all these chances to buy into AI.
Then, of course, what complicated it was the straight-up FOMO. I became a believer in AI. I think maybe one thing the market was mispricing was the fact that I’m no macro expert. I do a lot of professional national-security investing, so I have access to people who are experts and are excited to share their thoughts and opinions with me.
I think the market was mispricing the fact that the Strait of Hormuz being closed is actually relatively awesome for America. Why? Because, particularly for the goals of the current administration, electricity is a very important industrial or manufacturing input. The key input into American electricity prices, which feeds into AI, is NG1. Natural gas on Bloomberg was down 20%. Natural gas in Asia, Europe, and everywhere else doubled or tripled.
Our relative manufacturing competitiveness improved overnight. For better or worse, that is what the Trump administration seems to care about. They are very focused on America’s relative position. A lot of people had memories of the 1970s, and what made the ’70s so dramatic was that it wasn’t just that prices went up; there were actual gas shortages.
Then you go through: Okay, well, the U.S. economy is dramatically less energy-intensive than it was. The United States is now the world’s largest producer of oil and gas, and we’ve now become the world’s largest exporter of oil and gas. On top of that, there’s this relative manufacturing advantage.
That made it, I think, easier to stay focused on AI fundamentals and historically attractive valuations. On a relative basis, tech essentially got as cheap as it’s been versus the rest of the market at any point over the last 10 years. Think about that in the context of market efficiency: We have the most extraordinary moment in the history of capitalism, which is wildly bullish for AI, and you get a chance to buy AI at really attractive valuations.
What do you make of the multiples that Anthropic and OpenAI specifically trade at? In my mind, they’re the reference assets—the most pure-play takes on this trend—and the multiples really aren’t that crazy. If you just look at the sales multiple and compare it with what Databricks, Snowflake, and these companies traded at during their peaks, how do you make sense of it?
I do think OpenAI and Anthropic are pretty different animals from a capital-efficiency perspective. Anthropic clearly has a dramatically lower cost per token than OpenAI. They just do. You can see that in the amount of money they’ve burned to get to a roughly similar revenue scale. I think they’ve burned maybe 80% less than OpenAI.
As businesses, they clearly have very different structural ROICs. I think OpenAI is doing a lot. I think Sarah Friar is one of the most exceptional CFOs, and I think they’re doing a lot of things to try to improve this. They’ve secured a lot of compute—more than others.
They secured a lot of compute. That’s another big difference.
It turns out being aggressive really paid. But, yeah, Anthropic at $900 billion for $50 billion in ARR—
But growing 1,000%.
Yeah, growing at ridiculous rates. Maybe a true statement is that if Anthropic had all the compute, they’d probably be doing well north of $100 billion today. Maybe $150 billion. They have clearly degraded the intelligence of Claude. There’s an analysis that Claude, even on Opus, is generating 70% fewer tokens for the exact same question.
As we talked about last time, token quantity equals quality of answer and quality of thinking at some level. There is an intelligence density per token that also matters. I think I felt that as a user. So, I think they would be doing materially more—$100 billion, $150 billion, maybe $200 billion. You might be buying it at more like 5 times unconstrained—I’m going to make up a new number—URR, unconstrained run-rate revenue.
Why do you think they don’t raise $100 billion at a $3 trillion valuation or something like this? If you were the Anthropic CFO—Krishna’s awesome; we just had him on—or if you were Sarah, certainly, if the inbound I received following the Krishna episode is any indication, everyone I’ve ever met is trying to invest in both of these companies.
So, I think it’s wise. The future is uncertain. You are clearly in a very capital-intensive game. Even if you are Anthropic, I’m sure it is at very positive gross margins on inference today. Anthropic probably starts generating cash this year, if they’re not already generating cash, if they’re not already generating cash, which I think is probably the case.
Still, you probably want to be able to raise more capital and access more compute. The world is uncertain. Ukraine is starting to really, really win. How is Russia going to respond? I think there’s still a lot of uncertainty in Iran. All this uncertainty probably amplifies geopolitical uncertainty over Taiwan.
It’s an uncertain world. If I think about Elon, Elon has always made investors money. He treats it like a sacred covenant.
And as a result, because he’s made people money for 20 years now, he has a superpower: He can essentially raise as much capital as he wants, whenever he wants. I think it’s wise that these companies are taking this approach—I don’t know if that’s how they think about it—but I do think being focused on making investors money is wise and creates benefits that don’t just last for a year or 2. They can last for the next 20 to 30 years.
The way Elon did this was by systematically underpricing SpaceX or whatever else. What is the actual method? Just never being greedy on valuation. Never pushing valuation. Just that simple. My friend Antonio pointed out that SpaceX compounded it at the low 30% per year for whatever that was—a decade.
That was just because Elon was focused on preserving the superpower and trying to strike a fair balance between investors and employees. I think it’s wise. But could Anthropic raise money at probably at least a 100% premium to this rumored latest mark? Of course.
2. Orbital Compute and Data Centers in Space
Let’s get to the watts and wafers part of the discussion. Always my favorite thing to talk about with you: the importance of this infrastructure build-out. I feel like every time it’s getting overheated, and then the next time I talk to you, it seems like we should have done way more than we did. You’ve studied S-curves and the steepness of those S-curves a lot, and you know a lot about history. Talk us through how you’re thinking about watts and wafers today as the key inputs into this whole thing.
I think capitalism is going to solve the watt shortage, absent big regulatory or political blowback, which I think is a real possibility. The head of data center infrastructure investing at one of the big PE firms—think Blackstone, Apollo, or KKR—said, “It used to be energy and chips were our biggest gating factors. Now it’s zoning and approval. Much more important.” I think a lot of companies are waiting until after the midterms to take action in terms of maybe workforce reductions. Nobody wants to be a piñata during the midterms.
You’ve seen a lot of companies that make turbines announce plans to significantly increase capacity. There are 2 of these machines that can cast these big blades. We haven’t made 1 in 80 years in the West. We don’t know how to make them anymore. All of that is true, and by no means am I minimizing the industrial engineering magic and artistry that goes into those, but capitalism is very good at solving problems like these over time.
There are other sources of energy besides these turbines on a longer time frame. I think the watt shortage will probably begin to alleviate in 2027 or 2028. Then I think orbital compute will really solve that. I do want to reframe orbital compute because I think when people hear “data centers in space,” which we discussed in our last episode, they picture a Pentagon-sized building in space. They’re like, “Well, we can’t do that.” That’s not what it is.
A Blackwell rack weighs 3,000 lb. It’s 8 ft high, 4 ft deep, and 3 ft wide. It’s racks in space. SpaceX has shown you an illustration. It’s a rack. That’s the satellite. It’s probably about the size of a Blackwell rack. It has these solar wings that are probably 500 ft long on each side. You keep it in a sun-synchronous orbit so those solar panels are always facing the sun. Because it’s in an exact sun-synchronous orbit, the radiator extends behind it for hundreds of feet.
This is a common criticism, yeah. How are you going to cool it?
Yeah. I’ve spent a lot of time at Starbase over the years, and I’ve talked to a lot of SpaceX engineers. I do think it is the most talented group of engineers on planet Earth. They’re very confident they’ve solved this, and they’re not always confident. I think there’s probably some engineering that needs to happen to turn the Starship into a Mars colonial transporter. Will they do that? Absolutely.
What are they more focused on? I’d say probably the repair and maintenance. Those are the 2 big responses: the radiator, and how you repair whatever issue goes wrong in the rack. And the answer is, until you have, probably, a floating Optimus, you don’t.
Now, I do think Starship is going to change the space economy in ways we cannot imagine, particularly if regulation becomes a constraint on data centers. None of it’s going to matter. You’re going to sell as much orbital compute as you can make. And obviously, you link these racks using lasers traveling through a vacuum, which are already on every Starlink.
It’s mind-blowing to me that SpaceX operates the world’s largest satellite fleet, which is like 98% or 99% of all satellites in orbit. Every Starlink—they’re cooling it today. I think Starlink V3 is going to operate at 20 kilowatts. A Blackwell rack is only 100 kilowatts. People talk a lot about density.
Well, if you’re connecting the racks with lasers through a vacuum, you can make the rack bigger physically. You’re focused on weight, not size. In a data center on Earth, where you’re trying to connect racks, ideally using copper and minimizing lengths, cabling is a big cost. You do want that rack to be small because copper when you can, optics when you must. But in space, there are all sorts of things that SpaceX can do that I think maybe some of these naysayers are not contemplating.
They operate more satellites than anyone else. They have a 20-kilowatt satellite today, so maybe you just scale that up to 60 kilowatts to start. They seem very confident they’re going to go right to 100 to 120. And the same company now also operates the largest data center on Earth. They have the world’s best hardware engineers and all sorts of people, almost all of whom are not smart enough or practical enough to work at SpaceX. Are these armchair skeptics?
[Laughs.]
I don’t want to quote Larry Ellison, but somebody was being skeptical, and Larry was just like, “Listen, he’s out there landing rockets. I don’t see anybody else landing rockets.” The reality is that 10 years later, no other company is consistently capable of landing and fully reusing an orbital rocket. None of this works or makes sense without reusability. That means you have to land it.
I would like to redefine orbital compute as racks in space, not giant floating Pentagon-sized data centers in space, which is silly. But what makes a data center is that you’re connecting these racks with lasers. So it’ll be racks in space that are connected with lasers into a virtual data center.
If you think about that state of the world—let’s say that all happens and we’re really good at getting these things up economically, running matrix multiplication all over space—what does that mean for terrestrial data centers? Someone once said, “America was going to suck as hard as it can on every energy source it can get.” I just think the same is true of compute. It’s why I’m probably less worried about an edge AI bear case than I was.
We’re going to consume as much compute as we can. Inference, I think, is very sensible for orbital compute. Training will be done on Earth for a long time. I don’t think that this is super bearish for terrestrial data centers. I think those are going to be valuable for my lifetime.
But I do think if you are in this ecosystem of power production and cooling, and you are massively ramping capacity—and a lot of these capacity ramps are going to be hitting just as I think all of the silly skeptics start to understand that orbital compute is very real—it’s worth thinking long and hard about that if you’re one of those companies.
Then all sorts of cool stuff is happening in the interim. We’re getting really good at repurposing jet engines. There’s that Boom Supersonic that is doing this.
So, there’s a lot of capitalism hard at work on watts. On wafers, though, it’s just this group of plenty of older humans in Taiwan who are the most important humans in Taiwan. They are the overwhelming fraction of the country’s GDP, water usage, and electricity usage. They talk about the silicon shield. They all view themselves as inheritors of Morris Chang’s sacred legacy.
I vividly remember visiting Science Park more than 20 years ago and talking to them: “Do you think you could catch Intel?” And they said, “This is such a beautiful dream, but it’s a dream for our grandchildren.” And they did it. Partly because of Intel’s self-inflicted wounds, but they just think very differently.
You know, one reason Jensen flies over there so much is he wants them to expand capacity. I do think it’s wild that Jensen has never had a contract with Taiwan Semiconductor. They do business on what seems fair, in handshakes. Just fascinating. No contract. It’s going to be fair over time. We’re partners. We’re going to be fair to each other.
And the truth is, based on every prior market precedent for a foundational new technology like AI, you’ve always had a bubble. Carlota Pérez wrote this great book about this. Basically, markets are efficient. They correctly understand that this is a foundational new technology. There’s what he calls a breakdown in diversity. Everyone becomes bullish on this new technology.
I am beginning to worry a little bit about a diversity breakdown. And then you get a bubble. That bubble funds the build-out of this new technology, but supply gets ahead of demand. You get a crash, and it’s a particularly severe crash if it’s a debt-fueled build-out, like the year 2000.
One thing I’m really happy about—really good about—the current build-out is that it’s still overwhelmingly funded out of operating cash flows, which is a really important fundamental difference versus the year 2000. So is valuation, as is the fact that every GPU is running at 100% utilization, when 99% of fiber was unutilized. So there are all these fundamental differences.
But we do have to remember: history doesn’t repeat, but it rhymes. As investors, we have to be very cognizant of it and recognize that, based on the last 200 years—forget the internet bubble—we had a railroad bubble, a canal bubble. We should expect a bubble.
And that’s terrifying. Nobody wants a bubble. A bubble is terrible. The reason it’s terrible is that if you’re valuation-sensitive, you massively underperform. You get fired by probably all your clients.
George Vanderheiden, who is no longer with us, was a great Fidelity portfolio manager. He fought the bubble in ’99, and he retired in early 2000 because I think he just couldn’t take it. He knew it was wrong, and his clients were deeply skeptical: “George, you’re out of step. You don’t get it.”
He had white hair. He was a truly great man. I only overlapped with him briefly, but he was a very important mentor and friend to my good friend and mentor Jennifer Uris. So I have a lot of Vanderheiden DNA through her. He was the same person who said, “Being early is the same thing as being wrong.”
George retired because he couldn’t take the underperformance, and he couldn’t take clients saying, “What’s wrong with you? You don’t get it.” He had 40% of his fund in tobacco and 40% in homebuilders. Literally, he probably outperformed the Nasdaq by 20 or 30 times over the next 3 years.
I have been optimistic that this fundamental shortage of wafers, which today is really controlled by Taiwan Semiconductor, will prevent a bubble. If Taiwan Semiconductor did what Jensen wanted, I think NVIDIA could sell $2 trillion of GPUs in 2026 or 2027. Maybe $2.5 trillion. Maybe $3 trillion.
But there is a limit where consumers would consume so much that you probably would be in an overbuild. And so Taiwan Semiconductor—if we don’t get a bubble, we need to throw a party for them, because they will have single-handedly prevented a bubble.
You are starting to see companies go to Intel and Samsung. Let’s just assume TSMC stays super supply-constrained versus the latent demand. What happens? One of Intel and Samsung—I don’t know which one—is not going to stay disciplined. They will break. And then, at some level, that will force everyone else to break.
A lot of this may come down to the degree to which Taiwan Semiconductor can maintain a lead over Intel and Samsung. And you’ve got to remember, it’s whatever it is—9, 12, 15 months.
Sort of like the leading-edge node, you mean?
Exactly. The pace at which they expand capacity. If I were to watch one thing to understand where there’s a bubble, it’s Taiwan Semiconductor’s capacity decisions.
I think there’s a Goldilocks zone where they expand enough to make it hard for Intel or Samsung to really emerge as a second source at scale, with something well north of 30% market share, and yet they also keep this fundamental constraint on wafers that helps us avoid a bubble.
And then, obviously, I think the Terafab is going to play into this, too.
Say more about that, for people who are not familiar.
The Terafab is a SpaceX—I believe Tesla is involved as well—joint venture to build the world’s largest fab here in America. And I think they’re going to be successful.
First, they have a partnership with Intel, which is very important, because they’re getting access to 50 years of institutional knowledge. That’s just a 9-month, a few quarters, 12-month, three-to-five-quarter lag behind the frontier. That’s an advantage.
3. Terafab and the Future of US Manufacturing
It’s also an advantage that I believe the Terafab is going to get attention from the A-teams at all the semicap equipment companies. One big reason Taiwan Semiconductor caught up is that ASML, KLA-Tencor, Lam Research, and Applied Materials wanted them to catch up. They don’t like having a monopsony.
The A-teams were in Taiwan working with TSMC. Intel made some mistakes, and presto. The A-teams will be here because of Elon’s reputation in hardware engineering.
And then, to a degree that I think may be hard for people to imagine in America, where politics has replaced religion because Elon had his foray into politics, that makes it hard for some people in America to see him clearly, which is sad, because I do think he’s probably doing more for America than any other American.
He’s single-handedly bringing manufacturing back to America. He’s revived defense tech. SpaceX is, in some ways, the most important defense contractor in America. What he’s doing with Starlink is amazing for the world. He’s creating all these blue-collar manufacturing jobs, which is a goal, I think, of a lot of liberals, and it’s good for America.
He’s done more than any living human to decarbonize the world. And if you are upset about data centers on Earth for environmental reasons, well, here you go.
He’s a living deity in China, Taiwan, South Korea, and Japan. Having watched him for a long time, what he’s going to do is recruit the best people, because the best engineers want to work for Elon, especially in hardware engineering.
He’s going to recruit incredible engineers, and then, next to the Terafab, there’ll be a Taiwan town: “Oh, these are your favorite restaurants? I’m going to move them and their whole staff from Taiwan to Texas, and we’re going to make everything the way they like it.”
Then we’ll have Japan Town. Same thing. We’re going to have Korea Town. We’re going to have all these things exactly—but dialed to recruit the best engineers.
That’s just not the way that the people who run Intel and Samsung think. So he’s going to have the best talent. He’s going to have the A-teams at the wafer-fab equipment companies. He has Intel, which is important. It’s so good for all of any administration’s political goals. And I think it’s different enough that it will not alienate Taiwan Semiconductor.
And these have long lead times, right? So Terafab is going to be pumping out NVIDIA GPUs—whatever chips—quite a long time from now.
Elon tends to do things differently. Everybody else is taking 3 years to build a data center. He built one in 122 days.
You know, Samsung had to give him an office in their fab in Texas because he was so unhappy about the pace at which they were expanding and building. We’ll see.
Are you surprised by—you mentioned DeepSeek earlier—the simple reaction to that was, “Okay, these models are just going to get 95% as effective for some tiny fraction of the cost of these Chinese open-source models. We’ll be able to use these for most of what we want to do.”
Fast-forward a little bit of time—2 years from now, there’s no reason I have to spend $1 million a year in my small firm on tokens or something—but the actual reality seems quite different than this. I’m curious why there’s that dissonance in your mind.
I do think the returns to the frontier are fascinating. All the economic returns to AI at the model layer—not all of them, but an overwhelming amount of them—have been at the frontier, which is surprising to me. I think it’s been surprising to a lot of people.
4. Returns to the Frontier
I think this is one of the most important questions to be answered, and you need to have a hypothesis on it as an investor: Are frontier tokens going to continue capturing the overwhelming majority of economic value created at the model layer?
It is surprising. I just remember when Gemini 3.1 Pro came out. It was mind-blowing to me. It was so good. And today, it’s intolerable. Intolerable.
There’s probably a little bit of a dynamic where companies prototype with frontier models, then, when they put something into production, you’re hearing a lot of people do use for tasks or open source.
But still, it is a fact today that the overwhelming majority of these economic returns come from frontier tokens. That’s surprising, and whether or not it continues, I think, is a very interesting question. I’m much more open-minded to that, having had the experience I’ve had with Gemini 3.1 and then Opus. I do use Grok 4.3, and it is on the Pareto frontier.
This is, by the way, a big change and a consequence of what we talked about last time: Google losing its cost-per-token leadership as a result of making very conservative design decisions with TPU v8, while Broadcom and Nvidia continued to make aggressive choices. Google dominated the Pareto frontier 9 months ago. At every point on the Pareto frontier, OpenAI, xAI, and Anthropic were inside it. The Pareto frontier is intelligence first, cost second, and I think this is the most important thing to look at when analyzing AI labs.
Now, the Pareto frontier is dominated by Anthropic and OpenAI, and then Grok 4.3 is on the Pareto frontier. It’s clearly the best, lowest-cost 500-billion-parameter model. Gemini 3.1 is hanging onto the Pareto frontier. If I were to bet, I’d bet that they’re subsidizing that out of pride.
The biggest risk to this trade—to all of AI—is a violation of Richard Sutton’s Bitter Lesson. The closer someone is to AI, the more skeptical they are that this will occur. One thing I think contributed to weakness in March was a much more stupid version of DeepSeek, which is a thing called TurboQuant. “Turbo Quad” is some Google memory optimization that was written up in a paper a year ago.
Then, in the middle of an agreement, while Google was negotiating with Micron, Samsung, and SK Hynix to sign some LTA that would lock in really high prices for a long time, they released this. What people do is always more important than what they say, and they just publicized it on X. It went viral: “Oh my God, DRAM is cooked. Here’s this DRAM optimization.”
I was unable to find a single AI engineer on planet Earth who believed that “Turbo Quad” would have any impact on DRAM demand. But nonetheless, it was a violation of Richard Sutton’s Bitter Lesson: more compute will always outperform human algorithmic ingenuity. More compute and data beyond Chinchilla-optimal—that’s what people increasingly do today, I guess. That’s a real risk, man, and I think the people who are building these models are skeptical of that risk.
The reason I am a little less skeptical is that I think we are very close to ASI. Who knows if the Bitter Lesson holds for 400-IQ models? Or maybe we get a temporary period where, if you get to ASI, the first thing it wants is probably to be smarter and have more resources. How does it do that? It makes itself more efficient. I think that is an actual risk—that the Bitter Lesson, literally, includes humans in it.
We’re about to find out whether the Bitter Lesson applies to 300-IQ AIs, then 400, then 500, and 600. At some point, we may have a temporary violation of the Bitter Lesson based on AI and ASI. I do think this is the third big question: Are Bitter Lesson violations as a result of ASI less likely? Human ingenuity: Will frontier tokens still command the premium they do? And will you get continual learning, and if so, when?
So, I’m curious how you think about some other parts of the innovation around the model, continual learning and memory being 2 that people seem to be most focused on as things that might create yet another new paradigm that we would enter. What do you think about the role of those 2 things?
Yeah, well, I think we’ve done a lot with memory through these harnesses. It turns out that harness engineering is not as important as the model, but it really matters. These harnesses and these models are increasingly being co-developed. One of the big things a harness does—we used to think of it as a runtime that the model operates in—is that it knows where the tools are, creates context, memory, and state, has very specific prompts or instructions, and just makes a huge difference. Even simple versions make an incredible difference.
5. Continual Learning
I think the last time I was on here, or one of the other times, I just said, “Hey, as an investor, it’s very important that you pay for the $250-a-month version to get your own intuitive sense.” That’s no longer possible. To understand what frontier AI is capable of today, even for a non-coding use case, you need to have Claude Code or Codex, and you need to be on an enterprise plan.
The reason for this—and this is another dynamic that’s enabled by Google losing its cost leadership—is that these AI models just shifted to usage-based pricing. If you’re on that $250, $300, or $280-a-month plan, or whatever it is, you’re severely rate-limited. You’re getting a lobotomized version of the AI because, like we talked about, Claude now produces 70% fewer tokens. If you want the tokens that Claude and its harness really think it needs to produce to get you a good answer, you need to be on a usage-based plan.
By the way, this is so bullish for AI. I was a telecom analyst from 2005 to 2007. Cellular had been a great growth industry, really, for the last 10 years, and the reason was that you had a combination of fixed pricing and usage-based pricing over that. You had 900 minutes, or whatever it was, and then usage-based pricing over that. When did cellular stop being a great growth industry? When everybody just went to all-you-can-eat.
Long distance is the same thing. AI is just shifting from all-you-can-eat to pay-by-the-drink. It turns out people really like to talk to their friends long distance. They really like to talk to their friends on the phone. People really like to use AI, particularly now that 1 person can have 100 agents working.
I think this shift to usage-based pricing is probably why you will see OpenAI and Anthropic exceed well over $200 in ARR this year. Not only is more compute going to come online, but they’re going to be able to push frontier-token pricing with these usage-based enterprise models. It’s sad. It’s sad for the world because it just means if you can’t afford that, you’re not at the frontier.
Yeah, continual learning, man. If we solve that, how do you conceptualize it? There are so many mysteries about the human mind. We’re such sample-efficient learners relative to AI.
Orders of magnitude.
Yeah, many orders of magnitude. We have a crude variant of continual learning today when something is verifiable, and that’s just reinforcement learning during mid-training. Continual learning is when a model dynamically adjusts its weights, or adjusts in some way, in real time. As a human, that’s what you do. The first time I touch or put my hand in a fire, I learn never to put it in there again.
That model today needs to put its hand in the fire a million times and then have the designers effectively put a fire in the next training run or an RL gym for it to learn. I think it has to be dynamically updating the weights, but I think people are working on really smart techniques beyond this. If we get that, then we have a really fast takeoff. People seem confident that continual learning is just around the corner.
What is the role of new chip companies in all of this? We talked a lot about Nvidia and its relationship with TSMC and Intel and all these sorts of things. There are a thousand flowers blooming—literally, probably a thousand flowers blooming—trying to create a new chip to address some part of this bottleneck. I’m curious how you process this space and this opportunity, what role it will play, and what role they’ll play.
I think this is good and healthy for the world. It’s good for Jensen, too, because a different administration might take a different view. Competition, I think, is good for everyone.
In tank design, they talk about the iron triangle. The iron triangle of tank design is that all designers of a tank have to make trade-offs between attack, defense, and mobility, for obvious reasons. The more defense you have, which is your armor, the heavier the tank is and the less mobile it is.
So, you have to live in this triangle and make trade-offs. Okay? Like the Merkava in Israel, it's optimized for defense. Russian tanks and the Leopard are generally more optimized for mobility. Chip design is the same. There are these fundamental constraints imposed by the laws of physics, as embedded in the Taiwan Semi design rules, that you need to live within.
You have TPU, Trainium, and AMD, which are all essentially trying to be a better GPU. Today, I think probably Trainium is doing the best. Now, nobody's a better GPU. But Trainium is, I think, tugging on Superman's cape.
6. Watts, Wafers, and Infrastructure
Trainium 3 needs to ramp into production because it has a switched scale-up network, which you really need to economically run inference on MoE models. A lot of companies have a torus architecture. That's where Google was. And AMD, we'll see. The MI450, we don't know yet. We'll see. We probably know more about Trainium 3 than the MI450.
But that's a hard game to play. So, you have to do something different. And you have to do something different that is also hard to do. So, I think the best path for these startups—my rule of thumb is 1% market share is going to be worth $100 billion. $100 billion is a pretty good venture outcome.
I think what Jensen would say is, “Okay, if somebody does something different and it gets to 1%, 2%, or 3% share, we'll make that chip.” And that's coming for everyone. But if you're trying to make a better GPU, good luck. If you were doing something different, it also needs to be hard to do.
You can make different trade-offs. The disaggregation of prefill and decode has really opened the aperture for making these different trade-offs, because you can make very aggressive trade-offs for decode and aggressive trade-offs for prefill. Prefill is taking in the context; decode is writing the output.
I have a great colleague named Andrew Fox who said, “Picture an 18th-century British naval ship. Prefill is loading the cannon; decode is firing it.” What prefill literally is is just the model understanding the question, the prompt, and then keeping track of its own answer. That is fundamentally a memory-capacity-bound problem.
Decode is the process of generating new tokens, and that is memory-bandwidth-constrained. So, if you're a chip designer, this gives you a richer canvas to paint on. But even so, it needs to be hard, because if you make different trade-offs in that iron triangle to optimize for memory capacity, and they're not hard trade-offs to make, then Nvidia is going to make those same trade-offs.
They get better prices from Taiwan Semi than you're ever going to get. Good luck. Good luck. And they have the advantage of working with every model company and optimizing their designs.
By the way, another very funny thing is, if you're a VC and you're investing in a semiconductor company that is telling you they are going to have an advantage because of a Taiwan Semi process that they have special access to, I promise you that Jensen saw that process when it was a twinkle in Taiwan Semi's eye, and he knows more about it than this little company with 200 people can imagine.
Taiwan Semi—everybody in the supply chain is showing Jensen everything, the same way they're showing Amazon everything, AMD everything, and Google everything. And that's another reason: don't go try to make a better GPU. You can do something different. You can paint in the prefill canvas. You can paint in the decode canvas. But you also have to do something hard, because if it gets to scale, you're going to have those 4 companies as very fast followers.
7. New Chip Companies
My firm was a venture investor in Cerebras. What Cerebras has done is something hard and fundamentally different: wafer-scale computing. And it comes with a set of trade-offs. But that architectural decision they made was hard and lets them do something that no one else can do. And we'll find out how big that is.
They're working on really cool things. One of the problems Cerebras has, once you start needing to glue a lot of chips together and scale-up networks or scale-out networks, is that you need a lot of I/O. And I/O is bound by what's called the shoreline, the sides of the chip. So Cerebras has an overwhelming ratio of on-chip compute memory relative to shoreline I/O.
Well, they're really smart people. They did something really hard. They're trying to see if they can put an optical wafer right on top of that. And then that solves that problem. I'm sure they're looking at hybrid bonding of DRAM to get around these alleged limitations that are not true.
A Cerebras machine can theoretically run any size model. So there are types of models where they're much better than other architectures. So, Cerebras, what I think is interesting is they did something different that's hard to do. Really hard to do: wafer-scale computing.
So, I do think there's a role for these. And I would just encourage them all: make a different trade-off and try to do something hard. Because everybody's going to get funded after the Cerebras IPO. It's not going to be a problem.
But it took Cerebras 3 generations of chips to get it right, and it's really hard. Andrew Feldman, the CEO—you can just see how hard it was, what he did and what that whole team did, to get where they are today. And they need to have the grit to do that, the resilience. This first chip is a failure. It happens. Can you come back and make a second chip?
8. Extending GPU Lifespans and Private Credit
But the one last thing on this topic: this is going to be amazing for the useful lives of GPUs and may single-handedly save private credit.
Tell me about that. What do you mean by private credit?
Well, private credit is in pain from these SaaS loans. And however much they're marked down, they probably need to be marked down more. Because if the public companies are struggling to adapt, how's a debt-laden company going to adapt and invest in what is a very different margin-structure business?
There's a lot of private credit in GPUs, too. They were underwriting that to, I think, 3 or 4 years. But the disaggregation of inference means that I think these GPUs are going to have 10- or 15-year lives.
The AI skeptics are like, “Oh, these companies are all cooking their books. The useful life of a GPU is only a year or 2. The useful life of a CPU is only 4 years because of the rapid technological change.” No. What rapid technological change has done with the disaggregation of prefill and inference is mean that you can put a Cerebras system or Groq LPUs that NVIDIA acquired effectively in front of a Hopper or even an Ampere, use that Hopper and Ampere for prefill, and extend the useful life of that GPU until it melts.
Now, they do melt, so they have a time limit. But maybe you don't have to run them as fast. This is going to be really good for the whole private credit industry. It's going to help finance the AI buildout, because if you can start to finance GPUs at more like 5% or 6%, instead of—I think CoreWeave's lowest financing was like low 7s—that actually mathematically changes the cost of financing this buildout.
We had this technological innovation that's going to lower the cost of financing and extend the useful life of compute on Earth. And then I do think the one last thing that's interesting about that is my friend Jamin from Coatue just did a podcast, and Coatue had a deck. They talked about, “Hey, the sellers of scarcity are doing so much better than the buyers of scarcity.” Buyers of scarcity being the hyperscalers.
But if you own a giant installed base of what is currently in shortage, that's also a very, very good place to be. And we're hearing CPUs are way more important than they were in an agentic world. They do all these things around orchestration, tool calls, et cetera, et cetera, et cetera. The biggest CPU fleets in the world sit at the hyperscalers. So I think some of these hyperscalers may catch up a little bit to the sellers of scarcity.
9. The Application Layer
I want to talk about this idea of “different and hard” applied outside the infrastructure piece of this. So now you're starting to interact with new founders, existing CEOs, and founders that have to adjust to this new world. What are you seeing from the most AI-native founders that aren't building chips, infrastructure, or models, but are just using this technology to build other stuff? How do they feel the most different to you, if you've observed differences?
Well, one, I do think this is just as true for chip design. To me, it's always been a fundamental question for venture. There are different ideas that are obvious to everyone on planet Earth as soon as they hear them. And if that's where you are in venture—if it's not hard to do, if it becomes obvious to the world before you have built scale, and scale is the ultimate advantage—you're in trouble.
The great thing Amazon had was that it was obvious to a lot of people, but it wasn't obvious to the retail CEOs. And Amazon was very smart. Any e-commerce company that VCs invested in, they would destroy. They'd be like, “Oh, that's so cute. We're going to take our margins in that to negative 10,000%.” And the guys at Wayfair did something hard. Amazon tried to kill them, and they failed. Those were tough, operationally really competent CEOs.
For me in venture, I always look: Is this going to be obvious to the world before this company could build scale? Or is this both not obvious, different, and really hard to do?
I think a lot of founders are really struggling with this in AI. People are becoming worried. Today, in Jensen's 5-layer cake of AI, the profits are accruing to energy, they're accruing to data centers, they're accruing to chips, they're accruing to models, and not really accruing to the applications.
Cursor and Cognition got to a scale. They focused on coding. Eighteen months ago, people were focusing on coding. OpenAI was doing everything under the sun, while the people focused on coding were Cursor, Cognition, and Anthropic. It was a really righteous focus on code.
John Massaad, the founder of Replit, tweeted something that I thought was so smart. It was something like, “Bitter lesson adjacent is the fact that coding might be the shortest path to ASI and useful AI.” Because if you really go to coding, you can write yourself code to do anything. I think it was really smart of those companies to focus intensely on coding, and I think they all probably got to a scale where they have a place. I think Cognition is doing something really, really different.
A lot of founders are really struggling, man. They’re really struggling. I think they’re trying to get confidence that, in niche areas, they can get to scale and get a data moat before the model companies get to that niche. Or that it’s a small enough niche that the model companies won’t do it themselves, but it can still produce a different outcome.
Is this related to what you would call the token path? I know you’ve used that phrase with me before.
Yeah, it comes from a guy at Altimeter, Jamin Ball. He just said, “If you’re a software company or an AI company of any kind, you have to be in the token path.” So Databricks, that’s in the token path. Compute companies are in the token path. If you’re not in the token path and you’re not in some really niche thing, life may be hard.
Even for these vertical niches, I think if you talk to the people at the model companies, they’re skeptical of some of these because all of the data being generated in these niches comes from humans. But then you’re betting that you’re able to use that proprietary data in this narrow vertical to train a model that’s lower cost than the frontier labs can ever get to. Maybe that’s a good bet, but I just think you have to be very, very careful.
On the other hand, if the returns to these frontier tokens relative to other tokens come down, there’s going to be an explosion in value creation at the application layer. I think another really important point is that I have a belief that, whenever he wants, Jensen can probably get pretty close to the frontier.
With his own model.
With his own model. They’re doing some really cool things in Nemotron.
To monetize your compliment, as Sklansky would say.
I don’t think he wants to do that. That is what OpenAI and Anthropic are kind of trying to do to him, unsuccessfully. He’s a very logical thinker. This is the logical counter move.
I think you will see an open-source frontier, which today consists of Chinese models with stolen American tokens. Somebody told me that DeepSeek—the latest one, or maybe the original one—was only 150,000 reasoning traces. There are many ways to launder this if you’re a Chinese company. You can hit all these different APIs. You can make it hard.
Now, the American labs are working really hard on anti-distillation technology, but I just think Chinese open source is doing really impressive things in a very resource-constrained way, and there’s a lot of distillation. This is why I think, in addition to there not being enough compute to serve o3, they just did not want it to be distilled. They wanted to use o3, distill it themselves, and use it to RL their next model, whatever it is.
Eventually, I think, if OpenAI gets to economics that I feel good about, anyone on the frontier will do the same thing: just say, “There’s going to be some very interesting game theory.” It is a new kind of prisoner’s dilemma. We talked about the old prisoner’s dilemma being around, “Hey, you’re in a prisoner’s dilemma where you have to spend.”
10. The Token Path and Open-Source Dynamics
The new prisoner’s dilemma is going to be: If you’re at the frontier, do you release that model via API or not? If everyone at the frontier agrees not to do that, then Chinese open source catches up quickly. If one person defects, they’re going to have the best model, a lot of revenue and cash flow, and then, of course, resources equal intelligence, so they’ll start to pull ahead. That will lead to everybody else releasing it.
It’s a new game theory. It’s kind of the same game theory that you have with Taiwan Semi, Samsung, and Intel. The reality is, if a company like NVIDIA or AMD were ever really, really to use one of these other foundries, that foundry would get better really quickly. So I do think Jensen is going to keep open source a certain time frame behind the frontier. I think that’s going to be a very interesting thing to watch.
By the way, open source gets monetized. There’s this misnomer that open source is free. Open-source tokens cost energy to produce. You need to make it up on GPUs, and the open-source model companies almost always get a revenue share.
How are you preparing Atreides for the world of Mythos 3 and Mythos 4?
We’re just trying to overinvest in cybersecurity. Something I’ve said in multiple forums, and I really believe, is that everybody needs to have a safe word. Everybody needs to leave their digital devices behind—literally go to the ocean—and have a family safe word or a company safe word. It can’t be one that can be socially engineered.
This is just to avoid cybercrime, where what looks like your son, your daughter, your grandparents, your parents, or whatever FaceTimes you is an utterly accurate simulation of them. They know everything and can extrapolate, based on what they’ve said, what they’re likely to say, and say, “Wire me a million bucks.”
That’s defensive. What will you still be able to do that it won’t be able to do, I guess, on the analytical side?
That’s a good question. I just watched The Last Samurai, and I asked my firm to watch it. The Last Samurai, if you haven’t seen it, I highly recommend watching it. It’s actually a movie that’s aged really well.
It’s a Tom Cruise movie from 20 years ago. The conceit is that Tom Cruise is this bitter, washed-up Civil War veteran who’s actually a very good soldier. He’s bitter and washed-up because he feels like he participated in negative actions against the Native Americans. He’s hired by the modern elements of the Japanese government during the Meiji Restoration to train an army of peasants how to fight the samurai.
There’s a first battle. Of course, the samurai win even though they don’t have guns. He fights valiantly, so the samurai decide not to kill him and take him to their village. He becomes a samurai. It feels like the Civil War to him, so he fights on the side of the samurai. At the end, he’s massacred by a peasant with a machine gun.
The machine gun is here. If we do not all become masters of the machine gun, we’re going to get mastered. So I’m trying to become a master of the machine gun.
I’m optimistic there’s a long period of time where, just like if you were a 50-year-old samurai veteran of many wars who had fought many wars and mastered warfare, you would have advantages using the machine gun. I’m optimistic that, as a lifelong student of investing, I’m going to be able to master the machine gun—this new technology—and integrate it into my own process and our firm’s process in ways that let me contribute value as a human being for a long time.
Like everyone, I have agents running all the time now.
What’s your most useful agent?
Honestly, my single most useful agent is a really good summary of the points that would be interesting to me from podcasts. I think I told you this, and I don’t want to hurt your business, but there are 6 hours a day of stuff that I feel like is in my job description to watch.
11. Cybersecurity
Every time somebody from OpenAI, xAI, Google, Cursor, Fireworks, Scale AI—or, let’s say, nothing of Jensen, Elon, or Dario—I feel compelled to watch. I just don’t have that much time, and there are some real needles in haystacks.
There’s a set of things I always like to see. I’m very sensitive to management compensation. What are they incentivized to do? Do they have stupid RSUs, or do they have PSUs? If they have PSUs, what are those PSUs incentivized to do?
I think systems that do a very good first pass at that save people a lot of time. They free people up for more creative work than going through the proxy, pulling the PSU thing, and looking at how it’s changed versus all the proxies, because there’s signal in that. That’s very labor-intensive and so good for AI.
There are obviously all sorts of the same things within investing. This is the most exciting, thrilling time to be an investor, and I’m getting a little bit worried.
The diversity breakdown thing? Say just a little bit more about the kinds of people that are bullish on DRAM.
Yeah.
I don’t know of anyone like me who’s not really bullish on DRAM. No one. No one. There are all these interesting things happening with AI right now. Cross-sectionally, the valuations do not make sense. They just flat-out do not make sense.
You have semicap equipment companies trading at 40 times next quarter’s annualized earnings and DRAM companies trading at mid-single-digit multiples. At the peak of the last cycle, that was like 5 versus 12. At one point, it was like 3 versus 45. Those can’t both be true.
And yes, semiconductor capex business models have improved more than memory business models. We don't know how much HBM is going to improve memory business models yet. They have some element of recurring revenue with parts and maintenance, but it's not worth a 1,000% multiple gap. I think it's hard to square the valuation of something like NVIDIA, which, in early April, was essentially as cheap as it gets relative to the market over the last 10 or 12 years or whatever it is, and very cheap in absolute terms.
It's very hard to square that valuation with something like GE Vernova's valuation, because it builds in an unfathomable amount of share loss for NVIDIA. So, valuations are really different cross-sectionally. Because we are in shortages, the lowest-quality companies are doing the best.
12. Avoiding the AI Bubble
If you're an oil and gas investor or a mining investor, a natural-resources investor, and you're well-versed in thinking about costs, this is very intuitive to you. In a real bull market for a commodity, the commodity suppliers with the highest costs go up the most because it's the most beneficial to them. They go from being on the verge of bankruptcy to just gushing cash. This is, I think, one reason commodity investing is really, really hard, because quality outperforms during the cycles, but you get all of the outperformance during the downturns, when the high-cost guys that mooned during the shortages and the commodity bull markets go bankrupt or whatever.
You're seeing that happen in every industry. The lowest-quality players in these different industries that are hated and detested by the hyperscalers and the buyers because they have high costs, they're unreliable, the parts fail at a high rate, et cetera, et cetera, are sold out and raising prices. Then that activity gets the interest of these retail accounts on X, and these stocks get bid to the moon. Whereas some of the higher-quality expressions have actually really underperformed.
As an investor, it's hard because you know, without a shadow of a doubt, that the thing that's moved 10x in 3 months or 6 months is going to go right back down, subject to what they do with all the cash. But these low-quality companies really do smart stuff with cash. It worries me a little bit that people who were very skeptical a year ago are no longer skeptical.
13. Diversity Breakdown
But then I contrast that with the valuations of these high-quality companies, which are just not extended, and it makes me feel better. It does kind of feel like—I just thought it was funny in 2024 and 2025 that anyone asked about an AI bubble or talked about it. You have this nuclear bubble and this quantum bubble right here, right in front of you. What are we talking about? This is so real.
Some of that nuclear and quantum silliness has maybe spread into more speculative, lower-quality, smaller-cap names, where if you have a big presence on X or Reddit, it's easy to move them. That frightens me a little bit. But I just wish there were more AI bears. I wish there were more memory bears.
One reason I'm—Astera is a stock I've been close to for a long time—there are a lot of bears on that. I love that. Great, I first invested in the Series C. Good luck thinking you're going to price that differentially from me. Good luck thinking that's a copper loser.
Then there's also—you can feel the baskets in the market, in the leveraged baskets. What baskets you're in is really important: copper, optical, DRAM, NAND. A very interesting thing that's happened this year is that, in 2024 and 2025, the AI trade traded together.
So you could be long GPU compute, scale-up networking, and optical scale-out, and short power. That trade worked from a risk-management sense because I'm very factor-aware. That all blew out in January of this year. Scale-up networking would go crazy while scale-out was going down, or DRAM would massively underperform NAND and HDDs, which had not happened.
These cross-sectional correlations within AI really fell apart, and you had to get very fine-grained. You couldn't hedge your memory anymore with semicap equipment or NAND. Everything cross-sectionally really changed in January, and in a very interesting way. I think maybe one reason for that was that AI got to a scale where it was suddenly really easy for a bunch of people to get really smart on these different subsectors, start trading them, and then they get put into baskets.
Yeah, creating price efficiency.
Yeah, exactly. I think some of the biggest opportunities outside of these higher-quality names, which I think can compound for a long time and are safe, unlike these low-quality names, which are terrifying, are in names that are miscategorized.
Astera was in a lot of copper-loser baskets. Its biggest product is going to be a switch. You use both copper and optics to connect switches to accelerators.
I wonder if you could riff for just a sentence or two on each of the major companies. I feel like I always forget to ask you about Google, Microsoft, Amazon—the major players that all the conversation is centered around, versus these exciting new companies.
Google was incredible last year because they had that TPU advantage, which is now gone. The reason I think they're still in a great position is that they have the most compute of everyone. We talked about the value of installed bases being higher as a result of shortages. They have the biggest installed base of compute.
Google I/O is this week, and if they don't release something that even slightly leapfrogs OpenAI and/or Claude, that's interesting. It's not a disaster for Google; it's just interesting, and it means this Nvidia effect we discussed is even more powerful than maybe I'd imagined. I'm very curious to see what the Pareto frontier looks like literally in 5 days after Google announces its new stuff. This is a big card for them.
But Google, between the amount of data they have—the YouTube data is actually really genuinely valuable, and it is valuable in a world of robotics—the amount of compute they have, and the search business they have, is never not going to be in a good position. Then you see that with GCP going crazy.
You have to give Zuckerberg immense credit for what he's done in terms of making Meta an AI-first company internally, and I do think he is the only one of those true internet giants to have done that. I give him a lot of credit for that. I give him a lot of credit for paying up when he did for all those billion-dollar contracts to talent.
That model, I think, was a really big upside surprise. It was the first model from Meta Superintelligence Labs, or MSL. It's not on the Pareto frontier with xAI, Google's one entrant, and then OpenAI and Claude, but it's pretty close. That was very impressive to me.
So I think Meta is in a better position. It's still not as strong an absolute position as Google, but rates of change matter more than level, as you know, in markets, particularly over short, 3-year time frames. Over long time frames, the level of competitive advantages tends to dominate, but even within that, changes really matter.
Amazon, I think, is in a really strong position because of Trainium. You're going to see real P&L efficiencies from robotics over the next 18 months in their retail business. I actually think Nova, their internal models, are not where what sounds like “Muse” is, but they're better than they get credit for.
14. Assessing the Big Tech Players in AI
Microsoft—I think Satya is a really brilliant man, but in investor conversations, people just don't talk about him the way that they did. I like Satya. I admire him. I think he's an exceptional CEO, and I give him a lot of credit for the decisions he's made.
But he did go from, “We're going to make Google dance,” to being the product manager of Copilot in 3 years. I would love to know, during the coup attempt against OpenAI, whether Satya regrets his decisions. Does Satya wish that he had supported Ilya Sutskever instead of Sam Altman, and that Ilya and Mira Murati were really running OpenAI today?
In his heart of hearts, I would love to know, because I think the Microsoft-OpenAI partnership might look very different in that world. I think that's a very interesting question that we'll never know the answer to. But I give him a lot of credit. What he's doing now is taking risk.
This goes to the decisions you have to make in that cone of uncertainty: not only how much you spend, but what you're going to spend it on. I think Microsoft flinched for a moment in early 2025. They have this algorithm: “We spend this much in capex dollars, we get this return.” That algorithm was kind of off.
If you flinch, you lose position. You lose all these allocations, and it's difficult to get them back. So they flinched. Now the decision Satya is making—which the market has punished him for, but I think is the right decision—is, “We're going to use our compute.”
I mean, who knows how fast Azure could be growing if they're willing to just sell GPUs to OpenAI. We're going to use our compute internally to make our own products better. One reason Copilot is so bad, or has been so bad, is that there just wasn't enough compute available. They're fixing that. He's the product manager of Copilot.
I do think he's a great CEO. They're trying to use their compute to train their own models. I'm a little skeptical that they have the right team to succeed there, but they can certainly, just like Meta, afford to hire maybe a different team. I think he's making good decisions—risky decisions—to position Microsoft for this world where frontier models are no longer API-accessible. I think it's a really courageous decision, and I give him a lot of credit for it.
He's forgoing—I mean, Microsoft would probably be an $800 stock today if they were using their GPUs solely to serve OpenAI and Anthropic's capacity instead of using them for their own products. So I give him a lot of credit for making a great decision.
What's really interesting is the degree to which these companies are outward-facing in their decisions. The 2 companies who are the most deeply engaged with startups are Amazon and Nvidia by a mile. Then there's really intense engagement with Google; Google is next most intense. Broadcom is engaged in a different way. They're everybody's favorite ASIC supplier.
If you're a startup, it's considered a level up if you get to work with Broadcom for your 2nd-gen chip, and it's considered mana from heaven if Broadcom works with you for their 1st-gen chip. Then you see essentially zero engagement with startups from AMD, Microsoft, and Meta. I just wonder about that decision. Some of the best teams are no longer at big public companies; they're at these smaller startups. I think it's going to end up being a pretty big advantage for Nvidia, with AMD and Google right behind them, to have this engagement that you just don't see from these other hyperscalers.
As we wrap up, I'm curious for you to riff on any other out-there knock-on effects that you've started to think about for this giant trend. We've talked about the specific companies that this most impacts in a lot of detail. We've talked a little bit about the application layer and what would have to happen for there to be more value accruing to that layer of the stack. I'm curious about any other fun knock-on things that you've been thinking about as this world changes so quickly.
Yes, and it is wild. At the application layer, forget value accruing; just value has been destroyed. AI has net destroyed value. Even if you count Cursor and Cognition, the most successful AI natives, trillions of dollars of value have been destroyed by AI at the application layer.
In this context, I do think it's something we need to be aware of. The companies that are doing the best today, whose values increase the most and that are creating economic value, are the companies with the highest effective ratio of utilized GPUs per human. Maybe this just means that every human's going to get a lot of GPUs. But I think that's an interesting fact that we need to be cognizant of.
I will just say—and maybe this is a little dark—I am more and more worried about personal safety. I worry about this a lot more for people who have a much bigger public presence and are much more associated with AI. But I really worry about personal safety. I hope nothing tragic happens, but there is this upsurge in political violence here in America. As AI increasingly becomes political, I worry that's going to get directed at more and more AI political leaders.
Whatever I may think or may not think of OpenAI, I think it is terrible that someone threw Molotov cocktails at Sam Altman's house. I am worried that we are headed into a higher-variance, higher-beta, higher-risk world because of AI. That's for me as an individual, and then, for people who are big players on the chessboard, think about what it means geopolitically.
We're watching the Ukrainians really start to win. I think the reason they're winning is not really because they have better drones. I think they do have better drones; that's part of it. I think the reason Ukraine is really winning is that they have the best battlefield AI outside of probably America and Israel. As China and our adversaries begin to process that, how do they respond?
If you're America, it's great to have that edge in AI, but it is destabilizing for the rest of the world. Something I think a lot about is creating a charity to educate the world on how awesome the West has been. Slavery was endemic to almost every civilization, and slavery was really ended by the British Empire. Tell that story.
15. Geopolitics, Personal Safety, and the AI Horizon
But after 1945, we had the nuclear bomb; no one else had it. We could have controlled the world forever. Instead, we rebuilt Germany and Japan. Now, who are America's most reliable allies? Israel, South Korea, and Japan. That's a testament to the American spirit in our country. We didn't take over the world.
There were these fears, documented at the time, that the American generals—and MacArthur was a little bit of an American emperor in Japan—were just going to take over the world. They could have, and they didn't. They came home, we demilitarized, and then you had this period of great global stability. There were terrible wars, but you had the Pax Americana. So maybe it's not destabilizing. Maybe it leads to another Pax Americana informed by our AI dominance.
I'm so optimistic that AI is going to be amazing for the world. There's someone like me whose daughter was diagnosed with a very rare mutation. There's no cure. He was able to assemble a lot of resources and get a lot of compute from the labs. We were made aware of what was happening, spun up an immense number of agents, came up with, using AI, a drug on the market that can actually impact his daughter's disease, and then spun up a company to cure it. Her life is already immeasurably different because of AI.
So I'm an AI optimist, a maximalist, but I also acknowledge that it's an event horizon. I think it's for sure going to be a discontinuity that we need to navigate as a society. I think the Luddites are going to be wrong, but we need to be really thoughtful in how we address their concerns. We need to make sure that it's good for everyone. It is a little dystopian that now the best AI is only available to people with a lot of money. We need to solve that.
We need to approach this with humility, recognize there's a lot of uncertainty, and be thoughtful.
When I do this with you, I tell people afterward, I'm like, “May you find something that you love as much as Gavin loves markets and companies and capitalism and history.” That's on display today, as always. Gavin, thanks so much for your time.
Thank you. Thanks, Patrick.