万亿美元级 AI 建设内幕|Dylan Patel 访谈
- OpenAI–Nvidia 交易的核心是资产负债表问题,不是什么“无限资金漏洞”。 1GW AI 算力容量按5年合约出租,每年租金100–150亿美元,意味着每GW要先付出500–750亿美元;而 Sam Altman 要的是“超过10GW”。Nvidia 承诺投入1000亿美元股权(首期:1GW、100亿美元),实质上让 OpenAI 部分用股票支付:建设1GW约需500亿美元,其中约350亿美元流向 Nvidia,毛利率75%;这部分毛利约一半再循环成 OpenAI 股权。Patel 的说法是:“Nvidia 在降价,但没有降低报价”,Patrick 则称其为“史上风险最高的资本主义游戏之一”。
- 规模化不是边际收益递减,而是童工经济学。 算力每扩大10倍,就能买来一个能力层级;但层级之间的差距,是6岁小孩和13岁小孩的差距:“你能让13岁小孩完成的工作,其实非常有价值。”达到 Google 高级工程师水平的智能,意味着“2万亿美元的软件价值”——也就是全球软件工程师的工资总额。未对冲的尾部风险是:“如果模型不再进步,我们就彻底完了……美国经济会陷入衰退”,台湾和韩国大概率也会被拖下水。
- GPT-5 是一次服务能力决策,不是规模化失败。 Token 需求每2个月翻倍,但硬件不会同步翻倍,因此每个智能层级的成本必须坍缩:GPT-3 质量的 tokens 如今便宜约2000倍,DeepSeek 比 GPT-4 便宜约500–600倍,GPT-OSS 还要更便宜。GPT-4.5 证明自己“聪明得多”,却无法以合理成本服务后,OpenAI 让 GPT-5 “基本和4o一样大”,以服务所有用户、推动采用率,并把智能增量转移到思考层级。若能按下魔法按钮,他会优先解决容量/成本,而不是延迟,因为瓶颈在那里。
- 后训练还处于“我们刚投出第一球”的阶段。 约40家湾区初创公司正在搭建 RL 环境,包括虚拟 Amazon、数据清洗训练和让模型“疯狂爬坡”的数学题;后训练“最终会吞噬大部分算力”。基于文本的预训练仍“相当早期”,多模态规模化仍在推进,而机器人数据飞轮“甚至还没走出休息区”。
- “Nvidia 不承担风险,中间环节人人都有大量风险。” Neocloud 的经济模型是:Blackwell 按6年全包成本约2美元/小时,短期租金为3.50–4美元,直到下一代芯片做到“速度快10倍、成本高3倍”。真正的金鹅是与有真实资产负债表的客户签长期合约:Nebius 与 Microsoft 约190亿美元的交易,至少带来60亿美元毛利。Oracle 面向 OpenAI 签下3000亿美元合同,而 OpenAI 的 ARR 只有约150–160亿美元;如果成功,利润约1000亿美元;如果失败,就是用债务押注一个没有资产负债表的客户。
- 美中竞争的逻辑是:“不加速就会死。” 如果没有 AI 这轮繁荣,美国“到本世纪末很可能落后于中国,不再是世界霸主”。据节目所述,中国过去10年向半导体投入至少4000–5000亿美元,优先建设封闭供应链,而不是相信聚合理论;中国“几年内就能建出10GW 数据中心”——“Elon 的速度比中国慢”。而台湾末日风险会击穿传统组合逻辑:如果你相信这个风险,就“也不能投资 Apple”——“那就只能 YOLO 投 TSMC”。
- 投资判断已从看空转为“超级看多 Google”:凭借垂直 TPU 栈,Google 拥有最低 token 成本,并在每条战线上苏醒;Meta “手里有牌,可能把一切都拿走”——它是唯一同时拥有眼镜硬件、模型、服务容量、推荐系统能力,以及下一代人机交互所需资本的公司;相较 OpenAI,他更看好 Anthropic,因为其收入“加速得快得多”,且对应2万亿美元的软件市场;没有“Pornbot”之外的商业模式,XAI “正面临无法融资的真实风险”。
- SaaS 正迎来清算:AI 会压低软件构建成本,把 build-vs-buy 的天平推向 build;但每个 AI 产品都会增加“巨额 COGS”,而获客成本不变,因此“纯软件公司的时代在 AI 时代真的非常艰难”。 中国就是实验证据:开发者成本约低10倍,SaaS 和云服务从未在那里达到同等规模。已经完成规模化的平台反而获胜:“谁控制平台,谁就会一次又一次地赢。”
1. OpenAI–Nvidia 交易: “这笔账的资产负债表到底是谁来扛?”
- Patel 先拿那套梗开涮:OpenAI 付钱给 Oracle,Oracle 付钱给 Nvidia,Nvidia 再付钱给 OpenAI,“无限资金漏洞”——随后他直接否定:“实际情况不是这样。”真正的驱动力是“算力建设早于业务增长”:必须先建好集群,才能出租做推理,也才能训练出解锁下一批应用场景的模型。即便已有8亿用户,“OpenAI 规模太小、无法产生实质影响的风险仍然很高”——它面对的是 Zuckerberg、Google、Elon,是“世界上最富有的人”,也是他所谓科技巨头的“帕斯卡赌注”。
- 头条背后的算术是:数据中心容量每GW每年租金100–150亿美元,而 OpenAI 签的是5年合约——每GW要先流出500–750亿美元现金;Sam 想要“超过10GW”。Microsoft 已经失去兴趣,而“Oracle 的资产负债表根本不像 Google、Microsoft 和 Amazon 那么强”,所以 OpenAI 需要愿意提前投入资本开支、相信租金收入最终会到账的盟友。
- Oracle 的一侧,是一笔对着约150亿美元 ARR 的3000亿美元合同(“现在可能是160亿美元”,年底约200亿美元)。“如果赌赢,他们就赚到1000亿美元利润……纯现金利润。”如果赌输,Oracle 就只能自己扛下建设成本;而且它已经开始举债。
- Patel 简化后的 Nvidia 机制是:10GW 对应 OpenAI 1000亿美元分期股权投资(首期:1GW、100亿美元)。建设1GW约需500亿美元,其中约350亿美元直接流向 Nvidia,毛利率75%——可以粗略理解为400亿美元收入、100亿美元 COGS、300亿美元毛利,其中约一半循环投入 OpenAI 股权。Nvidia 先确认资本开支对应的收入,同时持有一家“可能值钱、也可能一文不值……可能有能力、也可能没能力支付数千亿美元算力合同”的公司的股票。Patrick 称这是“史上风险最高的资本主义游戏”。
2. 规模化不是边际收益递减——下行风险是经济衰退
- Patrick 直接追问:你确定那条双对数曲线还会继续吗?Patel 的答案是:“一切迹象都表明它会继续。”算力每增加10倍确实只买来一个能力层级,但层级之间的价值差距,是6岁小孩和13岁小孩的差距:每6个月换一批高中生组成的团队,最多只能“挖沟和做园艺活”;而25–30岁的人能搭建价值高得多的企业。“价值发生了剧烈变化”,并不是边际收益递减。
- 我们所处的位置取决于领域:在软件领域,“我们已经相当不错了”,这也是 Anthropic 能从“不超过10亿美元的收入增长到现在的70–80亿美元……这是我们见过最快的收入爬坡”,且“基本全都与代码有关”的原因。无限个 Google 高级工程师级别的智能,意味着“2万亿美元的软件价值,因为这就是世界今天支付给软件工程师的钱”——它是生产力倍增器,不是简单替代。
- 他的判断比“极度看多”更复杂:Sam 认为 AGI 会在“不到1000天”内出现,Dario 更加乐观,他的室友(一位 Anthropic 研究员和播客主 Dwarkesh)也是如此;但一些知名投资人觉得 Patel 的说法“听起来像疯话”。谈到上限,他属于“你能找到的最乐观的人”之一:最终可能出现数字上帝——“是10年?100年?1000年?我不知道。”但即便能力从现在起停滞6个月,也会是“在创造多少效率和价值方面的一份天赐之物”。他最喜欢的空头可能是 Yann LeCun:认为“在互联网文本上做自回归预训练无法带你走到 AGI”是对的,但把 RL 也一并否定则“完全错误”;此外,还有一些投资人“认为这全是胡扯,却仍然靠它赚了大钱”,比如在财报前基于市场预期买入 Oracle。
- Patrick 提出 Carlota Perez 的过剩产能论:每次短缺之后都会过度建设。Patel 不回避:“如果模型不再进步,是的,我们会过度建设……美国经济会直接陷入衰退”,台湾和韩国大概率也一样。但历史类比并不统一:郁金香和加密货币是“彻头彻尾的庞氏骗局”,英国铁路则连续10年消耗约6%的 GDP,是真实资产——“我们离 GDP 的6%还差得远”。而且最强的资产负债表可以踩刹车:Microsoft 曾经踩停,随后“又把插头插了回去”,还不得不从 Nebius 买容量。
3. Tokconomics:GPT-5 是服务决策,商业是变现终局
- 他创造了“Tokconomics”这个词,也就是 token 的经济学(“终于彻底干掉 crypto,一劳永逸”)。同样1GW算力,服务糟糕模型时能处理1000倍 tokens,服务优秀模型时是1倍,服务惊艳模型时只有0.1倍。需求每2个月翻倍,但“我不可能每2个月把硬件也翻倍”,所以给定智能水平的成本必须坍缩,而且事实确实如此:GPT-3 质量如今便宜约2000倍;DeepSeek 以比 GPT-4 便宜约500–600倍的成本震动市场;GPT-OSS 还要更便宜,而且“实际上比最初的 GPT-4 略好一点,因为它能调用工具”。
- GPT-5 的决策由此而来:OpenAI 曾用 GPT-4.5 尝试更大的跨级跃迁——“它确实聪明得多”,但“没人能以合理的成本或速度服务它”。到年底有效容量只有约2GW,速率限制已经在限制用户(Patel 自己要开多个 ChatGPT 账号来批量发起深度研究任务),因此 GPT-5 “基本和4o一样大”,成本大致相同甚至更低,能够服务更多用户,再把额外智能放进思考/Pro 层级。
- 如果可以按下一个魔法按钮——应该解决延迟还是容量?——他的答案仍然是:“我大概还是会说容量/成本比延迟更重要。”他自己的行为就是证据:虽然可以使用大概率为 Claude 4.1 Opus 的模型,但他使用 Sonnet 多得多——“它客观上更笨,但 [Opus] 太慢了……我的时间也是有价值的。” Anthropic 的收入主要来自 Sonnet 而不是 Opus,也是同一个原因:“没人想用一个慢模型。”
- 变现终局是代替用户购买:Etsy 超过10%的流量直接来自 GPT(Amazon 把它挡在外面,“否则比例会非常高”);OpenAI 的应用负责人曾打造 Shopify 的购物代理。“模型会替你完成购买……哪怕只占0.1%、1%、2%,也会像一笔信用卡交易。Visa 是世界上最惊人的生意,因为这个原因;聊天产品也可能做到这一点。”
4. 后训练: “我们刚投出第一球”
- 现在的问题不是模型更大,而是 grokking:模型会先记忆、再泛化,而参数过剩的模型“从来没有机会泛化”。真正的挑战是把数据放进有用领域:“互联网上没有任何地方会告诉你,如何只用键盘在电子表格里飞快操作……它连基本操作都不会,比如玩转电子表格”,尽管它已经读完了整个互联网。
- 因此才有 RL 环境——“现在湾区有40家初创公司在为 OpenAI、Anthropic 等搭建这些环境”:虚拟 Amazon 要求模型从一堆干扰项中买到正确的除臭剂;反复进行数据清洗训练;由模型评分的医疗病例;以及数学题。去年第四季度到今年第二季度,模型在数学题上“疯狂爬坡”,主要靠学会编写 Python 来完成计算。
- 按 Patrick 的说法,赛程进度是:文本预训练“还处于相当早期”(学习效率仍可提升,而且任何预训练收益都会传导到下游所有环节);多模态规模化仍在早期,V3 和 Banana Nano 还在推进(大概率指 Google 的 Veo 3 和 Nano Banana 视频/图像模型);至于后训练,“我认为我们刚投出了第一球”,而且“最终会吞噬大部分算力”。他用自己刚出生的侄子校准感官来说明这个阶段:孩子把手塞进嘴里——“强化学习还处于极早期,因为人类就是这样,我们是强化学习者。”
- 关于记忆:Transformer 在精确回忆上非常强(如今已经能很好处理大海捞针),但“真正糟糕的是拥有无限上下文”——人类会把世界压缩成稀疏信息(你对童年的记忆是被重新回忆出来的图像,“只是稍微变了形”)。模型不必像人类一样工作:深度研究已经会把笔记写在侧边,“用语言压缩信息”,跨数百万 tokens 运行45分钟——“你从别人那里读到的很多备忘录,至少初级员工写的,已经和深度研究不相上下。”这也是实验室需要数百万张 GPU 的原因:不是只跑一次巨型训练,而是“我得尝试无数种不同做法,因为我不知道哪种会奏效”。
5. 人才战争:机器学习研究就是半导体制造
- 数十亿美元的薪酬包是理性的:研究人员是在指导使用1000亿美元芯片的实验,哪怕浪费三分之一算力,影响也大得惊人。但增加人手会拖慢研究;Meta 在超级智能前的困境是“人太多了,而真正出色的领导层没有把他们带好”。他的朋友、OpenAI 的 Roon 可能曾发帖说:“每次想到 Meta 浪费了多少 H100,我都会产生一种肉眼可见、发自身体的愤怒”——不过“所有人都在浪费算力”。
- 他最近最喜欢的类比是:机器学习研究与半导体制造完全一样——每台设备有1000个工艺旋钮,搜索空间不可能穷举测试,所以只能靠直觉选点、读取模糊数据,“然后直接 YOLO”。研发晶圆厂“除了教会你如何做出下一代工艺节点之外,不产生任何经济价值”——训练中烧掉的算力也是如此。
- Roon 的另一个想法,是从深圳通过开出离谱报价来收购流程知识、招揽整个团队,Patrick 称这是“绝妙的主意”;Patel 强调这是 Roon 的想法,不是他的,并把它与 Intel 的衰落联系起来:最聪明的18岁年轻人不再去读20万美元的纳米化学博士,而是去 Google 拿80万美元、去 OpenAI 拿1000万美元、去 Meta 拿1亿美元——这与最高分医生都去皮肤科和麻醉科是同一类人才偏移。难点在于筛选:“有多少人不善言辞,却真的极其擅长做事?”他还指出 Sam 的自我安慰——“他们没拿到我们最好的人”——而这话是在内部竞价挽留人才时说的。
- 他一直记得 Jensen Huang 的一句话:“美国之所以富有,是因为我们把所有劳动力都出口了,但把所有价值都留了下来。” Nvidia 和 Apple 把制造外包给亚洲,却保留了毛利。
6. 权力关系: “史上最精彩的肥皂剧”
- Anthropic 是否把 Cursor 的牌全部拿在手里?Cursor 的年化收入接近10亿美元,其中大部分回流给 Anthropic(Patel 认为利润率“略为正数”);Anthropic 再把毛利投入算力,所以无论怎样,“毛利美元最终都会流向硬件层”。但 Cursor 掌握数据和用户,也有自己的 embedding 与自动补全模型,并且“我想什么时候切换到 OpenAI 都可以”——甚至可能训练出一个在特定细分场景击败 Anthropic 的模型。“所有人都是亦敌亦友。”
- Microsoft–OpenAI 是“世界上最疯狂的权力关系”:2023年时,“Microsoft 要拥有整个世界”;到2024年下半年,Amy Hood 等人开始后撤——“也许我们没必要为3000亿美元兜底”,于是暂停数据中心建设,并把算力排他性让给 Oracle。交易本身包括约20%的收入分成、49%封顶利润结构、IP 共享,以及 AGI 条款终止 API/IP 权利——“这他妈到底是什么意思?”重新谈判的 MOU 则是“史上最不像公告的公告”。他对 AGI 定义的看法是:如果把今天的 ChatGPT 展示给20年前的人,“这就是 AGI”;标准总会移动——对他而言,真正的门槛是把手塞进嘴里的感知时刻。
- Nvidia 的困境是:它无法通过收购扩张(当年还是行业中几乎无足轻重的角色时,ARM 收购就被拦下),而“如果你只能回购股票,你就是输家——这等于承认你无法让资本获得更高回报”。于是资产负债表变成武器:提供需求保证,为本来根本不会建设的 CoreWeave 集群提供兜底,并实质上预支 OpenAI 第一年的算力——“我有一年的1GW,可以用来想清楚商业模式。”与此同时,“风投给一家公司融资,结果这一轮70%都花在算力上——他们他妈喜欢死了。”
7. Neocloud:Nvidia 不承担风险,中间环节人人都有大量风险
- Neocloud 模式“做得好时绝对惊人,做不好时绝对糟糕”。一台 Blackwell 按6年计算的全包成本约2美元/小时,短期租金“超过3.50美元甚至4美元”——利润高得离谱,直到下一代产品以“3倍成本实现10倍速度”上市,短期租金随之崩塌。真正的金鹅是带资产负债表的长期合约。
- Nebius 刚签下典型交易:与 Microsoft 约190亿美元,“至少能从这笔交易赚60亿美元毛利……我愿意每天都做这种生意”。市场给 Microsoft 的付款义务定价甚至低于美国政府债务——“这在我看来很疯狂,但随便吧。”CoreWeave 也曾靠同样的交易起飞,直到 Microsoft 不再继续,再转而找到 Google 和 OpenAI;但 OpenAI 的合同,无论名义金额多大,最终都建立在一个“没有资产负债表”的客户身上。“那我怎么确定他们真的会付款?”
- Google 因数据中心容量紧缺,开始为加密矿工的交易提供兜底,其中包括 TeraWulf 和 Fluid Stack——Google 把 TPU 系统实体卖给第三方,由第三方部署并出租,“而 Google 仍然赚走所有钱”。推理服务商确实有真实需求:Roblox 想在游戏里接入 LLM,Shopify 需要客服,企业需要微调和服务开源模型;但也有一类人是“YOLO,我把 tokens 卖给一群想做 SaaS 应用的随机客户”,他们可能很快耗尽现金跑道。主线始终是:“Nvidia 不承担风险,中间环节人人都有大量风险。”
8. 建设确实存在:电力焦虑、并网怪相,以及一个由 AI 造出的业务
- 对能源恐慌,他的第一近似判断是:“我们现在有点太娇气了——目前还没用那么多电。”数据中心约占美国电力的4%,其中一半来自 AI:“这根本不算什么,老兄。只是我们差不多40年没建设电力了。”真正的约束是供应链和劳动力:GE、Mitsubishi 都在把涡轮产能翻倍;变压器线圈固化产能需要两年建设;有公司因为涡轮卖光,只能并行接线组装柴油卡车发动机;Elon 则从波兰运来电力设备。流动电工的工资已经翻倍,西德州“像回到了2015年,大家都去做压裂工”。但换个视角,数字仍然惊人:OpenAI 的2GW园区消耗“整个费城的电力”;如今听起来稀松平常的500MW,一旦装上 GPU,就是250亿美元资本开支。
- 真正有意思的是三阶影响:训练负载变化太快,足以扰乱电网——频率出现偏移,“你的冰箱会更早坏掉……而你甚至可能不知道,因为附近就有数据中心”。Texas 和 PJM 正在采用限电规则,提前24–72小时通知数据中心削减一半用电;这又迫使数据中心配备现场发电机,但如果每月运行“超过8小时”,就会触发空气排放许可问题。
- 面对 Patrick 关于“我们是否只是把旧事做得更快”的质疑,Patel 给出的反例是自己的损益表:他的第二大收入产品,是对全球每个数据中心的卫星照片做图像识别,再抓取 LLM 解析的许可证和监管文件,最终卖成双周表格(“这个 Amazon 数据中心的风扇开始转了,所以我们可以预测 Amazon 的收入”)——“没有 AI,这个业务根本不可能存在”,而且只用了3个人,不是过去需要的50–100人。再比如主机迁移:Amazon 离开 Oracle 数据库,“他妈花了20年”;现在可以快得多。
- 他手绘的看多光谱值得保留:Dario → 机器学习研究员 → Patel → Patrick → 纽约半导体投资人 → “那个自2023年以来就认为 AI 是泡沫的 Sequoia 人” → 仍然不肯建设电力的公用事业公司高管。
9. 美国 vs 中国: “不加速就会死”
- 在他看来,赌注是:“没有 AI,我们肯定会输。”供应链会更慢、更贵;债务不可持续;社交媒体上收入差距的可见性,以及算法信息流对旧有单一文化的切割,会放大社会不稳定(“你我其实很相似,但我们的信息流完全不同”)。“AI 必须大幅加快 GDP 增长。一旦开始讨论怎么分蛋糕,你就完了。”没有 AI 繁荣,“美国到本世纪末很可能落后于中国,不再是世界霸主”。
- 中国正在复制自己在钢铁、太阳能、手机和 PCB 上走过的长期路线:据节目所述,过去10年通过国企、税收政策、省级土地拨款和大基金,向半导体投入至少4000–5000亿美元;相比之下,美国的关税和 CHIPS Act 都只是“杯水车薪”。中国优先优化封闭供应链,而不是聚合理论;讽刺的是,美国“某种程度上正在做中国历史上做过的事——向某个方向倾注海量资本”,押注 ChatGPT 成为智能领域的 YouTube(永远亏钱,最终在拥有30–50亿用户时控制平台)。ByteDance 已是全球第三大 GPU 用户;DeepSeek 工程师收入高于同行,但没人支付1000万美元——那里没有人才挖角文化。
- 建设速度方面,“Elon 的速度比中国慢——我认为他自己也知道”。如果中国想建一个10GW数据中心,“我敢打赌几年就能建成”(必要时还可以走私芯片);同期 OpenAI 的全部容量,乐观估计也只能达到10GW。中国缺少最好的芯片和内存,但拥有最多电力,以及最快的建设速度。
- 台湾末日场景会击穿传统风险管理:封锁或 AI 加持的颠覆行动(“中国可以把 Cambridge Analytica 在台湾放大10亿倍”)意味着“美国经济会某种程度上自由落体——没有台湾芯片,我们连冰箱都造不出来”,新车造不出来,新的 AI 数据中心建不起来,云业务也无法增长。因此,投资组合经理禁止投资 TSMC 的规则并不自洽:如果你相信这种风险,“也不能投资 Apple”、Amazon 或 Google——“所以这就变成 YOLO 投 TSMC。”
10. 投资组合:做多 Google 和 Meta,看好 Anthropic 胜过 OpenAI——SaaS 正迎来清算
- 快问快答的结论是:Anthropic 胜过 OpenAI——“它们的收入加速得快得多,因为它们专注的方向与那2万亿美元的软件市场更相关”,而 OpenAI 同时分散在消费、科学和抽成模式上。AMD:“我喜欢它们,但它们挺一般”——这是他第一只多倍股,有感情,不是投资逻辑。Oracle:“如果你相信 OpenAI 会成功,它将赚到非常多的钱”——但在大多数 OpenAI 能支付3000亿美元的世界里,“OpenAI 本身会是一家10万亿美元或5万亿美元公司”。XAI “正面临无法按所需规模融资的真实风险”:Colossus 2 将拥有30–50万颗 Blackwell,成为全球最大的单体数据中心,但它需要“Pornbot 之外”的商业模式;Patel 的解决方案是与 OnlyFans 合作,把创作者用特洛伊木马的方式带进 X。
- 两家老牌巨头获得最强背书。Meta “手里有牌,可能把一切都拿走”:下一代人机交互是把你的需求告诉 AI,而 Meta 独有四项能力——硬件(新型显示眼镜)、模型、服务容量、推荐系统能力——“再加上资本”。Google 则是:“大约两年前我还相当看空 Google,但现在超级看多。”TPU 已对外销售,模型真正具备竞争力,基础设施投入激进,而且有机会同时拿下消费和专业两种界面;Meta 只能拿到消费端。
- 收尾框架据说来自一位同事,可能是 Doug O'Laughlin:SaaS 黄金时代建立在平坦的研发成本、极低 COGS 和可通过规模摊薄的高 CAC 之上。AI 从两方面打破它:竞争性软件的构建成本大跌,把 build-vs-buy 推向 build;AI 功能又增加“巨额 COGS”,而 CAC 不变,于是市场碎片化,公司“永远达不到逃逸速度”。中国是天然实验场:软件开发者的有效成本约低10倍,因此 SaaS 和云服务从未在那里达到同等规模。Google 再次受益——凭借垂直 TPU 栈,拥有“任何公司每个 token 的最低 COGS”;随着内容生成成本下降,已经完成规模化的平台将持续繁荣,YouTube 的“辉煌时代”还在前面,因为“谁控制平台,谁就会一次又一次地赢”。
- 考虑到节目关于“股票诅咒”的传说,最后的收束是:“泡沫正在此刻破裂,因为 AI 的上限是无限的。”不过 Patrick 提醒,他们查过数据:所谓诅咒“其实只是市场表现”——Patel 上次看多 Applied Materials 后的6个月里,该股上涨了约70%。
If the models don't improve, we're absolutely screwed. In fact, the U.S. economy will go into a recession.
It's about the highest-stakes capitalism game of all time.
Godsend in terms of how much efficiency and value can be created, and it doesn't ever have to get to digital-god level. Now, I do believe we're going to get to digital-god level eventually.
Eventually, if I could have an intelligence as smart as a Google senior engineer, that's $2 trillion of software value.
Is that the main bottleneck to be attacked?
We're popping the bubble right now because the limit of AI is infinite.
1. The OpenAI and Nvidia Deal: The Infinite Money Glitch
I was going to lay out this idea of going through the past, present, and future of compute as the big, big idea for our conversation. But since this just happened, and I don't think I've heard you talk about it anywhere, I'd love to start by asking about this whole OpenAI-NVIDIA thing, which sounds exciting, seems vague, and I'm not really sure what's going on. Maybe you could explain it to us as you see it and what the strategic implications are of the big announcement.
All right. So, I think it's very, very simple, right? You've got OpenAI paying Oracle lots of money. You've got Oracle paying NVIDIA lots of money. You've got NVIDIA paying OpenAI lots of money. It's a meme.
We've got the infinite-money glitch here.
No, no, no. That's not actually what's happening, right? What's really happening is OpenAI has an insatiable demand for compute. The compute precedes the buildup of the business. You have to have the cluster before you can rent it out for inference, or rather, run models on it for inference. You have to have the cluster to train the model that's good enough to unlock new use cases, which then can be adopted, and there's an adoption curve there for any new use case.
You have to have all these things sequenced, given this is a game of the richest people in the world, or rather, the biggest tech giants in the world. It's Zuck. It's Google—Larry and Sergey, or Sergey is constantly in the business now again. It's all the biggest people in the world. It's Elon.
There's very much a risk of OpenAI being too small to matter, which is crazy to say because they've got 800 million users. But where's the revenue? Where's the compute? They could easily get swamped in terms of how much compute they have. If they don't move fast enough, and if they don't have the most compute, or at least among the most compute, they will get beaten.
The magic of OpenAI was that they just spent way more compute on a single model run on GPT-3 and GPT-4. They had the foresight, the vision, and the execution. They made that bet, and they were able to secure it. At the time, it was like, meh, right? It was a few hundred million, whatever. That's a ton of money, but now it's sort of like, well, Mark Zuckerberg sees how much compute he's going to have to get, even though he has this insane cash flow.
He's like, “Oh, wait. I need to go sign a deal with Apollo for $30 billion on this data center in Louisiana, this megadata center I'm going to build.” It's like, “Wait, why didn't you just fund this with cash flows? You have so much cash flow.” It's like, “Because my plans—that's just the physical data center. Now, what am I going to put in it?” That is so much money.
2. OpenAI's Compute Challenge and Capital Requirements
The amount of capital that people are going to have and are dumping into this is insane. Google was slow to wake up, and then they were slow to pivot their data center operations. They were slow to do everything, and so while they could have way more compute than anyone by a humongous degree, they haven't been able to deploy as fast.
Google is still on the curve of how much they allocate to Search. Generative search is not really necessarily competing with OpenAI, right? It's the megamodels. If you have this tremendous vision of what's going to happen with AI, you know that it takes a ton of compute to build them. Pretty much, the amount of compute you could dedicate to these models is limitless, and they will get better.
Now, it's a log-log scale, right? You need 10× more compute to get to the next tier of performance. You might think of it as diminishing returns, but what if the next tier of performance is like a 6-year-old versus a 16-year-old? Child labor is quite effective versus a 6-year-old, whom you can't get to do much. This is not exactly the way to think of AI, but this is the conundrum that OpenAI is in.
They have to get more compute than anyone, or at least among the most. They have to race with the giants. These giants are trillion-dollar businesses.
So, how does OpenAI get there?
It's partnering with Microsoft. Well, that soured some. It's partnering with Oracle. Oracle can do a lot, but Oracle doesn't even have a balance sheet like Google and Microsoft and Amazon, et cetera.
It's Meta, Elon.
Sport of kings.
3. Oracle's $300 Billion Bet on OpenAI
Yeah. This is very much like the Pascalian-wager nature of all of this with the tech giants. Oracle can be part of it, but OpenAI needs allies. They need people to effectively spend the capex ahead of the curve and trust that they'll be able to pay the rental income, because that's what it is at the end of the day.
OpenAI is committing to 5-year deals. These 5-year deals cost X amount of money. It's $10 billion to $15 billion per gigawatt of data center capacity that you pay a year. You're paying that for 5 years. That's $50 billion to $75 billion of cash that goes out the door for OpenAI for 1 gigawatt of capacity. And you talk about what Sam's saying: “I need 10 gigawatts. I need more than 10 gigawatts.”
Then you end up with this really challenging aspect of, how do you pay for that? And hey, that's only the rental price. If I were to actually do the capex—or if I were to, because it's front-loaded, right—it becomes: who is the balance sheet for this?
That's the reason these deals are coming about. Oracle is making a massive bet. Larry is getting good margin off of it, but he's making a massive bet that this capex that he's going to pay for OpenAI will actually be paid, because he signed a $300 billion deal with OpenAI.
Where's that going to come from?
Yeah. Your revenue is like $15 billion ARR this month, maybe. On a run-rate basis, it'll get to $20 billion by the end of the year. Pretty clearly, maybe it's $16 billion now, but how do you pay $300 billion of revenue now?
4. Nvidia's Strategic Investment and Deal Mechanics
If the bet works out, they've just made $100 billion of profit—pure cash profit. It's crazy. But if it doesn't work out, they've got this huge liability, and they're starting to raise debt. There was a small deal they signed recently, but they're going to start raising more and more debt.
NVIDIA has the same conundrum. Google and Amazon are doing these deals, whether it's with other vendors for TPUs or for Trainium, whether it's Anthropic or others. They're trying to court OpenAI; they're trying to court other companies. How do I get into this game? Fine, I can rely on Microsoft somewhat. I can rely on Oracle somewhat. But at the end of the day, if I want GPUs to be king, part of it is just that my chip is the best. Part of it is also: who's going to pay the capex upfront?
Google and Amazon will pay the capex upfront if it's for TPUs or Trainium. They won't necessarily pay the capex upfront for that same capacity of GPUs. So you've got this challenging aspect, and that's where this NVIDIA and OpenAI deal comes from.
5. Understanding the Demand Dynamics
I want to dig into the underlying assumptions driving this on the training and inference side. Obviously, there's the willingness—Zuckerberg just needs to go down the hall to a CFO to get access to all this capital.
He doesn't even need to go down the hall. He can just make it so.
He's got the voting share.
Sam's got to fly to Norway and Saudi Arabia and other places, and we're at that tier of capital.
I think you're making it sound way easier than it is.
I don't mean to at all. I'm just saying, you know, Zuckerberg is—
Hold on. If it's this easy, let's raise $100 billion, dude.
We should do it. We can compete. But I want to make sure I understand your thinking on the underlying two sides of this. One is your view on the diminishing-return curve—just the return on this. I want to ask about inference too, and the growth in token inference demand, but—
I don't think it's a diminishing return, right? I think that's important to recognize.
Start there. I want to ask about inference too, and the growth in token inference demand, but—
6. Scaling Laws and Diminishing Returns Debate
Given it's a log-log chart, scaling laws are right. Given there's no model-architecture improvement, you just throw more compute, data, and model size at it, and it gets better at this pace.
But you're confident that that will continue?
I think everything has shown that it will continue, and it has continued. GPT-5 isn't necessarily that much bigger than GPT-4, right? And GPT-4 is smaller than GPT-3. What's changing is the paradigm of how you spend the compute. Also, if they made a bigger model, could they even serve it? No.
Right? They did GPT-4.5, and it was terrible. No one could serve it, right? It was actually quite a bit smarter, but they couldn't serve it at any reasonable cost and speed. This is why Anthropic has the same issue—or I wouldn't even call it an issue—but all of their revenue comes from Claude Sonnet 4; it doesn't come from Claude Opus 4.1, which is the better model. It's bigger, but it's slow because the hardware hasn't caught up in terms of inference speed for that, and no one wants to use a slow model, right? The user experience sucks.
Yeah.
But as far as whether the model gets better at each scale of hardware spend, I would say all the tech giants believe it. I believe it. I think a lot of people in the financial community are like, “This is freaking scary.”
Yeah.
Because the moment it stops, wherever you were on the rung—if we went from $50 billion of spend to $500 billion of spend, well, that $500 billion of spend is never going to have ROI, right? It was one thing if $50 billion didn't have ROI, but now this $500 billion doesn't have ROI. It's a big problem.
7. Why Bigger Models Aren't Always Better
One could think of it as diminishing returns, because when you go from $50 billion of spend to $500 billion of spend, you only move up, let's call it, 1 tier of model capabilities in the absence of major algorithmic improvements. I'm holding those off to the side for now, but that iterative performance improvement in the model is, like I mentioned earlier, a 6-year-old versus a 13-year-old, maybe. The amount of work you can get a 13-year-old to do is—I mean, if you do it right, we frown upon that now in this civilization—but the amount of work you can get a 13-year-old to do is actually quite valuable relative to a 6-year-old.
The same applies to a college intern versus someone who graduated and has even 1 year of work experience, because there's a learning curve for kids coming out of college all the time. So there's that learning curve, and I think while it may be incrementally the same—an order of magnitude more compute—the amount of value is much greater.
If we made a company full of high schoolers and had to refresh them every 6 months so they didn't learn too much and become really good, it would be really hard to create a valuable company. The most you could do is dig trenches and do yard work, but these kids wouldn't even show up. How valuable of a business could you build if you had unlimited high schoolers versus a business that refreshed them so they didn't build knowledge, versus college students, versus 25- to 30-year-olds? The value of the business you can build, even though incrementally it's just 5 years between each of them—
Yeah, it's drastic.
It's a drastic value change.
Where do you think we are today? Which level are we at, do you think?
8. The Economics of Tokens and Serving Capacity
It depends on the domain. For software developers, I think we're really pretty good. That's where we're seeing the most value creation happen, right? You see Anthropic go from $1 billion or less in revenue to $7 billion to $8 billion already. It's the fastest revenue ramp we've ever seen for anything of this scale.
And it's basically all code-related.
Right. I mean, some of it's their own Claude Code product, some of it's Cursor, some of it's GitHub Copilot, which has also offered Anthropic models since the beginning of the year. It's Windsurf. It's all these different avenues to access the same thing, and these companies aren't all doing the same thing. There are tweaks and nuances to how they're doing things differently, but it's all code.
In that sense, if I had a 30-year-old senior engineer at Google, and if I had an infinite number of those, all it cost was capex for chips and the operational cost was actually quite low, then you could build businesses worth insane amounts. You could have a replacement for the $2 trillion in wages that go to all the software developers in the world today. Or rather, you could augment them and build twice as much, 5 times as much, or 10 times as much if you could augment them, because these things don't just run on their own, right? They're more of a force multiplier to the existing person.
The value-creation potential is there. It's obvious if you've coded at all in your life. It even works for VBA. It's not that great for VBA, so I know a lot of people in this audience probably know VBA, but it's not even that terribly bad for making macros. The value-creation potential there is incredibly high.
So let's capture it. How do you capture it? This draws back to the OpenAI-NVIDIA deal, because I think most people in the market don't quite get it, right? They're like, “Oh, this is just round-tripping.” It is to some extent, right?
If OpenAI builds a gigawatt of capacity—they agreed to 10 gigawatts of capacity—NVIDIA will make a $100 billion equity investment into OpenAI in the form of cash, and NVIDIA gets returned capital. The first chunk of the deal in the press release is 1 gigawatt, $10 billion, right? So, pretty straight-line. But 10 gigawatts—1 gigawatt to build, as we established earlier, is like $50 billion.
So NVIDIA is paying $10 billion. OpenAI still has to come up with the other $40 billion somehow.
Right now, what they can do is go to the markets, get a loan, or get someone else to put up a loan. There are infrastructure funds that are trying to get into this. All these commercial real estate people are trying to get into this. There's some way where they'll be able to figure out other people to front the capital and then come up with a deal much like it is with Oracle. But OpenAI has to do more of the work in terms of setting up the cluster, the software, the networking, et cetera.
The nice thing for NVIDIA is that, of that $50 billion, they capture maybe $35 billion in capex that goes directly to NVIDIA. In year 0, OpenAI and its partner spend $50 billion on the data center. The timing isn't exactly that, but they spend $50 billion on the data center, and $35 billion goes to NVIDIA. NVIDIA's gross margin is 75%. Again, I'm going to use simple numbers: let's say it's $10 billion in COGS and $40 billion in revenue, for $30 billion of gross profit.
If we fix the numbers, effectively, half of their gross profit from that deal is going directly to OpenAI in the form of an equity investment. The 25% that's COGS is staying with NVIDIA—NVIDIA is paying for that—and they keep the other half of the gross profit on their balance sheet or do buybacks, whatever they want to do with it. NVIDIA is round-tripping some of this, but OpenAI effectively gets the opportunity to pay for a big chunk of it in equity.
Yeah. Right. So all they're really doing is saying, “Half of my money that's in this—sure, it does make its way to me somehow—but in reality, I still made half of that gross profit, and the other half is equity in a company that may or may not be worth something.”
A company that may or may not be able to pay hundreds of billions of dollars of compute deals that they've signed, right? In which case, they'd be bankrupt, right? So this is the mechanics of that deal.
It's about the highest-stakes capitalism game of all time. It's so interesting to think about when it might run out. You mentioned that if we hit that final point and we don't see the return, we're kind of toast and in a big hole.
But I'm also curious about the other side: the ability to serve and the demand for today's models by inference. The statistic I last saw is token demand doubling every 2 months or something crazy. Obviously, there are all these reasoning tokens that are really exciting for some of the longer-thinking models.
9. Rate Limits and the Adoption Curve Problem
How do you think about the growth of the pool of demand for inference tokens themselves, even in today's models? Even if we just stopped things and fixed things—and we'll leave that other side of the equation just for a second—what's your model for thinking about that today? What most interests you in the growth of just broad demand?
The thing I like to call it is “tokconomics.” I stumbled upon the word, actually. It's like a crypto—kill off crypto, finally, once and for all. I'm trying to make “tokconomics” SEO-direct to us talking about tokconomics, and then hopefully you talking about tokconomics—hopefully everyone using it.
Say “tokconomics” 20 more times.
It's the economics of the tokens, right? How much compute is being spent? How much is the gross profit? What's the value being created by these tokens? That's what's relevant here, right?
NVIDIA keeps saying “AI factory,” which produces intelligence. That intelligence has value. Let's say you have a gigawatt of capacity. What can I serve? Well, I could serve 1,000 times the amount with a model that's really shitty. I could serve 1 time the amount with a model that's good.
And I could serve 0.1 times a model that's amazing. Now, multiply that by whatever factor—how many users, what's the number of tokens output? I could do X number of tokens, X times 100, X times a million tokens, depending on the model quality.
This is sort of where the whole GPT-5 thing comes around. OpenAI had a challenging thing. They're like, “Hey, we effectively have a couple of gigawatts of capacity by the end of this year, roughly a couple of gigawatts of capacity as well—more or less, a little bit less right now. How do we maximize our serving capacity with this?”
One avenue is that we continue to serve big models and make bigger models, and the tokens are more expensive. But this log scale is really challenging because, yes, the value is an order of magnitude more, but the cost is way more. The real whammy is that the user experience is way worse.
Right. If I serve a massive, massive model, it's slow, and users are fickle. You need the response to be way faster. It's hard to calibrate.
Yeah. So there's this user-experience challenge, but really, in the end, for a given model level, I think there's a saturation point for how much demand for intelligence there is. You can only have such a large army of people digging trenches or whatever it is. This is very cancellable, but you could have a much larger army, or a much larger business, at a larger level of intelligence.
When you think about what I could have done with GPT-3, even if we paused there and paused the model capabilities, obviously the cost to serve the quality of GPT-3 has tanked.
99% or more.
Yeah, it's like 2,000 times cheaper now. It's so much cheaper now. GPT-4 is the same thing. People were freaking out about DeepSeek because it was 500 or 600 times cheaper. GPT-OSS came out, and that's even cheaper than that. For roughly the same quality, I would actually argue that the GPT-OSS open-source model is a little bit better than the original GPT-4 because it can do tool calling.
The cost of these things tanks rapidly with algorithmic improvement, not necessarily with the model getting bigger. As these algorithms get better, at X level of intelligence, you can only serve so much demand. The flip side is that it takes time for people to realize how to use it.
When GPT-3 launched, no one cared. When GPT-3.5 launched, most people still didn't care. ChatGPT launched with GPT-3.5, and people cared a little bit. GPT-4 launched on ChatGPT, and then people cared a lot. But a model tier like GPT-3.5 or GPT-3 can still be very useful in a lot of the world.
10. The Tokenomics of AI
It's not useful for a lot of use cases. For coding, it was terrible; for copywriting, it's okay. There are some levels of use cases, and the same happens with GPT-4, but it takes time for that adoption to happen. You've got this challenge: if I pause on a model capability, I end up taking way too long for adoption. Also, how can I get people to adopt it if I don't let people use it?
OpenAI had this tremendous problem with GPT-4. GPT-4 Turbo was smaller than GPT-4, and GPT-4o was smaller than GPT-4 Turbo. What OpenAI basically did was make the model as much smaller as possible while keeping roughly the same quality, or slightly better.
Going from GPT-4 to GPT-4 Turbo, the model was less than half the size. Going from GPT-4 Turbo to GPT-4o, GPT-4o's cost was way lower than GPT-4 Turbo's. They just kept shrinking the cost.
Now, with GPT-5, what could they have done? They could have gone for a big step. They actually tried that with GPT-4.5. They screwed up some things because it was really hard to get 100,000 GPUs to work properly. There are challenges there, and they also hadn't figured out the whole reinforcement-learning paradigm at that time.
The scaling laws are a chart of quality versus compute, but that compute breaks down into how much bigger I make the model and how much more data I put into the model. If the internet only has so many tokens, you're kind of screwed. There was potentially a cliff until reinforcement learning happened, where you could generate data and train the model to be better without the internet having that data.
They had this problem: you have X amount of compute and can service your users, but today, if people want to use my API, I rate-limit them because I can't actually serve them all.
Yeah.
If I want to use ChatGPT, I have to rate-limit the people who have the Free, Plus, and Pro plans—the $20 and $200 tiers. There are different rate limits. You can only do deep research so much.
I have multiple ChatGPT accounts because I use deep research. You kick off a bunch, read it, and you're like, “Wow, I learned a ton. Move on.” So you have this challenge where you can't actually serve your user base enough. How are they ever going to move up this adoption curve?
So then, as OpenAI, what's your choice? Do you go from GPT-4o to GPT-5? Do you make the model way bigger and not be able to serve anyone? Plus, because you can't serve anyone and it's slow to serve, the adoption curve doesn't really get going.
Or do you make the model the same size, which is what they did for GPT-5? It's basically the same size as GPT-4o and roughly the same cost—actually a little bit cheaper, potentially—and then you just serve way more users and get everyone up the adoption curve more.
And then, instead of putting them on a bigger model, you put them on models that do thinking. So, if you've used GPT-5 Thinking or GPT-5 Pro, there's more intelligence there. This is the whole conundrum they have, and this is where the whole tokenomics thing comes into play.
The question you had—I wanted to level-set it—which is: How do you serve these users? The demand is growing so much.
I'm not doubling my hardware every 2 months, but I'm doubling my tokens every 2 months. There has to be enough of a cost decrease, and there is, at a given level of intelligence.
11. Inference Latency vs Cost Trade-offs
If you could snap your fingers and change a dial somehow, what would most unlock and unleash more development? Is it just inference latency? Because then we could do bigger models and serve them much faster in a way that consumers would enjoy. Is that the main bottleneck to be attacked?
Inference is always a curve. All of these things are curves, and it's a trade-off. Everything in engineering is a trade-off. You have inference latency versus cost on any given hardware.
GPUs can do lower latency to a certain extent, but then the cost is way higher. Or you can do really high throughput, and the cost is way lower. The company just kind of YOLOs it: they set the dial where they think it makes the most sense.
There are other types of hardware that aim for their curve to be at a different spot. Maybe the GPU curve is here, but on latency over here, you're in very diminishing returns. Someone made a little curve right here: “Okay, maybe that's a useful point,” but actually the market cares about this point.
Anyway, there's a curve of who cares about latency. I think if I could just press a magic button—
Yeah. Is it capacity? Is it latency? What is it?
I think that's a tremendous question. I'd probably still say capacity and cost are more important than latency.
Really?
I think existing levels of latency are fast enough for a lot of things. Now, if the latency were 10 times lower for GPT-5, they could have made a model that was 10 times bigger and served it at this speed.
Yeah, that's what I'm wondering about.
But then you would have the same capacity issue. I guess if you could have your cake and eat it—which is all the capacity in the world and the lowest latency in the world—then you would just make the best models. You'd make the models way better.
I think it's the physical realities of, if I'm at OpenAI, what do I choose to do? Do I invest more in the model that people can use, or do I invest more in the fast one? Do I invest a lot in a model that most people won't use because it's expensive, first of all?
And even those who can afford it will often go back to the regular one.
Right. I have access to a likely Claude 4.1 Opus. I still use Sonnet way more.
Just because it's a better experience.
Right. It's dumber. It's objectively dumber, but it's slow.
Yeah.
And I don't know—my time's worth something, right? I think OpenAI wouldn't have been afraid to make a model way, way, way bigger with a terrible user experience.
12. Over-Parameterization and Model Learning
Yeah. As a result, we're probably just going to have to wait a little bit longer to see what the bigger models are like in practice, to see what consumers actually do with them, because it's just going to be too hard.
13. The Future of Reasoning and Compute Scaling
It’s not necessarily even bigger, right? There’s this whole concept of overparameterization. If you just throw more parameters into a neural network—and even when humans, I’ll equate it to humans—when you had a vocabulary test or some other test, you memorized before you understood. It wasn’t until you did multiple repetitions and encountered the material in different forms that you actually understood the content rather than just memorized it. It takes cycles.
When you do an LLM, it’s the same thing, right? If you throw some data at it, it will memorize it before it generalizes. It’s this concept called grokking, right? You grokked a subject; it’s like the aha moment.
Patrick O’Shaughnessy
The trick of understanding.
Yeah. The models do the same thing. They memorize it up until then, and they understand it at some point. If you make the model bigger and bigger and bigger without the data changing, you just memorize everything, and actually, it starts to get worse again because it never had the opportunity to generalize. The model was so big, and there were so many weights and so much capacity for information.
The challenge today is not necessarily to make the model bigger. The challenge is: how do I generate and create data that is in useful domains so that the model gets better at them? There’s nowhere on the internet to show you how to fly through a spreadsheet using only your keyboard and all these functions and all these things, right? That’s repetition. That’s sparse, but there’s no data on the internet about this.
So how do you teach a model that? It’s not going to learn it from reading the internet over and over and over again, which you and I could never do. And so, it has a level of intelligence that we can’t match. We can’t read the whole internet, but it can’t do basic stuff, like play with a spreadsheet.
So how do you get it to learn these things? That’s where this whole reinforcement-learning paradigm kind of happened.
Patrick O’Shaughnessy
Which is—
Giving it environments—specific environments—to learn in and then fold back in.
Patrick O’Shaughnessy
Right, exactly.
That’s where there’s sort of a challenge in terms of building those environments. There are, like, 40 startups now in the Bay Area doing these environments, and it’s questionable whether any of them will make it or what will happen. But there are 40 of them, and these companies are also making their own environments. These environments can be anything and everything.
Patrick O’Shaughnessy
Give me an example of one of the startups, just to get a sense of it.
These startups are just making environments for OpenAI, Anthropic, and others. It’s as simple as: here is a fake Amazon, because Amazon’s terms of service ban chat models and all these things. But here’s a fake Amazon full of items. Figure out how to click around and purchase items.
Figure out how to compare 2 items and pick one. I’ve generated a list of deodorants: 3 of them are fake, 1 of them is real, and 1 of them is not the one I want. Here’s the prompt. Figure out how to buy it. It tries many things, and you vary the prompt and all these things, but eventually it’s bought the right deodorant, and you’ve succeeded and fold it back in. That’s a simple thing.
Or it could be: clean this data. Here’s this table with a ton of dirty data in it. There are colons and stuff. There’s an address in 1 column. How do I separate out the columns so the address is street address, city, and ZIP code? It’ll try a bunch of stuff, but maybe it can’t do that yet. So you give it addresses in different formats, and you slowly and iteratively teach it.
There are all these challenges. That’s 1 example. Another example is that you’re in a game, whether it’s tic-tac-toe, Call of Duty, or a math puzzle—whatever the game is. A lot of these environments initially have been math puzzles. It’s: do this math puzzle. “Well, I can’t do this one because it’s too hard.” Here’s an easier one. “Oh, okay. I can work on this one. Okay, I’m good enough. Now I can learn this one.” It has iteratively stepped through those, and basically, from Q4 of last year to Q2 of this year, these things hill-climbed through math puzzles like crazy.
Patrick O’Shaughnessy
Yeah.
A lot of that was not, “Hey, I just know the math.” A lot of that was, “Here’s how I use Python to write something that does the math for me.” Now these things are actually quite good at math.
14. Building Environments for AI Training
These environments can be super varied. It doesn’t need to be something that’s clear-cut and dry. It can be: here’s a medical case; what’s wrong with it? Then you have another model say, “Here are your instructions on how you would grade the result of a case. If it looks like they didn’t even try this or didn’t even look up that, okay, you did that wrong.” You can feed these models into it, so these environments can be very, very complicated.
Building those out is a challenge, right? It was one thing to say, “I’m taking all the internet data, I’m going to filter it some, and I’m going to throw it to the model.” There are tons of engineering challenges there, for sure. There’s a different set of engineering challenges that take time to build out in those 2 worlds: the pure, raw internet pre-training world and this new environments world.
Patrick O’Shaughnessy
How far into the potential benefits have we gotten in each of those? What inning are we in?
This is where people say, “Well, then, Dylan, what you’re saying is you never need to make models bigger again, right? Because you’ve already run out of data, and until you figure out how to generate tons and tons of data, that’s great.” But actually, we haven’t run out of data.
We’ve seen another angle where it’s mostly just been pre-training scaling, right? That’s Veo 3 and Nano Banana, Google’s image and video models, and Genie, and all these Google image and video models. That’s purely scaling on multimodality, right? The models still aren’t that great at video, audio, and images. They’re fine, but they could be a lot better.
There are angles of scaling there, because when I said we’ve run out of the internet, we’ve run out of the text—there’s tons of video, image, and audio. It’s just so expensive. We didn’t get to that.
Patrick O’Shaughnessy
So maybe late innings on text, mid-innings on pre-training?
I think we’re early on text. Yeah, we’re quite early. The other angle is that just because you’ve used the text doesn’t mean you can’t learn faster, right? You take a class, you give everyone a book, you tell them to read it once, and you test them all. One kid is going to get 100, and one kid is going to get 40. That’s just the reality of life.
Maybe if you read the book out loud to them, the kid who got 100 might get 30, and the kid who got 40 might have gotten 60. There are these different parameters, and when we talk about model architecture, the same thing happens there. It’s not like you stop training new models. It’s not like you don’t have algorithmic improvements or smarter kids. It’s not like pre-training is done.
Patrick O’Shaughnessy
Yeah. In fact, it’s the base of everything. You want to keep having gains, because any gains on pre-training—I mean, the model learns a little faster, or the model’s a little bit smaller for the same quality—
Yeah, it feeds into the next stage, which is this whole post-training side, which will subsume the majority of the compute at some point.
Patrick O’Shaughnessy
Inning-wise, are we in the second inning of that? How far along is it?
I think we’ve thrown the first ball.
Patrick O’Shaughnessy
Wow.
Because think about how we need so many environments.
Patrick O’Shaughnessy
So many environments.
I think my favorite thing is that my brother just had a baby. This baby will literally stick his hand in his mouth, and I thought about it, and then it’s like, wait—he’s calibrating the senses on his fingers by sticking his hand in his mouth, because his tongue is the most sensitive thing. He doesn’t know he’s doing it, but that’s how he’s calibrating. He’s like, “Oh, that’s me. I can touch and feel,” right?
How does the model learn these sorts of things? You just have to try stuff and fail. We’re so, so early in this. Think about how much we see throughout our lives and how much of that information we throw away. We throw all of this information away. I don’t remember anything about what I had for lunch yesterday. But if it was amazing or bad, I would have remembered that: “I don’t like this,” or, “I like this.”
There’s all this information we throw away, and these models and these environments are generating tons of data and throwing most of it away while training the model. But it’s infinitesimal compared to what humans have done. I think there are so many environments you can put the model in.
There are people who even think you don’t get to the magical AGI until you embody it—that is, you put the model in something that can interact in the real world, like a robot. I think Elon and xAI are a bit more along that angle. They think embodiment is required to get to artificial general intelligence, because you need the model to be able to say, “Pick this up,” or, “Oh, wow, this is a rotating thingy,” which you could never get from just watching a video about it.
Patrick O’Shaughnessy
You wouldn't get the concepts of it, even. Yeah.
15. AI in Everyday Life
And so I think we're so early in reinforcement learning, because that's what humans are. We're reinforcement learners.
Patrick O’Shaughnessy
And the so what of, let's say, we fast-forwarded—we're in the 7th inning of that or something like this. What do you think the way that the average person will most feel that difference in terms of the utility of the model?
It'll be very different, like modes of using it, right? It's one thing to ask for information or ask it to organize information versus it just doing things. Those 12-year-olds, you need to really direct them how to dig a hole, because a lot of them haven't dug a hole. But you're talking about, “Order me this vitamin,” and it's just done, right?
We're actually not too far away from that. I think if you try and research electric toothbrushes—this is something, because you know, with my electric toothbrush, I lose it. I leave it at a hotel all the time. I've been obsessive about this. In 2021, I made a spreadsheet of all the electric toothbrushes based on how many ICs were in each one of them, right? This one has a Bluetooth IC. Why? I don't know. This one has a display IC. It has a color display IC. What's going on, right? So I made a spreadsheet of all this. It's this weird little thing that I do.
I've been finding that when I research which toothbrush I want to buy now—I bought an Oral-B iO Series 9 or whatever, right?—comparing them, these models can actually figure out exactly what you want. More than 10% of Etsy's traffic is straight from GPT.
Patrick O’Shaughnessy
Wow.
Amazon blocks GPT, but otherwise it would be really high. People make purchasing decisions through GPTs; they just don't make the purchase.
OpenAI's head of applications, or CEO of applications, was at Shopify and created the shopping agent, right? This is very clear. This is how they monetize. The models are going to purchase for you, right? They're going to do actions for you, and the model—and then the company that does those actions for you, the model—will be able to take some sort of take rate.
Even if it's 0.1%, even if it's 1%, even if it's 2%, it'll be like a credit-card transaction. Visa is the most amazing business in the world because of this, right? ChatGPT could be that, too. If I'm making my decisions on purchasing all sorts of things, I mean, I already almost outsource what I'm going to eat to the front-page recommendation of Uber Eats sometimes. I already outsource a lot of decisions. It's not too much further until I've completely outsourced a decision and a purchasing intent.
That's what's made Amazon and Google such amazing companies. They figured out how to get the thing you want to purchase in front of you as best as possible, right? All their work on recommendation systems is figuring out what you like and how to keep you on the platform longer, whether it's YouTube or Instagram or ByteDance with TikTok, or, “Hey, here's the ad of the thing you'll probably click on and buy, because that's how I get paid.” Everyone likes to claim they don't pay attention to ads, but you do, right?
Before asking even more holistically about your view on where we're going, there's a third category, which is the reasoning part of the equation. We've got pretraining, we've got RL and environments, post-training. What about just raw time spent reasoning, and where is that going as its own independent part of the overall scaling law?
The scaling laws, again, if you zoom out, that's not actually what the original paper is, but in spirit, sure: scaling laws are more compute, better intelligence. That could be bigger and bigger models; each iterative token is better—whatever word garbage I spew out. If I went back and wrote about everything I talked about in this, I could make it way more condensed. It could be way more clear, potentially. Right now, the benefit of a podcast is that a lot of times people find it more fun this way.
Driving, it's fun.
Yeah, exactly. They're walking their dog and listening, whatever it is. But the interesting and important thing here is that by putting it in these environments, you're teaching it like humans, right?
If I asked you to go figure something out, you might not necessarily know the answer right away, but I know you could probably figure it out in a given amount of time. That's reasoning. You're spending more brain cycles. The magic, again, of the intelligence of humans, of people, is not that they are the best at information retrieval. GPTs are amazing at information retrieval. We're really good at figuring out how to do things iteratively because we've been trained in these environments, which is our world.
And so reasoning and these environments are linked together, right? If I'm telling a model, “Hey, do this math puzzle,” it's not just spewing out, “Oh, the answer's 1. Oh, the answer's 2. Oh, the answer's 3. Okay, the answer was actually 7.” And when it got there, I trained it again. It's like, “Okay, now it knows: next time, the answer is 6, 7, or 8. Now it's 7. Okay, great.”
It's not like now it instantly knows the answer. It's actually like, “Oh, here's this puzzle. Oh, these numbers. Oh, this line—it's Sudoku. These numbers add up to this. Oh, it has 1 through 9, but it's missing 8. Okay, it's 8.” It's thinking through it, right? Like you and I would solve a Sudoku. Eventually, when you get good enough at Sudoku, you could probably just spit out an answer. You could do it in your sleep.
This reasoning time is a way of spending more compute, more brain cycles, on the task without actually scaling the model. Then the model becomes more versatile, right?
Humans have a rate. If I just held a match against you and you didn't notice it, you'd immediately jerk, right? The rate at which you operate is hundreds of hertz. Your body can actually take actions at hundreds of actions per second. If you look at a fighter pilot's reaction time—the peak of human reaction time—what reaction can they do? It's completely primal, instinctual, right? Very little thought is put into it.
If you think about this alien intelligence that we're trying to make, is it immediately going to one-shot the answer always? No. But at times it needs to be able to tell me exactly the answer in 2 seconds or half a second, or whatever action it needs to take immediately. But a lot of times it also needs to think through the problem and go and do stuff.
That's why you hire students. That's why you hire interns, because you're like, “Yeah, I know this data exists. Here's the format I kind of want it in, and go figure it out.” Then they spend a whole summer doing something you could have done in 3 days, but great, they learned a ton, right? These models need to go through that progression.
When I think about reasoning and RL, it's a lot about how the human psyche and intelligence work. I wouldn't say there's no caution in trying to make it too much like humans, because it's not—the fundamental substrate is not like humans. The processing is not like humans. Our brain is very different from how these ALUs on a chip work. The scaling of these things is very different. The raw speed, the amount of words they can process—everything is so different.
16. Memory and Context in AI: Short-term vs Long-term
But at the same time, it's important to reckon back to what actually makes people smart.
On the topic of embodiment, and continuing with the human analogy, how do you think about things like short- and long-term memory in a human versus just raw model capacity or something? What role does that analogy of memory—I don't mean literally semiconductor memory, but memory in a model—play? How do you think about the importance that that will play, and where are we in that?
The magic of transformers was attention. I calculate everything in my context length. I calculate the attention to each other, right? Basically, in a vector space, like king and queen, there are these vectors. There are dozens of vectors for each word. King and queen are actually exactly the same on a ton of stuff, but then it's the opposite on one dimension because one's male and one's female.
That will have a lot of other ramifications throughout other literary stuff, like what adjectives do you put with a male of these vectors? It's like regal and powerful and could be ruthless, whereas a queen could be dignitary or whatever. I don't know—stupid analogy. But when you think about how that applies to humans, what we're terrible at is exact recall. I could tell you a sentence and tell you to repeat it.
Yeah, it's like 6 numbers the average person can remember or something like that.
Right, but you get the gist of the sentence. If I told you a whole paragraph, you'd get the gist of it and you could repeat the meaning of it to someone. You could translate that meaning. So models are very different, right? Fundamentally, transformer attention has been calculating the attention to everything to each other and getting the models to actually be able to recall.
That’s been a training-data problem. You can get the model to repeat exactly what you want—anything in its context length. The needle-in-a-haystack problem was a benchmark that people used for a while because models had to get good at that, but now models are just amazing at it. You can say, “Tell me something from a random part of your context,” and it can do that.
What they really suck at is having infinite context, because when you have infinite context, it’s sparse. That’s the real word: sparse. You’ve taken this entire world and encoded it in such a small amount of data that lives in your brain, and it’s so sparse, but you understood how to grab the fundamental reason and put it down there. Models haven’t been able to create something sparse yet, right? How do you reason over the context of infinity?
Humans maybe have short-term memory and long-term memory. I think it’s a lot blurrier than that. There’s no clear line where you can say, “This was in my short-term memory, and this is in my long-term memory.” It’s much blurrier, but as we go back and back and back, it gets more and more sparse, right?
If we think about, “Hey, what do you remember as a kid?” the craziest thing in psychology—I remember when I learned it, I was like, “Wait, my memory of what I did as a kid with my dad at this thing is fake.” I remember it, invent the picture, and then remember that picture successively, but the actual memory of what happened has morphed a little bit over time.
There’s a ton of research going on in the domain of long context. How do I get longer and longer context without blowing up my model cost? This is a big challenge with reasoning. This is why we had this HBM-bullish pitch for a while, right? You need a lot of memory when you extend the context. It’s a simple thesis, right?
But the fundamental algorithm needs to change and improve iteratively over time to get to something like this short- and long-term context of memory. That doesn’t necessarily mean the model has to work like we do, right? Why can’t the model just reason and have a database that it writes stuff in, or a Word document that it writes stuff in, and then it takes it out of its context, works somewhere, and recalls it back? We don’t do that, right? You and I refer to our notes, our calendars, our texts, and anything else—the shopping list, right?
I know I need food for dinner, so I go to the store and I’m like, “I need a shopping list,” because otherwise I’m going to buy stupid shit, right? The model doesn’t necessarily have to fundamentally work the same way as humans. But there is that challenge of, how do I train the model to operate over the context length of a human? How do I train it to interact with these databases and Word documents that it writes to?
It’s never going to learn that from pretraining. It has to learn that from an environment. But these environments have to be architected in a way where the model knows it can write stuff down and refer back.
One of the first things OpenAI did was Deep Research, right? Everything is not in Deep Research’s context. Deep Research is working for 45 minutes, outputting millions and millions of tokens, and creating this amazing thing that it wrote. It’s pretty good research. I would say a lot of memos that you read from people are on par with Deep Research, at least a junior.
How did they do that? They enabled it to write something down elsewhere, have this recall, and effectively use language to compress the information that it looked at. It put that off to the side, used language to compress other information off to the side, used language to compress other information off to the side, and then looked at all this compressed information and wrote something. That’s sort of what Deep Research is.
So how do models get there? I’m not sure. I think it’s a fundamental research challenge. It’s why these companies need millions of GPUs to train on—not because they’re going to make a million-GPU model, but because they need to try a bajillion different things. They don’t know what will work, and what’s going to work for humans is so different from what works with models.
17. The Spectrum of AI Optimism
There are any number of parameters or things you could tweak that could end up changing how it develops, right? How good is it if I do it this way versus that way? That’s the whole point of ML research: you’re constantly trying stuff out and trying to get better and better.
If I add all of this up and hold the mirror up, it seems like I would put you in the category of unbelievably bullish on what these things are going to be able to do in 10 years, or something. Pick your timeframe.
Yeah.
Am I calibrated the right way? Among everyone you talk to who you respect and think is—
I’m much more bullish than a lot of people, actually, which is the crazy thing.
Help me understand that distinction. If you’re rating yourself from 1 to 10 among the people that you respect, with 10 being the most bullish, where are you? What’s the difference between you and the person who’s a 10?
I respect you, but I know I’m way more bullish than you. I respect Mark Zuckerberg, but I know he might be—he’s probably—maybe I don’t know if he’s more bullish than me. Sam Altman is definitely way more bullish than me, right? He says we have artificial general intelligence in less than 1,000 days.
Dario Amodei, I respect him immensely, but he’s way more bullish than me. My roommates are also way more bullish than I am. One of them is an Anthropic ML researcher, and one of them is another podcaster, Dwarkesh. They’re both way more bullish than I am.
Really?
Yeah. But even they’re not as bullish as some researchers in this field. If I go talk to someone I respect—some famous investor, right? Any of these famous investors; I don’t want to name one because I’m scared—but there are all these famous investors. They’re not more bullish than me, and the stuff I’m saying sounds like crazy shit.
Some of it, though, is timeline. I’m actually even more curious about the upper limit, the extent to which there is—
The upper limit? I think I’m among the most bullish you can get, because—
That’s what I mean.
The upper limit of this is that this will just be smarter than humans. I don’t think that will happen anytime soon. Even if that doesn’t happen anytime soon, there’s so much valuable stuff that can be done with these models that economically, we will skyrocket.
There’s so much value that can be created in the world just by saying, “Hey, if the models know how to do COBOL-to-C and Python migration of mainframes—”
Just migrate everything.
18. Timeline to AGI
Migrate everything from mainframes to the cloud. How much more efficient would the world be? Making all these random applications and automated reports, and stop using Excel as a database. Instead, you can make a real database and manipulate stuff in Excel.
There are all sorts of humongous business efficiencies or automation that could happen without the model ever being— We could literally just pause it at the 6-month-from-now timeframe of how good it is at software development, and it would be a godsend in terms of how much efficiency and value could be created for the economy. It doesn’t ever have to get to digital god.
Now, I do believe—
We’re going to get the digital god.
Eventually. Eventually—is that 10 years? Is that 5 years? Is that 100 years? Is that 1,000 years? I don’t know, because there are so many unknown unknowns.
Like I mentioned, babies are putting their freaking hand in their mouth to calibrate. Later, they put their foot in their mouth, and they’re like, “Oh, that’s my foot. Here are the senses on it.” Then they can pick up stuff in their hand, and they no longer have to put it on the most sensitive part of their body because they know what it is. Or they’re like, “Oh, this is a speck on the ground. What is it? It’s not food, but now I know what it feels like inside my hands, and I’ve calibrated.”
The models haven’t gotten there yet, right? They have no idea how to do this. Digital god is—well, 1, I kind of believe in embodiment, and you need a non-digital body. You need something physical, and you need the capability of having touch and feel and all that to truly have an experience like humans and be smarter than us in every way.
19. Physical Intelligence and Embodiment
But you know, that’s so far away. What do you think about what Physical Intelligence is doing, attacking the—whatever you want to call it—large movement model, large robot model, or something?
What they’re actually doing today is like, holy shit, it’s so simple in terms of what it is to a human.
Yeah. To models, picking this up is freaking hard. How much do I squeeze with my pinky versus this finger versus this finger versus this finger? I don't know. You pick up a glass of water and tilt it, and that's impossible for a model today.
At the level of dexterity, if I was swishing a wine glass, think about how simple that is. You don't even think about it, but you instinctually pick up a wine glass, swish it, let the aroma out, and smell it. That little swish requires so much tactile feedback and movement, and these models can't do that yet—nowhere close.
So, I think, yes, but it doesn't need to be that good. It doesn't need to be able to swish a wine glass and not break it, put it back down, tilt it perfectly, and not spill it. It doesn't need to do any of that to be tremendously valuable. What it needs to do to be tremendously valuable is pick this up and put it down here after knowing what it is.
Getting data. Yeah.
Yes. I think the robotics world is huge. I think we're literally warming up.
We haven't even left the dugout, right? We're nowhere close to scaling robotics. There's a ton of work to do; the data flywheel needs to get going there.
20. Talent Wars
One of the most interesting subplots of this whole world is the talent wars. A cool idea is that, as these things get better, maybe we begin to automate some of the research function that people formerly would have played. Do you see a world where we're squeezing down to fewer and fewer people who really matter, who will have all the impact on where we go in terms of net-new research? Does that mean all this crazy spending that's happening at Meta or elsewhere makes a lot of sense—maybe even those numbers should be higher or something like this?
I think it's tremendously hilarious that people are like, “Oh my God, this person's getting paid a billion dollars.” It's infeasible. How could this person possibly be worth that much? Well, they're running experiments on chips that cost $100B. If every wasted experiment they do used just a third of the compute, and their ideas and their impact on it wasted that compute if it was an idea that was already done, there's so much wasted compute.
I'll say I call it wasted; it's trying stuff and failing. But none of us know what to try and what not to try, and these things are so complicated. There's a group of people just trying different stuff on the existing data: How do you mix it? What order do you feed it into the model? How do you filter it? What's the architecture?
There are different people working on long context, and different people working on every single aspect of the model. If you just make them a little bit more efficient—if they come up with an idea that's 5% more efficient—fantastic. I just saved not only 5% of my compute time and training time, but I also save 5% across my entire inference fleet. And then I do it again and again and again and again, because we're so far away from these models being anywhere near as efficient as a human brain, and we know they can at least get as efficient as us. Maybe the compute substrate isn't the same, but whatever, right?
Adding more people to the problem doesn't make it faster, right? There are so many things you're trying. You run these experiments, you learn something, and then you implement it. You tweak the knobs in 100 different ways, and then you see the trend line and you're like, “Oh, so actually I should tweak it this way. Let's implement that.”
There's so much gut feel. There's so much reading data, understanding it, and reimplementing it into these things that, if you add people, you're going to slow it down. In a sense, a lot of Meta's problems before they did the superintelligence thing was that they just had too many people who weren't led by leadership that was amazing. They had a lot of failed experiments and wasted time doing things that didn't matter.
There's a tweet from one of my friends at OpenAI. He's pretty famous on Twitter; his name is likely Roon. He said, “I get visibly, viscerally angry every time I think about how many H100s Meta is wasting.”
It's such a funny tweet because, well, yeah, they're wasting a ton of compute. They were—maybe they still are—but everyone's wasting compute, right? OpenAI is wasting tons of compute because what's the Pareto-optimal model architecture? Who knows?
Another thing I saw likely Roon say recently, which was so interesting, was, “Why don't we just go make even more ridiculous offers to the people who have process knowledge for things that we want here in the US and in other countries? Why don't we, if we're getting pretty good at the Arizona fab that we've built and we think that we can sort of extract the process knowledge from the people, go acqui-hire all the best people in Shenzhen or all the best people in other places in the world?”
Do you think it starts to escalate to that level? So much is dependent on the process knowledge of a relatively small group of people, and the talent war should actually not be Meta and OpenAI; it should be the US, maybe through Meta and OpenAI, and people from all over the world. Do you think it starts to get that extreme, and should it?
That's almost a function of why Intel has fallen off a lot, right? You have all these geniuses in nanochemistry, with PhDs and all these random things, whether it be chemistry or physics—all these incredibly smart people. But there's a whole class of incredibly smart people that never went that way because they're like, “Oh, those guys are making $200K. Why would I do that? I'm going to go to Google and make $800K, and now I'm going to go to OpenAI and make $10M—or no, I'm going to go to Meta and make $100M.” Any smart 18-year-old is going to be like, “Fuck that. I'm doing this,” right?
Why do the smartest doctors—and I don't mean the smartest doctors in a general sense, but the really smart-skewed population of doctors—want to be dermatologists and anesthesiologists? Is that the most valuable thing for them to do?
No. But those are the 2 professions that give you good working hours and great pay.
Yeah. Not to say that the general doctor is not as smart as them, but if you took the population of general family doctors—just the random doctor—and you took the population of dermatologists, the newest coming out of school, the ones becoming dermatologists and anesthesiologists, are way smarter, or at least scored better and were able to get into the field that was coveted.
And so, yeah, the talent wars—it is truly like, we've sort of been through this process of human capital and capital goods, those 2 vying with each other. For a long time, with mechanization and industrialization, we had human capital decreasing as industrial capital increased. That got to a point where, especially in the ’70s, it really started to tank as the ability to globalize and all these things started to really hit the US. That's why we have a lot of the population-level dynamics and income inequality that we have today, which is very bad for the psyche of the US and its stability.
But now we're in such an age where manufacturing things is pretty commoditized. Most of the value doesn't come from the manufacturing of it; it comes from the creation of the idea. One thing Jensen told me, which I thought was amazing, right? He's like, “Dylan, the reason America is rich—people have it all wrong. The reason we're rich is because we've exported all the labor, but we've kept all the value.”
That's what Nvidia does, right? They've exported the labor of making their chips, and Apple, right—everyone. It's done in Asia.
And those companies make money.
Not as much money as Nvidia and Apple, right? All the gross profits are going to them. They're either reinvesting it or buying back stock or whatever. How they allocate the capital is a different concern.
If, as you said, the process knowledge is so valuable, why aren't we doing this? That's a great idea.
Roon's idea, not mine. Yeah, no, I mean, I think the challenge is how to choose people. It's really difficult. For some roles, if someone can talk the talk, they're great, right? People just automatically assume they're great because they can talk the talk. But how many people suck at talking and are really freaking good at doing?
Yeah.
Yeah, but then you don't know. You don't know, right? Because there are people who talk about being able to do better than the person who's doing it. These tests are never as good, right? Work trials—how do you select? And this was a big challenge for Meta. Some of the criticisms are that they didn't get all of the best people.
They actually got a lot of bad people. The cope from OpenAI and Anthropic and these kinds of companies is, “No, no, no, they didn’t get our best people.” That’s what Sam said, right? He’s like, “They didn’t get our best people.” Meanwhile, he did have to do counteroffers internally, right?
As far as process knowledge, I think ML researchers are an extreme example of how much value one can create. My favorite analogy that I came up with recently is that ML research is exactly the same as semiconductor manufacturing. There are a ton of jobs in semiconductor manufacturing that don’t exist in ML research, but it is a ton of tuning a thousand different knobs, right?
You put the wafer in this tool. You’re going to change the pressure of the chamber when you’re doing the deposition. You’re going to change the mix of the chemicals flowing in, which chemicals you’re putting in, what speed you do it at. Do you do it for 30 minutes? Do you do it for 31 minutes? There are so many knobs on every single tool, and you have a thousand input and process knobs, right?
Process knobs on each tool, plus the sequence of them all. You frankly cannot test everything, right? It’s impossible. It’s too large of a search space, just like designing a chip has too large a search space. You have 100 billion transistors. How are you going to possibly try every single thing? Impossible, right?
You just have to have enough intuition to pick that point, pick that point, pick that point, see the data, and say, “Oh, okay, I think the answer is here,” right? Then just YOLO, right? Once you think the answer is here, you test here and you’re like, “Okay, here.” But a different person might have seen these 3 and then said, “Okay, the answer is actually here, not here,” and the data is fuzzy. It’s somewhere in the center.
This whole idea of ML research is that you spend a lot of time on compute training, doing what were effectively useless things besides teaching yourself what’s the right thing to do and what’s the wrong thing to do. Semiconductor manufacturing is the same way, and actually all process manufacturing is the same way. If you’re iterating super fast and trying to get better and better and better, or you’re optimizing a process on a chemistry or whatever it is, you try, you fail, you learn, you do.
In semiconductor manufacturing, maybe it’s just running tens of thousands of wafers. Your R&D cost at Intel—or the cost of your main fab that is running the R&D—is very, very high, and it’s producing zero economic value besides teaching you how to do the next node, which then you can deploy at volume. That is what actually makes the money.
I want to go back all the way to where we started and ask about what I’ll call the wellspring or the fountain of power in this whole ecosystem. I want to understand how you think about who has the power and how to keep or generate power as a business.
21. Power Dynamics in the AI Ecosystem
It seemed like talent—maybe talent is the very beginning of the chain, and he who has the talent, on a long enough timeline, has the power or something like that. But there’s also structural stuff, like the scale, the industrial scale of some of these things, which just takes forever to build or whatever.
How do you think about even smaller, zoomed-in examples? Cursor is unbelievably popular, the revenue is insane, and so much of it goes back to Anthropic. Who has the power in that relationship? How does that dynamic change over time? The power dynamics are so fascinating in this world, and I’m curious where you think it comes from in the first place, where it exists today, and where it will go in the future.
When we think about the power structures, you mentioned a really interesting one: Does Anthropic hold all the cards in this Cursor relationship? Cursor has, I don’t know, nearly $1 billion of revenue now if you do current month times 12. That’s a ton, but again, their margins are what they are, and they’re sending most of it back to Anthropic. Some people say their margins may be negative. I think they’re slightly positive, but regardless, they’re sending most of it back to Anthropic.
The gross profit dollars are at Anthropic right now.
Anthropic is taking all the gross profit dollars and putting them into compute—
For training. Yep.
So then all those gross profit dollars are going to, like—
Well, maybe Jensen, or maybe Amazon, who’s then sending it to Nvidia, or Google, who’s sending it to Broadcom. The gross profit dollars are going to the hardware layer from all of this, for sure.
Does Anthropic have all the power? The common view is yes from a lot of people, but then it’s like, well, Anthropic only makes the model that’s generated the code. There’s a lot more in this system, right? Cursor gets all of the data. They get all of the users. They get how they interact with this. Anthropic doesn’t get that.
They get a prompt. They send a response. Prompt, response. Now they have Claude Code, which is taking share, and it’s very different from Cursor. But Cursor is like, “Oh, well, I’m training embedding models on your code database, and there are actually multiple models that I’ve made. I’ve made the embedding model. I’ve made the autocomplete model. I can switch the Anthropic model to the OpenAI model whenever I want to. I’m only using the Anthropic model because it’s the best one.”
Because Cursor has all this data, maybe it can train a model—not for everything better than Anthropic, but for its segment better than Anthropic. The power dynamics are—
22. Microsoft and OpenAI: A Shifting Power Balance
It’s weird. They’re frenemies, right? Everyone’s a friend, right? It’s the same as with OpenAI and Microsoft, the craziest power dynamic going on in the world. They signed an MOU that said they had an understanding of what the deal would actually be for them converting to a for-profit. What is going on here? This sounds like the most non-announcement announcement ever. The power dynamics of all this—it’s the most fascinating soap opera ever.
One of my friends was telling me about KPop Demon Hunters. I don’t know if you’ve heard of this.
I have a 9-year-old daughter, so it’s all I hear about.
You’ve seen it a lot. I had just heard about it, and they were like, “Oh, let’s watch it.” I’m like, “What? Whatever.” But there’s drama. This real-world power drama is way cooler than this—
At least for you and me. Which parts of the drama interest you personally the most? Where do you think the stakes are the highest in the various subplots?
The Microsoft-OpenAI one is absurdly interesting because, at one point—2023—it was like, “Microsoft’s going to own the world.”
Yeah. Right. 2024, a lot of it too.
And then in H2 2024, Microsoft backed down a lot. Right? They pulled back because Amy Hood and whoever else at Microsoft—Mikhail Parakhin, whoever—were like, “Maybe we don’t need to be on the hook for $300 billion. We’re not going to build out $300 billion worth of compute for OpenAI.” They can’t pay for it. That at least had to go through their head when they cut back.
They paused a bunch of data centers, right? They said, “Oh, you know, we don’t need to be the exclusive compute provider. You can go to Oracle. It’s fine.” They relinquished this power, right? Now Oracle has that deal.
But then OpenAI sends, like, 20% of its revenue to Microsoft—or API revenue, or something like this. Microsoft has this 49% capped profit structure on OpenAI, and then there’s this whole IP-sharing deal. It’s really hard to understand the mechanics of the OpenAI-Microsoft deal.
They’re trying to renegotiate this. OpenAI doesn’t want the whole deal to be, “When we have AGI, you no longer have API rights or IP rights.” What the fuck does that mean? If you asked someone 20 years ago and put them in front of ChatGPT, you know—
AGI.
This is fucking AGI. It knows everything and it can have a conversation. I can’t tell it’s not a human. Actually, I can tell it’s way smarter than a human.
Yeah.
But now it’s like, “Ah, whatever. I can’t do XYZ.” The bar always moves, no matter what the level of intelligence is. For me, it’s going to be like when the thing puts its hand in its mouth and says, “Yeah, this is me, because I’m a human.” That’s sort of the sentience, the consciousness of it all.
23. Nvidia's Dominance and Balance Sheet Strategy
Another power dynamic is the one around Nvidia and the hyperscalers, right? Nvidia is the king. All of the gross profit is going to them today—pretty much all of it. Sure, TSMC makes some. Sure, SK Hynix makes some, but they have to invest a ton in capex. Sure, Broadcom makes a bunch, and Broadcom makes a ton of gross profit off of these companies, but Nvidia makes by far the most gross profit in the industry, and it’s not even close.
And so, going back to the analogy of, well, they're king and they want to continue to be king, and they want to make sure GPUs continue to be the most used, but also they can't buy anything. They can't buy any companies. They weren't even allowed to buy ARM when they were a nobody, right? I don't mean nobody, but they were pretty much a nobody on the grand scheme of things, and they weren't allowed to buy ARM in 2020 or whatever, or 2021—whatever the time frame was.
They totally could not buy any major companies. They'll buy smart startups. I bought a startup that I was a seed investor in and an adviser in recently, but they can't buy a real company. So, what do they do with all this cash flow?
And, sorry, but you're a loser if you just do buybacks. That's admitting you can't get higher returns—yep, on your capital.
On your capital, which is fine. Meta, Apple, and Google were mature companies for a while. Guess what? Those companies aren't going to do buybacks ever again, right? Or not like ever again, but for a while—
Because they have way more—they think there's better ROI for their capital now.
And Nvidia, if you look at Jensen, he's always flirted with buybacks, but mostly he's been reinvesting in the business—
But you can't reinvest that much into the business.
So, like, how do you—
He's doing demand guarantees. He's doing all this crazy stuff now.
Yeah. Right. Right. He's using his balance sheet to—
Win.
Yeah, try and win more. Right. Which is an interesting dynamic. I don't know if there's ever been anything like this in terms of the non-anticompetitive nature of this, right? Where you backstop clusters.
CoreWeave recently got a deal with Nvidia where Nvidia backstopped a cluster. CoreWeave would have never built this cluster because it's for short-term demand, and renting GPUs on short term is a terrible business model, right? You want long-term contracts, and you want long-term contracts with people with balance sheets. That's the golden goose, but that doesn't exist so much.
So, you do long-term contracts with people who don't have a balance sheet, like OpenAI. And if you can't do that, then you'll do short contracts with people who do have a balance sheet, right? There's this whole matrix of who you rent GPUs to.
But from Nvidia's perspective, it's like, you know what I really love? When venture capitalists fund a company and then 70% of their round is spent on compute. They fucking love that, right? And that's what's happening with all these companies, whether it's Physical Intelligence, which is spending a lot on robot arms and shit too, but also spending a lot of compute, or it's any other startup that's raising—Cursor, whoever, right? And even if it's not directly—
It would be really good if it wasn't a 2-year deal or 3-year deal for that compute, if it was, oh, yeah, yeah, you can spend 70% of your round on 1 training run.
You know, leave a company with these ideas, gather the data, do the training run, and then you have a product. You try and show how good the model is, then you try and raise again. That's what would be really great for Nvidia, but no one wants to build a cluster that's predicated on that as the business model. That's crazy.
So, they have to backstop a cluster to do that. Or, hey, OpenAI might go to its own chip. It might go to some ASIC from another company, right? They might even buy TPUs. Yeah, you know, they might even go to Amazon, right? They don't really care. They're not beholden to Microsoft anymore—
Trying to serve a product to a customer. Yeah.
And they want to build the digital god, and they want to serve a product, right? Make revenue, right? So, they don't have to go to Nvidia. Nvidia is the best option.
But you know what'd be really, really helpful is if I could, going back to the earlier part in this discussion, if the first year I get the compute upfront and I don't have to pay for the compute for the first year, right? Like I was mentioning, the $10 billion for the—
I can, for a full year, do training. I can subsidize inference. I can do all these things that build up a user base, and then I can actually pay for it. I have a year of a gigawatt to figure out a business model, right?
Whether that is serving free tokens and then implementing purchasing of stuff for the free user, or it's, hey, I have to serve this model at worse gross margins or negative gross margins initially, but then eventually I can serve it at positive gross margins because the models keep getting cheaper.
Or it's, I train the next-generation model that's so much better than everyone else, and then I'll win all the business for that level of intelligence because I'm the only one with an 18-year-old. You guys all have 14-year-olds, right? Who are working for you.
So, this is a—they can do whatever they want with this allocation. It's not an allocation of capital, per se. It's an allocation of compute. They get to decide what they allocate that compute to.
Nvidia's helping them by effectively front-loading it if they can find the capital. That company is like, “Oh, yeah, Nvidia's backing this too. Oh, you know, there's all these other things.” It's much more reasonable for someone to say, “Oh, yeah, I'll pay the capex because I know the first year is already going to be paid because you've got that investment from Nvidia. What about the next 4 years?”
If you ask a bunch of investors who are students of economic cycles through history—likely Carlota Perez-type stuff—they'll say that the concern is that every shortage is followed by a glut, and we always overbuild on long-lead-time, big-capex projects. You've got multigigawatt power being installed. You've got all this crazy stuff in semiconductors, and at some point it just gets overbuilt.
All the stuff we talked about earlier feels like we're not really close to that. There's so much freaking demand.
If the models don't improve, yes, we will overbuild, right? It's pretty simple. There will be supply-chain things where it switches from one supplier to another, and that's a lot of the nitty-gritty stuff we focus on. At the end of the day, if the models don't improve, we're absolutely screwed.
In fact, the U.S., in another year—if this lasts another year and then it happens—the U.S. economy will go into recession, straight up, because of this, and probably Taiwan as well, and probably Korea as well, right? Because there's so much buildup and revenue flowing through to us for this.
But when you look at these other things, like the bubbles of the past, some of them were just silly nonsense, right? Tulips—silly nonsense, right? Crypto, a complete Ponzi scheme, right? But then there's other stuff that's like, no, this was real, right? The U.K. spent some absurd percentage of its GDP on railroads for like a decade.
6% or something crazy.
Yeah, we're nowhere close to 6% of our GDP. Holy shit. But that was like, okay, there's something tangible, but it's like, oh, well, we overdid it, overbuilt, because how many goods are there to transport?
But also, you must reduce—you must build these railroads to reduce the cost of transport so much, because you have no clue when the demand stops, and you've overbuilt. And because there's 10 people trying to do it at once, you're obviously going to overbuild at some point.
Same thing with fiber. A lot of the argument against this is like, well, no, but this time it's the strongest balance sheets in the world. It's the world's most profitable companies. They can all pull the plug at any point.
Yeah.
Microsoft pulled the plug at one point before they were like, “Oh, shit. No, no, plug it back in.” They recently plugged it back in. They're like, “Oh, wait. We're restarting this. We're going out into the market. We're signing deals with Nebius for GPUs.” I don't remember how big the deal was. It's like $10 billion-plus.
Yeah. It's like $19 billion for Nebius. It's like, well, if they had just not pulled the plug on their data centers, they wouldn't have had to do that. They wouldn't have to pay those gross-profit dollars to Nebius, right? But Nebius made the bet that the demand is there, and they were right.
So, when you think about this, it's like—
What is the level of demand where this stops, right? If scaling laws continue, right? How—I mean, of course, there's an adoption curve, there's a pace, there's realities with capital, there's realities with supply chains; things take time.
But if you boil it down to it, it's like, your demand for 30-year-old senior engineers at Google who know how to make and program anything is effectively—I don't want to say infinite, but it's $2 trillion of value.
Yeah. Right. If I could have an intelligence as smart as a Google senior engineer, that's $2 trillion of software value, right? Because that's how much the world pays software engineers today.
You just go down the list of every other use case, right? If you have a simple physical-intelligence robot that can recognize a headphone versus water, or versus a phone, pick up the right thing, manipulate it properly, put it in the right spot, and sort it, that's worth how much to the distribution supply chain? I don't know, but a lot. We don't need to get to digital God for there to be immense value.
But the interesting thing here is that it's human capital, not capital goods. All of these other revolutions have been capital goods that reduce the amount of human capital you need.
Whereas this is just creating human capital.
Right, in a sense. In a sense, right? If I get everyone bulled up, right, and we're on this podcast, there's this—I don't know if you've heard the curse, right? It's like, if you talk about a stock on this podcast, it goes down, right?
I've heard word of it.
We're popping the bubble right now because the limit of AI is infinite. For the record, we went and did the math one time because I was sick of hearing about this shit curse, and it's just market performance. It's not—
Oh, really?
So, last time—it wasn't your—it wasn't this podcast; it was your other podcast. I talked about Applied Materials, and the stock was up like 70% in the 6 months after.
There you go. Yeah, I broke the curse.
I was like, hell yeah.
24. The Middle Layer
What do you think about all the companies in the middle? We've talked a lot about NVIDIA and then people at the end serving applications. What about these companies like Together AI, Baseten, Fireworks AI, and you mentioned Nebius—all these interesting middle-layer players? Are there amazing businesses to be built there, or do you think they're temporary patchwork to make the system work and serve demand? What do you think about this middle layer?
The cloud business model—let's say the neocloud business model. You mentioned inference providers and neoclouds. The neocloud business model is absolutely amazing or terrible, depending on how you do it, right?
It's terrible if you sign short-term contracts and just hope and pray you have short-term contracts forever. Initially, your short-term profits have amazing cash flows, right? You bought a GPU, put it in a data center, and paid for the power and all that. The cost per hour over a 6-year period for Blackwell is $2. Let's just call it $2 for simplicity's sake. It's not exactly that, but if I sold it for 6 months, I could get north of $3.50 or $4. It's like, holy shit, that margin's insane.
But what happens 2 years from now, 3 years from now, when I'm still selling 6-month contracts or 1-month contracts and the next generation of NVIDIA chips is out and it's 10x faster for 3x the cost? Naturally, the price of this should tank.
The other way to do it is, I actually have a long-term contract. I'm selling to OpenAI. I'm selling to Microsoft. The other end of the spectrum is what Nebius just signed: a $19 billion contract with Microsoft. Microsoft will pay no matter what. The market literally believes Microsoft will pay its obligations before the U.S. government because it's literally a cheaper bond rate, which is insane to me, but whatever.
This $19 billion has a huge gross profit, because the price per hour—and it's not exactly $3, and it's not $2—but the margins here are really good. Nebius is going to make at least $6 billion of gross profit off of this. Obviously, they have their operational costs, but $6 billion of gross profit off of this deal is insane. I would do that all day.
25. The Risk Spectrum
CoreWeave did that until Microsoft stopped going to CoreWeave, right? CoreWeave turned around and found other customers, selling to Google and selling to OpenAI. But now OpenAI is definitely not a real—you can't rely on its balance sheet. I still have amazing margins when I sell to OpenAI, but it doesn't have a balance sheet. So how can I be sure that it's actually going to pay what it signed up to?
In theory, this contract is worth a ton of money, and in CoreWeave's books today, all the contracts they've signed are mostly Microsoft, mostly money in the bank, right? But the OpenAI contracts—what if they can't afford to pay for this? Now there's a bigger risk and a longer and longer tail of these businesses.
There have been more recent deals with crypto miners, Google, and Fluidstack because Google is really short on data center capacity. People want to use more TPUs, but Google can't serve them all itself, so it's going to sell TPU systems to providers. Google is backstopping the deals with TeraWulf, which is one of the companies. I can't remember the other one, but there are 2 companies they've signed deals with where Google is backstopping the data center, plus selling the TPUs physically to another company. Then they're being deployed and rented, and Google still makes all the money.
Those companies are great as well, but then there's the long tail of questions: Is the enterprise demand there? Who's taking the risk? OpenAI is taking the risk because it's betting its entire company could go bankrupt if it doesn't work. Oracle is taking a risk because it's signing up for $300 billion of contracts and $200 billion of hardware spend across data centers and chips. It's going to have to go get debt.
Oracle is on the hook, and it'll probably be able to pay for it if it happens, but its enterprise value will tank if OpenAI can't pay for all the hardware that Oracle bought. Luckily for Oracle, it phases in over time and whatever, right?
Then you go to the inference providers, and there is a business to be made here too. I'm serving models. Maybe Roblox comes to me and wants to put an LLM in its game for XYZ reasons. Roblox is a good customer. Or Shopify wants to put an LLM into customer service. Yes, it could do it itself, but inference is a hard thing, especially as you get to larger and larger models, more complicated models, and all the other requirements.
There are all these different use cases where people want to serve models. Maybe it's just open-source models, and maybe it's fine-tuning of those open-source models, which those companies can help you do—or you can do it—and they can serve for you. They have scalable, reliable capacity. There are businesses to be made here, but there's also the YOLO version: I'm selling tokens to random people who are trying to build SaaS apps in SF, and maybe they run out of runway, right?
That funding doesn't directly go to NVIDIA, but you go through some steps and it's going to NVIDIA after some value chain. NVIDIA is holding no risk. Everyone in the middle has a lot of risk.
I'd love to hear your thoughts on going back to the other side of the equation—the app side, the stuff we're going to use these models to do—and the significance of this switch from deterministic code to a much different thing.
It seems like we're doing the thing we always do. Apple used to call this the skeuomorphic era, where you basically use the new technology to do the old thing you used to do. We're making engineers better. That would be an obvious current example, but it seems like we haven't yet gotten into the world where we're going to start using this technology to do things that we couldn't do before with deterministic code. I'm curious how you think about pushing the envelope on that side.
Why is that? I feel like that's exactly what we do with it, right? The cost to develop things is so high that you can't do it, right? Or the cost to have someone go buy stuff for you. You might have an executive assistant, and you can tell them to do this, but the vast majority of people don't. Now GPT is on the cusp of doing that, right? Go buy this, go do this, and it'll find the best thing and buy it. You just trust it enough, right?
26. AI for Material Science and Hard Tech
It takes time to trust these systems, but they're proliferating across tech. Tech is the most deflationary thing in the world ever, right? In terms of quality of life, it gets cheaper way faster than the revenues go up, but the revenues still go up. That's the fundamental basis of semiconductors, of tech, of everything, right?
Are we doing things that we couldn't do before with tech, with AI? Sure. The COVID vaccine was created with AI.
Like, it was AI drug discovery. There are entire briefs about how it was done with AI. And guess what? If another pandemic happened, I bet it would be even faster to discover the vaccine, if there's a vaccine for it or whatever, right? There are all these protein-folding things. There are all these optimization things.
There's AI for materials science and AI for all these other aspects of society. There's optimization. Maybe it's not in your face, right? It's not like, “Oh my God, AI just made this drug.” It's like, no, AI worked with the researchers who made the COVID vaccine, so we didn't have to all be stuck inside forever or whatever, right?
Point being, it's already happening. The whole “use the new thing to make the old thing faster”—sure, but if I go back 3 years, how many people would it have taken to deploy an image-recognition model that looks at every data center in the world and looks at the pace of construction and what equipment they have?
Assuming this is something you do.
This is something we do, right? How many people would that have taken? I don't think it would have been possible. My business model—this is the 2nd-highest-revenue product for us—would not have been possible if it weren't for AI.
Vibe coding, being able to dig through permits and regulatory filings, and being able to run image recognition on satellite photos—this would not be possible. This business is not possible without AI. Am I using it directly? Oh yeah, sure. I'm scraping through the regulatory filings and permits with LLMs and then manually reviewing them with people, and doing the same with the images, the satellite images.
Yes, there's a lot of stuff that the image-recognition model does. We also look at them a lot. Then it's compiling them and selling a spreadsheet with biweekly reports on all the data centers and what's changed. It's like, “Hey, actually, this Amazon data center—the fans are starting to spin, so there's actually revenue going on from this Amazon data center, so we can forecast Amazon's revenue,” right? Okay, this is relevant.
I don't think this would have been possible just a few years ago, at least based on the proof right now. Especially because there's demand for it—everyone wants to track this, and it's so important. But it begets each other, and I think, at least in my daily life, I don't think I could have taken that step from where I was in a business which was still a research provider. That is a monumental jump.
Being able to do it with 3 people out of the gate versus 50 or 100—I don't know how many people it would have taken, but I don't think it's possible. Mainframe migration is something people have always wanted to do. Amazon leaving Oracle took 20 years, right? They wanted to do it 20 years ago, and their next 4 highest-revenue products after EC2 were database products at AWS. And yet they still freaking used Oracle's database because it's hard.
Now mainframe migration can be way faster, or migration from one tech stack to another can be way faster. You can make your business more efficient. You add more automation. Yes, the tech exists. Go to all the businesses around the world, and they aren't using the leading edge of what they could. They aren't using what a 2020 company could have done without AI, right?
Right? No one is doing that. And if they did, they'd be so much more efficient, right? But all of these things just take too long to build. They're too expensive to build. You have your existing processes. How do you hand them over? How do you switch them over? How do you teach people to do this?
AI can help you with all of this, right? So you can take the pessimistic view that we're just doing the same things, but the value here is humongous.
27. Building Infrastructure
If it's tokens on one end, we haven't talked much about watts at the very beginning and power. What are your thoughts on what's going on here and how humanity is responding to this crazy new demand for just raw power?
The first approximation is that we're being a bunch of pansies and it's not that much power yet, right? Data centers, period, are like 3–4% of U.S. power—not the economy. Of that, like, 2 is regular data centers and 2 is AI data centers. That's nothing, dude. That's literally nothing.
It's just we haven't built power in 40 years, right? Or we've transitioned from coal to natural gas more and more over 40 years. Mostly, we just don't know how to build power. There are these regulations, there's not enough labor, and the supply chains for GE Vernova and its combined-cycle gas turbines aren't there yet. The same goes for Mitsubishi.
There's this random VPI-curing process for transformer coils, and there's only so much capacity, and it takes 2 years to build them. It's just a supply-chain thing. It's a lack-of-labor thing. It's not that it's actually that much yet.
At the end of the day, it's like, okay, wait, wait: you're telling me OpenAI is making a data center with 2 gigawatts, and that's the entirety of the power consumption of Philadelphia.
Like, that is real. Yeah.
That's insane. That's insane, right? But we used to get excited about finding a couple hundred megawatts of new data center. Now it's like, if it's not a gigawatt—
I remember the guy who leads that team. He was like, “Oh, it's just 500 megawatts, whatever.” I immediately opined. I also agreed immediately, but afterward I was like, “Wait a second, dude. That's a lot of power. That's like—how much? 500 megawatts is $25 billion of capex. Come on.”
Once you put in the GPUs and everything, right? That's a ton of money, but it's a snore because there's so much of it happening, right? We're learning how to build power again. We're getting the supply chains to do it again. We're reshaping the grid.
There are all these challenges with these AI data centers with regard to demand response and making grids unstable, right? AI workloads, because they change so much, so fast—especially training—you can just cause brownouts or blackouts. Especially if the grid doesn't have enough inertia, or if you're not putting enough things to dampen it in between the workload and the grid.
Even if it's not destroying it, the grid runs at 60 hertz or whatever, right? If you skew it up and down too much, these transient power responses mean your refrigerator will break down sooner—the motors in it—and you might not even know it because the data center's nearby. There are all these things. There are so many third-order effects here with AI data centers.
The funniest one is just that we're building power, right? Whether it's gas, which is a lot of it, whether it's through efficient combined-cycle gas turbines, or whether it's random generators that are not nearly as efficient—simple-cycle or, even worse, diesel generators.
There's a company that's putting a bunch of truck engines in parallel—diesel truck engines—because the industrial capacity for diesel truck engines is huge. No one's tapped it yet. So why don't we just put a ton of them in parallel and create this power-generation thing right here? Then you're generating power with a bunch of diesel truck engines in parallel and are able to power a data center. Okay, great. Because I can't get turbines, right?
There are all these crazy things people are doing: Elon buying some power equipment from Poland and shipping it to America because he needed that power equipment but couldn't get it here because the supply chains were weird. Any lack of capacity in the supply chain is being eaten up immediately, and then everyone's like, “Okay, let's invest.”
GE is like, “I'm going to double my turbine production.” It's like, “Holy crap. Okay, that's awesome.” Mitsubishi is doing the same thing. You go down the list: my transformer supply chain is expanding like crazy, and they're fully sold out, so I'm going to go to the Korean guys. That's fully sold out, so I'm going to figure out how to get the Chinese stuff in, even though it's not exactly what people want to do, right?
There are all these weird things. Electrician wages have doubled for mobile electricians who can work on data center stuff—or rather, contract workers. If you're down to move to West Texas, it's like 2015 again, like being a fracking guy, right? You don't need to be super-duper skilled. You can go to West Texas and make a shitload of money off fracking. But there aren't enough of those people.
That's why, right? If there were enough electricians in West Texas, if there were enough electricians in America, we could build these data centers faster. So there are all these little supply-chain quirks and weirdities.
Everyone's supply chain is different because the way Google makes its data centers is different from the way Vantage makes its data centers, which is different from the way EdgeConneX makes its data centers, which is different from the way QTS makes its data centers, which is different from the way Amazon makes its data centers. Their supply chains are not exactly the same, so you get all this weirdness in all these different supply chains.
28. Grid Regulations and Backup Power Challenges
No one really knows it. You go talk to power people, and on one end of the spectrum is Dario Amodei. Then you take a few steps and it's the average machine-learning researcher, then it's me, and then it's you, in terms of how bullish we are on AI. There are only a few more people: the standard New York stock investor, the semiconductor investor, the New York non-semiconductor investor, and then the Sequoia guy who thinks AI has been a bubble since 2023.
Then there's the utility guy. This utility guy is like, “I'm not building power. Power doesn't go up,” or whatever. Then you have the regulations around it. It's like, “How can I build a data center at this density?” Okay, great. I'll build the data center at this density and have all these backup generators. Great—now all of a sudden the grid is like, “Yeah, so what we're going to do is…”
This has happened in Texas, or it's happening in PJM, which is the main Northeast kind of area. These 2 grids are putting in rules that say, “Hey, big loads, we can tell you 24 hours or 72 hours beforehand that we're going to cut off half your power.”
We need to, because we need it for something else.
Yeah, people need to have their homes powered. We're not fucking like Taiwan, where if we're in a drought, we limit people's water usage and not the fab. That's a real story. I think it was around 2021 or 2022. There were multiple cities where they were like, “Okay, yeah, we're going to limit the showers you can take to 3 a day or 3 a week,” which is fine because they're East Asian and they don't have the smelly gene.
If you did this in India, it'd be cooked. I mean, it's already cooked. They'll limit the water to these people before they'll limit the water to TSMC, because it makes sense: the economic value of TSMC is way above the economic value of people showering 3 times a week. But the U.S. grid is not going to work that way. We're not that authoritarian; people have more say.
In Texas and in PJM, you can cut half the power if you give them notice. If you do that, then you need to turn on the generators that are there on-site. They're often diesel generators; maybe they're gas, maybe they're hydrogen stuff. There are all sorts of weird things people try to do just to ramp up power for that period of time.
But then, all of a sudden, the density of my generators means that I fail the air permit if I run the generators for more than 8 hours a month. So now what do I do? There are all these weird regulations. Even in Texas, it's really fun to watch.
We get to watch it. You get to see and watch the supply chain and try to—at least from my perspective—provide the data so people can trade on it, or provide the data so people can adjust their supply chains industry-wide. People in your audience can trade on it, or they can see and invest and make money, and allocate capital more efficiently.
29. US vs China: Who Really Needs AI to Win?
If I were to line up all the stages of this between the U.S. and China—power, semiconductors, models, applications, and so on—where do you think the most interesting differences are? What are the storylines between the U.S. and China at those various layers of the AI stack that are the most interesting to you?
When you look at China, they're a very formidable competitor. I think if we didn't have the AI boom, the U.S. probably would be behind China and no longer the world hegemon by the end of the decade, if not sooner. A world where the U.S. is not the hegemon is a bad one for Americans, at least.
Without AI, we're definitely just going to lose. Our supply chains are slower, they cost too much, and we're sliding. Our debt is unsustainable. Our economy isn't growing fast enough to maintain the level of debt. We're overconsuming relative to what we produce.
I'm sort of like a fucking bald eagle carrying America. It's bad for the world. The financialization—there's all this darkness and dehumanization in the U.S. in terms of social instability, partly because of income inequality, but also largely because of the visual nature of income inequality and the tendency of people to flaunt their wealth more because of social media and how that hacks people's brains.
The algorithm also serves people different content, so we're drifting further and further apart in culture. Think about the monoculture of everyone watching the same movies in the ’50s, ’40s, and ’30s versus now. You and I are pretty similar, and our feeds are completely different. Think about someone who's not in this world, in our similar worlds—their feed is insanely different.
AI has to dramatically accelerate GDP growth. Once you start talking about dividing the pie, you're screwed, right?
It has to be growing the pie, and you know this whole thing. I'm saying the U.S. really, really needs AI. I have a very pessimistic view that if we don't accelerate, we die. If that's your worldview, then we really need to win AI.
China's view is a little bit different. They don't necessarily need AI to win. They've always played this long game. They did it with steel. They've done it with rare-earth minerals. They've done it with solar panels. They've done it with producing phones. They've done it with printed circuit boards. They've done it with so many industries. Incrementally, they're just going to continue to do that, and then they're going to win because they work harder and they're, on average, smarter.
If we don't have super-powerful AI systems, we'll run out of easily accessible nickel, cobalt, oil, and natural gas. We won't be able to make solar panels efficiently and fast enough, and everything will start to get more expensive. The pie will reduce, and we'll also tear each other apart in that way.
If we don't accelerate, we die. China wants to be the world hegemon, and who doesn't want to be the world hegemon? But there are only 2 countries in the world that can legitimately do it and are legitimately trying: the U.S. and China.
The way the Chinese AI ecosystem thinks about this is, “Well, we don't necessarily need to have the biggest compute cluster.” When OpenAI is trying to make a 2-gigawatt data center full of GB200s and GB300s—all these different chips—and those chips are way faster than the chips China can make itself, and China is deploying fewer of them, the girth of compute is huge.
We're kind of doing what China has done historically, which is dumping tons of capital into something, and the market becomes—
Interesting.
And the beneficiary is like, “Oh, if OpenAI has 800 million users today, when they have 3, 4, or 5 billion users across the world, which is possible, of ChatGPT and whatever applications they come up with, then they're on our system, and then they can start to make money.”
It's sort of like YouTube lost money forever, but now it's the platform for watching videos across the world. ChatGPT will be the same thing. There's that aggregation theory. China doesn't necessarily think of it the same way, but they are still incredibly focused on, “Well, we want to be able to make everything ourselves.”
We want to make all of the chips ourselves. We don't actually care that much about making all the chips ourselves. Sure, Trump's doing the tariffs, and sure, we had the CHIPS Act, but those were drops in the bucket compared to how much money China has been putting into the semiconductor ecosystem for the last 10 years.
They've dumped at least $450 billion into this ecosystem through state-owned enterprises, certain tax policies, certain land grants, provincial governments, and the Big Funds, which are almost like government venture funds. They've dumped so much more capital into semiconductors than we have, in an unprofitable way, because they want to build that ecosystem.
Over time, if you take any country in isolation, China is the one that has everything at the highest level on average. Sure, they're 30 years behind on jet engines—or 20 years or 10 years, whatever it is—but they don't need to go outside of China for any of the materials besides raw materials.
Whereas the U.S. needs titanium from here and other materials from there, the same applies to China’s semiconductor ecosystem. Sure, the U.S., Taiwan, and Korea are way ahead, but they also have the accumulated capital base of all the existing equipment and fabs. They need to import from all these different places because it’s a global supply chain.
China is much more concerned today about being insular than being the best at this sort of aggregation theory. But because they’re so talented and have an insular supply chain, yes, they purchase some things from the foreign world. They rent things. They have ByteDance, which I think is the third-largest user of GPUs in the world after OpenAI and probably Meta, although ByteDance may be bigger than Meta.
ByteDance and all the other major Chinese tech companies have these amazing graduates. They have a talent pool. Companies do poach from each other, right? DeepSeek engineers make a lot more than other engineers, but they’re not making $10 million, even though they may be worth it. There’s a really big perception difference.
China could build way faster than us. If they wanted to build a 2-gigawatt or 5-gigawatt data center, they could probably smuggle a lot of chips. It’s not like it’s purely derivative of them wanting to smuggle shitloads of chips, because if they wanted to build a 10-gigawatt data center, I bet they could build it in a few years.
The U.S. is not going to build a single 10-gigawatt data center for a while, right? OpenAI’s total capacity will be around 10 gigawatts in a few years, optimistically. They don’t have the best chips; they’re trying to get better and better and faster and faster. They don’t have the best memory; they’re trying to get better and faster there. They do have the most power. They can build stuff way faster, right?
We’re impressed by how fast Elon does stuff. Elon is slow compared to China, and I think he knows that. That’s probably why he’s the one who’s actually using the Chinese ecosystem more, in terms of the battery facilities being made in China and all these things. He probably recognizes it, too.
There are these major differences in viewpoint and approach because China wants an insular supply chain. They want supply-chain security. We talk about wanting that, but we don’t actually put the money behind it. Where is American capital being allocated? It’s building the biggest data centers and training the best models.
Whereas in China, the capital is being allocated to growing the EV supply chain, growing the semiconductor supply chain, and catching up in all these areas. The U.S., sure, wants to catch up, but actually, we’re just going to give terrible—
Maybe Jensen was right that what you want to own is the end-customer thing.
30. Favorite AI Bears
Yeah. Export production and import, doing the same thing they’ve done forever: prepare at the base level and be behind at the customer side. The value happens close to the customer.
But then you get to the point of, okay, what happens in 3 or 4 years? Even if U.S. AI is amazing, we have the doomsday scenario of China deciding to blockade Taiwan, or even invade it, or create some political instability. People talk about Cambridge Analytica and Russian trolls, whatever. China could do a billion times that into Taiwan, especially with AI and how good AI is now, and somehow subvert it, stage a coup, blockade it, or whatever.
Then we no longer have Taiwan, and the U.S. economy kind of free-falls, right? We can’t make refrigerators without Taiwanese chips. We can’t make cars. We can’t make AI data centers. We can’t grow any of the cloud. We can’t deploy any more SaaS applications. What the hell can we do?
Go back to acquiring all the talent, getting them over here, and doing that.
Right. I think that’s sort of the catch-22 of all this. If you push China too hard, they totally will start swinging. They have the talent. They could go crazy. If we no longer have Taiwan, China could build a way bigger cluster than us. If compute is all that matters, they could do all of these things, and they own the means of production for everything.
There’s this challenging aspect of geopolitical risk. That’s why people don’t want to invest in TSMC. But it’s almost like you can’t invest in Amazon, Apple, Google, or Microsoft if you have geopolitical risk. If you believe Taiwan has risk, then it’s like, YOLO, invest in TSMC. I know a lot of portfolio managers are like, “Oh, you can’t invest in TSMC because of geopolitical risk.” It’s like, “No, dude, you can’t invest in fucking Apple.”
Who is your favorite AI bear? Someone who is far distant from you in their perspective on the direction of this whole thing, but whom you nonetheless like and respect.
There are some AI-researcher gods, likely including Yann LeCun, who are AI bears. I respect them. I like their ideas. I think they’re completely wrong, but—
What’s their argument, if you had to sum it up?
The ways we’re doing this won’t work, right?
LLMs at scale, or—
Right, but it’s like, okay, autoregressive pretraining on the internet doesn’t work to get you to AGI. He’s completely right on that, but then he’ll turn around and be like, “Well, no, no, no, RL systems and all these things are not the right way either,” right? It’s sort of like the no-buts.
I think there are also some investors I know who think this is bullshit, but they’re making tons of money on it anyway. I would respect them to some extent, but with more and more evidence that this stuff is going to get super powerful, it’s hard not to, right? Again, this AI bubble is going to pop because of this podcast, man.
I assure you, it’s just a market return. It’s a coin toss. What startups interest you the most?
One of the startups—the most recent investment I’ve made—is called Periodic Labs. It’s mostly OpenAI people, a Google guy, and a couple of materials scientists.
The area of AI we’ve all been talking about is large-scale web training, RL, all text, all digital God, right? We want to make digital God.
Yeah.
But what would drive a shitload of value for the economy, besides automating the programming of everything, is if we came up with a battery chemistry that was 25% more efficient.
Holy shit. The main constraint against us all having AR glasses and things like that is that batteries are not good enough. There’s the power dissipation, but the battery is terrible, so you have to make all these compromises.
If I could have the processing power of a laptop on my face, we’d be way further ahead. If we all had these super-powerful machines attached to our faces, we could run inference on things and recognize and interact with AI at much higher speed and velocity. That would dramatically improve our productivity.
Things like this are gated by hard tech moving faster. What Periodic Labs is trying to do is take this RL paradigm and apply it to the real world. Test a chemistry for something. Here’s a chemistry, here’s an optimization, here’s something the model spit out—but then you also want to test it in the real world and feed that feedback back into the model.
You do this chain of circles, right? But instead of it being purely digital—which is why RL is really hard, because you need to generate a bunch of responses, test them, and then train the model—the flywheel is so freaking fast.
Yeah.
Right. The flywheel in the physical world is so slow. I need to make a chemistry, I need to try this, I need to test the thing, I need to input it back in, and I need to keep calibrating and keep doing this. It’s so much more expensive and harder to do. But actually, there’s a ton of low-hanging fruit there, I bet.
What about in the hardware world? Just in the pure hardware space, attacking some other interesting bottleneck?
When we talk about where we are in tech, semiconductor manufacturing is super space-age. It’s the most complicated set of tools we make in the world.
That includes tools that cost around $500 million, right? They’re super-amazing feats of engineering. Then the software behind them all is really shitty, right? You could accelerate all that.
But really, in the hardware world, the biggest challenge is that I’m not really a big bull on accelerator companies. I’ve never been.
Yeah, I got it.
Companies competing with NVIDIA, with TPUs, with Trainium, with AMD—I’m not a big bull on those kinds of companies.
Because it’s too hard.
31. Hardware Innovation Beyond Accelerators
It’s just too many things to do. It’s too capital-intensive. There’s not enough of a revolutionary leap, and there are too many predicated things. I wish it could happen, right? It’d be fun. Maybe it does happen, but it would take one hell of a badass thing.
I think there are a lot of individual parts of the supply chain that are not space-age, right? NVIDIA is space-age, yes. It’s the biggest value owner today, but its supply chain has so much old shit. Whether it’s NVIDIA’s supply chain or the hyperscale supply chain, transformers haven’t changed in 50 or 100 years.
There’s a guy building a company in that space.
Solid-state transformers, right? Things like this. There are all sorts of interesting things there. There are so many interesting companies in that space because there’s so much innovation to be done, and there wasn’t that much of a need to innovate before.
Another area is networking between chips. As we extend context length, the memory requirements become bigger and bigger. New memory technologies would be awesome, but DRAM as an industry has so much invested capital equipment and so many existing factories that it’s really hard to attack.
Networking is less constrained, and there are more breakthroughs that can be done in networking. Maybe you don’t have better memory technologies, but you’ve tied the chips closer together so they can use each other’s memory for the problem.
There’s so much more that you can do in the optics space, bridging the gap between electrical connectivity and optical connectivity. NVIDIA created Blackwell, and they had a ton of manufacturing problems and challenges with it in their supply chain. Balance sheets went up for various companies in the supply chain that were building servers and stuff because they were trying to figure it out. AI server and AI data-center deployments were slowed because of these challenges, and there are reliability challenges because these things are connecting to each other at absurd bandwidths.
Every chip in the rack can connect to every other chip in the rack at 1.8 terabytes per second, right? If you think about how much data that is, the amount of bandwidth for connecting these chips together is so high. You can’t fathom what a terabyte a second is. You can’t fathom what a gigabyte a second is.
A gigabyte a second is like a video, right? Or less than a video. Or a megabyte a second, but actually that’s a million bits of information. What’s a kilobyte a second? A byte a second. Okay, you can understand what a byte a second is because that’s 8 bits.
Okay, I’m transmitting 8 bits to you back and forth every second. That’s pretty fast.
That’s what used to exist. Where we are now, there’s still tons of innovation left to be done there.
I think part of the reason Intel is behind is also that data sharing internally was terrible.
And just within the fab, the lithography team doesn’t want to share its data with the etch team, and that data can’t leave the fab and go to an AWS data center to run correlations and all these other things. So you don’t learn from the experiments you do fast enough.
Right now, TSMC isn’t perfect here either. They won’t send their data to a cloud either. But this experiment, analyze-the-data, figure-out-the-new-experiments cycle is slow. How you break that is partially by changing these companies’ culture, which I think Lip-Bu Tan is trying to do.
But also, a lot of it is building better simulators and simulating the world more accurately. World models generally are like, “Hey, I’m going to simulate the world. I’m going to walk around in it.” One common one, I think, is Genie 3, which Google made, right? You can walk around in the world and see cars driving.
But a world model could also simulate molecules, not through classical methods, right? It’s not computational fluid dynamics. It’s the model experiencing this enough, then training a model on physics and feeding that back through, doing it through an AI method instead.
World models can do all sorts of things. You can make a world model to train robots how to pick up cups. You can make a world model that simulates some chemistry in a chemical reaction, or a fire. You can do all sorts of different things.
There are a lot of world-model companies out there. Some of them are really interesting, especially when they’re targeting the physics and reality of the world. Most of the cool innovation is just happening at big companies or already existing companies, right? That’s just the nature of it all. TSMC is doing the most cool innovation, NVIDIA is doing the most cool innovation, and Amphenol is doing cool innovation. All these companies are doing cool innovation.
32. Speed Round: Company Impressions
Could we do a quick speed round where I say a company and you give me a sentence or 2 on your impression of it, just how you feel about it in this moment?
Yeah.
Start with OpenAI.
Oh, yeah. Super awesome.
That’s it. I mean, we’ve talked about them all day. Anthropic.
I’m actually more optimistic on Anthropic than I am on OpenAI.
Why?
Their revenue is accelerating way faster because what they’re focused on is more relevant to that $2 trillion software market, whereas OpenAI is split between, “Yeah, they’re going to do that, but they’re also going to do these other things.” They’re also going to target AI for science, and they’re also going to target AI for the consumer app and do the take-rate thing.
All of these businesses could be amazing, and OpenAI may execute on all of them, but Anthropic is definitely executing on the software side better.
I love them, but they’re pretty mid.
Why do you love them?
When you grow up building computers and liking computers, and AMD is innovating, they’ve always fostered this underdog mentality against Intel and NVIDIA—evil Intel and evil NVIDIA. AMD is the nice company that’s the underdog, and they’ve always had the “Oh, they’re going to take share from them” thesis.
It’s hard not to love them, you know? I know so many people there, and I like all these major hardware companies. There’s not one that I don’t like in terms of the people. But AMD has a soft spot because I think it was my first multibagger as well.
Like, my first multibagger. I can’t own stocks anymore because of compliance. Sorry for the rant, but I fucking love AMD, you know? I also love NVIDIA.
But mid. xAI.
They’re in real danger of not being able to raise capital. Elon’s the best CEO. Of course, everyone’s going to give Elon capital, but the scale of capital required for him to keep up—he can get the next bet, he can get to Colossus 2, right? This megadata center that he’s building will be the largest data center in the world when he builds it: 300,000 Blackwells, 500,000 Blackwells. It’s going to be really great.
But if he doesn’t figure out a business model besides the porn bot, which is what Ani is, and I think he’s monetizing it the wrong way, I think he could monetize it so much better.
How?
You’ve captured the zeitgeist with a cute anime girl that talks to you in a cute voice and will rizz you up. You’ve got users who actually fall for it, and it’s not realistic enough yet, but it will slowly get more realistic.
You’re selling outfits for the same price. You should make it random: “Hey, you have a chance to buy the outfit that is actually her being nude.” Or, “Hey, you have the chance to buy the outfit of her looking like this one anime girl from this one anime.” Or, “Hey, you have the chance to buy this outfit that’s her in a nun suit.”
Obviously, people at xAI hate this, and a lot of them—and many of them—have left. But I think he has to figure out some business model beyond just this, although I think this could be a big business.
He should partner with OnlyFans and make AI manifestations of OnlyFans creators that are Ani, and then subsume the OnlyFans platform into X.
The everything app.
And be like X. You could just Trojan-horse OnlyFans away, because the discovery mechanism for OnlyFans is Instagram and Twitter, as far as I understand.
And you own one of them.
You could partner with the biggest OnlyFans creators to get them over. They don’t have to respond to all the losers; they can also just train a model that acts and looks like them and talks to them.
Anyway, there are all these different monetization methods, and I don't think that's what he should only focus on. To be clear, xAI can get to the next stage of compute. They won't have more compute than OpenAI. They won't have more compute than any individual company at Google, Meta, et cetera, but they will have the biggest individual data center.
What he does with that—and they'll have a very focused team—has to be something really big. Otherwise, they will fall behind in the race, and Elon will not let that happen. He doesn't want that to happen, but he can subsidize and fund this round. He can't go to a 3-gigawatt data center unless he gets capital, which he can't do unless he gets revenue and fundraising.
Oracle is going to make so much fucking money if you believe OpenAI is successful. But if you think OpenAI is going to be successful enough to pay $300 billion to Oracle, how many users do they have, and what's that IP worth? Maybe. There are also reasons you shouldn't own OpenAI, like the Microsoft stuff and the risks around Anthropic and all these things. But in most worlds where Oracle gets paid $300 billion by OpenAI, OpenAI is a $10 trillion or $5 trillion company, or something crazy.
We'll end with the OGs—the old, last-generation best 2 business models. First being Meta.
I think Meta's got the cards to potentially own it all. I don't know if you've seen these new glasses that they came out with, with the screen.
Yeah.
As we go through the history of computing, initially it was punch-card programming. Then it was DOS terminals, right? Then it was, “Oh, you have GUIs, mice, and keyboards.” Then you had touch. The next paradigm in a human-computer interface is that we don't actually have to touch it at all. We tell the AI what we want, and the AI will translate that into reality, right?
Whether it's, “Hey, send an email to this person” or “Send a text to this person,” that's basic stuff that you can already do with Siri or whatever, right? But, “Go buy this”—we're so close to all of these things. The input method into a computer is changing entirely.
The only company in the world that has the full stack—from good hardware, which is what Meta just showed with its glasses with the screen, plus good models, plus the capacity to serve them, plus the knowledge and know-how around recommendation systems to know what content to put in front of the user—is Meta. It's not just generating the content. It's not just interpreting the user's words and taking actions. It's also putting the right content in front of the user. It's all 4 of these things that you need to put in front of the user.
Plus the capital.
Plus the capital. I think Meta is so close to being the only company that can do that.
There are a lot of risks there too, right? So I like Meta a lot. Google to finish it off.
I was pretty bearish on Google about 2 years ago, but I'm super bullish on Google.
Why would that change?
They're waking up on every front. They're taking the TPUs and selling them externally. They're taking their models, and they're actually competitive on them. They're training much better and better and better. They're being aggressive on infrastructure investments.
There's still a lot of dysfunction throughout the company, but they do have the hardware business that they can pivot into this. They won't be as far ahead as Meta is. They won't be as good as Apple is, but they do have Android. They do have YouTube. They have all these IPs. They have search that can come together when we turn to that next consumer interface.
They can also potentially dominate the professional side, too, whereas Meta, I don't think, can dominate that professional side—only the consumer side. I think Google is well-positioned to go capture both markets, or a meaningful share of both.
I feel like we've covered an incredible amount of ground. Is there anything that we haven't talked about that you feel is really critical to what happens in the future that we didn't cover?
33. The Death of Traditional SaaS Business Models
I think the question that everyone constantly asks me is, “Okay, Dylan, you're lucky. Your obsession is that you loved hardware, and you followed it. You followed the supply chain, and you built this business on it. But you really don't follow the software side nearly as much, and all the value is going to get created there, right? When is that flip of the coin going to flip over?”
But I think the thing that most people don't realize is that software is not the same as it was 5 or 10 years ago. You've had dramatic changes in software, and the business model is going to change as well.
If we go back 5 years, 3 years, whatever, when SaaS was the darling—in November 2021, I remember SaaS started tanking. At the time, it was mostly that they were over-earning and all these other things. It doesn't matter. The interesting thing about the business model is that it was such a good business model when your R&D stays flat, right? You grow a little bit, but really R&D doesn't flex that much. Your COGS are super low. The flip side is that in a SaaS business, your customer acquisition cost is quite high.
When you look at what certain companies have done when they've acquired a business, they've just crushed the customer acquisition cost or crushed SaaS. They made the business amazing, whether it's Broadcom with VMware and stuff. It's not really customer acquisition; they just had a bunch of wasted SG&A. But that SG&A—this customer acquisition cost—was most of your cost. R&D was small, but not crazy. Once you hit critical mass, you just cash money.
But software changes a lot when the cost to build that software tanks like crazy. You look at non-U.S. markets, and the prevalence of SaaS is very different. I'll bring up China as an example and a counterpoint. China doesn't have that much of a SaaS business. Actually, their cloud business is pretty small relative to the U.S., despite importing tons of CPUs and storage historically.
Most people just did stuff on-premises and designed their own software because the cost of developing software in China was so much less than in America that the SaaS business model didn't work as well. People could just build rather than rent it out and buy it. That creates inefficiency in the market. I'm sure those weren't always the best-of-breed solutions, but that's what the software development cost may be.
Software developers in China in 2015 were getting paid maybe ⅕ of what they were paid in the U.S., and they were maybe twice as good or something like that. So, 10 times lower cost of software. I'm making up numbers, right? They had 10 times lower cost of software, and so SaaS never happened. Cloud never happened in as big a way as it did in the U.S. and around the world for all the companies that have that same economic reality. And that's despite the outsourcing to India, Eastern Europe, South America, et cetera.
You change all of this with AI software development and AI SaaS products generally—not just AI software development. There are 2 sides to this. AI software development tanks the cost of building a competing software stack. Do you now move to a world where xAI can just build, or I can just build, instead of buying or renting?
Second, if you are a SaaS business and your customer acquisition cost remains the same—and most businesses in AI and in SaaS are going to continue to have a high customer acquisition cost—sales is hard. Breaking into a company is hard. But now you add this AI part of it, and you've added a humongous COGS, right? Your cost of goods sold in any AI software is really hard and really big.
This is partially why I think Google also has an advantage. They have the lowest cost of goods sold for any token of any company because they have their own vertical stack on TPUs.
Coming back to this, because you have this high customer acquisition cost and this high COGS, and because anyone developing it themselves creates competitors in the market, you're going to have a very fragmented SaaS market, or people are just going to build it themselves. Therefore, you never hit the escape velocity where your customer acquisition cost and your R&D get amortized. Because you have such a high COGS, your amortization point means your gross and net profitability are actually much worse. I think the era of software-only businesses is really, really tough in the age of AI.
Already-scaled businesses can do great. I think YouTube is going to have its glory days, and I'm sure it'll always be amazing. But with the cost of generating content falling and falling, creating content, he who controls the platform is going to win and win and win and win.
The functionality you build within Salesforce is actually going to be way less than what you can build on your own, or whatever it is. I'm not saying this is specifically a take on Salesforce itself, but I think many software businesses will have a reckoning with the fact that their COGS are going to soar, their customer acquisition cost isn't going to fall, and they have a lot more competitors. So then they don't hit that escape velocity.
And I think that's the thing that maybe software—it's something I've thought about. There's a couple of people in my company, likely Doug O'Laughlin, and he's the one whose idea this actually is.
This has been incredibly fun. I love learning from you, listening to you, and reading what you put out. I think you're one of the most energetic and awesome thinkers in this whole space right now, so thank you for all the work you've done. When I do these, I ask the same traditional closing question: What's the kindest thing that anyone's ever done for you?
Done for me? It would have to be my brother. Everything he's done in my life. I've been an asshole my whole life, and I still am an asshole. Every time he pulls me back on the path, he corrects me. He loves me unconditionally. My brother has probably done the kindest things for me.
I've been an asshole for so much of my life—unconsiderate and everything. He's just always been there for me and always been—
Why were you an asshole?
Why?
Yeah. If you're aware of it, it makes it into—
No, no. It's terrible. Yeah. Maybe this is the M.O. of who I am, and maybe that's why I'm a good thinker, but I vibe really hard and I'm in the moment really hard. I digest tons of information, but I'm very bad at task orientation and remembering to do specific things. I'm very bad at those things.
Thankfully, I've been able to surround myself throughout my life, whether through birth or not, with people who help me with the things I'm bad at, because I'm very bad at a lot of things. I think of it as a radar plot of how good I am at things. When I'm just vibing and focused on a path, that path can end up hurting someone else—whether it's that I didn't call someone, or I didn't think about their feelings when I did something or said something. That makes me an asshole.
Yes, I should be more conscious of this, and I try to be, but it's just one of the things I'm going to wrestle with in my life forever. A lot of times, I don't even realize I'm being a freaking idiot—
—until my brother's like, "You're a freaking idiot."
God bless your brother.
34. The Kindest Thing
And so, the kindest thing anyone's ever done for me is my brother, throughout my whole life.
I love it. I love it. Wonderful place to close. Thanks so much for your time.
Thank you so much. Yeah.