[BidClub_]
Invest Like the Best · · 86 分钟

GPU、TPU 与 AI 经济学详解|Gavin Baker 访谈

Patrick O'ShaughnessyGavin Baker

YouTube
TL;DR
  • Gemini 3 的真正意义在于:预训练缩放定律依然成立。 这一点被“毫无疑问地”确认,意义在于“地球上没有人知道预训练缩放定律为何、也不知道它如何运作”。Baker 的逆向判断是:考虑到20万张 Hopper 的协同上限,以及 Blackwell 艰难至极的切换,“24年和25年本来真的不该有任何进展”——“推理能力算是救了 AI”,填补了18个月的空档(ARC-AGI 用4年从0提升到8%,随后3个月内从8%升至95%)。缩放定律是乘法关系,因此“Blackwell 模型会非常惊艳”。
  • Google 的低成本 token 优势只是暂时的。 Gemini 3 使用2024-25年一代 TPU v6/v7 训练——放在 Blackwell 旁边,它们就像“F-4 Phantom”,而 Blackwell 是 F-35。等 GB300(可直接兼容 GB200 机架)转向推理,垂直整合的 Blackwell 用户将成为低成本生产者;届时 Google 以负30%利润率“吸干 AI 生态经济氧气”的理性策略会变得痛苦:“这可能开始影响它们的股票”。等 Rubin 落地后,“差距将显著扩大”。
  • ASIC 赛道很可能收缩到 TPU 和 Trainium。 Broadcom 在 TPU 后端拿走50%-55%的毛利率——在约300亿美元的2027年项目中,约150亿美元落入 Broadcom,而其半导体部门全部运营开支约为50亿美元——因此内部自研“绝对不可避免”(MediaTek 是第一记警告枪)。做出一颗好芯片需要3代,而 Nvidia 对每一款内部 ASIC 的回应都是按年更新:“你们不可能跟上我们。”
  • 真正重要的只有4家实验室——OpenAI、Gemini、Anthropic 和 xAI——而且差距正在复利式扩大。 推理能力重新启动了数据飞轮(用户反馈成为可验证奖励),内部 checkpoint 用来训练下一代模型;Meta 的失败说明这有多难,Zuckerberg“错得不能再错”。中国拒绝 Blackwell 会把差距彻底拉开——DeepSeek 在 v3.2 论文中承认了算力短缺;到2026年末左右,中国才会意识到“糟糕,我们确实需要 Blackwell”,但届时稀土杠杆已经被解决。
  • 董事会正在围绕 token 成本重新定位。 OpenAI 是高成本生产者,还要为算力向别人支付利润——“从约1.4万亿美元的模糊乐观情绪,很快就会变成红色警报”;Anthropic“烧钱显著少于 OpenAI,增长却更快”,而其与 Nvidia 签下的50亿美元协议,让 Jensen 在对抗 Google 时从2名选手变成3名。
  • ROI 在“经验上、事实上、毫无歧义地”都是正的。 大型 GPU 买家的经审计 ROIC 已经上升,Q3 更是非科技类 Fortune 500 企业首次量化 AI 增益的季度:C.H. Robinson 现在能在几秒内为100%的入站请求报出卡车运价,而此前只有60%的请求能在15-45分钟内完成,股价上涨约20%。他关注的风险是 Blackwell 在芯片用于训练期间出现 ROI“断层”,因为“训练没有 ROI——ROI 来自推理”。
  • 应用 SaaS 正在重演实体零售业当年错过电商的错误。 它们执着于80%的毛利率,而 AI 原生公司用低于35%-40%的毛利率运行 agent——“这等于保证自己不会在 AI 上成功”。这是一项“生死攸关的决定”,而除了 Microsoft,几乎所有公司都在失败;Salesforce、ServiceNow、HubSpot、GitLab 和 Atlassian 都可以采用激进投资者的打法,单独披露低毛利率的 AI 收入。
  • 边缘 AI 是最可信、也最可怕的熊市情景。 大约3年后,一部更大、更厚的手机可能免费运行剪枝后的 Gemini 5 或 Grok 4.1,速度达到每秒30-60个 token,“这显然是 Apple 的策略”。自2020年以来,滚动泡沫已从电动车和 meme 股票转向核能和量子,但没有任何公开市场标的是这两个领域真正的领导者。未来3-4年最重要的事情是太空数据中心——“在各个方面……都优于地球上的数据中心”。
摘要 · 为研究而整理的核心内容

1. Gemini 3 证明缩放定律仍然成立,而 Blackwell 延迟期间是推理能力“救了 AI”

  • Baker 处理任何新产品发布的第一条规则是:亲自使用,而且要付费。他惊讶于有多少“知名且德高望重的投资者”只用免费版本就得出确定结论——“免费版本就像你在和一个10岁孩子打交道”,却据此推断成年人的表现;每月200美元的版本(Gemini Ultra、SuperGrok)则是“一个完整的、30多岁或35岁的成年人”。其余流程包括:AI 的前沿讨论发生在 X 上,关注全球真正理解 AI 的500-1,000人(“Andrej Karpathy 写的每一篇东西,你至少要读3遍”),以及只要4家重要实验室——OpenAI、Gemini、Anthropic、xAI——中的任何人上播客,就一定要听。
  • Gemini 3“毫无疑问地”确认,预训练缩放定律依然完整——这很重要,因为它并不是一条定律,而是一种经过精确测量、却尚未被理解的经验观察。他的类比是:我们像古埃及人面对太阳,能够把金字塔精准对齐到春分和秋分,却完全不理解轨道力学。“每次得到确认都非常重要。”
  • 他认为公众投资者误解的地方是:“24年和25年本来真的不该有任何进展。”协同运行的 Hopper 不可能超过约20万张,而 Blackwell 是“技术史上迄今为止最复杂的产品切换,远超其他任何一次”——从风冷转向液冷,机架重量从约1,000磅增至约3,000磅,功耗从30kW升至130kW(“想象一下,买一部新 iPhone 还要更换家里所有插座……并加固地板”)。“推理能力算是救了 AI”——RLVR 和测试时计算填补了约18个月的空档:ARC-AGI 用4年从0提升到8%,而第一款 OpenAI 推理模型出现后,3个月内就从8%升至95%。
  • 关键在于,缩放定律是乘法关系;Gemini 3 使用2024-25年一代 TPU v6/v7 训练。按他的战斗机分类,Hopper 是 P-51 Mustang,这些 TPU 是 F-4 Phantom,Blackwell 则是 F-35。把两条新的缩放定律应用到 Blackwell 训练的基础模型上,“Blackwell 模型会非常惊艳”。

2. Blackwell 改写低成本 token 竞争,Google 的算盘也随之改变

  • AI 是“我作为科技投资者整个职业生涯里,第一次看到低成本生产者真正重要”——Apple、Microsoft 和 Nvidia 的市值达到万亿美元级别,并不是因为它们便宜。作为 token 的低成本生产者,Google 一直在“吸干 AI 生态的经济氧气”;当竞争对手需要外部融资、而 Google 不需要时,以负30%利润率运行 AI“显然是最理性的决定”。
  • 第一批 Blackwell 训练的模型会在2026年初出现,而且第一个将来自 xAI:按照 Jensen 公开说过的话,“没有人比 Elon 更快地建设数据中心”;一款新芯片还需要6-9个月调优,才能仅仅超过上一代。实际上,xAI 会替所有其他人先把 Blackwell 调试好。
  • 随后局面翻转:GB300 可直接兼容 GB200 机架,“不需要新的电力墙”;运行 GB300 的公司,尤其是垂直整合的公司,将成为 token 的低成本生产者。Google 一旦失去这个头衔,负利润率就会变得痛苦——“这可能开始影响它们的股票”;等 Rubin 到来后,“相对于 TPU 和其他所有 ASIC,差距将显著扩大”。
  • 董事会开始从成本角度重新定位。OpenAI 为算力向其他公司支付利润——“也许替他们运行算力的人,并不擅长运行 GPU”——这解释了 Stargate,也解释了为什么“从约1.4万亿美元的模糊乐观情绪,很快就会变成红色警报”。Anthropic“烧钱显著少于 OpenAI,增长却更快”;它与 Nvidia 签下50亿美元协议,Dario 对 Blackwell 和 Rubin 相对于 TPU 的理解,让 Nvidia 在对抗 Google 时从2名选手变成3名。OpenRouter 的数据也说明了这一点:xAI 处理了1.35万亿 tokens,Google 约8000亿-9000亿,Anthropic 约7000亿。

3. ASIC 洗牌:Broadcom 的50%-55%毛利率撕开了故事的裂缝

  • Google 负责 TPU 前端,Broadcom 则负责后端,并以50%-55%的毛利率管理 TSMC。一个约300亿美元的2027年 TPU 项目,意味着约150亿美元收入流向 Broadcom,而其整个半导体部门的运营开支约为50亿美元。“Google 可以把 Broadcom 半导体部门的每个人都挖走,给他们翻倍薪酬,还能多省下50亿美元。”把 MediaTek 拉进来、接受低得多的台湾 ASIC 利润率,是“第一记警告枪”;一款真正出色的 SerDes 是基础设施核心,其价值大致在每年100亿-250亿美元区间。
  • “做出一颗好芯片至少需要3代。”TPU 直到 v3/v4 才“稍微具备竞争力”。Amazon 拥有“任何半导体公司里最好的 ASIC 团队”(Graviton、Nitro/SuperNIC),但仍然要等到 Trainium 3 才算“可以”。“如果最终除了 Trainium 和 TPU 之外还存在很多 ASIC,我会很惊讶”;而这两者最终都会转向客户自有工具——“不管这些公司怎么说……从经济学上看,这是绝对不可避免的”。
  • 与此同时,Lisa Su 和 Jensen 对每一款客户自研 ASIC 的回应都是按年更新:“我们只会继续加速……你们不可能跟上我们。”他调侃那些跃跃欲试者:“哇,你们做了自己的加速器。那 NIC 怎么办?扩展交换机呢?光学呢?……糟糕,我做出了这么一颗小芯片。我以为这很容易。”
  • Nvidia 自己无法独立维持这样的节奏——一台 Blackwell 机架有数千个零部件,而 Nvidia 只生产其中约200-300种。这也是半导体 VC 复兴具有基础性意义的原因:创始人平均约50岁,而这一轮复兴几乎完全由 Nvidia 的市值点燃,整个生态必须一起提速。“过去7年里,我这家小机构做的半导体交易,可能比前10大 VC 加起来还多。”

4. 推理能力重新启动数据飞轮,Meta 的失败说明这有多难

  • 他过去常引用 Eric Vishria [可能是] 的话:没有独特数据和互联网级分发能力的基础模型,是历史上贬值最快的资产。但“推理能力从根本上改变了这一点”,而且改变得非常深刻:当大量用户持续喜欢或不喜欢某个回答时,这就是一种可验证奖励,可以反馈给模型。曾让 Netflix、Amazon、Meta 和 Google 成为规模回报递增型企业的 Bezos 飞轮,已经开始在前沿实验室转动。“现在还早……但你已经能看到它开始转了。”
  • 证明这很难的实证案例是:Zuckerberg 今年1月说 Meta 会在2025年的某个时候拥有最好的 AI——“我不知道他是否排在前100名……他错得不能再错”。Microsoft 也失败了(Inflection),Amazon 的 Nova 也不在前20(Adept)。原因之一是运行 GPU 时存在巨大差异:如果你的集群正常运行率是30%,而竞争对手是90%,“你们根本不在同一个竞争里”;另一个原因是“品味”,即判断哪些实验值得投入5万张 GPU、运行数天的直觉。他用零售业打比方:把1,000家门店都经营得干净、明亮、库存充足,由“不会偷东西的友善员工”负责运营,你就能成为一家200亿-300亿美元的公司——历史上大约只有15家公司做到过。
  • 复利的下一层是:每个顶尖实验室都会运行一个更先进的内部 checkpoint,并用它训练下一代模型;没有这个 checkpoint,“追赶会变得非常困难”。这也是为什么“中国开源对 Meta 来说是上帝的礼物”:它提供了启动模型的 checkpoint。
  • 中国正在浪费这张牌:就在 Blackwell 落地之际,强迫本国实验室使用国产芯片。DeepSeek 的 v3.2 论文以“非常政治正确、但仍有一点风险”的措辞承认,它没有足够算力追上美国实验室。差距会被彻底拉开;到2026年末左右,中国才会意识到“糟糕,我们确实需要 Blackwell”。届时稀土问题会“比任何人想象得更快”解决——DARPA 正在推进酶精炼项目,友好国家也有矿藏;“它们显然没那么稀有,只是名字起错了”。如果 Blackwell 重新回到中国市场,“中国开源将卷土重来”,而他认为这是好事。

5. ROI 在“经验上、事实上、毫无歧义地”为正,Q3 是关键验证

  • 他对这场争论感到恼火:最大的 GPU 买家都是上市公司,有经过审计的季度财务数据,而它们的 ROIC 已经高于加大投入之前。原因一部分是运营开支节省,更大一部分是把广告推荐系统从 CPU 转移到 GPU,从而加快收入增长。“但那又怎样?ROI 一直都在。”每家大型互联网公司内部,收入负责人都会和研究人员争抢芯片:“这是一个非常线性的方程。你给我更多 GPU,我就能带来更多收入。”而对于 Blackwell——“Rubin 肯定如此”——经济学最终会压过推动资本支出的囚徒困境,其中一部分动力来自对 ASI 的宗教式信仰:“他们几乎都想永生。”
  • Q3 的里程碑是:非科技类 Fortune 500 企业首次给出量化 AI 增益。货运经纪商 C.H. Robinson 现在可以在几秒内为100%的入站请求报出价格和运力;此前只有60%的请求能在15-45分钟内完成。财报发布后,公司股价上涨约20%。
  • AI 的采用节奏与云计算类似:到2013年第一次 re:Invent 时,每家初创公司都已经运行在 AWS 上;约5年后,Fortune 500 才完成标准化。VC 比公开市场投资者更乐观,因为他们能直接看到 AI 原生公司的每名员工收入提升(ICONIQ 的图表,以及 a16z 的 David George 提到的“model busters”);年轻创始人“因为一直在和 AI 对话,所以更快变得成熟、表达更专业”。
  • 他原本担心出现 Blackwell ROI“断层”:资本开支高得难以想象,而芯片还在进行训练;“训练没有 ROI……ROI 来自推理”。Meta 已经公布过一个 ROIC 下降的季度,并因此受到股价惩罚。C.H. Robinson 这样的季度说明,这个断层是可以跨越的。顺带一提,VC 运营的 AI 控股公司不会比收购基金更擅长执行——“你不可能在私募股权擅长的游戏里击败私募股权”——但 PE 会系统性地应用 AI。

6. 从智能转向有用性:让他最担心的熊市情景是边缘 AI

  • 事件路径是:GB300(而且可能 MI450 的作用大于 MI355)会压低单 token 成本,让模型能够思考更长时间。Gemini 3 帮他订了一家餐厅,“这是它第一次真正替我做成一件事”;从订餐厅到订酒店、订航班、叫 Uber,“突然之间,你就有了一个助手”。在技术领先的大公司里,“超过50%的客户支持已经由 AI 完成”,这是一个规模达4,000亿美元的行业;而 AI 擅长说服,也就是销售:一家公司的3项职能——制造、销售、支持——到2026年末,AI“可能会做好其中2项”。
  • Karpathy 提出的组织原则是:软件领域,任何能够被明确描述的事情都可以自动化;“AI 领域,任何能够被验证的事情都可以自动化”。账目是否平衡?销售是否完成转化?“这就像 AlphaGo……最重要的职能之所以重要,正是因为它们可以被验证。”
  • 对非专家而言,进步正变得不可见——在付费版本里,你必须拿 PCIe 和 Ethernet 的扩展网络问题去测试,才能看出模型差异;不过在他的慈善梦幻橄榄球队阵容测试中,“这些新模型确实好得多”。因此,“我们需要从追求更高智能转向追求更强有用性”;然后,有用性还必须交棒给科学突破。有用意味着持续稳定地完成任务,同时保留全部上下文。他的旅行规划测试包括朝东阳台、Huberman 式晨间阳光、飞机上的 Starlink,以及全家人的偏好;METR 式任务长度必须不断延长。
  • 除了缩放定律失效,最可怕的熊市情景就是边缘 AI。大约3年后,一部更大、更厚的手机可能免费运行剪枝后的 Gemini 5 或 Grok 4.1,速度达到每秒30-60个 token。“这显然是 Apple 的策略”:成为 AI 的隐私安全分发者,只有在需要时才调用云端的“神级模型”。如果设备端约115 IQ 的能力已经足够,“我认为这就是一个熊市情景”——也是“最可信、最可怕的熊市情景”。

7. 太空数据中心:“在各个方面都更优越”

  • 他的判断非常明确:“未来3到4年,世界上最重要的事情就是太空数据中心”——这会对所有在地球上建设电厂或数据中心的人产生“非常深远的影响”。从第一性原理看,数据中心就是电力、冷却和芯片。在轨卫星每天24小时处于阳光下,光照强度比地面高30%,太阳辐照度达到6倍,而且不需要电池,直接消灭一项巨大的成本。地面机架重量中很大一部分来自冷却,而在太空中冷却是免费的:“只要把散热器放在卫星的背阴面……那里可以接近绝对零度。”
  • 网络也会改善:“唯一比激光穿过光纤更快的东西,就是激光穿过绝对真空。”这能形成比地球上更快、更一致的集群。太空训练会更慢,“只是因为它规模太大”;但对于推理,Starlink 已经展示了直连手机能力,手机→卫星→答案,跳过整个城市级光纤往返。阻力在于发射:“我们需要很多 Starship。”
  • 这几条线正在汇合:Elon 昨天表示,Tesla、SpaceX 和 xAI 正在融合——xAI 成为 Optimus 的智能模块,Tesla Vision 负责感知,SpaceX 的轨道数据中心为一切提供动力——“每一个都在为另一个创造竞争优势”。

8. 省电、供给过剩,以及转向核能与量子的滚动泡沫

  • “过剩跟随短缺”的铁律会不会适用于这里?AI 与软件不同:每一次使用都会消耗算力。每家实验室都可以消化多10倍的算力(Mark Chen 已公开说过);200美元的版本会变得更好,免费版本则会变成今天的200美元版本。Google 在 AI Mode 中加入广告,“会给所有其他人许可”,让他们也在免费版本中放广告;此外还有 agent 佣金——“这里有3个假期方案,你要不要我替你订一个?”
  • 两个自然的制约因素正在发挥作用:电力,以及 TSMC 的谨慎。TSMC——就是那群见过 Sam Altman、笑着说他是“播客哥”的人——正在犯一个错误:扩产不足,最终会填满 Intel 空置的晶圆厂;Lip-Bu Tan [可能是] 正在收获 Patrick Gelsinger [可能是] 的战略成果,而解雇 Gelsinger 是“可耻的”。第三个需要观察的制约因素是:自90年代末以来的第一个真正 DRAM 周期。当时一片晶圆的价值“像一颗5克拉钻石”,价格按倍数而不是百分比上涨——“那是完全不同的游戏”。他的结论是:制约因素是好事,“更平滑、更持久就是好事”。
  • 在功率受到约束时,“你为算力支付的价格无关紧要”。每瓦多产生3-5倍 token,就意味着多产生3-5倍收入;他的例子是,一座500亿美元的数据中心产生250亿美元收入,而一座350亿美元的 ASIC 数据中心只能产生80亿美元收入。因此,“无论价格如何,最好的产品都会赢”,并拥有“疯狂的定价权”。核电建设速度不够快——“一只蚂蚁”都可能拖延一座电厂,而它原本只需要被挪走——所以答案是天然气和太阳能;这也是 Abilene 的逻辑,Caterpillar 刚刚表示未来几年将把产能提高75%。
  • 自2020年以来,滚动泡沫依次从非 Tesla 电动车初创公司(下跌99%)转向 meme 股票,再转向核能、聚变、小型模块化反应堆和量子计算:这些都是具有变革性叙事的主题,但“所有公开市场投资方式都很可能不会成功,也没有真正的基本面支撑”。他认为真正的量子领导者是 Google、IBM 和 Honeywell 的量子部门;量子霸权只意味着某些计算是经典计算机做不到的——“不意味着量子计算会接管世界”。他还有一个令人不安的宏观观察:“过去2年里,AI 无论需要什么来继续增长和进步,它都能得到。”核能的舆论一夜之间翻转,太空也在地面电力受限时打开空间——Kevin Kelly 所说的 technium 正在成为现实。

9. SaaS 正在重演实体零售业当年错过电商的错误

  • 他的框架是:应用 SaaS 公司“正在犯下实体零售商当年犯过的完全相同的错误”——它们看到了客户对电商的需求,却讨厌电商的利润结构,因此没有投资;如今 Amazon 北美零售业务的利润率已经超过许多大众市场零售商。“如果有一种具有根本变革意义、且客户正在要求的新技术,不拥抱它永远是错误。”
  • 机制在于:SaaS 只需写一次、以低成本分发,毛利率可以达到70%-90%;AI 则每次都要重新计算答案,因此一家优秀的 AI 公司毛利率约为40%。但它们比 SaaS 更早产生现金,不是因为毛利率高,而是因为“几乎没有人类员工”。为了保护80%毛利率而定价的 agent“永远不会成功……如果你试图保住80%的毛利率结构,就等于保证自己不会在 AI 上成功。绝对保证”。
  • 投资者已经证明,他们能够容忍这种转变:云计算迁移期间,Adobe 从本地部署迁移出去时,收入和利润率都崩溃;Microsoft 早期股价表现艰难,随后收购 GitHub,把它作为 Copilot 的分发渠道,并以低得多的毛利率结构打造出“一项巨大的业务”。所以,“这是一个生死攸关的决定,而除了 Microsoft,几乎所有人都在失败——它们的平台正在烧钱”,这是对那份[可能是 Nokia]备忘录的有意呼应。
  • 他认为激进投资者——“或者建设性激进投资者”——应该迫使 Salesforce、ServiceNow、HubSpot、GitLab 和 Atlassian 采用这样的打法:单独披露 AI 收入和 AI 毛利率(“你知道那是真 AI,因为它的毛利率很低”),甚至暂时以零毛利率运行;同时由原本能够产生现金的核心业务提供支持,这是创业竞争者没有的优势。另一种结局很糟糕:“现在,另一个由别人制造的 agent 正在访问你的系统……把数据拉进它自己的系统,而你最终会被关掉。”

10. “投资就是寻找隐藏的真相”

  • 被问到会如何向 Patrick 的孩子推介投资这份职业时,他说:“投资就是寻找真相——如果你最先发现了真相,并且判断它确实是真相,这就是你创造 alpha 的方式……你在寻找隐藏的真相。”他的公式是:尽可能深入地掌握历史,再与对当前事件最准确的判断交叉,形成差异化观点;在“可以想象的最伟大的技能与运气游戏”中,寻找 Perry Mutual 系统里被错误定价的股票。这个过程始于大二时在投行实习、负责邮寄研究报告;到第3天,他已经开始狼吞虎咽地阅读报告,随后读 Peter Lynch、把 Buffett 的股东信读了2遍,并自学会计,最终放弃原本成为滑雪闲人、河流向导和小说家的既定人生规划。
  • 他最后的坦白很有分量:每支体育队伍选人时他都是最后一个,被选去参加私人滑雪课花了“一笔小财富”却始终没变好,也从没赢过公园里的国际象棋玩家——“这是我唯一一件勉强算有竞争力的事情。我希望自己还能擅长别的事,但我就是不擅长。”而在 Alta 清洁厕所的经历,“永久改变了我对待他人的方式”。
Patrick O'Shaughnessy

I’d love to talk about how you, in the nitty-gritty process, approach new things that come out in this whole AI world, because it’s happening so constantly. I’m extremely interested in it, and I find it very hard to keep up. I have a couple of blogs that I go read and friends that I call, but maybe let’s take Gemini 3 as a recent example.

When that comes out, what—literally, take me into your office—what are you doing? How do you and your team process an update like that, given how often these things are happening?

Gavin Baker

I think the first thing is, you have to use it yourself. I would say I’m amazed at how many famous and august investors are reaching really definitive conclusions about AI based on the free tier.

The free tier is like you’re dealing with a 10-year-old, and you’re making conclusions about the 10-year-old’s capabilities as an adult. You could just pay, and I do think you need to pay for the highest tier, whether it’s Gemini Ultra, SuperGrok, whatever it is. You have to pay the $200-per-month tier, whereas those are like a fully fledged 30- or 35-year-old. It’s really hard to extrapolate from an 8- or 10-year-old to the 35-year-old, and yet a lot of people are doing that.

The second thing is, there was an Insider post about OpenAI, and they said, to a large degree, OpenAI runs on Twitter vibes. I just think AI happens on X. There have been some really memorable moments, like there was a giant fight between the PyTorch team at Meta and the JAX team at Google on X, and the leaders of each lab had to step in publicly and say, “No one from my lab is allowed to say bad things about the other lab. I respect them, and that is the end of that.”

The companies are all commenting on each other’s posts. The research papers come out. If on planet Earth there are 500 to 1,000 people who really understand this and are at the cutting edge of it, and a good number of them live in China, I think you have to follow those people closely. I think there is incredible signal. Everything in AI is just downstream of those people.

Patrick O'Shaughnessy

Everything Andrej Karpathy writes, you have to read 3 times, minimum.

Gavin Baker

Yeah. He’s incredible. I would say anytime somebody from one of the 4 labs that matter—OpenAI, Google, Anthropic, and xAI, which are clearly the 4 leading labs—goes on a podcast, it’s so important to listen.

For me, one of the best use cases of AI is to keep up with all of this. Just listen to a podcast, and then, if there are parts that I thought were interesting, talk about them with AI. I think it’s really important to have as little friction as possible. I’ll bring it up. I can either press this button and pull up Grok, or I have this.

Patrick O'Shaughnessy

Oh, wow. I don’t touch that. That just brings it right up.

Gavin Baker

Yeah, it brings it right up. What do you think of Patrick O'Shaughnessy?

“Patrick O'Shaughnessy is one of my favorite voices in investing. His Invest Like the Best podcast is straight fire. He does deep dives with folks like Bill Gurley—”

Patrick O'Shaughnessy

Gurley?

Gavin Baker

Yes.

Patrick O'Shaughnessy

It’s so—can you believe we have this?

Gavin Baker

I know. It’s like we have—I think somebody said on X—we’ve imbued these rocks with crazy spells, and now we can summon superintelligent genies on our phones over the air. It’s crazy.

1. The Blackwell Transition

Patrick O'Shaughnessy

Crazy. So when something like Gemini 3 comes out, the public interpretation was, “Oh, this is interesting. It seems to say something about scaling laws and the pretraining stuff.” What’s your frame on the state of progress in frontier models in general? What are you watching most closely?

Gavin Baker

I do think Gemini 3 was very important because it showed us that scaling laws for pretraining are intact. They stated that unequivocally, and that’s important because no one on planet Earth knows how or why scaling laws for pretraining work.

It’s actually not a law. It’s an empirical observation, and it’s an empirical observation that we’ve measured extremely precisely and that has held for a long time. Our understanding of scaling laws for pretraining—and maybe this is a little bit controversial with 20% of researchers, but probably not more than that—is kind of like the ancient British people’s understanding of the sun or the ancient Egyptians’ understanding of the sun.

They could measure it so precisely that the east-west axes of the Great Pyramids are perfectly aligned with the equinoxes, and so are the east-west axes of Stonehenge. Perfect measurement. But they didn’t understand orbital mechanics. They had no idea how or why it rose in the east, set in the west, and moved across the horizon.

Patrick O'Shaughnessy

The aliens.

Gavin Baker

Yeah, our god in a chariot. It’s really important every time we get a confirmation of that. Gemini 3 was very important in that way.

I think there’s been a big misunderstanding, maybe in the public-equity investing community or the broader, more generalist community, based on the scaling laws of pretraining. There really should have been no progress in 2024 and 2025.

The reason for that is, after xAI figured out how to get 200,000 Hoppers coherent, you had to wait for the next generation of chips. You really can’t get more than 200,000 Hoppers coherent. Coherent just means you could think of it as each GPU knowing what every other GPU is thinking. They’re sharing memory, they’re connected, they scale up networks and scale out, and they have to be coherent during the pretraining process.

Everything in AI is a struggle between Google and NVIDIA. Google has a TPU, and NVIDIA has its GPUs. Google only has a TPU, and it uses a bunch of other chips for networking. NVIDIA has the full stack, and Blackwell was delayed.

Blackwell was NVIDIA’s next-generation chip, and the first iteration of that was the Blackwell B200. A lot of different SKUs were canceled, and the reason for that is it was by far the most complex product transition we’ve ever gone through in technology.

Going from Hopper to Blackwell, first you go from air-cooled to liquid-cooled. The rack goes from weighing, in round numbers, 1,000 lb to 3,000 lb. It goes from, in round numbers, 30 kilowatts—which is 30 American homes—to 130 kilowatts, which is 130 American homes.

I analogize it to imagining if, to get a new iPhone, you had to change all the outlets in your house to 220 volts, put in a Tesla Powerwall, put in a generator, put in solar panels—that’s the power—put in a whole-home humidification system, and then reinforce the floor because the floor can’t handle this. It was a huge product transition, and the rack was so dense that it was really hard for them to get the heat out.

Blackwells have only really started to be deployed in scaled deployments over the last 3 or 4 months. Had reasoning not come along, there would have been no AI progress from mid-2024 through essentially Gemini 3. There would have been none. Everything would have stalled, and can you imagine what that would have meant to the markets? For sure, we would have lived in a very different environment.

Reasoning bridged this 18-month gap. Reasoning saved AI because it let AI make progress without Blackwell or the next generation of TPUs, which were necessary for the scaling laws for pretraining to continue.

The reason we’ve had all this progress is that we have these 2 new scaling laws of post-training, which is just reinforcement learning with verifiable rewards. Verifiable is such an important concept in AI. One of Karpathy’s great things was, “With software, anything you can specify, you can automate.”

With AI, anything you can verify, you can automate. It's such an important concept, and I think an important distinction. Then there's test-time compute. All the progress we've had—and we've had immense progress since October 24 through today—was based entirely on these 2 new scaling laws.

Gemini 3 was arguably the first test since Hopper came out of the scaling law for pretraining, and it held. That's great because all these scaling laws are multiplicative. So now we're going to apply these 2 new scaling laws—reinforcement learning with verifiable rewards and test-time compute—to much better base models. Google came out with the TPU v6 in 2024 and the TPU v7 in 2025.

And in semiconductor time, imagine Hopper is like a World War II-era airplane. It was by far the best World War II-era airplane. It's the P-51 Mustang with the Merlin engine.

Patrick O'Shaughnessy

And 2 years later, in semiconductor time, that's like—you know, you're an F-4 Phantom. Okay. [Laughter]

Gavin Baker

Because Blackwell was such a complicated product and so hard to ramp, Google was training Gemini 3 on 2024- and 2025-era TPUs, which are like F-4 Phantoms. Blackwell is like an F-35.

Patrick O'Shaughnessy

It just took a really long time to get it going.

Gavin Baker

So I think Google for sure has this temporary advantage right now. From a pretraining perspective, I think it's also important that they've been the lowest-cost producer of tokens. And this is really important because AI is the first time in my career as a tech investor that being the low-cost producer has ever mattered.

Apple is not worth trillions because they're the low-cost producer of phones. Microsoft is not worth trillions because they're the low-cost producer of software. NVIDIA is not worth trillions because they're the low-cost producer of AI accelerators. It's never mattered.

Patrick O'Shaughnessy

Okay.

Gavin Baker

And this is really important because what Google has been doing as the low-cost producer is sucking the economic oxygen out of the AI ecosystem, which is an extremely rational strategy for them and for anyone who's a low-cost producer. Let's make life really hard for our competitors.

So what happens now, I think, has pretty profound implications. One, we will see the first models trained on Blackwell in early 2026.

Patrick O'Shaughnessy

Yeah.

Gavin Baker

I think the first Blackwell model will come from xAI. The reason for that is that, according to Jensen, no one builds data centers faster than Elon. Jensen has said this on the record.

Even once you have the Blackwells, it takes 6 to 9 months to get them performing at the level of Hopper because Hopper is finally tuned. Everybody knows how to use it. The software is perfect for it. Engineers know all its quirks. Everybody knows how to architect a Hopper data center at this point.

By the way, when Hopper came out, it took 6 to 12 months for it to really outperform A100, which was the generation before. So if you're Jensen or NVIDIA, you need to get as many GPUs deployed in 1 data center as fast as possible, in a coherent cluster, so you can work out the bugs.

This is what xAI effectively does for NVIDIA because they build the data centers the fastest. They can deploy Blackwells that scale the fastest, and they can help work with NVIDIA to work out the bugs for everyone else. So because they're the fastest, they'll have the first Blackwell model.

We know that the scaling laws for pretraining are intact, and this means the Blackwell models are going to be amazing. Blackwell is not an F-35 versus an F-4 Phantom, but from my perspective it is a better chip. Maybe it's like an F-35 versus a Rafale. [Laughter] Now that we know pretraining scaling is holding, we know that these Blackwell models are going to be really good.

Patrick O'Shaughnessy

And based on the raw specs, they should probably be better.

Gavin Baker

Then something even more important happens. So the GB200 was really, really hard to get going. The GB300 is a great chip. It is drop-in compatible in every way with those GB200 racks. Now, you're not going to replace the GB200s. No new power walls.

Patrick O'Shaughnessy

Yeah. Just any data center that can handle those. You can slot in the GB300s, and now everybody's good at making those racks. You know how to get the heat out. You know how to cool them.

Gavin Baker

You're going to put those GB300s in, and then the companies that use the GB300s are going to be the low-cost producers of tokens, particularly if you're vertically integrated. If you're paying a margin to someone else to make those tokens, you're probably not going to be.

I think this has pretty profound implications because I think it has to change Google's strategic calculus. If you have a decisive cost advantage and you're Google, and you have search and all these other businesses, why not run AI at a negative 30% margin?

It is by far the rational decision. You take the economic oxygen out of the environment. You eventually make it hard for your competitors, who need funding unlike you, to raise the capital they need. Then, on the other side of that, maybe you have an extremely dominant share position.

Well, all that calculus changes once Google is no longer the low-cost producer, which I think will be the case. The Blackwells are now being used for training. Then, when that model is trained, you start shifting Blackwell clusters over to inference, and then all these cost calculations and these dynamics change.

Patrick O'Shaughnessy

I do think it's very interesting—the strategic and economic calculations between the players. I've never seen anything like it. Everyone understands their position on the board, what the prize is, and what play their opponents are running. It's really interesting to watch.

Gavin Baker

So I just think if Google changes its behavior, because it's going to be really painful for them as a higher-cost producer to run that negative 30% margin, it might start to impact their stock. That has pretty profound implications for the economics of AI.

And then when Reuben comes out, we'll know—the gap is going to expand significantly versus TPUs.

Patrick O'Shaughnessy

Versus TPUs and all other ASICs.

Gavin Baker

Now, I think Trainium 3 is probably going to be pretty good. Trainium 4 is going to be good.

Patrick O'Shaughnessy

Why is that the case? Why won't TPU v8 or v9 be every bit as good?

Gavin Baker

A couple of things. For whatever reason, Google made more conservative design decisions. I think part of that is that, in round numbers, there's a front end and back end of semiconductor design, and then there's dealing with TSMC.

You can make an ASIC in a lot of ways. What Google does is mostly the front end for the TPU, and then Broadcom does the back end and manages TSMC. It's a crude analogy, but the front end is like the architect of a house. They design the house. The back end is the person who builds the house, and managing TSMC is like stamping out that house, like Lennar or D.R. Horton.

For doing those 2 latter parts, Broadcom has a 50% to 55% gross margin. We don't know what the gross margin is on TPUs. Let's say in 2027, TPU consensus estimates are maybe somewhere around $30 billion. Again, who knows?

Patrick O'Shaughnessy

Yeah, yeah, yeah, but $30 billion, I think, is a reasonable estimate.

Gavin Baker

At 50% to 55% gross margins, Google is paying Broadcom $15 billion.

Patrick O'Shaughnessy

Okay, that's a lot of money.

Gavin Baker

At a certain point, it makes sense to bring a semiconductor program entirely in-house. In other words, Apple does not have an ASIC partner for their chips. They do the front end and back end themselves, and they manage TSMC. The reason is they don't want to pay that 50% margin.

At a certain point, it becomes rational to renegotiate this. Just as perspective, the entire operating expense of Broadcom's semiconductor division is, in round numbers, $5 billion. So it would be economically rational. If Google is paying them $15 billion on $30 billion, Google can go to every person who works in Broadcom's semiconductor division, double their compensation, and make an extra $5 billion.

In 2028, let's just say it does $50 billion. Now it's $25 billion. You could triple their compensation. And by the way, you don't need them all.

Patrick O'Shaughnessy

Yeah.

Gavin Baker

Of course, they're not going to do that because of competitive concerns. But with TPU v8 and v9, all of this is beginning to have an impact because Google is bringing in MediaTek.

This is maybe the first way you send a warning shot to Broadcom: “We're really not happy about all this money we're paying.” They did bring MediaTek in, and the Taiwanese ASIC companies have much lower gross margins. So this is kind of the first shot across the bow.

Then there's all this stuff people say: “Oh, but Broadcom has the best SerDes.” Broadcom has really good SerDes, and SerDes is extremely foundational technology because it's how the chips communicate with each other. You have to serialize and deserialize.

But there are other good SerDes providers in the world. A really good SerDes is not a moat. Maybe it's worth $10 billion or $15 billion a year, but it's probably worth about $25 billion a year.

Because of that friction, and because of what I think are conservative design choices on the part of Google, maybe the reason they made those conservative design choices is because they were going to a bifurcated supply. TPUs are slowing down. I would say GPUs are accelerating.

This is the first competitive response of Lisa and Jensen to everybody saying, “We're going to have our own ASIC.” It's, “Hey, we're just going to accelerate. We're going to do a GPU every year, and you cannot keep up with us.” And then I think what everybody is learning is like, “Oh, wow, that's so cool.”

You made your own accelerator as an ASIC. Wow, what’s the NIC going to be? What’s the CPU going to be? What’s the scale-up switch going to be? What’s the scale-up protocol? What’s the scale-out switch?

What kind of optics are you going to use? What’s the software that’s going to make all this work together? Then it’s like, “Oh, shit, I made this tiny little chip.” Whether it’s admitted or not, I’m sure the GPU makers don’t love it when their customers make ASICs and try to compete with them. And then it’s like, “Whoops, what did I do? I thought this was easy.”

Patrick O'Shaughnessy

How do you know?

Gavin Baker

It also takes at least 3 generations to make a good chip. TPU v1—I mean, it was an achievement that they made it.

Patrick O'Shaughnessy

Yeah.

Gavin Baker

It really wasn’t until TPU v3 or v4 that the TPU started to become even vaguely competitive.

Patrick O'Shaughnessy

Is that just a classic learning-by-doing thing?

Gavin Baker

100%.

Patrick O'Shaughnessy

Yeah.

Gavin Baker

Even if you’ve made the first one, from my perspective, the best ASIC team at any semiconductor company is actually the Amazon ASIC team. They were the first ones to make the Graviton CPU. They have Nitro. It was the first—it’s called a SuperNIC. They’ve been extremely innovative, really clever.

With Trainium and Inferentia 1, they may be a little better than TPU v1, but only a little. Trainium 2 is a little better. Trainium 3, I think, is the first time it’s like, “Okay.” Then I think Trainium 4 will probably be good. I’ll be surprised if there are a lot of ASICs other than Trainium and TPU.

And by the way, Trainium and TPU will both run on customer-owned tooling at some point. We can debate when that will happen, but the economics of success that I just described mean it’s inevitable. No matter what the companies say, the economics and reasoning from first principles make it absolutely inevitable.

Patrick O'Shaughnessy

If I were to zoom all the way out on this stuff—because sometimes I find these details unbelievably interesting, and it’s like the grandest game that’s ever been—what is the global human dividend of all this crazy development? We keep making the loss lower on these pretraining scaling laws, but who cares? It’s been a while since I’ve asked this thing something where I wasn’t blown away by the answer, for me personally.

What are the next couple of things that all this crazy infrastructure war allows us to unlock because it’s so successful? If I were to posit an event path, I think the Blackwell models are going to be amazing. The dramatic reduction in per-token cost enabled by the GB300, and probably more by the MI450 than the MI355, will lead to these models being allowed to think for much longer, which means they’re going to be able to do new things.

I was very impressed that Gemini 3 made me a restaurant reservation. It’s the first time it’s done something for me, other than research something and teach me stuff. If you can make a restaurant reservation, you’re not that far from being able to make a hotel reservation and an airplane reservation and order me an Uber. All of a sudden, you’ve got an assistant.

2. The Prisoner's Dilemma

Gavin Baker

Yeah. You can just imagine—it’s on your phone. I think that’s pretty near-term. Some big companies that are very tech-forward already have 50% or more of customer support being done by AI, and that’s a $400 billion industry. What AI is great at is persuasion. That’s sales and customer support.

If you think about the functions of a company, they’re to make stuff, sell stuff, and then support the customers. Right now, maybe in late 2026, you’re going to be pretty good at 2 of them. I do think it’s going to have a big impact on media. I think robotics, which we talked about last time, are finally going to start to be real.

There’s an explosion in exciting robotics startups. I do still think that the main battle is going to be between Tesla’s Optimus and the Chinese, because it’s easy to make prototypes and hard to mass-produce them. But then it goes back to what Andrej Karpathy said about AI: it can automate anything that can be verified. Any function where there’s a right or wrong answer, a right or wrong outcome, can use reinforcement learning to make the AI really good at that.

Patrick O'Shaughnessy

Yeah. What are your favorite examples of that so far, or theoretically?

Gavin Baker

Does the model balance? They’ll be really good at making models. Do all the books globally reconcile? They’ll be really good at accounting because it’s double-entry bookkeeping. It has to balance. There’s a verifiable “You got it right or wrong.”

Did you make the sale or not? That’s very clear. It’s just like AlphaGo: did you win or did you lose? Did the customer convert or not? Did the customer ask for an escalation during customer support or not? Its most important functions are important because they can be verified, right?

So I think if all of this starts to happen, there’ll be an ROI on Blackwell, and then all this will continue. Then we’ll have Rubin, and that’ll be another big quantum of spend: Rubin, the MI450, and TPU v7. I think the most interesting question is, what are the economic returns to artificial superintelligence?

All of these companies in this great game have been in a prisoner’s dilemma. They’re terrified that if they slow down, they’re just gone forever. Their competitors don’t slow down, and it’s an existential risk.

Patrick O'Shaughnessy

Yeah.

Gavin Baker

Microsoft blinked for about 6 weeks earlier this year.

Patrick O'Shaughnessy

Yeah.

Gavin Baker

I think they would say they regret that.

Patrick O'Shaughnessy

Yeah.

Gavin Baker

But with Blackwell, and for sure with Reuben, economics are going to dominate the prisoner’s dilemma from a decision-making and spending perspective, just because the numbers are so big. This goes to the ROI-on-AI question, and the ROI on AI has empirically, factually, and unambiguously been positive.

I always find it strange that there’s any debate about this, because the largest bidders on GPUs are public companies. They report something called audited quarterly financials, and you can use those things to calculate something called a return on invested capital. If you do that calculation, the ROIC of the big public spenders on GPUs is higher than it was before they ramped spending.

You could say, “Well, part of that is opex savings.” At some level, that is part of what you expect the ROI to be from AI. Then you say, “Well, a lot of it is actually just applying GPUs”—moving the big recommender systems that power advertising and recommendation systems from CPUs to GPUs—and you’ve had massive efficiency gains. That’s why all the revenue growth at these companies has accelerated.

But so what? The ROI has been there. Every big internet company has people responsible for the revenue who are intensely annoyed at the amount of GPUs being given to the researchers. It’s a very linear equation: “If you give me more GPUs, I will drive more revenue. Give me those GPUs, we’ll have more revenue and more gross profit, and then we can spend money.” It’s this constant fight at every company.

3. The Bear Case: Edge AI

One of the factors in the prisoner’s dilemma is that everybody has this religious belief that we’re going to get to ASI. At the end of the day, what do they all want? Almost all of them want to live forever. They think that ASI is going to help them do that.

Patrick O'Shaughnessy

Right. Good return.

Gavin Baker

That’s a good return. [Laughter]

Patrick O'Shaughnessy

But we don’t know. And if, as humans, we have pushed the boundaries of physics, biology, and chemistry—the natural laws that govern the universe—I’m very curious about your favorite sort of throw-cold-water-on-this-stuff takes that you think about sometimes.

One would be the things that would cause this demand for compute to change, or even cause its trajectory to change. I’m curious what you think those things would be.

Gavin Baker

There’s 1 really obvious bear case, and it’s just edge AI, and it’s connected to the economic returns to ASI. In 3 years, on a bigger and bulkier phone to fit the amount of DRAM necessary—and the battery probably won’t last as long—you’ll probably be able to run a pruned-down version of something like Gemini 5, Grok 4, Grok 4.1, or ChatGPT at, I don’t know, 30 to 60 tokens per second.

Then that’s free. This is clearly Apple’s strategy: “We’re going to be a distributor of AI, and we’re going to make it privacy-safe and run it on the phone.” Then you can call one of the big models—the god models—in the cloud whenever you have a question.

If 30 to 60 tokens a second at a 115 IQ is good enough, I think that’s a bear case.

Patrick O'Shaughnessy

Other than just the scaling laws breaking?

Gavin Baker

But if we assume scaling laws continue, we now know they’re going to continue for pre-training for at least 1 more generation. We’re very early in the 2 new scaling laws for post-training, mid-training, RLVR—whatever people want to call it—and then test-time compute inference.

We’re so early in those, and we’re getting so much better at helping the models hold more and more context in their minds as they do this test-time compute. That’s really powerful, because everybody’s like, “Well, how’s the model going to know this?” Eventually, if you can hold enough context, you can just hold every Slack message and Outlook message and company manual in a company in your context.

Patrick O'Shaughnessy

Yeah.

Gavin Baker

And then you can compute the new task and compare it with your knowledge of the world—what you think, what the model thinks, all this context. It may be that really, really long context windows are the solution to a lot of the current limitations. That's enabled by some of these cool tricks, like KV-cache offload and stuff.

But I do think that, other than scaling laws slowing down and other than there being low economic returns to ASI, edge AI is by far the most plausible and scariest bear case.

Patrick O'Shaughnessy

I like to visualize different S-curves. I invested through the iPhone, and I love to see the visual of the iPhone models as it sort of went from this clunky, bricky thing up to what we have now. Obviously, we've sort of petered out on its form factor.

If you picture something similar for the frontier models themselves, does it feel like it's at a certain part of that natural technology-paradigm progression to you? If you're paying for Gemini Ultra or SuperGrok and you're getting the good AI, it's hard to see differences.

I have to go really deep on something like, “Do you think PCI Express or Ethernet is a better protocol for scale-up networking, and why? Show me the scientific papers.” If you shift between models and you ask a question like that, where you know it really deeply, you know the answers. Then you see differences.

Gavin Baker

I do play fantasy football. Winnings are donated to charity, but these new models are quite a bit better at helping with, “Who should I play?” They think in much more sophisticated ways.

By the way, if you're a historically good fantasy football player and you're having a bad season, this is why: you're not using it. I think we'll see that in more and more domains.

But I do think they are already at a level where, unless you are a true expert or just have an intellect that is beyond mine, it's hard to see the progress. That's why I do think we need to shift from getting more intelligent to more useful, unless more intelligence starts leading to these massive scientific breakthroughs and we're curing cancer in '26 and '27.

I don't know that we're going to be curing cancer, but I do think from an ROI S-curve, we need to hand off from intelligence to usefulness. Then usefulness will have to hand off to scientific breakthrough; that creates whole new industries.

Patrick O'Shaughnessy

What are the building blocks of usefulness in your mind?

Gavin Baker

Just being able to do things consistently and reliably. A lot of that is keeping all the context. There's a lot of context if someone wants to plan a trip for me. I've acquired these strange preferences. I follow that guy Andrew Huberman, so I like to have an east-facing balcony so I can get morning sun.

The AI has to remember that being on a plane with Starlink is important to me. Here are the resorts I've historically liked. Here are the kinds of areas I've liked. Here are the rooms that I would really like at each. That's a lot of context, and keeping all of that and weighting it is a hard problem.

So I think context windows are a big part of it. There's this METR task-evaluation thing—how long it can work, how long it can work for. You could think of that being related in some way to context, although not precisely.

Task length needs to keep expanding, because booking a restaurant is economically useful, but it's not that economically useful. Booking me an entire vacation and knowing the preferences of my parents, my sister, my niece, and my nephew, and what that means—that's a much harder problem. That's something that a human might spend 3 or 4 hours optimizing. If you can do that, that's amazing.

But then again, I just think it has to be good at sales and customer support relatively soon. I do think we're going to see an acceleration in the awesomeness of various products, just because engineers are using AI to make products better and faster.

Patrick O'Shaughnessy

We both invested in Fora, the hearing aid company, which is just absolutely remarkable. I think it's something I never would have thought of.

Gavin Baker

And we're going to see, I think, something like that in every vertical. That's AI being used for the most core function—

Patrick O'Shaughnessy

Yeah.

Gavin Baker

—of any company, which is designing the product. There are already lots of examples of AI being used to help manufacture the product and distribute it more efficiently, whether it's optimizing a supply chain or having a vision system watch a production line. I think a lot of stuff is happening.

The other thing I think is really interesting in this whole ROI part is that Fortune 500 companies are always the last to adopt a new technology. They're conservative; they have lots of regulations and lots of lawyers. Startups are always the first.

So let's think about the cloud, which was the last truly transformative new technology for enterprises: being able to have all of your compute in the cloud and use SaaS. It's always upgraded, it's always great, et cetera, and you can get it on every device. Those were dark days before the cloud.

The first AWS re:Invent, I think it was in 2013. Every startup on planet Earth ran on the cloud. The idea that you would buy your own server and storage box and router was ridiculous. That probably happened even earlier; it had probably already happened before the first re:Invent. The first big Fortune 500 companies started to standardize on it maybe 5 years later.

You see that with AI. I'm sure you've seen this in your startups, and I think one reason VCs are more broadly bullish on AI than public-market investors is that VCs see very real productivity gains. There are all these charts that, for a given level of revenue, a company today has significantly fewer employees than a company of 2 years ago.

The reason is AI is doing a lot of the sales, the support, and helping to make the product. ICONIQ has some charts. a16z, by the way—David George is a good friend, great guy—has his Model Busters thing. There's very clear data that this is happening. So people who have a lens into the world of venture see this.

I do think it was very important in the 3rd quarter. This is the first quarter where we had Fortune 500 companies outside of the tech industry give specific quantitative examples of AI-driven uplift. C.H. Robinson went up something like 20% on earnings. Should I tell people what C.H. Robinson does?

Patrick O'Shaughnessy

Let's just say a truck goes from Chicago to Denver. The trucker lives in Chicago, so it's going to go back from Denver to Chicago. There's an empty load. C.H. Robinson has all these relationships with truckers and trucking companies, and they match shippers' demand with that empty-load supply to make trucking more efficient.

They're a freight forwarder. There are actually lots of companies like this, but they're the biggest and most dominant. One of the most important things they do is quote price and availability.

Somebody, a customer, calls them up and says, “Hey, I urgently need 3 18-wheelers from Chicago to Denver.” In the past, they said it would take them 15 to 45 minutes, and they only quoted 60% of inbound requests. With AI, they're quoting 100% and doing it in seconds.

4. Meta, Open Source, and Model Depreciation

Gavin Baker

And so they printed a great quarter, and the stock went up 20%. It was because of AI-driven productivity that's impacting the revenue line, the cost line—everything.

I was actually very worried about the idea that we might have this Blackwell ROI air gap because we're spending so much money on Blackwell. Those Blackwells are being used for training, and there's no ROI on training. Training is making the model; the ROI comes from inference.

I was really worried that maybe we'd have this 3-quarter period where capex is unimaginably high, those Blackwells are only being used for training, margins are staying flat, and AI is going up.

Patrick O'Shaughnessy

Those Blackwells are only being used for training, margins are staying flat, and AI is going up.

Gavin Baker

Yeah, exactly. And so ROIC goes down. You can see that with Meta. Because Meta has not been able to make a frontier model, Meta printed a quarter where ROIC declined, and that was not good for the stock. I was really worried about this.

I do think those data points are important in suggesting that maybe we'll be able to navigate this potential air gap in ROIC.

Patrick O'Shaughnessy

Yeah, it makes me wonder about this market. I'm like everybody else: it's the 10 companies at the top that have all the market cap and get more than all of the attention. There are 490 other companies in the 500. You studied those too.

What do you think about that group? What's interesting to you about the group that now nobody seems to talk about and no one really seems to care about because they haven't driven returns and they're a smaller percentage of the overall market?

Gavin Baker

Well, I think people are going to start to care if more and more companies print these C.H. Robinson-like quarters. Companies that have historically been really well run—the reason they have a long track record of success is that you cannot succeed without using technology well.

So if you have a kind of internal culture of experimentation and innovation, I think you will do well with AI. I'd bet on the best investment banks to be earlier and better adopters of AI than maybe some of the trailing banks. Sometimes the past is prologue.

One thing that I have a strong opinion on: all these VCs are setting up these holding companies, and we're going to use AI to make traditional businesses better. They're really smart VCs, and they have great track records.

Patrick O'Shaughnessy

But that's what private equity has been doing for 50 years. You're just not going to beat private equity at their game. What Vista did in the early days, right?

Gavin Baker

Yeah. Private equity's maybe had a little bit of a tough run. Multiples have gone up, so private assets are more expensive. The cost of financing has gone up, and it's tough to take a company public because the public valuation is 30% lower than the private valuation.

Private equity's had a tough run. I actually think these private equity firms are going to be pretty good at systematically applying AI. We haven't spent much time talking about Meta, Anthropic, or OpenAI, and I'd love just your impression of everything that's going on on this infrastructure side that we talked about.

Patrick O'Shaughnessy

These are 3 really important players in this grand battle, this grand game. How does all of this development that we've discussed so far impact those players specifically, do you think?

Gavin Baker

The first thing, let me just say about frontier models broadly.

Patrick O'Shaughnessy

Yeah.

Gavin Baker

In 2023 and 2024, I was fond of quoting Eric Vishria, our friend, a brilliant man. Eric would always say, “Foundation models are the fastest-depreciating assets in history.” I would say he was 90% right. I modified the statement: “Foundation models without unique data and internet-scale distribution are the fastest-depreciating assets in history.” But reasoning fundamentally changed that in a really profound way.

There was a loop, a flywheel, to quote Jeff Bezos, that was at the heart of every great internet company. You made a good product, you got users, and those users using the product generated data that could be fed back into the product to make it better. That flywheel has been spinning at Netflix, Amazon, Meta, and Google for over a decade, and that's an incredibly powerful flywheel.

That's why those internet businesses were so tough to compete with. It's why they had increasing returns to scale. Everybody talks about network effects much more, and network effects were important for social networks. I don't know to what extent Meta is a social network anymore. It's more like content distribution, but they just had increasing returns to scale because of that flywheel.

That dynamic was not present in the pre-reasoning world of AI. You pre-trained a model, let it out in the world, and it was what it was. It was actually pretty hard. They would do RLHF—reinforcement learning with human feedback—and try to make the base model better. Maybe you'd get a sense from Twitter vibes that people didn't like this, and so you'd tweak it. There were the little up and down arrows, but it was actually pretty hard to feed that back into the model.

With reasoning, it's early, but that flywheel has started to spin, and that is really profound for these frontier labs.

Patrick O'Shaughnessy

So, 1, reasoning fundamentally changed the industry dynamics of frontier labs. Just explain why specifically that is. What is going on?

Gavin Baker

Because if a lot of people are asking a similar question and they're consistently either liking or not liking the answer, then you can use that as a verifiable reward. That's a good outcome. Then you can feed those good answers back into the model. We're very early in this flywheel spinning—it's hard to do now, but you can see it beginning to spin.

5. Geopolitics and Rare Earths

This is important fact number 1 for all of those dynamics. Second, I think it's really important that Meta—Mark Zuckerberg, at the beginning of this year, in January, said, “I anticipate—I’m highly confident”—I’m going to get the quote wrong—“that at some point in 2025, we're going to have the best and most performant AI.” I don't know if he's in the top hundred.

Patrick O'Shaughnessy

Okay.

Gavin Baker

So he was as wrong as it was possible to be. I think that is a really important fact because it suggests that what these 4 companies have done is really hard to do. Meta threw a lot of money at it, and they failed. Yann LeCun had to leave. They had the famous billion-dollar offers for AI researchers.

By the way, Microsoft also failed. They did not make such an unequivocal prediction, but they bought Inflection AI. There were a lot of comments from them that they anticipated their internal models quickly getting better and that they were going to run more and more of Copilot on their internal models. Amazon bought a company called Adept, and they have models called Nova. No, I don't think they're in the top 20.

So clearly, it's much harder to do than people thought a year ago. There are many, many reasons for that. It's actually really hard to keep a big cluster of GPUs coherent. A lot of these companies were used to running their infrastructure to optimize for cost rather than performance.

Patrick O'Shaughnessy

Yeah.

Gavin Baker

And if you're running—because of the laws of physics, maybe the most anybody can get is 200,000 or 300,000 Blackwells coherent. We'll see. But if you have 30% uptime on that cluster and you're competing with somebody who has 90% uptime, you're not even competing.

So, 1, there's a huge spectrum in how well people run GPUs. 2, then I think there are these AI researchers. They like to talk about taste, and I find it very funny. “Why do you make so much money?” “I have very good taste.” What taste means is that you have a good intuitive sense for the experiments to perform.

This is why you pay people a lot of money, because it actually turns out that as these models get bigger, you can no longer run an experiment on a 1,000-GPU cluster and replicate it on 100,000 GPUs. You need to run that experiment on 50,000 GPUs, and maybe it takes days. So there's a very high opportunity cost, and you have to have a really good team that can make the right decisions about which experiments to run on this.

Then you need to do all the reinforcement learning during post-training well, and the test-time compute well. It's complicated. It's really hard to do, and everybody thinks it's easy, but all those things are really hard.

I used to have this saying—I was a retail analyst long ago: “Pick any vertical in America. If you can just run 1,000 stores and have them clean, well-lit, stocked with relevant goods at good prices, and staffed by friendly employees who are not stealing from you, you're going to be a $20 billion company, a $30 billion company.” Like 15 companies have been able to do that. It's really hard.

It's the same thing. Doing all of these things well, and then reasoning with this flywheel, is beginning to create more separation.

What's even more important is that every one of those labs—xAI, Gemini, OpenAI, and Anthropic—has a more advanced checkpoint internally of the model. A checkpoint is just—you’re continuously working on these models, and then you release a checkpoint. The reason these models get better is that the one they're using internally is better, and they're using that model to train the next model.

If you do not have that latest checkpoint, you're behind. It's getting really hard to catch up. Chinese open source is a gift from God to Meta because you can use Chinese open source to try things, and that can be your checkpoint. You can use that as a way to bootstrap this, and I'm sure that's what they're trying to do, and everybody else.

The big, giant swing factor, I think, is that China has made a terrible mistake with this rare-earths thing. China has the Huawei Ascend, and it's a decent chip. Versus something like the deprecated Hopper, it looks okay, so they're trying to force Chinese open source to use their Chinese chips, their domestically designed chips.

The problem is that Blackwell is going to come out now, and the gap between these American frontier labs and Chinese open source is going to blow out because of Blackwell. Actually, DeepSeek, in its most recent technical paper, DeepSeek-V3.2, said one of the reasons it struggles to compete with the American frontier labs is that it doesn't have enough compute.

That was their very politically correct, still a little bit risky way of saying, because China said, “We don't want the Blackwells,” right? And they're saying, “Guys, that might be a big mistake. That might be a big mistake.”

If you just play this out, these 4 American labs are going to start to widen their gap versus Chinese open source, which then makes it harder for anyone else to catch up because that gap is growing. So you can't use Chinese open source to bootstrap.

Geopolitically, China thought it had the leverage. It's going to realize, “Oh, whoopsie-daisy. We do need the Blackwells.” Unfortunately for them, they'll probably realize that in late 2026. At that point, there's an enormous effort underway. DARPA has—there are all sorts of really cool DARPA and DoD programs to incentivize really clever technological solutions for rare earths, like using enzymes to refine them.

There are all sorts of really cool things happening. There are a lot of rare-earth deposits in countries that are very friendly to America, that don't mind actually refining them in the traditional way.

So, I think rare earths are going to be solved way faster than anyone thinks. They’re obviously not that rare; they’re just misnamed. They’re rare because they’re really messy to refine.

Geopolitically, I actually think Blackwell is pretty significant, and it’s going to give America a lot of leverage as this gap widens. Going back to the dynamics between these companies, xAI will be out with the first Blackwell model, and then they’ll probably be the first ones using Blackwell for inference at scale. I think that’s an important moment for them.

By the way, it is funny: if you go on OpenRouter, you can just look. It’s whatever it is—1% of API tokens—but it’s an indication. They process 1.35 trillion tokens. Google did 800 or 900 billion over the last 7 days or the last month, and Anthropic was at 700 billion. xAI is doing really, really well, and the model is fantastic. I highly recommend it. You’ll see xAI come out with this, and OpenAI will come out faster.

OpenAI’s issue that they’re trying to solve with Stargate is that they pay a margin to people for compute. Maybe the people who run their compute aren’t the best at running GPUs. They are a high-cost producer of tokens, and I think this explains a lot of their code red recently.

Patrick O’Shaughnessy

Yeah. Well, just the $1.4 trillion in spending commitments.

And I think that was just, “Hey, they know they’re going to need to raise a lot of money,” particularly if Google keeps its current strategy of sucking the economic oxygen out of the room. You go from $1.4 trillion to rough vibes, to code red, pretty fast. The reason they have a code red is because of all these dynamics.

6. Data Centers in Space

So then they’ll come out with a model, but they will not have fixed their per-token cost disadvantage relative to both xAI and Google, and almost Anthropic at that point. Anthropic is a good company. They’re burning dramatically less cash than OpenAI and growing faster, so I think you have to give Anthropic a lot of credit. A lot of that is their relationship with Google and Amazon for the TPUs and the Trainiums.

You can look at Dario and Jensen. Maybe there have been a few public comments made between them.

Patrick O'Shaughnessy

Jousting.

Gavin Baker

A little bit of jousting. Well, Anthropic just signed a $5 billion deal with NVIDIA.

That is because Dario is a smart man, and he understands these dynamics about Blackwell and Reuben relative to TPU. NVIDIA now goes from having 2 of the fighters—xAI and OpenAI—to 3 fighters. That helps in this NVIDIA-versus-Google battle.

If Meta can catch up, that’s really important. I’m sure NVIDIA is doing whatever they can to help Meta: “Let us know you’re running those GPUs this way. Maybe we should twist the screw this way or turn the dial that way.” If Blackwell comes back to China, which it seems like it probably will, that will also be very good, because then Chinese open source will be back.

Patrick O'Shaughnessy

What are the other poles of things? One pole would be the other breakthroughs that you have your mind on in the data center that aren’t chips, that we’ve talked about before, as one example.

Gavin Baker

I think the most important thing that’s going to happen in the world in the next 3 to 4 years is data centers in space. This has really profound implications for everyone building a power plant or a data center on Planet Earth. There is a giant gold rush into this.

Patrick O'Shaughnessy

I haven’t heard anything about this, so please explain.

Gavin Baker

Everybody thinks AI is risky, but they say, “I’m going to build a data center. I’m going to build a power plant that’s going to power a data center. We will need that.” But if you think about it from first principles, data centers should be in space.

What are the fundamental inputs to running a data center? There’s power, there’s cooling, and then there are the chips.

Patrick O'Shaughnessy

That’s the total, if you think about it from a total-cost perspective.

Gavin Baker

Yeah.

Patrick O'Shaughnessy

And just the inputs to making the tokens come out of the magic machines.

Gavin Baker

Yeah. So in space, you can keep a satellite in the sun 24 hours a day, and the sun is 30% more intense. You can keep it in the sun constantly. If the sun’s here, you can have the satellite always catching the light.

Patrick O'Shaughnessy

Catching the light.

Gavin Baker

Catching the light. The sun is 30% more intense, and this results in 6 times more irradiance in outer space than on Planet Earth. So you’re getting a lot of solar energy. Point 1.

Point 2: because you’re in the sun 24 hours a day, you don’t need a battery, and this is a giant percentage of the cost. The lowest-cost energy available in our solar system is solar energy in space.

For cooling in one of these racks, a majority of the mass and the weight is cooling. The cooling in these data centers is incredibly complicated—the HVAC, the CDUs, the liquid cooling.

Patrick O'Shaughnessy

It’s fucking gold.

Gavin Baker

And it’s as close to absolute zero as you can get. So all that goes away, and that is a vast amount of cost.

Let’s think about how these might work. Maybe each satellite is a rack. That’s one way to think of it. Maybe some people make bigger satellites that are 3 racks. How are you going to connect those racks? In a data center, the racks are connected over a certain distance with fiber optics. That just means a laser going through a cable. The only thing faster than a laser going through a fiber-optic cable is a laser going through absolute vacuum.

If you can link these satellites in space together using lasers, you actually have a faster and more coherent network than in a data center on Earth. For training, that’s going to take a long time—

Patrick O'Shaughnessy

Yeah, just because it’s so big.

Gavin Baker

But for inference—and I think even training will eventually happen—let’s think about the user experience. When I asked Grok about you and it gave the nice answer, a radio wave traveled from my cell phone to a cell tower. Then it hit the base station, went into a fiber-optic cable, and went to some sort of metro aggregation facility in New York, probably within 10 blocks of here. There’s a small little metro router that routed those packets to a big xAI data center somewhere.

The computation was done, and it came back over the same path. If the satellites can communicate directly with the phone, and Starlink has demonstrated direct-to-cell capability, you just go boom, boom. It’s a much better, lower-cost user experience. So in every way, data centers in space, from a first-principles perspective, are superior to data centers on Earth.

7. Power Constraints as a Governor

Patrick O'Shaughnessy

So if we could teleport that into existence, I understand that portion. What are the frictions to that? Why will that not happen? Is it launch cost? Is it launch availability?

Gavin Baker

We need a lot of SpaceX Starships. The Starships are the only ones that can economically make that happen.

Patrick O'Shaughnessy

We need a lot of those Starships.

Gavin Baker

Maybe China or Russia will be able to land a rocket. Blue Origin just landed a booster. It’s an entirely new and different way to think about SpaceX.

It’s interesting that Elon posted yesterday, or said in an interview, that Tesla, SpaceX, and xAI were converging. They really are. xAI will be the intelligence module for Optimus, made by Tesla. Tesla Vision has its perception system, and SpaceX will have the data centers in space that will power a lot of the AI, presumably, for xAI, Tesla, Optimus, and a lot of other companies.

It’s interesting the way that they’re converging, and each one is creating competitive advantage for the other. If you’re xAI, it’s really nice that you have this built-in relationship with Optimus. Now Tesla is a public company, so I cannot imagine the level of vetting that will go into that intercompany agreement. Then you have a big advantage with these data centers in space.

It’s also nice if you’re xAI that you have 2 companies with a lot of customers who you can use to help build your customer-support agents and your customer-sales agents, with built-in customers. They really are all converging in a neat way. I do think it’s going to be a big moment when that first Blackwell model comes out from xAI next year.

Patrick O'Shaughnessy

Hmm. If I go to the other end of the spectrum and I think about something that seems to have been historically endemic to the human economic experience—that shortages are always followed by gluts in capital cycles—what if, in this case, the shortage is compute?

Mark Chen is now on the record as saying they would consume 10 times as much compute if you gave it to them in a couple of weeks. There seems to still be a massive shortage of compute, which is all the stuff we’ve talked about today. But there also just seems to be this iron law of history that gluts follow shortages.

What do you think about that concept as it relates to this technology? Could this technology be a glut?

Gavin Baker

Yeah. You know, AI is fundamentally different from software in that every time you use AI, it takes compute in a way that traditional software just did not. I think it is true that every one of these companies could consume 10x more compute. What would happen is that the $200 tier would get a lot better, and the free tier would get like the $200 tier. Google has started to monetize AI Mode with ads.

Patrick O'Shaughnessy

And I think that will give everyone else permission to introduce ads into the free mode, and then that is going to be an important source of ROI. It seems like OpenAI is tailor-made to—

Gavin Baker

Yeah, absolutely. All of them. And actions, like, “Hey—”

Patrick O'Shaughnessy

Here are your 3 vacations. Would you like me to book one? And then they're for sure going to collect a commission.

Gavin Baker

Yeah. You know, there are many ways you can make money. I think we went into great detail on maybe a prior podcast about how inventory dynamics made these inventory cycles inevitable in semis. The iron law of semis is just that customer buffer inventories have to equal lead times, and that's why you got these inventory cycles historically.

We haven't seen a true capacity cycle in semis, maybe arguably since the late 1990s. And that's because Taiwan Semi has been so good at aggregating and smoothing supply. A big problem in the world right now is that Taiwan Semi is not expanding capacity as fast as its customers want. I think this is actually a pretty big mistake they're in the process of making, just because you do have Intel with these fabs, and they're not as good, and it's really hard to work with their PDK.

But now you have this guy Lip-Bu Tan, who's a really good executive and really understands that business. By the way, I think Pat Gelsinger was also a good executive, and he put Intel on the only strategy that could result in success. I actually think it's shameful that the Intel board fired him when it did. But Lip-Bu Tan is a good executive, and now he's reaping the benefits of Pat's strategy. Intel has all these empty fabs, and eventually, given the shortages we have of compute, those fabs are going to be filled.

Patrick O'Shaughnessy

So I think Taiwan Semi is in the process of making a mistake, but they're just so paranoid about an overbuild.

Gavin Baker

Yeah.

Patrick O'Shaughnessy

And they're so skeptical. You know, they're the guys who met with Sam Altman and laughed and said, “He's a podcast bro. He has no idea what he's talking about.” They're terrified of an overbuild. So it may be that Taiwan Semi, single-handedly, because they're cautious, is the brakes on the bubble—the governor.

Gavin Baker

And I think governors are good. It's good that power is a governor; it's good that Taiwan Semi is a governor. If Taiwan Semi opens up at the same time that data centers in space relieve all power constraints—but that's, I don't know, 5 or 6 years away, when data centers in space are a majority of deployed megawatts—yeah, I think you get an overbuild really fast. But we have these 2 really powerful natural governors, and I think that's good. Smoother and longer is good.

Patrick O'Shaughnessy

We haven't talked about power, other than alluding to it through the space thing. We haven't talked about power very much. Power was the most uninteresting topic because there was demand, and nothing really changed for a really, really long time. All of a sudden, we're trying to figure out how to get gigawatts here, there, and everywhere. How do you think about it? Are you interested in power?

Gavin Baker

I'm very interested. I do feel lucky that, in a prior life, I was the sector leader for the telecom and utilities team.

Patrick O'Shaughnessy

Okay.

Gavin Baker

I do have some base level of knowledge. One, having watts as a constraint is really good for the most advanced compute players because if watts are the constraint, the price you pay for compute is irrelevant. The TCO of your compute is absolutely irrelevant because if you could get 3x, 4x, or 5x more tokens per watt, that is literally 3x, 4x, or 5x more revenue.

If you're going to build an advanced data center, an advanced data center costs $50 billion. A data center with your ASIC maybe costs $35 billion, but if that $50 billion data center pumps out $25 billion of revenue and your ASIC data center at $35 billion is only pumping out $8 billion, well, you're pretty bummed. It's good for all of the most advanced technologies in the data center, which is exciting to me as an investor.

As long as power is a governor, the best products are going to win irrespective of price and have crazy pricing power. I think that's the first implication that's really important to me. Second, the only solutions to this are— we just can't build nuclear fast enough in America. As much as we would love to build nuclear quickly, we just can't. We just can't.

Patrick O'Shaughnessy

Yeah.

Gavin Baker

It's just too hard. NEPA, all these rules—it's just too hard. A rare ant that we could move, and it could be in a better environment, can totally delay the construction of a nuclear power plant. One ant. It's crazy, actually. Humans need to come first. We need to have a human-centric view of the world.

The solutions are natural gas and solar. The great thing about these AI data centers is that, apart from the ones that you're going to do inference on, you can locate them anywhere. I think we're going to see all this activity in Abilene, because it's in the middle of a big natural gas basin, and we have a lot of natural gas in America because of fracking. I think we're going to have a lot of natural gas for a long time. We ramp production really fast.

I think this is going to be solved. You're going to have power plants fed by gas or solar. I think that's the solution. And already, all these turbine manufacturers were reluctant to expand capacity. Caterpillar just said, “We're going to increase capacity by 75% over the next few years.” The system on the power side is beginning to respond.

Patrick O'Shaughnessy

One of the reasons that I always love talking to you is that you do as much in the top 10 companies in the world as you do looking at brand-new companies with entrepreneurs who are 25 years old trying to do something amazing. And so you have this very broad sense of what's going on.

If I think about that second category of young, enterprising technologists who are now the first generation of AI-native entrepreneurs, what are you seeing in that group that's notable, surprising, or interesting?

Gavin Baker

These young CEOs are just so impressive in all ways, and they get more polished faster. I think the reason is that they're talking to the AI. “How should I deal with pitching this investor? I'm meeting with Patrick O'Shaughnessy. What do you think the best ways I should pitch him are?” And it works.

“Do deep research.” And it's good. “Hey, I have this difficult HR situation. How would you handle it?” That's correct. And it's good at that. “We're struggling to sell our product. What changes would you make?” It's really good at all of that today.

And that goes to these VCs seeing massive AI productivity in all their companies. It's because their companies are full of these 23-, 24-, or even younger AI natives. I've been so impressed with young investment talent, and it's just part of it. Your podcast is part of that. Knowledge, and very specific knowledge, has become so accessible through podcasts and the internet.

Impressive young people come in, and they're just—I feel like they're where I was as an investor in my early 30s, and they're 22. I'm like, “Oh my God, I have to run so fast to keep up.” These kids who are growing up AI-native are just proficient with it in a way that I am trying really hard to become.

Patrick O'Shaughnessy

Can we talk about semiconductor VC specifically and what is interesting in that universe?

Gavin Baker

The one thing I think is so cool about it and so underappreciated is that your average semiconductor venture founder is about 50 years old.

Patrick O'Shaughnessy

Okay.

Gavin Baker

Jensen Huang, and what's happened with NVIDIA and the market cap of NVIDIA, have single-handedly ignited semiconductor venture. But the way it's ignited, it's ignited in an awesome way that's really good for NVIDIA and Google and everyone.

Let's just say you were the best DSP architect in the world and had done that for the last 20 years, every 2 years, because that's what you have to do in semiconductors. It's like every 2 years, you have to win—run a race. If you won the last race, you start a foot ahead, and over time those compound and make each race easier to win.

Maybe that person and his team—maybe he's the head of networking at a big public company—are making a lot of money and have a good life. Then, because he sees these outcomes and the size of the markets in the data center, he's like, “Wow, why don't I just go start my own company?”

But the reason that's important is that there are thousands of parts in a Blackwell rack, and there are thousands of parts in a TPU rack. In the Blackwell rack, maybe NVIDIA makes, I don't know, 200 or 300 of those parts.

And the same thing in an AMD rack. They need all of those other parts to accelerate with them.

Patrick O'Shaughnessy

So they couldn't go to this one-year cadence if the rest of everything was not keeping up with them.

Gavin Baker

The fact that semiconductor venture has come back with a vengeance is really important. Silicon Valley stopped being Silicon Valley long ago. My little firm may have done more semiconductor deals in the last 7 years than the top 10 VCs combined, but that's really important because now you have an ecosystem of companies who can keep up.

That ecosystem of venture companies is putting pressure on the public companies that also need to be part of this if we're going to go to this annual cadence, which is so hard. It's one reason I'm really skeptical of these AI startups that don't already have some degree of success. I do think that's a super important dynamic and one that's absolutely foundational and necessary for all of this to happen, because not even NVIDIA can do it alone.

AMD can't do it alone. Google can't do it alone. You need the people who make the transceivers. You need the people who make the wires, who make the backplanes, who make every—who make the lasers. They all have to accelerate with you.

One thing that I think is very cool about AI as an investor is that it's the first time when, at every level of the stack that I look at, at least, the most important competitors are public and private. NVIDIA has very important private competitors. Broadcom has important private competitors. Marvell has important private competitors. Lumentum and Coherent—all these companies have important private competitors.

There's even a wave of innovation in memory, which is really exciting to see, because memory is such a gating factor. By the way, something that could slow all this down and be a natural governor is if we get our first true DRAM cycle since the late '90s.

Patrick O'Shaughnessy

Say more about what that means.

Gavin Baker

If a DRAM wafer is valued like a 5-carat diamond in the '90s, when you had these true capacity cycles before Taiwan Semi kind of smoothed everything out and DRAM became more of an oligopoly, you would have these crazy shortages where the price would just go 10x—things that are unimaginable relative to the last 25 years.

A giant DRAM cycle—a good DRAM cycle—is when the price stops going down. An epic cycle is maybe it goes up 30%, 40%, or 50%, whatever it is. But if it starts to go up by Xs instead of percentages, that's a whole different game. By the way, we should talk about SaaS.

Patrick O'Shaughnessy

Yeah, let's talk about it. What do you think's going to happen?

Gavin Baker

Application SaaS companies are making the exact same mistake that brick-and-mortar retailers did with e-commerce. Brick-and-mortar retailers, particularly after the telecom bubble crashed, looked at Amazon and said, “Oh, it's losing money. E-commerce is going to be a low-margin business. How can it ever be more efficient as a business?”

Right now, our customers pay to transport themselves to the store, and then they pay to transport the goods home. How could it ever be more efficient if we're sending shipments out to individual customers? Amazon's vision, of course, was that eventually they were just going to go down a street and drop off a package at every house.

They did not invest in e-commerce. They clearly saw customer demand for it, but they did not like the margin structure of e-commerce. That is the fundamental reason that essentially every brick-and-mortar retailer was really slow to invest in e-commerce.

Now here we are, and Amazon has higher margins in its North American retail business than a lot of mass-market retailers. Margins can change, and if there's a fundamentally transformative new technology that customers are demanding, it's always a mistake not to embrace it.

8. The SaaS Mistake

That's exactly what the SaaS companies are doing. They have their 70%, 80%, and 90% gross margins, and they are reluctant to accept AI gross margins. The very nature of AI is that software you write once is written very efficiently, and then you can distribute it broadly at very low cost. That's why it was a great business.

AI is the exact opposite, where you have to recompute the answer every time. A good AI company might have gross margins of 40%. The crazy thing is that, because of those efficiency gains, they're generating cash way earlier than SaaS companies did historically.

They're generating cash earlier not because they have high gross margins, but because they have very few human employees. It's tragic to watch all of these companies. If you want to have an agent, it's never going to succeed if you're not willing to run it at a sub-35% gross margin, because that's what the AI natives are running it at.

Patrick O'Shaughnessy

Yeah.

Gavin Baker

Maybe they're running it at 40%. So if you are trying to preserve an 80% gross-margin structure, you are guaranteeing that you will not succeed at AI.

Patrick O'Shaughnessy

Absolute guarantee.

Gavin Baker

And this is so crazy to me because, one, we have an existence proof for software investors being willing to tolerate gross-margin pressure as long as gross-profit dollars are okay. It's called the cloud.

People don't remember, but when Adobe converted from on-premise to a SaaS model, not only did its margins implode, its revenues actually imploded too, because it went from charging up front to charging over a period of years.

Microsoft was less dramatic, but it was a tough stock in the early days of the cloud transition because investors were like, “Oh my God, you're an 80% gross-margin business, and the cloud is in the 50s.” And Microsoft was like, “Well, it's going to be gross-profit-dollar accretive. It probably will improve those margins over time.”

Microsoft bought GitHub and used GitHub as a distribution channel for Copilot for coding, which has become a giant business. It runs at much lower gross margins, for sure, but there are so many SaaS companies. I can't think of a single application SaaS company that could not be running a successful agent strategy.

They have a giant advantage over these AI natives in that they have a cash-generative business. I think there is room for someone to be a new kind of activist, or constructive activist, and just go to SaaS companies and say, “Stop being so dumb.”

All you have to say is, “Here are my AI revenues, and here are my AI gross margins. It's real AI because it has low gross margins. I'm going to show you that, and here's a venture competitor over here that's losing a lot of money. So maybe I'll actually take my gross margins to 0% for a while, but I have this business that the venture-funded company doesn't have.”

This is such an obvious playbook that you can run at Salesforce, ServiceNow, HubSpot, GitLab, and Atlassian. All of them could run this. The way those companies could or should think about using agents is just to ask the question: “What are the core functions we do for the customer now? How can we further automate that with agents effectively?”

Patrick O'Shaughnessy

100%. Just like if you're in CRM, what do our customers do? They talk to their customers.

Gavin Baker

Yeah, we're customer relationship management software, and we do some customer support, too. So make an agent that can do that, right? Sell that at 10% to 20%, and let that agent access all the data you have, right?

Because what's happening right now is another agent made by someone else is accessing your systems to do this job, pulling the data into their system, and then you will eventually be turned off. It's just crazy, and it's just because they want to preserve their 80% gross margins.

This is a life-or-death decision, and essentially everyone except Microsoft is failing it. To quote that memo from that Nokia guy long ago, their platforms are burning.

Patrick O'Shaughnessy

Burning platform. Yeah.

Gavin Baker

Yeah. There's a really nice platform right over there, and you can just hop to it. Then you can put out the fire in your platform that's on fire. Now you've got 2 platforms, and it's great.

Patrick O'Shaughnessy

Your data centers and space thing makes me wonder if there are other less-discussed, off-the-wall things you're thinking about in the markets in general that we haven't talked about.

9. Nuclear and Quantum

It does feel like, since 2020 kicked off and 2022 punctured it, we've had a series of rolling bubbles. In 2020, there was a bubble in EV startups that were not Tesla, and that's for sure a bubble. They all went down 99%.

There was also a bubble in more speculative stocks. Then we had the meme stocks, GameStop. Now it feels like the rolling bubble is in nuclear and quantum.

Gavin Baker

These are fusion and SMR. It would be a transformative technology. It's amazing. But sadly, from my perspective, none of the public ways you can invest in this are really good expressions of this theme, are likely to succeed, or have any real fundamental support.

The same thing is true with quantum. I've been looking at quantum for 10 years. We have a really good understanding of quantum, and the public quantum companies, again, are not the leaders.

From my perspective, the leaders in quantum would be Google, IBM, and Quantinuum. The public ways you can invest in this theme, which is probably exciting, are not the best. So you have 2 really clear bubbles.

I also think quantum supremacy is very misunderstood. People hear it, and I think they mean that quantum computers are going to be better than classical computers at everything. With quantum, you can do some calculations that classical computers cannot do.

The thought that I have had, which is maybe less related to markets than just AI, is that I have been fascinated that for the last 2 years, whatever AI needs to keep growing and advancing, it gets.

Patrick O'Shaughnessy

Have you ever seen public opinion change so fast in the United States on any issue as it has on nuclear power?

Gavin Baker

It just happened like that.

Patrick O'Shaughnessy

Like that. Why did that happen right when AI needed it to happen? Now we're running up against the boundaries of power on Earth. All of a sudden, we're talking about data centers in space.

Gavin Baker

It's just a little strange to me that whenever there is something—a bottleneck—

Patrick O'Shaughnessy

A bottleneck—

Gavin Baker

A bottleneck that might slow it down, everything accelerates. Reuben is going to be such an easy, seamless transition relative to Blackwell, and Reuben's a great chip. Then you have AMD getting into the game with the MI450. It's just whatever AI needs, it gets.

Patrick O'Shaughnessy

You're a deep reader of sci-fi, so—

Gavin Baker

Yeah, exactly.

Patrick O'Shaughnessy

You're making me think of Kevin Kelly's great book, What Technology Wants. He calls it the technium, the overall mass of technology that is supplied by humans to grow more powerful.

Gavin Baker

Absolutely. It just wants to grow more and more powerful. Now we're going into an insane state.

Patrick O'Shaughnessy

I have a selfish closing question. Speaking of young people, my kids are 12 and 10, but especially my older son is developing an interest in what I do, which I think is quite natural. I'm going to try to start asking my friends who are the most passionate about entrepreneurship and investing why they are so passionate about it, and what about it is so interesting and life-giving to them. How would you pitch what you've done, the career you've built, and this part of the world to a young person who's interested in it?

Gavin Baker

I do believe, at some level, investing is the search for truth. If you find truth first, and you're right about it being a truth, that's how you generate alpha. It has to be a truth that other people have not yet seen. You're searching for hidden truths.

The earliest thing I can remember is being interested in history. I loved looking at books with pictures of the Phoenicians, the Egyptians, the Greeks, the Romans, and pyramids. I loved history.

I vividly remember, in I think the 2nd grade, my dad drove me to school every day, and we went through the whole history of World War II in 1 year. I loved that. Then that translated into a real interest in current events very early.

As a pretty young person—I don't know if it was 7th, 8th, or 9th grade—I was reading The New York Times and The Washington Post. I would get so excited when the mail came because it meant that maybe there was an Economist, Newsweek, Time, or U.S. News. I was really into current events because current events is applied history: watching history happen and thinking about what might happen next.

I didn't know anything about investing. My parents were both attorneys. Anytime I won an argument, I was highly rewarded. If I could make a reasonable argument for why I should stay up late, my parents would be so proud, and they would let me stay up late. But I had to beat them.

I was just going through life. I really loved to ski, and I loved rock climbing. I went to college, and rock climbing was by far the most important thing in my life. I dedicated myself to it completely. I did all my homework at the gym. I got to the rock-climbing gym at 7 a.m. and would skip a lot of classes to stay there. I would do my homework on a big bouldering mat.

Every weekend, I went and climbed somewhere with the Dartmouth Mountaineering Club. As part of that, on climbing trips, maybe we'd play poker. The movie came out while I was in college, and we started playing poker. I also liked to play chess.

10. Gavin’s Investing Origins

I was never that good at chess or poker. I never really dedicated myself to either. My plan, after 2 or 3 years of college, was that I was going to leave. I was a ski bum at Alta in college. I was a housekeeper. I've cleaned a lot of toilets.

It was shocking to me how people treated me, and it has permanently affected how I treat other people. You'd be cleaning somebody's room, and they'd be in there reading the same book as you. You'd say, “Oh, that's a great book. I'm about where you are,” and they'd look at you like you were a space alien. Then they would get even more shocked that you read. It had a big impact on how I've treated everyone since then.

I was going to be a ski bum in the winters. I was going to work on a river in the summers, and that was how I was going to support myself. Then I was going to climb in the shoulder seasons. I was going to try to be a wildlife photographer and write the next great American novel.

Patrick O'Shaughnessy

I can't believe I never knew this.

Gavin Baker

That was my plan. This was my plan of record. I was really lucky. My parents were very supportive of everything I wanted to do. My parents had very strict parents, so of course they were extremely permissive with me. I'll probably end up being a strict parent. The cycle continues.

My parents were lawyers, and they had done reasonably well. They both grew up in very economically disadvantaged circumstances. My dad talks about how he remembers every person who bought him a beer because he couldn't afford a beer. He worked the whole way through college and was there on a scholarship. He had 1 pair of shoes all through high school.

They were very supportive of this plan, and I had been very lucky. They sent me to college, and I didn't have to pay for it. They paid for my college education. They said, “You know, Gavin, we think this plan of being a ski bum and river-rafting guide, being a wildlife photographer, climbing in the shoulder seasons, and trying to write a novel sounds like a great plan. But we've never asked you for anything. We haven't encouraged you to study anything. We've supported you in everything you've wanted to do. Will you please get 1 professional internship—just 1? We don't care what it is.”

The only internship I could get at the end of my sophomore summer at Dartmouth was with Donaldson, Lufkin & Jenrette, or DLJ. It was in the private wealth management division, and I worked for the guy who ran the office. Whenever DLJ published a research report, I would go through it and look at which of his clients owned that stock. Then I would mail the research to the clients.

That day, we wrote on General Electric, so I needed to mail the GE report to 30 people. Then I needed to email the Cisco report to 20 people. I started reading the reports, and I thought, “Oh my God, this is the most interesting thing imaginable.”

I conceptualized investing as a game of skill and chance, kind of like poker. There is obviously chance in investing. If you're an investor in a company and a meteor hits its headquarters, that's bad luck, but you own that outcome. There is irreducible chance, but there's skill, too. That really appealed to me.

The way you got an edge in this greatest game of skill and chance imaginable was by having the most thorough knowledge possible of history. You intersected that with the most accurate understanding of current events in the world to form a differential opinion on what was going to happen next in this game of skill and chance. Which stock is mispriced in the Perry Mutual system? That is the stock market.

That was day 3. I went to the bookstore and bought the books they had, which were Peter Lynch's books. I read those books in 2 days. I'm a very fast reader. Then I read all these books about Warren Buffett. Then I read Market Wizards. Then I read Warren Buffett's letters to his shareholders. This was during my internship. Then I read Warren Buffett's letters to his shareholders again.

Then I taught myself accounting. There's this great book, Why Stocks Go Up and Down. Then I went back to school. I changed my majors from English and history to history and economics. I never looked back.

Investing consumed me, although I continued to really focus on climbing. I would be in the gym and print out everything that the people at The Motley Fool wrote. They were early to talking about return on invested capital, and incremental ROIC is a really important indicator. I would read it and underline it. I'd read books, and then I'd read The Wall Street Journal.

Eventually, there was a computer terminal set up near the gym. I'd go there and read news about stocks. It was the most important thing in my life. I barely kept my grades up.

That's how I got into it: history, current events, skill and chance. I'm a competitive person, and I've actually never been good at anything else.

Okay, I got picked last for every sports team. I love to ski. I've literally spent a small fortune on private skiing lessons. I'm not that good of a skier. I like to play ping-pong, and all my friends could beat me. [laughter] I tried to get really good at chess, and this was before you actually had to play the games, before it was easy to do it on the phone. My goal was to beat one of the people. I'm sure there's a park somewhere.

Patrick O'Shaughnessy

It's literally right there. The famous one is right there.

Gavin Baker

Okay. Well, there's one in Cambridge, and I wanted to beat one of them. I never beat one of them. I've never been good at anything. I thought I would be good at this. The idea of being good at something other than taking a competitive test was very appealing to me. I think that's been a really important thing, too. To this day, this is the only thing I've been vaguely competitive at. I'd love to be good at something else. I'm just not. [laughter]

Patrick O'Shaughnessy

I think I'm going to start asking this question of everybody. The ongoing education of a person may be an amazing place to close. I love talking about everything so much.

Gavin Baker

This is great, man. Thank you. Thank you. Thank you.

GPU、TPU 与 AI 经济学详解|Gavin Baker 访谈 — 文字稿与摘要 | BidClub