Neil Movva——让 AI 成本降10倍——[Invest Like the Best,EP.488]
- Neil Movva 正在把 Sola Research 打造成一家“token 工厂”——这是一家只为一件事优化的推理公司:把单位 token 成本做到行业最低,而且是“遥遥领先”,服务于长时间运行的后台 Agent,而不是聊天机器人。 他的核心预测是:到今年年底,工作负载可能变成“后台与实时各占50/50”,但长期会变成“后台90%、实时10%”,因为后台 token 需求没有上限,而人机协同需求受人的注意力时长限制。“每当你把某样东西变便宜10倍,它就会成为一个全新的产品类别”(Whenever you make something ten times cheaper, it’s a new product category)。
- GPU 系统中延迟与吞吐量的权衡是根本约束,而 Cursor 把市场带向低延迟后,所有人都选择了延迟。 GPU 本质上是吞吐量机器,最适合处理大批量任务——他的比喻是公交车对私人交通;NVLink 和张量并行可以换取最低延迟,但扩展是次线性的,即用8×硬件换来可能4–5×的速度。如果不再在意延迟,那些每美元 FLOPS 击败 NVIDIA、但缺少互联能力的芯片就有了用武之地,这正是他的技术栈要捕捉的套利空间。
- 供给侧策略是有意为之的“捡漏”:从“没有坏芯片,只有坏价格”,一路延伸到别人不愿碰的电力和数据中心。 他愿意“在世界任何地方、以任何期限购买任何芯片”,把算力部署在1MW数据中心里——8个冰箱大小的液冷机架、单路供电、没有冗余——并表示“我会买95%的 uptime”,只要价格合适,甚至可能接受80%的 uptime,包括间歇性太阳能和风电,因为失败请求可以迁移,后台 Agent 也能吸收额外延迟。“不是大型一体化钢厂,而是小型钢厂。”
- 面对泡沫担忧——半导体如今约占 S&P 500 的20%,历史上只有2–4%——他的判断是:今天的资本开支并非投机,不同于互联网泡沫时期的网络设备,也不同于2023–24年的 Hopper 紧缺。 训练天然带有投机性;推理算力则因为能立即产生价值而被购买——“你不会囤积 token,而是立刻使用它们……推理不存在投机”(There’s no speculation on inference)——他预计推理支出将单调上升。
- 两个反共识硬件判断是:按同口径比较,NVIDIA 的 bfloat16 单位功耗性能从 Hopper 到 Blackwell 再到 Rubin 并没有大幅提升;即便失去 TSMC,情况“也不会那么糟”——Intel 最多可能只落后约2倍,远小于芯片战叙事所暗示的差距。 他认为真正值得攻克的瓶颈是 HBM,因为“博伊西那帮人不喜欢为周期性业务投入巨额资本”;他的判断是,iPhone 可能削减内存、提高售价,而不是短期内快速建成新的内存工厂。
- 在前沿实验室与开源的竞争上,他认为“实验室为领先3到6个月支付了巨额溢价”,但怀疑这种溢价能否持续——企业仍在使用 Opus 4.6/4.7,而通过 AI 生成 GitHub 仓库形成的“潜在蒸馏”,从根本上不可能被阻止能力扩散。 “训练出前沿级模型每天都在变得更容易。”
- 他认为最值得按优先级解决的低效环节是 KV cache 内存——“没有体现出应有的价值”,占用量可能高估了1到2个数量级——以及闲置算力。 NVIDIA 今年将生产500万颗 Blackwell,但他不相信这些芯片都在持续运行,并称全球其他地方的闲置利用率比人们嘲笑的 xAI 集群更糟。终局目标是每天处理1万亿个 token;按 OpenAI 对5.5或5.6的定价,价值至少约500万美元,这意味着他要求成本下降3到6个数量级。
1. Sola Research 押注丰沛供给,目标是成为 token 工厂
- Movva 对公司的字面描述是:以“市场上无可匹敌的价格”提供开源 LLM API,同时推出“Sola Boxes”——专为运行数小时、数天乃至数周的 Agent 设计的长时运行 Agent 虚拟机。核心主题是丰沛供给:“机器能够思考”这件事意义如此重大,而现在的任务就是让世界上尽可能多的机器投入思考。
- Token 是当前的北极星指标,但不是最终计量单位;终点是结果,让 Agent “自行管理 token 预算”,而不是由公司给工程师设定支出上限,token 数量则变成任务的从属变量。
- 他提到的两个顺风是开源崛起——客户“在乎拥有智能……你始终拥有模型权重”——以及现有服务商的定位:Baseten、Fireworks、Together AI 都因“一个非常重要的客户 Cursor”而围绕低延迟推理优化。一年前这是正确选择;“从6个月前开始,情况看起来可能是,低延迟并不是唯一需要的东西。”
2. 核心押注:后台 Agent 将吞噬实时工作负载
- Patrick 提出“我希望一切都尽可能快”,Movva 则给出了他的标志性重构:“最好的延迟就是没有延迟(the best latency is no latency at all)。当你早上醒来时,工作已经在夜里完成了,你甚至不需要提出请求。” 人在环路中才是瓶颈;Agent 应该按照人的管理节奏运行——你每周检查一次同事,而不是每5分钟检查一次。
- 证据是测试时算力扩展在去年年底随着 Claude Opus 4.5 变得可下注——那是“第一个多少适合长周期任务的 Agent。它刚发布时相当平庸”——而现在运行1小时已经“相当合适”。沿着这条曲线再走几个点,就能拉出指数增长。
- 规模逻辑很直接:“我喜欢这个市场,因为它没有上限。”他本人无法在 Codex 或 Claude Code 中消耗多10倍 token——已经饱和;后台消费没有这个天花板。他预计今年年底后台与实时为50/50,长期则变成后台90%、实时10%。
3. 丰沛后台推理释放的应用:1万来源研究与作为工作量证明的安全
- 当前最好的例子是对“不是100个来源、不是1,000个来源,而是1万个甚至更多来源”进行深度研究。客户 Parallel Web Systems 希望实时索引和监控整个互联网,这是“疯狂的 exabyte 级任务”,需要完全不同规模的智能。
- 网络安全的故事是:Fable 或 Mythos 刚推出时,社区就希望让它们针对所有曾经写过的代码、以20种不同方式运行,直到“安全变成了工作量证明”——软件有多安全,等于“你花了多少美元购买 Anthropic API,试图攻破自己的软件?”而且前沿能力并不平滑:Haiku 能找到 Fable 找不到的漏洞,反之亦然,因此多样化采样更有价值。
- 更理想化的版本是主动式 Siri:拥有“对你生活更百科全书式的理解”,在你打开手机时直接提示下一步行动。他说现在还没有完全做到。真正的突破在于“你必须愿意在没有任何回报承诺的情况下花费 token”。软件、形式化数学以及潜在的科学发现等可验证问题,成本正从数百万美元降到数千美元,甚至可能降到几十美元。对于不可验证的任务,他的保留意见是:“我们还没有解决人类品味的问题,我不知道这个问题从根本上是否能被解决。”
4. NVIDIA 的早期基因:张量核心证明了芯片面积的价值,“光速”才是信仰
- Movva 第一份工作是在2010年代中期的 NVIDIA。当时管理层在实验室笔记中记录了 ICML/NeurIPS 的研究生使用游戏 GPU 训练模型,并写下“我们应该深入研究这件事”。张量核心,即矩阵乘法加速器,最初需要与图形团队争夺“可能5%、10%”的芯片面积,应用场景首先是计算机视觉卷积。
- 这种文化遗产是:“他们总是在追逐任何一块硬件的光速上限”——追求理论上可能达到的边界。他今天的版本是:“我们追逐的是100%的光速。我不关心相对竞争对手的数字,只关心绝对数字。”
- NVIDIA 的节俭甚至体现在不提供免费牛奶:工程师每月支付1美元加入“milk club”,由俱乐部采购 Costco 牛奶。“我们在 Sola 不这么做。”
5. 延迟与吞吐量的权衡是底层约束,NVLink 本质上是延迟技术
- 机制在于,GPU 希望把多个用户的任务组成大批次处理,但大批次中的每个请求“会在 GPU 上停留更长时间,因为要和其他人的流量一起被承载”。他的比喻是市中心的公交车与私人交通:公交车服务所有人、频繁停站;汽车则走直达路线。“窄而快,还是宽而慢,是计算机科学中的经典权衡。”
- NVLink 可以把一次矩阵乘法拆分到最多8块 GPU 上,从而降低延迟,但扩展并不线性:“你会使用多8倍硬件……可能只得到4到5倍速度。” Patrick 的追问点明了战略核心:NVLink 首先是一项延迟技术。
- 由于 Sola 不在意低延迟,一块擅长矩阵乘法、但无法与其他芯片通信的芯片,“可能是非常、非常好的每美元算力选择”。Movva 优化的是芯片全生命周期内的 FLOPS/美元,用专家并行或流水线并行替代张量并行,并以低延迟服务无法采用的方式隐藏通信延迟。
6. 从物理原理理解 Cerebras 与 Groq,以及 KV cache 为何是它们的棘手问题
- SRAM/DRAM 的区别在于:SRAM 由逻辑裸片上的6T单元组成,速度极快但占用面积大——一块800mm²、尺寸接近 Blackwell 的裸片若全部由 SRAM 构成,也只能容纳个位数 GB;而 NVIDIA 的逻辑裸片上可能只有约500MB。DRAM 则由 Micron、SK Hynix、Samsung 等厂商把电荷写入电容,并且每约50ms刷新一次,就像“不断把数十亿个球抛在空中”。但堆叠为 HBM 后,Blackwell 可以拥有288GB。Cerebras 直接使用整片晶圆,提供约50GB SRAM,带宽达到21PB/s,相比 HBM 的约10TB/s,并获得“一条通往每秒1,000个 token 的路径”。
- 难题是 KV cache,即对话中每个 token 的存储表示,规模“经常比模型权重还大”。它会随着使用量动态增长,在以 SRAM 为核心的架构上很难处理。“KV cache 会成为你身边的一根大刺。”
- 他的预测是混合架构。“Transformer 的原罪,是把这个本质上受内存限制的层——attention——紧挨着一个受计算限制的层——MLP。”可以把权重和 MLP 放在 Cerebras 式 SRAM 上,把 attention 放在拥有更大容量的 GPU 上——“我相信 NVIDIA 和 Groq 正在发生的就是这件事”。应把晶圆级芯片视为与 GPU 搭配的加速器,而不是替代品。
7. Transformer 是海绵,下一块数据资源是 RL 训练场
- Transformer 胜出的原因是可扩展:“Transformer 是把参数规模从百万级推向万亿级的桥梁”,而 attention “可以表示你想要的任何成对关系”,所以“只要序列中存在任何关系,你就能找到它”。他一直记得 Karpathy 的经验:首先过拟合,以证明这种关系确实存在;“一旦能够过拟合,就可以进行压缩,而泛化正是由压缩产生的。”
- 数据方面,“互联网是一次性补贴”——大约有30万亿个高质量 token,扩大口径则有300万亿个,“而我们基本上已经把它全部看过了”。即便是随机的 ChatGPT 交互数据,如今也不再提供太多价值:“你从随机人类偏好、未经条件化的人类偏好中得到的信号,现在实际上已经不值钱了。”他认为当前真正重要的是专家级人类偏好。
- 下一代数据资源是带有可验证任务和自我评分机制的 RL 环境训练场——“环境本身就变成了数据”。他更进一步认为,只要每项任务都可验证,AGI 将通过“不断叠加专用智能,直到不再有空白需要填补”实现;实验室的支出也已经从数据转向 RL 环境。
8. Kernel 如今在白板上编写,软件护城河正在收窄
- Kernel 并没有被解决:任何模型偏差,例如新的 RoPE 变体,都会破坏 FlashAttention 的顺畅路径。但工作方式已经改变:“我们现在在白板上写 kernel”——人类负责概念设计,用自然语言描述调度策略,再由模型实现。Tri Dao 的观点是:“你不一定还应该手写 kernel。”
- 在 Patrick 追问下,他坦承 kernel 技能不是护城河。“像 Mythos 或 GPT-5.6 Sol 这样的涨潮力量,会把所有船都抬起来……在这项能力上,我们都处于前沿模型的下游。”面对 Patrick 关于煤炭的类比,他给出了更细致的回答:大批量矩阵乘法已经达到峰值的70–80%,限制来自功率节流而非软件;NVIDIA 报出的峰值 FLOPS“有些乐观”。真正的前沿是编程整个 NVL72 机架——GB300 出货形式是由72个单元组成的机架——“现在是一场开放竞争,看谁能找出编程整台机架级计算机的方法。”
9. “没有坏芯片,只有坏价格”:捡漏式芯片采购
- 关于 Blackwell 紧缺,NVIDIA 原本可以通过涨价出清市场,但 Movva 认为,公司也在考虑让最有钱的买家获得过多权力的长期风险;“他们明白,如今算力就是权力。”获取芯片要靠关系,或者“惊人的财力”。
- 他认为 AMD 的芯片“整体很出色”,问题在于人们不会编程,而其他厂商也没有 NVIDIA 那么多开箱即用的 kernel 工程。“其他人觉得 AMD 不如 NVIDIA,这对我来说是天籁之音。我很高兴他们低估这块芯片。”不过他也承认,这个判断“现在已经不完全正确”,Meta 和 OpenAI 都在大量采购。他还关注 Etched、SambaNova 和 d-Matrix;这些公司的约束在于 TSMC 晶圆配额。
- 他的立场是:“我们会在世界任何地方、以任何期限购买任何芯片”——TPU、Trainium,什么都可以。“每块芯片都有比较优势。我们要找到这个优势,然后沿着那个方向把它榨干。”
10. 泡沫问题:“与其说我是历史的学生,不如说我是历史的参与者”
- Patrick 的铺垫很精准:半导体历史上占 S&P 500 的权重为2–4%,如今已经升至19–21%;Micron 和 SK Hynix 的投资者担心均值回归,因为“长期来看,算力是一种商品”。
- Movva 从个人经历回答:他出生于1997年,母亲在互联网泡沫上涨期间任职 Intel,自己记得 Cisco 曾是全球市值最高的公司。区别在于,互联网泡沫时期的网络设备资本开支押注了从未到来的需求;即便2023–24年的 Hopper 紧缺,也主要面向训练,“天然具有投机性”。如今“人们购买 token,是因为它们立即有价值……你不会囤积 token,而是立刻使用它们”,企业也开始为 Claude Code 设定支出上限。“我确实认为推理支出会单调增加。推理不存在投机。”
11. 1MW 数据中心、95% uptime,以及别人不要的太阳能
- 市场结构是:数据中心原本按照训练需求建设,而训练覆盖的是更大的工作负载集合;同时,规模越大反而越不经济——在美国,GW 级站点基本不可能,100MW 很难,10MW 已接近极限,而“1MW,我认为供应非常充足”。总量上的电力到处都有,但集中式电力并不普遍。液冷可以让1MW容纳约8个冰箱大小的机架。
- 激进之处在于砍掉冗余:没有柴油备电,只有一条埋地光纤,也不提供故障切换 SLA。“我不会惊讶其中一些数据中心最终降到95%的 uptime……你几乎不可能找到95% uptime的数据中心买家,而我就是第一个买家。”故障会通过稳健的控制平面迁移;后台 Agent 的一次执行多花1分钟、几分钟,甚至10分钟,“我的客户并不在乎,因为他们的 Agent 本来就运行了数小时”。明确的交易条件是:平均吞吐量具备竞争力,P99 不受控制,但经济性无可匹敌。
- 电力遵循同一套逻辑:“如果价格合适,我甚至能接受80%的 uptime 吗?可能可以。”太阳能和风电的间歇性——最坏情况下可能连续数天或数周停供——“实际上高度可预测”;建模天气,再迁移工作负载。Patrick 给这种策略贴上了准确标签:捡漏。“我不想和 Anthropic 或 OpenAI 竞价争夺算力……我们要建造全球最好的钢厂,但它会通过小型钢厂实现,而不是大型一体化钢厂。”
- 关于垂直整合,他描述了创始人与 CEO 的分工:“我内心的创始人想把一切都做了……在从头到尾建成最高效系统之前,我永远不会停下”(Patrick 称之为“现实版 Factorio”);但 CEO 会先利用软件的杠杆,更倾向于签电力购买协议,而不是拥有发电设施,并通过实际表现赢得把更多环节收归内部的资格。
12. 按优先级盘点浪费:KV cache 与闲置芯片
- 他的低效清单中,算力扩展已经相当高效:前沿 MoE 激活的专家比例更接近1%,而不是10%。但 attention 的内存使用效率很差:“KV cache 压缩得远远不够……没有体现出应有的价值”,每个 token 占用的 KB 可能高估了1到2个数量级。DeepSeek 已经发表了进一步压缩 KV cache 的工作,并取得实质性进展。
- 更大的丑闻在调度。NVIDIA 今年将生产500万颗 Blackwell——“它们是不是一直都在被使用?我当然怀疑”——但他表示,大量 GPU 闲置在仓库或私人算力池中。“我们都拿 xAI 集群的总 FLOP 利用率问题开玩笑。现实是,世界其他地方的情况糟糕得多。”
- 晶圆厂方面,不要期待按下按钮就能解决问题——“芯片多生产20%,马上就会出现另一个瓶颈”。他的创造性设想是:TSMC 在收紧制程窗口上投入巨大,试图让最差的芯片匹配最好的芯片;如果把性能与价格的取舍透明化,他很乐意接收更多次品,并扩大美国电力供应,让边际芯片“在数据中心证明自己的价值”。对应到招聘,他筛选的是“100%的好奇心”和对性能工程的热爱;CUDA 经验是“一个巨大的烟雾弹”。
13. 开源与闭源、蒸馏异端,以及 Jensen 为什么不卖 token
- 他的框架是:“实验室为领先3到6个月支付了巨额溢价……这笔钱可能仍然值得”,但这种溢价能否持续值得怀疑,因为企业仍在使用 Opus 4.6 或4.7。对于把蒸馏视为窃取的观点,他的替代解释是:AI 生成的产物正在充斥公共知识库——“我们认为过去1年创建的仓库中,有多大比例是 Claude Code 写的?”——“潜在蒸馏”还会持续很久,他不认为从根本上可以阻止信息或模型能力扩散。开源也不可能消失:“如果一个领导者退出,就会有新的领导者补位……训练出前沿级模型每天都在变得更容易。”
- 两个反共识判断是:按同口径比较,bfloat16 单位功耗性能从 Hopper 到 Blackwell 再到 Rubin 几乎没有变化,在 TSMC 5/4/3/2nm 制程之间也类似;因此失去 TSMC 渠道“不会那么糟。供应肯定会受到冲击”,但 Intel “最差也可能只差2倍的单位功耗性能”。对硬件创业者而言,要掌握4个瓶颈——TSMC 晶圆、HBM、先进封装和电力——并且“形成尖峰优势”;HBM 是最值得攻克的环节,因为“博伊西那帮人不喜欢为周期性业务投入巨额资本,他们已经被反复教训过”。结果可能是 iPhone 削减内存、提高售价。
- NVIDIA 为什么不自己卖 token?“Jensen 很擅长让他的朋友们成为亿万富翁。”一批多元、渴望增长的 neocloud 足以构成替代供给:如果某家服务商转向 AMD 或垂直整合,就可以由其他服务商接替。与此同时,未来需求的规模极其庞大:每天1万亿个 token,按 OpenAI 对5.5或5.6的定价,至少价值约500万美元,意味着成本需要下降3到6个数量级;不过“对于某种规模的模型,我们正接近把1万亿个 token 的成本做到数万美元”。当被问到市场是否需要如此丰沛的智能时,Movva 回答:“我永远不会相信需求会消失。世界上总是需要智能。”
My guest today is Neil Movva, the founder of Sola Research. Sola is building what Neil calls a token factory, an inference company designed for a specific kind of future—one where AI agents run in the background for hours or days at a time rather than answering a human in real time. In that world, latency matters less and cost matters much more. Neil has built the entire company around driving the cost of a token as low as it can possibly go.
What makes this conversation special is that it’s one of the most detailed tours I’ve ever done through the full stack of intelligence: the software, the chips, the power, and how the three connect. Along the way, we cover the trade-off between speed and cost that lives inside every GPU, his scavenger strategy for buying the chips and power no one else wants, his contrarian view on NVIDIA, and why the premium that frontier labs charge for being 3–6 months ahead may not last.
I think it’s important early in these conversations to say the thing—literally, what you’re building and what it does today. Maybe orient us there with a brief description: literally, what is the system you’re building, and why should it exist?
1. Building A Token Factory
Sola Research is a token factory. We have an API where anyone can send us requests to use large language models—open-source large language models—for any task they want. We will serve those tokens to them at a price that is unbeatable in the market. We also support their ability to build agents on top of this. We host what we call Sola Boxes, which are long-running agent virtual machines hosted in the cloud, designed for agents that run for hours, days, or weeks.
So I should think about you as a peer company to others that serve different kinds of inference. You’re serving one specific kind of inference, and your goal is to be the absolute cheapest provider and enabler of a certain kind of use of intelligence.
Exactly. The theme of our company is abundance. We want to deliver this new commodity of intelligence to as many people as possible at a cost that is sustainable for almost every industry. We think that whenever you make something 10 times cheaper, it’s a new product category. We aspire to do that for tokens. We think it’s so profound that the machine can think, and now our job is to make as many machines as possible in the world work toward thinking.
So if you think about the theme of the day being token costs, is token cost the right way to think about this? Is there some other way you’d put it?
To start with, absolutely: token cost. Today, my North Star is to have the lowest cost per token in the industry and do that by a mile. I don’t think tokens are the final unit of work or intelligence, but they are what we use today. After tokens, you start to move more toward outcomes, which is a vague direction.
You can imagine, for example, that today, when you consume tokens through an agent, you don’t actually control how many tokens the agent reasons for. It can reason for a certain amount of time, or it can call a certain number of tools. Increasingly, I think we will have agents do some unit of work, take as many shots on goal as they can, and however many tokens they use to get there is going to be a dependent variable depending on the task. You think about agents that self-administer a token budget, as opposed to a company setting a budget for how many tokens engineers can spend per month.
Why is there an opportunity here that you can tackle? It seems like the entire world is oriented around more, better, faster, cheaper tokens right now. It seems like the world is trying to solve this problem very aggressively. What was the unique opening that you saw that maybe the market isn’t being efficient in its attempt to tackle this?
2. Open Source Enables Long Horizon
I think there are 2 things that are tailwinds for our company. One has got to be the rise of open source. I had to talk about that first. I think we are starting to see an increasing number of our customers and the broader market care about owning intelligence. They want to have control—sovereignty over the thing that they depend on.
That created a much more robust market for our customized models, or even just these vanilla open-source models that no one can ever take away from you. You always have the weights. You always have the right to deploy them however you like.
In that world, there’s been a reasonably robust market for the past couple of years serving these models at large scale. The challenge is that all those companies—you could take your pick: Baseten, Fireworks, Together AI—all focus on low-latency inference, and they were pulled in that direction by one very important customer: Cursor. I think that was the right choice about a year ago, and as of 6 months ago, it started to look like maybe low latency wasn’t the only thing you wanted from an agent.
You wanted more persistence, more long-horizon tasks. Now, it’s very obvious to me that the future of agentic inference is long-horizon tasks. You’re gonna run the machine for hours or days at a time. It doesn’t matter if it spits out tokens at 100 tokens per second. Maybe 10 is just fine if that comes with corresponding advantages in efficiency.
Why are you so confident in that? To me, it seems like I want everything as fast as possible.
When you’re waiting on it, you absolutely deserve the fastest answer possible.
Yeah.
My trick is that I don’t want you to be waiting on it. I want it to be proactive. I want it to be in the background. One way to say it is that the best latency is no latency at all. When you wake up in the morning, the work’s already been done overnight. You didn’t even have to ask for it. That’s the dream. We’re not quite there yet.
More importantly, I think that the more you’re in the loop as you prompt agents and wait for a response, the more you’re the bottleneck in having the agent do more or less work. What we’d like is for the agent to operate on more human time scales. You don’t manage your colleagues every 5 minutes. You ask them to do a high-level task, and you come back and check in maybe every day, but more likely once a week. And that, to me, is the future of human-agent collaboration, more like human time scales.
Say more about the early indications that this is happening and, therefore—
Yeah.
—you should be building this company.
So the first and most important thing is the idea of test-time compute scaling: the idea that you can give an agent more time, and it will give you a better answer. That was theorized about 2 years ago, but it wasn’t really something that we could actually bet on until, I would say, late last year with Claude Opus 4.5. Claude Opus 4.5 was the first agent that was at all suitable for longer-horizon tasks.
It was pretty mediocre when it first came out, but you look at the more recent models and what we’ve done in open source as well, and you see that agents are capable of running for an hour at a time. I wouldn’t say it’s days, but definitely an hour is quite suitable today. Seeing average task length get longer and longer, it doesn’t take many points to have you draw out the exponential and see that agents are worth running for longer periods.
What do you think will be the market share of long-running agents in 3 years or something like this?
I love this market because it’s unbounded.
There’s no human in the loop, so you can consume as many tokens as you like in the background versus the human attention span. If you tell me to consume 10× as many tokens at Codex or at Claude Code, I’m actually not sure if I can anymore. I’m already in the loop and locked in coding for most of the day that I’m at the laptop. What is unbounded is how many tokens can be consumed in the background or proactively.
Long term, I think we’re going to end this year at maybe 50/50 background and real-time workloads, but I see this going to 90/10 in favor of background.
What are your favorite examples of something that gets accomplished much better as a background task than as a human-in-the-loop task?
Most deep research—most questions where you want to have a definitive answer over not 100 sources, not 1,000 sources, but 10,000 sources or more. If you want to build an authoritative index of information, like, for example, one of our customers, Parallel Web Systems, seeks to do, they want to build an index over the whole internet, and they want to monitor the internet in real time for changes. That is the kind of crazy, exabyte-scale task that you need a very different kind of intelligence, or scale of intelligence, to achieve. Deep research is a top category for us.
And then, increasingly, we see cybersecurity following this direction. If you think about it, yes, there’s so much code you can generate, but there are exponentially more ways to break that same code than there are to generate that code. There are some great customers out there who are working very hard to find agents that can break any piece of software and proactively patch it.
When Fable first came out, for example, or when Mythos first came out, there was this push in the cybersecurity community to run Fable against every line of code we’ve ever written and look for bugs in 20 different ways. That means you’re looking for memory errors, business logic errors, network vulnerabilities—all these things are actually things that you would write specialized agents for. You wouldn’t just have Fable look at the source code once. You’d actually set up environments where you could penetration-test these applications.
At some point, people started to make this joke that security has become proof of work. When you want secure software, it’s really a question of how many dollars you spent on Anthropic’s APIs trying to break into your software. That is the best indication of how secure it is because that’s the best tool in the world.
Increasingly, we found that open-source models—the frontier of intelligence here is quite jagged. It’s not the case that Fable finds a superset of all bugs in software. You would find some bugs with a very small model that you don’t find with a large model. You’d find some bugs with Haiku that you wouldn’t find with Fable, and vice versa. So it encouraged this very diverse approach to sampling and trying to build cybersecurity agents that break software autonomously so that you can patch it.
If you were to get speculative and imaginative about the sorts of things that very cheap, very long-running agents can enable, we talked about some very practical examples—deep research, cybersecurity, et cetera. If you get a little bit dreamier about the use cases, the new product category that this sort of inference will unlock, I guess the question is just, so what? If you’re maximally successful, dream a little bit about what that might enable.
I think for individual users, what I’m most excited about is this idea of proactive intelligent agents. You can imagine Siri running in the background all the time to understand all the emails you received in a day, all the text messages you receive in a day, and have a much more encyclopedic view of your life and how to be helpful in that life. Right now, they’re still point solutions. You end up doing a lot of prompting. Siri is not very proactive. That’s something we can fix with abundant inference.
If you trust the machine enough—that it’s reliable and also trustworthy, as in private—you might even imagine the machine understanding how you interact with it and proactively surfacing your next action. Whenever you open your phone, can we build a good model of what you’re going to do next? My estimation is yes, we totally can.
And the key to that is incredibly cheap intelligence.
You have to be willing to spend tokens without any promise of return. That is the unlock.
The long-lens view to take on this is that we have a form of intelligence that can tackle any verifiable problem. Any verifiable problem means most software. It means a lot of formal math proofs and similar. It could also mean scientific discovery. These are all relatively verifiable problems, and all those things currently have a dollar cost attached to them.
Essentially, that’s a hidden cost. It’s like, how many tokens could you possibly harness to make this work? We have started to bring within view a dollar cost for these long-horizon tasks that is reasonable. It’s not millions; it’s thousands, and maybe it could be hundreds or even tens of dollars in the near future to have a definitive answer to any scientific question, to any research problem.
So if we dream about that future, we then become limited just by the questions that people can ask, basically?
Pretty much—the questions we can ask. The models are on the cusp of taking even a high-level question and chasing it down through every possible follow-up. You can have the model essentially take that on its own. The question is: What is your token budget? And we will solve the token-budget problem.
What about non-verifiable tasks?
Those are basically the entire category of human taste. We have not solved human taste yet, and I don’t know that it fundamentally can be. I’m excited to be surprised here, but we are focused on very quantitative problems. We leave the quality of writing and the beauty of art to people.
I think for individual users, what I’m most excited about is this idea of proactive, intelligent agents. You can imagine a Siri that is running in the background all the time to understand all the emails you received in a day, all the text messages you receive in a day, and has a much more encyclopedic view of your life and how to be helpful in that life. Right now, there are still point solutions. You end up doing a lot of prompting. Siri’s not very proactive. That’s something we can fix with abundant inference.
All right, now let’s talk about the very clever stack of solutions that you hope to build ultimately to have this giant token-factory supplier of extremely low-cost intelligence. I think you think about this in terms of software, hardware, and power. Talk through what your master plan is to approach this challenge that’s so different from what others are thinking about doing.
3. The Full Stack Strategy
We always had to start with software. Where is the opportunity in today’s data centers to improve efficiency? The first thing we did was try to build the entire LLM software stack around peak GPU efficiency. We’re using NVIDIA GPUs, and we wanted to squeeze out more tokens from the same chip than anyone else in the world.
That starts with the lowest level of programming: kernels. It’s actually my background. I spent my whole professional life working on GPUs and kernels. NVIDIA was my first job while I was in college. I got to see how the tensor cores earned their right to be on the chip.
What does that mean? What is a tensor core?
A tensor core is a specialized unit on the GPU that accelerates matrix multiplication. Simple as that. There’s been a long history of how we evolved that tensor core over time, which we’ll get into.
And why is matrix multiplication so important?
I cannot say that there is a divine truth of the universe that explains why matrix multiplies seem to be the atomic unit of computation. But one way I’ve heard it described to me is: It’s a really succinct way to mix two blocks of numbers together and have them interact in some interesting way. That’s as much as I can say about it.
Yeah.
It is really convenient that linear algebra turns out to be a very compact representation of arbitrary relationships in data.
NVIDIA, a great graphics company, had dominant market share in GPUs and gaming graphics for quite some time. Starting in the mid-2010s, they started these skunkworks projects to make the graphics processor more suitable for the machine learning tasks they were tracking. I remember reading some of the lab notebooks of my managers when I was at NVIDIA. They would visit small ML conferences like ICML or NeurIPS, and they would take note of papers like, “This deep learning thing seems to be catching on. What’s really interesting is that these grad students are using gaming NVIDIA GPUs to train their large models. We should double-click on this and figure out what’s going on here.”
By 2015 or 2016, at least Jensen had the conviction to double down on, “This usage of our chips is only going to grow. Let’s start allocating more and more precious silicon die area to this capability that seems to be emerging. Let’s put the first version of tensor cores on the chip.” When we’re talking about taking this gaming chip, which is designed for painting pixels on a screen, and adapting it to do matrix multiplies, it was early. You would be competing against the graphics teams, essentially. When you ask for more silicon area at any chip company, there’s always competition for that.
It is something that the designers guard so carefully. You don’t ever want to invest in the wrong technology, because that’s opportunity cost that you could have allocated to some other functionality. We fought tooth and nail and got just a tiny bit of die area, maybe 5% or 10%, for the first generation of these chips to get some amount of acceleration for basic convolutions, which were the fundamental operation for computer vision models of the day. Then we had a software team that was trying to squeeze all the performance we could out of the chip.
I think on that software team, which is where I worked, that’s what actually taught me the most about the ethos that NVIDIA has around this term called speed of light. They always chase the speed of light for any piece of hardware that they make. It is so ingrained in every engineer’s mind that if the machine can do it, we’re going to push the machine to the frontier until it does what we think it should do.
Because the speed of light is the edge of what’s possible.
The speed of light is the edge of what’s possible, exactly. If we think the chip can run at this frequency and produce this many multiplies per cycle, we’re going to get there. We’re going to break every bottleneck and get to that peak level of performance. To this day, I tell all my engineers, “We’re chasing 100% speed of light.” I don’t care about relative numbers versus the competition. I only care about absolute numbers. What are we able to do on the chip? How do we achieve that?
Before we leave that chapter of your time at NVIDIA, is there anything else beyond that cultural touchpoint that changed the way you think about things or that stood out the most about how the business ran back then or its culture?
I have a ton of stories about NVIDIA. I could tell you a few of them. One of my favorites is that, on the tenure side, a lot of people I worked with at NVIDIA in 2015 and 2016 are still there today. That company has incredible retention, and these are the best engineers, on the silicon side at least, I’ve worked with in my whole career. They’re extremely motivated and passionate. They believed in parallel computing as a concept through its various incarnations and have loved seeing the chip evolve. This is their life’s work, and they’re extremely competent in that direction.
They’re also a very frugal company. NVIDIA and all of the Silicon Valley companies after 2008 had some cutbacks in perks, so there was no free lunch, for example. NVIDIA took it one step further: There was no free milk in the fridge. If you wanted to drink coffee at NVIDIA and wanted some milk, you’d actually have to chip in $1 every month to the milk club. The milk club would stock Costco milk in the fridge, and I remember that distinctly. We don’t do that at Sail. It’s a frugality that permeates the company.
Coming out of this time there, you get this experience of what it’s like to develop more efficient usage of the underlying hardware through software. Link that to today’s environment.
4. Throughput Beats Latency
The GPU is fundamentally a throughput machine. The GPU is happiest when you give it a lot of work to do and let it chew through that work at peak utilization of its compute units. But that’s not actually the way that we’ve taken AI in the last couple of years. We’ve pushed AI to be an interactive chatbot tool, which is the most common form of AI usage today. In that world, you care a lot about spitting answers out to the person at the keyboard as quickly as possible. To your point, you don’t want to make the user wait. You want things as fast as possible.
That’s actually quite interesting for the GPU. It’s very difficult to put the GPU in its happy path of being fully compute-utilized when you’re trying to spit out tokens quickly. There’s a fundamental trade-off on the GPU between being throughput-oriented and latency-optimized, and everyone has chosen latency optimization because the shape of usage was chatbot-oriented. I believe that’s the most profound change we’re going to see in the next year. We’re going to move away from chatbots to more proactive or background agents. In that world, it makes a lot more sense to build a stack around throughput.
Can you explain technically why the trade-off between throughput and latency is unbreakable? Why can’t we have both from the same hardware?
It’s quite foundational in almost every system that you could ever possibly look at. There’s always a trade-off between getting a small amount of data through the system as quickly as possible and leaving a lot of buffer room for that, or trying to run wide and slow. Narrow and fast or wide and slow is a classic trade-off in all of computer science.
For GPUs specifically, I think there’s one thing to focus on, which is the concept of batching on the GPU. We want to group many users’ work together into a batch that we can run all at once on the GPU. That’s the parallel processing of the GPU. We’d like to have a lot of parallel work to do.
The thing is, though, you’re doing more work when you run a large batch of compute together, so you might be filling all the units. But every step along the way, as you carry a batch of work through the GPU, there’s more work to be done. Any individual token or any individual user’s request in that batch is going to spend a longer time on the GPU being carried with other people’s traffic.
Maybe the way to say it is that if you want to get downtown in San Francisco, you can take the bus or you can take private transit. Private transit is going to have its own direct path, like as the crow flies, or use exactly the roads that you want from point A to point B. A bus is going to have to serve many more people, and it has to fundamentally do something that works for everyone, so it takes a slower path and stops and waits for other people to get on and off. I think the bus-versus-car analogy is pretty accurate.
It’s a great analogy, and step 1 for what you’re trying to do is create the best possible bus on top of NVIDIA GPUs. That’s step 1 of your optimization?
That’s exactly right. It means we explore things like different parallelism schemes. Maybe another example I can give you is with NVIDIA GPUs. One of the things that NVIDIA has really innovated on and done a great job with is the NVLink interconnect between GPUs.
In fact, that NVLink system is so good that if you have a large matrix multiply that you want to perform faster, you can actually cut that matrix multiply in half and shard it across 2 or more, up to 8, NVIDIA GPUs, and have them all work on pieces of that larger matrix multiply. Then they can connect their results together at the end and reduce their results back together.
This is a great way to cut the minimum latency of an operation. Each GPU is now doing one-eighth as much work, let’s say, and therefore it can finish faster, but not 8 times faster. It’s sublinear scaling. You’ll use 8 times more hardware, but you won’t get 8 times the speed. You might get 4 to 5 times the speed. You’re not going to get strong scaling.
This is because of communication overhead. It’s because every GPU is going to be a little less efficient working on a smaller tile of work than a larger tile of work. It’s the only way to speed up. If you want the minimum latency possible, you can do that, but it is not the choice I would make, for example. I would prefer to use a different parallelism scheme, like expert parallelism or pipeline parallelism. We may do interesting things to overlap and hide the communication latency in a way that you would have less ability to do that for a low-latency service.
Is the right way to think about NVLink as a technology that improves latency performance—
Yes.
—and only latency performance?
That segues into the next segment of what we do differently as a company. But yes, NVLink is mandatory for low-latency inference. NVIDIA is excellent at low-latency inference, and I’m telling you that we don’t really care that much about low-latency inference. So where does that leave us?
I’m not holding my breath for other companies broadly to figure out NVLink quickly. It’s a challenging technology to figure out. It’s hard to scale. It’s hard to productionize. If I do have some other vendor’s chip and it is good at the foundational compute components—it can still do matrix multiplies really well, it just can’t communicate those results across its peers quickly—well, maybe there’s room for that other chip in my stack as a really good compute-per-dollar option.
That's what I actually optimize for in most cases: how many FLOPs this chip has and how much it's going to cost me per hour to own and operate. There are other chips that definitely rank higher than NVIDIA on FLOPs per dollar, but they may not have as much interconnect. So it's my job to figure out which parallelism scheme I'm going to use to make this chip suitable for inference. It's not going to be tensor parallelism. NVIDIA is basically mandatory for that, but other techniques may work well for me.
So before we leave the latency part of the story—
Yeah.
Can you comment on companies like Cerebras or others that can perform incredibly fast operations? I'm curious what you think about those approaches and what might happen in the future. What is your prediction for the future of very-low-latency-focused hardware?
5. Specialized Chips Bet On Memory
Cerebras, Groq, and a couple of others that are coming out of stealth now have made a very interesting bet on not just building another GPU, but actually building a different kind of accelerator that focuses on a different memory hierarchy. They want to maximize the amount of SRAM on the chip and use that as very, very fast memory for weights and KV cache.
There are two ways to make memory for a chip: SRAM versus DRAM. One is to integrate the memory on the logic die itself, meaning you tell TSMC, “I want this many megabytes of storage on my chip,” and there’s a way to build that. TSMC has a standard-cell library you can use, and you can just print out a bunch of SRAM cells. The problem with SRAM is that it takes a lot of area on the silicon die.
So if you want to build a large die, like, let’s say, the NVIDIA Blackwell at 800 square millimeters, if you made that whole die SRAM, it would be in the single-digit gigabytes, it feels like. It’s not a crazy amount of data storage. Compare that to if you’re willing to take a different process technology entirely—not TSMC anymore, but Micron, SK Hynix, or Samsung. They build DRAM, which is a whole different way to build memory that’s more focused on capacitors than transistor cells.
The standard way to build SRAM is what’s called the 6T transistor cell. It’s a stable transistor arrangement that allows you to write a bit to it, and then it holds that state as long as you apply some power, without any sort of active management. It’s static. Dynamic RAM, or DRAM, is dynamic because what you do to write some data is write a charge onto a capacitor. As soon as you write that charge into the capacitor, the charge starts dissipating. The dynamic part of DRAM is that you must refresh every bit you’ve written every 50 milliseconds or so.
So you’re constantly juggling billions of balls in the air, essentially. Billions of bits have to be managed by a memory controller, which is reading and refreshing every bit on the DRAM. The benefit of that is that you can get much, much higher density, and it’s a whole different process technology. There are a ton of different trade-offs, hence why we split DRAM manufacturing into entirely different companies like Micron, SK Hynix, and Samsung. These are the best companies in the world at doing this.
They build DRAM, and if you take DRAM from those companies and stack it into many layers and kind of print it or solder it around the main logic die that you get from NVIDIA, you can now get hundreds of gigabytes. Blackwell has 288 gigabytes of HBM capacity around the logic die, and the logic die itself may only have 500 megabytes of SRAM. So it’s possibly a difference of multiple orders of magnitude—three orders of magnitude—in density for DRAM versus SRAM.
So let’s go back to Cerebras. What are they doing? They see this problem. There’s not really an obvious way to increase SRAM density on the chip, but the thing with SRAM is that because it’s so physically close to the logic gates that actually do the computation, the arithmetic logic units are right next to the SRAM that they’re going to pull from. The compute units that are doing the matrix multiplies can pull data from SRAM at mind-boggling speeds.
Cerebras quotes petabytes per second: 21 petabytes per second for its Wafer Scale Engine 3. Compare that to HBM on an NVIDIA Blackwell, which is 10 terabytes per second or so, in that range. So once again, there’s a difference of many orders of magnitude: more capacity, but proportionally less bandwidth, essentially.
What Cerebras does is say, “We’re going to take as many of these dies as we can. We’re not going to limit ourselves to the 800-square-millimeter reticle limit that TSMC imposes on us. We’re going to take the entire wafer and have every die connect to every other die over scribe lines, and we’re just going to try to get as much SRAM as we can on the whole wafer.” We can get to, let’s say, 50 gigabytes of SRAM per wafer. Then we’re going to stack many wafers together in a pipeline or something similar, and now we can have up to a terabyte of very, very fast memory.
You do all that work just to get the ability to read data from SRAM at 21 petabytes per second per wafer. Therefore, you can now serve these language models at extremely high tokens per second because you can move the entire parameter count of a large model like Kimi. You can move all that data on and off the logic cores in about a millisecond or something like that.
Yeah.
So there you go. You have a path to 1,000 tokens per second.
What is your prediction for that segment of the market?
What happens to them is some hybrid. We have to pair the Cerebras chip, where it’s very strong and very, very good at fast access to memory, with something that has more capacity for memory. It’s true that you can take a 1 trillion-parameter model like Kimi and fit it on a large number of Cerebras wafer-scale engines, but you can’t do something about the KV cache very easily. The KV cache is something that grows as people use the model more, and that is always dynamic. You don’t even know how much KV cache you’re going to need. It depends on how many users you have and how many users you want to serve.
Can you explain KV cache?
KV cache: whenever you use a language model, every token you send through the language model stays in the context window of the language model for as long as you’re having a conversation. We talk for 100,000 tokens, and the 100,001st token is still in the conversation behind us. The model is referencing all the past conversation history in order to make better predictions about what the next thing we’re going to say is.
That KV cache is a bunch of memory. You have to store a representation for every token that we send through the language model, and it frequently gets to be larger than the weights of the model themselves. You have this crystallized knowledge in the model weights, and you have the dynamic knowledge of the exact conversation we’re having in the KV cache. That’s the way I like to think about it.
And this is why sometimes people observe that, deep in a conversation, things start to degrade because there’s some sort of technical problem?
The KV cache is quite interesting in that regard. The KV cache is an exact representation of everything that came before. We store all the information that we’ve seen in the conversation. However, during training, the model was not trained primarily on very long-context conversations. It was trained primarily on, let’s say, 8,000-token conversations or 16,000-token conversations.
So if you take the model to 200,000 tokens, there was some training that happened at that context length, but it’s not the model’s core strength. There’s always been a challenge for the frontier labs to figure out how to make the model exactly as intelligent at 10,000 tokens as we expect it to be at 200,000 tokens.
Mm-hmm.
And it’s going to be a perennial battle for us. We’ve had 1-million-token context windows as a concept for years now. Anthropic was, I think, the first to hit that 1-million-token context-window length. I still use /compact in my Claude Code well before 1 million context length. I don’t think it’s actually great to hit the full length.
These extremely fast, extremely low-latency approaches—
Yes.
—ultimately are limited by this factor.
Yes. You can do whatever you want for the weights. It’s very possible to have much better performance on weight storage. However, the KV cache is going to be a big thorn in your side.
So 3 years from now, 5 years from now, what role do you think these kinds of chips play? What sort of market share do they have in the heterogeneous chip market?
Cerebras and Groq, and maybe a couple of others—you should think of them as accelerators. What they are really good at is being used in conjunction with a more traditional, GPU-like device that critically has this off-chip memory built in. You want off-chip memory for capacity and on-chip memory for speed. We want to hybridize these two things.
So if you take transformers in the limit, if you take a transformer to a 1-million-token context length, what ends up happening is that you have this compute-bound stage, which is the actual matrix multiplies for what we call the MLP, where most of the model’s world knowledge is encoded. Then you have the attention layer, which is where we’re dynamically adapting to the current conversation.
Attention in the limit is usually memory-bound. The MLP in the limit is compute-bound at a large enough batch size. I would say the original sin of transformers is that you’ve taken this fundamentally memory-bound layer and juxtaposed it right next to a compute-bound layer.
It is very difficult to have a single chip that is good at both compute operations and memory operations. The GPU is quite balanced in this regard, but you had to choose one or the other. Cerebras has very fast memory access for something like a matrix multiply, and it’s really good to host the MLP—the weights, essentially—on the Cerebras chip.
But the GPU has the capacity to scale to really long context lengths. You would like to put the attention possibly on the GPU and the MLP on the Cerebras chip. I believe this is what’s happening with NVIDIA and Groq.
Can you riff for a minute just on transformers? For people who, again, aren’t deeply familiar with what this innovation was in 2017, what its strengths and weaknesses are, and whether or not you think it will remain the dominant architecture, or a dominant architecture, for the future of AI?
What it did was allow us to learn on unsupervised data really effectively. Because transformers, at the end of the day, are all about taking any sequence—any arbitrary sequence of data—and trying to find patterns in that data. Critically, the attention operation, which is the headline component of transformers, allows the model to dynamically adapt to what it thinks is the most relevant component of the sequence.
Every step you take through a transformer, you are essentially reweighting the input that you looked at before and figuring out which is most relevant for your next prediction. It’s extremely amenable to learning arbitrary sequence data, and the most interesting sequences of data that we produce on a regular basis are language. That’s how we got to dominance in the language regime.
To zoom out even further, I think what transformers really did well is that they scaled. Transformers make no such human prior. Transformers just say, “Well, there’s going to be a pattern in the sequence of data, and if there is a pattern, I’m going to find it. I’m going to throw more and more parameters at this problem until it works.”
Transformers benefit from a lot of the computer vision work, too. For example, one of the challenges in computer vision was that we had a hard time going from hundreds of thousands of parameters, which you get for linear models like support vector machines or other legacy machine learning models. Those had thousands of parameters. Then we got to deep learning and got to tens of millions of parameters with computer vision.
The biggest models were 150 million parameters, which was a huge model for computer vision. Now we routinely talk about trillions of parameters, and transformers are the link to go from millions to trillions of parameters.
If I think about the important units of scaling being data and compute, does it stand to reason that you think transformers will just stick around because those are the things that we’re good at getting more of?
Transformers are such great sponges. You increase the compute available to a transformer by 10×, and you’ll get some logarithmic improvement somewhere. So far, the scaling laws really work. They’re really quite beautiful.
To the point about what transformers do really well, they extend to almost any dataset you can throw at them. They’re extremely powerful general learners. I think what’s especially useful about transformers over other techniques that we’ve tried to use to replace attention is that transformers represent any pairwise relationship that you want. Any token in the sequence can attend to any other token in the sequence.
So if there’s any relationship that’s in the sequence at all, you’re going to find it with a transformer. It may be the case that you don’t need wall-to-wall modeling. You don’t need every token to look at every other token. But if you need to, transformers give you that option.
Until we know a better way to prune that space down, a better way to have information modeling be more selective, attention is a very, very good operation. This is another trick that we learned in the computer vision days.
One of the old Karpathy sayings is that if you have a new dataset that you want to train a model for, your first goal should be to overparameterize the model and try to overfit the data that you have. That proves there is a relationship that you can model or memorize, that your learning algorithm works, and that you can instill knowledge into the model.
Once you can overfit, then you can compress, and the compression is how you get generalization. You don’t want to actually memorize the data that you have in front of you. You want to generalize. Therefore, once you overfit the dataset, you can work backward and try to find the general patterns that fit into the smallest parameter count possible.
What’s your prediction for the future of data, and riff on the importance of data in this whole story?
6. The Next Data Frontier
I like the phrase that the Internet was a one-time subsidy on data. We got it for free. It’s extremely high quality. There are about 30 trillion tokens of high-quality text, and 300 trillion tokens if you take a wider view on what qualifies as good text. We’ve basically looked at it all already.
Models have seen the entire Internet many times over at this point, and there is not a whole lot more to be done on human data from the Internet. The next phase of data, in my mind, is model self-improvement through RL environment gyms, basically.
In fact, we don’t even benefit from getting more random user interactions with AI anymore. It used to be that the new type of data we cared about a lot was the interaction data from people using ChatGPT and giving ChatGPT signals on what they liked and didn’t like.
I like the argument now that the median model that we serve is so much more advanced than a random human giving feedback that the signal you get from random human preference—unconditioned human preference—is not actually worth anything anymore. You want expert human preference at this point.
Everyday Joe.
Yeah, everyday Joe, exactly. So the future of data, to me, is giving the model a hard, verifiable task and letting it run in this gym, where it’s isolated and just has a problem that it can make progress on, along with a measurement of whether it made progress on that problem or not.
You can imagine coding problems are in this category. Math problems are also in this category. More and more, we can just give the agent a computer, essentially, and have it act like it’s a human worker and give it feedback on whether it’s making progress toward the target outcome.
That environment becomes the data. I think this is not a super-differentiated take, but it’s been really, really productive from what I’ve seen so far.
Do you think that just goes on for a really long period of time, or is that another—if I think about the Internet as this one big block—is this another big block that will have its day in the sun, and we’ll kind of get it all—
Yeah.
—and then we’ll have to move on to something else?
I think it’s actually more profound than that. Basically, the idea is that if you want artificial general intelligence, the best way to get there is to just keep stacking specialized intelligences until you have no more gaps to fill.
The only thing you need to make sure you do to make this work is ensure that your task is verifiable. You need to give the model a self-grading system. If you have that, you have the recipe for self-improvement on any task you like.
I think you’ve seen this held up by the way frontier labs spend. They used to spend that much on data. Now they spend a lot more on RL environments. These environments absolutely capture that relationship of recursive self-improvement on a verifiable task.
Coming back to your initial task of making existing hardware more efficient by being more in control of what’s going on at the hardware level through—
Yeah.
—software, keep going on what you’ve done so far and what you want to do, and then we’re going to jump to hardware and then jump to energy, finally.
7. Automating Kernel Engineering
Sounds good. I mentioned kernels. It’s surprising. People think kernels are done. There are great people like Tri Dao who write excellent kernels, and they form the bedrock of all our modern deep learning, which is built on FlashAttention. Modern transformers are built on FlashAttention.
But if you deviate from the happy path at all—if there’s a new model that comes out that has a slightly different way to embed positional information, like a change to the RoPE system—suddenly, the kernel that we had is not suitable for this new model, and we may have to make a patch to this kernel.
I wouldn’t say we’re in the phase where we have to invent new kernels from scratch, but having the ability to quickly modify existing kernels is important. A kernel, by the way, is a general term for any program you run on the GPU.
Historically, kernels tend to be put into a library where every kernel has a very, very scoped purpose. Typically, you have a kernel for a matrix multiply. You have another kernel for something as simple as addition. If you want to add 2 tensors together, that’s another kernel.
Increasingly, we’ve started to fuse those kernels together. If I do a matrix multiply and then I want to add it to another matrix that I’ve also multiplied, maybe those 2 become 1 kernel. I just fuse the operations so that instead of writing the data out to DRAM and then reading it back in just to do the addition, maybe I can do this easily.
Why are humans still doing this? It seems like the sort of thing that AIs would be exceptionally good at: engineering more efficient kernels.
Maybe that's where we're going, and we're just not quite there yet. But if we aren't there yet, is that where we're going? If we're not there yet, why are humans still doing this? Why is Tri Dao so well known? You know, it's a name I know.
I don't want to speak for Tri, but what he taught me was that you shouldn't necessarily write kernels by hand anymore. I like to say we write kernels on the whiteboard. We go to the whiteboard and describe what we think the machine should be doing. Then we succinctly describe that in natural language to a model, and the model is able to execute: “Here is my input and output. Here is the strategy for how we want to dispatch this work onto the GPU. I'm gonna go implement it.”
So we're doing the conceptual design.
Exactly. I'm not sure exactly why models are not superb at doing this. I don't think this is our moat or anything like that. I'm sure in 6 months' time we'll have much better models for kernel engineering, and I'm sure the labs would tell you that they already do a lot of their kernel engineering in a fully automated way.
So software as an edge, if I think about software as maximally near-speed-of-light-efficient usage—
Yes.
—of an underlying piece of hardware is going to trend toward not being an advantage for a company like yours over time.
I think that's right. The rising tide of something like Mythos or GPT-5.6 Sol lifts all boats. It really does. I actually don't think there's a point in specializing to say we work on making the model better for just kernel engineering. I think that's not the most meaningful subset of coding in general, and kernel engineering in particular. Maybe there's some privileged information you inject into the prompt that's a useful way to steer the model to be better at writing kernels. But broadly speaking, yes, we're all downstream of the frontier in terms of this capability.
I always love this from the history of energy. There's always this pendulum between the raw source, let's say coal, and then, if there's a certain amount of energy available inside a hunk of coal, what percentage of it we can harness and use. A big part of the history of energy was getting that number from 10% to 95%, or whatever.
If I just think about a Blackwell or something, and Blackwell is the piece of coal, what percentage do you think we're at? How efficiently can we use—
Yeah.
—an existing piece today?
There are a lot of different ways to analyze that. At some level, we are really efficient at optimizing performance when the GPU is doing the thing that it's most happy doing, which is a large-dimensional matrix multiply. That operation runs at 70% to 80% of peak utilization, and it's limited not by software but by power. The way NVIDIA quotes peak FLOPs is a little optimistic. You never hit that because of power throttling.
Because of heat.
Exactly—thermals. In practice, you don't spend the majority of your time in a transformer in that happy path where you're doing a large-batch matrix multiply. Our job is to basically build the engine around the chip such that we are feeding the GPU these large batches of work at all times.
One of the most profound transitions we've had in the GPU world in the last year has been this move away from programming one GPU at a time. You should think about the whole rack, and maybe you should think about the whole cluster—the entire data center—at a time. With NVIDIA, they've started shipping not just a single GPU or a single motherboard, but actually the whole rack system is something that they prescribe. They call it NVL72. Their latest chip, the GB300, ships as a rack of 72 units, and it's an open race to figure out who can program the whole rack-scale computer as efficiently as possible.
Mm-hmm.
And my belief is that this shape of compute is the future of efficiency and speed. In fact, NVIDIA does a great job of saying that if you want the lowest possible latency, you should be using that chip, and if you want the highest possible throughput, you should probably also be using that chip as of right now. It all comes down to this being a very new paradigm of programming.
One of the things you hear is that the market for the best chips—Blackwells, let's say—is like a drug market or something right now.
Yeah.
There's all sorts of fascinating things happening to get as many of them as possible because everyone's so short. I would love you to react to that analogy. Is that what it feels like? And then also talk about what the market is like for not the bleeding-edge chips. If I am willing to accept a slightly or moderately inferior chip, what's that market like? Let us into that world.
8. The Scavenger Strategy
NVIDIA has a long-term view on all their chips. They see this immense demand for the Blackwell chips. They could do what other suppliers have done in the past, which is crank up prices to meet the market; supply and demand curves will correct. They'll intersect at some point, and everyone will be technically happier. But NVIDIA sees that if they just let the deepest pockets buy all the chips, that maybe hurts them in the long term if the customer ends up accruing a lot more power. They understand that compute is power today. They're quite strategic about how they allocate compute.
That's the first thought. The second thought is that relationships matter a lot. Nobody wants to have a huge order for chip rentals come in from this new startup that says, “Oh, yeah, I'm gonna rent 10,000 Blackwells for 3 years or 5 years.” The startup has only been operating for months, typically. Who knows that they're good for the money?
The way you convince someone to give you access to compute is quite challenging these days and requires some pretty great relationships or just incredible financial backing to make this happen on the NVIDIA side. It's all because the scarcity is so high and demand is just off the charts.
Now, for other chips, I wouldn't even call them inferior. I like to say there's no bad chips. There's only bad pricing. I will make any chip work at the right price. That's one of the company's ethoses.
Let's talk about AMD. AMD has great chips overall. The challenge is that people don't understand how to program them very well. I've been talking to you about how we have such a great kernel team. We're so serious about squeezing the performance out of the hardware. NVIDIA is pretty good at doing that for their own chips, frankly. There's some alpha that we can squeeze out, but there's a lot more to be done on other chips because the vendor does a little bit less work than NVIDIA does to make the best kernels out of the box.
Or even better, there's alpha because other people have this perception that AMD is not as good as NVIDIA. That's music to my ears. I'm very happy for them to sleep on this chip and for me to buy as much as I can. Now, I think that that's not actually super true anymore. I think AMD is actually somewhat popular among some large buyers. Publicly, Meta and OpenAI have bought a ton of AMD chips. We're increasingly seeing that all of AMD's supply is also being allocated.
But there's a long tail of other companies that are popping up. Net-new companies are great, such as Etched, SambaNova, or d-Matrix. All these companies are popping up. I think the main challenge for them is scale. Can they actually get enough wafer allocation from TSMC to pump out chips and make it into the market? If there's a new chip on the market, I'd like to know about it as quickly as possible and evaluate whether we can buy a good fraction of that supply.
And so it's fundamentally an arbitrage for you. If you can be much better at eking out performance from chips that have received less attention, you can then resell that at a margin, and it can be a great business.
Exactly. And I don't think it's the case that everyone else has a skill issue—that they can't make these chips work as well. I think we're quite competent at this. I think we're probably one of the best teams in the world at using multiple silicon architectures and being quite aggressive in chasing down performance in unlikely places.
But, yeah, I think it's the speed at which we're willing to build our stack around a new chip. We don't have a huge amount of incumbency around, “Well, our data center providers are only stuck with this class of chip, and it's gonna be a huge pain for us to deploy these net-new chips.” We have some very creative data center partners who are willing to move very quickly, and there's a new class of those that we can talk about.
Most importantly, we don't shy away from the challenge. That's frankly a big part of this: just saying, “Yes, we love TPUs. We're gonna make TPUs work. Yes, we love Trainium. We're gonna make Trainium work. And if it doesn't work easily, we're gonna find a way to fit it in with a heterogeneous serving system. It will have a place.” Every chip has a comparative advantage. We have to find that advantage and then squeeze it in that direction.
Just as an interlude before we get to hardware, data centers, energy, et cetera—which will be a really fun part of the conversation—I’d love you to talk about your perception of the investor class's worry. You look at memory stocks, or my current favorite is the chart that plots the percentage of the S&P 500 that's semiconductors.
Yeah.
Historically, it was 2%, 3%, 4%.
Right.
Now it's 19%, 20%, 21%. And it just looks like, if you're a student of market history, you get all these things through time that have reached some crazy near-term peak and then collapsed back to long-term norms.
That has all investors worried. A lot of people have made a lot of money in Micron and SK Hynix and companies like this, but everyone feels like, long term, compute’s a commodity. It will not represent a quarter or a fifth of the entire market capitalization of the world.
They’re scared, and that’s the setup. Everyone acknowledges that there’s a huge shortage. Everyone feels like, “Ah, we’ll figure it out, and these things will revert back down to their normal place in capital markets.”
Right.
I’m curious what you think about that narrative.
I’m less of a student of history and more of a member of history. I was born in 1997, and my mom worked at Intel in the run-up to the year 2000 and the dot-com boom and crash. I remember the time when Cisco was the most valuable company in the world, and Intel was close behind.
I mostly draw parallels to that period of history from 25 years ago to today. I think the main difference is that a lot of the investment in networking equipment historically was speculative. We anticipated this future demand for users that never came.
I think what’s interesting about token consumption, or AI consumption broadly, is that it’s no longer speculative. People buy tokens because they’re immediately valuable to them. You don’t hoard tokens; you use them immediately.
This is also different from what we had 2 years ago, when there was a supply crunch for Hopper-generation chips in 2023 and 2024. In that period, it was all training-oriented spend, and training is inherently speculative. Now everyone’s instituting caps on how much you can spend on Claude Code.
It’s a very, very different world to be talking about inference spend and predicting inference spend to go up. I do think inference spend monotonically increases. There’s no speculation on inference.
Coming back now to your take on hardware, the unit level is interesting to me. We’ve talked about chips, racks, and clusters. I’d love to talk about data centers. You said you’ve had some interesting partners doing some cool things.
Yeah.
Talk to us about the present and future of data centers as you see it, because this seems like an obviously critical thing for being able to serve all this inference. There’s lots of innovation in this part of the world, and obviously you’re focused on it.
Yeah.
I think one of the themes in our conversation has come back to what training versus inference is and what the difference is between these 2 categories. What was different about being training-oriented 2 years ago and being inference-oriented today?
I think the most conservative players in the entire AI stack have got to be the infrastructure players, whether that’s data centers or, even more conservative, TSMC—the chip-infrastructure people. Data centers were built for training. Training is a superset workload over inference. You can make any training cluster work for inference, but maybe not vice versa.
The difference there is networking: how much do you invest in bandwidth between chips, and how large of a cluster do you need? There’s actually a diseconomy of scale to data centers in some way. It’s way more expensive and difficult to build 100,000 GPUs in 1 data center than it is to build 10,000, than it is to build 1,000.
Now we just talk about how many megawatts or gigawatts you have. Basically, there’s no way to build a gigawatt data center in the United States easily anymore. Even 100 megawatts is increasingly hard. It’s basically impossible unless you’re a very special set of customers. Ten megawatts is probably on the edge of what’s possible today, and 1 megawatt, I would argue, is plentiful.
So there’s this incredible lore in the market where you can find lots of aggregate power, but it will not be concentrated. That was not interesting to anyone who was building training data centers because you just assumed—
It needs to be all up in 1 spot.
—it needs to all be in 1 spot for training. No one wants to deal with cross-data-center training. The market has some lag in it. I think the market still assumes that we had to go shake down those 100-megawatt and 10-megawatt data centers wherever we can find them.
That’s still the attitude I hear from a lot of data center developers. But increasingly, we’re seeing a few new thinkers realize that inference is going to be suitable for these distributed 1-megawatt data centers. We’re quite in agreement with that, and we’re very happy to buy small pools of compute across the United States and use that as our inference fleet.
Give us a sense of the literal physical size of—
Yeah.
—1 megawatt versus 10.
Yeah. This got really wonky with the advent of liquid cooling. Now you can pack insane levels of power density into a single physical rack. You’d imagine 1 megawatt of compute as a massive data hall, like a huge warehouse, basically.
Now you can actually pack that into around 8 racks’ worth of compute. Each rack is about the size of a refrigerator. You can just imagine 8 of them lined up. That’s 1 megawatt.
Your view would be that the future that you want to help build—
Yeah.
—is a whole bunch of different chips that can be used together that you can buy. You’re a buyer trying to eke out the most per chip, and those chips can then be coupled in very small data centers to just do inference.
Those 2 steps—a whole bunch of random compute, some of which is cheaper than it should be, and your ability to eke more out of it, and then small units of inference in a data center—equal way cheaper intelligence.
I certainly think so. Yes. There are a lot of ways to access cheaper FLOPS if you’re able to be creative with what you take. One of the ways that I describe what we do is, we will buy any chip anywhere in the world for any duration of time.
That is a level of flexibility and liquidity that no one else has right now. We’re very aggressive about putting our money where our mouth is, and we will take any capacity and find a way to make it work in our fleet. That’s a big part of our advantage today.
Long term, we have to create more of that advantage by investing in these data centers that other people are going to be skeptical of. What’s going to happen when you set up this army of 1,000 small data centers versus the 1 big gigawatt data center?
A few things: you’re not going to have power redundancy. Quite often, you’re not going to have backup diesel generators on-site. Those are all very expensive, and we cut all that overhead.
We’re not even going to have redundant networking in a lot of cases. We’re going to put these in facilities where we have good access to power, a single source of power, and we’re going to trench 1 line of fiber to these data centers.
But we’re not going to have 3 lines of fiber with redundancy, failover, and SLAs. It’s just going to go down sometimes. I won’t be surprised if some of them get down to 95% uptime.
Which is bad.
Very bad. That’s fatal, atrocious for anyone else in—
You couldn’t survive in a big gigawatt center.
You’d have basically 0 buyers for a data center that has 95% uptime. I’m that first buyer.
Yeah.
I will buy 95% uptime.
And the reason for that is because of this background-agent thing? There are things running in the background, and you don’t care?
Partially. It’s actually 2 things. One is that we have a really robust control plane that’s going to be fine handling any single failure in any single data center, as long as it’s not correlated with other data centers, and I can just move the workload somewhere else. I’m cool with that.
The failures happen at some rate, and I’m basically linearly happy with a data center that has 95% uptime versus 98% versus 99%. It’s just linearly good or bad for me.
You do need that async piece that I mentioned. We serve these long-horizon agents because what happens when a request fails is that I’m going to have to go find a new GPU to put that request on.
That means that, for that single turn of the agent’s work, it’s working for an hour, but then it hits a roadblock because its GPU got pulled away. At that moment in time, that agent is going to experience maybe an extra minute or 2 or 3, maybe even 10, of latency.
My argument is that my customers don’t care because their agent was running for hours.
They’re sleeping.
It doesn’t matter. It doesn’t matter if a single turn occasionally becomes a little bit longer. We tell our customers, “Look, our average throughput is going to be very competitive, but our P99, our 99th-percentile latency, is not going to be controlled. It cannot be. In return, I’ll give you unbeatable economics.”
I think that’s the right fit for background agents.
Talk about power as a category—what the innovation you’re seeing is and where you think it goes from here. What are you seeing that’s interesting and innovative? Where do you think this goes?
I said I want 95% uptime on my data centers. Could I even take 80% uptime at the right price? Probably.
And what does that mean? I’m a son of California. I love solar and wind. I think solar and wind power is way undertapped in the United States, and the challenge has always been this intermittency.
You would even consider solar and wind unsuitable for data centers because you have a persistent base load and an intermittent power source. What are you going to do? I think we’re actually not that far from solving that problem.
I am totally capable of tolerating an outage from a data center that’s measured in even days or weeks, which is the worst-case nightmare scenario for a data center—that we’re going to have a long-term outage because the wind doesn’t blow and the clouds are in the sky.
Fog is hanging over the valley for some time. That's the worst-case scenario. It's, in fact, highly predictable, and I can just call in capacity in some other place in the world whenever that happens. I'll model the weather and figure out when my data centers are going to be offline, move my workload somewhere else, and it's fine.
The trick is that it's going to give me better access to power that no one else is going to touch because it is so annoying to deal with that kind of outage. And if my chips are cheap enough, they're probably not going to be NVIDIA racks, but if my chips are cheap enough, I don't mind the capital cost of having idle chips.
I've heard you describe this entire system as a scavenger strategy.
Yeah. Yeah, that's right.
Unpack that analogy a little bit.
Well, first we scavenge chips, and then we scavenge power for those chips. The idea is, in both cases, I do not want to be bidding against Anthropic or OpenAI for compute capacity. I'm not going to win against them, and I don't want to. I want to be more creative and use the supply that they don't find legible today.
Over time, I amass enough aggregate supply. I'm never going to get concentrated supply; I will only get aggregate supply. And over time, I build my aggregate factory that is unbeatable in economics. We are building a factory. We're trying to build the best steel factory in the world, but it will come through mini-mills, not through large, monolithic steel plants.
And if I imagine the different versions of this—how vertically integrated you can be—the extreme would be you own everything, so it's a very capital-intensive business. You own the power source. You build the data centers. You design your own chips. You control the software that ekes the most out of those chips, and you sell the finished token to your user. Your user is me, and you just own the whole stack.
But you can imagine many other permutations of the business where you draw the line anywhere. You could be incredibly capital-light, own nothing, and just be the coordination plane across all this stuff—the virtual scavenger. How do you think about that question of which type of these businesses to be?
You know, there's actually two parts of me receiving that question. One is the CEO of a company that needs to work every day and grow sustainably and as quickly as it possibly can. The other is the founder, and the founder is much more imaginative and loves this stuff.
The founder in me wants to do everything. This is my entire life. I spent my entire life thinking about chips, power, and energy. All I care about is this stuff, so of course I want to be maximally ambitious. I don't want to stop ever. I will never stop until I have built the most efficient system from soup to nuts.
You're doing real-life Factorio, basically—
Very much so.
—for intelligence.
Very much so. So that's the emotional, from-the-heart answer. On the CEO side, I think we have to be more pragmatic. The capital we're looking at for owning everything is insane. Software has high leverage, so we have to start with software.
But ultimately, do we own power generation, or can we get great power-purchase agreements with utilities? I'm more inclined to pursue letting other people specialize in the things that they're historically good at and then see if we can get to the scale.
I think of it as, like, I want to get to the scale where I earn the right to take this under our wing. I absolutely think that there's efficiencies to be gained everywhere in the stack if you can break the assumption that the people I'd be buying from made assumptions about who their customers would be, and I maybe break those assumptions.
It's a pretty optimistic view. I think it's only possible because we're actually trying to underwrite the largest market for compute in the history of computing. We're actually going to build so many millions, trillions of dollars of investment into inference. Because of that focus, it makes sense to build a lot of things that are custom for inference, and it's my job to seek all the places where that's possible.
And then, as they become obvious to me and my partners, I'll look to my partners to build custom things for me, and if they can't do it for me, I will do it myself.
If you had to just zoom out on this entire system—
Yes.
—software, hardware, energy, et cetera—and stack-rank the places that you think we're the most inefficient today at producing useful, intelligent tokens, what does that list look like?
I think compute scaling is actually very efficient. You give me more FLOPS, and I will use more FLOPS. I would say we're actually fairly judicious already with our use of FLOPS.
If you look at a modern MoE model, there are very few models that are more than 10% dense, meaning 10% of the possible number of experts you can activate are activated, and I think the frontier models are closer to 1%. They're fairly sparse already. I don't think that we're wasting too much on the MoE side. People have been working with MoEs for quite some time. They're pretty good at squeezing MoEs.
Where we are not good is attention and its use of memory, specifically. The KV cache is quite uncompressed right now. I think if you look at the entropy in a KV cache, it's not earning its keep. We're storing many kilobytes of data in the KV cache per token, and that's probably off by an order of magnitude or two.
I don't know what the frontier labs do, but DeepSeek certainly publishes really interesting work to compress that further and further, and they're making good progress. And I think the fact that they're able to make an order of magnitude in progress here every year or so signals that there's a lot more room to go.
If you zoom out further, I think that we actually don't marshal our compute effectively at all. We have all this compute in the world. NVIDIA's pumping out 5 million Blackwell chips this year. Where are they all going? Are they all being used all the time? I certainly doubt it.
I think that, at some level, we just need better orchestration of compute across the world. This is very difficult to do because a lot of the compute disappears into private pools of compute that will never see the light of day, and those GPUs sit very sadly idle.
It pains me physically to see that. Silicon and power went into those GPUs, and they're just sitting idle, and I want to fix that. How we organize and orchestrate the world's compute as a shared resource and pack it more efficiently—I would estimate that we all make fun of xAI for having some challenges with total FLOP utilization on its clusters. The reality for the rest of the world is far worse. A ton of GPUs just sit in warehouses or sit in private pools allocated to a specific customer and don't get utilized.
You're attacking the efficiency of that very directly.
That's way more effective, yeah.
What about fabs? What do you think is the future of fabs themselves? I think everyone is wondering: Will the memory companies, TSMC, Intel, and others expand capacity? Will we do it here in the U.S.? Riff on fabrication of chips themselves. If we could just snap our fingers and have 100 times the chips in stock today, we'd probably have way cheaper tokens.
Yep.
That seems like an important part of the universe to hear your view on.
Everything grows in balance with each other. If we snap our fingers and double all those things, you might fix a TSMC bottleneck, but you're just going to run into another bottleneck. You make 20% more chips, then you have another bottleneck immediately.
I will say, though, it is interesting what they consider to be a must-deliver, like an invariant that their customers—me—are always going to want, versus what I think of as a more fluid relationship. I think that if a fab exposes more of its trade-offs to me, I'm able to make more intelligent decisions about what I can do.
One of the most interesting examples here is that any fab has a lot of spread between the worst chip that comes out of the production line and the best chip that comes out of the production line. There's a lot of variance in how chips are made.
And the question is, if you have a company like TSMC, they work very, very hard to tighten what we call these process corners. We want to keep the worst chip as close in characterization to the best chip as possible, and they go to great lengths to make that possible. But that means that they are adding a lot of controls in the process that maybe I don't need.
Maybe I'm actually willing to find a place for that worst chip. You don't need to tighten the process control as much, which takes more time and costs more. Maybe I'm willing to take a lot more rejects.
And I think for us, it's a more holistic optimization around the cost of the dies, the supply of the dies, and then the cost of power and the places we can put them. My whole goal is to so dramatically expand the supply of power across the United States that I have a home for a lot of chips that otherwise would not have earned their place in a data center.
Can we talk about how you design the system of your own business? What lessons have you learned? You talked about some of the interesting NVIDIA lessons. Bring me into the culture and how you structure a team and a business where this is the North Star.
There's a lot of in-the-limit thinking. We don't worry about the immediate nature of when we start working on a model—the efficiency's not going to be very good. We think about where we could end up in a month, 6 months, or a year's time.
We don't accept the state of the machines we work on as fixed. Even something like the Blackwell chip—if we think that there's some bottleneck holding us back from achieving this performance, it's very important to me that we understand and characterize that very well and write it down, so we can both, A, tell NVIDIA about it, who are friends, and also basically keep this in mind for future chips that we buy.
We want to learn things that are invariant for us or the company long term and fold that into future decisions that we make. We're very collaborative. I think one of the most important traits that we look for are people who are both good students and great teachers. A lot of our people on the team were TAs in college and loved the experience of sharing knowledge in this way.
We do whiteboard sessions all the time. The collegial environment, where everyone has something to teach and something to learn, is extremely important for us.
What are the attributes of people that you would want to hire that you think will be resilient to the work environment three years from now—
Yeah.
—when more stuff is handled by machines?
Curiosity. 100% curiosity. The one thing I cannot teach is love for performance—love for digging into every microsecond the machine is working and understanding what's happening on the machine at that time. That, to me, is the most important trait for a performance engineer. It's what I look for.
I don't look for lots of AI experience. I don't look for CUDA experience at all. That's actually a huge red herring. CUDA as a concept, or GPUs as a concept, have evolved so much in the last 5 years. There's no point asking for 10 years of experience. I want to teach that, but I cannot teach the love for performance engineering. That is what I seek.
Can you give your assessment of the major labs one by one, but also then the relationship of closed source as a category to open source—what you think is happening and will happen?
9. Closed Models Meet Open Source
In a line, I would say the labs pay an immense premium to be 3 to 6 months ahead of everything else. I think that's probably still worth it. I think it makes perfect sense for Anthropic to do what they do.
There's a sensitive topic around distillation, which I think is a very core piece of the relationship between closed and open frontier models. I'd like to offer an alternative view on that. There is a sense that distillation is theft—that you are taking something from the frontier models when you distill on their outputs.
Even if that's not your intent, and even if you don't ever try to scrape data from Anthropic, one thing I'll offer is that an increasingly large percentage of the artifacts we put out on the internet are AI-generated. You just look at GitHub alone. What percentage of repositories created in the last year do we think were created by Claude Code? Do we consider that to be distillation? Because that's probably all we need.
I would not be surprised if you could train a frontier-class model only on the outputs of code you consider good on GitHub that's open source. Certainly, if we take the position that users own the outputs of their interaction with AI, and they choose to put that up on GitHub—which a lot of them do—we're going to have latent distillation for a long time. It seems fundamentally impossible to me. I don't think it's fundamentally possible to prevent the diffusion of information or model capabilities. It will happen. The question is just how fast.
And so then the question becomes: Do scaling and improvement laws hold forever, or for a really long period of time? If they do, then there's value to being 3 to 6 months ahead, and that will just last as long as it lasts.
Mm-hmm.
And they can charge a huge premium for those tokens relative to a very cheap open-source token. Is that the right way to think about it?
I think it's possible. I don't know that the premium for being 3 to 6 months ahead is going to last that long. If you look at enterprise deployments, they don't move at 3- to 6-month speed. A lot of enterprises are probably still on, like, Opus 4.6 or Opus 4.7. They don't adopt the bleeding edge rapidly. There are a lot of questions that people have around rolling out any change at all.
We're just so early in scratching the surface that I don't think there's any way to call a winner in this race. Certainly, I don't even think this is a race that can be decided ever. It's a continual process. Fundamentally, I don't think open source ever goes away. If there's a vacuum because one leader steps out, a new leader will step in. There's too much incentive and too much tailwind, too. It gets easier every day to train a frontier-class model.
So your hope of what the future looks like is what balance between closed and open? What balance between model companies doing everything because they have the advantage of owning the stack or whatever? Anthropic can do that. There's, like, the new Google will just do that or something. What do you hope the future looks like?
I want abundant tokens and diverse harnesses. I want everyone to build their own harness—every company, every user even. Make the agent yours, your own. We're not that far away from that level of customization and capability.
I want people to own their intelligence, and I want that intelligence to be customized, probably not through weight fine-tuning, but probably through more in-context learning. That's a more technical detail, but the underlying input to this abundant future is about cheap tokens. My job is to make the tokens as cheap as humanly possible. I will achieve that, and I will do it through every layer in the stack available to me.
I love the supply-side levers. I will use every chip, every source of power, and every piece of land in the United States that's suitable for this. In return, people will have the incentive to explore what it's like to have abundant intelligence. We still treat the agent as a person that's expensive to consult, and you should ask it when you have a hard question. That's not the way to think about intelligence. It's incredible that the machine can think, and we should try to get that into as many hands, and to as many people, as possible.
You sit in such a unique seat, and you have such a unique perspective on what you're trying to do to make this future a reality. What do you think are your most divergent views of the world versus your friends who are really well-informed and interested in this stuff? What ideas of yours make your friends look at you like you have 3 heads?
Most of the ideas on chips, I would say. When I talk about building custom chips and they ask me, “Oh, so what's different?” it's about sidestepping the HBM shortage and focusing on more extreme offload to other forms of memory, such as flash. I'm quite passionate about that idea.
Everyone on my team knows that I keep banging the drum around what we would have to change about the model architecture to make offloading KV cache to flash work at a much greater level. That's in the community of inference people. We have some divergent views on what you can do if you design a system around serving at 1 to 10 tokens per second, which is our whole North Star. More broadly, I think there is this larger sense around how people consume 1 trillion tokens per day. That's the world we want to create—the capability for them to do that.
What's 1 trillion tokens? Ground us in how much that is.
Yeah, 1 trillion tokens. Well, okay, at OpenAI pricing, that's at least $5 million—
Wow.
—at the very least for 5.5 or 5.6. I think the dollars is probably the most—
Yeah, that's a good metric.
Yeah. It's millions of dollars.
Yeah. So what's the world in which we consume what currently costs $5 million per person per day?
We're asking for at least 3 to 6 orders of magnitude improvement in cost per token. Get that under $5,000, and you probably have some customers. In fact, I would argue that for some size of model, we are approaching 1 trillion tokens being measured in tens of thousands of dollars. That's something that you could imagine running for a single job.
Mm.
I never will believe in that. There is always demand for intelligence in the world. The on-ramps to that intelligence are our challenge as a product community. I'm not a product person, so I cannot say I have the best vision on this.
Yeah, you want to enable those people.
I want to enable those people. I want them to never be held back by the sense that, “My free-tier users—I can't afford to give them this many tokens.” I hear that from my customers all the time. We want to fix that.
What about the inverse question? Not what you think is craziest, but what consensus thing do you think is wrong?
One of the things I keep coming back to is this question of NVIDIA. I am bullish on NVIDIA in the short term, and you should never bet against them. They're always going to reinvent themselves.
Fundamentally, I think one thing that surprises people is when I tell them, “Hey, if you look at Hopper to Blackwell to Rubin, and you compare like for like, what is the performance per watt of bfloat16 multiply? It hasn't improved all that much.” Or you take that one step further and go to TSMC. If you look at TSMC 5-nanometer versus 4 versus 3 versus 2, the performance per watt on these chips doesn't change by a dramatic amount.
The consequence of this is that people lose their minds over geopolitics—what would happen if we lost access to TSMC for any reason? My contrarian take is that it wouldn't be that bad. Supply would take a shock, for sure, but the best processes that we have in the West, like Intel, are not that far behind—at worst, maybe 2× worse performance per watt. The gap is just far smaller than you would make it out to be if you follow the chip-war dialogue.
What else is happening in the AI world that is not in your path—meaning it's not a component of this whole system that you would end up doing something in—that interests you most?
Well, we're fully downstream of models. The model people get to decide how to design their architectures. I don't have any input to OpenAI or Anthropic, but I can only pray that they go in a direction that's amenable to me, or I have to do my best to predict where I think they're going to go and build my serving architecture accordingly, both software and hardware choices.
They have, I think, the most interesting game to play in some ways. Once again, this goes back to the profundity of machine thinking and how consequential it is to decide to use something like sparse attention versus dense attention, or how consequential it is to use a different data type. We were training in bfloat16, but now we can train in FP8 or FP4 or lower-precision data types. That is just an arbitrary choice, it feels like, but it has profound implications for what chips I can use and how I should build my hardware and think about the future of compute.
If you had 100 entrepreneurs in a room, all of whom wanted to create some new compute startup, and let's say they specifically wanted to make hardware—chips or systems or racks or whatever—what advice would you give them on how to orient their companies, or the type of company? Not the specific choice they're making on a tech bet or something like this, because it seems like we're going to try everything, and that will be great for the world. Some stuff will work.
But if you had to give them advice on how to orient their business to be successful in this coming world, what advice would you give them?
It's all about the bottlenecks of the supply chain. So you need to first convince me, or convince an investor, that you understand the 3 to 5 bottlenecks that dictate modern chip supply. There's TSMC wafer capacity, there's HBM capacity, and there's advanced packaging. Maybe a fourth one would be power. Where will you get the power? How will you build these racks?
You should have a great answer to each of those 4 bottlenecks and how you're going to work around them, because it's all arbitrage at the end of the day. You're building a chip because you think that Nvidia has made some choices that are difficult for them to change, which is true. Nvidia makes a lot of choices that are difficult for them to change. They're not perfect; they're just really well-balanced.
You want to be spiky. You want to pick something and say, "I think they've underpriced the impact of how short we're going to be on HBM. We're going to push really hard in this other direction instead," which I do think is probably the thing to attack most.
Why?
There's no easy way to bring on a lot more memory fabs.
So it's just going to be a while until we have—
It's going to be a while.
Yeah.
Yeah. The boys in Boise don't love huge CapEx for something cyclical. They've been burned on that many times.
Conceivably, because of that shortage, the world is just going to route around it by making everything else in the system more efficient?
I think they're going to make everything else more expensive. I think that iPhones will cut their memory. iPhones are going to go up in price, and we're just going to deal with it.
Why doesn't Nvidia go all the way to the end and sell tokens, do you think?
Nvidia is really smart about this. They don't compete with their customers. Nvidia takes the long view on everything. Why don't they even start with the neocloud? Why don't they just sell compute out the back door?
Well, Nvidia is really good. Jensen is really good at making his friends billionaires. He's made CoreWeave a multibillion-dollar company, and there's no need for him to destroy that goodwill. He wants to create a diverse community of neoclouds and inference providers who are all jockeying to create demand for Nvidia, such that if any one of them decides to, I don't know, vertically integrate or go with AMD or any other option, he's got 3 more people hungry to fill that position.
Mm-hmm.
It's great to have competition among his buyers.
My favorite closing question for everyone is: What is the kindest thing that anyone's ever done for you?
My immediate first thought is all my mentors that I've had over the years. It's a rare person who takes a lot of time out of their schedule and makes it their personal interest, essentially, to make sure that you understand something, teach you something, or ingrain some value in you that they think you're on the cusp of understanding, but just push you over the line to understanding.
A lot of the people at NVIDIA that I mentioned earlier instilled that love of performance engineering in me. But also my professors in college, whom I remember. My advisor in sophomore year—I was a very impatient student—so I showed up at his office hours and said, "I want to build AI chips. I know what I want to do. Why am I wasting time taking all these other basic classes in networking and operating systems?"
He just laid out basically the whole stack and showed me the beauty of understanding every piece in the puzzle. He took my entire path of trying to focus on one piece of the system and said that it's so rare that someone can actually understand the entire stack, from gate-level silicon all the way to building a great Internet-scale service.
You should aspire to be someone who, over the course of your lifetime, achieves that level of understanding. It is such a rare trait. That level of expertise is so noble to chase. That stays with me quite a bit.
Well, amazing conversation. Thanks so much for your time.
Thank you so much for having me.