[BidClub_]
20VC · · 81 分钟

Jonathan Ross,Groq 创始人兼 CEO:NVIDIA 对决 Groq——训练与推理的未来 | E1260

Harry StebbingsJonathan Ross

YouTube
TL;DR
  • Groq 的 headline number 被误读了:Jonathan Ross 表示,“我们没有融资 15 亿美元——那是收入”,规模约为 OpenAI 收入的 30%。 Aramco/沙特交易方承担资本开支,Groq 负责部署并按约定 IRR 偿还,之后再切换分成——“我们不再受资本限制”。公司目标是在 2027 年底前至少拿下全球一半的 AI 推理算力,并坚持一个信条:“当你的增长速度超过指数增长时,赚多少利润都不重要。”
  • Ross 拒绝把 Groq 与 Nvidia 定义为竞争关系,反而将训练视为“已经解决的问题”,并乐于让给 Nvidia——他甚至告诉自己的客户:“能买到的每一块 GPU 都买下来”。 Groq 接手 Nvidia 70-80% 毛利业务之外的低毛利(前端约 20%)、高规模推理需求。其宣称的经济性包括成本低逾 5 倍、每个 token 能耗仅为三分之一,以及“最新 GPU 光是内存成本,就高于我们每颗已部署芯片的全包资本开支”。
  • 结构性突破口在于:Nvidia 是 HBM 的买方垄断者——只有 SK Hynix、Samsung 和 Micron 能生产 HBM,而 Groq 的 LPU 完全绕开 HBM。 LPU 采用与手机芯片相同的硅工艺,因此“我们的扩容实际上几乎没有上限”:2024 年初投产 640 颗芯片,年底达到 40,000 颗,今年目标超过 200 万颗,几乎需要占满晶圆厂产能。
  • Ross 最大的宏观担忧是 3-4 年后出现电力危机:潜在电力需求约 20GW,而全球现有数据中心约 15GW,当前正在过度建设。 一旦亏损的建设者收缩,而芯片数量仍每 18-24 个月翻倍、推升芯片驱动的需求至 120-240GW,“电力就会变成硬瓶颈”。与此同时,大量“假数据中心”——只有地产、没有发电机和水的项目——正在虚增表面供给。
  • 关于 scaling laws,所谓渐近线只是把所有数据视为同等质量的结果;LLM 生成并筛选自己的合成数据后,可以突破这条表面渐近线,再与测试时推理叠加,形成“几何级增长的改进”。 算力只是“软瓶颈”,DeepSeek 做的是算法改进;Nvidia 在 DeepSeek 后下跌 15%,只是“市场的 popularity contest……与 weighing machine 毫无关系”。
  • 对泡沫的判断,两半都成立:“我可以保证会有大量资金被烧光,但我也押注最终赚到的钱会多于投入的钱。” 本轮的新鲜之处,是凯恩斯的选美竞赛“已经彻底失控”:多个竞争者手里都握有数十亿美元,资本不再决定赢家,“真正拥有最佳产品的人最终会赢,因为所有人都能获得资本”。
  • 这套定位论,是在长达 7 年没有产品市场匹配的经历中磨出来的:“你的工作不是追随浪潮,而是为浪潮做好位置。” Groq 靠“Groq bonds”活了下来:80%的员工用薪资换股权,其中一半降至法定最低工资。他对下一轮浪潮的判断指向 4 类定义时代的公司:解决幻觉、拆解 agentic 子目标、实现“发明”,以及替人做决策。
  • 中国真正的短板不在 Blackwell 是否能拿到,而在审查制度:云端租赁以及马来西亚/新加坡“wink wink”的 GPU 部署,可能让物理接入不再决定胜负;但如果中国不允许“更开放、更真实的模型”,就天然处于劣势。 对 Nvidia 下一个 10 年,Ross 也没有单边判断:“如果它变成现在的 3 倍大,我不会惊讶;如果它仍大致维持原状,我也不会惊讶。”
摘要 · 为研究而整理的核心内容

1. 合成数据延长 scaling

  • Ross 对 scaling laws 讨论的修正是:OpenAI 的曲线假设所有训练数据质量相同,但事实并非如此。他用训练孩子作比喻——先问“1+1 等于多少?2×3 等于多少?双曲正切函数平方的二阶导数是多少?”——这正是模型的训练方式:先吃互联网残渣,把高质量数据留到最后。解决方案来自 AlphaGo Zero:让 LLM 生成合成数据,删掉错误部分,再重新训练,“不断往上走”,因此“真实的 scaling laws 并不像这些渐近线”。
  • 合成数据胜过真实数据的原因在于:“Reddit 很好,但质量未必比得上和一个博士交流。”更聪明的模型能够生成更好的数据,而离线筛选又能让留下的数据优于生成它的模型本身,形成棘轮效应,而不是撞上天花板。
  • 数学上仍有一个下限:Big-O 复杂度。LLM 很难完成大数乘法,因为乘法不是线性的,必须经过中间步骤,这是“数学上的必然要求”。训练购买的是直觉,即 system one;推理是叠加其上的算法,即 system two;两者结合,便得到“多项式线性……几何级增长的改进”。而直觉非常吃数据:3位数乘法需要的样本量是2位数乘法的 10 倍。
  • 对瓶颈,他拒绝单一化描述:瓶颈同时包括算力、数据和算法,但算力是“软瓶颈”,也是最容易加杠杆的一项,因为它可以替代和调配;更多算力能够压过较弱的数据或算法。DeepSeek 并没有证明算力不再必要,它只是做了算法改进——“把答案写进一个框里这种看似愚蠢的事情”——从而降低了训练数据的生成成本。

2. Groq 围绕 HBM 设计

  • Groq 的创立押注来自两个观察:在 Google,每个新模型最终投入在推理上的算力都是训练的 10-20 倍,“推理始终是关键基础设施”;同时,AI 的进步速度超过 Moore 定律,因为芯片数量也每 18-24 个月翻倍,增长是 4 倍而非 2 倍。因此 Groq 按无限扩容设计:所有模型参数都放在芯片上,让计算像流水线一样流经 600 或 3,000 颗芯片;GPU 则是反复运行这条流水线的 1/100。
  • Ross 以 Seven Powers 解读 Nvidia:它是 HBM 和 interposer(可能是 CoWoS)的买方垄断者。只有 SK Hynix、Samsung 和 Micron 能生产 HBM;HBM 成本高、扩产极其困难,而 GPU 必须使用它,因为普通内存会像“用鸡尾酒吸管喝水”。GPU 的逻辑 die 是商品化部件,采用与手机芯片相同的工艺,“Nvidia 实际上排在 Apple 之后拿到产能”;唯一稀缺的是内存。绕开 HBM 后,“我们的扩容实际上几乎没有上限”。
  • 约 3 倍的能效优势来自物理规律:芯片与 HBM 之间长而宽的连线充放电,是能耗发生的地方;片上内存只需通过短而细的连线传输。由此得到一个多数人弄反的结论:边缘计算的能效低于数据中心,就像用轻便摩托车而不是货运列车,把一吨煤运过城。
  • Groq 不使用交换机,“我们直接把芯片接到芯片上——我们的芯片就是交换机”,也不需要网络调优。沙特项目从签约到生产环境首次输出 token 只用了 51 天;相较 GPU 的波动性,Groq 的方案被描述为“100% 可预测”,Ross 用巴黎交通与火车作比喻。

3. “能买到的每一块 GPU 都买下来”

  • Ross 不接受竞争叙事:“如果你在竞争,说明你做错了什么严重的事情……因为已经有人解决了这个问题。” Nvidia 的训练能力“比任何人都强,而且强出很远”,训练已经是解决的问题;它只是没有提供快速、便宜的 token。当一次演示让客户问是否应该停止购买 GPU 时,他的回答是:“不——能买到的每一块 GPU 都买下来……因为我希望你的模型在我们这里运行得足够好。”
  • 他主动给出利润率对比:Nvidia 为 70-80%,Groq 前端约 20%,合作伙伴达到 IRR 后再进行后端分成。Groq 成本低逾 5 倍——“最新 GPU 光是内存成本,就高于我们每颗已部署芯片的全包资本开支”——而在每个 token 能耗仅为三分之一的情况下,GPU 单是运营成本就等于 Groq 的资本开支加运营成本。
  • 因此他把双方关系描述为柔道:“你几乎可以说,我们是 Nvidia 有史以来遇到的最好的事情之一。”Nvidia 可以把生产的每一块 GPU 都以高毛利卖给训练客户,Groq 则接手低毛利、高规模的推理业务。推理越多,训练需求越大,反之亦然;Groq 甚至测试过把 LPU 作为客户现有 GPU 集群的“nitro boost”出售。
  • 对 Nvidia 未来 10 年,他仍然保持双向判断:“如果它变成现在的 3 倍大,我不会惊讶;如果它仍大致维持原状,我也不会惊讶。”投资逻辑原本假设 Nvidia 会把推理也一并拿下,但“它们就是没有做出适合推理的东西”。估值倍数可能合理,但“popular­ity contest 会扭曲一切”。他还嘲讽 GTC 上“快 30 倍”的图表:那只是挑选曲线端点的规格战营销,“只要告诉我每美元能跑多少 token、每瓦能跑多少 token,其他都不重要”。

4. 15亿美元是收入,商业模式是第二项发明

  • 他希望外界记住的修正是:“我们没有融资 15 亿美元,那是收入……约为 OpenAI 收入的 30%。”结构上,Aramco 与新成立的沙特实体承担资本开支,Groq 从盈利中偿还,达到“不错的 IRR”后再切换分成;这是一种带上行参与的类债务结构,而且 Groq 从一开始就能盈利。“我们不再受资本限制……我们能赚多少钱,取决于能部署多少。”
  • 扩张轨迹是:去年在 51 天内部署 19,000 颗芯片;今年尽可能卖出全部产能,合同规模将达到“数十亿美元”;明年可提供的硬件产能达到数百亿美元——“如果按 GPU 价格计算,那会是数千亿美元,只是我们没有收那么多钱。”
  • 针对那篇认为 Groq 无法以最低价格盈利的论文,他回应称:“我们现在的 contribution margin 非常正……据我们所知,我们是唯一真正靠运行这些开源模型赚钱的人。”与此同时,竞争对手像 Uber 一样烧掉 VC 资金争夺份额。专有模型也开始以收入分成方式加入,Play AI 的语音模型已在 LEAP 上线。
  • 他反复对团队强调的经营信条是:“我们的增长速度超过指数增长,而当你的增长速度超过指数增长时,赚多少利润都不重要——重要的是先拿到立足点。”目标是在 2027 年底前拿下全球至少一半的 AI 推理算力,在维持毛利率的同时继续降价,“然后我们进入 Jevons 悖论,日子就好过了”。

5. 电力将成为硬瓶颈

  • 今天的需求信号带有回声效应:一家 hyperscaler 告诉 60 家数据中心建设者需要 1GW,“突然之间就变成了约 60GW 的需求”。Ross 盘点的实际数字是:他知道有人希望提供约 20GW 电力,而全球数据中心现有规模约 15GW,潜在供给已经超过当前容量的 2 倍。
  • 他真正担心的是来回摆动:当下略微过度建设,亏损的建设者开始收缩——“我建了这么多电力,却没人使用……我们再也不会做这种事了”——但芯片数量仍每 18-24 个月翻倍,把 15GW 的需求推到 120GW,再推到 240GW。“3-4 年后,电力会变成硬瓶颈。”
  • 供应链中相当一部分是假的,来自把数据中心当地产项目的人。行业笑话是:3 个月交付 100MW,签这里。你的 uptime 是多少?发电机在哪里——“你知道现在发电机交付期是 90 个月吧”?水在哪里?“等等,数据中心需要水?”Amazon 不会被这种项目骗到,这些所谓项目可能根本不会落地。
  • 整个生态建立在期限错配之上:模型按约 6 个月摊销,芯片使用 3-5 年,数据中心使用 10-15 年(现在的要求降至 7 年),电厂使用 15-20 年。但期限最长的资产反而风险最低,因为它最通用:“如果不用来做 AI,就拿来给所有电动车供电。”这也是 Aramco 这类主权级、长周期伙伴成为自然交易对手的原因。

6. 凯恩斯的选美竞赛已经彻底失控

  • Ross 复述的资本开支数字是:Meta 每年 650亿美元,Google“700亿或750亿”,Microsoft 800亿美元,另加 Stargate,且都包括芯片和系统。“从来没有过这样的事情”,但也从未有过终局价值如此清晰的案例:Google 当年选择保持私有,是为了不让 Microsoft 看清搜索业务的经济性;而“它一上市——就有了 Bing”。
  • 对泡沫的判断两半都成立:“我可以保证会有大量资金被烧光,但我也押注最终赚到的钱会多于投入的钱。”骗子和投机者的阶段可以预见——“AI T 恤……AI 导热硅脂……接下来你会看到 AI 公寓”——而解决办法是教育。
  • 他用凯恩斯选美竞赛解释 VC:“这会解释你需要知道的关于 VC 的一切。”你押注的是资金会聚集到哪里,而不是谁更漂亮;SoftBank 的策略就是通过出价更高来取胜。本轮前所未有之处在于:“你会看到有人融资数十亿美元,而他们的竞争者也融资了数十亿美元。”竞赛“已经彻底失控”,所以“真正拥有最佳产品的人最终会赢,因为所有人都能获得资本”。代价是人才分裂到“一个本不该存在的竞争者”那里。
  • 在价值分配上,他是 power law 的信徒:经济规模越大,结果越容易向单一主导者剧烈倾斜。因此当下反而让他觉得异常:各家 hyperscaler 的市值“靠得太近……你本来会预期其中一家已经遥遥领先。我无法理解为什么没有”。

7. 7 年没有产品市场匹配:为浪潮做好位置

  • 他的核心观点是:“你的工作不是追随浪潮,而是为浪潮做好位置——这最难,因为所有人都在劝你上岸。”如果在拨号上网时代打造 Uber,用户可以叫到车,却无法回家;今天做医疗或法律 AI,幻觉会杀死产品。但一旦问题被解决,而你已经提前站好位置,“你就会像 Groq 一样处在最佳位置”。几乎所有人都劝 Groq 不要做 LLM,而他们的反应是:“这就是我们打造出来要做的事情。”
  • 公司一度接近现金归零,Groq 仿照二战战争债券发行了“Groq bonds”,在全员会上由管理层选择展示脆弱,而不是假装实力强大。员工没有离职,约 80% 用薪资换取股权,约 50% 将工资降至法定最低水平;当 3亿美元融资的第一笔资金到账时,公司剩余现金甚至少于这批债券节省下来的金额。对他而言,PMF 的情绪变化是:“世界更明亮了,鸟儿在歌唱……我能睡觉了。”否则,创始人只能靠第三种幸福生活——“未来的幸福”。
  • 他是否怀疑过?“有过怀疑,但从未停顿。”因为使命早于公司:AI 是“最重要的技术”,如果不受约束,它会把“超出比例的控制权”交给少数人。“我们的目标是在 AI 时代保留人的自主性。如果做不到,我们就失败了。”
  • 他与 Dario 在安全问题上的分歧是:“我担心的是不同的事情……人们主动放弃自己的决策权。”他把这称为“financial diabetes”,源自父亲一次次输掉财富的经历,甚至沦落到在一座 2,000万美元豪宅外说服中餐外卖员赊账请自己吃饭。他预计充裕时代会把这种病扩散到整个社会:“如果人可以不用工作就过一生,会发生什么……我们如何让人仍然做出自己的决定?”

8. 4 家公司将定义这个时代,另有 2 个 moonshot

  • 当被问到愿意押注哪家非 Groq 公司时,他依次描述了 4 类未具名、但将定义时代的公司:首先是解决幻觉的人,从而解锁医疗和法律;其次是最擅长为 agent 拆解子目标的人。“agentic 要等解决幻觉问题之后才会出现,否则你会得到很长的链条,并在其中引入幻觉。”
  • 第三和第四阶段分别是“invent stage”和 proxy stage。LLM 写作“很糟糕,因为它是可预测的”,而发明需要“看起来不明显,但看到之后又显而易见”的输出;proxy stage 则是你开始信任模型像 EA 或 chief of staff 一样替你做决定。现在投入 agent 的资金未必会烧光:Perplexity 即使有幻觉也“运行得很好”,因为风险不高。
  • 他提出一个被视为疯狂的预测,并保留全部限定条件:自己靠 Mounjaro 减掉 70 磅后,他认为“如果有可能”显著放慢或停止衰老,那么未来大约 10 年内会出现一次 Mounjaro 时刻——突然发生,毫无预兆。Harry 而不是 Ross 说寿命可能延长 60 年;Ross 重申,“如果可能的话”。
  • 他最期待的单一方向,是把 prompt engineering 变成新的可及性浪潮:硬件需要晦涩知识,软件需要机器时间,但“语言——你已经会了,不需要再学任何东西”。如果给非洲的 13-14亿人一款只需通过说话就能构建应用的工具,“那就会再增加 13-14亿名潜在创业者”。

9. 用 Big-O 管理:300 人、problem units、反创始人模式

  • Groq“只做需要次线性员工数量的事情”:如果客户翻倍就必须员工翻倍,说明设计错了,应该自动化。最终 300 人完成了自研芯片、网络硬件和软件、runtime、编排层、编译器及云服务。他用 Walmart 与 Amazon 说明颠覆性模板:Walmart 通过门店翻倍来服务翻倍的客户,Amazon 不会把网站数量翻倍。“不要只说我需要更多人,应该关注你的业务算法。”
  • 增长用“problem units”计量:任何事情每增加到 3 倍,都会带来同样数量的问题,而管理带宽是有限的。把 LPU 从 640 颗扩到 40,000 颗,增加了 4 个 problem units;如果同时把员工数量增加到 3 倍,就又增加一个。因此团队有意保持精简,并把“人才密度”设为常数。
  • 他的反共识信念是:“我选择 anti-founder-mode,相信授权。”微观管理意味着“这个人不适合这份工作”,或者你没有让他与目标对齐。对齐是可以被设计出来的:每个人都带着一枚每秒 2,500万 token 的 challenge coin,开会时任何人都可以把它敲在桌上,说:“不,不,不,事情不会这样发展。”
  • 面对 100-200万美元的薪资大战,Groq 刻意不报最高价:“如果我们通过竞价赢下一个人,下一次有人拿出更高薪资,一切就结束了,没有忠诚可言。”受股权驱动的员工“更容易管理,因为他们以使命为导向……他们不是为了 kombucha 才来这里”。

10. 中国与欧洲走向分化

  • 中国真正领先的地方,是愿意跨越红线——“他们蒸馏了 OpenAI 模型”,这是大多数供应商不愿触碰的红线——以及粗放规模。被问到 Stargate 与中国 1,280亿美元投入的比较时,Ross 说:“如果他们想部署 150 座核反应堆,没什么大不了的,他们直接做就是了。”因此,即使芯片效率较低,也可以在本土用规模压过去。DeepSeek 是“一针强心剂”,也是“Sputnik 2.0”;但要向海外输出更难,因为全球没有足够电力运行低效加速器。
  • 决定“中国到底有没有机会”的信号是审查制度:“他们最大的噩梦之一就是言论自由……你能想象 Xi Jinping 说——国家,我们已经失去 AI 优势,我需要你们的帮助?永远不可能。”如果中国不允许“更开放、更真实的模型”,就天然处于劣势;每个科技创始人都害怕成为 Jack Ma:“如果你的专业是 AI,我会开始寻找出口。”
  • Blackwell 的获取可能没那么关键:“大多数云服务商都很乐意让你刷卡租用”,而马来西亚/新加坡的 GPU 建设则依靠“wink wink,我们不会把它租给中国”的默契;否则,“那个地区为什么会有这么多 GPU”。
  • 欧洲的解法不是 EU 所谓的 1,500 名 AI 安全雇员——“我不会浪费时间去监管一个并不存在的东西”——而是建立一个承担风险的飞地。他与 Xavier Niel 和 Station F 的 Roxanne 勾勒出“City F”:先容纳 10,000 人,再扩大到 100万人,给予特殊经济安排,并允许次日换工作。因为 6 个月的等待期“会压低工资……而且公司反正要为这 6 个月付钱,完全没有道理”。面对这会惩罚现有保险公司的公平性质疑,他回答:“没有人有权永远做 incumbent,尤其是一个懒惰的 incumbent。”
Jonathan Ross

We did not raise $1.5 billion. That’s revenue. That’s actually about 30% of the revenue of OpenAI. Your job is not to follow the wave; your job is to get positioned for the wave.

You can almost say we’re one of the best things that ever happened to NVIDIA, because they can make every single GPU that they were going to make and sell it for training. High margin gets amortized across deployment, and we’ll take the low-margin, high-volume inference business off their hands. They won’t have to sell either margin.

We are growing faster than exponential, and when you’re growing faster than exponential, there is no amount of profit that you can make that matters. What matters is getting a toehold in the market and becoming relevant.

Harry Stebbings

Jonathan, thank you so much for agreeing to do this in Paris. You look fantastic, by the way. I feel so underdressed.

Jonathan Ross

Thank you. I could take the tie off if you want, but I’ll never be able to tie it again. I don’t know how to tie a tie. My chief of staff has to tie it for me, and it’s a struggle because he’s putting it on himself and tying it.

I literally only bought this suit recently.

Harry Stebbings

You look fantastic. I don’t think I have a suit, so you’re one up on me.

I want to split the show into 2 parts. I want to talk about the landscape and where we’re at, and then I want to dive specifically into Groq and where you’re at. You’ve announced a massive new deal that I think everyone is slightly misunderstanding, which is what we were just talking about.

1. Scaling Laws and AI Model Training

I want to start with where we’re at in terms of scaling laws. Everyone says we’re at the limits of scaling laws, and then there seems to be exponential innovation happening with the likes of DeepSeek and others. Where are we in terms of the limits of scaling laws?

Jonathan Ross

Scaling laws is a paper that was published by OpenAI, and what it effectively says is that the more parameters your model has, the better it can absorb information. You’ll see these curves that they draw, and they’re amazing. You should show it if you can.

Effectively, you have these asymptotic drop-offs where you keep getting better and better, but you get a logarithmic improvement when you put a linear number of tokens in. This is why you see people doing 15 trillion tokens of training and whatnot. But they’re misunderstood because the assumption is that all of the data is the same quality.

You have a kid now, right? Eventually, you’re going to be training your kid, and you’re going to say—and play along with me here—what’s 1 + 1?

Harry Stebbings

2.

Jonathan Ross

What’s 2 × 3?

Harry Stebbings

6.

Jonathan Ross

What’s the second derivative of the square of the hyperbolic tangent?

Harry Stebbings

Good question.

Jonathan Ross

That’s how we train these models. We give them really simple problems to solve, and then we give them really hard ones. We don’t really train them up; we don’t do it smart.

Some people will train on the dregs of the internet, and then save some high-quality data for the end to make them better. But what you can do—and this is where I think everyone’s getting confused—is something like AlphaGo Zero, where it generated its own data and trained on it. You could have an LLM generate synthetic data, and when it generates the synthetic data, the data is better. You then train on that synthetic data.

Harry Stebbings

Why is synthetic data better than real data?

Jonathan Ross

Because the model is smarter. Reddit is great, but it’s not necessarily as high-quality as talking to someone with a PhD in a topic.

Just like with more expert people who are more knowledgeable and capable, if you have a better model, it generates better data. So you train the model, it gets better, you produce better data, and you produce a range of data here. You get rid of all the parts that are wrong, so now it’s the best part.

It’s a little better than the model itself because you’re pruning it, and you get to do this offline. Then you train the model and the model comes up here. You do this again, keep the better data, train it again, and you just keep moving up.

When you do that, the actual scaling laws don’t look like these asymptotes. They actually—

Harry Stebbings

But there has to be a ceiling on efficiency, no?

Jonathan Ross

There is a mathematical limit. If you study computer science, you’ve probably heard of something called Big O complexity.

Big O complexity is, if I’m solving a problem and I look at how I solve it, I might need to take more steps if I solve it with 1 algorithm versus another. For example, quicksort versus bubble sort: with quicksort, I need n log n steps; with bubble sort, I need n².

What’s the difference? If I’m sorting 1,000 numbers, n log n is 10,000 steps, but with n², that’s 1 million steps. It’s either 10 × 1,000 or 1,000 × 1,000.

One of the reasons these LLMs struggle to multiply large numbers is that multiplication is not linear. These LLMs can do anything linear without needing to think, but just like on a piece of paper, where you need to write out all those intermediate steps, these LLMs need that intermediate space and those steps in order to compute these things.

It’s a mathematical requirement. There’s nothing you can do to train a model enough that it will see any arbitrarily large number and just be able to multiply it. But you can choose bigger and bigger groupings of numbers to memorize, in which case it can do it in fewer steps.

Effectively, as you’re training the model on more and more data, it’s seeing more and more examples. Now it just has the answer for more specific situations, so it doesn’t need to do as much reasoning. But it still needs to do reasoning for some of these problems.

Harry Stebbings

What does that mean for the next step? If we have no efficiency ceiling, what does that actually mean?

Jonathan Ross

You need both. Training the model makes it more intuitive. It means that it can come up with the answer like that, with more stream-of-consciousness thinking. The reasoning part is different. Reasoning is the algorithm on top: the Big O complexity portion.

It’s System 1 and System 2 thinking, or “Thinking, Fast and Slow,” like Daniel Kahneman’s book. When you pair them together—when you make it more intuitive—you get better this way. But when you start adding in the System 2 portion, you start to get this.

2. Synthetic Data and Model Efficiency

The volume is very low, but when you do this, you get what they call polynomial—or you could think of it as geometrically increasing—improvement in the model when you combine that improved training with what they call test-time compute, or runtime compute.

Harry Stebbings

I totally get that. Just so I understand, when we think about bottlenecks, if synthetic data powers training and makes the model more intuitive—if it gets to the answer more quickly, like a grandmaster in chess seeing the right moves—synthetic data isn’t constrained in terms of its supply side.

If we think about the other bottlenecks, there’s hardware, energy efficiency, and algorithmic limits. What is the bottleneck?

Jonathan Ross

If your job is to get better at multiplying numbers, and I tell you that I want you to be able to do it with fewer steps and more intuitively, for you to be able to multiply 3-digit numbers versus 2-digit numbers, you need 10 times the data and 10 times the examples. As you get better on the intuitive part, you need more examples to train on.

Harry Stebbings

That makes sense. What is the bottleneck, then? Is it hardware quality? Is it compute? Is it algorithms? Because it’s not data.

Jonathan Ross

It is the compute, it is the data, and it is the algorithms. It’s all 3 of them.

People misunderstand the concept of a bottleneck. Compute has been less of a bottleneck and more of a soft bottleneck. When you provide even more compute, you can overpower the lack of data or the lack of improvement in algorithms. It’s not a hard bottleneck; it’s a soft bottleneck.

Ideally, you would improve all 3. You would be getting better data, better algorithms, and the algorithm improvements are going to be there. The data improvements are going to be there. But compute has always been the easiest lever because it’s so fungible. If I just give you more compute, it works better.

Harry Stebbings

Has DeepSeek not shown us that we don’t need the compute, and that you can do more with less?

Jonathan Ross

Not exactly. There was an algorithmic improvement. The algorithmic improvement, as I explained, is this seemingly silly thing where they just wrote the answer in a box, and then they knew what to look for rather than having to have a human being check it or something like that.

It was very simple, but that was an algorithmic improvement, and it made it easier to generate the data that was then trained on.

3. Inference vs. Training Costs: Why NVIDIA Loses Inference

Harry Stebbings

Can I ask about what I think are some misconceptions around compute, data—especially synthetic data—and algorithms? When you think about the biggest misconceptions people have around AI, and specifically inference, what do you think they are?

Jonathan Ross

The first misconception, which people don’t hold anymore, is that training was more expensive than inference. At Google, anytime we would train a new model, we would end up using 10 to 20 times as much compute on inference as on training.

Inference was always the critical infrastructure piece that we needed. But then, after getting past that, now everyone understands that inference is important.

Harry Stebbings

Do you think they fully do? When you look at NVIDIA’s stock price after DeepSeek, it was down 15%. If you understood the value of inference, it shouldn’t have been down 15%, with Jevons’ paradox and all that.

Jonathan Ross

I don’t agree that NVIDIA stock should have gone down because of that. I think that was a misunderstanding on most people’s part. But it also shows something else: everyone keeps saying NVIDIA stock can’t possibly go higher, and they were looking for an excuse to say, “Now that’s it. That’s why we were wrong, and we need to sell now.”

That has nothing to do with it. That’s just the popularity-contest side of the market. It had nothing to do with the weighing machine of the market.

Harry Stebbings

If a founder is building a startup today, should they build with the assumption that scaling laws will continue? Should they build with what we have today? How do you advise them?

Jonathan Ross

I would advise you to build based on things getting better, but I would also focus a little more on the big quantum steps.

The analogy I like is that if you look at the information age, we went through the printing press, the telephone, the telegraph, the internet, and smartphones. If you had built Uber when we only had the internet, it wouldn’t have worked. You’d book a ride, go somewhere, and then ask, “How do I get home?”

Harry Stebbings

Exactly.

Jonathan Ross

We’re in the same sort of space now. The models hallucinate, so it would be hard to build a medical-diagnosis company. It would be hard to build a legal company.

However, if you were doing that and the algorithmic enhancements happened that got the hallucination rate down, you would be perfectly positioned. Just like Groq, we were around for 7 years before we had product-market fit.

We were around because our bet was on scaled inference: that inference was going to be the bottleneck, and that we were going to need to run really big, heavy models. Everyone was assuming you would have a single PCIe card running inference because training was the complicated part.

The reality was that we made the right bet ahead of time, and then we were perfectly positioned. Your job is not to follow the wave; your job is to get positioned for the wave. That’s the hardest thing to do, because everyone is trying to talk you into coming ashore again. Almost everyone was telling us, “Don’t do LLMs. They’re going to be terrible for you.”

We were saying, “This is literally what we built for.”

Harry Stebbings

Did you ever doubt yourself? 7 years is an incredibly long wait time.

Jonathan Ross

There was doubt, but there was never a pause. Even before starting the TPU, I was concerned that AI was going to be a technology that would allow some people to have outsized control and outsized influence.

If you allow that to just happen in potentially not the best hands, it doesn’t really matter how rich you are. Nothing matters. It’s the most important technology.

It didn’t matter how hard it got. There was no choice but to be successful. Our goal is to preserve human agency in the age of AI. If we don’t do that, we have failed. It wouldn’t matter whether there was doubt or not.

There was plenty of doubt. There was a point where we were so close to running out of money that we did this thing we called Groq Bonds.

Harry Stebbings

War bonds from World War II, of course. For anyone who doesn’t know, what is a war bond?

Jonathan Ross

World War II was funded with bonds from the U.S. government. They had these posters saying, “Fund your troops,” and people would buy the bonds and receive a return. That funded the war effort.

We were very close to running out of money at one point. Rather than trying to pretend to be strong, we were vulnerable with our employees. We said, “We’re going to run out of money. We need you to trade equity for salary.”

We literally took pictures of the war bonds, put “Groq Bonds” on them instead, and had an all-hands meeting where we explained it. We were worried everyone was going to leave.

Instead of leaving, about 80% of the employees participated. I think 50% went to the statutory minimum salary required by law. When we finally raised the first bit of our $300 million round, we had so little money left in the bank that it was less than the money we saved through Groq Bonds.

Had we not done that, we would have literally run out of money. There were some really hard times. I know every founder has these moments, and from the outside, it’s so hard to understand. It’s like watching a TV show: you’re not in it.

When you’re there, everything is 10 to 100 times more intense, because people left their jobs and careers, and their families are banking on this. You have to make decisions like, “What would have happened if we went out there and asked everyone to do Groq Bonds and everyone quit?”

The shareholders would have said, “You have all of these people depending on you.” But if you lean toward vulnerability, people are often going to go with you.

4. The Future of AI Inference: Efficiency and Cost

Harry Stebbings

What is a world where inference is so crucial and 20 times more important than training? What does that world look like?

Jonathan Ross

The simplest way to understand it is to equate an LPU or a GPU to an employee. If you have enough of them, the LPUs or GPUs can do work, just like an employee.

It’s a little different in the sense that they can’t quit and take another job. You don’t have to retrain them. Once you get a model to a certain capability, it will always be at least that capable. It’s not going to regress, so you get consistency out of it.

Imagine that you’re a startup, and rather than having to go out and hire 100 people, you hire 10 and buy the amount of compute equivalent to 90 employees. That’s a very different way of thinking about the world, because now capital expenditure—or, in some cases, different types of operating expenditure—can be used instead of just employees.

In terms of inference, to give you a sense of our scaling, we started 2024 with about 640 chips in production. We ended with more than 40,000. This year, we want to be at more than 2 million, and next year the number is much, much larger.

5. Chip Supply and Scaling Concerns

Harry Stebbings

Are we seeing constraints on chip supply? That’s an unbelievable scaling story.

Jonathan Ross

For us to hit our numbers next year—which I’m not sharing publicly—we’re going to need almost all of the capacity of the fab that we’re using.

The biggest issue is that you don’t normally think of tech companies as having a cornered resource, but NVIDIA has a cornered resource. They’re a monopsony—the opposite of a monopoly, a single buyer—for HBM and the interposer, likely CoWoS.

Harry Stebbings

What is HBM?

Jonathan Ross

HBM is high-bandwidth memory.

Harry Stebbings

Which GPUs use it, and who produces it? Sorry for the dumb questions.

Jonathan Ross

There are 3 companies in the world that do this: SK Hynix, Samsung, and Micron. It’s specialty memory that’s only used in high-end servers.

There’s a limited quantity that’s built, it’s very expensive to ramp up, and it’s a very technically challenging type of memory to build. There’s a very limited supply, and GPUs are so fast computationally that if you were using regular memory, it would be like drinking out of a martini straw. It would take forever.

This is why you see people preferring to do inference, but especially training, on GPUs rather than CPUs, because the memory bandwidth is too limited. CPUs rarely use HBM; they mostly use regular memory.

Our architectural observation when we started Groq was that everyone knows Moore’s law: every 18 to 24 months, like clockwork, the number of transistors doubles, which means double the compute.

But we noticed that AI was getting better faster. It clearly wasn’t the algorithms, because algorithms have discontinuous jumps. It didn’t seem to be the data, because there wasn’t that much more data. The transistors were only doubling every 18 to 24 months, so where was all of this capability coming from?

It turned out that the number of chips was also doubling every 18 to 24 months. Rather than 2 times, it was 4 times. The question we asked was, if you’re effectively going to have an unlimited number of chips, do you do something architecturally different?

The answer is absolutely. Rather than using external memory, we use a large number of chips and keep all of the parameters of the model live in the chips. Then we have this pipeline where the computation flows through it, sort of like an assembly line.

Imagine you were trying to build a factory, and the factory was only 1/100th of the size needed for the assembly line. You would run a bunch of cars through 1/100th of it, tear it down, set up the next 1/100th of the assembly line, and do that over and over again. That’s the way a GPU works.

LPUs are very different. We have the computation flow through a whole bunch of chips. Rather than using 8 chips, we’ll use 600 or 3,000 for a model.

6. Energy Efficiency in AI Computation

Harry Stebbings

How does that change energy efficiency? How does it improve when you use more? You use more per token, so the footprint is higher.

Jonathan Ross

Think of it as the difference between a factory and a backyard garage. The backyard garage is not going to be as efficient, but it has a lower energy footprint.

Another example would be if you were trying to transport a ton of coal from one side of the city to the other, and you did it on mopeds or with freight trains. Which would be more efficient?

The moped would use less energy per trip, but it would need more trips, and therefore it would use more energy overall. In fact, this is one of the things most people misunderstand. They think edge computing uses less energy.

Actually, edge computing is less energy-efficient than computing in the data center. When you’re computing in the data center, it’s a little bit like that freight train. You’re getting to do a whole bunch of jobs simultaneously. The fact that we don’t have to read from that external memory means we don’t have to spend the energy doing that. Even with GPUs, you get to batch.

Going back to why it’s so energy-efficient, the amount of energy used in a chip involves physical wires. The physical wires have a width, and when you look at the width and the length, you charge that wire up to set it to a 1 and discharge it to set it to a 0. It’s like charging and discharging a capacitor, and you’re using energy.

The longer the wire, the more charge is required. When you have HBM here and another chip here, you’re charging a wire between the chips and discharging it every time you send a bit. That’s a long distance to travel, and the wires are wider than the wires inside the chips.

You use a lot more energy when we keep that memory in the chip. It’s only traveling a short distance using much thinner wires, so it uses a lot less energy.

Harry Stebbings

Do we see a world of LPU and GPU usage in combination? How does that distribution look between LPU usage and GPU usage?

Jonathan Ross

There are a couple of things. First, training should be done on GPUs. I think NVIDIA will sell every single GPU they make for training.

Right now, about 40% of their market is inference. If we were to deploy a lot of much lower-cost inference chips, what you would see is that the same number of GPUs would be sold, but the demand for training would increase. The more inference you have, the more training you need, and vice versa.

The other use case is that we’re so much faster than GPUs that we’ve experimented with taking some portions of the model and running them on our LPUs while letting the rest run on a GPU. It actually speeds things up and makes the GPU more economical.

Since people already have a lot of GPUs deployed, one use case we’ve contemplated is selling some of our LPUs to nitro-boost those GPUs.

Harry Stebbings

People have bought GPUs so far ahead of time that, by the time they get them, they’re deployed and installed, and they’re almost out of date.

Jonathan Ross

We’ve spoken with some customers that put orders in more than a year in advance, paid a year in advance, and still haven’t received them.

The recent deployment we did in Saudi Arabia went from contract to the first tokens being served in production in-country in 51 days.

Harry Stebbings

How were you able to do it so quickly? 51 days is astonishing.

Jonathan Ross

Part of it is that, architecturally, things are much simpler for us. We don’t have a bunch of other hardware components. We don’t use switches to communicate between our chips; we just plug our chips into our chips. Our chips are the switch.

We also don’t have all of this network tuning. Think about it this way: when you’re going across town in Paris, how long does it take to get from one side to the other?

Harry Stebbings

A long time.

Jonathan Ross

A long, variable amount of time. If you do it in the middle of the night, it might be fast. If you do it in the middle of the day during an event like the one we have going on with the AI Summit, it’s slow.

Harry Stebbings

Exactly.

Jonathan Ross

It’s unpredictable. However, certain modes of transportation, like trains, can be predictable. With what we’re doing, it’s 100% predictable given the energy efficiency and the predictability.

Harry Stebbings

Why is NVIDIA not being more proactive on LPUs?

Jonathan Ross

What makes you think they don’t want to be more proactive on it?

Harry Stebbings

They don’t talk about it. Why would they not talk about LPUs? If you wanted to protect shareholder value and a Wall Street image of dominance and being ahead of the game, you’d at least say, “Of course, we’re working on LPUs as well.”

Jonathan Ross

Until they had that ability, they would effectively be exposing that there was something missing. If you look at the last GTC, there was an announcement that the latest GPUs were 30 times faster than the previous generation.

When you look at how it was done, there was this curve that looked like this, and then it ended here. Then there was another curve that looked like this. That 30 times was from the end of this curve to the end of the next curve.

If you moved it here, it would have been less than 30 times. If you moved it here, it would have been infinite. Their chip is infinitely faster than the previous one, but that wouldn’t have sounded reasonable.

7. Why Most Dollars Into Datacenters Will Be Lost

There’s a history in this market of specmanship, because it’s so hard to get access to chips. This is a lesson in enterprise sales. People rely on specmanship: “My specs are better than your specs. My chip is faster than your chip. I get more teraflops per second than you do.”

Who cares? Tell me what the tokens per dollar are and what the tokens per watt are. Nothing else really matters.

People will find all of these other weird things to measure that they might be better on. It’s like saying, “I’ll sell you a car with better RPMs.” RPMs don’t matter. What matters is miles per gallon, and maybe the speed you can drive, although speed limits render that moot.

In enterprise sales, people used to buy or market soap by saying, “Our soap has more bubbles than this other brand’s soap.” Who cares? What they figured out was to put really happy people on a billboard after they used the soap, so people might associate that happiness with the product.

Harry Stebbings

Lifestyle marketing.

Jonathan Ross

Exactly. For some reason, enterprise still hasn’t learned that lesson. It’s still, “We have more bubbles. We have more teraflops. We have more of whatever”—things that people literally don’t care about.

Harry Stebbings

You think NVIDIA saying it’s 30 times faster isn’t good marketing?

Jonathan Ross

I think it worked because it’s what people are used to. Our counter was that we did a press release saying, “Groq still faster.” That was it, and people went crazy over it because it was simply, “We’re still faster.”

Harry Stebbings

Do you think Wall Street views it that way?

Jonathan Ross

I think they’re starting to. But again, I don’t think there’s real competition here. If you’re competing, you’ve done something seriously wrong.

If you’re competing, it means you haven’t found an unsolved customer problem. If someone else has already solved the problem, why are you spending time on it?

Harry Stebbings

So you don’t view NVIDIA as a competitor?

Jonathan Ross

They don’t offer fast tokens, and they don’t offer low-cost tokens. It’s a very different product. What they do very well is training. They do it better than anyone else, and by such a wide degree that it’s a solved problem.

Why would we bother trying to solve a problem that’s already been solved?

Harry Stebbings

So you cede the training market to them and own the inference market?

Jonathan Ross

They’re saying they also want the inference market, of course. It’s the way it always works.

Harry Stebbings

So what do we do now? Are we competing in the inference market?

Jonathan Ross

We don’t really have people saying, “We’re going to buy GPUs instead of you.” We do have people saying, “We’re going to buy both.” That happens, but we don’t care.

I showed a demo to someone, and he said, “Should we just not buy any more GPUs?” I said, “No. You should buy every single GPU you can get your hands on.”

He looked at me very perplexed. I said, “How are you going to do training? We don’t do training. Buy the GPUs. Get every single one you can, because I want your models running on us to be really good.”

For inference, they don’t need to buy NVIDIA anymore. They don’t need to buy GPUs for inference. But if you can get them, they’re a little expensive, and if you’re used to them, why not? Plenty of people still sell mainframes.

If you want lower cost and faster performance, you want an LPU.

Harry Stebbings

How much lower cost is it? More than 5 times lower?

Jonathan Ross

More than 5 times lower. The memory alone in the latest GPUs costs more than our fully loaded capital expenditure per chip deployed.

On top of that, we talked about energy efficiency. We use about 1/3 of the energy per token. Over a 3-year period, 1/3 of our cost is operating expenditure, which is mostly energy and data-center rent, and 2/3 is capital expenditure.

Since we use 1/3 of the energy, the cost to run that GPU to produce the same number of tokens is the same as our total cost. Just the GPU’s operating expenditure is the same as our capital expenditure plus our operating expenditure.

Harry Stebbings

Why is 40% of NVIDIA’s revenue inference, then? Why haven’t you taken more of that?

Jonathan Ross

At the beginning of 2024, we only had 640 chips. At the end, we had 40,000. We’re not at that scale yet.

You have to provide quality, low cost, speed, and capacity. This is where the most important part of not using HBM came in. It means that we effectively have no scale limits.

The GPU itself is manufactured using the same process you use for a mobile phone. The same silicon that’s in your mobile phone is the same silicon used for the GPU. In fact, they build the mobile-phone chips first because they’re smaller, and NVIDIA gets them after Apple.

The difference is the memory. That’s the only difference. But that memory is the hard part to manufacture, and it’s limited in scale. By avoiding it, we effectively have almost no limit on how much we can scale up.

That’s important for inference.

Harry Stebbings

What is NVIDIA’s margin? 70% to 80%?

Jonathan Ross

70% to 80%.

Harry Stebbings

So they can take 70% to 80% off and be radically more cost-effective than you?

Jonathan Ross

Comparatively, yes.

Harry Stebbings

You could destroy their margin. Why wouldn’t you?

Jonathan Ross

In that same vein, you can almost say we’re one of the best things that ever happened to NVIDIA. They can make every single GPU that they were going to make and sell it for training, where it has a high margin. That gets amortized across the deployment.

We’ll take the low-margin, high-volume inference business off their hands, and they won’t have to sell either margin.

Harry Stebbings

What does low margin mean?

Jonathan Ross

Depending on the deal, we do get some on the back end, but up front it’s about 20%.

Harry Stebbings

So there’s 80% for NVIDIA and 20% for you, but you’re looking at a 20-times advantage.

Jonathan Ross

Then we get more later.

8. Meta, Google, and Microsoft's Data Center Investments

Harry Stebbings

What do you mean, you get more later?

Jonathan Ross

The deals we do are structured so that the partner puts up the money for us to deploy. We pay it back with a decent IRR, but we split the revenue, and most of it goes to the partner. Once we hit the IRR, it flips the other way.

Harry Stebbings

So others are putting up the capital expenditure for you?

Jonathan Ross

Yes.

Harry Stebbings

What does it look like at the end?

Jonathan Ross

It’s not like other business models. We didn’t just innovate on the chip; we also innovated on the business model.

We’re limited in how much money we can make based on how much we can deploy, not how much money we have, because the partners are putting that money up. When I’m looking at what we can do, it’s all about how much we can scale.

Harry Stebbings

What are the limits to your deployment? Is it purely chip constraints?

Jonathan Ross

Mostly.

You asked about misconceptions in AI. I think one of them is around power. It’s true that there’s a mismatch in the market between people with chips and people with power, but that’s partially because you need a data center in the middle, and there aren’t enough data centers.

Data centers aren’t the hardest thing in the world to build. They’re not easy, but they’re not the hardest thing. It’s harder to build up the power.

Because of that mismatch, you have big hyperscalers going around saying, “I need 1 gigawatt of power,” and they’ll say this to 60 different potential data-center builders. All of a sudden, you hear an echo: “I heard there’s a gigawatt here, a gigawatt there, and a gigawatt somewhere else.”

Suddenly, there’s 60 gigawatts of demand. It’s just an echo from that first gigawatt.

I’m aware of about 20 GW of power that people want to make available for data centers. Right now, there are about 15 GW of data centers worldwide, so that’s more than double the current capacity.

My concern is that people are now building up more power. In the next 3 to 4 years, people will say, “I built up all this power and no one is using it. This was a complete waste, and we’re never going to do this again.”

But remember that doubling of chips every 18 to 24 months. Over 3 to 4 years, you double that 15 GW twice, and now you’re talking about 120 GW. There isn’t that much power available. Then you double it again, and now you’re at 240 GW.

What’s going to happen is that we’re going to overbuild slightly right now because of the mismatch and miscommunication. Then we’re going to dampen our building and close down on that, and then we’re going to have the real need for the power.

That power will become a hard bottleneck in 3 to 4 years.

Harry Stebbings

Why will we have data-center oversupply when we’re moving into a world of inference that will be 20 times larger than training?

Jonathan Ross

The problem with data centers is that everyone thinks data centers are real estate. A lot of people do real estate data centers, but data centers are not real estate.

The common joke in the industry now is that someone says, “I’m going to have 100 megawatts of capacity for you in 3 months. Are you willing to sign?”

Then you ask, “What’s your uptime?”

They say, “I don’t know. Whatever the power grid is.”

You ask, “Where are your generators?”

They say, “I haven’t ordered those. I’ll order them now.”

There’s a 90-month lead time on generators right now.

Harry Stebbings

Really?

Jonathan Ross

Then there’s the next question: “Where are you getting the water from?”

They say, “Data centers need water? I thought it was just a bunch of chips.”

There are a lot of people who have no idea what they’re doing going into this because they think it’s real estate. Those people are building an oversupply of data centers, but they’re not really building them. They’re fake data centers that people think are real.

Harry Stebbings

What happens to those data centers if they’re not utilized? Is Amazon going to pay for a data center that doesn’t work?

Jonathan Ross

Amazon doesn’t fall for this. Amazon has really good people. Whoever the buyer is isn’t going to pay for a data center with no water or power.

Harry Stebbings

So these projects will never be developed? Will we build them fast enough?

Jonathan Ross

It does take time to build the data center, so it’s almost okay. If you train a model, you really want to amortize it over about 6 months. If you deploy chips, you really want to amortize them over 3 to 5 years. We’re more on the 3-year side; others are more on the 5-year side.

If you build a data center, you’re probably talking about 10 to 15 years. For a power plant, you’re talking about 15 to 20 years.

The problem we have in the industry is this mismatch between the financing and the needs. Someone wants to train a model, and they’re going to be doing that for 6 months. They don’t understand why people want 3-to-5-year commitments on the chips.

The people deploying the chips don’t understand why someone wants a 15-to-20-year commitment on the data center.

Harry Stebbings

It’s at 7 years now on the data centers, and then the people building the data centers need a long, 7-year commitment?

Jonathan Ross

That’s the kind of commitment they’re asking for.

You have this complete mismatch throughout the ecosystem. The funny part is that, while they all want to take zero risk and have a sovereign-wealth-level credit rating on the other side with long commitments, the longer the payoff time, the more generic the infrastructure is.

A model has a very specific use, but accelerators like LPUs and GPUs can be used for other things besides generative AI or LLMs. The data center can be used for other things besides the accelerators. The power can be used for anything.

While they’re looking for the least risk, they’re looking for it in the place where there is the least risk. If we don’t use the power for AI, we’ll use it to power all of the electric cars.

Harry Stebbings

Is this a case where incumbents win because they’re one of the only ones able to match the durations required by data-center providers?

Jonathan Ross

This is why we’ve partnered with Aramco and this new entity in Saudi Arabia. They have an enormous ability to fund this over the long term, a very long-term perspective, and an amazing credit rating.

Harry Stebbings

When you say they have the ability to fund it, this is why there was a misconception. People think it’s a funding round of $1.5 billion.

Jonathan Ross

We did not raise $1.5 billion. That’s revenue. That’s actually about 30% of the revenue of OpenAI.

Harry Stebbings

Can you walk me through how that deal is structured?

Jonathan Ross

We started off last year, and we got to 19,000 of our chips deployed. We did that in about 51 days. The question was what we could do this year.

They’ve gone off and collected a bunch of power in the country. The deal is structured so that they will put up the capital expenditure for us to deploy our chips in that data center or those data centers, and we pay it back based on the money we make.

It’s a little different from debt in that they participate in the upside, but it’s similar in nature. It is revenue because we actually make a profit upfront.

Harry Stebbings

How does that change what you can do?

Jonathan Ross

We’re not limited by capital anymore.

There’s another misconception around Groq. There was a paper that said we couldn’t be profitable while being the lowest price. It said we could charge more, but we actually have a very positive contribution margin right now.

As far as we know, we’re the only ones making money running these open-source models. With the open-source models, everyone is competing with venture-capital dollars, trying to take market share in an Uber-style model.

Meanwhile, we’re sitting here saying, “We could do this all day long,” because we’re making money. We’re able to pay off an IRR and make our partners money.

There’s another part of the model. We’re also working with some proprietary model providers. We showed off the first one at LEAP on Sunday, where we did a voice model with PlayAI. That one is also a revenue share.

The difference is that they get to make money off it, whereas most others in the industry are losing money because of the commoditization of the models.

Harry Stebbings

Do you have cheaper pricing over time as you have less monopoly power, or do you have higher prices as your monopoly increases?

Jonathan Ross

We want the margin to stay about the same, but we want prices to go down. Then we get into Jevons’ paradox, and life gets great because we’re going to scale.

Our focus is on getting to scale. To preserve human agency in the age of AI, we need to be one of the most important compute providers in the world.

Our goal by the end of 2027 is to provide at least half of the world’s AI inference compute. We think we could be further than 2 times that, given that we don’t have all the constraints.

To get there, we need to be aggressively building out, and we need to give people no excuse for not running their models on us and using the models that are on us. We do that by charging extra.

What I keep telling the team over and over again—you have to remind them sometimes—is that we are growing faster than exponential. When you’re growing faster than exponential, there is no amount of profit you can make that matters.

What matters is getting a toehold in the market and becoming relevant.

Harry Stebbings

What would prevent that?

Jonathan Ross

We used to worry that someone would try to price below us. Then we realized that wasn’t a concern, because so much money is going into this that people are going to want to lose less money by running on us.

That isn’t a concern. It was the big one early on, until we realized that when we see Mark Zuckerberg investing $65 billion in data centers, he’s internalizing all of the margins he would have had to spend on data centers with the providers we mentioned earlier.

Harry Stebbings

Meta is doing $65 billion a year. I think Google said $70 billion or $75 billion, and Satya said Microsoft is doing $80 billion. Then you’ve also got Stargate.

These are crazy sums of money. Is all of this for data-center building?

Jonathan Ross

It also includes the things that go in them, including the chips, the systems, and everything else.

Harry Stebbings

We’ve never seen money like this.

Jonathan Ross

No. There’s never been anything like this. But there’s never been a case where it was so clear that there was going to be value at the end.

If you knew how successful search was going to be, remember that Google stayed private as long as it did because they were afraid Microsoft would figure out how much money search was making and try to replicate it.

The moment Google went public, Bing appeared. They called that perfectly. Everyone knows how much money there is in AI, so everyone is going after it.

9. Distribution of Value in the AI Economy

Harry Stebbings

Do you think that value is distributed among many players or concentrated toward 1 or 2? I completely agree with you in terms of the clear value when assigned, but is it distributed somewhat evenly or concentrated?

Jonathan Ross

It’s a power law. The more value there is in the economy, the more risk there is of a single entity being so far on one end that it just dominates.

You see this with the Magnificent 7. The bigger the economy gets, the more you’ll have big swings in the economic outcomes.

Right now, the hyperscalers are all roughly even in their market caps. It’s strange. You would expect one of them to be killing it and taking it much further, so I don’t understand why they’re so closely grouped.

10. Stages of Startup Success

Harry Stebbings

When we think about that distribution, how do we think about changing it? With Groq, you want to be one of the Magnificent 7 and one of the most important companies in the world. How do you see that happening?

Jonathan Ross

The way you get there and the way you stay there are 2 very different things.

There’s a circle of life that happens in startups. The first stage is to solve an unsolved problem. That’s how you go viral and do well.

The second stage is the marketing stage, where other people are trying to copy what you’ve done because they can’t think of something themselves. Now you have to fight it out in advertising, marketing, and so on. You see consumer-packaged-goods companies often get stuck there, and it becomes more about where they are on the shelf than anything else.

The final stage is the 7 Powers. It’s once you’ve found some of those powers, started improving, and developed systemic advantages.

Then someone solves an unsolved customer problem and the whole cycle of life continues. Google has to redo this now because LLMs are better than search.

11. The AI Investment Bubble

The way you become one of the Magnificent 7 is by solving that unsolved problem. The way you stay there is by finding one or more of those 7 powers, and then being ready for when you get disrupted so you can continue fighting back and solving customer problems.

Harry Stebbings

We mentioned the huge amounts of money being spent here. Is this a good bubble that lays the foundations for an incredible next 10 to 20 years, where the capital actually turns out to be productive even if it doesn’t seem so on paper? Or is it a case where a huge amount of money is incinerated on depreciating assets?

Jonathan Ross

I can guarantee you that a huge amount of money will be incinerated. But I also bet that, in total, more money will be made than will be put in.

That’s the problem. You have to look at it either in aggregate or as individual bets. When everyone is making investments in the market, some people are going to lose money because not every company is going to be successful.

What you always see when there are real technology improvements or things coming is that you have the early things that people invest in heavily, and they’re super successful. Then everyone else wants to get in on it.

You go from AI chips and AI models to AI T-shirts, and next thing you know you have AI thermal grease. People start applying AI to everything. Next thing you know, you’ll have an AI condo.

Harry Stebbings

Sure.

Jonathan Ross

The trick is discerning what’s real and what isn’t. You’re always going to have obnoxious charlatans coming in whenever there’s something real. That’s unfortunate, but eventually they get cleared away once people understand the technology and what’s real.

The job is to start educating. The more educated people are, the less they’ll invest in AI thermal grease.

Harry Stebbings

What is the largest individual bet that will lead to the largest incineration of cash?

12. The Keynesian Beauty Contest in VC

Jonathan Ross

I’m not going to call anyone out in particular, but I actually think it will happen across every single discipline.

Are you aware of the Keynesian beauty contest?

Harry Stebbings

No.

Jonathan Ross

John Maynard Keynes, the economist, had this great idea. It explains everything you need to know about venture capital.

Harry Stebbings

I’m nervous, but keep going.

Jonathan Ross

Take a magazine full of models—human models, good-looking models—and have a whole bunch of VCs in the room. They’re allowed to make bets on who the most beautiful model is.

In the end, whoever has the most money on them is the winner. Based on the proportion that you put on that particular model’s face, you get a share of all the money.

If you put money on one that isn’t the most beautiful by dollars, you lose your money to the people who bet on the one that was. That was the bet SoftBank was making: “I can win the Keynesian beauty contest. I’m just going to put more money in, and I’m going to win.”

That’s problematic when you have true technological advantages as opposed to marketing. When you’re solving customer problems, it’s a weighing machine. Once the customer problem has been solved, you get into this popularity contest of marketing.

Something unusual has happened this time around that I don’t think has ever happened in venture capital before. You see people raising billions of dollars who have competitors that have also raised billions of dollars.

Usually, there’s a clear winner in the Keynesian beauty contest. You don’t have this fight where it’s, “I have to put a little more money in. I have to put a little more. I’m going to put in $10 billion. I’m going to put in $20 billion. I’m going to put in $500 billion.”

The Keynesian beauty contest has gone completely amok. This has never happened before, so people don’t even understand how to react. It used to be that if someone had raised $1 billion, you would say, “They’re the winner.” Now there are 3 or 4 competitors who have $1 billion, so who wins and who loses?

Harry Stebbings

Is Masayoshi Son going to incinerate the largest amount of cash ever?

Jonathan Ross

I think the Keynesian beauty contest no longer applies because there’s so much money available, spread out across the market.

I think the people who have the best products are actually going to be the winners. Everyone can be capitalized, but there will be problems for the winners because of this.

The problems will be things like the employee you were going to hire, where someone offered them a ridiculous amount of money. You see this all the time now. They could have gone and contributed to the winner, but now they’re contributing to a competitor that shouldn’t exist, or that’s equally likely to win. Now you’re splitting the talent.

Harry Stebbings

What do you do when you have such high salaries? We’ve seen $1 million or $2 million for junior-to-mid-level people at some of these companies, and they’re living an amazing life in great places.

Jonathan Ross

Do you think they’re living that amazing life in Guangdong when they’re working for DeepSeek or another Chinese alternative? I don’t think so. They’re getting paid much less, working their asses off 20 hours a day, not getting kombucha, and not being paid $2 million a year.

Harry Stebbings

Fair.

Jonathan Ross

Not only fair—we have a policy that we never offer the highest salary. We want people to choose us, not choose the salary.

If we win in a bidding war, the next time someone comes along with a higher salary, that person is just going to take the other job. There’s no loyalty, and they don’t believe in the mission.

Instead, we focus on saying, “We’re going to build this. This is your opportunity. You’ll get to work with amazing people.” Spend some time with the team. Are these the people you want to be working with?

Frankly, you’re going to make so much cash that it doesn’t matter. Bet on the equity and the outcome. Help us make this thing valuable.

People who buy into that are much easier to manage because they’re mission-oriented. They all want to do the same thing. They’re not there because they want kombucha, and they’re not going to complain because the cappuccino machine is broken. They’ll just go and buy their coffee next door.

13. NVIDIA's Role in the AI Ecosystem

Harry Stebbings

Will you and NVIDIA move into the model layer? Everyone talks about model builders becoming application providers. Will infrastructure providers become model providers?

Jonathan Ross

We’ve decided that we’re not going to train our own models. We’ll do a little fine-tuning for specific cases, but we don’t want to compete.

That’s important because people are putting their models and weights on us, and they don’t want us to learn from and take that information for our own benefit.

That’s the problem you have when you work with a hyperscaler: they’re also doing everything you’re doing. We’ve decided that model providers should make the model. We don’t do that.

There’s also the data side—the users and the queries. Another thing we could do, but do not do, is log the queries and then use that data if we wanted to train.

We don’t train, and we have no reason to hold the data. We only temporarily store things in DRAM, so there’s no persistent storage. If the power went out, everything would be gone. DRAM is limited, so we can’t hold things for a long time.

You know that we don’t have your data. People who are building businesses on top of us can obviously keep the data from their customers if they want. We have no control over that, and that’s fine. But we don’t take any data.

Harry Stebbings

Do you think NVIDIA will move into model provision?

Jonathan Ross

It’s possible, but if I were them, I would avoid it. I wouldn’t want to give my customers to a company that I was competing with.

NVIDIA is great at training. It’s crazy. It would be like being an automotive company and then creating your own taxi service. You’re competing directly with your customer.

Tech companies love to do this. We have a management philosophy based on Big O complexity, and we only do things that require a sublinear number of employees.

If someone comes to me and says, “I need 10 people to go do this,” a lot of people would say, “Why can’t you do it with 5?”

I would say, “You’re supporting customers. If we double the number of customers, do you need 20 people or 11?” I want to know the growth rate. Are they automating everything?

We completely automated our compiler. We automated large portions of our cloud and everything else. That means we can scale with a small team.

We have 300 people. We built our own chip, networking hardware and software, runtime, orchestration layer, compiler, and cloud. We built all of this with 300 people.

We can only do this with a small number of people because you don’t have the communication overhead. You have to decide what your constants and variables are—what are the things you want to preserve?

One of our constants is talent density. We want to stay small and nimble.

The other side is that growth is a problem. We measure our growth in what I call problem units. Every time you triple something, you have about the same number of problems as the last time you tripled it.

Going from 100 employees to 300, from 300 to 1,000, and from 1,000 to 3,000: each of those has roughly the same number of problems.

We scaled from 640 LPUs at the beginning of last year to 40,000. That’s 4 problem units—4 triplings of the number of chips. If we were also tripling the number of employees, that would be another problem unit.

Management bandwidth is limited. You can only solve so many problems, so you have to decide where you’re going to allocate them.

If you build things really well from the beginning and can scale up with the number of employees you have, then you can scale over here. If you want to triple the number of customers, there’s another problem unit that you have to solve.

Harry Stebbings

What’s the biggest challenge when you’re scaling at that rate but the team isn’t scaling in conjunction with it?

Jonathan Ross

There’s a common belief that the people you have early on are right for the job, and that the people you get later might be better in a more corporate environment. I don’t think that’s the case.

You should always try to get generalists. Otherwise, you get stuck in a particular way of doing things because that’s what one person knew how to do.

There are people who burn out. Being in a startup is hard, and some people literally burn out. There are also people who were the best you could get at the time, and people who are unmanageable wild children. They should go off and start another startup; they shouldn’t be scaling with you.

That happens, but it’s rarer. Saying you’re going to hire B-players because you’ve gotten large enough is laziness and an excuse. It’s a lack of creativity in your business model and in the algorithm of how you’re going to scale.

Think of it this way: Walmart versus Amazon. Walmart wants to double the number of customers, so it has to double the number of stores and employees. Amazon doesn’t need to double the number of websites.

That’s a fundamental advantage, but Amazon still has to double and improve its logistics. It doesn’t have as many problems that have to scale linearly, but it has some.

If you wanted to disrupt Amazon, you would build a completely robotic logistics system and bring the complexity and overhead down. Then you could outmaneuver them.

That’s how you need to think. Don’t just say, “I need more people.” Focus on the algorithm of your business.

14. China's AI Strategy and Global Implications

Harry Stebbings

The last time we spoke, we discussed DeepSeek. I think more has come out over the last few weeks about their innovations and some of the distillation they used. Where is China better than us today?

Jonathan Ross

As we discussed, they’re more willing to use things that perhaps they shouldn’t be using. They distilled the OpenAI model.

A lot of people have the opinion that OpenAI was scraping the internet, so, “Good for DeepSeek.” But whether that’s right or wrong, most model providers considered it a red line. They didn’t want to cross it.

I don’t know if that’s going to change, but it might.

Harry Stebbings

What about the open-source nature of DeepSeek? Does OpenAI now benefit from the innovations they made?

Jonathan Ross

They also probably have all the data that DeepSeek paid them to generate.

Harry Stebbings

But DeepSeek was clever. They innovated.

Jonathan Ross

I think the biggest thing is that this is a shot in the arm for morale in China. It gives them a sense that they’re in the race.

But, as I said, it’s Sputnik 2.0. It has also woken up the United States.

Harry Stebbings

How do you compare Stargate with the $128 billion China has now committed?

Jonathan Ross

China has a more complicated situation and a simpler one at the same time.

The problem is that they don’t have the chip efficiency we have. On the other hand, they have scale. If they wanted to deploy 150 nuclear reactors—I think that’s the plan—it’s no big deal. They just do it.

If the chips aren’t as efficient, they can deploy more of them. On the other hand, if they want to go out into the world and deploy chips the way they did with Huawei and networking equipment, that’s going to be complicated.

People around the world aren’t going to have the power to run more expensive accelerators. At home, I don’t think anything is a problem. As they try to expand, it’s going to be an issue.

Harry Stebbings

China is quite opaque in everything. What do we not know about China that we would like to know?

Jonathan Ross

The most important thing to understand is where they’re going to end up on the censorship and privacy of these models.

We come from democratic countries, and we have an expectation that companies can build something that says anything. Are they going to be permissive and allow models to make mistakes and hallucinate, or are they going to shut them down?

If you know that, you know whether China has a shot.

One of the biggest nightmares they have is free speech. It’s the exact opposite of the vulnerability we talked about earlier.

Can you imagine Xi Jinping going out and saying, “Country, we’ve lost our advantage in AI. I need your help”?

Harry Stebbings

Never.

Jonathan Ross

It’s always going to be, “We’re the greatest. We’re the best.” Everyone will know differently, but they’ll all have to toe the party line.

Because of that, I think it’s hard for them to allow these models to say anything. For them, it’s a bad thing if the model says the United States is great and better at something.

That’s going to tell you a lot about the AI story in China. If they aren’t permissive of more open and truthful models, then they’re inherently disadvantaged.

Harry Stebbings

You’re saying that if they aren’t more permissive, and you’re running a Chinese tech company, your fear is that you become Jack Ma.

Jonathan Ross

That’s really going to stifle innovation. If I were in China right now, I’d be looking for the exit. If your craft is AI, you’d want to do that somewhere supportive.

Harry Stebbings

Do you really buy that they don’t have access to Blackwell? This is China. I can imagine Xi Jinping saying, “Sorry, no Blackwell.”

Jonathan Ross

I don’t think it matters whether they physically have it. Right now, most cloud providers are happy if you swipe a credit card to rent it to you.

Harry Stebbings

But there are limits to renting.

Jonathan Ross

I think one of the concerns right now is Malaysia or Singapore—or that region—being a place where people are deploying GPUs with a wink and a nod: “We’re not going to rent them to China.”

A lot of people believe that’s happening. Otherwise, that’s a lot of GPUs for that region.

Harry Stebbings

It feels like an even bigger safety net in case the tap is turned off at the hyperscalers, because right now you can write a check to any of the hyperscalers and say, “I need these chips.” They’ll deploy them and you can run on them. It doesn’t really matter where you’re coming from.

Jonathan Ross

If you’re a sanctioned country, it matters.

Harry Stebbings

China isn’t sanctioned.

Jonathan Ross

China isn’t sanctioned.

15. Europe's Potential in the AI Revolution

Harry Stebbings

So we have China, which is obviously proving that it’s in the race. We have the United States, and then we have Europe, which feels like it’s languishing. Is this the ultimate nail in Europe’s coffin?

Jonathan Ross

We talked about how Groq almost died, but we had the right technology all along. We were just waiting for LLMs to arrive. I think Europe is very similar.

Europe has amazing talent—amazing talent—but that talent leaves and goes to the United States or other places. The question is, how do you have Europe’s LLM moment?

It’s not that complicated. When you surround yourself with people, you become the average of your 5 closest friends. If your 5 closest friends say, “That’ll never succeed. You should just keep your job. Startups are terrible,” you’re going to be risk-averse.

If your 5 closest friends say, “You should do it. That’s great. I support you,” you’re more likely to start a company.

Even in Silicon Valley, people make the transition from a big tech company to a startup, and it’s hard. They’re comfortable, making those crazy salaries. The big companies take care of them, and they have a fiduciary obligation to their families.

How do they make that leap? It’s because they have entrepreneurs constantly trying to hire them. They hear the pitch all the time, and they get used to it. They see success around them, and VCs come in and try to close candidates in the early stages.

Europe needs the same thing. You need a place where people are surrounded by entrepreneurial people who are risk-on and aren’t going to try to talk them out of joining a startup.

Harry Stebbings

From a regulation perspective, Europe is unbelievably efficient in the mastery of regulation. I was speaking with someone the other day, and the EU has supposedly hired 1,500 people for AI safety and policing.

What would you do if I put you in charge of European AI regulation?

Jonathan Ross

I wouldn’t waste my time regulating something that doesn’t exist. Instead of regulating, what are you going to promote?

You want to promote risk-taking. You want to promote an enclave of people who are risk-on.

I was visiting Station F yesterday. It was amazing. Macron was there, and it was full of vibrant people. You could feel it.

I was talking to the person who runs Station F, Roxanne Varza, and Xavier Niel. We were talking about City F: a place where you start with 10,000 people in the center, within a small radius. Once it’s full, you expand it, and once that’s full, you expand it again, until you get to perhaps 1 million people in Europe who are all risk-on.

It would be a little Silicon Valley. I would give it special economic dispensations. I would allow everything that employers need, make it simple, and say, “If you don’t want to buy into that, go to other regions in France or other regions in Europe. But if you want to participate in what’s going to be the biggest technological revolution in human history, this is the city for you.”

Harry Stebbings

You’re inherently punishing incumbents. If we’re talking about AI insurance-underwriting startups, there are many companies going after insurance underwriting with AI. If you give them benefits like that, you’re inherently punishing some of the biggest insurance providers in your region.

You’re punishing people who hire 200,000 people. That feels unfair.

Jonathan Ross

There is no right to be an incumbent, especially a slothful incumbent that isn’t reacting to disruption. You want to encourage disruption.

One of the things in Silicon Valley is that you can move from one place to another. There are no non-solicitation agreements anymore. When I started, we had them, but even those are gone.

That free movement of people is very important.

Harry Stebbings

Are you allowed to start work straight away?

Jonathan Ross

Straight away.

Harry Stebbings

But not before?

Jonathan Ross

If you start before, that’s a problem.

Harry Stebbings

We have to wait 6 months.

Jonathan Ross

There’s no such thing in Silicon Valley.

Harry Stebbings

If you’re a company right now, it feels like it’s harder to poach people. But what does that do? It suppresses wages. It’s harder to hire someone, people are less likely to move, and there’s less competition.

It suppresses wages, and the company has to pay for the 6 months anyway. It makes no sense at all.

Jonathan Ross

I totally understand that.

Harry Stebbings

You mentioned what you would promote. A lot of people would promote safety and regulation. Being European, I thought first about safety, specifically.

All Dario talks about these days is safety. Is he losing a step by being so focused on safety when, bluntly, his competitors are talking about product?

Jonathan Ross

Safety matters in AI. It’s a little bit like nuclear power: there are lots of pros and lots of cons.

I’m worried about different things than Dario is worried about. I’m more worried about people voluntarily giving up their decision-making authority because it’s so easy. This is what I mean by preserving human agency in the age of AI.

A good analogy is that you probably know plenty of wealthy people and the struggles they have bringing up children with wealth. I refer to it as financial diabetes.

You have children who aren’t incentivized to strive to succeed. I was very fortunate when I was growing up. My father lost all of his money multiple times.

He would sell a billion-dollar life-insurance policy and get all the commissions from it. You would have tons of money, and then you would spend it all.

There was a time when we were living in a $20 million mansion. There were a couple of times when we ordered food, and he would talk to the delivery guy and convince him to give us the food and let us pay him back later, because we would get money later.

One time, he was so despondent that he locked himself in his office and wouldn’t come out. My little brother came to me and said I had to go talk to the Chinese-food delivery guy and convince him to give us the food.

I was mentally preparing how to convince him. I walked out, walked up to him, and started getting ready to make my whole speech. He handed me the food.

I said, “I don’t have the money right now.”

He said, “Pay me later.”

I didn’t have to do anything. He trusted us because we were living in a $20 million mansion.

That happened multiple times. I have a friend who was homeless once for a couple of weeks, and he had almost been homeless a couple of times. He said the best thing that ever happened to him was being homeless for a couple of weeks because he survived it.

He said, “I’ve been through it. I always viewed this as the worst thing that could ever happen in the world, but now that I’ve been through it, I can survive it. I’m not worried anymore.”

I think we live incredibly comfortable lives—way too comfortable. Most people don’t have to go through things like that, so we have financial diabetes as a society.

I think it’s going to get worse with AI. We’re going into an age of abundance. Very few people have to worry about food security now, but what happens if you don’t need to worry about housing or anything else? What happens if you can live a life without working?

What is that going to do to your psychology? As we enter an age of abundance, how do we get people to keep making their own decisions and have a fulfilled life?

Harry Stebbings

Do we get better, or do we become accepting of good enough?

Bluntly, with the majority of shows, we start with OpenAI and do deep research. Then we use different prompts depending on the guest, and supplement that with a huge amount of research from speaking to ChatGPT, speaking to Claude, and speaking to everyone in between.

We care about it being good enough first and great later, with all the references. Most people will just be happy with good enough and get away with it.

Do we, as a human society, become happy with good enough?

Jonathan Ross

When we hire, we hire for something we call “booking the win early.” One of the most important driving forces for people is loss aversion. When you have something, you don’t want to lose it. People are less likely to go after something they’ve already had.

You grew up in a family that was well-off and then lost that. That might be part of your drive, because you want to get back to it.

When we’re hiring an engineer and there’s a room full of people saying, “If we do this thing, we could be twice as fast,” I want that engineer to hear, “If we don’t do that, we’re going to be half the speed we could have been.”

That’s the loss aversion. Book the win early. Because it’s possible, it must be done.

I think that’s a smaller segment of the population. Those are the people who deliver amazing things that no one else is going to do, because everyone else says, “That’s good enough.”

With AI, it’s so easy to create a prototype that, to stand out, you’re going to need to do more.

One of the things that has happened with the ability to communicate more freely and see what other people are doing is that the average restaurant is better than high-end restaurants were 20 years ago.

People see everything others are doing and start to expect that. You have less localization and more globalization, so you have to compete at the highest level.

AI is no exception. There are going to be 40 people creating the app you’re creating. You have to polish it in order to stand out.

Harry Stebbings

Listen, Jonathan, I could talk to you all day. I do want to do a quick 5.

What do you believe that most people around you disbelieve?

Jonathan Ross

I’m going to go with anti-Founder Mode. I’m anti-Founder Mode.

I believe in delegation. When you’re telling people how to do their jobs, that’s an indication that it’s not necessarily a problem with you. It could just be that the person isn’t right for the job, and it’s easier to direct them than to find someone else competent.

But it also means you probably haven’t aligned them.

We align people through this challenge coin. Everyone at Groq carries a 25-million-tokens-per-second challenge coin. It tells everyone what we’re doing. It’s alignment.

I can’t tell you how many people I’ve shown it to who say, “That’s awesome,” and yet no one else is making them.

Harry Stebbings

It’s heavy.

Jonathan Ross

It is heavy. But you like to say that the greatest things in life—the heaviest things—aren’t gold.

Harry Stebbings

I know. Gold, but made decisions.

Jonathan Ross

Exactly. This was a made decision, because I had to consolidate everything we were doing into 1 very simple message: “We’re going to get to 25 million tokens per second.”

I engraved it on a coin in this tiny amount of space and gave it to everyone at Groq. Whenever we’re in a meeting and something doesn’t help with this, they can tap their coin on the table and say, “No, no, no. That’s not the way this is going to go.”

Harry Stebbings

So is everyone wrong on Founder Mode?

Jonathan Ross

I think that’s what you do when you don’t have the quality of people working for you. You need the right gearing ratio between you and your direct reports.

Harry Stebbings

It’s a really unfair question, but I have to ask it. How do you analyze Elon’s attempt to buy Twitter—not buy Twitter, to buy OpenAI?

Jonathan Ross

I was sitting at the Élysée Palace, or however you pronounce it, at the dinner with Macron, Sam Altman, and JD Vance.

Frankly, I think Elon was a little jealous that Sam Altman was sitting next to JD Vance and it wasn’t him. It was right around the time Sam Altman was speaking that Elon announced it.

I thought Sam’s tweet response was pretty good. I would have probably said, instead of whatever he said about $9 billion, “I’m going to take Twitter public at $420 a share.”

It was attention-grabbing. Some people can’t stand not getting attention. My revenge on this is to give as little attention as possible.

Harry Stebbings

Let’s move on.

What would you do if you knew you couldn’t fail?

Jonathan Ross

I would put in 100% of the orders for every single chip we could possibly manufacture, because right now the demand is unlimited.

Every time you triple, you find the same number of problems, so you have to do it judiciously. But if I knew that, no matter what problem came up, we didn’t need to be safe at all, I would say, “Great. We’re going to go build 20 million chips.”

Harry Stebbings

In 10 years, is NVIDIA 3 times bigger, 10 times bigger, or 50 times bigger?

Jonathan Ross

I think they’ll be bigger. Training will become more important. I wouldn’t be surprised if they were 3 times bigger, and I also wouldn’t be surprised if they stayed around the same.

It’s hard to tell where things are going, because a lot of assumptions in the investment in NVIDIA were that they were going to run away with the entire market, including the inference market. They just haven’t built the right thing for inference.

As a weighing machine, I do think they should increase in value. But so much of it is a popularity contest that I don’t know whether they’ll need to grow to get to where they are.

16. Future Predictions and AI's Impact on Society

They might need to grow their revenue to justify where they are, but it’s a pretty fair multiple given everything going on. I couldn’t tell you. The popularity contest skews everything.

Harry Stebbings

What’s a crazy AI prediction you have that everyone else thinks is science fiction?

Jonathan Ross

I would assume that, in the next 10 years—and I know this is going to be crazy—you’ll see a Mounjaro moment for aging.

Harry Stebbings

You saw that picture of me and my weight loss. Unbelievable.

Jonathan Ross

70 pounds.

Harry Stebbings

70 pounds. I was on Mounjaro.

Jonathan Ross

If you know anyone who’s overweight and it’s hurting their health, get them on Mounjaro as soon as you can. It works.

Harry Stebbings

What is Mounjaro?

Jonathan Ross

It’s one of those GLP inhibitors, one of the weight-loss drugs that have become popular recently. It works.

My crazy AI belief is that, if it’s possible to significantly slow or stop aging, you’ll have a Mounjaro moment in perhaps the next 10 years. It came out of nowhere. All of a sudden, you could lose weight. Something finally worked, and it’s worked for a lot of people.

I don’t know whether it’s possible to slow or stop aging. Wear and tear is a real thing, and it might be impossible. But if it is possible, then in the next 10 years we will do it, and it will be sudden. It will be like the Mounjaro moment.

Harry Stebbings

I don’t see how it’s not possible. When you look at the advances that will come in medical research, I don’t see how it’s not possible that we’ll at least extend longevity by 60 years.

Dario will live to 150. I don’t see why that’s impossible.

Jonathan Ross

I don’t either, but I also don’t know that it is possible. Until I know that, I’m going to keep the conditional in there and say, “If possible.”

Harry Stebbings

What have you changed your mind on in the last 12 months?

Jonathan Ross

This is less of a mental one and more of an emotional one. We didn’t have product-market fit for 7 years.

That’s terrible. The morale changes when you find product-market fit. The world is brighter, the birds sing, I feel like hugging people, and you sleep. Life is better.

I forget whether it was you or someone else who was talking about type 1 and type 2 happiness. I think there’s a third type.

The 2 common ones are that the present is happy, and that you went through some really terrible things but the memories make you happy. There’s past, present, and future.

As a founder, the only type of happiness you get is this third type: future happiness. When you get product-market fit, you start to get past happiness. When you start to get revenue and everything else, you get present happiness.

It changes everything.

Harry Stebbings

If you had to bet on 1 company other than Groq to define the AI era, who would it be?

Jonathan Ross

I would focus more on the companies you haven’t heard about. I don’t know what those companies are, but I can tell you what they’re going to do.

The first will be the one that solves the hallucination problem.

The second will be the one that’s best able to break down subgoals for agentic systems. I think agentic systems come after you solve the hallucination problem, because otherwise you get these long chains where you can introduce hallucinations. It will work, but it will work much better after that problem is solved.

The next one is what I call the invent stage. Right now, the way LLMs work is that they make the most probable prediction.

It’s amazing. Imagine taking an entire novel and deleting the ending. You have a detective murder mystery, and you get to the point where the detective says, “The murderer is…”

The model can actually predict the answer. It had to understand everything, but it’s going to give you the most probable answer. That’s not good for invention, and it’s not good for art.

The reason the writing from LLMs is terrible is that it’s predictable. How do you say something that’s non-obvious but obvious when you see it? We don’t even have the right word for it: non-obvious but obvious.

That is going to unlock invention.

The final one is what I call the proxy stage. Someone will make it so models can make decisions for you. You can proxy your decisions, like the decision to do this interview.

Other things had to happen: the flight had to be booked, we had to get a ride over, and other things had to be canceled. You would trust an executive assistant or a chief of staff to make those decisions, but you wouldn’t trust an LLM yet.

That’s the final stage before you get to general AI. Each company that does one of those things is going to be a defining company.

Harry Stebbings

You said we have to fix hallucinations before we get efficient agents. Does that mean money going into agentic AI will be burned?

Jonathan Ross

No. Take hallucination as an example. The examples I gave were medical diagnosis and law—2 areas that will be unlocked once we get rid of hallucinations.

But there are startups like Perplexity that are doing just fine even though there are hallucinations, because it’s not high-risk. It’s for entertainment.

If you click the links, you can check them, and it works reasonably well. It depends on how risky the industry you’re in is.

You can start trying to position for the wave early and generate something. If you’re in the right position, as we were for 7 years, the wave comes.

That money isn’t necessarily incinerated. In fact, the recent deal we just announced is more revenue than the money we’ve raised.

Harry Stebbings

How does that cash hit? I know it’s an ARR throughout the year.

Jonathan Ross

Throughout the year, yes. But it’s this year, and there could potentially be more next year.

Harry Stebbings

A lot more?

Jonathan Ross

A lot more.

Harry Stebbings

What is that contract in 3 years?

Jonathan Ross

If we sell everything we possibly can this year, it’s many billions. From the capacity alone, there are tens of billions of dollars of hardware that we could build next year through these types of deals.

We’re doing it at high volume and low margin. If we were talking about GPU sales and GPU prices, we would be talking about hundreds of billions. We’re just not charging that much.

Harry Stebbings

The thing I’m most excited about is disease discovery and drugs. My mother has MS, and it was always taught to me that it was incurable. Now it’s actually, “Maybe it’s not.”

What are you singly most excited for?

Jonathan Ross

We went from a phase where people were hardware engineers to software engineers.

Being a hardware engineer is ridiculously difficult. The training you need is significant, you have to get things right, and there’s a real expense if you get it wrong.

Becoming a software engineer is much easier. All you have to do is get some time on a machine and teach yourself. Nowadays, you can download manuals, tutorials, or whatever from the internet.

I think prompt engineering is going to unlock a huge swath of human society. There are 1.3 billion or 1.4 billion people in Africa who know how to speak.

If you gave them access to a tool that allowed them to create applications live just by speaking to it, that would be another 1.3 billion or 1.4 billion potential entrepreneurs.

There are 8 billion people on the planet, and the difference is that hardware was ridiculously difficult. It was arcane knowledge that was hard to acquire.

Software was plentiful. Language is something you already know; you don’t have to learn a thing.

Harry Stebbings

What’s that going to do for venture? What’s that going to do for entrepreneurship?

Jonathan, I love talking to you. It’s always such a broad and wide-ranging discussion. Thank you so much for putting up with me in person. I’ve loved it.

Jonathan Ross

Awesome. I’m so glad to be here.