[BidClub_]
20VC · · 81 分钟

Groq 创始人 Jonathan Ross:OpenAI 与 Anthropic 将自研芯片,NVIDIA 能否涨到10万亿美元

Harry StebbingsJonathan Ross

YouTube
TL;DR
  • Ross 的核心判断是:「5年后,如果 Nvidia 的市值还不到10万亿美元,我个人会感到意外。」但更尖锐的版本是:Nvidia 只需销售少数芯片,却持续拿走大部分收入——可能是10%的芯片贡献51%的收入;随着贡献90%-99%总支出的35-36家客户开始按性能而非品牌采购,格局会变化。而他更希望你问的是:「5年后 Groq 的市值会到10万亿美元吗?有可能。」

  • 如果把 OpenAI 和 Anthropic 现有的推理算力都翻倍,它们的收入将在1个月内接近翻倍。Anthropic 最大的抱怨之一是速率限制;OpenAI 通过降低聊天速度来限流,牺牲了用户参与度。2周前,一位客户要求 Groq 提供其总容量5倍的算力,没人能接下这个需求。

  • OpenAI、Anthropic 以及所有超大规模云厂商都会自研芯片,但目的未必是击败 Nvidia。真正的奖品是「掌控自己的命运」,摆脱 Nvidia 对 HBM 的买方垄断——Nvidia 每年可以制造5000万颗 GPU die,但实际出货量约550万颗 GPU。至于新芯片创业公司,Ross 认为「已经太晚了,这艘船已经开走」:即使一切顺利,从设计到流片也要3年,而首次流片成功率只有14%。

  • 判断泡沫与否,应该看聪明钱在做什么:它们都在加码。在 Goldman Abu Dhabi 的一场活动上,管理着100亿美元以上资产的50多位经理中,没有1个人确信10年后 AI 不会取代自己的工作。「它们当然会像醉酒水手一样花钱」——而 AI 也打破了 SaaS 的经济学,因为每次查询投入更多算力,就能直接改善产品。

  • 中国模型的运行成本约为 GPT OSS 的10倍。它们优化的是低成本训练,而封闭市场的定价让供应商可以围绕稀缺性收费,导致人们把价格和成本混为一谈。中国可以赢下本土主场——拥有150座核反应堆和补贴;但在无法自行建设电厂的盟友国家这场「客场赛」中,美国还有2-3年的窗口期。

  • 「掌控算力的国家将掌控 AI,而没有能源就不可能拥有算力。」欧洲选择「通过立法」竞争,而不是建设基础设施;如果不能以足够速度行动,「欧洲经济将变成旅游经济」。拥有风电、且风电规模达到水电5倍的 Norway,理论上可以提供与美国同等的能源;日本的2nm晶圆厂、650亿美元 AI 投资和重启核反应堆,则展示了所需的行动速度。

  • AI 带来的将是大规模劳动力短缺,而不是失业。它会对整个经济施加通缩压力,让人们更早退出工作,并创造今天无法想象的新职业。「我们将通过生产更多算力,把更多劳动力加入经济体系……这在经济史上从未发生过。」

  • 行业格局上,长期对决仍是 OpenAI 对 Google;Anthropic 则在做不同的事——编码。5000亿美元的 OpenAI 和1800亿美元的 Anthropic 都「被严重低估」,因为这些实验室会成为「Mag 9、Mag 11、Mag 20」。Microsoft 短期内会重置与 OpenAI 的关系;CUDA 的锁定效应「对训练成立,但对推理不成立」。

摘要 · 为研究而整理的核心内容

1. 别问「是不是泡沫」,问聪明钱在做什么

  • Ross 重新定义了问题:如果一个问题始终得不到答案,就换个问法——Google、Microsoft、Amazon 和各国政府在做什么?答案是:全部在加码,每次发布新消息后都继续增加支出。他最有力的证据是,Microsoft 在1个季度内部署了 GPU,随后却没有把它们放进 Azure,因为自用带来的收入高于出租。

  • 他把当前市场类比为石油钻探早期:「很多干井,只有少数喷油井」,约35-36家公司贡献了99%的收入,或者至少贡献了大部分 token 消费。结果高度离散,直觉仍然胜过科学;「现在是投资者最好的时点」,一旦市场变得可预测,优秀投资者反而赚得更少。

  • 这些支出并不完全出于经济理性。在 Goldman Abu Dhabi 的首届活动上,他问50多位管理着100亿美元以上资产的人:谁能100%确定 AI 10年后不会取代自己的工作?没有人举手。他说:「很好,这就是超大规模云厂商的感受。」因此它们当然会像醉酒水手一样花钱,因为另一种结果是被彻底锁出自己的业务。

  • Stebbings 追问:无论是不是 Mag 7,回报最终都必须兑现。Ross 承认「没错」,但坚持认为价值已经在落地:一位客户提出一个功能需求,他「在给工程师做提示工程」,4小时后功能便已进入生产环境,且没有一行人工编写的代码。6个月后,这种事情可能在一场会议结束前就完成——这是质的差异,能够拿下竞争对手无法赢得的订单。

2. 无止境算力论:推理翻倍,收入翻倍

  • 这期节目的核心赌注是:「如果把 OpenAI 现有的推理算力翻倍,把 Anthropic 的也翻倍……从现在起1个月内,它们的收入将接近翻倍。」机制很直接:速率限制是 Anthropic 最大的抱怨之一;OpenAI 通过放慢聊天服务来控制流量,牺牲用户参与度。

  • Groq 自己的客户漏斗证明了约束所在:100%的客户最初都要求速度,但一旦看到供应短缺,就没人继续追问速度——「真正的价值主张是,你能不能提供更多算力容量」。2周前,一位客户要 Groq 总容量的5倍;Groq 和任何一家超大规模云厂商都无法承接。Harry 的本能反应是「要不要狠狠干脆买 CoreWeave?」得到的回答是,CoreWeave 是一家优秀公司,但和所有人一样,GPU 配额有限。

  • 过去10年的数据中心预测始终错在同一个方向:所有人都建设不足,于是把预测上调到高于最乐观情景,最后仍然建设不足。算力是「最容易调节的旋钮」——算法很少改善,数据很难获得,但算力每年都在进步,只要开出足够大的支票,就能相对可靠地买到。可即便如此,「我们还是会低估自己需要多少」。

  • Ross 承认自己在推理算力上的意外是:「我从没想到 AI 会建立在语言之上……我以为它会更像 AlphaGo。」语言让 AI 变得极易使用——「我原以为 AI 会更早到来、增长更慢;它来得更晚,增长速度却远超我的想象。」全球人口中有10%被描述为每周活跃的「GBT」用户。

3. 速度会复利成品牌:认为未来用户不在意延迟是「完全错误」

  • 对于用户会愉快地发出提示词然后离开的说法,他的回答是:「完全错误。」他按毛利率排列 CPG:烟草、嚼烟草、软饮料,最后是水;毛利率对应的是成分作用于人体的速度。多巴胺循环会建立品牌黏性,Google 和 Facebook 都建立在这一点上;速度每提升100毫秒,价值约相当于8%的转化率。

  • Groq 早期展示速度时,观众曾问:既然输出速度比阅读速度还快,为什么还需要更快?他的反问是:「网页为什么要比你阅读它的速度更快地加载?」人们始终不擅长预测什么会驱动参与度,这是早期互联网留下的教训。

4. 所有人都会造芯片:为了掌控命运,而不一定为了超越

  • TPU 被视为超大规模云厂商自研芯片可行的证明,但它只是 Google 同期约3个芯片项目之一,也是唯一击败 GPU 的项目;Dojo 刚刚被取消。造一款芯片挑战 Nvidia,「有点像说 Google Search 很不错,我们复制一个吧。这是疯了」。

  • 但他「毫不怀疑」OpenAI、Anthropic 和所有超大规模云厂商都会造芯片,关键在于动机。他讲过一次 Google 实验室参观经历:Google 造了1万台 AMD 服务器,随后把芯片拆下来扔进垃圾桶——这从一开始就是安排好的,因为真正目的,是用 AMD 芯片去压低 Google 实际采购的 Intel 芯片价格。

  • 真正的瓶颈是 Nvidia 对 HBM 的买方垄断:GPU die 使用的是手机芯片工艺,Nvidia 每年可以制造5000万颗,但 HBM 和 interposer 产能把今年的出货量限制在约550万颗 GPU。当一家超大规模云厂商要求100万颗 GPU、被拒绝并威胁自研时,「突然之间,那些 GPU 就找到了」。自研芯片购买的是「掌控自己的命运」,即使成本更高也值得。

  • Sarah Hooker 所说的「硬件彩票」是新进入者的陷阱:模型是围绕现有硬件设计的,attention 在 GPU 上运行良好,因此更好的架构「反而跑不好」,也就不再是更好的架构。现有厂商可以提前2年规划,挑战者则需要更快的迭代循环。

5. 1000亿美元「无限资金循环」与10万亿美元的 Nvidia

  • 对于 Nvidia 向 OpenAI 投入1000亿美元,Ross 的说法是:「如果实际产生了有生产力的结果,就不叫资金空转」——至少40%会流向建设基础设施的供应商。Harry 说,那就是一个不完整的循环:60%回到 Nvidia,再加上几千亿美元的市值增长。Ross 的回答是,从经济角度看,「为什么不一直这么做」;只有在锁定效应成立时,这种溢价才合理,「而我会说 Nvidia 很可能确实如此」,因为全球根本没有足够的算力。

  • 5年后的图景是:Nvidia 可能用约10%的芯片拿走超过50%的收入。品牌仍然有价值——「买 Nvidia,没人会因此被解雇」——但贡献90%-99%总支出的35-36家客户将拥有按商业价值选择的能力,因此其他芯片也会被采用。

  • Nvidia 5年后市值超过还是低于10万亿美元?Ross 回答:「5年后,如果 Nvidia 的市值还不到10万亿美元,我个人会感到意外。你真正应该问的是,5年后 Groq 的市值会到10万亿美元吗?有可能。」Groq 的主张是:由于供应链约束更少,「我们几乎可以生产无限数量」的芯片,满足市场上最稀缺的资产。

  • 关于 Nvidia 最大的误解,他的快速回答是:「Nvidia 的软件构成护城河。」CUDA 的锁定效应「对训练成立,但对推理不成立」——Groq 已有220万名注册开发者,而 CUDA 自称有600万名。

6. 芯片经济学:跨过2道门槛,新进入者已经错过时机

  • 一款芯片的生命周期有2套计算:部署时只需战胜资本开支,持续运行时只需战胜运营开支。全行业的赌注是,新芯片不会把旧芯片的价值压到运营成本以下,这也是接近5年机龄的 H100 仍能盈利出租的原因:「今天你绝不会部署 H100,但它们运行起来仍然有利润。」只有算力短缺还能维持这一点。

  • Groq 拒绝做长期赌注。Ross 认为,人们对摊销周期的设想比他更长;他先说「5到6年」,Harry 说「那大概就是3年」后,他又改口说「短一点」。Groq 大约每年升级一次芯片,因为芯片到了5年机龄,收入会低于电力和数据中心成本。持有长期合约的人还要做第3个计算:违约成本是否低于继续亏损运行?那之后会发生什么?「我不能告诉你,因为我们正在努力避免这种情况。」

  • 如果今天创办 Groq,他「不会做芯片」。「这艘船已经开走了。」即使完美执行,从设计到量产也要3年,而首次流片成功率只有14%。Ross 做过3款芯片,全部一次流片成功;V2 甚至已经安排了重新流片,结果令他们震惊的是,根本不需要。Groq 现在采取一年一代的节奏,从 V2 到 V3 再到 V4;Nvidia 每款芯片则采用流水线方式,需要3-4年。

  • 真正的护城河是供应链:相比受 HBM 约束的 GPU 供应链,「你给我们开一张100万颗 LPU 的支票,第一批 LPU 6个月后就开始交付」——GPU 则要2年。向一家超大规模云厂商的基础设施负责人推销时,速度和成本只得到礼貌性点头;6个月供应周期「是他唯一在意的事情」。

7. 破解中国误解:主场赛与客场赛

  • 「我们干脆把所有误解都逐一拆掉。」中国模型的运行成本并不更低,而是 GPT OSS 模型的「约10倍」。它们优化的是低成本训练,封闭市场的定价让供应商可以围绕稀缺性收费:「人们把成本和价格混为一谈。」OSS 模型压缩出的智能说明「美国仍然拥有训练优势」,而这种优势建立在芯片获取能力之上。Ross 承认:「一开始连我也有点被蒙住了。」

  • 中国可以赢下自己的「主场赛」——计划建设150座核反应堆,并用补贴覆盖低效率芯片的成本。但「客场赛」不同:一个拥有100兆瓦电网的盟友,不可能随手建一座核反应堆,因此能源效率更高的芯片会胜出。「未来2-3年,美国在这场客场赛中拥有明确优势」——前提是行动足够快,把盟友带进来。

  • 开源策略方面,他曾仅凭品牌力量预测 OpenAI 会开源:「它们大概可以使用 Llama 2……人们仍然会用。」Anthropic 应该「开源上一代模型」,因为提示词兼容性就是新的软件兼容性:用户选择 OpenAI 的 OSS 模型而不是中国模型,是因为原有提示词可以复用;低成本用户日后还会升级到高级模型。

  • GPT-5 强调效率,并不意味着放弃规模化。要赢得印度市场,价格必须做到每月99卢比(约1.13美元),因为替代方案根本不是另一款 AI,而是完全没有 AI。

8. 「掌控算力的国家将掌控 AI」:欧洲旅游经济警告

  • 他希望所有人带走的一句话是:「掌控算力的国家将掌控 AI。而没有能源就不可能拥有算力。」这并不是只有核能一条路:Norway 的风电利用率约80%,风电规模达到水电5倍,因此「Norway 一个国家就可以提供与美国同等的能源」——而且是持续提供。

  • 欧洲的失败路径,是拥抱「不作为的错误」,通过立法而不是建设基础设施来竞争:制定数据驻留规则,却不增加兆瓦级电力;在增长型经济中,「错过机会的代价高于把事情搞砸」。许可审批的腐化也发生在美国:一位核电公司董事告诉他,公司在审批上的支出是电厂建设成本的3倍。为什么不建核电?「恐惧。」在这一点上,Trump「确实有帮助」。

  • Japan 是他衡量行动速度的基准:决策慢,行动快;一座2nm晶圆厂已经开始生产晶圆,虽然良率还未达到量产标准;政府已拨出650亿美元用于 AI,核反应堆也在重新上线。「当 Japan 要重启核反应堆时,欧洲需要听进去。」

  • 如果欧洲不能以足够速度行动,「欧洲经济将变成旅游经济」。人们会来到这里参观古朴的老建筑,仅此而已。模型主权救不了欧洲——「你可以拥有一个比 OpenAI 聪明10倍的模型,但如果 OpenAI 拥有10倍算力,OpenAI 的模型就会更好。」他与 Mistral 合作,也喜欢这家公司;解决方案是增加算力,这样 Mistral 很可能就能竞争。过渡方案是 Saudi 的「数据大使馆」:在其3-4GW的建设规模上,叠加主权数据监管。

9. AI 打破 SaaS 经济学:创造的是劳动力短缺,而非失业

  • AI 不等于 SaaS,因为产品质量并不会在交付时固定下来:「我可以运行2个提示词实例,然后选出更好的答案,来提升产品质量。」这也是为什么 token-as-a-service 的账单几乎接近收入,以及为什么 OpenAI 刚刚宣布面向有限用户、价格更高且消耗大量算力的产品——「我们想看看,当给 AI 更多算力时会发生什么」。

  • 在毛利率问题上,他的观点并不主流:高毛利存在的唯一理由,是为波动提供缓冲;「你的毛利率就是我的机会」。他希望 Groq 的毛利率「在保持业务稳定的前提下尽可能低」,因为「信任会产生利息」,而 Jevan 悖论会推动剩下的增长:「如果我们生产10倍算力,就会实现10倍销售额。」

  • 宏观链条是这样的:「经济中最有价值的东西是劳动力。现在,我们将通过生产更多算力和更好的 AI,把更多劳动力加入经济体系。这在经济史上从未发生过。」结果有3个:大规模通缩压力——机器人种植、基因工程咖啡、更便宜的住房;人们选择退出工作——减少工时、更早退休;以及今天无法想象的新职业——农业人口占美国劳动力的比例从98%降至2%,「100年前,人们也无法想象 influencer 这种职业」。最终结果是「大规模劳动力短缺」,而不是大规模失业。

  • 对于 S&P 接近7000点是否过热,他把称重机和人气投票分开看。他从未买过 Bitcoin,因为「我无法参与人气投票」。AI 正在创造称重机式价值的信号,是私募股权公司纷纷追逐 Groq,寻求低成本算力来改变被投企业的利润底线。经济下行在原则上无法预测——「如果一个预测会影响这个预测本身,你就无法预测它」。他判断过热的标准是,经济是否开始阻碍企业成功;他指出的唯一扭曲是,每个优秀工程师都在筹集1000万、2000万、1亿美元甚至10亿美元,去创办又一家实验室,而不是加入现有实验室。「请停止这么做。」

10. 行业格局:OpenAI 对 Google,两家实验室都被「严重低估」;差异化才能生存

  • 长期对决是 OpenAI 与 Google——「Anthropic 做的是不同的事」:编码。Google 是他见过最令人印象深刻的翻身案例,工程师驱动的文化构成「系统性优势」;Gemini 按用户采用数据看是成功的,但在 Gmail 中的消费者整合「几乎不可用」。这些都是值得下注的尝试,类似 Google TV 做 Chrome 的路径;但在全球人口中有10%被描述为每周活跃的「GBT」用户后,他承认:「Google 可能已经太晚了。」

  • 5000亿美元的 OpenAI 和1800亿美元的 Anthropic,选哪一家?「我想两家都投……它们都被低估了,而且是严重低估。」错误在于把市场当成有限蛋糕:这些实验室会通过研发扩大市场,最终只是加入指数,成为「Mag 9、Mag 11、Mag 20」。在现有巨头中,Microsoft 会短期重置与 OpenAI 的关系,但会保留已经部署的算力——「算力就像黄金」;Amazon 没有 AI DNA,但拥有算力;Meta 和 Google 则一直拥有 AI DNA。

  • 说到 Elon,xAI 大概能做成,但路径不同:它有编码模型,却「没有编码分发渠道」;Anthropic 的高明之处,是拒绝什么都做。「如果你不差异化,就会死。」

  • Groq 的底线是:不会创建自己的模型,让客户可以放心在其平台上构建——「我在这个判断上可能犯了一个巨大的错误。」Groq 融资7.5亿美元,估值接近70亿美元;原计划融资3亿美元,最终超额认购4倍,并拥有正向硬件毛利率。公司内部要求工程师使用 AI,但工具由他们自己选择:先是 Sourcegraph,然后是 Anthropic,现在是 Codex,「下个月可能又会回到 Sourcegraph」。他最后借用 Galileo 的比喻:LLM 是「心灵的望远镜」——它们现在让我们感到渺小,但「100年后,我们会意识到,智能比我们想象的广阔得多,并且会觉得这件事无比美好」。

Jonathan Ross

The countries that control compute will control AI. You cannot have compute without energy. Now we're going to be able to add more labor to the economy by producing more compute and better AI. That has never happened in the history of the economy before.

Harry Stebbings

What is that going to do?

Jonathan Ross

I personally would be surprised if, in 5 years, NVIDIA wasn't worth $10 trillion. But I can't predict the outcome. The demand for compute is insatiable. I would wager that if OpenAI were given twice the inference compute that they have today, and if Anthropic were given twice the inference compute that they have today, within 1 month from now, their revenue would almost double.

Harry Stebbings

I'm sorry. Can you unpack that for me?

So, I'm thrilled to welcome Jonathan Ross, founder and CEO of Groq, back to the hot seat. Jonathan, you've just been told by our team that our last show was the most successful of the year when it came out, so there's no pressure at all for this to be the most successful of this year. But welcome to the studio, man.

Jonathan Ross

Thank you.

1. Analyzing the Current Market Landscape

Harry Stebbings

It's great to have you here, dude. I wanted to start with an understanding of where we are. It seems the world moves faster than ever before, and honestly, I think a lot of us are trying to understand where everyone lies in a new market. If we look at the current state of the market today, how do you analyze it?

Jonathan Ross

Are you asking whether there's a bubble?

Harry Stebbings

Relatively.

Jonathan Ross

Okay. In terms of whether or not there's a bubble, my answer is: if you ask a question and you keep not getting an answer, maybe you should ask a different question. Instead of asking, “Is there a bubble?” you should ask, “What is the smart money doing?”

What is Google doing? What is Microsoft doing? What is Amazon doing? What are some nations doing? They're all doubling down on AI. They're spending more. Every time they make an announcement about how much they're spending, it goes up the next time.

One of the best examples of the value that's coming from this spend is Microsoft. In one quarter, Microsoft deployed a bunch of GPUs and then announced that they weren't going to make them available in Azure because they made more money using them themselves than renting them out. So there's real money in the market.

The best way that I think to explain this market is like the early days of oil drilling: a lot of dry holes and a couple of gushers. I think the statistic that I heard was that 35 or 36 companies are responsible for 99% of the revenue, or at least the token spend, in AI right now.

Harry Stebbings

I'm surprised it's not less when you look at NVIDIA. No, but I mean, seriously, NVIDIA has such a heavy concentration of revenue with 2 clients.

Jonathan Ross

Yeah. Maybe NVIDIA represents 98% of that. But when it's that lumpy, what that's an indication of is that it's like the early days of oil drilling, where people didn't know how to find oil. They were going off instinct—almost vibe investing. People who had a good instinct would make a fortune, and everyone else would lose their shirts.

Over time, it becomes a science. It becomes very predictable, and there's less lumpiness. There's more predictability, but investors make less money at that point. The good investors make less money. So right now is the best time for investors. Right now, people are making more money than they're spending. It's just very lumpy.

2. Why the Hyperscalers Have to Keep Spending Recklessly on AI

Harry Stebbings

I'm sorry. They're making more money than they're spending? But as an aggregate, plenty of people are going to lose their shirts, but overall, less money is going to go in than is going to come out.

When we look at the capex spend today by the big providers, everyone is going, “Okay, okay, okay,” because there's something coming at the end of it.

Jonathan Ross

Yeah.

Harry Stebbings

And the trouble is, the capex spend is going up and up and up.

Jonathan Ross

Okay. You're thinking of it purely financially, and I think that the financial returns will be positive, but that's not why people are motivated.

I was in Abu Dhabi at the inaugural Goldman Sachs Abu Dhabi event, and as you now know, we're sponsoring McLaren. Zak Brown was talking, I was talking, and it was a fun event. But I was asked a similar question: “Is AI a bubble?”

I asked the following question. This is a bunch of people who manage $10 billion-plus in AUM, right? The entire 50-plus people who manage $10 billion-plus. I asked, “Who here is 100% convinced that in 10 years AI won't be able to do your job?” No hands went up. Great. That's how the hyperscalers feel.

Of course, they're going to be spending like drunken sailors because the alternative is that they're completely locked out of their business. So it's not a purely economic framework that they're using. It's a question of whether they get to maintain their leadership.

When you look at it, the next step is that there are these scaling-law sort of outcomes. You want to remain in the top 10. We keep talking about the Mag 7. If you're not a member of the Mag 7, you're not going to be able to get anywhere near the valuation. So what do you do to stay there? You spend, and it's worth it because the stock value stays up because you're in the top 7 or 10.

Harry Stebbings

At some point, the returns have to be delivered, though. The spend has to materialize into actual, tangible revenue, and if it doesn't, whether you're in the Mag 7 or not doesn't matter.

Jonathan Ross

Correct. But right now, AI is returning massive value already. It's very lumpy in the applications, but it's returning massive amounts of value.

Let me talk about an example that actually happened for us. I've tried a little bit of vibe coding. I'm not the best in the world at it. We've got some interns who are amazing at it. We had a customer visit us, and I had a meeting with them. They asked for a feature, and I specced it out at a very high level, vibey. I was prompt-engineering the engineers, and 4 hours later, it was in production.

Not a single line of code was written by a human being. There was no debugging done by a human being. It was all prompting. I think we even have Slack integration now, where you commit things through Slack. So all that was done, and 4 hours later, it was in production.

Think about the value there. Now imagine, fast-forward 6 months from now, when that could happen before the customer meeting is over. It's a qualitative difference. It's not even just a dollar-amount difference.

Harry Stebbings

Yes.

Jonathan Ross

When you're able to do it that fast, you spend less to get the feature into production. That's real ROI. However, qualitatively, when you can do that before the customer meeting is over, you're going to be able to win deals that your competitors won't.

Harry Stebbings

Can I ask you, just going back to the Mag 7, to stay in the Mag 7: do you think everyone realizes that they will need to move into the chip layer and own the full vertical end-to-end?

Jonathan Ross

I don't think you're going to see too many successfully moving into the chip layer. People look at the TPU as a big success, and what they don't realize is that there were about 3 chip efforts at Google at the same time, and only 1 of them ended up outperforming GPUs.

When you look around the industry, you've got a bunch of people building chips. Some of them are getting canceled, like Dojo recently got canceled. Building chips is hard. Going off and saying, “I'm going to build my own AI chip to compete with NVIDIA”—it's a little bit like saying, “Google Search is pretty nice. Let's go replicate it.” It's insane.

The level of optimization, the level of design and engineering that goes into that—you're not going to be able to replicate it with a high probability of success. However, if there's a bunch of players out there trying to do it and you have optionality, and 1 of them succeeds, then you have another chip.

We mentioned earlier that you have to spend if you want to stay in the Mag 7.

Harry Stebbings

Mhm. NVIDIA investing $100 billion into OpenAI for OpenAI just to go and buy back NVIDIA chips. Is this not just an infinite money loop?

Jonathan Ross

That would be the case if they weren't spending it with suppliers to build those chips. It's not round-tripping if actual productive outcomes are occurring.

Think of it this way: what percentage of the spend is going to building that infrastructure? 40%? So at least 40% of those dollars are actually going out into the ecosystem. That is not an infinite loop.

Harry Stebbings

Okay, so it's a partial loop. 60% is going back to NVIDIA.

Jonathan Ross

Sure.

Harry Stebbings

And then they get a bump in their stock price of a couple hundred billion dollars.

Jonathan Ross

Yes.

Harry Stebbings

How did you analyze that?

Jonathan Ross

Let's analyze it in a couple of different ways. From an economic point of view, it makes perfect sense. Why not do that all day long?

The value occurs if there is lock-in. When revenue increases result in stock-price increases that are greater than the amount of the revenue, it's because you believe that that revenue is going to continue. That's the belief.

I would actually say with NVIDIA, that's probably true. However, it's not just because NVIDIA is good—and NVIDIA is very good—it's also because there isn't enough compute in the world. There isn't. The demand for compute is insatiable.

Right now, one of the biggest complaints about Anthropic is the rate limits. People can't get enough tokens from them.

If they had more compute, they could produce more tokens and charge more money. With OpenAI, it's a chat service. So how do you regulate your chat service? You run it slower. You get less engagement.

Harry Stebbings

How important is speed, do you think? There's a lot of people who think, actually, it's fine. I'm very happy to have latency, and I'm very happy to have a prompt, and then I go away, do something else, and something happens when I'm away.

Jonathan Ross

Those are interesting opinions. Let's look at CPG, so consumer packaged goods. I want you to rank the CPG goods by margin. At the very top is tobacco—smoking tobacco. Below that is chewing tobacco. Below that is soft drinks. Below that, you keep going down, and you get to water and other things like that.

What is the number one thing that a high margin correlates to in CPG? It's the speed at which the ingredient acts on you. So that dopamine cycle—how quickly something occurs—determines your brand affinity. When something has a very quick response, you associate with that brand, and then you accrue brand value.

This was the entire basis of Google focusing on speed, and Facebook focusing on speed. Every 100 milliseconds of speed-up results in about an 8% conversion rate. So that is wrong in terms of people's assessment of the future, where they think, "Oh, it's fine. We'll actually just have lots of prompts going on in the background, and we'll be happy to let them run for long periods of time."

100% wrong. In fact, when we first started working on getting speed on our chips, we knew what speed we could get. We even made a video example of how fast we could be. People would look at that video example and say, "Why does it need to be faster than you can read?"

I would respond to that by saying, "Why does a web page need to load faster than you can read?" There's just this mental disconnect where people couldn't grasp the sort of visceral importance of speed. People are very bad at determining what's actually going to matter in terms of engagement and outcome, but we know this from building the early internet companies.

3. Why OpenAI and Anthropic Will Have to Build Their Own Chips

Harry Stebbings

Do you think OpenAI will be able to move into the chip layer? At some point, NVIDIA must be concerned that OpenAI will want to verticalize and own the chip layer as well. Do you think they will be able to make that successful transition?

Jonathan Ross

I think one of the problems in building your own chip is, first of all, everyone thinks that building the chip is the hard part. Then, as you do it, you start to realize building the software is the hard part. Then, as you do it, you realize keeping up with where everything is going starts to become the hard part.

I have no doubt that OpenAI will be able to build its own chips. I have no doubt that eventually Anthropic will be building its own chips, and that every hyperscaler will build its own chip.

One of the things—I had this experience when I was at Google, where I got a lab tour. This was before AMD was doing a great job, right? AMD was struggling for a little while, and now they're doing great. They had built 10,000 servers, and those 10,000 servers with AMD chips—I was walking through the lab, and they were pulling the servers out of the racks, taking the AMD chip, popping it off, and throwing it in a trash can.

The funny thing was, it was almost preordained because everyone knew that in that generation, Intel was going to win. So why did Google build 10,000 servers? Because they wanted to get a discount on the Intel chips they bought.

When you're at that scale, the cost to design your own server—they had to design their own motherboard in order to fit the AMD chip—and to build that out and test it, versus the discount that you get, is totally worth it. So you have to think of what all the motivations are when people are building their own chips. It's not just because they're going to deploy that chip in mass production.

The thing is, NVIDIA effectively has a monopsony on HBM. A monopsony is the opposite of a monopoly. When you're a single buyer and there's a finite amount of HBM capacity—which is the high-bandwidth memory that goes into the GPUs—the GPU itself is made using the same process that's used to build the chip that's in your mobile phone.

If NVIDIA wanted to, they could build 50 million of those GPU die per year, but they're going to build about 5.5 million GPUs this year. The reason is because of that HBM, because of the interposer that it goes on. There's just a finite amount of capacity.

What happens is, a hyperscaler comes in and says, "I want 1 million GPUs." NVIDIA is like, "Sorry, I've got other customers." The hyperscaler says, "No problem. I'm going to build them myself." Then, all of a sudden, those GPUs are found by NVIDIA to give to the hyperscaler.

There's just a finite amount of capacity. By building your own chip, what you really get isn't your own chip. It's that you get control over your own destiny. That's the unique selling point of building your own chip.

Harry Stebbings

What does that mean, control over your own destiny?

Jonathan Ross

NVIDIA can't tell you what your GPU allocation is. It may cost you more to deploy your own chip because it's not going to be quite as good as NVIDIA's.

Let's think about why NVIDIA's GPUs, with a slight edge over AMD's GPUs, dominate. If your total cost to deploy is a huge multiple of the cost of the chips and the systems, then a small percentage increase in the cost of the chip is negligible.

Think about it this way: if I'm going to deploy a CPU, and that CPU is 20% of the BOM, and I get a 20% increase in the speed of the chip, that is a 20% value increase in the entire system versus a 20% increase in the chip cost, right? It's negligible.

You get these huge multiples when you improve the chip performance. Small differences in performance make a huge difference in the value of the product. So a small edge gives you a massive edge in selling that product.

Harry Stebbings

Can I ask you—you mentioned the monopsony?

Jonathan Ross

Yes.

Harry Stebbings

Yeah. Is it possible for OpenAI, Anthropic, any of the Magnificent 7, or any of the other providers to move into the chip layer if there is a monopsony in the HBM market?

Jonathan Ross

It's very hard. However, there is an incentive from those building HBM to spread that around, because NVIDIA gets to negotiate very good rates because they're such a large buyer.

However, if you're building an HBM fab and packaging house and all of this other part of the ecosystem, if NVIDIA comes in and writes a big check, then you're going to build the fab for them. So NVIDIA is always going to get the amount of supply that they want in advance.

The problem is, you have to write that check more than 2 years in advance. With where AI's gone—just absolutely honking, hockey-sticking—even when you have the cash flow of NVIDIA, it's hard to actually write the checks for the amount of demand that's going to be there in advance.

There is going to be a supply constraint, and it's not purely based on being a monopsony. Part of it is based on just the sheer capital costs, and the memory suppliers are very conservative. There's also this situation where the margin on HBM is so high that no one wants to actually increase the supply, because then the margin goes down.

4. OpenAI and Anthropic Will be $5BN Companies: The Bull Case

Harry Stebbings

I totally understand that. Can I ask you, when you look at that, and when you look at OpenAI, when you look at Anthropic having their own chips, is that why they're raising the money they are? Sam said they're going to need hundreds of billions of dollars. Is that factoring that in?

Jonathan Ross

No. Most of the spend is on buying systems. Buying a system is expensive. Buying a data center is more expensive. The reason is, you're amortizing that data center over a longer period of time.

Even if a data center was going to be one-third of your cost per year, if you're amortizing that data center over 10 years and the chips over 3 to 5 years, the data center is going to end up costing you more per year.

When you hear the hyperscalers talking about that $75 billion to $100 billion a year investment because they're building out the capacity for data centers, they're putting a lot of money up for returns that they're expecting over the next 10-plus years. So it's actually not that much money when you think about it.

Harry Stebbings

Are we thinking about amortization in the right way in a 3-to-5-year cycle if chip cycles are actually faster than that?

Jonathan Ross

I think that the amortization—people are definitely thinking about it over a longer period than I would. We use a more conservative number internally. I think 5 to 6 years.

Harry Stebbings

Which would be like 3 years?

Jonathan Ross

A little bit less.

Harry Stebbings

Yeah.

Jonathan Ross

We're looking at upgrading chips about once a year.

Harry Stebbings

Yeah.

Jonathan Ross

Now, here's the way to think about it: there's 2 phases of the value of a chip. There's, "Am I willing to buy it and deploy it?" and there's, "Am I willing to keep it running?" They're 2 very different calculations.

When you deploy it, you have to be able to cover the capex. When you keep it running, you just have to beat the opex. So if I deploy a chip today, I have to beat the capex. I have to earn all my capex back, make a profit, and produce a return once I've deployed it.

As long as I'm beating my operational costs, I'm going to keep that thing in production. So you're okay with the price—the value of that chip—going down over time. Now, the bet that everyone is making is that those new chips that come out aren't going to reduce the value of the old chips below the opex.

Harry Stebbings

That's it.

Jonathan Ross

That's right. And in our case, we actually don't think that 5 years makes any sense.

Because they will be so much less performant that, actually, the value will be lower than the operating cost for the electricity and for paying for the data center.

So what happens then?

We just have this excess supply of wasted chips, which are going to waste because a lot of these people have entered into really long contracts. They have a third point that they have to consider in their calculation: Is breaking this contract cheaper than running the chip at a loss?

Harry Stebbings

Yeah. Can you see this? So what happens then?

Jonathan Ross

Then I can’t tell you what happens because we’re trying to avoid that situation. By having a much faster payback period in all of our calculations, I would not want to make a bet that far out. The shorter the time frame that you’re making the bet, the clearer your outcome is.

Harry Stebbings

So, essentially, you want to minimize the payback period as much as possible and then minimize operating costs so that you can shed less performant chips faster.

Jonathan Ross

Yes. But also, here’s another crazy part: When you look at the math this way, you’re like, “If I’m approaching it as an accountant, I’m going to be like, this is a terrible idea.” But if I look at it empirically, people are still renting H100s. How old are those chips? They’re getting close to 5 years old, and they’re still operating well. They’re still earning more than their operating cost by quite a bit.

You would never deploy an H100 today, but they’re still profitable to run, right? They’re in that second phase. The reason is that people can’t get enough compute. If that wasn’t the case, H100s would be renting for a fraction of what they’re renting for today. As long as you can’t get enough compute, that’s going to be true.

The question is, is there an alternative out there that isn’t supply-constrained? This is where we’re hoping to come in. Let’s talk about our value proposition. You started off asking me about speed. Do you know how many customers come to us asking for speed?

Harry Stebbings

No.

Jonathan Ross

100%. Do you know how many customers keep asking about that once they realize the supply constraint out there?

Harry Stebbings

None.

Jonathan Ross

So they start with speed because they know the value of that to their end customer, and then they’re like, “Oh, wait a second. I can’t even get enough compute.” The real value proposition is: Can you provide more compute capacity?

Two weeks ago, we had a customer come to us and ask for 5x our total capacity. They couldn’t get that capacity from any hyperscaler. They couldn’t get it from anyone else. We couldn’t give it to them. No one can. We couldn’t get that customer, the hyperscalers couldn’t get that customer, and there isn’t enough compute.

When you’re in a market where there isn’t enough compute, your choice is: I buy this compute and I get the customer. This is where I was going when I said, if OpenAI or Anthropic were to double their compute, they would double their revenue, right?

If you’re someone who can’t get enough compute to serve your customer, then you’re going to be willing to pay whatever it takes to get those customers because you feel that there’s lock-in value by getting that customer now. The number-one value proposition that we have is that our supply chain is not like a GPU supply chain. You have to write a check 2 years in advance to get GPUs. For us, you write us a check for 1 million LPUs, and the first of those LPUs starts showing up 6 months later.

Harry Stebbings

Wow. So you’ve got an 18-month chasm of difference.

Jonathan Ross

That’s right.

Harry Stebbings

Wow.

Jonathan Ross

I had a meeting with the head of infrastructure of one of the hyperscalers, and I talked about speed, cost, and all this stuff. But when I talked about the supply chain and how we could do something in 6 months, he just stopped the conversation for a moment and wanted to dig into that. That was the only thing he cared about.

Think about it this way.

Harry Stebbings

Given the speed of progression of the landscape of models, does 2 years make sense?

Jonathan Ross

Well, do you know Sarah Hooker?

Harry Stebbings

No.

Jonathan Ross

She wrote this paper, The Hardware Lottery: The Devil and the Details of AI. My TL;DR on that one is that people are designing the models for the hardware.

There are architectures that could be better than attention. However, attention works really well on GPUs. So, if you are the incumbent, you have an advantage because people are designing their models for your hardware. It doesn’t even matter if there’s a better architecture out there. It’s not going to run well, so it’s not a better architecture. There’s a little bit of a loop there.

If you are building 2 years out and you’re the incumbent, that’s okay. But if you’re trying to enter the market, no one’s going to design for your chips 2 years out. So you have to have a faster loop.

Harry Stebbings

When you see everyone moving into the chip layer, as you said—OpenAI will have their own, Anthropic will have their own—what does NVIDIA do in that world?

Jonathan Ross

NVIDIA still keeps selling chips.

Harry Stebbings

To whom, given the concentration of their buyers?

Jonathan Ross

No one is successfully predicting how fast AI is progressing. We started off talking about whether AI is a bubble. If you look at infrastructure for data centers, you’re planning that out 2, 3, 4, or 5 years in advance, right? What happens is everyone’s predictions are wrong. They end up building too little. This has just been what’s happened for the last 10 years.

If you don’t build enough for 10 years, what do you do? You try to overbuild. You try to build more than your most optimistic projections, and then once again, you haven’t built enough. So you increase your projections, and you just keep doing this. That’s what’s been happening. Yet people still aren’t building enough compute.

Where people’s instincts are off—and this just hasn’t, I think, been recognized yet—is that AI doesn’t work the way SaaS does. In SaaS, you have a bunch of engineers who go out and build a product, and the quality of that product is determined based on what those engineers did. That’s not the case in AI.

In AI, I can improve the quality of my product by running 2 instances of the prompt and then picking the better answer. I can actually spend more to make my product better on each query. I can even decide, “This customer is more valuable, and I’m going to give them a better result.”

That’s kind of what OpenAI announced—and they did this this week—when they said, “We’re now going to release some products where we can’t really afford the compute, so we’re going to give it to a limited set of users and charge more because we want to see what happens when we give more compute to the AI. We want to see what that product looks like and how much better it is.”

5. Why China is Behind the US in AI and Deepseek is More Expensive to Run

That is going to be our future. Every time you give more compute to an application, the quality increases. This is why it’s not coincidental that you see people’s tokens-as-a-service bill almost matching their revenue, because they’re competing for customers, and if they just spend more, their product gets better.

Harry Stebbings

I totally understand that. But bluntly, the assumption when you look at GPT-5 and the focus on efficiency is that Sam’s transition from performance to efficiency is because compute does not equal a parallel level of performance improvement. Do you think that is fair and true, and does that not go against what you just said?

Jonathan Ross

No. You have to think of the different outcomes that they’re looking for. If you are OpenAI, you have moved into markets that are incredibly cost-sensitive.

Let’s talk about India for a second. If you want to go win India, what’s the one thing you need? 99 rupees a month. That’s about $1.13 with current conversion rates. You need to charge your customer $1.13 for your product. So they’re going after a market whose alternative is, “I have no AI.”

Harry Stebbings

You’ve got OpenAI. I mean, they can use DeepSeek.

Jonathan Ross

This is another misconception in the market. Let’s just start busting every misconception.

Harry Stebbings

Sure. Great.

Jonathan Ross

When the Chinese models came out, everyone reacted by saying, “Oh my God, they’ve trained models that are almost as good as the US models.” We had a podcast on this, right? Even I was snookered a little bit at first. I thought, “Oh my gosh, aren’t these models so much cheaper to run?”

Now that I know more about the foundation models that people are using versus the Chinese models, no, they’re not cheaper to run. They’re about 10x as expensive.

Actually, let’s just take the GPT-OSS model that was released. It’s optimized for something different than the Chinese models, but the quality is very high. I would argue it’s clearly a better model for what it focuses on than the Chinese models. Now, the Chinese models focus on different things.

However, the cost to run the GPT-OSS model is about 1/10 that of the Chinese models. So why was everyone charging less? When you have a captive market for a model because people say, “I want this model, and there’s only 1 provider of it,” you can charge 10 times as much. The price was higher, and people were confusing the cost with the price.

The Chinese models were optimized to be cheaper to train as opposed to being cheaper to run. When you see how much intelligence has been squeezed into the GPT-OSS model versus the equivalent Chinese models, it’s clear that the US still has a training advantage.

The economics work out such that you have to amortize that training over every inference, which means that you want to charge more. There’s still a balance there. But as you scale out into larger and larger numbers of people, being able to afford to train a model starts to be a payoff. As you deploy more inference capacity, you want to spend a bit more on the training to get your inference cost down. In the US, we have a massive compute advantage, and so people train the models harder, bringing the cost down.

Harry Stebbings

Why do we have a compute advantage in the US, just in terms of access to chips?

Jonathan Ross

That’s correct.

Harry Stebbings

Yeah, and so will China not just subsidize the inference and the running, though?

Jonathan Ross

Yes.

Harry Stebbings

So does it matter if their cost of running is higher, but the Chinese—the CCP—will just subsidize it? Does it matter?

Jonathan Ross

There's a home game and there's an away game. The home game is that we want to build enough compute for the United States. The away game is that we want to build it for our allies, right? Europe, South Korea, Japan, India, and so on.

The advantage that the U.S. has is in the away game, so China can win its own home game. They're going to build 150 nuclear reactors, so they're going to have enough energy even though their chips aren't as energy-efficient. They can subsidize it, as you mentioned.

But the away game is different. If a country only has 100 megawatts of power, what are they going to do? Build another nuclear power plant? That's just not a realistic thing. You can do that in China; you can't do that elsewhere. So having a better chip gives you an advantage in the away game.

My expectation is that, right now, for the next 2 to 3 years, the United States has a clear advantage in that away game over China. If we move very quickly, then we're going to be able to bring a bunch of allies into the AI race.

Harry Stebbings

Do you think we should have open models to allow China to distill in the effective ways that they have done already?

Jonathan Ross

I think the model itself is not a clear advantage. So the first time you had me on your podcast, I predicted that OpenAI was about to open-source its model.

Harry Stebbings

You remember that?

Jonathan Ross

Yeah. My prediction was based on their branding strength. Frankly, OpenAI could probably use Llama 2, the old model from 2 years ago, and people would probably still use it. So there's a brand advantage there.

Now, they do have very good models, but they don't necessarily need it because of that brand advantage. I think Anthropic should open-source its previous generation in order to get people using those models instead of the Chinese models, because if someone is willing to use a Chinese model, then they would at least be using the Anthropic model and their prompts would be recyclable. Just like you have software compatibility, you have prompt compatibility.

For example, when the OpenAI GPT-OSS model was released, one of the main reasons people started adopting it over the Chinese models was that they could reuse their prompts. When someone has a low-cost application and they can't afford the premium for OpenAI, they want to use one of these open-source models.

6. How Europe Could Compete in AI and Why the US is More Risk Averse Than Europe

Eventually, they start doing really well, they make more money, and they start wanting to get access to the premium model. Their prompts are reusable. So there's a win by open-sourcing these models, and you're also getting all of these infrastructure providers to drive the cost down on that model as well. There's a lot of innovation that goes into that.

Harry Stebbings

Totally get that. Can I ask you—there are so many different areas that I want to take this—but we said that we should just build as much compute as possible. The energy requirements are intense. Is the only way to provide the energy required for this compute wave, tsunami, whatever you want to call it, nuclear?

Jonathan Ross

No. Nuclear is efficient and cost-effective, but renewables are efficient and cost-effective. I'll give you my simple hack.

All the allies of the United States have to do in order to have more energy than China is be willing to locate their compute where energy is cheap. So right now, let's compare Europe to the United States. The United States is incredibly risk-averse compared to Europe.

Harry Stebbings

Wow.

Jonathan Ross

Yeah.

Harry Stebbings

In energy?

Jonathan Ross

No. In everything. But you have to ask what kind of risk. There are 2 kinds of risk: mistakes of commission, where you do something that's a mistake, and mistakes of omission, where you don't do something and it's a mistake.

The United States is terrified of making mistakes of omission. When you are in a massive growth economy, missing out is more expensive than fumbling something. Europe is incredibly willing to embrace the risk of omission.

The way that Europe is trying to compete is through legislation, by saying things like, "I want to keep this data in Europe," or, "I want to keep this data in this country." If Europe wanted to compete in AI, all you'd need to do is say, "Norway, please deploy an enormous number of wind turbines."

Why? Norway has about an 80% utilization rate of wind, so about 80% of the time you can be generating energy. They have enough hydro that, if you deployed 5 times the wind power of the hydro, Norway itself could provide as much energy as the United States and could do it consistently. The entire United States—that's 1 country in Europe.

7. Why We Have to Have Nuclear Energy and How to Bring it Back

How much other energy is out there that could be unlocked that isn't nuclear? And by the way, let's also deploy nuclear. Nuclear is incredibly safe these days. Why do we not, then?

Harry Stebbings

Fear. Is that really it?

Jonathan Ross

Yeah.

Harry Stebbings

When you speak to European governments, what do they say to you?

Jonathan Ross

I don't bring up nuclear because I'm not going to push an energy source that everyone's going to push back on. But when I was in Japan recently, they were talking about bringing their nuclear reactors back online.

Japan has a reputation of being very slow. There's a lack of subtlety and nuance in that perception. The reality is that Japan is slow to make a decision, but when they decide something, they move really fast.

Let's take an example. Japan decided to build a 2-nanometer fab. When I was there last, they were showing off these 2-nanometer wafers that they had produced. Now, the yield isn't where it needs to be. This is not production-grade, but they built a 2-nanometer fab, and they are producing wafers out of it. They're going to start getting that defect density down. They're going to move quickly.

They've allocated $65 billion for AI, and they're going to spend it. They're going to spend it quickly. They're going to turn their nuclear reactors back on. When Japan turns its nuclear reactors back on, Europe needs to listen to that and go, "Gosh, we need to catch up in energy."

Harry Stebbings

Catch up is exactly what I was thinking, because what I'm thinking is the speed it takes to build out. You said about Norway's latent capacity of wind and how we could utilize it. Dude, it takes years to build a huge supply of turbines.

Jonathan Ross

Does it?

Harry Stebbings

Yeah. Why do you think the Norwegian government is going to let you shell out and have 10,000 wind turbines on the ground?

Jonathan Ross

Why does the Norwegian government need to pay for it?

Harry Stebbings

Who should? How about the hyperscalers? How about other governments that want to locate there?

Jonathan Ross

In Saudi Arabia, there are gigawatts of power, and they're building out data centers for that. Why doesn't Europe work with Saudi Arabia to say, "You know what?" Saudi Arabia wants to do a program of data embassies where you have sovereign oversight over your data, but you get to use their energy. Why not use that? Problem solved.

They're going to build out 3 to 4 gigawatts in the very near future. So the hyperscalers would pay Norway to use its renewable energy sources and then leverage that. The complaint that the hyperscalers have is all of the paperwork and the slowness.

I was talking to someone who was on the board of a major energy company that builds nuclear power plants. He said they spend 3 times as much on permitting in the United States as on the nuclear power plant. I don't know about Europe, but typically the United States is better than Europe on this.

Harry Stebbings

How much does it cost to build a nuclear power plant in Europe versus the actual cost of the infrastructure versus the permitting?

Jonathan Ross

Here's what everyone needs to walk away from this with: the countries that control compute will control AI, and you cannot have compute without energy.

Harry Stebbings

How far behind is Europe? Is there a way for us to get back? Is it too late? I don't want to be negative. I'm not overly pessimistic, but is there a chasm that we can catch up on?

Jonathan Ross

I don't think there's a problem right now if Europe acts now. China is ahead in action, but there are 500 million people in Europe. There are over 300 million in the U.S.

If you start bringing all the allies together—South Korea, which, by the way, knows how to build nuclear power plants—the power plant in the UAE was built by South Korea, and they could build power plants here. France knows how to build power plants. How about a little bit of a Manhattan Project for building enough energy?

When I'm walking around in Europe in the summer, it's incredibly hot, and when I'm walking around in the winter, it's incredibly cold. That is not an experience you have anywhere else in the world. Build more energy.

Harry Stebbings

I'm with you, Jonathan, but I'm also realistic. I know how slow we are as governments, both singularly and in collaborating together. It's not going to happen at the speed at which this needs to be done. What happens if that does not happen at the speed with which it needs to be done?

Jonathan Ross

Then Europe's economy is going to be a tourist economy. People are going to come here to see the quaint old buildings, and that's going to be it. You cannot compete in a new economy if you don't have the resources that the new economy is built on. The new economy is going to be AI, and it's going to be built on compute.

Harry Stebbings

Is model sovereignty enough to win?

Because if you don't have compute, you can't run the AI. It doesn't matter how good your model is. You could have a model that is 10 times smarter than OpenAI's model, and if you have 10 times the compute, OpenAI's model is going to be better.

So, for Mistral, who says, “Hey, we’re going to have sovereignty within Europe, and the German healthcare system and the Croatian transport ministry are going to use Mistral because we’re a European alternative,” that’s not a reason to win. What’s the USP? What’s the unique selling point?

Jonathan Ross

It’s a European model, and it doesn’t have ownership in the U.S. under a Trump administration. What does it have to do with giving you enough compute? What you’re solving for there is removing someone else’s ability to control you.

Harry Stebbings

Yeah.

Jonathan Ross

But what you’re not solving for is having enough of it. By the way, I’m not saying don’t use Mistral. We have a partnership with Mistral; we love Mistral. The thing I’m saying is: build enough compute so that Mistral can compete.

If you listen to this, are you not just like, “Shit, I should just buy the fuck out of CoreWeave”? Seriously, when you look at what they provide on demand, CoreWeave is a great company. But they have a finite allocation of GPUs. Everyone has a finite allocation.

Harry Stebbings

When we chatted before, you said to me that GPUs are not the best infrastructure for inference.

Jonathan Ross

Correct. We are moving more and more into a world of inference as we move further along the maturation cycle of training models.

Harry Stebbings

Yes. Does that not mean that NVIDIA’s power hold weakens further?

Jonathan Ross

No. NVIDIA is going to sell every single GPU that they build. Even if we end up supplying 10 times as many LPUs as GPUs, all that’s going to do is increase the demand for GPUs and allow them to charge an even higher margin.

Harry Stebbings

Why is that? Sorry.

Jonathan Ross

Because the more inference you have, as mentioned before, the more you need to train the model to optimize for the inference. The more training you have, the more inference you want to deploy to optimize for the cost of that training—to amortize the cost of the training. There’s a virtuous cycle between the two.

Harry Stebbings

Is the inference market playing out as you expected it to in terms of maturation and deployment speed?

Jonathan Ross

What I never expected was that AI was going to be based on language. What that’s done is made it trivial to interact with AI. I thought it was going to be more like AlphaGo. I thought it was going to be intelligent in some weird, esoteric way. The fact that it’s language means anyone can use it.

So, I expected AI to come sooner and grow slower. It came later, and it’s growing faster than I ever imagined. It is so easy to interact with AI that anyone can do it. 10% of the world’s population is a ChatGPT weekly active user.

Harry Stebbings

Isn’t that astonishing?

Jonathan Ross

Yes. But what’s holding it back?

Harry Stebbings

Compute.

Jonathan Ross

So, compute is holding it back in terms of its quality. More people would use it; they just wouldn’t get as much out of it. More people would use it if more languages were supported.

Harry Stebbings

Well, this is the number 1 complaint we hear around the world. You know what would solve that?

Jonathan Ross

More compute, more data. If you have more data, then you can train more, but you need more compute. By the way, if you have more compute, you can generate more synthetic data, so you can train more.

Every one of these—so you have data, algorithms, and compute—if you improve any one of them, it’s not a bottleneck. It’s not like if the compute doesn’t get better, I can’t use more data, or if the data doesn’t get better, I can’t use more compute. Any one of these that gets better improves AI, and that makes it really easy to improve AI because you can improve 1 dimension of it.

It just turns out the easiest knob to improve in AI is not the algorithms. Algorithms rarely improve. It’s not the data, because it’s really hard to get more data, and we haven’t fully figured out synthetic data generation. We’re good at it, but we’re not at the point yet where we can just directly turn compute into more data. We’re getting there.

Compute is the easiest knob because it just keeps getting better and better and better every year. If I write a check for enough money and I’m willing to wait a little while, I’m going to get more compute. It’s the most predictable part of the pipeline.

Harry Stebbings

Given it’s the most predictable part of the pipeline—

Jonathan Ross

And yet we still underestimate how much we need.

Harry Stebbings

Do you think we are dramatically underestimating how much we need today?

Jonathan Ross

Yes. Yes.

Harry Stebbings

By what scale?

Jonathan Ross

Going back to what I said about how every time you add more compute, a product gets better, there is no limit to the amount of compute that we can use. It’s different from the Industrial Revolution. In the Industrial Revolution, you couldn’t use energy unless you had the machinery to use it, and you had to build machinery, and that took time. If I wanted to have more cars on the road, I had to build the cars. It wasn’t enough to just pull more oil out of the ground.

AI is not like that. Yes, if I make my model better, I can actually do more with the same amount of compute. But if I double my compute, I double the number of users. I improve the quality of the model. This is different. I can literally just add more compute to the economy, and the economy gets stronger.

We’ve never had that before, where it wasn’t a bottleneck. It was more of a rubberneck where you could just force more of 1 component through, and then everything improves.

8. Deflationary Pressures and New Job Markets

Harry Stebbings

You said the economy gets stronger. When we think about what that’s predicated on, that’s predicated on the $10 trillion labor spend in GDP shifting to AI and us taking a portion of that. Do you think that we will see significant shifts in the GDP, or the spend on labor moving toward AI, in the next 5 years?

Jonathan Ross

I believe that AI is going to cause massive labor shortages. I don’t think we’re going to have enough people to fill all the jobs that are going to be created.

There are 3 things that are going to happen because of AI. The first is massive deflationary pressure. This cup of coffee is going to cost less. Your housing is going to cost less. Everything is going to cost less, which means people are going to need less money.

Harry Stebbings

So, how is it going to cost less to have a cup of coffee because of AI?

Jonathan Ross

Because you’re going to have robots that are going to be farming the coffee more efficiently. You’re going to have better supply chain management. You’re going to be able to genetically engineer the coffee so that you get more of it per watt of sunlight, right? It’s just going to be across the entire supply chain, across the entire spectrum.

You’re going to have massive deflationary pressure. That’s number 1. What that means is people will need to work less. That’s going to lead you to number 2, which is people are going to opt out of the economy more. They’re going to work fewer hours, fewer days a week, and fewer years. They’re going to retire earlier because they’re going to be able to support their lifestyle working less.

Then number 3 is we’re going to create new jobs, new companies, and new industries that don’t exist today and are going to need labor. Think about 100 years ago. 98% of the workforce in the United States was in agriculture. 2% did other things. When we were able to reduce that to 2% of the population working in agriculture, we found things for those other 98% of the population to do.

The jobs that are going to exist 100 years from now, we can’t even contemplate. 100 years ago, the idea of a software developer made no sense. 100 years from now, it’s going to make no sense, but in a different way, because everyone’s going to be vibe coding, right? Influencers wouldn’t have made sense 100 years ago, but now that’s a real job. People make millions of dollars off of it.

So, what jobs are going to exist 100 years from now? Number 1, deflationary pressure. Number 2, opting out of the workforce because of that deflationary pressure. Number 3, jobs and companies that couldn’t exist today that are going to exist and are going to need labor. We’re not going to have enough people.

Harry Stebbings

It’s fascinating, the counternarrative, isn’t it? Everyone’s being like, “Ah, millions and millions of people will be unemployed,” and you’re like, “No, we’re actually not going to have enough people for the jobs.”

Jonathan Ross

Well, what was the famous prognostication 100 years ago that there was going to be massive famine because we weren’t going to be able to feed ourselves? People always underestimate what’s going to change in the economy when you improve technology.

Harry Stebbings

When you think about the requirements from an energy perspective, and then also what you just said there about labor, do you think Trump and the Trump administration are doing more to help or to hurt the advancement of AI in the U.S.?

Jonathan Ross

Definitely help. All of the moves that have been made are things that are going to help with AI. For example, the permitting issues. Overall, it’s been a very positive experience on AI.

9. The Future of Vibe Coding

Harry Stebbings

You mentioned vibe coding. I do just have to ask about it. Do you think this is an enduring and sustainable market? When you look at a lot of the use cases today, they’re quite transient. How do you analyze the future of the vibe coding market, having played with it a little bit and having seen interns, as you said, who are very good at it internally use it?

Jonathan Ross

Well, let’s take reading. Reading and writing used to be a career. If you were a scribe, you were 1 of the small percentage of people who knew how to read and write, and people would hire you just to record things. You did much better than the average person in the economy because of that, because it was a specialized skill.

Coding has been the same thing. A very small percentage of the population did it. It took a couple of years to learn how to do it well, and some people were really good at it. Now everyone reads and everyone writes. It’s not a special skill; it’s expected in every job.

Coding is going to become the same thing. For you to be in marketing, you’re going to have to be able to code. For you to be in customer service, you’re going to have to be able to code.

I was having dinner with someone who runs a chain of 25 coffee shops and has never coded in their life, and they vibe-coded a supply-chain tool that allowed them to check inventory. They didn’t write a single line of code; they got it to work. It was funny because they discovered all the problems that we software engineers discover over time. They started getting feedback from their employees: “This feature doesn’t work. This thing doesn’t work when I do this.” All the little edge cases came up, and then he just started fixing them, all through vibe coding.

10. Why AI Companies Should Strive to Have Low Margins

Harry Stebbings

Do margins matter in a world of exponential growth? When we look at the demand for your products, and when we look at the demand for Lovable or Replit, both, bluntly, have bad margins. Does it matter having bad margins when growth demands are so high?

Jonathan Ross

I would say that, first of all, you do have to have profitability in the end, or at least break-even, to be an ongoing concern. At some point, you can’t just keep raising money. Even Amazon had to start making some money.

The real reason why you need higher margins is volatility. If you have a razor-thin margin and the market moves, you may not be able to raise more money, and you may not be able to get a loan. What a margin does is give you stability and staying power in the market.

On the other hand, what it does is also give competition the ability to enter. Your margin is my opportunity. What you’re trading is stability for a competitive moat. That’s the decision that you have to make.

Harry Stebbings

How do you think about margin internally today?

Jonathan Ross

I think you want the ability to have margin, and you want to give it to your customers and give them an advantage. If you have the ability to take that margin when it’s needed, then you’re in a great position.

We hired this amazing CFO recently, but I remember talking to a previous candidate. When we were talking about margin, they said that we should price so that our supply met our demand. In other words, they wanted to increase the price in order for the demand to come down.

Harry Stebbings

Makes sense. Does it make economic sense?

Jonathan Ross

Yeah. Economic sense. Logically and rationally, yes.

But then, logically, why not use up your brand equity? Why not use the trust that your customers have to sell them things that aren’t good? Brand value and brand equity have value. You want to keep your brand equity as high as possible because trust pays interest.

Similarly, you want to keep your margins low enough that you’re building up this sort of equity value with your customers, where they know that you’re giving them a good deal. When you charge a high margin, you are at odds with your customer, and you want to do everything that you possibly can to align with your customer.

I want my margin to be as low as I possibly can make it while keeping my business stable. I’m going to make my cash flow by increasing the volume.

One of the things that I love about the compute business is that the need for compute is insatiable. It’s Jevan’s paradox. If we produce 10× the compute, we will have 10× the sales. That’s just the way it works. As long as we keep bringing the cost down, people are going to buy more.

I want to keep bringing that cost down. I want to keep increasing the volume. I want to keep selling more for less so that people get more value out of their business, they buy more, and that cycle continues.

Harry Stebbings

How far are we on the journey to bring the cost down? I look back at some of the shows, dude, and I would cringe at myself because I’m talking about Canva implementing AI and it hurting their margins because they’re implementing AI and it’s going to cost them more. It’s such a naive approach to ask that question, because now the cost of implementation has gone down by 98%. How far are we in terms of that cost-reduction cycle?

Jonathan Ross

Let’s step back and use your Canva example. Successful businesses don’t watch the bottom line. They watch their customers. They solve problems that their customers have.

If you are competing, you’re doing it wrong. You want to differentiate. You want to solve a problem that your customer has not solved yet and can’t solve any other way, and then they’re happy to pay you money. That’s how it works. You solve their problem, and then your cash flow is solved.

If you just look at the balance sheet, someone spending on AI doesn’t make sense. But when the customer is very happy and they’re solving a problem that they couldn’t solve otherwise, first of all, you’re increasing the TAM, usually with AI, because it makes the product so much easier to use.

Did you use Photoshop 2 years ago? Impossible. Now, if you want to generate an image, you just explain what you want. That increases the TAM. You may be able to charge less per photo, but your total revenue increases and your total market increases.

11. S&P 7000, Mag 7 & Market Choppiness

Harry Stebbings

Forgive me for this financial question, but we see the S&P about to hit 7,000. We see this ripping of the MAG 7 like we haven’t seen a concentration of value in many, many years, and people suddenly start to feel like, “Wow, it’s getting toppy.” I listen to you and I hear all of this, and I think it’s just the start. How should I think about the duality of those 2 thoughts?

Jonathan Ross

There are 2 components to the value. One is the weighing machine, and one is the popularity contest. There are some products that are a pure popularity contest, like crypto. I have never bought a Bitcoin. I missed out. Why? Because I can’t play in the popularity contest. I’m not good at it. I don’t know what’s going to be popular and what isn’t.

All I can do is see value. When I look at AI, I see real value being delivered. The best example is that private equity firms are all over us. They want access to cheap AI compute because every time they get more cheap AI compute, they can change the bottom line of their businesses. It has real value.

When PE firms go after something and see value in it, it’s not a popularity contest. It’s pure value. The reason companies get a large multiple is that people see that the actual value is going to accrue, or they get hype-cycled on it. There are different participants in the market. Some of them are just playing the popularity contest. Others are looking at the value, and they may come to the same conclusion for different reasons.

They’re coming at it from the value point of view, the weighing-machine point of view. The most valuable thing in the economy is labor. Now we’re going to be able to add more labor to the economy by producing more compute and better AI. That has never happened in the history of the economy before. What is that going to do?

Harry Stebbings

Do you worry that if we have a speed bump in the short term, it will derail significant parts of the economy, given the concentration of value? Everyone rips today. But if NVIDIA, Meta, Google, and Microsoft suddenly hit speed bumps and the AI speed train is just slowed down, the consequent multiplier effect is mega. Do you worry about that?

Jonathan Ross

Yeah. This is independent of the value of AI. This is the sort of control-system theory of what’s going on.

A stock market could inherently be on an upward trajectory. It can overheat, and that overheating causes it to run away. People bid things up, they realize they’ve made a mistake, and then it has to come back down. Then it dips below where it should be. Spending retreats, and people don’t have the funds they need to build their businesses.

A lot of good businesses can die during one of these downward trends. But this is also where the best businesses are made. How many times do you see a downturn and a ton of amazing businesses come out of it?

Harry Stebbings

Do you think we will have a downturn in the next year?

Jonathan Ross

I can’t predict whether or not there’ll be a downturn. The ability to predict something is largely dependent on whether or not predictions affect predictions. If a prediction affects the prediction, you cannot predict it because whatever your prediction is changes the outcome.

The only things that are predictable are things where the predictions don’t change the outcome. If an asteroid is headed toward the Earth and we see that, if we don’t have the technology to stop it, then it’s going to happen. But if we see that happen and we can predict it, then we might develop the technology to stop it. Do you see the problem?

Harry Stebbings

I do.

Jonathan Ross

In the economy, you don’t have to do anything other than move dollars around. You have these very fast twitches in the economy based on people’s ability to predict, which makes it unpredictable.

I can’t tell you what’s going to happen in the economy. All I can tell you is that right now, the biggest problem I see in AI is that if you see a good engineer, one that you would have hired before, they can go out and raise $10 million, $20 million, $100 million, or $1 billion. Rather than contributing to one of the other AI startups, they go create their own, which means that you have difficulty getting critical mass of talent in any one of these AI startups.

On the other hand, AI is making everyone at one of these startups more productive. In terms of whether or not the economy is overheated, I think one of the best predictors of that is: Is the economy getting in the way of the success of the companies?

If it's not getting in the way, then I don't think it's overheated.

Harry Stebbings

Do you not think it is getting in the way? Because fundamentally, the capital supply side is so large that we are actually preventing you from being able to get great engineering teams together because we're funding talent to the extreme, where they can raise huge amounts of money rather than join Groq.

Jonathan Ross

Yes. Please stop doing that. No, but AI is making people more productive. So, it might be possible for the economy to keep ripping and for all of the companies to continue being very successful. We don't know. We've never been through this before.

Harry Stebbings

Is the war for talent insane today?

Jonathan Ross

It's definitely much more aggressive than it's ever been in history, but only in tech. When you look at sports, sports have always been insane, or at least recently been insane. You look back 20 or 30 years ago in sports, the salaries looked a lot like tech salaries.

Harry Stebbings

Sure.

Jonathan Ross

People are just realizing the value. The problem is, in sports, you have a limited number of teams. You might even institute a salary cap and things like this. In technology, we're not doing that. You have an unlimited number of teams, an unlimited number of startups, right?

Just imagine if anyone could go create their own football team. What would that do to salaries? And what would that do to the value of the franchise?

Harry Stebbings

Which incumbent are you most impressed by, and which are you most worried or concerned for?

Jonathan Ross

I would say Google has probably done the biggest turnaround, and they had a structural advantage in that. Google historically has depended more on their engineers to come up with good ideas, and as long as management gets out of the way, great things happen at Google. I just think, from a cultural perspective, that's a systemic advantage.

Harry Stebbings

You think Gemini has been a success for them ultimately?

Jonathan Ross

I do. You just look at the numbers of the adoption; it's been great.

Harry Stebbings

How do you feel about the implementation into consumer products?

Jonathan Ross

Less so. You see random Gemini introductions into each product. It's in Gmail, but it's practically unusable. It's in pretty much every product, and it seems thrown in, kind of half thought through. But you shouldn't judge that yet, because at least they're getting exposure to how people are using it, and they can use that to figure out what they should actually do.

12. Why OpenAI and Anthropic are so Undervalued

I mean, what happened with Google Chrome, right? It was originally Google TV. It was a total flop, and then they iterated and turned it into Google Chrome. This is the classic problem where someone puts something out there, everyone throws darts at it, and you don't realize that they're just willing to take those darts in order to build a better product.

Harry Stebbings

And it's fine to take those darts as long as the window of distribution advantage remains. But what's challenging is OpenAI has closed that chasm so significantly.

Jonathan Ross

That's true. Google may be too late.

Harry Stebbings

Do you see what I mean? It's like a classic question: can the incumbent attain innovation before the startup acquires distribution? The startup's acquired distribution to 10% of the world. It's pretty impressive.

Jonathan Ross

Yeah. At this point, it would be hard to imagine a scenario where OpenAI goes away. I just don't see how that happens. So, at the very least, you have 2 competitors from this point on going at it.

Harry Stebbings

Which is OpenAI and Anthropic, or OpenAI and Google?

Jonathan Ross

OpenAI and Google. Anthropic does something different. Anthropic is doing coding, right? OpenAI is doing a chatbot. Google's doing a chatbot. Google's also doing coding. Google's doing everything.

Harry Stebbings

Well, OpenAI is doing coding, too.

Jonathan Ross

That's—well, yes. And actually, our engineers recently started using Codex more than using the Anthropic tools.

Harry Stebbings

Wow.

Jonathan Ross

Yeah. And it's funny because it's almost on a monthly basis. We have a philosophy: we don't tell our engineers what tools to use. We do tell them they must use AI, because otherwise they're just not going to be competitive.

But we saw them using Sourcegraph. We saw them then using Anthropic. We saw them then using Codex. Next month, it'll probably be Sourcegraph again. It just keeps going around and around in a circle.

Harry Stebbings

Do any of these have enduring value, then, if the switching cost is so low and if they're just being used so promiscuously?

Jonathan Ross

Our engineers are cutting-edge engineers who will switch to the best tool the moment it's the best tool. Not everyone is like that. A lot are like that, though.

Harry Stebbings

A lot of the people you interact with are like that. Enterprises make these long-term deals, and they stick with whatever deal they made a year ago.

Would you rather invest in OpenAI at $500 billion or Anthropic at $180 billion?

Jonathan Ross

I'd want to invest in both.

Harry Stebbings

Would you?

Jonathan Ross

Yeah. They're both undervalued. Highly undervalued. You're still looking at them as if they're competing in a finite market for a finite outcome, when they're actually increasing the value of the market with the more R&D that they do.

Harry Stebbings

Play this out for me, then. If we do the bull case for them, what does that look like? I think the current tech companies can increase their value significantly, but I don't know why they couldn't increase their value significantly while the AI labs catch up to where the current technology leaders are.

The Mag 7 is going to increase in value, and what's going to happen is the AI labs are going to achieve the same amount of value as the current Mag 7, but the Mag 7 is going to be more valuable. The question is, will the AI labs overtake the Mag 7?

What will determine that?

Jonathan Ross

I don't know. Frankly, I think they're just going to become the Mag 9, the Mag 11, the Mag 20.

Harry Stebbings

Do you think the AI labs move very significantly into the application layer and subsume the majority of it?

Jonathan Ross

That is the natural tendency of a very successful tech company. They start to do what their customers do and they move up the stack, and then they subsume what their customers did. Then there are new people who build on top of them, right?

OpenAI—I think on your show Sam Altman said something about how, if you're just doing something like a small refinement on top of OpenAI, you're going to get overrun or whatever. He was just being very honest; that's what they do.

In our case, we found an area where we will not compete with our customers, which is we will not create our own models. So, we just won't do it. By putting that line in the sand, we're saying it's safe to build on our infrastructure, right? Because we're not going to go after what you do.

That may be the wrong call. We may find that we're subsumed by one of our customers. But it also means that you can trust that you can build on us. I could be making a huge mistake on that call.

Harry Stebbings

You could be. You would also need a lot of cash to do that—to build your own models. Speaking of cash, how much did you just raise?

Jonathan Ross

We raised $750 million.

Harry Stebbings

$750 million at—what was it?—$6 billion.

Jonathan Ross

Yeah, almost $7 billion.

Harry Stebbings

Okay, got you. This sounds really unfair. And that's amazing. Is that enough money?

Jonathan Ross

It is. In fact, we were only going to raise $300 million. You brought up the question of profitability and all that. The hardware companies are in a good position because, unlike these other companies, we actually make money off of what we sell.

When we sell hardware, those hardware units actually have positive margin.

Harry Stebbings

I thought you had negative margin.

Jonathan Ross

When we sell hardware now—

Harry Stebbings

Versus when you sell software.

Jonathan Ross

When we sell software, it depends on the model. Our most popular models on the chip that we're ramping up now are positive margin. But we do have some models that we run that beat the opex, but we're not happy with the capex. Others would be happy with the capex, but we're more conservative.

It's just easier to say, when we sell hardware, we have positive margin because you know it at that moment. We might have positive margin on even our least profitable models because we just don't know how long the hardware is going to last.

Harry Stebbings

What are the margins, and where do they go over time?

Jonathan Ross

Well, one of the benefits of being private is I don't have to tell you.

Harry Stebbings

You don't. But it'd be lovely if you did.

Jonathan Ross

It's the only advantage of being private.

Harry Stebbings

No, no, no. There are many, many advantages. You don't have a lockup period. You can sell much more easily.

Jonathan Ross

Yeah, but I don't sell shares, so—

Harry Stebbings

You've never sold a share, have you? Never.

Jonathan Ross

No.

Harry Stebbings

Yeah, you clearly don't understand how this game works. Don't worry, I will teach you.

But margins over time, do they get significantly—how do you think about that? I'm not asking necessarily—

Jonathan Ross

No, no, no. I'm going to say what I said earlier, which is I want our margins to be as low as our business remains nonvolatile. Like I said, the only reason for a high margin is because you want to have the ability to bring in cash when you need it.

All you need is the ability to price higher if you need to in order to be able to lower your margin. The demand for compute is so high that if someone came to us and said, “I need this compute,” and we have it, they will pay a higher margin, which allows us to charge a lower margin.

13. The Chip Market in 5 Years

Harry Stebbings

Can you help me understand what the chip market looks like in a 5-year timeline? You said we'll have OpenAI, we'll have Anthropic, we'll have all the providers having their own chip infrastructure.

You'll also have NVIDIA. What does that look like?

Jonathan Ross

My prediction is that in 5 years, NVIDIA will still have over 50% of the revenue. However, they will have a minority of the chips sold. They might have a minority share. They might have 51% of the revenue, and they might have 10% of the chips sold.

Harry Stebbings

Can you help me understand that?

Jonathan Ross

Yeah, there is huge value in being a brand. You get to charge more. However, it makes you less hungry, and you're going to start charging high margins, and some people are going to pay it because no one’s going to get fired for buying from NVIDIA. It's a great place to be in. That business is going to remain incredibly valuable. If you're invested in NVIDIA, you're probably going to do okay.

However, if you're looking at it from the customer point of view, when you have customer concentration like we're seeing, where 35 or 36 customers are 90% to 99% of the total spend in the market, they're going to make decisions less on brand and more on what makes their business successful because they're going to have more power to make those decisions. So, you're going to see other chips being used because those companies are going to have enough power to make decisions themselves.

Harry Stebbings

You said you won't do badly if you're an NVIDIA investor. One of my friends says, “The thing I love about Harry is that he's wonderfully charming, but at the end of the day he goes, ‘That's great, that's great, but what about me?’” Which is very true.

Over-under on NVIDIA in a 5-year timeline: $10 trillion?

Jonathan Ross

I personally would be surprised if in 5 years NVIDIA wasn't worth $10 trillion. The question you should ask is, will Groq be worth $10 trillion in 5 years? Possible. We don't have the same supply chain constraints. We can build more compute than anyone else in the world.

The most finite resource right now is compute, the thing that people are bidding up and paying these high margins for. We can produce nearly unlimited quantities of it.

Harry Stebbings

What do you think the market does not understand about Groq that you think they should understand?

Jonathan Ross

Oh, it changes every month. It used to be we couldn't have multiple users. Then we demoed multiple users to people on the same hardware, right? They used to think that we—

Harry Stebbings

This is because of the SRAM structure.

Jonathan Ross

Because of the SRAM.

Harry Stebbings

Actually, here's another one. I'm still impressed with my learning from last time. Thank you so much. I learned so much from you, genuinely. I was genuinely learning so much. But okay, the question I get asked the most is, isn't SRAM more expensive than DRAM?

Jonathan Ross

The answer is yes. A good way to think of it is that SRAM is inherently 3 to 4 times as expensive per bit. SRAM has 3 to 4 times as many transistors or capacitors—just transistors for SRAM—as DRAM. DRAM is a capacitor and a transistor. SRAM is 6 to 8 transistors.

So SRAM is inherently larger per bit, which means it uses more silicon and therefore it's more expensive. You're also deploying it on a more expensive chip, like a 3-nanometer chip, so it costs you more per unit of area than DRAM. There's a multiple. Maybe it's 10 times as expensive per bit.

The thing is, when we're running a model like Kimi and we're running it on 4,000 of our chips, and you're running that Kimi model on 8 GPUs, we're using 500 times as many chips. That means the GPUs have 500 copies of that model, which means they're using 500 times as much memory, which means that their cost is higher because even if the SRAM is 10 times more expensive, they're using 500 times as much memory in the DRAM.

This is one of those classic problems of looking at it from a chip point of view rather than a system point of view. Everything that we did was actually from a system point of view, and now it's world point of view.

We actually load-balance things across our data centers. We're now at 13 data centers. We have data centers in the United States, in Canada, in Europe, and in the Middle East. When you have a world-scale distribution, you don't just make decisions at the data center level.

We actually will have more instances of some models in some data centers, with different compile optimizations for input or output based on what's going on in a geography. We may not even have an instance of a model in a particular data center. We may have it elsewhere, and we can load-balance that. So we're optimizing at the world level, not at the data center level.

Harry Stebbings

What would you do if you weren't scared, Jonathan?

Jonathan Ross

I'll rephrase that to: Where could I increase risk in the business?

Harry Stebbings

Yeah, same question.

Jonathan Ross

Where we haven't, we could double our orders in our supply chain. We have a 6-month supply chain, so we can respond to the market faster than anyone else.

Harry Stebbings

How overweight demand are you in supply?

Jonathan Ross

Like I said, last week someone came to us and asked for 5 times our total capacity.

Harry Stebbings

If you're not supply-constrained, why can't you just do that?

Jonathan Ross

Because there are thresholds. So, for example, if we had double the capacity, we wouldn't have won that customer. They needed 5 times as much. So it's not enough to have twice as much. We have to have enough. If we double the capacity, do we have enough for those customers?

Harry Stebbings

The risk that you could take is to what? Sorry, just specifically.

Jonathan Ross

We could just double the rate at which we're building out supply. With this fundraise, we ended up raising more than twice what we were expecting to raise. And then we were 4 times oversubscribed over what we did raise.

So we could have raised a lot more money. It would have been more dilutive, and I'm trying to be dilution-sensitive for investors and everyone else. But, on the other hand, we could have just raised more money and built a ton of compute.

The other advantage that we have, versus anyone else, is that our cost per token, especially at a given speed, is very advantageous. So we know that we can charge less than the rest of the market, which matters when you're trying to build these businesses—not because people are spend-conscious.

If we lower what we charge by 50%, people are going to buy twice as much. They're spending as much as they're making because whatever they spend increases the quality of the output.

Harry Stebbings

Do you think about going public at all?

Jonathan Ross

Our focus is purely on execution right now. Whether or not you go public, that's a completely different game than we're playing right now. Right now, all that matters is: Can we satisfy the demand for compute?

Harry Stebbings

Why do you think Cerebras decided to go public?

Jonathan Ross

Well, they recently decided not to go public.

14. Quick-Fire Round: Biggest Fear, Nvidia: $10TRN, Zuck Buying AI: Work or Not

Harry Stebbings

That answers that question. Dude, I could talk to you all day. I do want to discuss a quick-fire round. I say a short statement, and you give me your immediate thoughts. Does that sound okay?

Jonathan Ross

Yeah.

Harry Stebbings

What's the biggest misconception about NVIDIA today?

Jonathan Ross

That NVIDIA's software is a moat.

Harry Stebbings

CUDA lock-in is bullshit?

Jonathan Ross

Yeah. It's true for training, but it's not true for inference. I mean, we have 2.2 million developers on us now. That's how many have signed up.

Harry Stebbings

Wow.

Jonathan Ross

Yeah.

Harry Stebbings

How many does CUDA have?

Jonathan Ross

They claim 6 million.

Harry Stebbings

If you were founding Groq today, with NVIDIA at $4 trillion and the AI boom in full swing, what would you do differently?

Jonathan Ross

I wouldn't do chips. That ship has already sailed. It takes too long to build a chip.

Harry Stebbings

Does it? So, for the chip providers today that are coming out, we're seeing new chip providers come out where they're raising a lot of money from good people.

Jonathan Ross

It's too late. Yeah, so the reason that I decided to go into chips—so, I did the Google TPU, but also, before I left, I set a record on the best classification model, like ResNet-50, with someone in Google Brain. We did an experiment. We beat everything.

I could have gone on the algorithm side, especially in formal reasoning, which is good that I didn't. The main motivation to go into chips was the moat—the temporal moat.

A question we get asked by VCs a lot is, what prevents someone from copying what we're doing? The answer to that is, if you copy what we do, you're 3 years behind us because it takes that long to go from the design of a chip to a chip in production, if you execute perfectly.

I've done 3 chips now that are in production or ramping to production. All 3 were A0 silicon. Only 14% of chips that are taped out for the first time work the first time; they're A0 silicon. So that means there's an 86% chance each time that you're going to have to respin it.

When we built our V2 chip, we had already scheduled a respin for it, and we ended up not having to do it because, to our shock, the first one worked. You shouldn't expect that. That 3 years is if everything goes perfectly.

NVIDIA typically takes 3 to 4 years per chip, and they just have multiple being done at a time.

Groq is now in a 1-year cycle. So, a year after our V2 is our V3, and a year after that is our V4.

Harry Stebbings

How do you evaluate the meteoric rise and reacceleration of Larry Ellison and Oracle?

Jonathan Ross

Brilliant business decisions and the willingness to move fast. Most people right now keep asking themselves, “Is AI overheated? Should we double down on this?” They just went for it. They’re aggressive, and that’s what it takes to win.

When everyone else is fearful, you should be greedy. And when everyone else is greedy, you should be fearful. Right now, there’s a lot of fear around AI. What you’re seeing, though, is a couple of greedy, really smart people making tons of money, and it looks like there’s a lot of greed out there. It’s just a handful of people that are moving fast.

Harry Stebbings

Where should I be greedy and where should I be fearful? I’m an investor today, obviously.

Jonathan Ross

Wherever there’s a moat. Hamilton Helmer, 7 Powers, right? Wherever you see a moat, you should be greedy.

Harry Stebbings

Very few people have a moat.

Jonathan Ross

Yeah, and especially at the stage that you invest in.

Harry Stebbings

Yeah. So, you have to predict that there’s going to be a moat.

Jonathan Ross

And if there is a moat, it’s a billion-dollar valuation for a pre-moat.

Harry Stebbings

I mean, there’s a billion-dollar valuation for a pre-moat. That’s what you should call it: pre-moat. That’s what the investors should denote it as. Pre-moat.

What have you changed your mind on in the last 12 months?

Jonathan Ross

Oh my gosh. It’s not so much that I’ve changed my mind; it’s that I’ve changed how much—what percentage of our business—doubles down where. Every month, we become more focused. We say yes to fewer things, and what happens is the business just does better.

I would say I used to think that the most important thing was preserving optionality, and now I think it’s focus. However, I think having that optionality early on was crucial so that we could play where we would be most successful, and now it’s about focus.

Harry Stebbings

We’ve spoken a lot about OpenAI and Anthropic. Do you think Elon Musk is able to pull it off with Grok and xAI?

Jonathan Ross

Yes, although it’s probably going to be different. Whenever a new area emerges, a bunch of people think that they’re competing, and they’re not. All of these people creating foundation models think that they’re competing for the exact same thing.

What did Anthropic do that was brilliant? They decided to stop competing by doing everything and focus on coding. And that’s worked great for them, right?

If you look at xAI, they have a social network and they’ve integrated their chatbot with that. I’m not going to use that chatbot for solving deep analysis or deep research problems. I’m not going to use it for coding. Now, they do have a coding model, but they don’t have a coding distribution. Can they use that distribution to get into coding? Maybe, but then they’re not going to be as focused. So, what are they doing?

Eventually, the markets will diverge. The Magnificent 7—all of those companies have some overlapping business, but the primary business of each of those Magnificent 7 companies is different. If you do not differentiate, you die.

Harry Stebbings

When you look at Google, Microsoft, and Amazon, you can buy one and you can sell one. Which do you buy? Which do you sell?

Jonathan Ross

It depends on the time frame. In the short term, I think Microsoft is resetting a little bit because of the OpenAI relationship. Long term, they’re probably going to do fine again, I think.

Harry Stebbings

Do you think that’s material damage to them?

Jonathan Ross

No, that’s why I’m saying in the short term I think it’s going to hit them, and in the long term it’s not.

Harry Stebbings

Have they not done massively well from that? They have the financial ownership of OpenAI, and then they have the flexibility to use Anthropic for most of the suite.

Jonathan Ross

And they’ve deployed an enormous amount of compute. So, if OpenAI diversifies and gets its compute elsewhere, they have that compute now. Compute is like gold, right? If you have it, you have AI.

And then Amazon, I think, doesn’t have AI DNA. If you compare them—you didn’t mention Meta, right?—Meta and Google always had the AI DNA, and Microsoft bought it with OpenAI, but that bought them time. Amazon still doesn’t have that DNA, but they do have compute.

Harry Stebbings

Final one: What are you most excited for when you look forward? I like to end on an element of positivity. What are you most excited for when you look forward over the next 5 to 7 years?

Jonathan Ross

I think the things that scare most people are what excite me. What I mean by that is, everyone’s afraid of what AI is going to do. I think there’s a good historical analogy here, which is Galileo.

A couple hundred years ago, Galileo popularized the telescope, right? He got in a lot of trouble for that. The reason he got in so much trouble was the telescope allowed us to see some truths and allowed us to realize that the universe was larger than we imagined. It made us feel really, really small. Over time, we’ve come to realize that while we may be small, the universe is grand and it’s beautiful.

I think over time we’re going to realize that LLMs are the telescope of the mind. Right now, they’re making us feel really, really small. But in 100 years, we’re going to realize that intelligence is more vast than we could ever have imagined, and we’re going to think that’s beautiful.

Harry Stebbings

Jonathan, dude, I always end up taking copious notes in our conversations. Thank you so much for doing this with me, man. So lovely to do it in the studio, and you’ve been fantastic.

Jonathan Ross

Thank you.

Groq 创始人 Jonathan Ross:OpenAI 与 Anthropic 将自研芯片,NVIDIA 能否涨到10万亿美元 — 文字稿与摘要 | BidClub