Cerebras CEO Andrew Feldman:为何融资 10亿美元、推迟 IPO,以及 Nvidia 为什么担心增长
- Cerebras 融资 10亿美元,创下该类别历史最大规模和最高估值,由 Fidelity(“投资界的牛津或剑桥”)领投,Tiger Global、Valor 和 1789 参投,Feldman 表示 Cerebras 仍计划上市:“我们仍然完全打算上市。” Feldman 称,这轮融资之所以能够完成,是因为 Cerebras 有毛利率,而同期寻求融资的竞争对手毛利率为负;他认为 S-1 显示,阿联酋可能在 2024年上半年贡献了约 75-80% 的收入,订单之大甚至“吞噬了我们全部的制造产能”。
- Nvidia 正显露出一家担心增长的巨头才会有的迹象:“更多使用资产负债表,少用技术”——通过买下业务而不是赢得业务(正如 Cisco 在 1992-2001年所做的那样),在客户还拿不到 B200 之前就“掠夺式预告” B300,并对“规模巨大的”现场故障率保持沉默。 对 OpenAI 的 1000亿美元投资“就是设计成没人能看懂……这根本不是一个可分析的东西”,除了 Nvidia 借此锁定 OpenAI 部分需求之外,几乎无法解读。
- 芯片按两年折旧在经验事实上是错的:“两年折旧”与现实不符。 H100 过了两年仍在赚钱,A100 已经用了三到四年,“可能长达五到六年”。真正决定折旧的变量是代际提升;在同口径比较下(8-bit 对 8-bit),Feldman 估计每个有意义的世代提升约为 2-2.5x,因为对 GPU 架构上的推理而言,瓶颈是内存带宽,而不是算力。
- 即使买方也无法知道需求会有多大——客户向 Cerebras 要求的每秒查询数在 500万到4000万之间,Harry 指出这相差了一个数量级——因此应把超级大单看成对未来的期权,把“未来五年最高 1000亿美元”理解为“营销史上最伟大的免责措辞”。 他仍可能低估需求的概率是:“100%。我一直都错。”
- Mag7 集中的风险不在估值(Nvidia 4.5万亿美元:“可能还低了”),而在被错误定价的分散化:标普“不是全球经济指数——它是 7家公司占 30% 或 50% 的指数”,持有人因此承担了从未主动选择的行业风险。 “金融市场中的风险,来自人们从根本上低估风险。”
- 能源稀缺叙事“完全错误”:“我们有充足的电力,只是它在错误的地方。” 西得州天然气和纽约州北部水电远离人口、建筑和光纤;核电是合理方案,但并非不可替代。相较于中国的中央规划,美国碎片化的审批制度才是真正短板——当地消防条例曾让 Samsung 得州晶圆厂延误 8-10个月。
- 对扩散速度的逆向判断:Groq 的 Jonathan Ross 预测 5年内出现 AI 劳动力短缺,“完全错误——15年后可能是真的”;AlphaFold 虽然赢得诺贝尔奖,但“说出一款由它带来的药,一款也没有”。 只有当社会围绕 AI 重新组织起来,生产率才会跃升——电力和电动机的历史已经说明这一点;如果 AI 只是替代 Google,生产率不会出现同等规模的变化。
1. 融资 10亿美元:Fidelity 的背书,以及毛利率买来的入场券
- Feldman 称,这是 Cerebras 所属类别有史以来最大、估值最高的一轮融资,由 Fidelity 领投——“投资界的牛津或剑桥……他们选择领投一轮融资,会给华尔街带来很大信心”——共同领投方是“那些条约国家”(很可能指 Atreides),Tiger Global、Valor 和 1789 参投。Harry 还引用 HubSpot CEO 的说法佐证:在 IPO 前和 IPO 时都能拿到 Fidelity,所携带的信号价值,是风险投资人普遍低估的。
- 为什么不直接上市?“我们仍然完全打算上市”——只要“能非常快地完成,而且不会分散精力”,上市前融资就是标准做法。这轮融资之所以能够以更高估值、从更好的投资人那里拿到更多资金,原因在于“我们有毛利率,而其他正在找钱的人毛利率为负”。毛利率是“从一个想法走向一家真正公司”极其重要的一环。
- 这笔弹药将用于扩充制造能力、增加数据中心(今年已在美国新增 5座)以及推进“更多大想法”;他同时讽刺竞争对手:“从 8-bit 降到 4-bit 所实现的假想收益,无法带我们抵达 AI 的应许之地。”
2. 看清小字:超级大单是对未来的期权
- Feldman 拆解 headline 数字:“最高”是“营销史上最伟大的免责措辞——未来五年最高 1000亿美元……可能是 300亿美元、120亿美元,也可能是 400亿美元,不会更大。”而且没人审计这些承诺:“8个月过去了,有没有人拿出一张小表格——9个项目加 1座工厂?谁来追究任何人的责任?没人。”
- 这些公告真正释放的信号,是需求大到买方无法准确规划。客户向 Cerebras 要求的每秒查询数在 500万到4000万之间,Harry 指出这相差了一个数量级,因为“6、8、12个月之后,所有人都不确定。变化太快,规模太大”。正确的理解方式,是把它们看成对未来产能的付费期权——“如果未来朝不利于你的方向发展,你损失的就是期权费”。在这种环境下做规划“极其残酷”:面对 5年、7年的数据中心承诺,需要的是“规则不断变化下的良好规划”,而不只是良好规划。他仍在低估需求的概率是“100%”;至于 OpenAI 估值何时变得可以想象,答案是“前一天”。
- 需求的有力证据来自海湾地区:Feldman 认为,S-1 显示阿联酋可能在 2024年上半年贡献了约 75-80% 的收入,订单大到“吞噬了我们全部的制造产能”——“你可以在硅谷做 20年或 30年的职业销售,却从没见过一笔 5亿美元的订单”。他们“大胆,而且足够早”。
- Harry 追问:面对最大客户所在地区,你难道不是必须说些好话?“问题问得公平。我作为一个犹太人,在我们还没有做成任何生意之前,就去那里谈生意。”随后他给出一个快速的逆向判断:我们有生之年中东会实现和平,依据是阿联酋证明了温和路线能够带来回报——“我们忙着建设,没空仇恨……我们太忙于建设了。”
3. Nvidia 正显露出一家担心增长的巨头才有的迹象
- 被问到 Nvidia 是否坚不可摧(Groq 的 Jonathan Ross 告诉 Harry,Nvidia 5年内能达到 10万亿美元),Feldman 回答:“那我希望他一直做多 Nvidia。”他的判断是:“我们看到了一些大公司开始担心增长时会做的事情——更多使用资产负债表,少用技术。”公司会“开始买业务,而不是赢得业务”,就像 Cisco 在 1992-2001年前后所做的那样。
- 第二个迹象是“掠夺式预告”:“在任何人拿到 B200 之前先宣布 B300,在 B200 技术上还没完成之前就开始谈 Rubin”,与此同时,绝口不提“产品的现场故障率,而且故障率规模巨大”。目的就是让买家等待,而不是选择“更好且已经存在”的技术。
- 对 OpenAI 的 1000亿美元交易“就是设计成没人能看懂……金额上限明确、期限未指定、估值没有给出——这根本不是一个可分析的东西”,唯一明确的是:“Nvidia 选择通过投资 OpenAI,试图锁定其部分需求。”
- 说到毛利率,Nvidia 的毛利率“处于硬件公司历史最高水平之一”——整体为 78%,高端芯片“可能达到 85%”——这正是 AWS 自研 Trainium(听起来像“cranium”)的原因。怨气还会累积:“Intel 摔倒时,突然从各个角落冒出来、落井下石的人多得惊人。”
4. 芯片按两年折旧在经验事实上是错的
- 人们在 H100 上获得价值的时间已经超过两年,A100 也是如此——“更接近 3年或 4年,而且可能长达 5年或 6年”。因此:“如果你说它们按两年折旧,那在经验事实上就是错的。”
- 真正的问题是:“未来世代比当前世代快多少。”只有当替代芯片在每瓦性能上快到足以让同一个 50MW 外壳赚到更多钱时,淘汰一块已经完全折旧的芯片才有意义。如果行业不再交付非凡的代际提升,“它们就会使用更久。折旧期也会拉长。”
- 撇开营销、做一点工程层面的深挖后,真实的代际提升“可能是每个有意义的世代 2-2.5x”,而且要同口径比较(8-bit 对 8-bit)。如果“内存带宽没有提升超过 2x,你就无法用上更多算力”,更多 flops 只会被浪费。对推理而言,内存带宽“是 GPU 架构的根本限制”。
5. 晶圆级芯片下注:听起来显而易见,却 75年无人解决
- 内存的取舍是:SRAM 速度极快,但容量很小;HBM(一种 DRAM)容量大,但速度慢。GPU 选择 HBM,是因为图形处理很少频繁访问内存。Cerebras 的做法是制造一块“餐盘大小的芯片”,把高速 SRAM 塞到极限,用巨大的硅面积突破 SRAM 的容量限制。一块普通尺寸的 SRAM 芯片运行万亿参数模型,需要“四五千块芯片,简直是一团乱。”
- Harry 反问:这不是显而易见吗?“确实如此,不是吗?”但在计算机 75年的历史中,没有人制造出超过约 840mm² 的芯片:Gene Amdahl 失败了,IBM 失败了,TI 失败了;“我们做到之后,Elon 也在 Dojo 上尝试过,但他们失败了。”
- 这场赌注一度濒临死亡:从 2017年到 2019年初,约 15个月里他们始终造不出一块可用芯片,每月烧掉 600-700万美元,每次尝试后都进行正式的失败分析。第一片晶圆终于运行起来时,创始人盯着设备看了半小时:“我们刚刚解决了一个 75年来这个行业最聪明的人都没能解决的问题。”
- 训练和推理谁更快?“不,我们两者都更快。”但训练意味着要移植以 GPU 为中心的配方;推理则不同,“没人关心 CUDA,甚至没人关心 PyTorch……他们要的是一个 API。从基于 GPU 的 OpenAI OSS 120B 方案迁移到我们的方案,真的只需要 10次击键。”再加上推理用户远多于训练用户,这就是推理更适合率先抢占市场的原因。
6. 电动机的教训:不重组,就没有生产率跃升
- Feldman 提到了很可能是 Robert Solow 在 1988年提出的悖论——“除了生产率统计数据,到处都是计算机”——以及 Paul David 的《计算机与电动机》:电力从 1880年起被制造业采用,却几乎没有带来任何变化,直到车间围绕电力重新组织起来,生产率才跃升;90年代中期计算机联网后也是如此。“如果你像使用 Google 一样使用 OpenAI,只会看到非常有限的提升……如果我们围绕 AI 重新组织自己,就会看到巨大的生产率增长。”
- 人口代际差异也说明了这一点,与 Sam Altman 的判断相呼应:年长用户把 ChatGPT 当作 Google 的替代品;年轻用户则把它当作“生活的操作系统”来使用。这种消费模式此前从未存在,生产率跃升将从这里发生。“我知道的是,这种转型需要时间。”
- 推理增长本身由 3个变量相乘:用户数 × 使用频率 × 每次使用的算力。“问题在于这 3个变量都在快速增长,结果会产生一些让人头脑发麻的影响……我们一开始就知道这一点,但它仍然让人倒吸一口气。”
7. 扩散很慢:没有劳动力短缺,AlphaFold 带来的药在哪里?
- Ross 预测,AI 将在 5年内制造大规模劳动力短缺。Feldman 回应:“完全错误。经济失调不可能在非常短的时间内解决。15年后可能是真的。”AI 会“逐步啃进”经济体系。
- 他的证据是 AlphaFold:它解决了化学领域最难的开放问题之一,并让发明者获得诺贝尔奖——“说出一款由它带来的药,一款也没有。”至于理论上会被取代的 X 射线晶体学家?“对他们的需求反而更多了。”
- 真正会改变的是教育:“自 Alexander the Great 接受 Aristotle 训练以来,我们教育孩子的方式几乎没变。”AI 可以把一个学生的错误模式与数千名其他学生进行比对,再开出能够修补特定知识漏洞的练习册。咨询公司和银行的入门级工作也会变化——“擅长做表格、撰写他人研究的摘要:AI 会做得更好。”而 Feldman 一直认为,让 22岁的人做这些本来就是对人才的糟糕使用。
8. “我们有充足的电力,只是它在错误的地方”
- 稀缺叙事“完全错误”:西得州有大量天然气,纽约州北部有大量水电,只是它们不在人员、建筑或电信光纤所在的地方。问题是错配,不是供应。核电在数十年维度上是“非常合理且具成本效益的策略”,但并非不可替代;加拿大的水力资源可能提供“全世界最便宜的电力”,芬兰和冰岛则拥有地热能。
- 中国“长期、深入地思考过电力基础设施”,并进行战略规划;美国“分散式政府结构让我们留下了一套拼接式电力基础设施”。当地消防条例曾迫使 Samsung 重新设计得州晶圆厂,让一个数十亿美元项目延误 8-10个月。
- 政治评分方面:“拜登政府被误导且充满恐惧”;总体而言,特朗普可能更有帮助,因为他身边聚集了聪明的 AI 人才,也放松了一些令人痛苦的监管。对于美中竞争叙事,他不接受简单的竞赛框架——“军备竞赛当然没有帮助美国或俄罗斯中的任何一方”——但现实政治是:“他们更擅长制造无人机,也更擅长制造机器人”,北京还为 AI 基金的亏损提供了兜底。Feldman 在 2019年、出口管制出台前,曾因担心技术用途而放弃一笔大型中国交易。
- 与电力消耗相伴而来的责任是:“如果我们要消耗这么多电力,交付价值的责任就在我们身上。”这些价值可以是药物、医疗和延长寿命。至于可能由 Ghibli 图片带来的算力需求?“一个市场需要大量坏想法,才能筛出少数好想法。”混乱本身就是机制;政府资金和审批便利应当导向真正重要的项目。
9. Mag7:风险在于被错误定价的分散化,而非估值过高
- 集中风险“不在于它们值这么多钱——我认为它们值这么多钱,是因为未来经济会奖励它们”。真正危险的是认知模型错配:“人们仍然把标普看作全球经济指数,但它不是,它是 7家公司占 30% 或 50% 的指数。”持有人“以为自己已经分散投资,实际上高度依赖一个非常狭窄的行业”。
- 贯穿本期节目的风险框架是:“金融市场中的风险,来自人们从根本上低估风险。当风险被正确计价时,你的结果就不会令人意外。”
- 说到 Nvidia 的 4.5万亿美元估值:“21世纪第一个 25年最伟大的公司……我不知道 4万亿美元是否正确,但一个非常大的数字——可能还低了——是正确的。”他不会选择公开市场股票——“你可以在好公司上亏钱,也可以在烂公司上赚钱,这对我来说无法接受”——而押注最大的巨头,在这件事上“没有 alpha”。
- 这轮繁荣能否持续?外推存在边界:“如果 Nvidia 继续按当前增速增长,11年后地球上的每个人都为它工作——自己算一算。”但一个围绕 AI 重组、且 AI 占比更高的经济体“不仅可能,而且几乎确定会出现”。
10. 瓶颈——以及投资人会在哪些地方输掉衬衫
- 首先是专业人才:大学“培养不出足够多”的 AI 从业者,美国移民政策的挑战“没有帮助”;Feldman 主张打通 J-1 到 H-1B 的路径,并称美国必须认真对待移民人才,而大学则“缺算力”。这正是他不担心人才争夺推高薪酬的原因:“没有公司会因为给杰出人才太多钱而破产。想破产,就给平庸的人太多钱。”毕竟,美国曾给 Charlie Sheen 每集 250万美元的片酬。
- 物理瓶颈包括:TSMC 和 Samsung 无法足够快地建设各自 300-500亿美元的晶圆厂,限制了所有人的芯片供应,并使成本居高不下。至于承诺中的吉瓦级数据中心?“所有人都在承诺。它们在哪里?还没建起来。”Elon 可以在 6-8个月内完成建设,“世界其他地方要 1年半,甚至更久”。
- 华尔街喜欢数据中心,因为它“看起来像债券……你能拿到一个投资级租户”;CoreWeave 的金融工程展示了这条路径。但“建数据中心并不适合所有人”:最好的建设成本是每兆瓦 800万美元;“如果你花 1200万或 1400万美元,这就是你亏钱的方式”,而且还会叠加电力接入、审批、成本控制和租户问题。
- 对于 OpenAI 或 Anthropic 自研芯片(Ross 认为肯定会发生),他的判断是:“软件公司造芯片失败,有着悠久的历史。”微软规模的公司也失败过;Google 是最成功的案例,但“他们已经做了 10年”。成功案例通常来自收购,如 Apple 收购 PA Semi、Amazon 收购 Annapurna。OpenAI 多年来托管在 Azure 上,Anthropic 同时依赖 AWS 和 Google,如今都没有实现垂直整合。“造芯片是 MBA 的噩梦。”Intel 在 2000-2010年拥有全球最好的架构师和晶圆厂,却“完全无法造出可用的手机芯片”;ARM 赢下了本世纪最大的算力市场。硅片“不是 25岁 CEO 应该涉足的领域”。真正投入不足的环节是数据清洗和数据管道——“许多 AI 项目失败……是因为数据一团糟”。数据供应商(Surge、Mercor、Invisible、Turing、Handshake)构成了一个“非常有意思的市场”:它们现在显然重要,但这种需求是否持久、机器能否把这些工作做得和人一样好,“可能有两种走向”。
- 对于芯片行业整体,Feldman 不接受 90%垄断的判断:Intel 曾主导 x86,但在手机市场占有率为零;Broadcom 则主导交换芯片。他不认为市场最终会集中到 1家或 2家公司手中。
- 说到主权 AI,他认为 Mistral 的主权战略,加上 Cerebras 通过其所谓地球上最快的硬件提供推理服务,会让其潜在的 Le Chat 产品极具吸引力;除此之外,欧洲真正做有趣工作的 AI 实验室太少。
Things are moving at a rate that 6, 8, 12 months out, everybody's unsure. It's so fast. It's so big. There is unbelievable demand, and nobody knows where it will go in the future.
The question of depreciation is: How much faster are future generations than the current generation? That's the actual question on depreciation. People often say we don't have enough power in the U.S., and this is strictly wrong. We have plenty of power. It's in the wrong places.
Risk comes in financial markets, where people fundamentally underestimate risk. No company ever went bankrupt by paying extraordinary people too much.
Andrew, dude, it is so lovely to have you back on. I so enjoyed our first show. You put up with my naive questions enough to agree to do a round 2, man. I must be charming.
Harry, I'm okay with any questions, naive or otherwise, so I'm happy to do it anytime. I read your LinkedIn posts and your Twitter posts. I'm rooting for your mom. All good. All good.
1. Why We Did Not IPO and Raised $1BN From Fidelity
Dude, you are too kind. Listen, I want to start with the billion-dollar raise that you just announced yesterday. Can you talk to me about the billion-dollar raise, why it's important, why now, and what it means for the company?
Well, look, it was the largest raise ever done in our category. It was done at the highest valuation and with the premier investors. In late-stage investing, you're looking for the likes of Fidelity. They are the—what would the English call it?—the sort of Oxford or Cambridge of investing, right? They are the premier public-market investors, and when they choose to lead a round, it brings Wall Street a great deal of confidence.
We were really happy to partner with them and with likely Atreides to lead the round, and then we were able to get enormous participation from Tiger Global, Valor, and 1789. So that's point 1.
I think point 2 is that we now have the dry powder to really push and take the opportunities in front of us: to build out our manufacturing to the scale and scope we want, to add new data centers—we added 5 this year in the U.S.—to add more data centers, and we have more big ideas.
I think incremental improvements, make-believe gains achieved by dropping from 8-bit to 4-bit, aren't going to get us to the promised land in AI. We've got real work to do as a community, and I think this funding puts us in the catbird seat for that.
On the Fidelity side, it's actually interesting. I had Brian Halligan, the CEO of HubSpot, on the show recently, and he taught me the importance of specifically getting Fidelity in both your pre-IPO round and when you IPO, just because of the signal that it sends.
I didn't realize, as a venture guy, the weight that's placed on it. You're like, "Who are the public guys? Okay, Fidelity, whatever. Sure, they're all the same," right? And I was like, "No, no, they're not. Fidelity are the monster in the room, and the importance of getting them is very high."
Can I ask you, dude: Why not go public? It was rumored that you guys were going to go public. Why do this pre-IPO round?
We still have every intention of going public. I think it's very common in late-stage investing to do a pre-IPO round if you can get it done very quickly, if it doesn't distract you, and if you keep moving. There were so many opportunities in front of us that gathering the capital so that we could continue to prosecute these opportunities was sort of a no-brainer.
2. Analysis of Chip and Compute Landscape Today
You said there is real work to be done. I think it's quite difficult for everyone who's not really in the market to understand what the hell's going on, given all the news that we see. Can you help us with a lay of the land over the last 3 months? Where are we at now? What's changed? Let's start there.
The first thing, Harry, is that we are in a stage of the market where the claims are enormous, right? Tens of billions of dollars are being committed here and there, and nobody's reading the fine print that it's over 5 years and it's "up to" this.
The great sort of CYA word in marketing history is "up to $100 billion over 5 years," right? "Up to" means it could be $30 billion, it could be $12 billion, it could be $40 billion. It won't be bigger than that.
As you read these deals, you have to really think about the time frame over which they're being done, and you have to think about whether anybody is actually counting. Lots of people are saying they're going to bring hundreds of billions of dollars and jobs to the U.S., and this and that. In 8 months, has anybody got a little spreadsheet like, "9 jobs plus 1 factory"? Who holds anybody to account? The answer is nobody.
I think that's number 1. Number 2, what this signals more than anything is that there is unbelievable demand and nobody knows where it will go in the future. It's so big and happening so quickly that they don't know.
We have customers coming to us and saying, "We would like between 5 million and 40 million queries per second." How do you not know by 35 million queries per second where your demand's going to be? The answer is that things are moving at a rate that 6, 8, 12 months out, everybody's unsure. It's so fast. It's so big.
I think you should think about these announcements as options on the future. That's really the way to think about it. In an unknown environment, how can I take an option on the future? I don't know if I'll use it all, but I'll pay something for the future rights to have some capacity. That's a way to think about it.
Given that it's so fast and so big, how do you think about planning for that uncertain future?
It's brutal. I think there's a very interesting question about, in extraordinarily rapidly moving environments, what the right planning cadence is. What you really need is good planning with changing rules rather than good planning, right?
We have to make big bets. We're making 5- and 7-year investments in data-center capacity. We are making hundreds of millions, now in terms of billions of dollars of bets, in supply chain. Those are not 3-month bets.
I think what you need to do is use different rules than have historically been used. You plan more frequently, you have a shorter view, and you take options on the future. If the future moves against you, you lose the premium on the option. You pay a little price to secure some capacity, and if you don't use it, you just go, "All right, that was a way to manage uncertainty about the future."
What do you think the chances are that you are still underestimating even your wildest demand expectations?
100%. I've been wrong. Look, if you had said a year ago, 2 years ago—pick a time—that it would have been conceivable that OpenAI would get the valuations they're getting, pick a time. It wouldn't have been conceivable 3 months ago, 6 months ago, 9 months ago. When was it a reasonable idea? The day before, right?
I think that's true with the demand we're seeing. It's true with valuations on companies we're seeing. It's true with the rate of ideas entering the community.
How much of this do you think is sustainable? Everyone argues that it's not sustainable, that a lot of it is experimental and not enduring. How much do you think is sustainable?
I would say this: There are always grumpy people who say it'll never work, you'll never beat Goliath, and truth is, most things don't work and most of the time Goliath wins, right? But there's no alpha in that. There's no money made for you or me betting on the biggest of the big dogs to continue not to lose. How uninteresting is that?
Of course, if NVIDIA keeps growing at the rate they're currently growing, 11 years from now everybody on Earth works for them. Do the math, right? However, is it possible that our economy looks very different in 5 years? Is it possible that the things we value are very different, that we have reorganized around AI?
We've seen a major bump in labor productivity. We've benefited dramatically, and the economic pie is much larger. I think that's not only likely; it's almost certain.
You mentioned that if NVIDIA continues to grow the way that they do, everyone will work for them. You just keep doubling at that rate; you multiply. I mean, you can't keep doing that.
To what extent is it completely unshakable for them at this point, given the scale and the size of the money? Jonathan Ross from Groq said on the show that they will unwaveringly get to $10 trillion within a 5-year timeline.
I hope he's long on them, then. I don't pick public-market stocks. I think in the public market you can lose money on good companies, you can make money on shitty companies, and that, for me, doesn't sit well.
3. Mag7 Value Concentration: Feature or a Bug
As an entrepreneur, as a David in the battle with Goliath, I want to make money when we build a great company, period. But can they continue to grow? I think we are seeing some things that big companies do as they begin to worry about growth.
I think they use their balance sheet more and their technology less. Right? This is something that historically large companies have done as they feared for their technical prowess.
And when you say that, you're kind of referring to investments in your OpenAIs of $100 billion, ElevenLabs, and everyone in between. You start buying businesses as opposed to winning business, and I think that we saw that with Cisco, which emerged in a dominant position, you know, from 1992 to 2001. That has been one of the strategies.
Another strategy you see is this predatory pre-announcement, where you announce B300s before anybody can get B200s. You start talking about Rubin before B200s are technically finished. You don't talk about the field-failure rates of your products, which are massive; rather, you keep talking about the future in an effort to convince people to wait, to make a good decision, rather than go with technology that's better and present.
4. NVIDIA Showing Signs They Are Running Out of Ideas
I think these are the strategies of very large companies using their strengths, and I think that's what you're beginning to see unfold with NVIDIA. Specifically, I do want to talk about the speed of chip development, but speaking of the $100 billion into OpenAI, how did you analyze that? For me, reading that, I didn't really know how to analyze it. It's so unprecedented.
Well, I think it was designed for nobody to understand it. If one wants to make something very clear in an investment—that we've invested this amount at this valuation—the deal is done. Now, if you want to make something more difficult, that's up to this amount over an unspecified amount of time, at no valuation given, or a valuation specified, but it can change, right?
It wasn't designed for you or other analysts to anchor on different things, and that's a very reasonable thing for both of them. But it's just not an analyzable thing. What beyond the fact that NVIDIA has chosen to try and lock up a portion of OpenAI's demand by investing in them? That's about as much as you can say.
Totally get you. And I'm glad that it's meant to be confusing, because I was confused looking at it, going, “What price was this? How much are they buying?”
I don't know if it was meant to be confusing.
You know, my mother goes shopping, and I ask, “It's a lovely dress, Jules. How much is it?” “Well, it doesn't matter. It doesn't matter.”
You call your mom by her first name. Hold on, wait a second. Let's go back to the important thing. You call your mom by her first name?
Oh, yeah. Jules.
Okay. You don't call her Mom? I've never called my mom, surely. I mean, I never call her. I mean, that would be “Mom” or something else, but not.
No, no. But when I get the “price on application,” I'm like, “Oh, Andrew. Oh, dear. Oh, that's what this felt like.” That's what I was like: really? No, nothing there.
I certainly don't think it was. The truth is, there may be—and there likely are—a huge number of moving parts that make it impossible to clearly describe without giving out more than they wanted to give out.
5. The Real Questions to Ask on Chip Depreciation
But you mentioned pre-announcements, like B300s and B200s, and timings of such. Are we thinking about chip depreciation in the right way? Again, I just did a show with Jonathan, and he's like, “Hey, we actually think about them on an 18-month time cycle to maybe 2 years.” I was like, wow, that's quite quick. Are we thinking about it the right way, and how should we be thinking about the amortization of chips?
We are in unprecedented waters. I think people are clearly still getting value from H100s.
And that's more than 2 years, right?
So, if you say it's a 2-year depreciation, you're empirically wrong. I mean, they are—and I think people are still getting value from A100s, though not on the cutting edge, and so that's closer to 3 or 4 years and could be as long as 5 or 6.
The question of depreciation is: how much faster are future generations than the current generation? That's the actual question on depreciation, because with depreciation, you're saying that at some point it's no longer worth using a part that's fully paid off because there's a new part that's so much faster, uses so much less power, that it's better for me to retire it. That's the actual underpinning to the depreciation question.
If I have a data center and it's 50 megawatts and I have this much capacity in it, at some point, even though my chips in it have been depreciated and I'm running them at zero cost—power plus zero depreciation cost, right?—it makes sense to move them out because the new chips are so much faster, so much better, use so much less power, I get so much more dollars per watt, and so that's the question.
If we don't, as an industry, continue to build extraordinarily better parts generation after generation, then people don't move from one generation to the next. They last longer. You depreciate them longer.
Where are we? I feel very naive for asking. Where are we in the performance-improvement pathway for chips? Are we in the “we've got 90% and we're seeing incremental gains” phase, or are we at the “we are still at the super-early stage and we have 90% of the gains to be made” phase?
I think the question in that case is whether you read people's marketing material or the actual performance results. Certainly, people's marketing material would lead you to believe that, generation over generation, there are huge gains. A little bit of engineering digging probably leads you to the conclusion that you're getting 2 to 2.5x per meaningful generation move, not more.
Right? If you compare apples to apples—8-bit to 8-bit, 4-bit to 4-bit—if you compare actual performance, you might have more FLOPs on the chip, but your memory bandwidth didn't improve more than 2x, so you can't get to them. These chips are a system. If you make one part fast and the other part doesn't move as far forward, it becomes the new bottleneck.
It doesn't matter how many FLOPs your chip has. If you can't get data onto and off the chip, those are wasted. The question isn't, “How much faster is the chip?” It's, “How much faster is the solution?” That includes memory, which for inference is the fundamental limiter for the GPU architecture. It doesn't matter how much faster the chip goes; it matters how much faster the memory bandwidth is.
On this, I was chatting to a founder in the space, and he said that what everyone fails to understand is that, although SRAM sounds great in terms of having memory—and SRAM is obviously, you'll describe it much better than me and hate me for this, but SRAM is obviously memory on-chip versus off-chip—it seemingly is great, but it's completely unable to handle scale. Although it may be quicker, for anyone who wants to do large-scale work, it is incapable at present of doing that, and that's a fundamental need and requirement of any of the large providers. Do you think that's fair, and how do you think about it?
Well, not only is it fair, it's the reason we went to wafer scale.
So let me explain. What your friend said is strictly true in that SRAM is blazing fast and low capacity. HBM is a flavor of DRAM. It has high capacity and it's very slow.
Now, NVIDIA and all GPUs, including AMD's, chose a big-capacity memory that's slow because it's perfect for graphics. You don't have to go to memory very often. You can hold a lot; you don't go very often. SRAM is blazing fast, but it can't hold very much.
So the problem on traditional chips is that if you put memory on the chip, you're using space that could otherwise be used for compute; you have a fixed amount of real estate. If you put half of it into memory, then you have half your real estate available for compute.
And so our idea was that if we built a chip that was the size of a dinner plate, we could stuff it to the gills with fast SRAM, overcoming the limitation of SRAM, which is that it doesn't store very much, by putting a huge amount down, by using more silicon area.
Now, if you're an SRAM solution today in a normal-sized chip and you're trying to do a trillion-parameter model, you use 4 or 5,000 chips. What a mess. Do you know how many cables that is? Do you know the impact to the AI? It's a horrible mess, right? And it limits you from doing things you want to do with the AI, like speculative decoding. It has all sorts of painful challenges.
On the other hand, use 1 of these, or 2, or 4, right? And it's simple. It's easy. And this is what your friend said, exactly, right? The reason we went to build a bigger chip was so we could fill it with this fast SRAM, so we could get over the traditional limitations of SRAM—that it couldn't store very much—by using a lot of space, by using a huge amount of silicon area. So your friend is exactly right. You might listen to him again in the future.
Question for you. No offense, but that seems a little obvious: okay, increase the real estate, shove more on.
It does, doesn't it? Yeah.
Is it as obvious as it seems? Am I missing something here?
Well, what we were missing is that for 75 years, nobody could do it. Building a bigger chip had proven impossible before we did it. Nobody in the history of the computer industry had been able to build a chip bigger than about 840 square millimeters in the 75-year history of the computer industry. Many people had tried and failed. After we did it, Elon tried at Dojo, and they failed.
And so it’s really, really hard. Our strategy had never been done before, never been successfully yielded, and so, while it was obvious, it was hard.
So when we think about where the market is today in terms of training and inference, do we agree that NVIDIA’s chips are much better for training than yours are, but yours are much better for inference, and that the market splits in that respect?
No, we’re faster on both, but the software challenges in training are real.
What does that mean?
It means that when a new model is built and everybody reads about it in a publication, it was done on a GPU. Everybody, to train it, takes the recipe that was originally done for the GPU and has to move it to the recipe for their hardware, whether that’s a TPU, an AMD GPU, or another dedicated chip like ours. You have to move it, and that’s a harder software lift than inference.
The truth is, nobody cares about CUDA. Nobody even cares about PyTorch. What they want is an API. So it’s literally 10 keystrokes to move from a GPU-based solution on OpenAI OSS 120B to our solution. It’s 10 keystrokes. That’s it. It’s nothing.
I think the answer is that while we are faster at training and faster at inference, it’s easier to demonstrate inference. You just put up a side-by-side to show that you’re faster than 1,000 B200s. You have to get B200s, train the model for 4 weeks or 6 weeks, stand up a cluster of our machines—it’s a bigger lift.
I think the market right now is finding it easier to move people off GPUs in inference, and the number of people doing inference is vastly higher than the number of people doing training.
Can I ask you, when you look at the inference market today, how has it developed in a way that you did not expect?
I think it’s really hard for the mind to wrap itself around geometric growth or exponential growth. There is nothing confusing about the rate of growth of inference. The rate of growth of inference is the number of people who use it times the frequency of use times the amount of compute needed per use. It is 3 different variables multiplied by each other.
The problem is they’re all growing fast, and that produces some mind-numbing effects. More people are using AI. Once they start using AI, they use it more frequently, and what they want to do with it is bigger and more complicated, so it uses more compute. You have 3 variables. The size of the market is the product of the 3, all growing fast.
We knew that going in. We see that, and it still takes your breath away.
I don’t think we’ve seen anything yet.
I agree with that. I think the reason I’m 100% sure that we’re underestimating the market is because of that premise. I think Sam Altman said it very well in terms of how people use ChatGPT, but he said, essentially, the majority of people use it like Google—a Google replacement—and actually, younger people use it as an operating system for the future, which is the right way to do it in his mind.
Absolutely right.
In 1988, likely Robert Solow, who won a Nobel Prize in economics, asked this question. He said, “We see computers on every desktop and everywhere we look except in the productivity statistics.”
There was another economic historian who jumped into the fray. His name was Paul David, and he wrote a very famous paper called The Computer and the Dynamo. What he studied was the adoption of electricity in the manufacturing sector between about 1880 and 1955.
What he showed was that, at the beginning, electricity produced very little productivity gain. It was basically used as a backup for belt-driven systems. It wasn’t until they reorganized the shop floor to take advantage of electricity that you got this huge jump in productivity.
If you roll that forward to the computer, what he was saying is, “Look, we used computers to do things we were already pretty good at.” We replaced a typewriter; we replaced general ledger accounting with spreadsheets. We were good at those things. You didn’t get a big jump.
What he predicted happened immediately thereafter: by the mid-1990s, you had a huge jump in productivity. We had begun tying them together. We built the internet. We had the first parts of a cloud, and all these things used compute differently, in ways that had never been consumed before, and you got this massive jump in productivity.
6. Energy Requirements for AI: Is it Feasible?
If you use OpenAI and its various competitors the way you use Google, you’ll see a very modest jump in productivity. If you use them in a fundamentally different way—that was Sam’s point—you’ll see a huge jump. If we reorganize ourselves around AI, you’re going to see massive productivity gains. If we use AI to replace things we’re already doing—Google or something else—you’re not going to see very big jumps at all.
What I know is that transition takes time. What he pointed to was a demographic: younger users are using it in a different way than older users. Older users are replacing something they already had. Younger users are using it in a way that never existed before, as an operating system for life.
If we’re going to see that transition you mentioned, the energy requirements are just insane. Sam said a $1 trillion spend. He needs the energy of Japan or more, to be blunt.
Yeah.
Is this feasible for the country?
Yeah, it’s feasible. Whether it’s desirable or good for society is a different question. It’s feasible. People often say we don’t have enough power in the US, and this is strictly wrong. We have plenty of power. It’s in the wrong places, right? It’s not where we have people or where we have fiber-optic cable.
We have a ton of power in West Texas in natural gas. We have a ton of power in upstate New York in hydro. We have a ton of power in lots of places. We don’t have people there.
The problem is one of a mismatch between where all the power is and where the people are, where the buildings are, or where the telco fiber is that we need to get data to and from the data center. So that’s the first observation.
The second observation is one of community: to the extent we consume this extraordinary amount of power, we have an obligation to deliver amazing things. And that’s not all of us.
I think we have an obligation to deliver drugs that are more efficacious, to deliver better health care, to make aging less painful, and to make looking after aged parents or sick parents less painful. You go through society’s ills and woes. If we are going to consume this amount of power, the burden is on us to deliver value for it. If we use it and don’t do that, then it’s not a gain for society.
Do you think that’s controllable? Creating Ghibli images—I get it wrong—isn’t particularly value-inducing, but it burns a huge amount of compute and energy. Can we control that?
It’s a very hard question. I’m one voice in this. The problem with markets is that they do a lot of things that aren’t productive in order to get one that is very productive.
Ghibli may or may not have been a net societal gain, but maybe the technology that is used in Ghibli is used for X-ray crystallography and later is fundamental to finding major scientific breakthroughs. That’s the messiness at any given point in time.
You can point to a thousand-poppy strategy, which is what a market is, right? A market has a lot of bad ideas to get a few good ones, right? That’s your business. Your business is investing behind a lot of big ideas, most of which fail.
In that environment, you can always point to, “Oh, that was a bad investment, Harry. Why’d you invest with them? They blew up.” You can always say it after the fact: “Oh, look at that. They’re using a ton of energy for that. That’s not useful.”
I think the answer is that we need to be sure that, at a societal level, where we use government dollars, tax breaks, or permitting breaks, we are giving these disproportionately to projects that matter to society.
Do you think Trump’s done more to help or to hurt the US AI effort?
I think it’s confusing. On net, it’s probably done more to help. The Biden administration was misguided and afraid. I think, to his credit, the Trump administration surrounded itself with some smart people in the AI space, and on net it’s been positive.
When you look at what is required in terms of energy, is nuclear unavoidable, or is it the sole solution for providing energy for this next generation of AI?
No, it’s not unavoidable. It’s a very reasonable decision for countries that don’t have lots of alternatives. Canada has more falling water than anywhere else on Earth, right? The opportunity for Canada to develop the cheapest power on Earth is mind-boggling.
There is cheap power in lots of places. But for countries that wish to pursue this and don’t have the natural resources of Finland, which has geothermal, or Iceland, which has geothermal, or Canada, which has falling water, nuclear is a very reasonable and cost-effective strategy, especially over a several-decade view.
What worries you most today, Andrew?
I do think about this idea that, to consume the resources we’re consuming, we have to be sure that we produce some extraordinary outcomes.
I worry the opportunity is so big that, as a community, we're running helter-skelter at it. Sometimes, instead of running, where you trip and fall and graze your knee and chip a tooth, if you stopped and thought and marched, you might get further over a 30-, 60-, or 90-day period.
Do you worry about the concentration of value in the Magnificent 7? They now make up more of the S&P than they pretty much ever have done in history, and that concentration of value is very real. If AI hits a speed bump in any way, the market could derail significantly, and the multiplier effect of that is felt by everyone.
The risk there is not that they consume that much or that they are that much value. That's not the risk. I think they're that much value because the future economy that we believe the future economy will reward that.
I think the issue is that people then think the S&P is a safe investment, or a safer investment than it might be. The risk is the mismatch in the mental model people have. Risk comes in financial markets where people fundamentally underestimate risk. When risk is priced properly, your outcomes are not surprising.
But if people continue to think the S&P is an index of the global economy—and it's not; it's 30% or 50% seven companies—then they're exposed to sector risk that they weren't signing up for. They thought they were diversified, and in fact they're heavily dependent on a very narrow sector. That's a risk.
Do you know what I mean?
I totally—
Yeah. That seems to me to be a challenge in the new world order and all the advice that the pundits give: a diversified portfolio. When the world changes and that portfolio is not diversified anymore because of consolidation, if you keep holding it, then there's real risk.
Do you think the risk is priced properly when you look at Nvidia at $4.5 trillion?
Look, I think they've proven themselves to be the greatest company of the first quarter of the 21st century. They've proven themselves to be an extraordinary company in the first quarter of the century. I don't know if $4 trillion is right, but I think a very big number—maybe it's too low—is right because of what they've achieved.
7. Talent is the Bottleneck and Trump Makes it Worse
When we look at what we've said before about the insatiable demand—the demand that we cannot predict or anticipate—what are the bottlenecks today in your mind?
We had Jonathan McGroarty on, and he was like, "Actually, I had someone come and demand 5 times the supply that I have in total, and that was from 1 customer." Supply is mine. How do you think about the bottlenecks that we have to reach the insatiable demand that you mentioned?
I think if you go back to planning, if you've got customers demanding 5 times your capacity, you probably didn't get your planning right. You probably should have planned better.
I think there are bottlenecks at every level that are meaningful. I think the first one is expertise. We have fundamental limitations in AI expertise. We're not making enough AI practitioners. We're not making enough data scientists who understand data pipelines.
Our universities aren't minting enough, and our challenges in the US with immigration don't help that. Historically, we've sucked the best and the brightest first on J-1s to come to our schools and then H-1Bs to stay. If that is not our policy, we need to make it our policy. If the government decides that that is not the way they want to build a workforce, and instead they want to build it out of people who live here, we need to do a better job of training those people.
We need to do a better job of teaching them in K through 12. We need to do a better job of educating them in our universities in order to make the number of engineers we need to meet this demand. That's a bottleneck, and it's why the best and the brightest are getting such extraordinary compensation.
Is the war for talent completely out of control? You're seeing your Zucks of the world spend hundreds of millions on 1 person. Do you think that's a blown-up anomaly, or do you see the war for talent being unprecedented?
There are engineers who have skills that no number of other engineers working together can achieve. There are scientists who have ideas and brains that can't be replicated by lots of other talented people working together. Ought they to be paid more than world-class soccer players? I have no idea. Maybe, maybe not.
I mean, inherently, yes. From an economic rationale standpoint, yes. The value generated from a chief scientist at OpenAI, if they add $50 billion of enterprise value, to pay them $1 billion is worth it.
That's what we have to think about. I don't know. We paid Charlie Sheen $2.5 million an episode for Two and a Half Men. I'm pretty sure that there are lots of people whose net productivity to society is above that.
And he still spent it all.
I just saw the show—was it Netflix or Prime? What a sad story of somebody who was so self-destructive and so talented.
But should we be paying soccer players or basketball players? I have no idea, and I don't spend a minute worrying about whether we're paying extraordinary people too much. I think no company ever went bankrupt by paying extraordinary people too much. If you want to go bankrupt, pay mediocre people too much. That's how you mess up.
8. Evaluating the Data Centre Economy: Many Will Lose Money
Nobody's ever struggled by paying truly extraordinary people too much.
What's the other bottleneck? You said expertise is 1. What's another?
TSMC can't build fabs fast enough. I think the truth is that, for both TSMC and Samsung, these fabs are the most amazing manufacturing plants on the planet. These are $30 billion to $50 billion factories, and their ability to build them quickly enough is very much limited.
I think that, in turn, limits and keeps the supply below where it would like to be of chips—not just our chips or NVIDIA's chips, but everybody's chips—below where it might otherwise be. It keeps the cost up.
Right now, there's a shortage of data center capacity. I think there's a huge amount of investment that has gone into that. There's a lot of talk, but where are these gigawatt facilities that everybody's been talking about? Everybody's committing to them. Where are they? Well, they're not up yet.
How long does that take? For somebody like Elon, who's the fastest in the world and maybe the best at building plants and large construction projects, it takes 6 or 8 months. For the rest of the world, it takes a year and a half, maybe longer.
Are we investing enough in data center builders? It is one of the most insanely hot categories now in terms of investment properties. I'm coming from a pure Wall Street mindset. Every Wall Street guy wants to be in data centers.
Yeah. It has a structure that they really understand.
Right. It looks like a bond to them.
It looks like a piece of real estate. You get a tenant, they pay rent every month, and you can loan against that. You get an investment-grade tenant that's basically a bond.
It has the advantage of falling into a category or pattern that is really well understood in the debt market and in the capital markets. That's an advantage, and I think CoreWeave and some of their financial engineering and innovations there help the world see that.
Like many things, lots of people will enter. The smart will make money; the less sophisticated will lose money. I think building data centers is not for everyone.
How will you lose money building data centers?
Look, I think if the best can build them for $8 million a megawatt and you're spending $12 million or $14 million, that's how you lose money.
You lose money because it begins with: Can you get access to low-cost power? It then continues to: Once you have access, can you get permitting? Does that take a long time, or do you have real access that gets you fast permitting? Once it becomes a construction project, can you keep control of your costs?
Once it's finished, can you keep good tenants in it? The ways to lose money in property are large and many, and there's no free lunch there either. When you are trying to go unbelievably quickly, it's harder and harder to be disciplined and not make mistakes.
To what extent is it important to be fully horizontal? We hear about Zuck wanting the data center buildout to be immense in terms of size and scale. To what extent does it need to be horizontal versus vertical?
It is completely unclear. The 2 most successful companies to date, OpenAI and Anthropic, are neither vertical.
OpenAI used Azure for 100% of its infrastructure for years, and Anthropic has used a combination of AWS and Google. Neither are vertically integrated to date.
Whether that's the right strategy going forward, whether they'd make those decisions again, who knows? But it's clear that it's not the only strategy. There are plenty of working models where you are not fully integrated from chip through system, through data center, through software, all the way to the top.
Again, sorry to cite it, but it’s kind of handy having just done it. Jonathan said that you would definitely have OpenAI and Anthropic build out their own chips because then they would have control of their own destiny. Do you think OpenAI and Anthropic build their own chips so they don’t have self-reliance on NVIDIA in the way that they do today?
I think there is a long history of software companies failing to build chips. The list is very large. I think whether OpenAI can do it, and whether they can do it through partnerships with other vendors, with Broadcom, or with smaller, more innovative companies, is an open question. But I think companies the size of Microsoft have been unable to deliver chips.
There are plenty of examples as you look across the FAANG group where chips were tried. Probably the most successful is Google, and they’re 10 years in, maybe longer. Modern software does not fit well in a chip-making framework. Weekly sprints don’t work well on 2-year-long projects.
Move fast and break things often is not the way you think in the chip world. The way you think in the chip world is, “Measure twice before you cut once,” because your bugs cost you 6 months and tens of millions of dollars. It’s a very different mentality.
Where there’s been success, it has frequently been acquired. Apple got into the chip business through buying PA Semi. Amazon got into the chip business through acquiring Annapurna Labs. Google acquired the talent from a collection of companies and then set it in a BU that was set aside and under somebody who had enormous respect in the organization and who had a 10- or 15-year view. These are things that have been challenging in many companies.
Chip building is an MBA nightmare, right? Your analysis says, “Look, Intel had, between 2000 and 2010, some of the world’s leading architects and the world’s leading fabs, and proved completely unable to build a working cell phone part.” You ask yourself why. They had everything they needed, and you do an MBA chart and it’s like, you cannot—it’s impenetrable.
The answer is this is really hard, and the very small mental-model differences produce tremendously different results. How did every leader miss the largest compute market right in the first part of the 21st century? How did AMD miss it? How did—how did, I mean, how did ARM win it? All the leaders missed it. Then you say, “All right, maybe there’s something in the guts here that I don’t understand.”
You’ve got to really get in there. It’s not on a PowerPoint, it’s not in a 2-by-2, and it’s not at some sort of consultant level. It is deep in the DNA of the small number of people who can build these things. We are lucky at Cerebras. We’ve got one of the top 6 or 8 teams in the world, and other startups don’t.
What does that market look like, do you think, in 10 years’ time? I know 10 years is a huge amount of time given where we’re at, but in 10 years’ time, is it a monopoly market with one taking 90%? Is it like cloud?
Which market?
Specifically, the chip market.
Which part of the chip market? Are we talking about AI silicon, or are we talking about silicon in general?
I would say silicon in general.
Absolutely not one takes 90%. Even at Intel’s strength, they had dominance in x86 and zero market share in cell phones, and almost no share in the switching market. Broadcom had dominance in the switching-silicon market, which is a form of processor and silicon, and no share in x86 or other forms of compute. It will not all accrue to 1 or 2 companies.
How do you think about the importance of margin today as a business at Cerebras?
I think the reason we were able to raise at a higher valuation, from better investors, and with more money is because we had them, and others who were out looking for money had negative margins. I think, as you prepare for being a credible public company, people do look at your margins, and I think that’s a really important part of moving from being an idea to being a real company.
What are NVIDIA’s margins today? Extraordinary—some of the highest in history for a hardware company. How do you think about that? Is that just pricing power, which they are taking advantage of?
Absolutely. The short answer is, why does it make sense for AWS to build a Trainium part? Because they want to get rid of the 78% gross margin that NVIDIA is charging. That’s why it makes sense. On the high-end chips, it might be 85%.
People don’t like that historically. Historically, people sort of put it in the back of their mind and remember it. When Intel stumbled, the number of people who came out of the woodwork to kick them when they were down was extraordinary. Years of pent-up frustration came out when the giant stumbled, and I think we’ve seen that again and again.
Speaking of that giant stumbling and being built, do you think sovereignty will be a big enough reason why incumbents are built? We have Mistral, a model provider in Europe, and sovereignty is their core play. Do you believe that is a sufficient enough core play to be a giant?
Right now, sovereignty, plus the fact that we deliver their inference through the fastest hardware on Earth, makes their product—the Le Chat product—really compelling. I think they are using their advantages to compete. There were too few, if you want my opinion, too few AI labs in Europe that were doing interesting work, and they looked around and used a strategic advantage: “We want to be Europe’s leader.” They played that card really well. Hats off to them. Then they raised at a huge valuation.
9. Three Changes the US Could Make to Beat China in AI
Final one, just in terms of geography. DeepSeek obviously had their moment, and it kind of solidified the concerns around China. How do you feel about China today as a pressing concern toward the US in terms of the race toward AGI between the two? Do you hate the way that it’s posited as China versus the US, the AI race? How do you feel about that?
I think it benefits neither—the position we’re in, right? The arms race certainly didn’t help either the US or Russia in the ’80s and ’90s. We both spent money on weapons that we wish would have been spent on infrastructure, people, or other things.
I think we will be much stronger if we can find ways to peacefully engage before these issues. We knew the guys at DJI, ByteDance, Alibaba, and Baidu extremely well. They’re talented engineers trying to build cool stuff. I think our governments were at loggerheads, and that’s a problem.
Of course, we made choices. We had a huge opportunity in China in 2019, and I decided to pass because I didn’t think it was the right thing to do, long before the Department of Commerce limited exports to China. I didn’t think it was right, and I was concerned about how the technology would be used.
But I think the realpolitik right now is that they’re better at making drones and better at making robots. Their government has an extraordinarily aggressive policy in AI. For years, they backstopped their venture groups, right? So if you lost money in an AI company, the government would make you whole.
Imagine that, Harry. Imagine how much money you could make if the government of the UK offset some of your losses from AI companies that didn’t work out. We have real work to do in the US.
What work do you have to do that you haven’t done? What would you like to see?
China thought long and hard about its power infrastructure, and its form of government allowed it to plan strategically. Our decentralized form of government has left us with a patchwork of power infrastructures, where even if the federal government wants to support you, there are local regulations at the city and county level in towns that can interfere with a project and set it back billions of dollars.
Samsung built a fab in Texas, and they had to change the design of a fab because of a local fire ordinance. The US government worked for years to get deployment of billions of dollars in Texas, and a local fire ordinance set them back 8 or 10 months and caused them to redesign the fab. That’s a problem, and that’s a challenge that we have to collectively work through.
I think we have the premier universities. We have historically drawn talent from around the world. If you look at, say, the great CEOs in our industry—likely Jensen, likely Hock, Lisa—I mean, you go down the list: Sundar, at Microsoft. They came; their parents came. We’ve got to take that really seriously.
You don’t buy the whole, “Well, actually, a load of people just abused H-1Bs, and we’ll just move to O-1s, which people were using anyway, and the average salary for an H-1B was $120,000, and so it’s a good thing, and people will just use O-1s”?
We have H-1Bs and we have O-1s. I am sure that in every government program there’s an amount of abuse. I’m not saying there’s no abuse. Was there more abuse in the H-1B than in other areas? I don’t think so.
Having the best and the brightest come to your universities and, once they benefit from our great institutions, wanting them to stay and contribute—first with a J-1, which is the student visa, and then entering the H-1B lottery through the approved process to get a green card and become citizens. This is how my parents did it. I think it’s one way to bring an extraordinary amount of talented people to the US.
Is there anything else you’d change? You said the power infrastructure and the permitting around it.
Power infrastructure ends up at the local level, which is not necessarily where big ideas and strategy are well knitted together. I think we’ve starved our universities of compute. If you want to do interesting training work at a university, it’s very hard to get enough compute to do that. We’re just not set up for that.
We have power and people. Those are 2 dimensions. I think the Trump administration has done a good job generally relaxing some of the regulations that were painful.
10. Quick-Fire Round
Andrew, I want to do a quick-fire with you. I’m going to pummel you with quick questions, and you’ve got to give me your immediate thoughts.
That’s hard because I only have long answers, Harry.
It’s totally fine. You’ll be honest. What do you believe that most around you disbelieve?
We will have peace in the Middle East in our lifetimes.
Why do you believe that?
Because I believe, having visited and spent time now in the UAE, Saudi Arabia, and Qatar, that the returns to moderation—the economic gains—are enormous. Someone said, “We’re too busy to hate right now. We’re too busy building.” I think those gains have been writ large so clearly in the UAE, with the rise of Dubai and the UAE, in return for making peace with Israel and in return for a more moderate position.
I really believe that that is the path to the future.
How much of your revenues are from the UAE?
I think in the S-1 it says—and that was for maybe the first half of 2024. We haven’t published the others, but a lot, I think—75%, 80%.
I mean this in the nicest way: do you not have to say nice things about it then? Like, if someone’s giving you—
Fair question. No, I went there to do business as a Jewish guy before we had any business done, right? What I found surprised me.
We don’t do much in Saudi Arabia, and I think they’re making great strides. We don’t do anything in Qatar right now, and I think they’re making great strides. So I don’t think it’s just—it may well be colored by the fact that I spent time in Abu Dhabi, I spent time in Dubai, I spent time in Riyadh, and I spent time in Doha. Sure, it’s colored by those things.
Why are your revenues concentrated there? Is it just because they’re more willing to embrace innovation, new relationships, and new vendors?
No, I think they bought so much that they consumed—and the data I gave you was through the first half of 2024. They placed such big orders that they consumed all our manufacturing capacity.
I mean, they were building at such extraordinary rates that, through the first half of 2024, they consumed an enormous amount of our manufacturing capacity.
Did their orders exceed your expectations?
I think their orders exceeded everybody’s expectations. You can be a professional salesperson in Silicon Valley for 20 or 30 years and not see a $500 million order.
Really, I think you can go around the Valley right now and talk to VPs of sales or EVPs of sales at dozens of public companies who’ve never seen an order of that size. They were bold, and they were early. When we started doing business with G42, nobody had heard of them, but now everybody in the world has heard of them.
Do you think that was a resource-planning mistake on your part? I mean that nicely, but, like you said, was it a resource-planning mistake?
They’re all mistakes in retrospect, right? If we hadn’t won them and we had the resources for it, that would have been a resource-planning mistake. I’m in the business of making big bets and making lots of mistakes, Harry.
What’s the biggest bet you’ve made with Cerebras that didn’t work out?
My bets here have been pretty good. We went to wafer scale to solve a problem that nobody had previously solved. Gene Amdahl, one of the fathers of our field, failed. IBM failed. TI failed. Everybody failed at this.
We had a period of about 15 months, between about 2017 and early 2019, where we couldn’t make one. We were running a burn of about $6 million to $7 million a month, and we stayed with it. Our board stayed with it.
Did you have signs that it would work?
Yeah, we did. We weren’t running around like chickens without our heads. We were going through the engineering process. Each failure was—you know, we did a full FA, a failure analysis. Each time, we fixed the cause.
We did another one; it didn’t work. We did another one; it didn’t work. Each time, we got a little better, and we got better and better and better. Then we solved it.
When the first one worked, the founders were in a tiny little lab that was a converted conference room. For cooling, we had the windows open, and we’d blown a hole in the wall so we could get an external chiller outside and pipe it in.
When we had it running, the founders stood there together and stared at the box running, which is about as interesting as watching paint dry. We stood there, and we couldn’t believe it. It was like, “We have just solved a problem that, for 75 years, the smartest people in our industry have been unable to solve—and we have done it.”
We stood there for about half an hour, and it was one of the highlights of my career.
Pretty cool. All right, I’ll give it to you. That was a big, big bet. That was fair enough—the $6 million to $7 million a month burn. I’m like, “All right, fair.” I’m almost picturing angels singing and tears coming down your face.
You know what? It felt like that, and it was the brainchild of my co-founders—Gary, Sean, J.P., and Michael. It was their invention and a physical manifestation of their ideas.
Where are people investing today where they will completely lose their shirt?
The silicon industry is not a place for 25-year-old CEOs, no matter how smart you are. The returns to having built parts before in what we do are enormous, and the number of different relationships that are necessary is huge.
You need a relationship with the fab, a relationship with the EDA toolmaker, back-end design engineers, logic design engineers, and IP relationships with IP providers. It has been an extremely difficult road for young CEOs.
On the other hand, young CEOs in many of the markets you invest in have done extraordinarily well, particularly where they look like their customer. The reason that the entire social-networking world was built by young founders is that they were building a product for their friends, and that is an advantage.
The reason that AI startups doing tools for other students and coders are young is because they understand the needs and demands of their target customer base extraordinarily well. I think that’s an area where people are going to get clobbered: taking a mentality that says it’s enough to be smart in this field to build a good chip.
That has historically not been the case. There are real returns to having done 15 or 20 of these in the past.
Where are people not investing enough, where they should be investing more?
I think there’s this collection of extremely unsexy things that are causing tremendous pain across the industry: data cleaning, your data pipeline. Nobody puts “data pipeline expert” on their LinkedIn profile, and yet these are some extraordinarily valuable people.
Nobody leads with “a leader in the cleaning and tokenization of data,” and these are extraordinarily important roles. I think many AI projects fail on those fronts. They have nothing to do with the AI. They fail because the data was a disaster. They fail because everything except the AI was a failure.
I think that’s an area that’s profoundly underinvested in.
What do you think the data-provisioning market looks like? We see Surge AI, Mercor, Invisible, Turing, and Handshake moving into it more and more. What does that market look like in 5 years? All of them are above $100 million. What the fuck happens to that category?
I’m different from some of your guests. There’s a lot of stuff I don’t know, and I’m not afraid to just tell you. That is a very curious market.
Scale sort of pioneered it. I think Turing was in a completely different market and pivoted to it and found great success in it. There are these collections of others. I think clearly the provisioning of value-added, tagged, or evaluated data is really important.
Whether it’s durable, whether we get machines that do it every bit as well as people, is a question that’s really hard to answer right now. That’s why it’s curious: it’s clearly important now, and the question is, will it clearly be important in 3 years? That’s a question I don’t know the answer to. Maybe—it could go either way.
What’s your craziest prediction in terms of how AI reshapes the future in 5 years? For example, Jonathan said, “Hey, I think AI will create massive labor shortages.”
It will create so many jobs for so many people that we will have massive labor shortages in 5 years.
Absolutely wrong. Economic dislocation isn't resolved in very short periods of time. That might be true in 15 years, but I certainly don't believe that will be the case in the 3- to 5-year time frame.
I think the adoption of AI, or the diffusion of AI into the economy, will nibble its way in. Let's ask this question: AlphaFold solved one of the hardest problems in chemistry, a problem that had been open for years. Name a drug that's resulting from it. Not one.
Now, I believe there will be one, but AlphaFold is, what, 4 years old now? 3 years old, right? This was a massive breakthrough for which the inventors were given Nobel Prizes. Where's the drug now? Show me the medical benefit. It will get there. Continuations of the model will have fundamental impact, but where are the X-ray crystallographers who were displaced because of it?
That was what X-ray crystallographers were doing, only physically. They're not out of work. In fact, there's more demand for them.
I think it will have really interesting effects on the way we educate children, and that's an interest of mine. We've been educating children the same way since Alexander the Great was tutored by Aristotle, right? It's like: get a smart person. They're older. They stand behind you. They tell you what to do. You read. You talk to them about it. They correct your paper.
This form of instruction has been unchanged. Maybe YouTube changed it a little bit in that you had different instructors. But imagine a system where, for example, you made a set of mistakes in your math work and the result wasn't read on a paper: “Oh, look, you got it wrong here.” Instead, they compared the type of mistakes you made to the type of mistakes thousands of other students made and said, “For this group of mistakes, we have found that the following workbook is extremely effective at remedying this hole in their thinking.”
Nobody differentiates and modifies the training based on the type of error the students are making. That's exactly what you ought to do. So I think the way we teach will change a great deal.
I think what it means to be entry-level in a company will change a great deal, because what entry-level has generally meant at consulting firms and at investment banks has been doing shit work—in particular, being really good at spreadsheets and writing summaries of other people's research. AI will be better at that. That will change a great deal.
I always thought that was a terrible way to spend the extraordinary years when you're 22 through 24. You're coming out of top schools. There's so much to be learned, and there's so much you can contribute, but to do huge hours of spreadsheets, I think there's vastly more productive thinking that those students, those young people, are capable of, and more learning that they can do. Therefore, they can be vastly more productive in the following years. I think AI will change that.
Final one for you, Andrew. What would you do if you knew you wouldn't fail or couldn't fail?
I guess I don't—I’ve never thought of that. I look at it the other way: every day, I go to battle with Goliath. Every dollar we sell is a dollar that, if we didn't work at it, if we didn't think, if we didn't invent, if we weren't 10 times better, would default to NVIDIA.
Before I competed with NVIDIA, I competed with Cisco for 15 years. Every dollar that we sold there, if we didn't build a better product, if we weren't more aggressive, if we weren't more creative, would have defaulted to the market-share leader.
I take great pride in facing every day the most wicked curveball pitcher. For your cricket example, the scariest spin bowler, the scariest speed bowler. I enjoy that. In a quiet moment, you sit back and say, “I'm competing with every disadvantage against the absolute best in the world every single day at work.” I love that.
What's more, everybody's betting against me except a very small group of people who you named, who stood up early on and said, “Maybe he can beat them.” That's the life I've chosen and the career I love.
Dude, I absolutely love that. That is such a good ending as well. I do many shows, and there are some endings where you think, “Oh, we can't end like that. That's a depressing ending.” That's a fantastic ending.
I so appreciate that. I so appreciate you. You're my go-to when I'm trying to understand what the shit is going on. Thank you so much for explaining this to me today, dude.
Well, look, I'm happy to jump on. My view of your LinkedIn posts, by the way, is that there are exactly 2 things I've read in my life that feel like they understand what entrepreneurship is.
The first is Ben Horowitz's book, The Hard Thing About Hard Things. The other is your tweets and your LinkedIn posts. I think they have the feel of what my life is.
This notion that somehow you can achieve greatness, that you can build something extraordinary by working 38 hours a week and having work-life balance, is mind-boggling to me. It's not true in any part of life.
Your willingness to jump in and say, “No, that's not how it's done, guys,” I mean, you can have a great life, you can do many really good things, and there are lots of paths to happiness. But the path to building something new out of nothing and making it great isn't part-time work. It isn't 30, 40, 50 hours a week. It's every waking minute.
Of course, there are costs. It's probably true for world-class athletes, too. If you listen to what Ronaldo talks about, he worries about everything he puts in his body. He trains every single day. They work on rest, right? Rest isn't rest. Rest is something you work on so your body rejuvenates faster. These guys are the best in the world at everything.