Jonathan Ross:DeepSeek 专题——OpenAI 和美国政府应如何回应|E1253
- Ross 的核心判断是:DeepSeek 是「Sputnik 2.0」——NASA 太空笔对俄罗斯铅笔的故事“刚刚又发生了一遍”,但 600万美元是营销话术。 “他们在训练上确实花了大约 600万美元,但在蒸馏或抓取 OpenAI 模型上花得多得多”;既然 OpenAI 据报道每个 API token 都在亏钱,它实际上“无意中补贴了这个模型的训练”。真正的创新是全自动、可验证奖励的 RL,整个过程不需要人类介入。
- 模型如今已经赤裸裸地商品化——“如果之前还有任何疑问,现在疑问已经结束了”——而 LLM“完全没有切换成本”,云计算的类比因此失效。 如果 Ross 是 Sam Altman,他会选择“为 OpenAI 的模型开源做准备”来回应——“开放总是赢,永远如此”,Linux 已经证明了这一点——因为“很明显你会输掉这一局,不如试着赢得所有用户和他们的喜爱”,再退回 OpenAI 真正的七大力量护城河:品牌。
- 5000亿美元的 Stargate“不是花得太多,而是花得不够”。 Ross 在 TPU 时代看到 Google 的推理投入是训练的 10–20 倍;Harry 认为 Jenson 说过 Nvidia 的收入已有一半来自推理;Ross 则认为推理占比最终可能达到 95%——“你不会为了成为心血管外科医生而训练,然后一生只有 5% 的时间真正做手术”。测试时算力进一步放大这一趋势:DeepSeek 的一次回答烧掉了 18,000 个中间 token。
- 可交易的判断是:Harry 在 16% 的暴跌中“刚买了一大笔 Nvidia”,称这是“本世纪最明显的买入机会”。 Ross 也从称重机视角认同这一判断:Nvidia“实际上因为 DeepSeek 更有价值了,而不是更没价值”。Jevons 悖论意味着,算力成本每十年下降约 1,000 倍,消费量却上升约 100,000 倍,因此支出增长 100 倍。训练是高毛利的“主机”利基市场;推理才是更大的市场,而 Groq 承接低毛利规模业务“可能是 Nvidia 股票有史以来发生过的最好事情”。
- CCP 的风险在数据,而且情况很快会变得更糟。 “现在 CCP 可能会把武器上的保险栓拿掉……现在我们要数据了”;Harry 认为北京把 DeepSeek 当成另一个 TikTok 的概率是 100%。与此同时,出口管制只是做戏:“你完全可以登录、刷信用卡、租 GPU”——“这就像马奇诺防线,绕过去就行了。”
- 接下来所有人都会复制 DeepSeek 的极稀疏混合专家架构(约 671B 参数、约 250 个 2B 参数专家,每次只激活少数几个),以及基于合成数据的再训练——Llama 3.3 70B 已经通过在更好数据上的微调击败了 3.1 405B。DeepSeek 只允许中国手机号注册,说明“他们的推理算力用完了”:训练规模取决于 ML 研究员数量,推理规模取决于终端用户数量。
- 给基础模型公司的建议是:“转向,接受现实——直接转向。” 模型是引擎,不是汽车;当幻觉率下降时,Perplexity“处在完美位置”(在今天的模型基础上构建应用,就像“还没有智能手机时试图创建 Uber”);而欧洲的处方是:今年年底建成 100 个 Station F,明年达到 1,000 个。
1. Sputnik 2.0——但 600万美元是营销话术
- Ross 毫不掩饰 DeepSeek 的意义:“是的,它就是 Sputnik 2.0。”但成本故事是包装——相比约 600万美元的 GPU 时间,他认为 DeepSeek 在“蒸馏或抓取 OpenAI 模型”上花了更多钱;GPU 时间大致相同,都是 4,000 张 GPU 跑 30 天,与最初的 Llama 70B 相当(Llama 第一代模型的 GPU 时间成本约 500万美元,“而且它以一种好的方式让世界为之震动”)。他的总结是:“他们真的很擅长营销。”
- 机制很清楚:Scaling law 假设数据质量均匀,但更好的数据胜过更多 token。AlphaGo Zero 展示了这条路径——训练、生成更好的棋局、重新训练、继续升级。DeepSeek 的捷径是:“既然已经有一个非常好的模型摆在那里,就让它生成数据,然后你就‘嗖’地一下直接追到它所在的位置。这就是他们做的事。”
- Ross 拒绝“China 只会复制”的框架——RL 的创新确实来自极简设计:不是由人类给输出打分,而是“这是一个盒子,把答案输出到这里,然后检查它……完全不需要人类介入,整个过程自动化”。至于 DeepSeek 在奖励建模上的创新,他保留了一个诚实的“不知道”:“这块我不太熟……你不如告诉我你看到了什么,我可以判断它是否合理。”
- 他指出其中的讽刺:OpenAI 可能每个 API token 都不赚钱,却“无意中补贴了这个模型的训练”——每个 token 亏一点钱,DeepSeek 则拿走了训练数据。OpenAI 很可能仍然拥有这些数据,可以直接用来训练;没必要反过来蒸馏 DeepSeek,因为“他们其实仍然更强”。
2. 出口管制是一条马奇诺防线
- 最大的漏洞是:没人需要走私芯片,因为“你完全可以登录、刷信用卡,从任何云服务商那里租 GPU”。“这就像马奇诺防线——绕过去就行了。还需要把它封得更严一点。”
- Groq 会封锁中国 IP 地址——“我相信我们可能是唯一这么做的公司”——但 Ross 也承认这“多少有点徒劳”,因为任何人都可以在其他地方租一台服务器,再从那里登录。他对整个管制体系的结论是:“这是一堵巨大的瑞士奶酪墙”,IP 封锁可能本来就不是正确工具。
3. CCP 会把 DeepSeek 当成另一个 TikTok——100%
- 数据问题“可能是最重要的”:即使是善意的公司也不会真正删除数据——“他们只是在你的数据旁边写上‘删除’……数据仍然在那里。你真的认为 CCP 没有你的所有数据吗?”而且暴露是集体性的:邻居的投诉或配偶的健康数据,都可能让你陷入风险。
- Groq 在 2016 年决定不进入 China,出发点是商业而非地缘政治,并由此形成一条公式:“你必须给 China 的钱多于从 China 拿走的钱”——同时还要交出所有数据,并接受对答案的塑形。如今的信号很明显:低温度设置下的 DeepSeek 不会讨论 Tiananmen。Ross 描绘了更可怕的版本:“TikTok 要不要禁?绝对不要,理由如下——而且它还能给出一个有条理的理由。这就有点吓人了。”
- 他的预测是:DeepSeek 原本像一只自主行动的对冲基金,但“现在 CCP 可能会把武器上的保险栓拿掉……他们会问,为什么要把这个模型开源?现在我们要数据了。”当被直接问到北京是否会把它视为另一个 TikTok 时,Harry 的回答是:“100%。”Harry 的反驳是:TikTok 明天可以被禁掉;“但这里是开源的”——没有关闭开关。
- Groq 为什么打破自己的规则托管 R1:DeepSeek 登上 App Store 第一名后,人们无论如何都会把数据交进去——因此不如提供一个替代方案,并承诺“我们什么都不存”——“我们甚至没有硬盘……电源一断,一切都消失。”
4. 商品化已经赤裸裸——Sam,开源吧
- “这件事已经把模型商品化的事实彻底暴露出来……如果之前还有任何疑问,现在疑问已经结束了。”Ross 用 Hamilton Helmer 的七大力量分析每一笔投资,也要求 Groq 的每个人都填写这份分析;他认为 OpenAI 真正拥有的力量是品牌——“这个领域没有第二家”,而 Stargate 是 Sam 试图从品牌跨向规模经济。
- 他的建议是:“如果我处在那个位置,我会开始准备开源自己的模型作为回应——很明显你会输掉这一局,不如试着赢得所有用户和他们的喜爱。”担心自我蚕食?人们仍然愿意为 Dell 而不是 Super Micro 付费,因为信任仍有价值;品牌会在产品免费后存活。唯一的问题是时机:现在做看起来像是在回应,而不是有意为之……而且它确实是在回应。
- 开放为什么会赢:当年所有人都认为开源更不安全、bug 更多,但 Linux 仍然赢了;“现在人们反而期待开源产品更安全、bug 更少、功能更多——专有软件怎么可能赢?”LLM 完全没有切换成本,因此云计算的类比“完全站不住脚”,尽管 Linux 存在切换成本。Meta 依靠网络效应,甚至可以把一切免费送出去——“越是走向开源,他们的优势就越大。我对此完全嫉妒。”
- 他想象中的 OpenAI 内部状态是:基层员工在问“我的股权还值钱吗”,高层则在努力维持士气。值得保留的一句话是:“糟糕决策的第一驱动力是恐惧。”他们必须做出选择、坚定投入,并且勇敢执行。
5. 5000亿美元 Stargate 还不够——推理占比将升至 95%
- Harry 问,DeepSeek 难道不是在嘲笑 5000亿美元的公告吗?Ross 的回答是:“实际上,我觉得这还不够花。”先例来自 Jeff Dean 在 2011–12 年左右给 Google 管理层做的两页演示:第一页是好消息,机器学习终于有效了;第二页是坏消息,我们付不起这个成本——仅仅为了语音识别,就要把 Google 全球数据中心规模扩大 2–3 倍,投入 200亿–400亿美元。
- Ross 在 Google 的那个时代看到,推理投入是训练的 10–20 倍;Harry 认为 Jenson 说过,推理如今已经占 Nvidia 收入的一半;Ross 则认为未来可能达到 95%。“你不会为了成为心血管外科医生而训练,然后一生只有 5% 的时间真正做手术——情况恰恰相反。”测试时算力会进一步加速这一趋势:DeepSeek 的一次查询在给出答案前消耗了 18,000 个中间 token。
- 5000亿美元是否真实?Gavin Baker 在 Twitter 上算过一笔账,Ross 独立计算后得出了“惊人相似的数字”;但业内人士说他们确实有这笔钱——“可你继续追问,就会发现,可能里面有些巧妙之处。”他的判断是:Stargate 承认模型已经商品化,基础设施才是护城河;但资本开支推进缓慢,“真正的胜负手在品牌。我会尽可能聘请最好的品牌公司。”3年后的 OpenAI 品牌会“强大得多”。
6. Jevons 悖论——DeepSeek 让 Nvidia 更有价值
- Harry 透露自己的实时仓位:“我刚买了一大笔 Nvidia”,是在 16% 的暴跌中买入的,称这是“本世纪最明显的买入机会”。Ross 借用 Buffett/Munger 的说法回应:短期市场是人气投票机,长期是称重机;从称重结果看,Nvidia“实际上因为 DeepSeek 更有价值了,而不是更没价值”。抛售逻辑假设算力主要用于训练,而便宜模型意味着需要更少芯片;这两个前提都是错的。
- Jevons 悖论是,蒸汽机效率提高后,煤炭购买量反而增加,因为“当运营成本下降,更多活动就会变得有利可图”。Ross 说他在 Sacha 之前发过这条观点:“正如 Sacha 喜欢说自己让 Google 跳舞,我让 Sacha 跳舞。”过去 5–6 个十年里,算力成本每十年下降约 1,000 倍,消费量却上升约 100,000 倍,因此支出每十年增长 100 倍。每次 token 价格下降,Groq 看到的开发者数量都会“飙升”。
- 关于“你的利润就是我的机会”与利润是否构成防御,Ross 的判断是:“训练是一个高毛利的利基市场”——主机业务,每年规模仍有数千亿美元——而推理是更大的市场。Groq 承接“低毛利、高规模的推理,让 Nvidia 可以继续保持漂亮的利润率”,这“可能是 Nvidia 股票有史以来发生过的最好事情”。即便在 2024 年底融资时,他仍然需要解释为什么推理优于训练;这一直是 Groq 自 2016 年以来的核心论点。
7. 接下来会复制什么:稀疏 MoE、合成数据,以及 DeepSeek 算力耗尽
- Ross 认为,R1 大约有 671B 参数,而 Llama 只有 70B;其结构大致是 250 个 2B 参数专家,每次查询只激活其中少数几个——“你的大脑也不是每个神经元都会放电”。更多参数可以从更少的数据中保留更多信息,稀疏结构则跳过不必要的计算。Ross 回忆称,GPT-4 据报道曾经大约有 16 个专家,后来缩减到 8 个;DeepSeek 反其道而行之,“聪明之处之一就是想清楚如何拥有这么多专家”。
- 即将到来的复制浪潮已有先例:Meta 的 Llama 3.3 70B 击败了自家的 3.1 405B,而且不是从头重新训练,只是在少量高质量数据上做微调。现在,拥有数十万张 GPU 的公司都会围绕这一架构大量生成合成数据并训练:用户少的场景使用更大模型,用户多的场景使用更便宜的模型。
- DeepSeek 限制新用户只能用中国手机号注册,只有一种解读:“他们的算力用完了”——具体说,是推理算力。底层规律是:“训练规模取决于你有多少 ML 研究员;推理规模取决于你有多少终端用户。”这也是为什么芯片初创公司“会做得很好”。
8. 兴奋剂、研发窃取,以及欧洲的 1,000 个 Station F
- China 实行“RDT——research, development, theft(研究、开发、窃取)……这就是文化的一部分,而且不只是针对西方公司,彼此之间也一样”。他提到,Huawei 的交换机曾经启动时显示 Cisco 的标志,“而且还带着所有 bug”。西方是否应该反过来窃取?Ross 说:“这在直觉上让我感到恶心——我真的对这个想法感到排斥。”Harry 的反驳是:和一个服用兴奋剂的人比赛,就必须服用兴奋剂。Ross 让步称,也许政府必须介入——他愿意和 DeepSeek 那些“非常聪明的人”进行公平竞争,但“政府一直在用手指压着秤”。
- Harry 提出,CCP 可能会像补贴 BYD、摧毁欧洲汽车市场那样,为免费访问提供补贴,以换取数据捕获。Ross 希望建立一种类似冷战威慑的自动反制机制:“如果你补贴这个产业,我们就自动补贴对应产业……所以不要这么做。”Ross 直言,Xi 只关心权力维持,因此“理性讨论游戏规则,说白了不现实”。
- Ross 认为 China 自身也在焦虑:China 最大的优势是人口,但“如果一块 GPU 等同于劳动力中的一个贡献者……China 的优势会不会被削弱?”因此,他推动 Europe 的 500 million 人口加入竞争,并给 EU 开出具体处方:“今年年底前,你们应该有 100 个 Station F;明年年底前达到 1,000 个”——每个 Station F 聚集 3,000 名创业者,再由其他风险承担者围绕他们形成网络。
- 他最担心的是 AI 自动化网络战。Google 刚宣布首个由 LLM 发现的 zero-day;民族国家如今可以自动化扫描漏洞,而攻击又可以否认来源——“到底是 China?Russia?North Korea?还是某个友好国家伪装成其中之一?”这与核武器时代的相互确保摧毁不同,“我只是在黑你——而这可能会失控升级”。甚至声誉攻击也一样:诋毁一个公众人物“在某些方面可能比开枪杀死他更糟”,但你却可以逃脱惩罚。
9. 基础模型要转向,应用要靠工艺
- 对于那些为基础模型亏掉“数亿美元”而哀叹的 VC 朋友,Ross 的回答是:“有多少公司在不转向的情况下取得了巨大成功?很少。转向,接受现实——直接转向。”他以 Suno 创始人(可能是他)为例:“他从一开始就看到了——模型会商品化……模型是引擎。汽车是什么?”至于 Mistral 能否活下来,他只说:“每家公司都必须找到自己的方向……转向是可能的”,这是一个对冲式回答,而不是肯定判断。真正会输的是“只想沿着直线继续走的人”。
- Perplexity “处在幻觉——更准确说是虚构——率下降那一刻的完美位置”:届时,医疗诊断和法律工作都会打开。在那之前,这就像“还没有智能手机时试图创建 Uber”;但人们已经愿意为 Perplexity 付费,因此它可以一边等待海啸,一边乘着当前的浪。关于套壳应用与模型,他说:“所有人都说这些套壳应用没有价值;所有人也都说这些基础模型没有价值——价值到底在哪里?这正是令人兴奋的部分,我们正在发现它。”他的答案是工艺:“细节不是细节,细节就是一切。”
- 对于行业可能触及平台期的说法,Ross 认为自动驾驶的门槛高得多,因为机器对死亡事故的容忍度为零;诗歌和代码则不同。生成式 AI 时代会被快速推进,因为“我们就是智能手机”——我们知道这项技术会走向哪里,所以所有人都在提前配置资本。如果他是 Elon/xAI,他会更看好硬件下注,却不太会选择自建模型——“Elon 为什么要这么做?直接从地上捡一个不就行了。”
Everyone’s seen the news about DeepSeek today. Is it as big a deal as everyone is making of it?
1. Scraping OpenAI Models for Higher Quality Output
Yes, it is. It’s Sputnik 2.0. It’s true that they spent about $6 million, or whatever it was, on the training. They spent a lot more distilling or scraping the OpenAI model.
I can’t speak for Sam Altman or OpenAI, but if I were in that position, I would be gearing up to open-source my models in response, because it’s pretty clear you’re going to lose that. You might as well try to win all the users and the love from open-sourcing. Open always wins.
Always ready to go, Jonathan. I’m so excited for this. I’ve heard so many good things from so many different people, so thank you so much for doing this emergency show with me today.
No problem. Before we start, can I just say one thing? I think you have the most amazing, unique go-to-market strategy that I’ve ever seen in my life for a podcast. I’ve never seen this before.
I think your strategy is that you’re literally interviewing every single audience member, forcing them to watch videos and get addicted to you.
I thought you were going to say my accent, but I’m totally going to take that. That’s wonderful. Yes, you’re absolutely right: sometimes the biggest benefits of your business you don’t actually see until you do them at scale. It’s totally true.
2. Concerns About US Customer Data Going to China
I do want to start with DeepSeek. For a little bit of context, why are you so well placed to speak about DeepSeek? Let’s just start there.
My background is that I started the Google TPU, the AI chip that Google uses, in 2016. I started an AI chip startup called Groq, with a Q, not with a K, that builds AI accelerator chips, which we call LPUs.
3. Is DeepSeek News as Big a Deal as It Seems?
Fantastic. I wish everyone was as coherent as you in terms of their introductions.
Everyone’s seen the news about DeepSeek today. I want to start off by saying: is it as big a deal as everyone is making of it?
Yes, it’s Sputnik. It is Sputnik 2.0, and even more so. You know that story about how NASA spent $1 million designing a pen that could write in space, and the Russians brought a pencil? That just happened again. It’s a huge deal.
Why is it such a huge deal? Let’s unpack that.
Up until recently, the Chinese models had been behind Western models. I say “Western” including Mistral and some other companies. It was largely focused on how much compute you could get.
Most people don’t realize this, but most companies have access to roughly the same amount of data. They buy it from the same data providers, then churn through that data with a GPU, produce a model, and deploy it. They’ll have some of their own data, which will make them subtly better at one thing or another, but they’re largely all the same. The more GPUs, the better the model, because you can train on more tokens. That’s the scaling law.
This model was supposedly trained on a smaller number of GPUs and a much tighter budget. I think the way it’s been put is that it cost less than the salary of many of the executives at Meta.
That’s not true?
It’s actually an element of marketing involved in the DeepSeek release. It is true that they trained the model on approximately $6 million worth of GPUs. They claim that was GPU usage for, I think, 60 days, which, by the way, was also about the same amount of GPU time—4,000 GPUs for 30 days—as the original Llama 70B, I believe.
More recently, Meta has been training on more GPUs, but Meta hasn’t been using as much good data as DeepSeek, because DeepSeek was doing reinforcement learning using OpenAI’s model.
4. Distillation & DeepSeek's Use of OpenAI Data
What is distillation, just so I understand? Can you help me and the audience understand what distillation is in this regard, and how DeepSeek has been using distillation to get better-quality output through OpenAI data?
It’s a little bit like speaking to someone who’s smarter and getting tutored by someone who’s smarter. You actually do better than if you’re speaking to someone who’s not as knowledgeable about the area or is giving you wrong answers.
Before we get into any of this, I need to start with the scaling laws. These are like the physics of LLMs. There’s a particular curve, and the more tokens—which are sort of the syllables of an LLM, although they don’t match human syllables exactly—the more tokens that you train on, the better the model gets.
There are asymptotic returns where it starts trailing off. The thing about this scaling law that everyone forgets—and that’s why everyone was talking about how it’s the end of the scaling law because we’re out of data on the internet—is that it assumes the data quality is uniform. If the data quality is better, you can get away with training on fewer tokens.
Going back to my background, one of the fun things I got to witness, although I wasn’t directly involved, was AlphaGo, when Google beat the world champion Lee Sedol at Go. That model was trained on a bunch of existing games, but later they created a new one called AlphaGo Zero, which was trained on no existing games. It just played against itself.
How do you play against yourself and win?
You train a model on some terrible moves. It does okay, and then you have it play against itself. When it does better, you train on those better games, and then you keep leveling up like this. You get better data. The better your model is when it outputs something, the better the result and the better the data.
You train a model, use it to generate data, train a model, use it to generate data, and keep getting better and better. That lets you beat the scaling-law problem.
One quick hack for getting past all of that is, if there’s a really good model already right here, just have it generate the data and go straight up to where it is. That’s what they did.
It is true that they spent about $6 million, or whatever it was, on the training. They spent a lot more distilling or scraping the OpenAI model. They scraped the OpenAI model, got this higher-quality data from that and from refining it, and then got higher-quality output.
Correct?
Correct. All of that said, they did a lot of really innovative things. That’s what makes it so complicated. On the one hand, they kind of just scraped the OpenAI model. On the other hand, they came up with some unique reinforcement-learning techniques that were so simple and so impressive.
A lot of people wanted to say, “The Chinese copy and duplicate, as they always have done.” No, they came up with innovative stuff.
The best way to describe it is this: have you ever taken a test, gotten an answer right, and had your professor mark it wrong? Then you go back to the professor, argue with them, and everything, and it’s a pain, right?
If there’s only one answer, it’s a simple answer, and you say, “Write that answer in this box,” then there’s no arguing. You either get it right or you don’t.
What they did was, rather than having human beings check the output and say yes or no, they said, “Here’s the box. Output the answer here,” and then checked it. If it’s correct, we have the answer; if it’s not, we don’t. There’s no need to involve a human. It’s completely automated.
I read about reward-modeling stage, and that they innovated on this in a unique way. Did they not? Can you explain that area for me?
I’m not as familiar with that, so I’m probably not going to. You’ve been doing the research, so why don’t you tell me what you saw, and I can tell you if it tracks?
Essentially, they combined 2 different types of reward models to get higher, more accurate output. That was what I didn’t understand.
Yes, that’s not an area where I’ve dug too deeply into it.
Can OpenAI not just do distillation on DeepSeek’s model and then get better?
They don’t need to, because they’re actually still better. They’re a little bit better.
Could they buy the GPU usage, or is that questionable?
I don’t think you have to distill it because of the quality delta. However, why would they try to smuggle in GPUs when all they have to do is log into any cloud provider and rent GPUs?
This is the biggest gaping hole in the whole way export control is done. You can literally log into a cloud provider, swipe a credit card, and pay to use GPUs.
So are the export-control laws unnecessary, then?
They’re good, but the problem is that it’s like the Maginot Line: you just go around it. You need to seal it up a little more. There’s a little bit of room left to go here.
The other thing is, keep in mind that OpenAI was effectively subsidizing the training of this model accidentally, because DeepSeek was using OpenAI. Rumors are that OpenAI may not be completely profitable yet in terms of every token in the API—maybe on the subscriptions, but in the API. Each token they generated was effectively losing OpenAI a little bit of money while DeepSeek was getting training data.
OpenAI probably still has that data. In theory, they could just train on it.
George Krizan said in a tweet today that this would likely be a violation of U.S. export laws. Do you think that’s not true?
I’m not aware of where it would be an export issue. I do know that many people log into cloud providers and use them remotely.
One of the problems is that we actually block IP addresses from China, and I believe we might be unique in doing that. It’s also a little bit fruitless, because someone can just rent a server anywhere and log into us from there. There’s nothing we can check.
I don’t know that IP addresses are really the right way to do it. We need something more sophisticated.
You mentioned blocking IP addresses from China. There’s a lot of concern about U.S. customer data going back to China. Do you think that’s a legitimate and justified concern?
Yes. It’s probably the most significant concern. There are other concerns, but that’s probably the most significant.
People are so used to using these services that they might be shocked to hear this: when you use one of these other services and say “delete,” what they do is write “delete” next to your data. They don’t actually delete it. They just mark it as deleted. When you later come back and ask for your data, they give it to you with the word “delete” next to it. It’s still there.
These are well-meaning companies. Do you really think the CCP doesn’t have all your data and isn’t going to look it up later?
Some governments are more aggressive than others. If they have access to your data, it’s not even necessarily your data. It could be your next-door neighbor’s data. Your next-door neighbor might put something in there that accidentally gives information away and makes you more vulnerable.
Maybe they had a package delivered and put a complaint somewhere. You might not even do it yourself, but other people around you might. Think about the health data of a spouse.
5. DeepSeek and Its Potential Use by the CCP
Jonathan, I’m going to avoid the British indirectness: do you think DeepSeek is an instrument that will be used by the CCP to increase control on less democratic countries?
Yes, but I don’t think it’s DeepSeek that’s doing it.
You have to understand that any company operating in China and Hong Kong—the “one country, two systems” thing didn’t quite work out as anticipated, or maybe as anticipated but not as stated—has no choice.
When Groq started in 2016, we decided that we weren’t going to do business in China. This wasn’t a geopolitical decision; it was purely commercial.
We kept seeing companies like Google and Meta fail over and over again trying to win in China. The formula is actually pretty simple: you’re not allowed to make net money. You’re allowed to spend more money in China, but the moment you start to become profitable, or anywhere near profitable, all of a sudden there’s a thumb on the scale.
Companies that manufacture a lot in China and send more money to China can be successful there. They can sell things there. It’s a pretty simple formula: you must send more money to China than you take out.
At the same time, they require you to hand over all data. They also require that certain answers be in a form they find acceptable.
One of the more common things you see about DeepSeek right now is that, when you ask about Tiananmen Square, if the temperature is low on the model—and temperature is how creative it is; we don’t need to get into that—it will give you an answer that basically says, “I don’t want to talk about that. It’s a sensitive topic.”
You can ask it about other things that are sensitive topics elsewhere in the world, and it’ll just answer. But what happens if the CCP requires that they start to say, “What about TikTok? Should it be banned? Absolutely not. Here’s why,” and it gives you a cogent reason? That’s scary.
What do we do from here? I share your concerns completely. My challenge is that TikTok can be banned and shut off. It’s a closed-end product that we can ban tomorrow if we really want to. Here, it’s open source.
And worse. Until recently, we refused to run any Chinese models. We had to make a very difficult decision on DeepSeek. We now have it on our API at Groq.
Why did you decide to break the rule for DeepSeek?
When we saw DeepSeek become the No. 1 app on the App Store, the realization was that people were going to be putting their data in there. We want to make sure that there’s actually an option.
We store nothing. There’s no “delete” or anything like that. We store nothing. We don’t even have hard drives. We just have DRAM, and when the power goes off, everything goes away.
We wanted to make sure there was an alternative where, when you use DeepSeek’s model, your data isn’t going to the CCP.
Right now, the CCP is probably going to be taking the safeties off the weapons. They’re going to be asking, “Why are you making this model open source? Please direct your data toward us. Go win a bunch of customers this way.” But now they want the data.
They’re going to change the strategy. Remember, DeepSeek is a hedge fund. They’re doing this themselves, and they’re just influenced by the CCP. Now that the CCP has seen the success of this, it might see it as yet another TikTok.
They will see it as another TikTok. My question to you is: how long is it before the U.S. reacts to prevent this?
It should be. The first question to ask is whether we’re going to be talking about DeepSeek, or R1, for the next 6 months. The answer is absolutely not. We might be talking about R2, R3, and R4, but R1 was a one-shot.
The question is whether they’re going to keep coming up with interesting things, whether we’re going to play cat and mouse, and whether everyone is going to learn from this.
The biggest problem is that this has made it absolutely, nakedly clear that the models are commoditized. You’ve been asking the question: if there was any doubt before, that doubt is over.
What is the moat? I love Hamilton Helmer’s 7 Powers. It’s one of my favorites. I do it for every single investment we make. Every person at Groq has to fill it out.
Marketing is the art of decommoditizing your product, and the 7 Powers are 7 great ways to decommoditize your product: scale economies, network effects, brand, counter-positioning, cornered resource, switching costs, and process power.
The question is, who’s going to do what?
OpenAI—and you have to give Sam Altman and that team credit—has amazing brand power, like no one else in this space. That’s going to serve them for a really long time.
What you see Sam trying to do is scale. He’s trying to scale. That’s why we hear about Stargate and $500 billion. That’s the power he would like to have, but the power he has right now is brand. He’s trying to bridge that.
Doesn’t this news ridicule the $500 billion announcement, at a time when we’ve seen increasing efficiency on a scale like never before with DeepSeek today?
The $500 billion doesn’t ridicule it. Actually, I don’t think it’s enough spending.
We saw this happen at Google over and over again. We built the TPU, so why did we do it? The speech team trained a model that outperformed human beings at speech recognition. This was back in 2011 or 2012. It was the first time that had happened.
Jeff Dean, the most famous engineer at Google, gave a presentation to the leadership team. Slide No. 1: good news, machine learning finally works. Slide No. 2: bad news, we can’t afford it.
We’re Google, and we’re going to need to double or triple our global data-center footprint, probably at a cost of $20 billion to $40 billion, just to get speech recognition. Do you also want to do search and ads?
It turns out there’s always this giant “mission accomplished” banner every time someone trains a model. Then they start putting it into production, and they realize, “Oh, this is going to be expensive.”
This is why we’ve always focused on inference. Think about it this way: at Google, we always ended up spending 10 to 20 times as much on inference as on training.
Now the models are being given away for free. How much are we going to spend on inference? I guarantee it’s going to be enormous. With test-time compute, I’ve asked questions of DeepSeek where it took 18,000 intermediate tokens before giving me the answer.
I think Jensen Huang said that now half of NVIDIA’s revenue is from inference. What does that look like in the future?
I think it’s 95%. It just makes sense. You don’t train to become a cardiovascular surgeon and then do that for 95% of your life and perform for 5%. You train for a little while, and then you do it for the rest of your life.
Do you think the U.S. will put sanctions on DeepSeek to prevent the CCP from using it for data capture on U.S. citizens?
I don’t know what the solution is. There’s a carrot and there’s a stick. You can use a stick and block it. That might be effective, although I don’t know that the U.S. has really done that before. I’m not aware of a case, although it may be possible that it’s happened.
There’s also the carrot. It’s interesting how DeepSeek is being offered for free in China, and not just in China, but to anyone else. Others are doing that, too.
Is it possible that the CCP is underwriting that because it wants the data? They’re doing it with the car industry. The subsidization of Chinese cars by BYD is destroying the European car market.
Absolutely. The thing is, we have a lesson from the Cold War, which was mutually assured destruction.
The problem is that we do some sort of tariff, and then China does a tariff back. There needs to be some sort of automated response: if you do this, we will respond. If you subsidize this industry, we will automatically subsidize the equivalent industry. Make it automatic, so don’t do it, because there’s no benefit to you.
How does the fact that it’s open source change everything?
It’s the only reason people are using it. If it wasn’t open source, it wouldn’t have gotten the excitement. Open always wins.
Keep in mind that Linux won back when people didn’t trust open source. They thought it was less secure, that the features were worse, and that it was more buggy. It still won.
6. Is DeepSeek Diminishing OpenAI's Distribution Advantage?
Now people expect open source to be more secure, less buggy, and to have more features. How is proprietary software ever going to win?
Everyone always says that distribution is one of the major advantages that ChatGPT and, hence, OpenAI has, especially over the other providers. Every single day that DeepSeek is out and being used so pervasively, it diminishes the value of OpenAI’s distribution.
I agree, especially for pricing, because they’re losing their pricing power.
I can’t speak for Sam Altman or OpenAI, but if I were in that position, I would be gearing up to open-source my models in response. It’s pretty clear you’re going to lose that, so you might as well try to win all the users and the love from open-sourcing. Otherwise, you’re already at a point where you’re going to be using your other powers, like brand.
Would that be possible? Wouldn’t it cannibalize one core main line of revenue?
How would it cannibalize it? Remember, people like distribution.
How many people are going to buy something because they trust Dell? People trust Dell because Dell has earned its reputation over the course of decades. Supermicro builds interesting hardware, but look at what they’ve been going through recently. There are pros and cons. It’s cheaper, it’s trusted—you have to make a decision.
OpenAI has been around for a while. Most people think of them synonymously with AI. They could just switch to DeepSeek and people would still use them. That’s brand. It’s one of the 7 Powers.
If you were OpenAI, on day 1, would you switch to open and offer it for free?
I would. There’s probably more cleverness they could use. They could probably strike some deals before they do it, or whatever, but that would be the move I would make.
It would also be a position of strength. It would simply say, “Look, the only problem is the timing. If it happens right after DeepSeek, it looks like a response as opposed to an intentional thing.” I don’t know how you do that, but it is a response.
Why not just own that it’s a response?
Maybe that’s a good one. You just say, “Look, we had to respond. We’re better. Let’s see which model people choose.”
7. Perplexity in 3 Years
What do you think the internal discussion is within OpenAI today?
I would imagine it depends on where you are. If you’re senior, you’re going to have very different concerns than if you’re at the foot-soldier level.
At the foot-soldier level, you’re going to be worried: is my equity going to be worth anything? Is there any longevity here? How do I do my job? Am I going to have a job?
If you’re further up, it’s going to be more like: how do I keep everyone? How do I keep morale up? What is my response?
You’re going to have a lot of very difficult decisions in front of you. The No. 1 driver of bad decisions is fear. What they have to do is pick something, commit to it hard, and be brave about it.
So many different decisions work if you commit and align. It’s all about alignment.
How should we think about Meta? Meta shares the open-source values that DeepSeek espoused. Does this help or hurt Meta?
That’s a good question. One of the ways we’ve been looking at LLMs is a little bit like looking at an open-source software project, like Linux.
8. The $500BN Stargate Project
The thing is, Linux has switching costs, and I think what we’ve discovered is that LLMs have no switching costs whatsoever. That’s why the analogy to cloud doesn’t hold up at all. Everyone says, “There are going to be a couple of cloud vendors,” but you don’t switch your cloud very often.
Let’s map the 7 Powers to the top tech companies. I would say Microsoft’s biggest strength is switching costs. I love Microsoft as a company, but you go into a room full of people and ask, “Who uses Microsoft?” A bunch of hands go up. Then you ask, “Who likes using Microsoft?” The hands go down. It’s largely switching costs.
With Meta, it’s network effects. They could literally give every piece of technology away for free. I’m completely jealous of that, because if I had that right now, I would open-source everything. You don’t have to worry about it, and you get everyone helping you.
Meta is always in a position where open source is to its advantage because of the network effect. It almost doesn’t matter where it comes from.
I’m sure Meta would prefer to have the Linux of LLMs, but the more it goes open source, the more of an advantage Meta inherently has.
If you were Meta, would you do anything differently?
Meta is an amazing competitor. Normally, if this were something proprietary—a social mechanism, for example—they would try to replicate it and compete. They would say, “Come join us,” or not. I don’t think “come join us” works here.
The beautiful thing is that all the information for this model is available. Meta has already been doing this. It has way more compute. The question is whether it’s willing to scrape OpenAI like DeepSeek did. I don’t think it is.
Meta has been super careful about everything it’s been doing, and that’s the disadvantage.
I’m not being rude, but do you put morals aside to win? This is the AI arms race.
I think that’s going to happen. You cannot lose. What this has done is change the game.
Let’s talk about Europe for a minute. We almost forgot about Europe.
For me, watching everything, it feels like Europe lacks a willingness to take risk. There’s a black mark if you get it wrong. Everything is about downside protection, whereas in the U.S. it’s, “That was a great effort. You failed, but I’m going to fund you again.”
Then you look at China. China practices IP theft. It’s just part of the culture, and it’s not only against Western companies; it’s against each other, too.
The difference is that if you’re a Western company, the government steals from the Western company and provides it to Chinese companies, which is less fair. There are famous stories of turning on Huawei switches and seeing Cisco’s logo, with all the bugs.
Does the West have to adopt a more theft-on attitude?
I really hope not. For Europe to compete with the U.S., Europe has to adopt a more risk-on attitude. But adopting a more theft-on attitude is viscerally disgusting to me. I’m literally repulsed by the idea.
Are we not being idealistic? If you’re running in a race with someone who’s willing to take steroids, and you want to win, you’re going to have to take steroids, too. Then everyone is taking steroids, whereas if no one were taking them, everyone would be healthier and you’d have a real competition.
It’s a real problem. The question is whether governments can get involved.
I would love nothing more than to compete directly with Chinese companies on a fair footing. They have really smart people. DeepSeek has proven this. But when the government keeps putting its thumb on the scale, we’re going to try to avoid that competition wherever we can.
Now there’s no avoiding it, so maybe governments just have to get involved. I’m being blunt: Xi Jinping cares about 1 thing—power retention and growth. That’s the only thing that matters to him, and AI is central to that. He will do whatever it takes to win.
Having rational discourse about rules of play is, bluntly, unrealistic.
China has a lot of advantages, but the chief advantage is the number of people it has. Number of people is not sufficient, though. You also have India, and India has an advantage from the number of people. China has out-executed it. In fact, India was asking China for some time to help build out roads and infrastructure. China has really mastered that.
China has people, organization, discipline, and alignment. The concern with AI is: what if an LPU or GPU becomes the equivalent of a contributor to the workforce? You could literally add more to GDP by creating more chips and providing more power.
If that becomes the case, does China’s advantage erode? China is concerned that, in terms of workforce, the U.S. or the West could catch up. At the same time, China has a huge population advantage.
9. Advising the EU on Europe's Stance Today
This is why I want Europe to get into the fight on AI. If there are 500 million people who could be jumping into this.
If you were to advise the EU today on Europe’s AI response, what would you say?
Have you ever seen Station F?
Of course. I was there last week. We hosted there.
I would say that by the end of this year, Europe should have 100 Station Fs, and by the end of next year, it should have 1,000.
You’re basically collecting 3,000 people and surrounding them with other risk-taking entrepreneurs. They support each other, and they’re risk-on. When you surround yourself with other people who are risk-on, you’re going to be risk-on, and you’re going to take the entrepreneurial leap.
What does this space look like in 3 years’ time? How fearful should I be? I’m obviously a venture capitalist for a living, and all of my friends are saying, “Oh my God, we just lost hundreds of millions of dollars on these foundation-model companies.”
How many companies are you aware of that have become incredibly successful without pivoting?
Few. Most pivot.
Exactly. Pivot. Get over it.
Frankly, I’ve been talking to a lot of the LLM companies, and they have some good ideas. I really like the Suno founder. I think he saw it from the beginning: models are going to be commoditized, and that’s why he’s focused on the product.
He got it from the beginning. What is your product, not what is the model? The model is a piece of machinery. It’s an engine. What is the car? What is the experience?
What do you think Perplexity is in 3 years?
The question I used to get asked when we were raising money a little while ago was, “Is AI the next internet?” I said, “Absolutely not.”
The internet is an Information Age technology. It’s about duplicating data with high fidelity and distributing it. That’s what the telephone does, what the internet does, and what the printing press did. They’re all the same technology, just at a much different scale, speed, and capability.
Generative AI is different. It’s about coming up with something contextual, creative, and unique in the moment. The LLM is just the printing press of the generative age. It’s the start of it, and there are going to be all these other stages.
Imagine trying to start Uber before we had mobile. “Great, I’m going to book a trip over to here. How do I get home?” You couldn’t carry a desktop with you. You need to be at the right stage.
When I look at Perplexity, I see it as being perfectly positioned for the moment when the hallucination—or, really, confabulation—rate comes down. The moment these models get good enough that you don’t have to check the citations anymore, it will open up a whole set of industries.
All of a sudden, you’ll be able to do medical diagnosis from LLMs. You’ll be able to do legal work from LLMs. Until then, it’s like trying to create Uber before we had smartphones. It doesn’t make sense.
However, people are willing to use Perplexity today, even though you have to check the citations. It has an actual business that gets to ride the wave. The moment that tsunami of a lack of confabulation, or hallucination, comes along, Perplexity is perfectly positioned.
Does Mistral survive?
Each company has to find its own thing. I would look at Suno as a great example of how things are being done around the product as opposed to just the models.
Is it possible to pivot when you are OpenAI, Anthropic, or one of the very large providers? You’ve ingested billions of dollars. If disruption happens and you’re not able to pivot now, you’re not going to be able to pivot later when you get disrupted anyway.
10. Commoditization of Models & Big Tech's Stock Struggles
Wouldn’t one think that, with commoditization of models and cheaper inference, big tech actually wins? Have you seen the stock market today? NVIDIA and the others have been hit hard. How do you think about that?
What you see is a bunch of people who are concerned about training and the need for it, with everyone still thinking that most compute is training. They see someone training a model on 2,000 GPUs—the nerfed H800 version with slower memory, or whatever it is—and they say, “People aren’t going to need as many chips.”
But think about Jevons’s Paradox: the more you bring the cost down, the more people consume.
For the last 5 or 6 decades, like clockwork, once a decade, the cost of compute has gone down by a factor of 1,000. People buy 100,000 times as much compute while spending 100 times as much. Every decade, they spend 100 times as much.
You make it cheaper, and people want more. Every time one of these models gets cheaper, we see our developer count skyrocket. It goes up, comes back down a little bit, but the slope is higher than when it started.
Better models create more demand for inference. More demand for inference leads people to say, “I should train a better model,” and the cycle continues.
I just bought a whole lot of NVIDIA because the stock dropped 16%, on the thesis that increasing efficiency obviously means we won’t need as many NVIDIA chips. I thought exactly what you said: you’ll still need NVIDIA for inference, and you’ll just have much higher usage.
To me, it’s the most screaming buy of the century. Do you share my optimism on NVIDIA, given what you just said about Jevons’s Paradox?
Over the long term, I’d say the only thing I can say is what Warren Buffett and Charlie Munger said: in the short term, the market is a popularity contest; in the long term, it’s a weighing machine.
I can’t tell you about the popularity contest. But in terms of the weighing-machine part, there’s a misunderstanding. NVIDIA is actually more valuable thanks to DeepSeek, not less valuable.
Jevons’s Paradox was discovered by William Stanley Jevons and was recently made famous in Sacha’s tweet. However, I beat him to it by quite a bit.
Just as Sacha likes to say that he made Google dance, I’m going to say that I made Sacha dance. He might take exception to that, but less than 1 month before he posted that, I did a cute little tweet on it.
What was really happening in the 1860s was that Jevons wrote a treatise on steam engines, which I guess is what you did for fun back then in England. He realized that every time steam engines became more efficient, people would buy more coal. That’s the paradox.
But if you think about it from a business point of view, when the opex comes down, more activities come into the money. People do more things.
Every time we’ve seen the cost of tokens for a particular level of model quality come down, we’ve seen demand grow significantly. Price elasticity.
11. Nvidia's High Margins and the Strength of Their Moat
A lot of people suggest that NVIDIA’s incredible high-margin status—and I’m going to butcher this; I can’t remember what it was in the latest release, but it was 45% or whatever it was, and very, very high—relates to your margins as my opportunity.
Do you think your margin is my opportunity, or do you think that defensibility is that margin today?
There’s a wonderful business selling mainframes with a pretty juicy margin because no one seems to want to enter that business.
Training is a niche market with very high margins. When I say niche, it’s still going to be worth hundreds of billions of dollars a year long term. But inference is the larger market.
I don’t know that NVIDIA will ever see it this way, but I do think that those of us focusing on inference and building things specifically for it are probably the best thing that’s ever happened for NVIDIA stock, because we’ll take on the low-margin, high-volume inference so that NVIDIA can keep its margins nice and high.
Do you think the world sees this?
No. We raised some money in late 2024, and in that fundraise we still had to explain to people why inference was going to be a larger business than training.
Remember, this was our thesis when we started 8 years ago. I struggle to understand why people think training is going to be bigger. It just doesn’t make sense.
For anyone who doesn’t know, training is where you create the model, and inference is where you use the model. You want to become a heart surgeon, you spend years training, and then you spend more years practicing. Practicing is inference.
12. The Future of Efficiency After Nvidia's Success
I’m thrilled to hear you share your optimism around NVIDIA. Where does efficiency go from here? Everyone was shocked by how much more efficient R1 is and what we’ve seen from it. What’s next?
What you’re going to see is everyone else starting to use this mixture-of-experts approach.
Just so I understand, is that the segmentation of where information goes, so that it’s routed to the optimal part of the model?
Yes. It’s called MoE, which stands for mixture of experts.
When you use Llama 70B, you use every single parameter in that model. When you use Mixtral 8x7B, you use 2 of the roughly 8 experts, although there are some shared weights on top of that. It’s much smaller, and while it doesn’t correlate exactly, the number of parameters correlates very closely with how much compute you’re performing.
Take the R1 model. I believe it’s about 671 billion parameters, versus 70 billion for Llama. There’s also a 405 billion-parameter dense model, but let’s focus on 70 versus 671.
I believe there are roughly 250 experts, each of which is somewhere around 2 billion parameters. Then it picks a small number—maybe 8, 16, or 32 of them; I’m forgetting exactly which—and only needs to do the compute for those.
That means you get to skip most of it, sort of like your brain. Not every neuron in your brain fires when I say something to you about the stock market. The neurons about playing football don’t fire. That’s the intuition.
Previously, it was famously reported that GPT-4 had, I believe, something like 16 experts, and they got it down to 8. I forget the exact numbers, but it started off larger and they shrank it a little.
With the DeepSeek model, they’ve gone in the opposite direction. They’ve gone to a very large number of experts. The more parameters you have, it’s like having more neurons: it’s easier to retain the information that comes in.
By having more parameters, they’re able to get good results with a smaller amount of data. Because it’s sparse, because it’s a mixture of experts, they’re not doing as much computation.
Part of the cleverness was figuring out how they could have so many experts, how it could be so sparse, and how they could skip so many of the parameters.
If we take that as where we are—how DeepSeek became so efficient—what’s the next stage? All the experts can be routed so efficiently. What happens now?
Here’s a fun one: Meta recently released Llama 3.3 70B, and it outperformed its Llama 3.1 405B. Its new 70B outperformed its 405B.
What was surprising to me was that I thought they had retrained it from scratch. It turns out, when you read the paper, they just fine-tuned it. They used a relatively small amount of data to make it much better.
Again, this goes to the quality of the data. They had higher-quality data, took their old model, trained it, and made it much better. That new 70B outperforms their previous 405B.
What you’re going to see now is that everyone has seen the DeepSeek architecture and is going to say, “I have hundreds of thousands of GPUs. I’m now going to use a lot of them to create a lot of synthetic data, and then I’m going to train the hell out of this model.”
The other thing is that, while it’s sort of asymptotic, the question is where you stop on this curve. It depends on how many people you have doing inference.
You can either make the model bigger, which makes it more expensive and means you train it on less data, or you make it smaller and cheaper to run, but you have to train it more.
DeepSeek didn’t have a lot of users until recently, so it would never have made sense for them to train it a lot. They would much rather have a bigger model. Now what you’re going to see is all these other people either making smaller models or trying to make higher-quality models of the same size by training them more.
We’ve seen DeepSeek now say that only Chinese phone numbers can log in. That’s a new sign-up restriction. What has happened, and what’s the result?
They ran out of compute. This is another reason chip startups are going to do just fine. You train it once, but then you need inference compute.
You spend money to make the model, like designing a car, but then each car you build costs you money. Each query you serve requires hardware.
Training scales with the number of machine-learning researchers you have. Inference scales with the number of end users you have.
Do you think DeepSeek is truly astonished by the response it’s received from the global community, or did it know this would happen?
I think it marketed very well. You look at some of the publications, and they make it sound like it’s a philosophical thing. They talk about spending $6 million on the GPUs, and everyone zoomed in on that, neglecting the fact that Llama’s first model was trained on about $5 million worth of GPU time and set the world on fire in a good way.
They ignored the fact that DeepSeek spent a ton generating the data and doing all of this. They’re really good at marketing. I think they were probably surprised at how well it worked, but I think this is what they were going for.
Is there anything I haven’t asked, or that we haven’t spoken about, that we should?
Maybe ask what’s up with the $500 billion Stargate effort.
What’s up with the $500 billion Stargate effort? Do you buy those numbers?
I’ve gone back and forth on that. Gavin Baker tweeted some math, and before I saw that tweet, I came up with very similar math—spookily similar math.
However, talking to some people in the know, some of the comments are that they’ve got it. Then you keep pressing, and it’s like, maybe there’s some cutesiness to it.
What I think it is, is an acknowledgment that the models have been commoditized and infrastructure is what’s important in terms of maintaining elite scale. Scale is one of the 7 Powers.
What you’re seeing is an attempt to move from having a cornered resource or something like that into scale economies.
Do you think it will work?
I don’t think you get there in a short period of time with GPUs, because most of the compute is inference. If you’re talking about building out all the power and all the infrastructure, it’s going to take time. It’s infrastructure. It’s capex.
I think the real win here is brand. That’s what I would be doubling down on. I would hire the best brand firms I could and do a complete makeover.
Will OpenAI have a stronger or weaker brand in 3 years’ time?
Much stronger. I think they’re going to double down on that and focus on it.
Who will lose?
People who can’t adapt to disruption. Anyone who just wants to keep going in a straight line and do what they were doing before is going to lose.
The rate of disruption is probably going to increase. Think about it this way: going back to the analogy of LLMs being the printing press, imagine if there were a couple of smartphones left over from an ancient civilization.
All of a sudden, the printing press is invented, and you say, “Uber’s coming. I want to position for it.” You know where this is going. We are the smartphones. We know where generative AI technology goes.
Now everyone is saying, “We know how big this gets. Let’s put money into it. I can’t be the one who doesn’t spend money on this, because I know how big an advantage it’s going to be.” It’s like getting to add more workers to the workforce.
I think the generative age is going to be speed-run faster than whatever comes next, because we know what it looks like.
Is there any chance we see a plateau? We saw it in self-driving, where we went through this desert of slower progress and suddenly, all at once, it came. Will we see that, or will we just see this continue to accelerate?
With self-driving, the problem was that the threshold had to be way higher. If you look at the number of miles driven by these self-driving vehicles, it’s an enormous number, and the number of fatalities and incidents is lower per mile.
But we have no tolerance whatsoever for that when it’s a machine. When you’re writing poetry or code, it’s very different from doing surgery or driving a car.
How are you feeling, and do you feel better or worse post this?
I would probably feel both better and worse. I would feel better about my bet on building out more hardware. I would feel worse about trying to build out my own model. Why is Elon doing that? There’s plenty to choose from—just pick one up off the ground. Why are you making your own?
13. Excitement or Nerves in the AI Arms Race?
Are you excited when you look forward at the next few years, or are you quite nervous? You could say this is a time of heightened international warfare in terms of this new AI arms race: China stealing everything, the U.S. forced to steal back.
Long ago, I stopped having good days and bad days. It’s how many good things; it’s how many bad things. When you run an organization, I’m both excited and nervous, and I’m excited and nervous about different things at the same time.
The thing that I am most nervous about is that, unlike nuclear war, you can use AI tools to attack each other. Google just announced recently the first zero-day exploit found by an LLM that was previously unknown.
Yeah, that’s a scary one. So now, just for anyone who doesn’t know what zero-days are, how would you like me to have access to your phone?
Not ideal. How would you like the CCP to have access to your phone? Even less so. That’s a nation-state, and nation-states have a lot of resources. If they stand up a bunch of compute and start scanning for vulnerabilities in all the open source that’s out there—and not even the open source, just scanning ports on the internet and trying to figure out if they can break in—they can automate that now. They don’t need to hire people to do that.
Now the defense has to be automated because there’s no way to keep up with automated attackers. What happens if this gets out of control? But worse, it’s small enough—it’s not killing anyone—and it’s also deniable. That’s the hardest part about it, because is it really China? Is it Russia? Is it North Korea? Is it a friendly that’s making it seem like it’s one of them, or vice versa?
You go from where we had a Cold War, because having a war was unconscionable—it was unthinkable because of the consequences—to, “Yeah, I’m just hacking you,” and that could spiral out of control. I’m worried that we’re going to have more back-and-forth.
Think of it this way: If you are a nation-state and, let’s say, Harry, you’re a beacon to the venture community and you want to rally the European entrepreneurs to be risk-on, and I’m someone who doesn’t want that because I don’t want the competition, a country that doesn’t want that could sully your reputation. Maybe I make you persona non grata. How is that any worse than shooting someone? It could be worse in some ways, but you can get away with it. That has me nervous—really nervous.
But I’m also really excited, because we are seriously going to be able to innovate as fast as we can come up with ideas now. You’re not going to have to implement things; you’re going to be able to prompt-engineer your way through things. Just as we moved from hardware engineers to software engineers and sped up productivity, you’re now going to be able to have a prompt engineer who doesn’t even write software.
One of our engineers made this app where you can just describe what you want built, and it builds it. Because we’re so fast, you just iterate, and it’ll build an app for you—crazy things.
I just don’t understand where the value accrues then, because you mentioned that they created this tool, which allows you to prompt and build the app. I’m sure you’ve seen Bolt.new. I’m not sure if you’ve seen Lovable, where it’s basically ChatGPT but for website creation, in its bluntest terms. Is there value in that?
Everyone was like, “There’s no value in these wrapper apps.” Everyone’s like, “There’s no value in these foundation models.” Where the fuck is that value?
That’s part of the exciting part—it’s discovering that. I think people will always prefer to use the highest-quality, most polished product. I think there is an opportunity for artisanship and craftsmanship, and just perfecting it and getting to a certain number of nines in the details.
There’s the Eames quote: “The details aren’t the details. The details are the thing.” I used to be a little concerned with the quote, “If you’re not ashamed of the quality of your first release, you’ve waited too long,” because there’s a subtlety and nuance there. There’s soundness, and then there’s completeness.
What you want is an incomplete product—something that doesn’t do everything. That’s why you should be embarrassed. But it shouldn’t give you the blue screen of death. That’s not a good embarrassment. What you’re going to see now is that, because it’s so easy to come up with something that just kind of works, people are really going to value well-crafted, high-quality products.
Jonathan, I cannot thank you enough for breaking down so many different elements for me and putting up with my basic questions. You’ve been fantastic, and honestly, I so appreciate the short notice.
No problem. Good luck, and have fun out there. This is a brand-new age. It really is.