[BidClub_]
Sohn Conference Foundation · · 16 分钟

推理革命:Groq、Nvidia 与 AI 的未来

Jonathan RossJohn Yetimoglu

YouTube
TL;DR
  • 在 Irish Open 大会上,Groq 创始人兼 CEO、Google TPU 发明者——主持人介绍他时称其为 Nvidia 首席软件架构师,并说他“最近有点上新闻”——与对冲基金经理 John 坐下来交谈。 他的核心框架是:AI 不存在永久性瓶颈——“每当一个瓶颈变得足够严重,人们就会解决它”,所以零部件在成为问题时可以收取溢价,但问题不大时不行;而当它们变成更大的问题,人们就会开始解决它们。
  • 对今天最紧俏交易的直接推论是:内存的定价权存在自我约束。 内存曾是“半导体供应链中最商品化的环节”,这里讨论的不是奢侈品,而是吉芬商品:就像大米,价格可以一直涨到消费者转而吃玉米为止。DeepSeek V4 发布就是算法效率的一个例子,据称将很可能是 KV cache 的部分压缩了90%:“它就是那根长得过高的罂粟;一旦长得太高,就会被砍掉。”
  • 他不接受以机会成本为基础、用 Jevons 悖论为内存短缺辩护的说法。 解决内存问题的工程师意味着机会成本——“如果他们有足够的内存,就会去做别的事情。”他的处方非常直接:“开始建设更多内存晶圆厂。”
  • 两人长期以来的分歧,正是与投资论点最相关的地方:John 认为博士水平以上的智能存在边际递减。 他认为开源模型大约落后闭源权重模型实验室6个月,因此只要边际递减的判断成立,开放权重模型最终就会追上闭源权重模型。Jonathan 的反驳是:“对智能的需求不可能被满足”——癌症尚未被攻克、人类仍会因衰老死亡、运行某些 AI 模型的算力仍然不足;再加上人与人之间的竞争,即使模型差异细微到人类无法察觉,也会“体现在我的回报里”。
  • Agentic AI 会进一步巩固最聪明模型的护城河。 “AI 喜欢使用 AI”,即使人类分辨不出差异,AI 也会识别并调用更聪明的模型。一个支持性案例是:用 LLM 筛选简历的招聘人员更偏好由自己所用模型写出的简历——所以“用 Claude Opus 47 写一份,再用 ChatGPT 写一份”。
  • 他对能力不会进入平台期的判断,为持续扩产提供了支撑。 如今模型会用自己的数据生成数据、筛选数据并重新训练,“以相当线性的速度持续改进”;AI 的直觉已经超过人类。他表示,Waymo 每天产生的数据量开始接近一个人一生的驾驶数据,但不知道是否已经达到这一规模。还有一个值得保留的定义:智能是一种静态能力,即作出预测或影响结果的能力;感知性则是“智能自我提升的速度”——它是“文明的属性”,也是社会中不断加速的反馈回路。
摘要 · 为研究而整理的核心内容

1. 两场登月式豪赌:AI 技术栈大于任何单一瓶颈

  • Jonathan 开场说,人类大约60年前把人送上了月球,而 AI“即便放到今天仍是前沿技术”——算法早在1970年代就已存在,只是终于拥有了足够的算力。投资者总盯着某一层,但要让 AI 真正运转起来,需要“芯片、封装、系统、网络、数据中心……电力,以及所有这些东西”。
  • 在此之上,训练和推理是“两个完全不同的问题”——就像抵达月球和在月球上着陆的区别。John 开场时的框架是:跨架构的可移植性正在下降,移植“不再只是工程上的不便”,而是一个经济劣势。

2. 瓶颈一旦过大就会被解决:内存就是那根长得过高的罂粟

  • 核心规律是:“每当一个瓶颈变得足够严重,人们就会解决它。”有限但尚可容忍的组件可以收取高价;一旦问题大到无法忽视,人们就会开始解决它。
  • 内存是当前 AI 最大的瓶颈,过去却是“半导体供应链中最商品化的环节”。这里用的是吉芬商品而非凡勃伦商品的类比:就像大米,涨价后消费者反而会在大米上花更多钱,直到“有人说,别再吃大米了,改吃玉米”。
  • John 提到 DeepSeek V4 发布,称其将很可能是 KV cache 的部分压缩了90%,并引出了围绕这一点的 Jevons 悖论论证。Jonathan 的回答着眼于机会成本:“如果他们有足够的内存,就会去做别的事情……它就是那根长得过高的罂粟,一旦长得太高,就会被砍掉。”他的解决方案是:“开始建设更多内存晶圆厂。”

3. 长期分歧:智能的回报是否会递减?

  • John 的观点是,超过博士水平后,人类“实际上无法真正理解模型之间的差异”;而如果开源模型大约落后闭源权重模型实验室6个月,那么在智能回报递减的前提下,开放权重模型最终会追上闭源权重模型。
  • Jonathan 的反驳是——“我可能会让你有点意外”:“对智能的需求不可能被满足。”只要癌症尚未被治愈、人类仍会因衰老死亡、运行某些 AI 模型的算力仍然不足,智能就不够用;竞争同样不会停止——“在座各位大概都已经可以退休了……但你们仍在投资”——因此,即使模型在人类看来无法区分,它们仍会在你的回报中拉开差距。
  • Agentic AI 进一步强化了这一点:“AI 喜欢使用 AI”,即使人类无法分辨,AI 也会识别并调用更聪明的模型。简历研究的说法是,LLM 更偏好由自己生成的简历,而招聘人员现在也用 LLM 进行筛选——所以要“用 Claude Opus 47 写一份简历,再用 ChatGPT 写一份”。

4. 能力为何不会饱和:直觉与合成数据飞轮

  • 借用 Kahneman 关于快思考与慢思考的框架,AI 属于直觉型系统,而且由于数据规模更大,“它在直觉方面实际上比我们更强”。他说,Waymo 每天产生的数据量开始接近一个人一生的驾驶数据,但不知道是否已经达到这一规模——你只要看到工字钢从货运卡车上掉下来“2到3次”,就会准确知道该怎么做。
  • 训练如今已经形成自举循环:模型生成数据,因为“能够判断什么是好的”而筛掉数据,再用筛选后的数据重新训练,能力不断上移;其改进速度“相当线性”。因此,“没有理由相信它们不会变得更聪明”,即使人类感知不到差距。

5. 感知性是文明的属性,而不是模型的属性

  • 他提出了一套自己的定义:智能是一种静态能力,即作出预测或影响结果的能力;“感知性是智能自我提升的速度”,存在于一个连续光谱上,而不是非黑即白。世界顶尖围棋手的能力最终趋于平台,是因为没有更强的棋手可以与之对弈。
  • 语言让信息得以在文明中传递:“智能是有机体或个人的属性,感知性是文明的属性。”AI 正在参与社会中不断加速的反馈回路——“我预计我们的孩子会比我们曾经聪明得多。”

John

Thank you to the Irish Open conference for having me back again this year. This is really a wonderful conference for a great cause, and I'm honored to be invited back with Jonathan. Jonathan, who you guys might have seen, has been in the news a little bit lately. He's the chief software architect of Nvidia, founder and CEO of Groq, and inventor of Google's TPU. He's had 4 first-time-right silicon designs. He's one of my dearest friends and someone I'd consider to be one of the greatest engineers of our time.

Jonathan

Well, thank you. For those of you who don't know John, although probably most of you do, John's one of those rare hedge fund managers who both gets great returns and provides immense entertainment value. Watching you hold those shorts like you've got diamond hands—I don't know how you do it. Everyone else is sweating for you.

John

Not everyone's supposed to know I'm a crazy short seller, so—

Jonathan

Oops.

John

Yeah. I can't believe you just doxxed me like that. Jonathan called everything that's happening today over 10 years ago. Everyone thought he was crazy. I was one of the lucky few who believed him back then. Up until a few years ago, nobody even knew what inference was, and so I think that's a good place to start, given it's now probably the most important thing about AI today.

1. Inference Economics Take Center Stage

Jonathan

For once, I don't have to explain what inference is. That's nice. Do you want me to get into the economics, or what do you want?

John

Yeah. One thing I think is very important is that portability across architectures is diminishing, and porting is no longer really an engineering inconvenience but has evolved into more of an economic handicap. Maybe we can talk through some of the tokenomics of inference.

2. Bottlenecks Keep Moving

Jonathan

A good parallel here is that we landed people on the moon about 60 years ago, roughly, and that was 60-year-old technology that got us there. What we're doing in AI is brand-new technology. It's cutting-edge even for today. We finally have enough compute to do these things with models. We've had the algorithm since the 1970s.

What I'm seeing a lot is that people will look at 1 small part of the stack and think that it's the most crucial part, that it's the bottleneck. I don't think people understand just how much you have to build to make AI work. It's the chips, the packaging, the systems, the networking, the data centers, the servers, the racks, the power—all the stuff. On top of that, it's both training and inference, which are 2 completely different problems. It's a little bit like getting to the moon and then landing on the moon: 2 very separate problems.

In terms of bottlenecks and all this, the focus I see a lot in the questions investors are asking is, "What is the bottleneck? What's the next thing that's going to be constrained? What should I do next?" That's not how this industry works. Every time a bottleneck gets big enough, people solve it. As someone running a business, you're always going, "Hey, what is my biggest problem, and can I solve it?"

When you look at some of the components that are limited, when they're a problem but not a huge problem, people can charge a lot of money for them. But as they start to become a bigger problem, people start solving that problem. It's this constant-shifting dynamic of where in the supply chain the biggest problem is. Just don't become too big of a problem, because then it'll get solved.

3. The Memory Supply Problem

John

Yeah. Memory is the big theme today, right? It's the biggest bottleneck in AI. But what you're saying is that if memory continues to become a larger and larger bottleneck and remains extremely supply-constrained, you think it's going to be solved?

Jonathan

Yeah, exactly. Memory used to be a commodity. It was the most commoditized segment of the semiconductor supply chain.

John

Maybe a good way to explain this is that there are 2 kinds of goods you can charge a lot for. One is a Veblen good, and the other is a Giffen good. Veblen goods become more desirable the more you charge for them, like luxury goods. Giffen goods are a little bit different. An example is rice. If you start to charge more for rice, some people can no longer afford steak, so they actually spend more money on rice. By raising the price—not for a luxury good, but for a staple—you can actually increase its value.

That happens to some extent until someone says, "Let's stop eating rice and eat corn or something else," right? There's an inverse headwind: if memory is too expensive and people don't build enough of it, they're going to solve that problem technologically. So do you think algorithmic efficiencies—for example, Deep Seek's latest model release, the V4 release, compressed what was likely KV cache by 90%—will solve this? The argument around that has been, "Jevons paradox, Jevons paradox, Jevons paradox." What do you think?

Jonathan

Those engineers working on that problem represent an opportunity cost. They could have been working on something else. If they had enough memory, they would have worked on something else. If it wasn't so expensive, they would have worked on something else. If you make it a big enough problem, it's the tall poppy. As soon as it gets too tall, it gets chopped down.

John

Got it. Yeah, cool. So let's—

Jonathan

Start building more memory fabs. That's the solution.

4. Smarter Models Keep Winning

John

Yeah. So what do you think? We talk about diminishing returns on intelligence a lot, right? The way we prepared for this presentation is we literally just went through our conversations and thought, "That was a good topic. That was a good topic," and then we picked—

Jonathan

A long-standing disagreement between John and me.

John

Yeah. I think intelligence has diminishing returns, and at a certain point the models get so smart—above PhD level—that human beings can't really understand the differences between 1 or the other. Couple that with the closed-source versus open-source frontier landscape, where open-source models are roughly 6 months behind the closed-weight model labs, and you can make the assumption, if you buy my argument that there are diminishing returns over time, that the open-weight landscape will catch up to the closed-weight landscape. There are a lot of nuances around that, so I want to hear your thoughts. Share your thoughts with us.

Jonathan

Well, I might surprise you a little. There are a lot of areas in the economy where, if you produce more of something, it becomes less desirable. Then there are others where it becomes more desirable. I would argue that with intelligence, there's no way to satiate the appetite for intelligence. The more intelligence you get, the more intelligence you want.

Let me break it down. First of all, is intelligence going to plateau? That's important for decisions. Second of all, if it didn't plateau, would we get to enough of it that we'd say we don't need any more? In terms of having enough intelligence, the first thing is that as long as cancer isn't cured, as long as people still die of old age, and as long as we don't have enough compute to run some of these AI models, we don't have enough intelligence. There's an economic incentive to keep building smarter and smarter machines to help us solve bigger and bigger problems.

The second is competition. Everyone in this room probably could retire. You probably don't need to make more money, but you keep investing and competing with each other. Why? It's competition. We just do it. We're human, and competition isn't going away. If my AI is less intelligent than your AI, and I can't tell the difference directly, I'm still going to be able to tell the difference in my returns based on which one I'm using. I'm going to want the better AI.

5. AI Starts Using AI

And so, whenever I program, I actually use—

John

It's a gift to the Earth that Jonathan is programming and writing code again. Everyone should say thank you to him.

Jonathan

But I'm actually writing code that's being used for real things now, again thanks to AI, right? I'll use multiple models. Each one is better at different things. Even though 1 model might be better than another, I'll still have a use for that model. I think AI will know that this other AI is smarter, even if we can't tell the difference, and will use that smarter AI. This is the whole agentic thing.

John

So what is agentic?

Jonathan

Agentic is this: do you get better productivity by using AI? Yes. Well, so does AI. AI likes to use AI. It calls on AI to do some task for it and return the result, just like you do. It's just more of that. The AI is going to recognize smarter AI and use that smarter AI. There was actually a conversation we had outside with 1 of the hosts. I don't know if you remember, about resumes.

John

Oh, yeah. Funny.

Jonathan

Yeah, so I didn't know this. This is something I just learned. Someone did a study and showed that resumes generated from 1 LLM are preferred by that same LLM over resumes from another. Recruiters are now using LLMs to determine who to interview. But you've got to figure out which LLM the recruiter is using. So you should build 1 resume with Claude Code or Claude Opus 47 and 1 with ChatGPT, and you'll have the highest probability of being selected, basically.

John

So then the other question is: Is intelligence going to saturate, or are we just going to need more and more intelligence, and this build-out is going to make total sense?

Jonathan

My argument for why it's not going to saturate is as follows. There are 2 components to intelligence, and there's a great, easily digestible book called Thinking, Fast and Slow that many of you have read by Daniel Kahneman that explains exactly what AI does. Thinking fast is the intuitive part. You're given a problem: Do you have an answer immediately? Thinking slow is when you iterate on it.

If you think about chess, speed chess is thinking fast, and regular chess is thinking slow. You're evaluating multiple opportunities and recognizing better moves when you string moves together, right? Even though AI is coming from computers, and so this is a little bit hard for us to see, AI is very intuitive. It's actually better at being intuitive than we are, and that's because it's been trained on so much data.

When Waymo is sending all their cars out, the amount of data that they get in a day is about—I don't know if it's now at the amount of experience that a human being gets in a lifetime of driving, but it's starting to approach that, at least. When you're getting a lifetime of driving data in a day, you're seeing every possibility. You don't need to figure out how to deal with the fact that some I-beam is falling off the back of a freight truck and about to hit you. You've seen it happen 2 or 3 times, and you know exactly what to do.

The thing is, the more these models produce data, the more they're able to intuitively deal with the situation because they've already seen it, right? And that's intuition: when you just have the answer. When you train these models, they used to be trained on just data pulled from the world that human beings were producing. Now what we do is use the models to generate the data that they get trained on.

You have a model at this level of capability, and it produces data at these levels of capability. It's gotten good enough that it can tell what good is. It keeps this data, trains, moves up to here. Then it produces data of this quality, prunes it to here, trains, and goes up to here, and just keeps moving up.

At this point, we're seeing these models improve at a pretty linear rate. There's no reason to believe that they're not going to get smarter. We may not recognize the difference between 2 really smart models, but 1 will be much smarter than the other. And that matters in the context of competition—competition and solving big unsolved problems.

John

Yeah, right. Do you want to talk about sentience?

Jonathan

Oh, jeez.

[Laughter]

6. Sentience Gets A Definition

Okay, I have a hobby. My hobby is to take words that people have used for centuries that don't have a good, concrete meaning and try to ascribe a definition to them that helps me understand the world. I'm going to give you my definition of sentience.

First of all, intelligence is your ability to make a prediction or influence an outcome to what you want to have happen. But it's stationary. It's like you have an amount of intelligence that is fixed in these models. But sentience is your rate of improvement in your intelligence.

That makes sense because when we talk about sentience, we're talking about the ability to self-reflect and get better, and that's an important part. So, I say that intelligence is your capability and sentience is your rate of change. But rate of change doesn't have to be binary. You're not sentient or not sentient. It's: How sentient are you? Are you linearly sentient? Are you asymptotically sentient?

You look at the world's best Go player, and he asymptoted. He stopped getting better because he didn't have better players to play against. People often conflate LLMs with just being intelligent, but there's something else that language gives you beyond intelligence. It gives you the ability to transfer information.

Going back to the example I gave with Waymo, if you have an entire civilization producing information, each participant in that civilization gets to benefit from that distillation of knowledge. While intelligence is a property of an organism or an individual, sentience is a property of a civilization.

AI is producing more intelligence. It's getting smarter. You interact with it, you get smarter. You ask better questions. You're making the AI smarter. And so there's this feedback loop of sentience that's accelerating in our society, and AI is contributing to that. As that happens, I would expect our kids to get much smarter than we ever were, just like we're probably smarter than our parents were because we had the internet.