[BidClub_]
Sohn Conference Foundation · · 16 min

The Inference Revolution: Groq, Nvidia and the Future of AI

Jonathan RossJohn Yetimoglu

YouTube
TL;DR
  • At the Irish Open conference, the Groq founder/CEO and Google TPU inventor — introduced as Nvidia's chief software architect, having "been in the news a little bit lately" — sat down with hedge-fund manager John. His core framework: there is no permanent bottleneck in AI — "every time a bottleneck gets big enough, people solve it," so components can charge a premium while they're a problem but not a huge one; as they become bigger problems, people start solving them.
  • The direct application to today's tightest trade: memory's pricing power is self-limiting. Memory was "the most commoditized segment of the semiconductor supply chain" and is discussed through a Giffen-good analogy rather than a luxury good — you can raise the price of rice only until people switch to corn. Deep Seek's V4 release was cited as an example of algorithmic efficiency, reportedly compressing what was likely KV cache by 90%: "it's the tall poppy. As soon as it gets too tall, it gets chopped down."
  • He rejects the Jevons-paradox defense of memory scarcity on opportunity-cost grounds: engineers working on that problem represented an opportunity cost — "if they had enough memory, they would have worked on something else." His prescription is blunt: "start building more memory fabs."
  • The pair's long-standing disagreement is the thesis-relevant one: John argues intelligence has diminishing returns above PhD level and that open-source models are roughly 6 months behind closed-weight model labs, so the open-weight landscape catches up to the closed-weight landscape if his diminishing-returns argument holds. Jonathan's rebuttal: "there's no way to satiate the appetite for intelligence" — unsolved problems (cancer, aging), insufficient compute, and human competition mean even imperceptible model gaps show up "in my returns."
  • Agentic AI hardens the moat for the smartest model: "AI likes to use AI" and will recognize and use the smarter model even when humans can't tell the difference. Supporting anecdote: LLM-screening recruiters prefer resumes written by their own model — so write "one resume with Claude Opus 47 and one with ChatGPT."
  • His case against a capability plateau underwrites the build-out: models now generate, prune, and retrain on their own data, "improving at a pretty linear rate," and AI's intuition already exceeds ours — Waymo's daily data, he said, was starting to approach a human lifetime of driving data, though he didn't know whether it had reached that amount. Bonus definition worth keeping: intelligence is a stationary capability—the ability to make a prediction or influence an outcome; sentience is "your rate of improvement in your intelligence" — "a property of a civilization," and an accelerating feedback loop in society.
Digest · the substance, structured for research

1. Two moonshots in one: the stack is bigger than any single bottleneck

  • Jonathan's opener: We landed people on the moon about 60 years ago, while AI is "cutting edge even for today" — the algorithm has existed since the 1970s, and there is finally enough compute. Investors fixate on one layer, but making AI work takes "the chips, the packaging, the systems, the networking, the data centers... power, all the stuff."
  • On top of that, training and inference are "two completely different problems" — like getting to the moon versus landing on it. John's framing to open: portability across architectures is diminishing, and porting is "no longer an engineering inconvenience" but an economic handicap.

2. Bottlenecks that get too big get solved — memory is the tall poppy

  • The core dynamic: "Every time a bottleneck gets big enough, people solve it." A limited-but-tolerable component can charge a lot; become too big a problem, and people start solving it.
  • Memory — today's biggest AI bottleneck, formerly "the most commoditized segment of the semiconductor supply chain" — is discussed through a Giffen-good analogy rather than a Veblen good: like rice, raising the price makes buyers spend more on it, until "someone goes, let's stop eating rice and let's eat corn."
  • John raised Deep Seek's V4 release, saying it compressed what was likely KV cache by 90%, and the Jevons-paradox argument around that. Jonathan's answer is opportunity cost: "If they had enough memory, they would have worked on something else... It's the tall poppy. As soon as it gets too tall, it gets chopped down." His fix: "start building more memory fabs."

3. The long-standing disagreement: do returns to intelligence diminish?

  • John's side: past PhD level, humans "can't really understand the differences" between models — and with open-source models roughly 6 months behind closed-weight model labs, the open-weight landscape catches up to the closed-weight landscape if returns diminish.
  • Jonathan's rebuttal — "I might surprise you a little": "there's no way to satiate the appetite for intelligence." As long as cancer isn't cured, people die of old age, and there isn't enough compute to run some AI models, there isn't enough intelligence; and competition — "everyone in this room probably could retire... but you keep investing" — means indistinguishable models still separate in your returns.
  • The agentic kicker: "AI likes to use AI," and it will recognize and use the smarter model even when humans can't. The resume study as told: LLMs prefer resumes generated by themselves, and recruiters now screen with LLMs — so build "one resume with Claude Opus 47 and one with ChatGPT."

4. Why capability won't saturate: intuition plus the synthetic-data flywheel

  • Via Kahneman's thinking fast/slow: AI is the intuitive kind, and "it's actually better at being intuitive than we are" because of data scale. Waymo's daily data, he said, was starting to approach a human lifetime of driving data, though he didn't know whether it had reached that amount — you've seen the I-beam falling off the freight truck "two to three times and you know exactly what to do."
  • Training now bootstraps: models generate data, prune it because they "can tell what good is," retrain, and move up — improving "at a pretty linear rate," so "there's no reason to believe that they're not going to get any smarter," even if we can't perceive the gap.

5. Sentience is a property of a civilization, not a model

  • His definitional hobby: intelligence is your stationary ability to make a prediction or influence an outcome; sentience is "your rate of improvement in your intelligence" — a spectrum, not a binary. The world's best Go player asymptoted "because he didn't have better players to play against."
  • Language transfers information across civilization: "intelligence is a property of an organism or an individual, sentience is a property of a civilization." AI is contributing to that accelerating feedback loop in society — "I would expect our kids to get much smarter than we ever were."

John

Thank you to the Irish Open conference for having me back again this year. This is really a wonderful conference for a great cause, and I'm honored to be invited back with Jonathan. Jonathan, who you guys might have seen, has been in the news a little bit lately. He's the chief software architect of Nvidia, founder and CEO of Groq, and inventor of Google's TPU. He's had 4 first-time-right silicon designs. He's one of my dearest friends and someone I'd consider to be one of the greatest engineers of our time.

Jonathan

Well, thank you. For those of you who don't know John, although probably most of you do, John's one of those rare hedge fund managers who both gets great returns and provides immense entertainment value. Watching you hold those shorts like you've got diamond hands—I don't know how you do it. Everyone else is sweating for you.

John

Not everyone's supposed to know I'm a crazy short seller, so—

Jonathan

Oops.

John

Yeah. I can't believe you just doxxed me like that. Jonathan called everything that's happening today over 10 years ago. Everyone thought he was crazy. I was one of the lucky few who believed him back then. Up until a few years ago, nobody even knew what inference was, and so I think that's a good place to start, given it's now probably the most important thing about AI today.

1. Inference Economics Take Center Stage

Jonathan

For once, I don't have to explain what inference is. That's nice. Do you want me to get into the economics, or what do you want?

John

Yeah. One thing I think is very important is that portability across architectures is diminishing, and porting is no longer really an engineering inconvenience but has evolved into more of an economic handicap. Maybe we can talk through some of the tokenomics of inference.

2. Bottlenecks Keep Moving

Jonathan

A good parallel here is that we landed people on the moon about 60 years ago, roughly, and that was 60-year-old technology that got us there. What we're doing in AI is brand-new technology. It's cutting-edge even for today. We finally have enough compute to do these things with models. We've had the algorithm since the 1970s.

What I'm seeing a lot is that people will look at 1 small part of the stack and think that it's the most crucial part, that it's the bottleneck. I don't think people understand just how much you have to build to make AI work. It's the chips, the packaging, the systems, the networking, the data centers, the servers, the racks, the power—all the stuff. On top of that, it's both training and inference, which are 2 completely different problems. It's a little bit like getting to the moon and then landing on the moon: 2 very separate problems.

In terms of bottlenecks and all this, the focus I see a lot in the questions investors are asking is, "What is the bottleneck? What's the next thing that's going to be constrained? What should I do next?" That's not how this industry works. Every time a bottleneck gets big enough, people solve it. As someone running a business, you're always going, "Hey, what is my biggest problem, and can I solve it?"

When you look at some of the components that are limited, when they're a problem but not a huge problem, people can charge a lot of money for them. But as they start to become a bigger problem, people start solving that problem. It's this constant-shifting dynamic of where in the supply chain the biggest problem is. Just don't become too big of a problem, because then it'll get solved.

3. The Memory Supply Problem

John

Yeah. Memory is the big theme today, right? It's the biggest bottleneck in AI. But what you're saying is that if memory continues to become a larger and larger bottleneck and remains extremely supply-constrained, you think it's going to be solved?

Jonathan

Yeah, exactly. Memory used to be a commodity. It was the most commoditized segment of the semiconductor supply chain.

John

Maybe a good way to explain this is that there are 2 kinds of goods you can charge a lot for. One is a Veblen good, and the other is a Giffen good. Veblen goods become more desirable the more you charge for them, like luxury goods. Giffen goods are a little bit different. An example is rice. If you start to charge more for rice, some people can no longer afford steak, so they actually spend more money on rice. By raising the price—not for a luxury good, but for a staple—you can actually increase its value.

That happens to some extent until someone says, "Let's stop eating rice and eat corn or something else," right? There's an inverse headwind: if memory is too expensive and people don't build enough of it, they're going to solve that problem technologically. So do you think algorithmic efficiencies—for example, Deep Seek's latest model release, the V4 release, compressed what was likely KV cache by 90%—will solve this? The argument around that has been, "Jevons paradox, Jevons paradox, Jevons paradox." What do you think?

Jonathan

Those engineers working on that problem represent an opportunity cost. They could have been working on something else. If they had enough memory, they would have worked on something else. If it wasn't so expensive, they would have worked on something else. If you make it a big enough problem, it's the tall poppy. As soon as it gets too tall, it gets chopped down.

John

Got it. Yeah, cool. So let's—

Jonathan

Start building more memory fabs. That's the solution.

4. Smarter Models Keep Winning

John

Yeah. So what do you think? We talk about diminishing returns on intelligence a lot, right? The way we prepared for this presentation is we literally just went through our conversations and thought, "That was a good topic. That was a good topic," and then we picked—

Jonathan

A long-standing disagreement between John and me.

John

Yeah. I think intelligence has diminishing returns, and at a certain point the models get so smart—above PhD level—that human beings can't really understand the differences between 1 or the other. Couple that with the closed-source versus open-source frontier landscape, where open-source models are roughly 6 months behind the closed-weight model labs, and you can make the assumption, if you buy my argument that there are diminishing returns over time, that the open-weight landscape will catch up to the closed-weight landscape. There are a lot of nuances around that, so I want to hear your thoughts. Share your thoughts with us.

Jonathan

Well, I might surprise you a little. There are a lot of areas in the economy where, if you produce more of something, it becomes less desirable. Then there are others where it becomes more desirable. I would argue that with intelligence, there's no way to satiate the appetite for intelligence. The more intelligence you get, the more intelligence you want.

Let me break it down. First of all, is intelligence going to plateau? That's important for decisions. Second of all, if it didn't plateau, would we get to enough of it that we'd say we don't need any more? In terms of having enough intelligence, the first thing is that as long as cancer isn't cured, as long as people still die of old age, and as long as we don't have enough compute to run some of these AI models, we don't have enough intelligence. There's an economic incentive to keep building smarter and smarter machines to help us solve bigger and bigger problems.

The second is competition. Everyone in this room probably could retire. You probably don't need to make more money, but you keep investing and competing with each other. Why? It's competition. We just do it. We're human, and competition isn't going away. If my AI is less intelligent than your AI, and I can't tell the difference directly, I'm still going to be able to tell the difference in my returns based on which one I'm using. I'm going to want the better AI.

5. AI Starts Using AI

And so, whenever I program, I actually use—

John

It's a gift to the Earth that Jonathan is programming and writing code again. Everyone should say thank you to him.

Jonathan

But I'm actually writing code that's being used for real things now, again thanks to AI, right? I'll use multiple models. Each one is better at different things. Even though 1 model might be better than another, I'll still have a use for that model. I think AI will know that this other AI is smarter, even if we can't tell the difference, and will use that smarter AI. This is the whole agentic thing.

John

So what is agentic?

Jonathan

Agentic is this: do you get better productivity by using AI? Yes. Well, so does AI. AI likes to use AI. It calls on AI to do some task for it and return the result, just like you do. It's just more of that. The AI is going to recognize smarter AI and use that smarter AI. There was actually a conversation we had outside with 1 of the hosts. I don't know if you remember, about resumes.

John

Oh, yeah. Funny.

Jonathan

Yeah, so I didn't know this. This is something I just learned. Someone did a study and showed that resumes generated from 1 LLM are preferred by that same LLM over resumes from another. Recruiters are now using LLMs to determine who to interview. But you've got to figure out which LLM the recruiter is using. So you should build 1 resume with Claude Code or Claude Opus 47 and 1 with ChatGPT, and you'll have the highest probability of being selected, basically.

John

So then the other question is: Is intelligence going to saturate, or are we just going to need more and more intelligence, and this build-out is going to make total sense?

Jonathan

My argument for why it's not going to saturate is as follows. There are 2 components to intelligence, and there's a great, easily digestible book called Thinking, Fast and Slow that many of you have read by Daniel Kahneman that explains exactly what AI does. Thinking fast is the intuitive part. You're given a problem: Do you have an answer immediately? Thinking slow is when you iterate on it.

If you think about chess, speed chess is thinking fast, and regular chess is thinking slow. You're evaluating multiple opportunities and recognizing better moves when you string moves together, right? Even though AI is coming from computers, and so this is a little bit hard for us to see, AI is very intuitive. It's actually better at being intuitive than we are, and that's because it's been trained on so much data.

When Waymo is sending all their cars out, the amount of data that they get in a day is about—I don't know if it's now at the amount of experience that a human being gets in a lifetime of driving, but it's starting to approach that, at least. When you're getting a lifetime of driving data in a day, you're seeing every possibility. You don't need to figure out how to deal with the fact that some I-beam is falling off the back of a freight truck and about to hit you. You've seen it happen 2 or 3 times, and you know exactly what to do.

The thing is, the more these models produce data, the more they're able to intuitively deal with the situation because they've already seen it, right? And that's intuition: when you just have the answer. When you train these models, they used to be trained on just data pulled from the world that human beings were producing. Now what we do is use the models to generate the data that they get trained on.

You have a model at this level of capability, and it produces data at these levels of capability. It's gotten good enough that it can tell what good is. It keeps this data, trains, moves up to here. Then it produces data of this quality, prunes it to here, trains, and goes up to here, and just keeps moving up.

At this point, we're seeing these models improve at a pretty linear rate. There's no reason to believe that they're not going to get smarter. We may not recognize the difference between 2 really smart models, but 1 will be much smarter than the other. And that matters in the context of competition—competition and solving big unsolved problems.

John

Yeah, right. Do you want to talk about sentience?

Jonathan

Oh, jeez.

[Laughter]

6. Sentience Gets A Definition

Okay, I have a hobby. My hobby is to take words that people have used for centuries that don't have a good, concrete meaning and try to ascribe a definition to them that helps me understand the world. I'm going to give you my definition of sentience.

First of all, intelligence is your ability to make a prediction or influence an outcome to what you want to have happen. But it's stationary. It's like you have an amount of intelligence that is fixed in these models. But sentience is your rate of improvement in your intelligence.

That makes sense because when we talk about sentience, we're talking about the ability to self-reflect and get better, and that's an important part. So, I say that intelligence is your capability and sentience is your rate of change. But rate of change doesn't have to be binary. You're not sentient or not sentient. It's: How sentient are you? Are you linearly sentient? Are you asymptotically sentient?

You look at the world's best Go player, and he asymptoted. He stopped getting better because he didn't have better players to play against. People often conflate LLMs with just being intelligent, but there's something else that language gives you beyond intelligence. It gives you the ability to transfer information.

Going back to the example I gave with Waymo, if you have an entire civilization producing information, each participant in that civilization gets to benefit from that distillation of knowledge. While intelligence is a property of an organism or an individual, sentience is a property of a civilization.

AI is producing more intelligence. It's getting smarter. You interact with it, you get smarter. You ask better questions. You're making the AI smarter. And so there's this feedback loop of sentience that's accelerating in our society, and AI is contributing to that. As that happens, I would expect our kids to get much smarter than we ever were, just like we're probably smarter than our parents were because we had the internet.

The Inference Revolution: Groq, Nvidia and the Future of AI | BidClub