Jonathan Ross
We did not raise $1.5 billion. That’s revenue. That’s actually about 30% of the revenue of OpenAI. Your job is not to follow the wave; your job is to get positioned for the wave.
You can almost say we’re one of the best things that ever happened to NVIDIA, because they can make every single GPU that they were going to make and sell it for training. High margin gets amortized across deployment, and we’ll take the low-margin, high-volume inference business off their hands. They won’t have to sell either margin.
We are growing faster than exponential, and when you’re growing faster than exponential, there is no amount of profit that you can make that matters. What matters is getting a toehold in the market and becoming relevant.
Harry Stebbings
Jonathan, thank you so much for agreeing to do this in Paris. You look fantastic, by the way. I feel so underdressed.
Jonathan Ross
Thank you. I could take the tie off if you want, but I’ll never be able to tie it again. I don’t know how to tie a tie. My chief of staff has to tie it for me, and it’s a struggle because he’s putting it on himself and tying it.
I literally only bought this suit recently.
Harry Stebbings
You look fantastic. I don’t think I have a suit, so you’re one up on me.
I want to split the show into 2 parts. I want to talk about the landscape and where we’re at, and then I want to dive specifically into Groq and where you’re at. You’ve announced a massive new deal that I think everyone is slightly misunderstanding, which is what we were just talking about.
1. Scaling Laws and AI Model Training
I want to start with where we’re at in terms of scaling laws. Everyone says we’re at the limits of scaling laws, and then there seems to be exponential innovation happening with the likes of DeepSeek and others. Where are we in terms of the limits of scaling laws?
Jonathan Ross
Scaling laws is a paper that was published by OpenAI, and what it effectively says is that the more parameters your model has, the better it can absorb information. You’ll see these curves that they draw, and they’re amazing. You should show it if you can.
Effectively, you have these asymptotic drop-offs where you keep getting better and better, but you get a logarithmic improvement when you put a linear number of tokens in. This is why you see people doing 15 trillion tokens of training and whatnot. But they’re misunderstood because the assumption is that all of the data is the same quality.
You have a kid now, right? Eventually, you’re going to be training your kid, and you’re going to say—and play along with me here—what’s 1 + 1?
Harry Stebbings
2.
Jonathan Ross
What’s 2 × 3?
Harry Stebbings
6.
Jonathan Ross
What’s the second derivative of the square of the hyperbolic tangent?
Harry Stebbings
Good question.
Jonathan Ross
That’s how we train these models. We give them really simple problems to solve, and then we give them really hard ones. We don’t really train them up; we don’t do it smart.
Some people will train on the dregs of the internet, and then save some high-quality data for the end to make them better. But what you can do—and this is where I think everyone’s getting confused—is something like AlphaGo Zero, where it generated its own data and trained on it. You could have an LLM generate synthetic data, and when it generates the synthetic data, the data is better. You then train on that synthetic data.
Harry Stebbings
Why is synthetic data better than real data?
Jonathan Ross
Because the model is smarter. Reddit is great, but it’s not necessarily as high-quality as talking to someone with a PhD in a topic.
Just like with more expert people who are more knowledgeable and capable, if you have a better model, it generates better data. So you train the model, it gets better, you produce better data, and you produce a range of data here. You get rid of all the parts that are wrong, so now it’s the best part.
It’s a little better than the model itself because you’re pruning it, and you get to do this offline. Then you train the model and the model comes up here. You do this again, keep the better data, train it again, and you just keep moving up.
When you do that, the actual scaling laws don’t look like these asymptotes. They actually—
Harry Stebbings
But there has to be a ceiling on efficiency, no?
Jonathan Ross
There is a mathematical limit. If you study computer science, you’ve probably heard of something called Big O complexity.
Big O complexity is, if I’m solving a problem and I look at how I solve it, I might need to take more steps if I solve it with 1 algorithm versus another. For example, quicksort versus bubble sort: with quicksort, I need n log n steps; with bubble sort, I need n².
What’s the difference? If I’m sorting 1,000 numbers, n log n is 10,000 steps, but with n², that’s 1 million steps. It’s either 10 × 1,000 or 1,000 × 1,000.
One of the reasons these LLMs struggle to multiply large numbers is that multiplication is not linear. These LLMs can do anything linear without needing to think, but just like on a piece of paper, where you need to write out all those intermediate steps, these LLMs need that intermediate space and those steps in order to compute these things.
It’s a mathematical requirement. There’s nothing you can do to train a model enough that it will see any arbitrarily large number and just be able to multiply it. But you can choose bigger and bigger groupings of numbers to memorize, in which case it can do it in fewer steps.
Effectively, as you’re training the model on more and more data, it’s seeing more and more examples. Now it just has the answer for more specific situations, so it doesn’t need to do as much reasoning. But it still needs to do reasoning for some of these problems.
Harry Stebbings
What does that mean for the next step? If we have no efficiency ceiling, what does that actually mean?
Jonathan Ross
You need both. Training the model makes it more intuitive. It means that it can come up with the answer like that, with more stream-of-consciousness thinking. The reasoning part is different. Reasoning is the algorithm on top: the Big O complexity portion.
It’s System 1 and System 2 thinking, or “Thinking, Fast and Slow,” like Daniel Kahneman’s book. When you pair them together—when you make it more intuitive—you get better this way. But when you start adding in the System 2 portion, you start to get this.
2. Synthetic Data and Model Efficiency
The volume is very low, but when you do this, you get what they call polynomial—or you could think of it as geometrically increasing—improvement in the model when you combine that improved training with what they call test-time compute, or runtime compute.
Harry Stebbings
I totally get that. Just so I understand, when we think about bottlenecks, if synthetic data powers training and makes the model more intuitive—if it gets to the answer more quickly, like a grandmaster in chess seeing the right moves—synthetic data isn’t constrained in terms of its supply side.
If we think about the other bottlenecks, there’s hardware, energy efficiency, and algorithmic limits. What is the bottleneck?
Jonathan Ross
If your job is to get better at multiplying numbers, and I tell you that I want you to be able to do it with fewer steps and more intuitively, for you to be able to multiply 3-digit numbers versus 2-digit numbers, you need 10 times the data and 10 times the examples. As you get better on the intuitive part, you need more examples to train on.
Harry Stebbings
That makes sense. What is the bottleneck, then? Is it hardware quality? Is it compute? Is it algorithms? Because it’s not data.
Jonathan Ross
It is the compute, it is the data, and it is the algorithms. It’s all 3 of them.
People misunderstand the concept of a bottleneck. Compute has been less of a bottleneck and more of a soft bottleneck. When you provide even more compute, you can overpower the lack of data or the lack of improvement in algorithms. It’s not a hard bottleneck; it’s a soft bottleneck.
Ideally, you would improve all 3. You would be getting better data, better algorithms, and the algorithm improvements are going to be there. The data improvements are going to be there. But compute has always been the easiest lever because it’s so fungible. If I just give you more compute, it works better.
Harry Stebbings
Has DeepSeek not shown us that we don’t need the compute, and that you can do more with less?
Jonathan Ross
Not exactly. There was an algorithmic improvement. The algorithmic improvement, as I explained, is this seemingly silly thing where they just wrote the answer in a box, and then they knew what to look for rather than having to have a human being check it or something like that.
It was very simple, but that was an algorithmic improvement, and it made it easier to generate the data that was then trained on.
3. Inference vs. Training Costs: Why NVIDIA Loses Inference
Harry Stebbings
Can I ask about what I think are some misconceptions around compute, data—especially synthetic data—and algorithms? When you think about the biggest misconceptions people have around AI, and specifically inference, what do you think they are?
Jonathan Ross
The first misconception, which people don’t hold anymore, is that training was more expensive than inference. At Google, anytime we would train a new model, we would end up using 10 to 20 times as much compute on inference as on training.
Inference was always the critical infrastructure piece that we needed. But then, after getting past that, now everyone understands that inference is important.
Harry Stebbings
Do you think they fully do? When you look at NVIDIA’s stock price after DeepSeek, it was down 15%. If you understood the value of inference, it shouldn’t have been down 15%, with Jevons’ paradox and all that.
Jonathan Ross
I don’t agree that NVIDIA stock should have gone down because of that. I think that was a misunderstanding on most people’s part. But it also shows something else: everyone keeps saying NVIDIA stock can’t possibly go higher, and they were looking for an excuse to say, “Now that’s it. That’s why we were wrong, and we need to sell now.”
That has nothing to do with it. That’s just the popularity-contest side of the market. It had nothing to do with the weighing machine of the market.
Harry Stebbings
If a founder is building a startup today, should they build with the assumption that scaling laws will continue? Should they build with what we have today? How do you advise them?
Jonathan Ross
I would advise you to build based on things getting better, but I would also focus a little more on the big quantum steps.
The analogy I like is that if you look at the information age, we went through the printing press, the telephone, the telegraph, the internet, and smartphones. If you had built Uber when we only had the internet, it wouldn’t have worked. You’d book a ride, go somewhere, and then ask, “How do I get home?”
Harry Stebbings
Exactly.
Jonathan Ross
We’re in the same sort of space now. The models hallucinate, so it would be hard to build a medical-diagnosis company. It would be hard to build a legal company.
However, if you were doing that and the algorithmic enhancements happened that got the hallucination rate down, you would be perfectly positioned. Just like Groq, we were around for 7 years before we had product-market fit.
We were around because our bet was on scaled inference: that inference was going to be the bottleneck, and that we were going to need to run really big, heavy models. Everyone was assuming you would have a single PCIe card running inference because training was the complicated part.
The reality was that we made the right bet ahead of time, and then we were perfectly positioned. Your job is not to follow the wave; your job is to get positioned for the wave. That’s the hardest thing to do, because everyone is trying to talk you into coming ashore again. Almost everyone was telling us, “Don’t do LLMs. They’re going to be terrible for you.”
We were saying, “This is literally what we built for.”
Harry Stebbings
Did you ever doubt yourself? 7 years is an incredibly long wait time.
Jonathan Ross
There was doubt, but there was never a pause. Even before starting the TPU, I was concerned that AI was going to be a technology that would allow some people to have outsized control and outsized influence.
If you allow that to just happen in potentially not the best hands, it doesn’t really matter how rich you are. Nothing matters. It’s the most important technology.
It didn’t matter how hard it got. There was no choice but to be successful. Our goal is to preserve human agency in the age of AI. If we don’t do that, we have failed. It wouldn’t matter whether there was doubt or not.
There was plenty of doubt. There was a point where we were so close to running out of money that we did this thing we called Groq Bonds.
Harry Stebbings
War bonds from World War II, of course. For anyone who doesn’t know, what is a war bond?
Jonathan Ross
World War II was funded with bonds from the U.S. government. They had these posters saying, “Fund your troops,” and people would buy the bonds and receive a return. That funded the war effort.
We were very close to running out of money at one point. Rather than trying to pretend to be strong, we were vulnerable with our employees. We said, “We’re going to run out of money. We need you to trade equity for salary.”
We literally took pictures of the war bonds, put “Groq Bonds” on them instead, and had an all-hands meeting where we explained it. We were worried everyone was going to leave.
Instead of leaving, about 80% of the employees participated. I think 50% went to the statutory minimum salary required by law. When we finally raised the first bit of our $300 million round, we had so little money left in the bank that it was less than the money we saved through Groq Bonds.
Had we not done that, we would have literally run out of money. There were some really hard times. I know every founder has these moments, and from the outside, it’s so hard to understand. It’s like watching a TV show: you’re not in it.
When you’re there, everything is 10 to 100 times more intense, because people left their jobs and careers, and their families are banking on this. You have to make decisions like, “What would have happened if we went out there and asked everyone to do Groq Bonds and everyone quit?”
The shareholders would have said, “You have all of these people depending on you.” But if you lean toward vulnerability, people are often going to go with you.
4. The Future of AI Inference: Efficiency and Cost
Harry Stebbings
What is a world where inference is so crucial and 20 times more important than training? What does that world look like?
Jonathan Ross
The simplest way to understand it is to equate an LPU or a GPU to an employee. If you have enough of them, the LPUs or GPUs can do work, just like an employee.
It’s a little different in the sense that they can’t quit and take another job. You don’t have to retrain them. Once you get a model to a certain capability, it will always be at least that capable. It’s not going to regress, so you get consistency out of it.
Imagine that you’re a startup, and rather than having to go out and hire 100 people, you hire 10 and buy the amount of compute equivalent to 90 employees. That’s a very different way of thinking about the world, because now capital expenditure—or, in some cases, different types of operating expenditure—can be used instead of just employees.
In terms of inference, to give you a sense of our scaling, we started 2024 with about 640 chips in production. We ended with more than 40,000. This year, we want to be at more than 2 million, and next year the number is much, much larger.
5. Chip Supply and Scaling Concerns
Harry Stebbings
Are we seeing constraints on chip supply? That’s an unbelievable scaling story.
Jonathan Ross
For us to hit our numbers next year—which I’m not sharing publicly—we’re going to need almost all of the capacity of the fab that we’re using.
The biggest issue is that you don’t normally think of tech companies as having a cornered resource, but NVIDIA has a cornered resource. They’re a monopsony—the opposite of a monopoly, a single buyer—for HBM and the interposer, likely CoWoS.
Harry Stebbings
What is HBM?
Jonathan Ross
HBM is high-bandwidth memory.
Harry Stebbings
Which GPUs use it, and who produces it? Sorry for the dumb questions.
Jonathan Ross
There are 3 companies in the world that do this: SK Hynix, Samsung, and Micron. It’s specialty memory that’s only used in high-end servers.
There’s a limited quantity that’s built, it’s very expensive to ramp up, and it’s a very technically challenging type of memory to build. There’s a very limited supply, and GPUs are so fast computationally that if you were using regular memory, it would be like drinking out of a martini straw. It would take forever.
This is why you see people preferring to do inference, but especially training, on GPUs rather than CPUs, because the memory bandwidth is too limited. CPUs rarely use HBM; they mostly use regular memory.
Our architectural observation when we started Groq was that everyone knows Moore’s law: every 18 to 24 months, like clockwork, the number of transistors doubles, which means double the compute.
But we noticed that AI was getting better faster. It clearly wasn’t the algorithms, because algorithms have discontinuous jumps. It didn’t seem to be the data, because there wasn’t that much more data. The transistors were only doubling every 18 to 24 months, so where was all of this capability coming from?
It turned out that the number of chips was also doubling every 18 to 24 months. Rather than 2 times, it was 4 times. The question we asked was, if you’re effectively going to have an unlimited number of chips, do you do something architecturally different?
The answer is absolutely. Rather than using external memory, we use a large number of chips and keep all of the parameters of the model live in the chips. Then we have this pipeline where the computation flows through it, sort of like an assembly line.
Imagine you were trying to build a factory, and the factory was only 1/100th of the size needed for the assembly line. You would run a bunch of cars through 1/100th of it, tear it down, set up the next 1/100th of the assembly line, and do that over and over again. That’s the way a GPU works.
LPUs are very different. We have the computation flow through a whole bunch of chips. Rather than using 8 chips, we’ll use 600 or 3,000 for a model.
6. Energy Efficiency in AI Computation
Harry Stebbings
How does that change energy efficiency? How does it improve when you use more? You use more per token, so the footprint is higher.
Jonathan Ross
Think of it as the difference between a factory and a backyard garage. The backyard garage is not going to be as efficient, but it has a lower energy footprint.
Another example would be if you were trying to transport a ton of coal from one side of the city to the other, and you did it on mopeds or with freight trains. Which would be more efficient?
The moped would use less energy per trip, but it would need more trips, and therefore it would use more energy overall. In fact, this is one of the things most people misunderstand. They think edge computing uses less energy.
Actually, edge computing is less energy-efficient than computing in the data center. When you’re computing in the data center, it’s a little bit like that freight train. You’re getting to do a whole bunch of jobs simultaneously. The fact that we don’t have to read from that external memory means we don’t have to spend the energy doing that. Even with GPUs, you get to batch.
Going back to why it’s so energy-efficient, the amount of energy used in a chip involves physical wires. The physical wires have a width, and when you look at the width and the length, you charge that wire up to set it to a 1 and discharge it to set it to a 0. It’s like charging and discharging a capacitor, and you’re using energy.
The longer the wire, the more charge is required. When you have HBM here and another chip here, you’re charging a wire between the chips and discharging it every time you send a bit. That’s a long distance to travel, and the wires are wider than the wires inside the chips.
You use a lot more energy when we keep that memory in the chip. It’s only traveling a short distance using much thinner wires, so it uses a lot less energy.
Harry Stebbings
Do we see a world of LPU and GPU usage in combination? How does that distribution look between LPU usage and GPU usage?
Jonathan Ross
There are a couple of things. First, training should be done on GPUs. I think NVIDIA will sell every single GPU they make for training.
Right now, about 40% of their market is inference. If we were to deploy a lot of much lower-cost inference chips, what you would see is that the same number of GPUs would be sold, but the demand for training would increase. The more inference you have, the more training you need, and vice versa.
The other use case is that we’re so much faster than GPUs that we’ve experimented with taking some portions of the model and running them on our LPUs while letting the rest run on a GPU. It actually speeds things up and makes the GPU more economical.
Since people already have a lot of GPUs deployed, one use case we’ve contemplated is selling some of our LPUs to nitro-boost those GPUs.
Harry Stebbings
People have bought GPUs so far ahead of time that, by the time they get them, they’re deployed and installed, and they’re almost out of date.
Jonathan Ross
We’ve spoken with some customers that put orders in more than a year in advance, paid a year in advance, and still haven’t received them.
The recent deployment we did in Saudi Arabia went from contract to the first tokens being served in production in-country in 51 days.
Harry Stebbings
How were you able to do it so quickly? 51 days is astonishing.
Jonathan Ross
Part of it is that, architecturally, things are much simpler for us. We don’t have a bunch of other hardware components. We don’t use switches to communicate between our chips; we just plug our chips into our chips. Our chips are the switch.
We also don’t have all of this network tuning. Think about it this way: when you’re going across town in Paris, how long does it take to get from one side to the other?
Harry Stebbings
A long time.
Jonathan Ross
A long, variable amount of time. If you do it in the middle of the night, it might be fast. If you do it in the middle of the day during an event like the one we have going on with the AI Summit, it’s slow.
Harry Stebbings
Exactly.
Jonathan Ross
It’s unpredictable. However, certain modes of transportation, like trains, can be predictable. With what we’re doing, it’s 100% predictable given the energy efficiency and the predictability.
Harry Stebbings
Why is NVIDIA not being more proactive on LPUs?
Jonathan Ross
What makes you think they don’t want to be more proactive on it?
Harry Stebbings
They don’t talk about it. Why would they not talk about LPUs? If you wanted to protect shareholder value and a Wall Street image of dominance and being ahead of the game, you’d at least say, “Of course, we’re working on LPUs as well.”
Jonathan Ross
Until they had that ability, they would effectively be exposing that there was something missing. If you look at the last GTC, there was an announcement that the latest GPUs were 30 times faster than the previous generation.
When you look at how it was done, there was this curve that looked like this, and then it ended here. Then there was another curve that looked like this. That 30 times was from the end of this curve to the end of the next curve.
If you moved it here, it would have been less than 30 times. If you moved it here, it would have been infinite. Their chip is infinitely faster than the previous one, but that wouldn’t have sounded reasonable.
7. Why Most Dollars Into Datacenters Will Be Lost
There’s a history in this market of specmanship, because it’s so hard to get access to chips. This is a lesson in enterprise sales. People rely on specmanship: “My specs are better than your specs. My chip is faster than your chip. I get more teraflops per second than you do.”
Who cares? Tell me what the tokens per dollar are and what the tokens per watt are. Nothing else really matters.
People will find all of these other weird things to measure that they might be better on. It’s like saying, “I’ll sell you a car with better RPMs.” RPMs don’t matter. What matters is miles per gallon, and maybe the speed you can drive, although speed limits render that moot.
In enterprise sales, people used to buy or market soap by saying, “Our soap has more bubbles than this other brand’s soap.” Who cares? What they figured out was to put really happy people on a billboard after they used the soap, so people might associate that happiness with the product.
Harry Stebbings
Lifestyle marketing.
Jonathan Ross
Exactly. For some reason, enterprise still hasn’t learned that lesson. It’s still, “We have more bubbles. We have more teraflops. We have more of whatever”—things that people literally don’t care about.
Harry Stebbings
You think NVIDIA saying it’s 30 times faster isn’t good marketing?
Jonathan Ross
I think it worked because it’s what people are used to. Our counter was that we did a press release saying, “Groq still faster.” That was it, and people went crazy over it because it was simply, “We’re still faster.”
Harry Stebbings
Do you think Wall Street views it that way?
Jonathan Ross
I think they’re starting to. But again, I don’t think there’s real competition here. If you’re competing, you’ve done something seriously wrong.
If you’re competing, it means you haven’t found an unsolved customer problem. If someone else has already solved the problem, why are you spending time on it?
Harry Stebbings
So you don’t view NVIDIA as a competitor?
Jonathan Ross
They don’t offer fast tokens, and they don’t offer low-cost tokens. It’s a very different product. What they do very well is training. They do it better than anyone else, and by such a wide degree that it’s a solved problem.
Why would we bother trying to solve a problem that’s already been solved?
Harry Stebbings
So you cede the training market to them and own the inference market?
Jonathan Ross
They’re saying they also want the inference market, of course. It’s the way it always works.
Harry Stebbings
So what do we do now? Are we competing in the inference market?
Jonathan Ross
We don’t really have people saying, “We’re going to buy GPUs instead of you.” We do have people saying, “We’re going to buy both.” That happens, but we don’t care.
I showed a demo to someone, and he said, “Should we just not buy any more GPUs?” I said, “No. You should buy every single GPU you can get your hands on.”
He looked at me very perplexed. I said, “How are you going to do training? We don’t do training. Buy the GPUs. Get every single one you can, because I want your models running on us to be really good.”
For inference, they don’t need to buy NVIDIA anymore. They don’t need to buy GPUs for inference. But if you can get them, they’re a little expensive, and if you’re used to them, why not? Plenty of people still sell mainframes.
If you want lower cost and faster performance, you want an LPU.
Harry Stebbings
How much lower cost is it? More than 5 times lower?
Jonathan Ross
More than 5 times lower. The memory alone in the latest GPUs costs more than our fully loaded capital expenditure per chip deployed.
On top of that, we talked about energy efficiency. We use about 1/3 of the energy per token. Over a 3-year period, 1/3 of our cost is operating expenditure, which is mostly energy and data-center rent, and 2/3 is capital expenditure.
Since we use 1/3 of the energy, the cost to run that GPU to produce the same number of tokens is the same as our total cost. Just the GPU’s operating expenditure is the same as our capital expenditure plus our operating expenditure.
Harry Stebbings
Why is 40% of NVIDIA’s revenue inference, then? Why haven’t you taken more of that?
Jonathan Ross
At the beginning of 2024, we only had 640 chips. At the end, we had 40,000. We’re not at that scale yet.
You have to provide quality, low cost, speed, and capacity. This is where the most important part of not using HBM came in. It means that we effectively have no scale limits.
The GPU itself is manufactured using the same process you use for a mobile phone. The same silicon that’s in your mobile phone is the same silicon used for the GPU. In fact, they build the mobile-phone chips first because they’re smaller, and NVIDIA gets them after Apple.
The difference is the memory. That’s the only difference. But that memory is the hard part to manufacture, and it’s limited in scale. By avoiding it, we effectively have almost no limit on how much we can scale up.
That’s important for inference.
Harry Stebbings
What is NVIDIA’s margin? 70% to 80%?
Jonathan Ross
70% to 80%.
Harry Stebbings
So they can take 70% to 80% off and be radically more cost-effective than you?
Jonathan Ross
Comparatively, yes.
Harry Stebbings
You could destroy their margin. Why wouldn’t you?
Jonathan Ross
In that same vein, you can almost say we’re one of the best things that ever happened to NVIDIA. They can make every single GPU that they were going to make and sell it for training, where it has a high margin. That gets amortized across the deployment.
We’ll take the low-margin, high-volume inference business off their hands, and they won’t have to sell either margin.
Harry Stebbings
What does low margin mean?
Jonathan Ross
Depending on the deal, we do get some on the back end, but up front it’s about 20%.
Harry Stebbings
So there’s 80% for NVIDIA and 20% for you, but you’re looking at a 20-times advantage.
Jonathan Ross
Then we get more later.
8. Meta, Google, and Microsoft's Data Center Investments
Harry Stebbings
What do you mean, you get more later?
Jonathan Ross
The deals we do are structured so that the partner puts up the money for us to deploy. We pay it back with a decent IRR, but we split the revenue, and most of it goes to the partner. Once we hit the IRR, it flips the other way.
Harry Stebbings
So others are putting up the capital expenditure for you?
Jonathan Ross
Yes.
Harry Stebbings
What does it look like at the end?
Jonathan Ross
It’s not like other business models. We didn’t just innovate on the chip; we also innovated on the business model.
We’re limited in how much money we can make based on how much we can deploy, not how much money we have, because the partners are putting that money up. When I’m looking at what we can do, it’s all about how much we can scale.
Harry Stebbings
What are the limits to your deployment? Is it purely chip constraints?
Jonathan Ross
Mostly.
You asked about misconceptions in AI. I think one of them is around power. It’s true that there’s a mismatch in the market between people with chips and people with power, but that’s partially because you need a data center in the middle, and there aren’t enough data centers.
Data centers aren’t the hardest thing in the world to build. They’re not easy, but they’re not the hardest thing. It’s harder to build up the power.
Because of that mismatch, you have big hyperscalers going around saying, “I need 1 gigawatt of power,” and they’ll say this to 60 different potential data-center builders. All of a sudden, you hear an echo: “I heard there’s a gigawatt here, a gigawatt there, and a gigawatt somewhere else.”
Suddenly, there’s 60 gigawatts of demand. It’s just an echo from that first gigawatt.
I’m aware of about 20 GW of power that people want to make available for data centers. Right now, there are about 15 GW of data centers worldwide, so that’s more than double the current capacity.
My concern is that people are now building up more power. In the next 3 to 4 years, people will say, “I built up all this power and no one is using it. This was a complete waste, and we’re never going to do this again.”
But remember that doubling of chips every 18 to 24 months. Over 3 to 4 years, you double that 15 GW twice, and now you’re talking about 120 GW. There isn’t that much power available. Then you double it again, and now you’re at 240 GW.
What’s going to happen is that we’re going to overbuild slightly right now because of the mismatch and miscommunication. Then we’re going to dampen our building and close down on that, and then we’re going to have the real need for the power.
That power will become a hard bottleneck in 3 to 4 years.
Harry Stebbings
Why will we have data-center oversupply when we’re moving into a world of inference that will be 20 times larger than training?
Jonathan Ross
The problem with data centers is that everyone thinks data centers are real estate. A lot of people do real estate data centers, but data centers are not real estate.
The common joke in the industry now is that someone says, “I’m going to have 100 megawatts of capacity for you in 3 months. Are you willing to sign?”
Then you ask, “What’s your uptime?”
They say, “I don’t know. Whatever the power grid is.”
You ask, “Where are your generators?”
They say, “I haven’t ordered those. I’ll order them now.”
There’s a 90-month lead time on generators right now.
Harry Stebbings
Really?
Jonathan Ross
Then there’s the next question: “Where are you getting the water from?”
They say, “Data centers need water? I thought it was just a bunch of chips.”
There are a lot of people who have no idea what they’re doing going into this because they think it’s real estate. Those people are building an oversupply of data centers, but they’re not really building them. They’re fake data centers that people think are real.
Harry Stebbings
What happens to those data centers if they’re not utilized? Is Amazon going to pay for a data center that doesn’t work?
Jonathan Ross
Amazon doesn’t fall for this. Amazon has really good people. Whoever the buyer is isn’t going to pay for a data center with no water or power.
Harry Stebbings
So these projects will never be developed? Will we build them fast enough?
Jonathan Ross
It does take time to build the data center, so it’s almost okay. If you train a model, you really want to amortize it over about 6 months. If you deploy chips, you really want to amortize them over 3 to 5 years. We’re more on the 3-year side; others are more on the 5-year side.
If you build a data center, you’re probably talking about 10 to 15 years. For a power plant, you’re talking about 15 to 20 years.
The problem we have in the industry is this mismatch between the financing and the needs. Someone wants to train a model, and they’re going to be doing that for 6 months. They don’t understand why people want 3-to-5-year commitments on the chips.
The people deploying the chips don’t understand why someone wants a 15-to-20-year commitment on the data center.
Harry Stebbings
It’s at 7 years now on the data centers, and then the people building the data centers need a long, 7-year commitment?
Jonathan Ross
That’s the kind of commitment they’re asking for.
You have this complete mismatch throughout the ecosystem. The funny part is that, while they all want to take zero risk and have a sovereign-wealth-level credit rating on the other side with long commitments, the longer the payoff time, the more generic the infrastructure is.
A model has a very specific use, but accelerators like LPUs and GPUs can be used for other things besides generative AI or LLMs. The data center can be used for other things besides the accelerators. The power can be used for anything.
While they’re looking for the least risk, they’re looking for it in the place where there is the least risk. If we don’t use the power for AI, we’ll use it to power all of the electric cars.
Harry Stebbings
Is this a case where incumbents win because they’re one of the only ones able to match the durations required by data-center providers?
Jonathan Ross
This is why we’ve partnered with Aramco and this new entity in Saudi Arabia. They have an enormous ability to fund this over the long term, a very long-term perspective, and an amazing credit rating.
Harry Stebbings
When you say they have the ability to fund it, this is why there was a misconception. People think it’s a funding round of $1.5 billion.
Jonathan Ross
We did not raise $1.5 billion. That’s revenue. That’s actually about 30% of the revenue of OpenAI.
Harry Stebbings
Can you walk me through how that deal is structured?
Jonathan Ross
We started off last year, and we got to 19,000 of our chips deployed. We did that in about 51 days. The question was what we could do this year.
They’ve gone off and collected a bunch of power in the country. The deal is structured so that they will put up the capital expenditure for us to deploy our chips in that data center or those data centers, and we pay it back based on the money we make.
It’s a little different from debt in that they participate in the upside, but it’s similar in nature. It is revenue because we actually make a profit upfront.
Harry Stebbings
How does that change what you can do?
Jonathan Ross
We’re not limited by capital anymore.
There’s another misconception around Groq. There was a paper that said we couldn’t be profitable while being the lowest price. It said we could charge more, but we actually have a very positive contribution margin right now.
As far as we know, we’re the only ones making money running these open-source models. With the open-source models, everyone is competing with venture-capital dollars, trying to take market share in an Uber-style model.
Meanwhile, we’re sitting here saying, “We could do this all day long,” because we’re making money. We’re able to pay off an IRR and make our partners money.
There’s another part of the model. We’re also working with some proprietary model providers. We showed off the first one at LEAP on Sunday, where we did a voice model with PlayAI. That one is also a revenue share.
The difference is that they get to make money off it, whereas most others in the industry are losing money because of the commoditization of the models.
Harry Stebbings
Do you have cheaper pricing over time as you have less monopoly power, or do you have higher prices as your monopoly increases?
Jonathan Ross
We want the margin to stay about the same, but we want prices to go down. Then we get into Jevons’ paradox, and life gets great because we’re going to scale.
Our focus is on getting to scale. To preserve human agency in the age of AI, we need to be one of the most important compute providers in the world.
Our goal by the end of 2027 is to provide at least half of the world’s AI inference compute. We think we could be further than 2 times that, given that we don’t have all the constraints.
To get there, we need to be aggressively building out, and we need to give people no excuse for not running their models on us and using the models that are on us. We do that by charging extra.
What I keep telling the team over and over again—you have to remind them sometimes—is that we are growing faster than exponential. When you’re growing faster than exponential, there is no amount of profit you can make that matters.
What matters is getting a toehold in the market and becoming relevant.
Harry Stebbings
What would prevent that?
Jonathan Ross
We used to worry that someone would try to price below us. Then we realized that wasn’t a concern, because so much money is going into this that people are going to want to lose less money by running on us.
That isn’t a concern. It was the big one early on, until we realized that when we see Mark Zuckerberg investing $65 billion in data centers, he’s internalizing all of the margins he would have had to spend on data centers with the providers we mentioned earlier.
Harry Stebbings
Meta is doing $65 billion a year. I think Google said $70 billion or $75 billion, and Satya said Microsoft is doing $80 billion. Then you’ve also got Stargate.
These are crazy sums of money. Is all of this for data-center building?
Jonathan Ross
It also includes the things that go in them, including the chips, the systems, and everything else.
Harry Stebbings
We’ve never seen money like this.
Jonathan Ross
No. There’s never been anything like this. But there’s never been a case where it was so clear that there was going to be value at the end.
If you knew how successful search was going to be, remember that Google stayed private as long as it did because they were afraid Microsoft would figure out how much money search was making and try to replicate it.
The moment Google went public, Bing appeared. They called that perfectly. Everyone knows how much money there is in AI, so everyone is going after it.
9. Distribution of Value in the AI Economy
Harry Stebbings
Do you think that value is distributed among many players or concentrated toward 1 or 2? I completely agree with you in terms of the clear value when assigned, but is it distributed somewhat evenly or concentrated?
Jonathan Ross
It’s a power law. The more value there is in the economy, the more risk there is of a single entity being so far on one end that it just dominates.
You see this with the Magnificent 7. The bigger the economy gets, the more you’ll have big swings in the economic outcomes.
Right now, the hyperscalers are all roughly even in their market caps. It’s strange. You would expect one of them to be killing it and taking it much further, so I don’t understand why they’re so closely grouped.
10. Stages of Startup Success
Harry Stebbings
When we think about that distribution, how do we think about changing it? With Groq, you want to be one of the Magnificent 7 and one of the most important companies in the world. How do you see that happening?
Jonathan Ross
The way you get there and the way you stay there are 2 very different things.
There’s a circle of life that happens in startups. The first stage is to solve an unsolved problem. That’s how you go viral and do well.
The second stage is the marketing stage, where other people are trying to copy what you’ve done because they can’t think of something themselves. Now you have to fight it out in advertising, marketing, and so on. You see consumer-packaged-goods companies often get stuck there, and it becomes more about where they are on the shelf than anything else.
The final stage is the 7 Powers. It’s once you’ve found some of those powers, started improving, and developed systemic advantages.
Then someone solves an unsolved customer problem and the whole cycle of life continues. Google has to redo this now because LLMs are better than search.
11. The AI Investment Bubble
The way you become one of the Magnificent 7 is by solving that unsolved problem. The way you stay there is by finding one or more of those 7 powers, and then being ready for when you get disrupted so you can continue fighting back and solving customer problems.
Harry Stebbings
We mentioned the huge amounts of money being spent here. Is this a good bubble that lays the foundations for an incredible next 10 to 20 years, where the capital actually turns out to be productive even if it doesn’t seem so on paper? Or is it a case where a huge amount of money is incinerated on depreciating assets?
Jonathan Ross
I can guarantee you that a huge amount of money will be incinerated. But I also bet that, in total, more money will be made than will be put in.
That’s the problem. You have to look at it either in aggregate or as individual bets. When everyone is making investments in the market, some people are going to lose money because not every company is going to be successful.
What you always see when there are real technology improvements or things coming is that you have the early things that people invest in heavily, and they’re super successful. Then everyone else wants to get in on it.
You go from AI chips and AI models to AI T-shirts, and next thing you know you have AI thermal grease. People start applying AI to everything. Next thing you know, you’ll have an AI condo.
Harry Stebbings
Sure.
Jonathan Ross
The trick is discerning what’s real and what isn’t. You’re always going to have obnoxious charlatans coming in whenever there’s something real. That’s unfortunate, but eventually they get cleared away once people understand the technology and what’s real.
The job is to start educating. The more educated people are, the less they’ll invest in AI thermal grease.
Harry Stebbings
What is the largest individual bet that will lead to the largest incineration of cash?
12. The Keynesian Beauty Contest in VC
Jonathan Ross
I’m not going to call anyone out in particular, but I actually think it will happen across every single discipline.
Are you aware of the Keynesian beauty contest?
Harry Stebbings
No.
Jonathan Ross
John Maynard Keynes, the economist, had this great idea. It explains everything you need to know about venture capital.
Harry Stebbings
I’m nervous, but keep going.
Jonathan Ross
Take a magazine full of models—human models, good-looking models—and have a whole bunch of VCs in the room. They’re allowed to make bets on who the most beautiful model is.
In the end, whoever has the most money on them is the winner. Based on the proportion that you put on that particular model’s face, you get a share of all the money.
If you put money on one that isn’t the most beautiful by dollars, you lose your money to the people who bet on the one that was. That was the bet SoftBank was making: “I can win the Keynesian beauty contest. I’m just going to put more money in, and I’m going to win.”
That’s problematic when you have true technological advantages as opposed to marketing. When you’re solving customer problems, it’s a weighing machine. Once the customer problem has been solved, you get into this popularity contest of marketing.
Something unusual has happened this time around that I don’t think has ever happened in venture capital before. You see people raising billions of dollars who have competitors that have also raised billions of dollars.
Usually, there’s a clear winner in the Keynesian beauty contest. You don’t have this fight where it’s, “I have to put a little more money in. I have to put a little more. I’m going to put in $10 billion. I’m going to put in $20 billion. I’m going to put in $500 billion.”
The Keynesian beauty contest has gone completely amok. This has never happened before, so people don’t even understand how to react. It used to be that if someone had raised $1 billion, you would say, “They’re the winner.” Now there are 3 or 4 competitors who have $1 billion, so who wins and who loses?
Harry Stebbings
Is Masayoshi Son going to incinerate the largest amount of cash ever?
Jonathan Ross
I think the Keynesian beauty contest no longer applies because there’s so much money available, spread out across the market.
I think the people who have the best products are actually going to be the winners. Everyone can be capitalized, but there will be problems for the winners because of this.
The problems will be things like the employee you were going to hire, where someone offered them a ridiculous amount of money. You see this all the time now. They could have gone and contributed to the winner, but now they’re contributing to a competitor that shouldn’t exist, or that’s equally likely to win. Now you’re splitting the talent.
Harry Stebbings
What do you do when you have such high salaries? We’ve seen $1 million or $2 million for junior-to-mid-level people at some of these companies, and they’re living an amazing life in great places.
Jonathan Ross
Do you think they’re living that amazing life in Guangdong when they’re working for DeepSeek or another Chinese alternative? I don’t think so. They’re getting paid much less, working their asses off 20 hours a day, not getting kombucha, and not being paid $2 million a year.
Harry Stebbings
Fair.
Jonathan Ross
Not only fair—we have a policy that we never offer the highest salary. We want people to choose us, not choose the salary.
If we win in a bidding war, the next time someone comes along with a higher salary, that person is just going to take the other job. There’s no loyalty, and they don’t believe in the mission.
Instead, we focus on saying, “We’re going to build this. This is your opportunity. You’ll get to work with amazing people.” Spend some time with the team. Are these the people you want to be working with?
Frankly, you’re going to make so much cash that it doesn’t matter. Bet on the equity and the outcome. Help us make this thing valuable.
People who buy into that are much easier to manage because they’re mission-oriented. They all want to do the same thing. They’re not there because they want kombucha, and they’re not going to complain because the cappuccino machine is broken. They’ll just go and buy their coffee next door.
13. NVIDIA's Role in the AI Ecosystem
Harry Stebbings
Will you and NVIDIA move into the model layer? Everyone talks about model builders becoming application providers. Will infrastructure providers become model providers?
Jonathan Ross
We’ve decided that we’re not going to train our own models. We’ll do a little fine-tuning for specific cases, but we don’t want to compete.
That’s important because people are putting their models and weights on us, and they don’t want us to learn from and take that information for our own benefit.
That’s the problem you have when you work with a hyperscaler: they’re also doing everything you’re doing. We’ve decided that model providers should make the model. We don’t do that.
There’s also the data side—the users and the queries. Another thing we could do, but do not do, is log the queries and then use that data if we wanted to train.
We don’t train, and we have no reason to hold the data. We only temporarily store things in DRAM, so there’s no persistent storage. If the power went out, everything would be gone. DRAM is limited, so we can’t hold things for a long time.
You know that we don’t have your data. People who are building businesses on top of us can obviously keep the data from their customers if they want. We have no control over that, and that’s fine. But we don’t take any data.
Harry Stebbings
Do you think NVIDIA will move into model provision?
Jonathan Ross
It’s possible, but if I were them, I would avoid it. I wouldn’t want to give my customers to a company that I was competing with.
NVIDIA is great at training. It’s crazy. It would be like being an automotive company and then creating your own taxi service. You’re competing directly with your customer.
Tech companies love to do this. We have a management philosophy based on Big O complexity, and we only do things that require a sublinear number of employees.
If someone comes to me and says, “I need 10 people to go do this,” a lot of people would say, “Why can’t you do it with 5?”
I would say, “You’re supporting customers. If we double the number of customers, do you need 20 people or 11?” I want to know the growth rate. Are they automating everything?
We completely automated our compiler. We automated large portions of our cloud and everything else. That means we can scale with a small team.
We have 300 people. We built our own chip, networking hardware and software, runtime, orchestration layer, compiler, and cloud. We built all of this with 300 people.
We can only do this with a small number of people because you don’t have the communication overhead. You have to decide what your constants and variables are—what are the things you want to preserve?
One of our constants is talent density. We want to stay small and nimble.
The other side is that growth is a problem. We measure our growth in what I call problem units. Every time you triple something, you have about the same number of problems as the last time you tripled it.
Going from 100 employees to 300, from 300 to 1,000, and from 1,000 to 3,000: each of those has roughly the same number of problems.
We scaled from 640 LPUs at the beginning of last year to 40,000. That’s 4 problem units—4 triplings of the number of chips. If we were also tripling the number of employees, that would be another problem unit.
Management bandwidth is limited. You can only solve so many problems, so you have to decide where you’re going to allocate them.
If you build things really well from the beginning and can scale up with the number of employees you have, then you can scale over here. If you want to triple the number of customers, there’s another problem unit that you have to solve.
Harry Stebbings
What’s the biggest challenge when you’re scaling at that rate but the team isn’t scaling in conjunction with it?
Jonathan Ross
There’s a common belief that the people you have early on are right for the job, and that the people you get later might be better in a more corporate environment. I don’t think that’s the case.
You should always try to get generalists. Otherwise, you get stuck in a particular way of doing things because that’s what one person knew how to do.
There are people who burn out. Being in a startup is hard, and some people literally burn out. There are also people who were the best you could get at the time, and people who are unmanageable wild children. They should go off and start another startup; they shouldn’t be scaling with you.
That happens, but it’s rarer. Saying you’re going to hire B-players because you’ve gotten large enough is laziness and an excuse. It’s a lack of creativity in your business model and in the algorithm of how you’re going to scale.
Think of it this way: Walmart versus Amazon. Walmart wants to double the number of customers, so it has to double the number of stores and employees. Amazon doesn’t need to double the number of websites.
That’s a fundamental advantage, but Amazon still has to double and improve its logistics. It doesn’t have as many problems that have to scale linearly, but it has some.
If you wanted to disrupt Amazon, you would build a completely robotic logistics system and bring the complexity and overhead down. Then you could outmaneuver them.
That’s how you need to think. Don’t just say, “I need more people.” Focus on the algorithm of your business.
14. China's AI Strategy and Global Implications
Harry Stebbings
The last time we spoke, we discussed DeepSeek. I think more has come out over the last few weeks about their innovations and some of the distillation they used. Where is China better than us today?
Jonathan Ross
As we discussed, they’re more willing to use things that perhaps they shouldn’t be using. They distilled the OpenAI model.
A lot of people have the opinion that OpenAI was scraping the internet, so, “Good for DeepSeek.” But whether that’s right or wrong, most model providers considered it a red line. They didn’t want to cross it.
I don’t know if that’s going to change, but it might.
Harry Stebbings
What about the open-source nature of DeepSeek? Does OpenAI now benefit from the innovations they made?
Jonathan Ross
They also probably have all the data that DeepSeek paid them to generate.
Harry Stebbings
But DeepSeek was clever. They innovated.
Jonathan Ross
I think the biggest thing is that this is a shot in the arm for morale in China. It gives them a sense that they’re in the race.
But, as I said, it’s Sputnik 2.0. It has also woken up the United States.
Harry Stebbings
How do you compare Stargate with the $128 billion China has now committed?
Jonathan Ross
China has a more complicated situation and a simpler one at the same time.
The problem is that they don’t have the chip efficiency we have. On the other hand, they have scale. If they wanted to deploy 150 nuclear reactors—I think that’s the plan—it’s no big deal. They just do it.
If the chips aren’t as efficient, they can deploy more of them. On the other hand, if they want to go out into the world and deploy chips the way they did with Huawei and networking equipment, that’s going to be complicated.
People around the world aren’t going to have the power to run more expensive accelerators. At home, I don’t think anything is a problem. As they try to expand, it’s going to be an issue.
Harry Stebbings
China is quite opaque in everything. What do we not know about China that we would like to know?
Jonathan Ross
The most important thing to understand is where they’re going to end up on the censorship and privacy of these models.
We come from democratic countries, and we have an expectation that companies can build something that says anything. Are they going to be permissive and allow models to make mistakes and hallucinate, or are they going to shut them down?
If you know that, you know whether China has a shot.
One of the biggest nightmares they have is free speech. It’s the exact opposite of the vulnerability we talked about earlier.
Can you imagine Xi Jinping going out and saying, “Country, we’ve lost our advantage in AI. I need your help”?
Harry Stebbings
Never.
Jonathan Ross
It’s always going to be, “We’re the greatest. We’re the best.” Everyone will know differently, but they’ll all have to toe the party line.
Because of that, I think it’s hard for them to allow these models to say anything. For them, it’s a bad thing if the model says the United States is great and better at something.
That’s going to tell you a lot about the AI story in China. If they aren’t permissive of more open and truthful models, then they’re inherently disadvantaged.
Harry Stebbings
You’re saying that if they aren’t more permissive, and you’re running a Chinese tech company, your fear is that you become Jack Ma.
Jonathan Ross
That’s really going to stifle innovation. If I were in China right now, I’d be looking for the exit. If your craft is AI, you’d want to do that somewhere supportive.
Harry Stebbings
Do you really buy that they don’t have access to Blackwell? This is China. I can imagine Xi Jinping saying, “Sorry, no Blackwell.”
Jonathan Ross
I don’t think it matters whether they physically have it. Right now, most cloud providers are happy if you swipe a credit card to rent it to you.
Harry Stebbings
But there are limits to renting.
Jonathan Ross
I think one of the concerns right now is Malaysia or Singapore—or that region—being a place where people are deploying GPUs with a wink and a nod: “We’re not going to rent them to China.”
A lot of people believe that’s happening. Otherwise, that’s a lot of GPUs for that region.
Harry Stebbings
It feels like an even bigger safety net in case the tap is turned off at the hyperscalers, because right now you can write a check to any of the hyperscalers and say, “I need these chips.” They’ll deploy them and you can run on them. It doesn’t really matter where you’re coming from.
Jonathan Ross
If you’re a sanctioned country, it matters.
Harry Stebbings
China isn’t sanctioned.
Jonathan Ross
China isn’t sanctioned.
15. Europe's Potential in the AI Revolution
Harry Stebbings
So we have China, which is obviously proving that it’s in the race. We have the United States, and then we have Europe, which feels like it’s languishing. Is this the ultimate nail in Europe’s coffin?
Jonathan Ross
We talked about how Groq almost died, but we had the right technology all along. We were just waiting for LLMs to arrive. I think Europe is very similar.
Europe has amazing talent—amazing talent—but that talent leaves and goes to the United States or other places. The question is, how do you have Europe’s LLM moment?
It’s not that complicated. When you surround yourself with people, you become the average of your 5 closest friends. If your 5 closest friends say, “That’ll never succeed. You should just keep your job. Startups are terrible,” you’re going to be risk-averse.
If your 5 closest friends say, “You should do it. That’s great. I support you,” you’re more likely to start a company.
Even in Silicon Valley, people make the transition from a big tech company to a startup, and it’s hard. They’re comfortable, making those crazy salaries. The big companies take care of them, and they have a fiduciary obligation to their families.
How do they make that leap? It’s because they have entrepreneurs constantly trying to hire them. They hear the pitch all the time, and they get used to it. They see success around them, and VCs come in and try to close candidates in the early stages.
Europe needs the same thing. You need a place where people are surrounded by entrepreneurial people who are risk-on and aren’t going to try to talk them out of joining a startup.
Harry Stebbings
From a regulation perspective, Europe is unbelievably efficient in the mastery of regulation. I was speaking with someone the other day, and the EU has supposedly hired 1,500 people for AI safety and policing.
What would you do if I put you in charge of European AI regulation?
Jonathan Ross
I wouldn’t waste my time regulating something that doesn’t exist. Instead of regulating, what are you going to promote?
You want to promote risk-taking. You want to promote an enclave of people who are risk-on.
I was visiting Station F yesterday. It was amazing. Macron was there, and it was full of vibrant people. You could feel it.
I was talking to the person who runs Station F, Roxanne Varza, and Xavier Niel. We were talking about City F: a place where you start with 10,000 people in the center, within a small radius. Once it’s full, you expand it, and once that’s full, you expand it again, until you get to perhaps 1 million people in Europe who are all risk-on.
It would be a little Silicon Valley. I would give it special economic dispensations. I would allow everything that employers need, make it simple, and say, “If you don’t want to buy into that, go to other regions in France or other regions in Europe. But if you want to participate in what’s going to be the biggest technological revolution in human history, this is the city for you.”
Harry Stebbings
You’re inherently punishing incumbents. If we’re talking about AI insurance-underwriting startups, there are many companies going after insurance underwriting with AI. If you give them benefits like that, you’re inherently punishing some of the biggest insurance providers in your region.
You’re punishing people who hire 200,000 people. That feels unfair.
Jonathan Ross
There is no right to be an incumbent, especially a slothful incumbent that isn’t reacting to disruption. You want to encourage disruption.
One of the things in Silicon Valley is that you can move from one place to another. There are no non-solicitation agreements anymore. When I started, we had them, but even those are gone.
That free movement of people is very important.
Harry Stebbings
Are you allowed to start work straight away?
Jonathan Ross
Straight away.
Harry Stebbings
But not before?
Jonathan Ross
If you start before, that’s a problem.
Harry Stebbings
We have to wait 6 months.
Jonathan Ross
There’s no such thing in Silicon Valley.
Harry Stebbings
If you’re a company right now, it feels like it’s harder to poach people. But what does that do? It suppresses wages. It’s harder to hire someone, people are less likely to move, and there’s less competition.
It suppresses wages, and the company has to pay for the 6 months anyway. It makes no sense at all.
Jonathan Ross
I totally understand that.
Harry Stebbings
You mentioned what you would promote. A lot of people would promote safety and regulation. Being European, I thought first about safety, specifically.
All Dario talks about these days is safety. Is he losing a step by being so focused on safety when, bluntly, his competitors are talking about product?
Jonathan Ross
Safety matters in AI. It’s a little bit like nuclear power: there are lots of pros and lots of cons.
I’m worried about different things than Dario is worried about. I’m more worried about people voluntarily giving up their decision-making authority because it’s so easy. This is what I mean by preserving human agency in the age of AI.
A good analogy is that you probably know plenty of wealthy people and the struggles they have bringing up children with wealth. I refer to it as financial diabetes.
You have children who aren’t incentivized to strive to succeed. I was very fortunate when I was growing up. My father lost all of his money multiple times.
He would sell a billion-dollar life-insurance policy and get all the commissions from it. You would have tons of money, and then you would spend it all.
There was a time when we were living in a $20 million mansion. There were a couple of times when we ordered food, and he would talk to the delivery guy and convince him to give us the food and let us pay him back later, because we would get money later.
One time, he was so despondent that he locked himself in his office and wouldn’t come out. My little brother came to me and said I had to go talk to the Chinese-food delivery guy and convince him to give us the food.
I was mentally preparing how to convince him. I walked out, walked up to him, and started getting ready to make my whole speech. He handed me the food.
I said, “I don’t have the money right now.”
He said, “Pay me later.”
I didn’t have to do anything. He trusted us because we were living in a $20 million mansion.
That happened multiple times. I have a friend who was homeless once for a couple of weeks, and he had almost been homeless a couple of times. He said the best thing that ever happened to him was being homeless for a couple of weeks because he survived it.
He said, “I’ve been through it. I always viewed this as the worst thing that could ever happen in the world, but now that I’ve been through it, I can survive it. I’m not worried anymore.”
I think we live incredibly comfortable lives—way too comfortable. Most people don’t have to go through things like that, so we have financial diabetes as a society.
I think it’s going to get worse with AI. We’re going into an age of abundance. Very few people have to worry about food security now, but what happens if you don’t need to worry about housing or anything else? What happens if you can live a life without working?
What is that going to do to your psychology? As we enter an age of abundance, how do we get people to keep making their own decisions and have a fulfilled life?
Harry Stebbings
Do we get better, or do we become accepting of good enough?
Bluntly, with the majority of shows, we start with OpenAI and do deep research. Then we use different prompts depending on the guest, and supplement that with a huge amount of research from speaking to ChatGPT, speaking to Claude, and speaking to everyone in between.
We care about it being good enough first and great later, with all the references. Most people will just be happy with good enough and get away with it.
Do we, as a human society, become happy with good enough?
Jonathan Ross
When we hire, we hire for something we call “booking the win early.” One of the most important driving forces for people is loss aversion. When you have something, you don’t want to lose it. People are less likely to go after something they’ve already had.
You grew up in a family that was well-off and then lost that. That might be part of your drive, because you want to get back to it.
When we’re hiring an engineer and there’s a room full of people saying, “If we do this thing, we could be twice as fast,” I want that engineer to hear, “If we don’t do that, we’re going to be half the speed we could have been.”
That’s the loss aversion. Book the win early. Because it’s possible, it must be done.
I think that’s a smaller segment of the population. Those are the people who deliver amazing things that no one else is going to do, because everyone else says, “That’s good enough.”
With AI, it’s so easy to create a prototype that, to stand out, you’re going to need to do more.
One of the things that has happened with the ability to communicate more freely and see what other people are doing is that the average restaurant is better than high-end restaurants were 20 years ago.
People see everything others are doing and start to expect that. You have less localization and more globalization, so you have to compete at the highest level.
AI is no exception. There are going to be 40 people creating the app you’re creating. You have to polish it in order to stand out.
Harry Stebbings
Listen, Jonathan, I could talk to you all day. I do want to do a quick 5.
What do you believe that most people around you disbelieve?
Jonathan Ross
I’m going to go with anti-Founder Mode. I’m anti-Founder Mode.
I believe in delegation. When you’re telling people how to do their jobs, that’s an indication that it’s not necessarily a problem with you. It could just be that the person isn’t right for the job, and it’s easier to direct them than to find someone else competent.
But it also means you probably haven’t aligned them.
We align people through this challenge coin. Everyone at Groq carries a 25-million-tokens-per-second challenge coin. It tells everyone what we’re doing. It’s alignment.
I can’t tell you how many people I’ve shown it to who say, “That’s awesome,” and yet no one else is making them.
Harry Stebbings
It’s heavy.
Jonathan Ross
It is heavy. But you like to say that the greatest things in life—the heaviest things—aren’t gold.
Harry Stebbings
I know. Gold, but made decisions.
Jonathan Ross
Exactly. This was a made decision, because I had to consolidate everything we were doing into 1 very simple message: “We’re going to get to 25 million tokens per second.”
I engraved it on a coin in this tiny amount of space and gave it to everyone at Groq. Whenever we’re in a meeting and something doesn’t help with this, they can tap their coin on the table and say, “No, no, no. That’s not the way this is going to go.”
Harry Stebbings
So is everyone wrong on Founder Mode?
Jonathan Ross
I think that’s what you do when you don’t have the quality of people working for you. You need the right gearing ratio between you and your direct reports.
Harry Stebbings
It’s a really unfair question, but I have to ask it. How do you analyze Elon’s attempt to buy Twitter—not buy Twitter, to buy OpenAI?
Jonathan Ross
I was sitting at the Élysée Palace, or however you pronounce it, at the dinner with Macron, Sam Altman, and JD Vance.
Frankly, I think Elon was a little jealous that Sam Altman was sitting next to JD Vance and it wasn’t him. It was right around the time Sam Altman was speaking that Elon announced it.
I thought Sam’s tweet response was pretty good. I would have probably said, instead of whatever he said about $9 billion, “I’m going to take Twitter public at $420 a share.”
It was attention-grabbing. Some people can’t stand not getting attention. My revenge on this is to give as little attention as possible.
Harry Stebbings
Let’s move on.
What would you do if you knew you couldn’t fail?
Jonathan Ross
I would put in 100% of the orders for every single chip we could possibly manufacture, because right now the demand is unlimited.
Every time you triple, you find the same number of problems, so you have to do it judiciously. But if I knew that, no matter what problem came up, we didn’t need to be safe at all, I would say, “Great. We’re going to go build 20 million chips.”
Harry Stebbings
In 10 years, is NVIDIA 3 times bigger, 10 times bigger, or 50 times bigger?
Jonathan Ross
I think they’ll be bigger. Training will become more important. I wouldn’t be surprised if they were 3 times bigger, and I also wouldn’t be surprised if they stayed around the same.
It’s hard to tell where things are going, because a lot of assumptions in the investment in NVIDIA were that they were going to run away with the entire market, including the inference market. They just haven’t built the right thing for inference.
As a weighing machine, I do think they should increase in value. But so much of it is a popularity contest that I don’t know whether they’ll need to grow to get to where they are.
16. Future Predictions and AI's Impact on Society
They might need to grow their revenue to justify where they are, but it’s a pretty fair multiple given everything going on. I couldn’t tell you. The popularity contest skews everything.
Harry Stebbings
What’s a crazy AI prediction you have that everyone else thinks is science fiction?
Jonathan Ross
I would assume that, in the next 10 years—and I know this is going to be crazy—you’ll see a Mounjaro moment for aging.
Harry Stebbings
You saw that picture of me and my weight loss. Unbelievable.
Jonathan Ross
70 pounds.
Harry Stebbings
70 pounds. I was on Mounjaro.
Jonathan Ross
If you know anyone who’s overweight and it’s hurting their health, get them on Mounjaro as soon as you can. It works.
Harry Stebbings
What is Mounjaro?
Jonathan Ross
It’s one of those GLP inhibitors, one of the weight-loss drugs that have become popular recently. It works.
My crazy AI belief is that, if it’s possible to significantly slow or stop aging, you’ll have a Mounjaro moment in perhaps the next 10 years. It came out of nowhere. All of a sudden, you could lose weight. Something finally worked, and it’s worked for a lot of people.
I don’t know whether it’s possible to slow or stop aging. Wear and tear is a real thing, and it might be impossible. But if it is possible, then in the next 10 years we will do it, and it will be sudden. It will be like the Mounjaro moment.
Harry Stebbings
I don’t see how it’s not possible. When you look at the advances that will come in medical research, I don’t see how it’s not possible that we’ll at least extend longevity by 60 years.
Dario will live to 150. I don’t see why that’s impossible.
Jonathan Ross
I don’t either, but I also don’t know that it is possible. Until I know that, I’m going to keep the conditional in there and say, “If possible.”
Harry Stebbings
What have you changed your mind on in the last 12 months?
Jonathan Ross
This is less of a mental one and more of an emotional one. We didn’t have product-market fit for 7 years.
That’s terrible. The morale changes when you find product-market fit. The world is brighter, the birds sing, I feel like hugging people, and you sleep. Life is better.
I forget whether it was you or someone else who was talking about type 1 and type 2 happiness. I think there’s a third type.
The 2 common ones are that the present is happy, and that you went through some really terrible things but the memories make you happy. There’s past, present, and future.
As a founder, the only type of happiness you get is this third type: future happiness. When you get product-market fit, you start to get past happiness. When you start to get revenue and everything else, you get present happiness.
It changes everything.
Harry Stebbings
If you had to bet on 1 company other than Groq to define the AI era, who would it be?
Jonathan Ross
I would focus more on the companies you haven’t heard about. I don’t know what those companies are, but I can tell you what they’re going to do.
The first will be the one that solves the hallucination problem.
The second will be the one that’s best able to break down subgoals for agentic systems. I think agentic systems come after you solve the hallucination problem, because otherwise you get these long chains where you can introduce hallucinations. It will work, but it will work much better after that problem is solved.
The next one is what I call the invent stage. Right now, the way LLMs work is that they make the most probable prediction.
It’s amazing. Imagine taking an entire novel and deleting the ending. You have a detective murder mystery, and you get to the point where the detective says, “The murderer is…”
The model can actually predict the answer. It had to understand everything, but it’s going to give you the most probable answer. That’s not good for invention, and it’s not good for art.
The reason the writing from LLMs is terrible is that it’s predictable. How do you say something that’s non-obvious but obvious when you see it? We don’t even have the right word for it: non-obvious but obvious.
That is going to unlock invention.
The final one is what I call the proxy stage. Someone will make it so models can make decisions for you. You can proxy your decisions, like the decision to do this interview.
Other things had to happen: the flight had to be booked, we had to get a ride over, and other things had to be canceled. You would trust an executive assistant or a chief of staff to make those decisions, but you wouldn’t trust an LLM yet.
That’s the final stage before you get to general AI. Each company that does one of those things is going to be a defining company.
Harry Stebbings
You said we have to fix hallucinations before we get efficient agents. Does that mean money going into agentic AI will be burned?
Jonathan Ross
No. Take hallucination as an example. The examples I gave were medical diagnosis and law—2 areas that will be unlocked once we get rid of hallucinations.
But there are startups like Perplexity that are doing just fine even though there are hallucinations, because it’s not high-risk. It’s for entertainment.
If you click the links, you can check them, and it works reasonably well. It depends on how risky the industry you’re in is.
You can start trying to position for the wave early and generate something. If you’re in the right position, as we were for 7 years, the wave comes.
That money isn’t necessarily incinerated. In fact, the recent deal we just announced is more revenue than the money we’ve raised.
Harry Stebbings
How does that cash hit? I know it’s an ARR throughout the year.
Jonathan Ross
Throughout the year, yes. But it’s this year, and there could potentially be more next year.
Harry Stebbings
A lot more?
Jonathan Ross
A lot more.
Harry Stebbings
What is that contract in 3 years?
Jonathan Ross
If we sell everything we possibly can this year, it’s many billions. From the capacity alone, there are tens of billions of dollars of hardware that we could build next year through these types of deals.
We’re doing it at high volume and low margin. If we were talking about GPU sales and GPU prices, we would be talking about hundreds of billions. We’re just not charging that much.
Harry Stebbings
The thing I’m most excited about is disease discovery and drugs. My mother has MS, and it was always taught to me that it was incurable. Now it’s actually, “Maybe it’s not.”
What are you singly most excited for?
Jonathan Ross
We went from a phase where people were hardware engineers to software engineers.
Being a hardware engineer is ridiculously difficult. The training you need is significant, you have to get things right, and there’s a real expense if you get it wrong.
Becoming a software engineer is much easier. All you have to do is get some time on a machine and teach yourself. Nowadays, you can download manuals, tutorials, or whatever from the internet.
I think prompt engineering is going to unlock a huge swath of human society. There are 1.3 billion or 1.4 billion people in Africa who know how to speak.
If you gave them access to a tool that allowed them to create applications live just by speaking to it, that would be another 1.3 billion or 1.4 billion potential entrepreneurs.
There are 8 billion people on the planet, and the difference is that hardware was ridiculously difficult. It was arcane knowledge that was hard to acquire.
Software was plentiful. Language is something you already know; you don’t have to learn a thing.
Harry Stebbings
What’s that going to do for venture? What’s that going to do for entrepreneurship?
Jonathan, I love talking to you. It’s always such a broad and wide-ranging discussion. Thank you so much for putting up with me in person. I’ve loved it.
Jonathan Ross
Awesome. I’m so glad to be here.