Steeve Morin
The thing with NVIDIA is that they spend a lot of energy making you care about stuff you shouldn't care about, and they were very successful. Who gives a shit about CUDA?
OpenAI is amazing, but it's not their compute. Ultimately, if you don't own your compute, you're starting with something at your ankle. In 5 years, I would say 95% inference and 5% training. You have the products, the data, and the compute. Who has all 3? Google. That's Android, Google Docs, whatever—they have everything they can sprinkle everywhere. This is the sleeping giant in my mind.
Harry Stebbings
Steeve, dude, I am so grateful to you for joining me today. I've wanted to make this one happen for a while, but when we were discussing who would be best for this topic, I was like, "We've got to have Steeve on." Thank you for joining me today.
Steeve Morin
Well, thank you. I feel humbled. I appreciate it. Thank you.
1. How Will Inference Change and Evolve Over the Next 5 Years
Harry Stebbings
I want to start: can you just give a quick overview of ZML, and specifically your role in the infrastructure strategy today and where you sit at the very bottom of things?
Steeve Morin
ZML is an AI/ML framework that runs any model on any hardware, and it does so without compromise. We sit ultimately at the infrastructure layer. We enable anybody to run their model better, faster, and more reliably, but on any compute whatsoever. It doesn't really matter; it could be NVIDIA, AMD, a TPU, or whatever.
We do all that without compromise. That's the key point, because if there's a compromise, then it's not really agnostic, right?
Harry Stebbings
Can I ask you, then, if we think about sitting between any model and any provider, whether AMD or NVIDIA, do you think we will exist in a world where people are using multiple models simultaneously and running them concurrently?
Steeve Morin
Yes. You can actually see it; it's been happening for a while. Models now are not the right abstractions. At least if you look at closed-source models, they're not really models—they're more like backends.
There are a lot of tricks that make you feel like you're talking into 1 model, but ultimately you're talking to a constellation, an assembly of backends, that produces an AI response. Probably the number 1 obvious thing would be that if you ask a model to generate an image, then it will switch to a diffusion model, not an LLM.
There are many, many more tricks. The Turbo models at OpenAI do a lot of tricks. Definitely, models in the sense of getting weights and running them is something that is ultimately going away in favor of full-blown backends.
You feel like you're talking to a model, but ultimately you're talking to an API. The thing is, that API will be running locally—or locally meaning in your own cloud instances and so on.
Harry Stebbings
Okay, so we will have a world where we're switching between models, and there's this kind of trickery around it. Perfect. So we've got that at the top, then we've got ZML in the middle, and then you said on any hardware. Will we be using multiple hardware providers at the same time, or will we be more rigid in our hardware usage?
2. The Importance of a Top-Down Strategy for Microsoft and Google
Steeve Morin
No, absolutely. You can get probably an order of magnitude more efficiency depending on the hardware you run on. That is substantial. Not a lot of people have that problem at the moment because things are getting built as we speak, but a simple example is that if you switch from NVIDIA to AMD on a 70B model, you can get 4 times better efficiency in terms of spend. That is substantial—very much substantial.
Harry Stebbings
If there's such a cost efficiency—4 times—why does everyone not do that?
3. Challenges and Innovations in AI Hardware
Steeve Morin
There are a few reasons. Probably the most important one is the PyTorch-CUDA duo, and that's very, very hard to break. These 2 are very much intertwined.
Harry Stebbings
Can you explain to us what PyTorch is?
Steeve Morin
Yes, absolutely. PyTorch is the ML framework that people use to actually build and train models. You can do inference with it, but by far the most successful framework for training is PyTorch.
PyTorch was very much built on top of CUDA, which is NVIDIA's software. The way PyTorch works makes it ultimately very, very bound to CUDA. Of course, it runs on AMD, it runs on Apple, and so on, but there are always tens of little details that don't run exactly as you would expect. There's work involved.
Then there's also supply. Probably that's the number 1 thing. The second thing is that there are a lot of GPUs on the market, and pretty much all of them are NVIDIA.
The reason is that if you think in layers and say, "All right, I'm going to buy GPUs and sell them to folks to maybe not even do training, but just do inference," then most likely, if you look at it that way, you'll end up buying NVIDIA because everybody will want to run on NVIDIA. Nobody really knows how to do otherwise, and they've trained on NVIDIA, so they're thinking, "I can just reuse my code," and so on.
There's this self-perpetuating circle of people buying NVIDIA because they want to resell it, and people using NVIDIA because it's there. But it's by far not the most efficient platform, and arguably, even in terms of software, it's not the best software platform. Those are probably 2 of the most important reasons.
4. Nvidia's Market Position and Competitors
Harry Stebbings
Before we move on, we were chatting about NVIDIA and AMD when DeepSeek happened and the stock crash happened. Why did NVIDIA rebound, do you think, in a way that AMD didn't?
Steeve Morin
There are a lot of things, but in my opinion, there's always going to be a need for inference. It's very hard to say whether it will be worth everybody's money to do it on an H100. That is a bubble that I think will blow sometime. I'm kind of afraid of that, to be honest.
Harry Stebbings
Why do you think that's a bubble that will blow sometime? Why is that not legitimate?
Steeve Morin
Because it was built on the A100 financial model, which was at generation 0: we do training, but when it's the last generation, we do inference. It worked beautifully for the A100.
Then the H100 comes along, and inference is worth 5 times the price, but it maybe runs twice as fast in terms of inference. On training, it's a lot better, but on inference, it's maybe twice as fast. When it actually came out, it ran at the same speed as the A100.
There's a money gap that's going to have to be bridged sometime. The part that worries me is that I see amortization plans of 6 or 7 years with the GPUs as collateral, and I'm not sure how it's going to work. When they came out, they were worth 5 times the price and were only 2 times faster, so something has got to give.
Harry Stebbings
Is the speed of development trumping chip-development speeds? Is it now becoming a real problem where models are far outpacing the speed of chip deployment?
Steeve Morin
Not much, ultimately. The 2 things that could really shake the chip industry, in my opinion, are agents and reasoning.
Harry Stebbings
Why does that change things?
Steeve Morin
For agents and reasoning, you need to wait until the end of the request to get whatever it is you came for. You don't really care about the speed at which the text outputs, which is what you want in a chat. You only care about how much time it takes between the beginning of your request and the end.
That fundamentally changes the incentives from throughput-bound to latency-bound. If you're running GPUs at, let's say, 10,000 tokens per second, you very much like to do it 100 times in parallel. They can do that, but they cannot give you 10,000 tokens per second only to you, per stream, as we say.
In terms of agents or reasoning, this is exactly what you want, because you don't want to wait 50 seconds for whatever thinking is happening. Agents are the same. Those 2 things are the shot that might at least make NVIDIA change its course with respect to chips. They're not idiots, right?
Harry Stebbings
How should agents change NVIDIA's strategy?
Steeve Morin
They're bound by the latency between the beginning of the request and the end of the response. You don't want your agents to take 30 seconds to generate whatever their response will be.
Harry Stebbings
How should they change, then? How should they change their response?
Steeve Morin
Oh, you mean NVIDIA? It's hard to say because NVIDIA has a very, very vertical approach. They do more of more.
If you look at Blackwell, it's actually crazy what they did. They assembled 2 chips, but the surface was so big that the chip started to bend, which further perpetuated the problem because it then didn't make contact with the heatsink and so on.
They are very much pushing the pedal in terms of GPU scaling. The power envelope is pushed to 1,000 watts, it requires liquid cooling, and so on. They are very much operating at the limit in terms of GPU scaling.
The thing is, GPUs are a good trick for AI, but they're not built for AI. It's not a specialized chip; it's a specialization of a GPU, but it is not an AI chip.
Harry Stebbings
Forgive me for continuously asking stupid questions, but why are GPUs not built for AI? If they're not, what is better?
Steeve Morin
The way it worked is that a screen can be thought of as a matrix. If you have to render pixels on a screen, there are a lot of pixels and everything has to happen in parallel, so you don't waste time.
It turns out that matrices are a very important thing in AI. There was this trick in which we essentially tricked the GPU into believing it was doing graphics rendering, while we were actually making it do parallel work. It was called GPGPU at the time, probably 20 years ago.
It was always a cool trick—very cool and very successful, mind you—but it was not dedicated to this. The pioneers were probably, of course, Google with the TPU, which is much more advanced on the architectural level.
5. Challenges of Incremental Gains in the Market
Essentially, the way GPUs work kind of works for AI, but for LLMs that starts to crack because they're so big and there's a lot of memory transfer and so on. That's why Groq, Cerebras, and all these companies achieve very high single-stream performance: the data is right on the chip. They don't have to get it from memory, which is slow, as a GPU has to do.
There are a lot of things that ultimately make it a good trick, but not a dedicated solution per se. That said, the reason NVIDIA probably won, at least in the training space, is Mellanox—not because of the compute.
You need to run a lot of these GPUs in parallel, so the interconnect between them is ultimately what matters. How fast can they exchange data? When you do a matrix multiplication, the matrix is read hundreds of times during the multiplication, so there are a lot of transfers going on.
So far, Mellanox with InfiniBand has the best technology. That's why a lot of people use it. When you do training, by the way, it is the name of the game: the interconnect. When you do inference, not so much. You don't care as much.
6. The Economics of AI Compute
Harry Stebbings
Before we move to inference, I do want us to stay on chips and ask: we have TPUs, NVIDIA, and AMD. In terms of the distribution of gains, is this a winner-take-all market? Is it the cloud, where you have several providers who are dominant? What does the distribution of gains look like in the chip market?
Steeve Morin
I would divide it into 3 categories: the GPUs you can buy or rent, the TPUs you can rent, and the TPUs you can buy. This is how the market is structured today.
Right now, if you want to go dedicated, there are at least 2 options: TPUs and Trainium—TPUs on Google and Trainium on Amazon. These are available chips; you can rent them today if you want.
If you want to buy or rent GPUs, there are GPUs everywhere—we know that all the time. There's this new way of computing, which is dedicated chips that you can actually buy: Tenstorrent, Etched, and likely VSORA.
I think it will be a mix of whatever you can get. For instance, if you're in Google Cloud, of course you don't want to use NVIDIA because you get ripped off. Here's the dirty secret: NVIDIA, like TSMC, sells to you at a 60% margin. NVIDIA sells to you at around a 90% margin, and on top of that there's Amazon, which takes, let's say, a 30% margin. So you are a very thin crust on a very big cake.
That's why, to me, it's a bit of a losing game if you go all-in on 1 provider. You want optionality, with increasing competitiveness within each of those layers.
Harry Stebbings
Do we not see margin reduction?
Steeve Morin
Absolutely, yes. But here's the problem. Let's say you're on Google Cloud and you're on TPUs. Suddenly, you remove that 90% chunk from the spend.
The problem is that, for multiple software reasons—which we are solving at ZML—they're not really a commercial success. They're very successful inside Google, but not much outside of Google. Amazon is pushing very, very hard for its Trainium chips as well.
7. Training vs. Inference: Infrastructure Needs
I would say the future I see is that you use whatever your provider has, because you don't want to pay an outrageous 90% margin and try to make a profit out of that.
Harry Stebbings
I totally get you. When we move to inference and training, I think everyone's focused so much on training. I'd love to understand the fundamental differences in infrastructure needs when we think about training versus inference.
Steeve Morin
These 2 obey fundamentally different tectonic forces, if you will. In training, more is better. You want more of everything, essentially, and the recipe for success is speed of iteration. You change things, see how they work, and do it again. Hopefully, it converges. It's like changing the wheel of a moving car, so to speak.
On inference, this is the complete reverse: less is better. You want fewer headaches. You don't want to be waking up at night because inference is production. You could say that training is research and inference is production, and it's fundamentally different in terms of infrastructure.
The number 1 difference between the 2 is the need for interconnect. If you're doing production and you can avoid having an interconnect between, let's say, a cluster of GPUs, of course you will avoid that.
This is why models have the sizes they have: so people can run them without needing to connect multiple machines together. It's very constraining in terms of the environment. That is probably the fundamental difference—the need for interconnect.
Number 2 is: do you really care what your model is running on, as long as it's outputting whatever you want it to output?
Harry Stebbings
Can you help me understand why training is more is more, while in inference less is more? Why do we have that difference?
Steeve Morin
Think of it like doing 1 painting versus doing 1 million paintings. The tools you use and the process you follow will be different. If you do 1 painting, what you favor is the speed at which you can make a stroke and iterate.
If you do 1 million paintings, what you want is a process—a process that is reliable and can deliver 1 million paintings efficiently. That's the same for training versus inference.
Harry Stebbings
How do people then put inference into production? We've seen with training that NVIDIA has dominated so heavily. How do people put inference into production?
Steeve Morin
There's a lot of duct tape. One of the problems is that training, from first principles, is actually 2 passes: forward and backward. It's called the forward pass and the backward pass. Inference is running only the forward pass. That's how things are today, mostly.
There are people who are trying to specialize a bit because, at some point, duct tape doesn't really work out. When you're at big scale, that creates a problem. It's a problem that's growing because a lot of people are coming to the market with needs for inference that weren't there a year or a year and a half ago.
OpenAI had this problem, and maybe Anthropic had this problem, but it wasn't a universal problem yet. Now it's becoming one.
Harry Stebbings
Can you articulate what problem OpenAI and Anthropic had with regard to inference?
Steeve Morin
For instance, probably the number 1 thing, depending on how you deploy, is what's called autoscaling. As your systems get more and more loaded, you want to provision compute because these things are tremendously expensive.
You don't want to say, "I have 1,000 GPUs 24 hours a day, and even if nobody is in production, I will pay for them," which, mind you, is what people are doing today. This is crazy.
What you want to do is provision compute as your needs grow. You want to scale up and scale down. That's probably the number 1 thing that gives you a lot of efficiency in terms of spend. We're talking about multiples—sometimes a 5x or even 10x improvement.
In regular backend engineering, this is a problem everybody knows. Everybody is doing it because the savings are so huge. But in AI, nobody really had that problem. Now they're coming up against it.
Harry Stebbings
So the problem is that they're not doing provisioning. They're paying a ton more because they're fully in production all the time, versus provisioning as needed?
Steeve Morin
That's 1 example. Another one is choosing the right compute. It's a vicious circle because provisioning compute is very hard. If you lose compute, it's very bad, so you're essentially incentivized to overbuy.
In the case of Amazon or Google, that would mean buying reserved compute, which you're not going to use, because if you buy it on demand, you will get tremendously ripped off. That creates this scarcity of compute because people buy preemptively, pay a ton of money, and don't use it.
Harry Stebbings
When you buy compute preemptively, does it not become outdated by the time you use it?
Steeve Morin
It might well be. Judging by the pace of development, it might well be. We are being spared a bit because Blackwell is late and H100s are getting canceled, so the H series is still active.
But yes, absolutely. What choice do you have?
Harry Stebbings
We have a moment in time where there's this massive overhang, or oversupply, of compute that we've proactively bought ahead of time. But then the hyperscalers say, "We'd rather just burn it, buy fresh, and we have the money to do that."
Steeve Morin
I think it already started. I'm getting cold emails for discounts from services I've never heard about. I started getting these emails probably around October or November, so some people are left with a lot of CapEx that they don't know what to do with.
It's very hard to build a cluster and run a training run. That's a different thing from literally building a cloud provider, hyperscaler, or whatever you want to call it.
There are a lot of people who do their training runs on the regular providers, but then move to a regular hyperscaler when they go into production. I very much worry there will be an oversupply of these chips.
8. The Future of AI Chips and Market Dynamics
The problem is that the chips are the collateral. Somewhere in the United States or wherever, there could be a data center with 1,000 GPUs that people may buy for 30 cents on the dollar. I don't know, but this is what might happen.
Harry Stebbings
What's the time frame for that?
Steeve Morin
Probably this year.
Harry Stebbings
Jensen has made it very clear that inference opens up more revenue opportunity for NVIDIA. He said that 40% of its revenue today comes from inference. To what extent is that correct? Or, as Jonathan Ross at Groq said on the show, is NVIDIA not meant for inference and that market won't be won by NVIDIA?
Steeve Morin
Technically speaking, he's right, but realistically speaking, I'm not sure I agree. These chips are on the market. They're here. I can open a tab in Chrome and get one. Availability is something I don't take lightly.
I think NVIDIA is here to stay, at least if the H100 bubble doesn't burst. These chips are going to be on the market, and people will buy them and do inference with them. What remains to be seen is the OpEx, electricity, and so on, but that is a complicated question.
As far as I know, the only chips that are really frontier in that sense are probably TPUs and the upcoming chips. They're great chips, but they're not on the market, or they're available at outrageous prices—millions of dollars to run a model.
Harry Stebbings
What chips are great, and why aren't they on the market?
Steeve Morin
If you look at, for instance, Cerebras, it's incredible technology and incredibly expensive. How will the market value the premium of having single-stream, very high tokens per second? There is value in that, as we saw with Mistral and Perplexity, but I'm not sure that was done profitably. I don't know the details, but I think Cerebras put it out at a loss.
Today, there are 3 actors on the market that can deliver this. I think this will be the pushing force for change in the inference landscape: agents and reasoning. That means very high tokens per second only for you, not for an aggregate of people.
Harry Stebbings
What is forcing the price of a Cerebras to be so high? You heard Jonathan at Groq on the show say that they're 80% cheaper than NVIDIA.
Steeve Morin
There's a trick, because there's no magic. This little trick is called SRAM. SRAM is memory directly on the chip, so it's very, very fast memory.
Here's the problem with SRAM: it consumes surface area on the chip, which makes it a bigger chip and makes yield very difficult because the chances of problems are higher.
SRAM is very, very fast memory, which gives you a lot of advantage when you do very-high-speed inference, but it's terribly expensive. If you look at Gro, on this generation they have 230 MB of SRAM per chip. A 70B model in BF16 is 140 GB, so you can do the math.
Cerebras has 44 GB of SRAM in what they call their Wafer-Scale Engine, which is a chip the size of a wafer. Most likely it's interconnected, but it's huge, and it has to be water-cooled. They have copper needles that touch the chip to cool it. It's crazy stuff—very impressive technology, mind you, but very, very expensive.
My bet is that there will be chips on the market that do that at a much lower price. Two companies I see going in that direction are Etched and likely VSORA. If you can deliver this at a price comparable to GPUs, you've won.
Harry Stebbings
Is minimizing SRAM the only way to reduce the unit cost on these chips?
Steeve Morin
It's hard to say. You need some SRAM, but if you can have a smaller process node and hook yourself up with external memory, then yes, you can do that a lot better.
But if you go full-blown SRAM, there's no magic: you will have to pay the price.
Harry Stebbings
I'm so enjoying this. I'm also learning. My notes here are just expanding by the day. How do you think the inference market evolves over the next 3 to 5 years, pushed by reasoning?
Steeve Morin
Reasoning—not in the sense that you see in DeepSeek and whatever, but reasoning and what's called latent-space reasoning. Latent-space reasoning and agents will push the market toward different types of compute.
Harry Stebbings
Can I just ask, what is latent-space reasoning?
Steeve Morin
The way models reason today is in tokens. It's as if, when you think to yourself, you say out loud what you're thinking.
It works, but it's a bit inefficient, and you lose information by doing this. Latent-space reasoning is reasoning without going into English or whatever language you're using—staying in what's called the latent space, which is where all the information of an LLM or an LM lives.
This is very much how we work as humans. We are moving toward what likely Yann LeCun calls energy-based models, in which we have different types of longer or shorter thinking times, if you will.
That fundamentally cannot be delivered by GPUs at scale.
Harry Stebbings
Why can't GPUs deliver it?
Steeve Morin
Because access to external memory prevents it. HBM is all the rage, but compared with SRAM, HBM is absolutely dirt slow. That's the problem you get. HBM is the best we can do, but it's still slow versus SRAM.
Harry Stebbings
When I had Jonathan on, he said that NVIDIA has such a stronghold because it's one of the only buyers of HBM, which gives it this unique position. Is being the sole buyer of HBM irrelevant if the world needs SRAM instead?
Steeve Morin
No, you want HBM, to be clear. SRAM will not deliver this. It's a dead end in terms of scaling. SRAM means consuming the surface area, which creates yield problems and makes everything explode.
You need some SRAM, so we will have bigger amounts of SRAM in chips and, of course, bigger amounts of external memory in chips. The issue with HBM is that it's still slow. Maybe NVIDIA has a stronghold and can prevent you from getting some, so I call it the Nutella situation: Nutella owns 80% of the hazelnut market. You can build a competitor, but who will you buy the nuts from?
This is a bit above my league, but there will be a need for HBM and a need for SRAM. Better, more dedicated architectures will be able to deliver these things. Then there's the next frontier after that, which is called compute-in-memory.
There are 2 companies that I know of, at least, that are in that market. One is called Rain AI.
Likely Sam Altman is one of the investors, so there’s no surprise. The other one is called Fractile; I think it’s in the UK, actually. This is the next frontier. The idea is that instead of transferring the data between external memory and the CPU and doing the compute there, you bring the CPU to the memory and do everything. It’s crazy stuff, but it’s coming—maybe not this year.
Harry Stebbings
How does that change the situation?
Steeve Morin
It makes it much more efficient. What does that actually mean in reality? It means you get maybe not SRAM-level performance, but a lot faster performance in terms of compute. If you translate that to LLMs, let’s say you get much, much higher tokens per second in a single stream, which is exactly what you want when you go into reasoning.
You want your model to maybe think for half a second and then, boom, produce the answer. You don’t want to wait 50 seconds and context-switch to some other thing, which is the problem everybody has today, mind you. I think inference will be major. At least, this is my thesis: the compute landscape will be pushed to change because of these 2 constraints.
Harry Stebbings
If you were to ascribe value between training and inference out of a pool of 100, is it 80 inference and 20 training? What does that look like in 5 years?
Steeve Morin
I would say 95% inference and 5% training.
Harry Stebbings
Wow. Do you think NVIDIA owns both of those markets in 5 years’ time?
Steeve Morin
It depends on the supply. I think there’s a shot that they don’t, because here’s the thing: even if we take the same amount of compute, let’s imagine we have a new chip from Amazon that’s the same amount—wait, we do. It’s called Trainium. Why would I pay a 90% margin to NVIDIA if I can freely change to Trainium?
My old production runs on NVIDIA GPUs anyway. If you’re in production on dedicated chips, of course, there’s the issue of commoditization. If I’m on AWS, I can just click and, boom, it runs on AWS’s chips. Who cares? I just run my model like I did 2 minutes ago.
Harry Stebbings
With that realization, do you think we’ll see NVIDIA move up the stack and also move into the cloud and models?
Steeve Morin
They are. They have a product called NIM that sort of does that. The thing with NVIDIA is that they spend a lot of energy making you care about stuff you shouldn’t care about, and they were very successful.
Who gives a shit about CUDA? I’m sorry, but I don’t want to care about that. I want to do my stuff. NVIDIA got me into saying, “Hey, you should care about this,” because there was nothing else on the market. That’s not true, but ultimately, this is the GPU I have in my machine, so off I go.
If tomorrow that changes, why would I pay a 90% margin on my compute? That’s insane. This is why I believe it ultimately goes through the software, because the software is my entry point to the ecosystem. If the software abstracts away those idiosyncrasies, as it does on CPUs, then the providers will compete on specs and not on fake moats or circumstantial moats.
This is where I think the market is going. Of course, there’s the availability problem. If you piss off Jensen, you might need to kiss the ring to get back in line. Ultimately, I don’t see this as being sustainable.
Harry Stebbings
Can I ask, when we chatted before, you said something about AMD? I said, “Hey, I bought NVIDIA and AMD. NVIDIA—thanks, Jensen—I’ve made a ton of money. With AMD, I think I’m up 1%, versus the 20% gain I’ve had on NVIDIA.” You said that AMD had basically sold everything to Microsoft and Meta and had a go-to-market problem. Can you unpack that for me?
Steeve Morin
All chipmakers have a go-to-market problem. It’s all of them, whether it’s Google, AMD, or Tenstorrent. There are probably 2 fundamental problems.
The number 1 problem is that maintaining multiple stacks today is very, very hard. Let’s say I buy AMD. That means I’m going to abandon NVIDIA. Then I think, “Oh, crap. I have a 6-year amortization plan on that. What do I do?” Do I need to support both stacks? Maybe, until AMD tells me, “Hey, you have, I don’t know, 1,000 NVIDIA GPUs, and you’re about to buy 100,000 AMD GPUs.” I mean, come on. Then I’m like, “Okay, that makes it worth my while.”
Ultimately, the fundamental problem is that the stakes are very high. I need to have a lot of incentives to buy into that ecosystem, so I need to buy a lot of them. If you’re AMD, that is already a problem. Then Microsoft comes along and buys it all, which, by the way, puts OpenAI—or at least the inference side of OpenAI—in the green because of the efficiency gains.
Harry Stebbings
Can I just try to understand? Are you saying the switching costs are really high from one provider to another, or are you saying that to get into one of these buying processes, you have to buy so much that it prohibits you?
Steeve Morin
It’s actually both. The buy-in is very high, so to make it worth it, you have to buy a lot. If you buy a lot, this is what everybody asks. We talk to all of them, and they always have the same questions. It’s completely understandable. They say, “This is great, great, but who’s the customer?”
Take Amazon, for instance, with Trainium. Anthropic just came and said, “Hey, we’re going to buy 100,000 of them.” You want to buy 10,000 and feel like the big shot? But go back in the queue, because Anthropic is before you. They have to have very high commitments to make it worthwhile.
You cannot be incrementally better. It’s very hard. You have to convince a lot of people. I can give you 1 metric, if you want. I know for a fact that being 7 times better in whatever metric you want—whether it’s spend or whatever—is not enough to get people to switch. People will choose nothing over something. I have stories. This is a very hard market to enter because you cannot compete on incremental gains.
9. The Zero Buy-In Strategy
Maybe you can go the Middle East route, where they sprinkle everything around and evaluate everything, but that’s not a very sustainable strategy in the long term, or at least in the medium term.
Harry Stebbings
What is the right sustainable strategy, then? You don’t want to go so heavy that you can never get out and have those switching costs, but you also don’t want to sprinkle it around. What’s the right approach?
Steeve Morin
The right approach, to me, is making the buy-in zero. If the buy-in is zero, you don’t worry about this. You just buy whatever is best today.
Harry Stebbings
How do you do that?
10. Switching Between Compute Providers
Steeve Morin
By renting. This is our overall promise—at least, ZML’s thesis: if the buy-in is zero, then you completely unlock that value. When you say the buy-in is zero, what does that actually mean? It means that you can freely switch compute to compute. You just say, “Now it’s AMD,” and boom, it runs. You say, “Now it’s Tenstorrent,” and boom, it runs.
Harry Stebbings
How do you do that? Do you have agreements with all the different providers?
Steeve Morin
Yeah, but not agreements. We work with them to support their chips. As a user myself of AI technology, my view is that if it’s free for me to switch or choose whichever compute provider I want—AMD, NVIDIA, whatever—then I can take whatever is best today and whatever is best tomorrow. I can run both. I can run 3 different platforms at the same time. I don’t care. I only run what’s good at the moment.
That unlocks a very cool thing for me, which is incremental improvement. If you’re 30% better, I’ll switch to you.
Harry Stebbings
Are you taking the risk on that hardware, then? If you’re the one providing on-and-off, on-demand provisioning, you name it, who takes the risk?
Steeve Morin
This is a great question. I think that if you’re doing it bottom-up, from infrastructure to applications, you’ll lose, because nobody will care, as they don’t today. If you look at TPUs, they’re available and they’re great, but nobody cares.
Harry Stebbings
Why does nobody care about TPUs?
Steeve Morin
Because the cost of buying in is always the same. You have to spend 6 months of engineering to switch to TPUs. Mind you, TPUs do training. They’re the only ones with training now, but AMD can do training. In terms of maturity, by far the most mature software and compute is TPUs, and then it’s NVIDIA.
The buy-in is so high that people say, “Well, we’ll see. I’m not on Google Cloud. I have to just sign up.” These are tremendous chips and tremendous assets, but people don’t want to make that commitment.
11. Microsoft's Strategy with AMD
In terms of the risk, I think if you want to do it, you have to do it top-down. You have to start with whatever it is you’re going to build and then permeate downwards into the infrastructure.
Take Microsoft with OpenAI, for example. They just bought all of AMD’s supply and run ChatGPT on it. That’s it. That puts them in the green. That’s actually what makes them profitable on inference, or at least means they’re not losing money.
Harry Stebbings
I’m sorry, how does Microsoft buying all of AMD’s supply make them not lose money on inference? Help me understand that.
Steeve Morin
I can give you actual numbers. If you run 8 H100s, you can put 2 70B models on them because of the RAM. That’s number 1. Number 2, if you go from 1 GPU to 2, you don’t get twice the performance. Maybe you get 10% better performance.
That’s the dirty secret nobody talks about on inference. You go from, let’s say, 100 to 110 by doubling the number of GPUs. That is insane. Would you rather have 2 machines with 1 GPU each than 1 machine with 2 GPUs?
With 1 machine of 8 H100s, you can run 2 70B models if you do 4 GPUs and 4 GPUs. If you run on AMD, there’s enough memory inside the GPU to run 1 model per GPU, so you get 8 GPUs’ worth of throughput. On the other hand, you get AMD GPUs at 2 to 2.5 times the throughput. That’s a 4× right there, just by virtue of this.
That’s the compute part. If you look at all of these things, there are tremendous amounts of memory. We talk to companies whose chips are coming with almost 300 GB of memory on them. That means 1 chip per model, which is the best thing you want if you’re running 70B models.
It’s not state of the art, but it’s the regular stuff people will use for serving. If you look top-down at what you’re going to build with them, it’s a lot better to go after the efficiency gains, because 4× is a big deal. These chips are 30% cheaper than NVIDIA’s. It’s a no-brainer.
But if you go bottom-up and say, “I’m going to rent them out,” nobody will rent them. That’s why I think it’s a good way to attack it from the software, because ultimately, do you really care whether your MacBook is an M2 or an M3? You just say, “Oh, it’s the better one,” and that’s it. Imagine if you had to care about these things. That would be insane.
Harry Stebbings
When I listen to you now, I’m thinking, “I should sell my NVIDIA and buy more AMD.” If you were forced to buy 1—I’m not saying sell the other, I’m saying buy 1—which would you buy, and why? I mean the stock.
Steeve Morin
The thing is, I’m long. I used to think the market was efficient, so this is not investment advice. Probably, today, I would still go with NVIDIA because of the supply. If we play our cards right and ship our stuff, hopefully I’ll come back and tell you to buy as much AMD as you can—or Tenstorrent, if they go public—or whoever else. These chips are amazing, by the way.
Harry Stebbings
What does everyone think they know about inference that they actually don’t?
Steeve Morin
Probably not a lot of people are accustomed to what it entails to run production. Inference is production, and production means somebody has to wake up at night. I used to be that guy. I don’t want to do it again.
Production is hard. Thankfully, we have a lot of software nowadays to do that a lot better, but there’s not a lot of reuse because the AI field, at least, isn’t really accustomed to that yet. It’s changing, but the discussions I had 1 year ago and the discussions I have today are not the same. They’re going in the right direction, but they’re not there exactly yet.
That would probably be the number 1 thing. Inference isn’t just training code running with only the forward pass. That’s not what it is.
12. Data Center Investments and Training
Harry Stebbings
Can I ask how you evaluate the data center investment we’re seeing? When you look at Meta doing $60 billion to $65 billion, Microsoft doing $80 billion, and some of the intense capital expenditure you’re seeing, how do you think about that on the data center side?
Steeve Morin
They’re still going after training, so there’s still this frontier. It’s probably also why NVIDIA is the better buy right now. On the NVIDIA side, if you do training, it’s incremental. If you’ve bought 1,000 NVIDIA GPUs and buy 1,000 new NVIDIA GPUs, that gives you 2,000 GPUs. But if you buy 1,000 NVIDIA and 1,000 AMD, that gives you twice 1,000. It’s a bit different.
They’re still going after training, definitely, and they’re very pragmatic in doing so. They have the capital expenditure to spend. They’re not making their money out of it.
Google is probably the only one, by the way, that owns its compute. There’s this triangle of winning—that’s my mental model. You have the product, the data, and the compute. Who has all 3?
Google. Amazon doesn’t have products. They have Amazon and AWS, but they don’t have actual products. Google has Android, Google Docs, and whatever else. They have everything, and they can sprinkle it everywhere. This is the sleeping giant in my mind. If they’re not busy doing a reorganization, they might—
Harry Stebbings
It’s fascinating, because if you’re a shallow thinker, you think OpenAI challenges the golden goose, which is search, and that Google is threatened more than ever now. OpenAI is amazing, but it’s not their compute. It’s Microsoft’s compute. If you own your compute, you own your margin, is essentially what you’re saying.
Steeve Morin
Yes. Even Microsoft, when they were running on NVIDIA, bought NVIDIA at outrageous margins. I talk to a lot of people who build data centers. These people buy tens of thousands of GPUs, and I ask them, “Do you at least get a discount or something?” They say, “No. The only thing we get is access to supply.”
13. How to Succeed in AI: The Triangle of Products, Data, and Compute
Ultimately, if you don’t own your compute, you’re starting with something around your ankle. This is why I like to think in this triptych, or at least this triangle: product, data, and compute. You can see where everybody sits and their weaknesses and strengths.
14. Scaling Laws and Model Efficiency
Harry Stebbings
Can I ask you, if we move a little bit—you said it’s totally rational that everyone’s focusing on training still. When we think about that, it’s rational if you think that efficiency and scaling laws continue to place such emphasis on it. How do you think about model scaling and scaling methods coming into play?
Steeve Morin
There’s a brute-force approach to this. It’s a very American approach: more and more and more. But you look at, for instance, the xAI cluster, and it’s not 100,000 GPUs. It’s 4 times 25,000.
You’re starting to see that because of InfiniBand and, in the case of likely RoCE, the technology they use to bridge their GPUs together, you have a network bound. At some point, you’re fighting physics. You can push, but it’s like trying to get to the speed of light. As you approach it, the amount of energy you need gets higher and higher, and it grows and grows.
There are 2 approaches to that—or, sorry, 2 counters to that. The first is that we still scale, but there’s a lot of waste and excess spending on the engineering side. That’s the DeepSeek approach, which was very successful. They said, “If we do this and this differently, we get multiples sometimes.” Virtually, you increase your compute capacity because you’re more efficient.
15. Future of AI Models and Architectures
The other approach is likely Yann LeCun’s approach, which is that this is not scaling, and at some point we need to look the problem in the face and do something better. Of course, we push and push and push because there’s still capital, but between these 2 approaches, I think you can do more with less.
Harry Stebbings
At what point do we stop and say, “There’s a lot of wastage, and we could do this better”? How far away are we from that?
Steeve Morin
Until somebody does it. DeepSeek was a good wake-up call. Suddenly, efficiency is in. That’s number 1.
Number 2 is until there’s a new architecture that comes out and changes the game. In the case of LLMs, for instance, you have what are called non-transformer models, which fundamentally change the compute requirements. That might be a frontier that completely obsoletes the transformers.
The transformers are the building blocks by which current models work. The way they work is that for each token, or syllable, if you will, the model looks at everything behind it. You can see that as you add more text, you have more work to do.
There are new architectures that don’t require this. That might change these things and probably shift the amount of compute needed to do training or inference. Then there’s the new thing, which is likely Yann LeCun’s thesis: the world model. LLMs are, at the end of the day, what we need is something that fundamentally understands the world.
Harry Stebbings
What is it called?
Steeve Morin
It’s likely JEPA. I’m very bullish on this, but it’s very frontier.
Harry Stebbings
Why are you bullish on it, and why is it so frontier?
Steeve Morin
It’s hard to explain. He explained to me how it worked, and I was blown away. It’s as simple as this, but it makes a lot of sense.
We’re creeped out because the machine talks back to us. That’s it. It’s not a new thing. When it exploded, it was new technology, but suddenly it was talking back, and that freaked us out, so we got crazy about it. Language is 1 form of communication, but it’s ultimately a very narrow window into the world.
We use language to describe the world, arguably with some loss. The JEPA approach, long story short, is that you have essentially 2 things you want to do, and you try to minimize the energy to do them. From this, understanding emerges. Physics emerges, and so on, because you’re trying to minimize the amount of energy to go from 1 state to another.
That actually makes sense. If you try to pick up this AirPods case, you’re not going to take a round trip around the block to get it. You just get it. In my brain, I’m wired to do the thing. If I go and talk to myself out loud—put my hand down, move to the left, and whatever—that feels very inefficient.
This will probably be something that changes. In the case of LLMs, there’s also good work on what are called diffusion-based LLMs. Instead of thinking autoregressively—that means you get a new token, reinject it, redo it, and so on—they think more like we do, which is in patches. Imagine a paragraph of text, and words appear until it’s done.
Harry Stebbings
Wasn’t that distillation what DeepSeek did on OpenAI models, basically copying them?
Steeve Morin
Oh, distillation. They did distillation on OpenAI models.
Harry Stebbings
If we’re all progressively moving toward a better future for humanity and more efficient models, is distillation not effectively open source in another wrapper?
Steeve Morin
I think it’s fair game, to be honest. There were some people who tried to ask—I don’t remember if it was an OpenAI model—but they asked it to generate an image from a Star Wars movie at a particular timestamp, and it came out with a screenshot from the Star Wars movie. Obviously, it was trained with it.
I think it’s fair game because there’s no free lunch. It was trained with data. You had a good ride, somebody was sneaky and took it, but you took it from the beginning, too. Let’s just accept that it’s fair game.
Harry Stebbings
You also learn from their advancements.
Steeve Morin
Absolutely. I take my cup and enjoy it very much, every single day.
Harry Stebbings
You mentioned training there. Obviously, data quality dictates a lot of training ability. When you think about the future of the data that feeds into training, how do you think about the balance between synthetic data and real data?
Steeve Morin
I’m a bit split on this. There’s a part of me that says if you reinject data into the system, the system deteriorates. That feels intuitive. But if you look at AlphaGo, for instance, the moment it ramped up its skills was when they started generating synthetic games.
I’m a bit split, but there are some verticals that benefit very much from this. Code LLMs, for instance—we can run code. That’s the Poolside thesis.
Harry Stebbings
Why does it work for coding and not for other things?
Steeve Morin
You don’t use the AI model to generate output. You use the machine. You run the code, see what it makes, run all this code, and create data out of it.
Whereas if you run an LLM and say, “Generate me 2 trillion tokens of text,” it will do it with its own output. You may reinject that data and stuff, so there are a lot of tricks. Ultimately, my gut tells me it feels wrong because you reinject data that was already there, so it will deteriorate. There’s loss.
I’m a bit bullish, but I’m not sure exactly on which verticals. Code is one. We’ll see. Distillation is, in some sense, a bit like that. You distill and create synthetic data from a bigger model into a smaller one.
The most mind-blowing thing about distillation is that sometimes the smaller models become better than the bigger models through distillation.
Harry Stebbings
I love that smaller models become better than bigger models purely because of the quality of the data that’s input into them.
Steeve Morin
One theory is that the smarter model is better at generating output that you would want it to generate. It’s not better in the general sense. It’s better at the task at which you were measuring it, because that’s what it learned to imitate.
Harry Stebbings
How do you think about the future in terms of large, monolithic models versus more dynamic architectures?
Steeve Morin
Sometimes it’s wasteful to run big models. A lot of the time, it’s actually wasteful to run big models. I think there are going to be a lot of smaller models for efficiency reasons. Less is better.
There’s a but, though. You talk to people at DeepMind, and they don’t even fine-tune anymore because they have such a big context window, which is the data you inject into the model at runtime. Nowadays, they just dump data into it and say, “Do whatever that data tells you to do,” instead of fine-tuning as we used to do.
If the efficiency gains aren’t there yet, then we’re not there yet. But if the efficiency gains pass that threshold, we’ll just do it at runtime. We’ll have 1 great model that specializes itself for each request. That’s not for tomorrow, I think.
Harry Stebbings
What is retrieval-augmented generation?
Steeve Morin
First, it’s a very clever trick. What you do is represent knowledge in what’s called vector space, or latent space. Imagine you have a 3D space that represents all knowledge, all of everything. A cat sits here, and a dog sits close because it’s an animal, but it’s far from some other property, and so on.
You run the user’s request through this same system. It’s called an embedding, and that gives you a vector. You take whatever is closer to you—that’s what’s called semantically close. Then you insert those pieces of text before the request.
It’s as if you said, “Knowing the following”—and you give the data, let’s say it’s law or whatever—“please answer my request.” That’s it. It’s a clever trick.
It’s a bit dirty because you’re limited by the amount of data you can input. There’s a problem around how you chunk the data that you input.
Harry Stebbings
In a lot of the things we do, when we say, “Here’s a link; summarize it to the key points,” is that not retrieval-augmented generation? We’re inputting the data, and then it’s doing that.
Steeve Morin
Sometimes it is, yes. It depends on how it works. Think of it as a preamble to your question: “Knowing the following,” and the following is a tiny window into the content, “please answer my question.” Of course, as you talk more and more, it will forget, because that window is fixed.
Harry Stebbings
How does that shift the movement from large, generalized models to smaller, more advanced models?
Steeve Morin
What pushes smaller models is efficiency—roughly, speed. Less is better. If we can do the same thing with less, then less it is. It’s as simple as that.
In terms of RAG, the key frontier is what we call attention-level search. This is something we’re working on. You have the exclusive now; I’m putting it out there. It doesn’t push model sizes. What really pushes model sizes is efficiency rather than specialization.
Harry Stebbings
So, if you can get the same performance with a smaller model that’s fine-tuned with RAG or whatever, you’ll use the smaller one because, again, less is better.
Before we move into a quick-fire round, I do want to ask you: when we had DeepSeek, to what extent were you surprised that such innovation, I would argue—and I think many would agree with me—came from a Chinese competitor and not from a Western competitor?
Steeve Morin
I love it. Constraint is the mother of innovation. We can troll a bit about the Singapore gray market and all of these things, but ultimately, they had no choice.
Here’s the thing: if you can buy more, why would you give a shit? You can just buy more. If you’re pushed toward efficiency, then you will deliver efficiency. These are very, very skilled people.
The coolest thing to me about AI is that geography doesn’t matter anymore. You can just do things, appear out of nowhere, and, boom, you’re on the map. I’m very glad that they did it. I found the reaction very entertaining, to be honest. Constraint is a very good driver of efficiency.
16. Why OpenAI’s Position is Not as Strong as People Think
Harry Stebbings
Do you think it’s a meaningful threat to OpenAI and ChatGPT? Bluntly, they still have the consumer loyalty and the consumer brand. To what extent is it actually a long-term threat?
Steeve Morin
I’m not sure who’s a threat to OpenAI at the moment. Here’s why: you look at the numbers, but we live in a bubble. We follow every new episode, every new model, who said what, and so on.
I go to my mother and ask her, “Do you know ChatGPT?” She says yes. Then I ask, “Do you know some other model?” I don’t want to dunk on anybody, but she says, “What is it?” Even Gemini—Google, right? They have a strong brand and a strong product.
Arguably, there’s a balance between the product and the models. Gary from Fluidstack told me that his mental model for model providers is that they’ll be like carmakers. There’s no winner-take-all. Everybody will have their own because, ultimately, human knowledge is—everybody has everything. We’re converging, but I like that analogy.
DeepSeek made very good waves, but the waves were amplified by the media, the narrative, and the drama.
Harry Stebbings
Do you think export regulations inhibit China’s ability to compete in any way today, or tomorrow?
Steeve Morin
Maybe tomorrow. I’m not sure. They’re a bit late in terms of ASICs. They’re at the A100 level, but one of their unfair advantages is that they’re exercising under constraint. It’s like when you exercise in water. That’s their state.
They’re bound to do better. They just can’t buy their way into better compute. I think it hinders their success, but I think it’s short-term to think that way.
Harry Stebbings
Are you fearful that Europe is going to regulate itself into constraints in a world of AI?
Steeve Morin
No, I don’t care. I have ZML. This makes me wonder sometimes. I understand the narrative and so on, but I’m absolutely not fearful. Let’s be successful first, and then we’ll talk about the politics.
I’m not Mistral. I’m not building gigawatt data centers and so on. If you build gigawatt data centers, you run into these problems. But if you’re successful, everything flows from there.
Harry Stebbings
Steeve, I’m being direct here, but I’m asking you for the pros. Everyone says Mistral just doesn’t have enough money to compete. That’s the word on the street. To what extent is that fair?
Steeve Morin
They’re very competent. I don’t know. It’s easy to spread FUD. There’s a lot of FUD going around, especially about regulation and everything.
Here’s the thing: I look around me, and I don’t see what I read. I’m hardly convinced. Everybody was saying they were dead, and then, boom, they came out with their release, and it was insane.
I don’t know. What I know is that I hope they don’t have too much money, that’s for sure. You want to be clear-eyed, right?
Harry Stebbings
Final one before we do a quick-fire round. Stargate was a $500 billion announcement. How did you evaluate that?
Steeve Morin
My first impression was that I didn’t buy it. It’s American style: you start with the claim and figure it out later. I don’t buy it. Ultimately, I’m not sure I care that much about it.
Let’s imagine it’s true. I don’t know if it is or not, to be clear, but let’s imagine it’s true. Congratulations, amazing. But it’s more of the same. It’s vertical scaling.
As you know, my days are spent on efficiency, so I look at these things and think, “All right, this is a bigger American car of AI. It’s big and consumes a lot of gas, but ultimately it’s not a good car,” or at least not to my liking.
There has to be sufficient capital, but at some point, I’m not sure it’s really a differentiator. That was my thesis before DeepSeek came, and it’s still my thesis. You need money, infrastructure, and so on, but the 2 limiting factors today are probably talent and energy. That’s it.
The rest—you can buy $500 billion worth of GPUs, by the way, at a 90% margin. If we work on that margin, we can shrink that number. That’s my view of the world. I’m probably wrong, but I’m not easily entertained by these numbers. I’ve seen how the sausage is made way too many times.
Harry Stebbings
I want to do a quick-fire with you. I’ll say a short statement, and you give me your immediate thoughts. If you had to bet on 1 major shift in AI infrastructure over the next 5 years, what would it be?
Steeve Morin
Latency in reasoning, definitely this year. The shift from throughput to latency—how quickly my answer starts to appear and how long it takes for my answer to complete—is probably 1 of the fundamental shifts this year.
Longer term, I’m really rooting for non-transformer models, which will also change the compute landscape, and, of course, world models and energy-based models.
Harry Stebbings
What’s 1 piece of advice you’d give to AI startups navigating the changing landscape of training, inference, and hardware?
Steeve Morin
The number 1 thing I’d say is: do not resell compute if you can avoid it. A lot of AI startups building on top of AI are trying to make a margin on top of a very big cake, and ultimately what they sell is compute.
If you look at a dollar of spend, maybe 98% of it goes to somebody else’s margin. If you do AI, try to verticalize on the product as much as you can, but not on the compute. If your business model implies buying a lot of tokens, it’s a very hard circle to square to put that into $20 a month. Look at it from that angle, and if you can, try to avoid it.
17. Challenges in AI Hardware Supply
Harry Stebbings
What’s the biggest challenge that Jensen Huang faces today?
Steeve Morin
The highs are very high, but they don’t last forever. We’re seeing it, actually. It’s probably how to navigate the downslope. Blackwell is probably something that keeps him awake at night.
Harry Stebbings
Why would that keep him awake at night? Would that not reenergize him—more orders, new enthusiasm, a new product?
Steeve Morin
Because orders are getting canceled. They have a lot of problems with these chips, so a lot of people are canceling their orders. These chips are on the frontier of scaling, and they were supposed to come out last summer, but they have heat-dissipation and motherboard-bending problems.
The people who are very pre-silicon told me, “This is what we call a pretty big problem,” end quote. Probably navigating the downslope is the challenge.
You may not know this, but the supply of H100s was actually smoothed out over the year so that they didn’t have a big spike in deliveries followed by a quarter with fewer deliveries, which pissed a lot of people off—especially people who bought a lot of them. Some still haven’t received their orders from last year, and they’re already seeing the new chip, the B200, and then the one after that. They’re super pissed.
I think navigating the downslope will be important at some point. The question is when and how. If there’s an H100 bubble, of course it will impact NVIDIA. I’m probably going to get a lot of flak for this, but I’ve seen some very worrying numbers about Blackwell and varying testimony from people who operate these things. I don’t know. Maybe that train will stop, or at least slow down.
Harry Stebbings
Steeve, I’m not sure I’ve ever learned quite as much in 1 episode. I love what I do because I’m able to ask anything to the smartest people in their businesses, and I so appreciate you unpacking so much of it for me today. I’m thrilled to say that I finally get what you do after speaking with you. You’ve been a star. Thank you.
Steeve Morin
Thank you. I appreciate it.