1. What Will Be the Ratio of Synthetic to Human Data Used in 5 Years?
Andrew Feldman
Our AI algorithms today are not particularly efficient. In a GPU, most of the time it's doing inference, it's 5% or 7% utilized. That means it's 95% or 93% wasted.
We won't be as dependent on transformers in 3 years or 5 years as we are now—100%. The fundamental architecture of the GPU with off-chip memory is not great for inference. Now, they will continue to do well in inference, but they can be beaten, and I think they know it.
Harry Stebbings
Andrew, it is such a pleasure to meet you. I've wanted to do this one for a while, and I've heard so many good things from Eric for a long time, so thank you so much for joining me.
2. Where Was AI Landscape in 2015 When Cerebras Founded
I have my pen ready. I feel like this is going to be a learning experience for me. I want to go back to 2015. What did you and the team see in the AI landscape in 2015 that led to the founding of Cerebras?
Andrew Feldman
We saw the rise of a new workload, and this is every computer architect's dream. We saw a new problem to solve, and what that means is maybe you can build a new machine better suited to that problem.
In 2015—and the credit goes to Gary, Shan, JP, and Michael, my co-founders—they saw on the horizon the rise of AI. What that meant was there'd be a new problem for computers, and what the AI software would ask from the underlying chip or processor would be different. We came to believe that we could build a better machine for that problem.
That's what we saw. Obviously, we didn't see it exactly right. I underestimated it. This is my 5th startup, and the first time I underestimated the size of the market by a lot. What we did get right was that this was going to be big, that it would put a different type of pressure on a processor, that it would put pressure on the memory bandwidth, and that it would put pressure on the communication structure.
3. NVIDIA’s Biggest Strength Has Become Their Biggest Weakness
That's what we saw. We dove in, and it's been an extraordinary 9 years.
Harry Stebbings
Can you help me understand how the movement into an age of AI changes the requirements from a chip perspective of what is needed for a provider, and how that resulted in how you built Cerebras?
Andrew Feldman
The way to think about a chip is that it does 2 things: it does calculations and it moves data. Sometimes, along the way, it stores data. That's what a chip does.
What AI presented was a very unusual combination of challenges. First, the underlying calculation is trivial. It's a matrix multiplication, and an FMAC can be developed by any 2nd-year electrical engineering student. You say to yourself, “Holy cow, this has a huge number of very, very simple calculations.”
The hard part with AI work is that results and intermediate results have to be moved a lot. They have to be moved to memory and from memory, and they have to be broken up and moved among GPUs. What we saw was that this was going to be the hard problem, and that if we could solve for that problem, we would build an AI computer that was faster and used less power.
4. What Happens to the Cost of Inference?
Harry Stebbings
When we think about what we're going to build and what we're building for, there are a couple of core elements. Where are you going to focus? Are you focusing on fine-tuning, training, or inference?
You chose all 3. Why? I'm sorry for my basic questions, but I thought GPUs were specialized toward training and weren't specialized toward inference. Can you have a mono-architecture that does all 3 best?
Andrew Feldman
The first step in computer architecture is deciding what you're not going to do. What are we not going to be good at? That's really the first important question to answer.
To answer your question, is the computational work for training from scratch different from fine-tuning? The answer is that it's not different. It's approximately the same.
Inference and training have some different requirements, and generative inference in particular has some very challenging requirements on exactly the communication dimension that I mentioned. In generative inference, you have to move all the weights from memory to compute to generate a single word. You have to move them again to generate the next word, and again and again.
If you have a 70-billion-parameter model—not a giant model—and each weight is 16 bits, you're moving 140 gigabytes of data to generate 1 word. This is an enormous amount of data movement across memory, and that needs memory bandwidth.
If you have an architecture like we saw in the GPU, that's your fundamental limitation. It's a fundamental architectural limitation. That was what we went to wafer scale to solve.
5. Why Are AI Algorithms So Inefficient?
They use a memory called HBM, a type of DRAM, and it's phenomenal memory, but it's slow. It's slow and high-capacity. When they set the architecture for graphics, that's what you wanted. You didn't have to go back and forth to memory very often.
SRAM, on the other hand, is unbelievably fast but has low capacity. We wanted to use SRAM, but if you build a normal-sized chip, you can't hold a model. By going to wafer scale, we were able to put down a huge amount of SRAM and get the benefits of speed and enough capacity.
If you build a normal-sized chip with SRAM and you want to do a 400-billion-parameter model in inference, you might need 4,000 chips. If you want to do a DeepSeek 671, you might need 6,000 or 8,000 chips. What an administrative nightmare.
You can keep it on 1 wafer, 2 wafers, 4 wafers, or 10 wafers. You get all the benefit of the SRAM, and because you've been able to use the wafer, you get this tremendous capacity as well.
Harry Stebbings
I totally get you on HBM and the slowness of it. Why is it, then, that so much of the market just continues to use it? Forty percent of NVIDIA's revenue is using those chips for inference.
Andrew Feldman
There wasn't really, unless you went to wafer scale, a credible other choice. This is the way GPUs had always been made. It's called a graphics processing unit. That's the way they were built, and it was part of their advantage against a CPU: they were built this way.
But now there are dedicated chips like ours, and what used to be their advantage is now a weakness. That's a fun market to be in, when over a very short period of time what you're good at becomes your weakness.
Harry Stebbings
With a market cap like theirs, and with Jensen as good as he is—which I'm sure we both agree with—they must know this.
Andrew Feldman
They do know this. There aren't a lot of choices. They don't make memory, so they're a consumer of other people's memory. That's SK hynix, Samsung, Micron. There are only 3, 4, or 5 companies that make huge amounts of memory. There aren't many choices.
It's part of a complex architectural tradeoff. The flip side is that it's worked really well for them. Look at where it's taken them.
In comparison to those of us who are wafer scale, it's a small set. It's a set of 1: us. We have a real advantage against them on inference.
Harry Stebbings
How do LPUs fit into this? We've got HBM, we've got SRAM with you, and, bluntly, we have many more of them to make it work and scale. Where do LPUs fit into this mix?
Andrew Feldman
There are a lot of ways to skin a cat. Our way is different from NVIDIA's way. It's different from the TPUs, and it's different from Trainium. They're different.
Right now, and every day since August 26, when we launched inference, our way has been the fastest way across a whole set of models tested by Artificial Analysis and others.
Harry Stebbings
When we think about that speed, you said that you're 1 of 1 with wafer scale and the associated architecture. What does that mean in terms of cost? With such efficiency, is it inherently more expensive, and what does that look like from a cost profile?
Andrew Feldman
This isn't our first dance. We've been building computers for a long time, and when you make a choice like wafer scale, you have to weigh the tradeoffs.
We use less power. We use less power because 1 of the most power-hungry things on a chip is the I/O, moving data off-chip. If you're moving data off-chip frequently, you're using more power than if you can keep it in the silicon domain, on-chip.
We knew we would use less power. We knew that if you went to wafer scale, you had to solve some problems that people said were impossible to solve, like yield. We had to invent techniques that allowed us to yield wafers. In fact, we invented techniques that allow us to yield as well as, or better than, others who are building much smaller chips.
Harry Stebbings
Can I interrupt and ask what yield is, and why is it impossible to solve?
Andrew Feldman
A wafer begins as a 12-inch-diameter circle, a slice of silicon, and your chip is punched out of this. It's the way your mother might take a cookie cutter and cut out cookie dough.
During the process, at some point, just like your mom might have done, she lifts up the edges and all the little bits are removed. What's left are just the cookies. Those are your chips.
What happens is there are a set of naturally occurring flaws. It's like your mother closing her eyes and throwing up a handful of M&M's. The bigger the cookie, the higher the probability you hit an M&M. The bigger the chip, the higher the possibility that you have a flaw.
Traditionally, when you had a flaw, you threw away the chip or sold it as a less valuable part. You shut down part of the chip and sold it as a less valuable part, something called binning.
Every wafer is going to have flaws. The bigger your chip, the higher the probability you hit a flaw, and the more of the silicon is wasted when you throw it away. This is what everybody thought was known truth.
One of the things our team realized was that there are other ways to handle flaws. What if, instead, you built your computer—your processor—out of hundreds of thousands of identical tiles? If there was a flaw, you could shut down that tile and work around it. You could have a row or a column of redundant tiles that, when you needed them, you could pull in.
That had traditionally been the technique used in memory-making, and memory yields are extraordinary. It occurred to us that if we could build a processor out of hundreds of thousands of identical tiles, we could use redundancy. When there was a flaw, we could leave it there, shut it down, work around it, and pull in 1 of the redundant tiles.
That had never been done in a computer before, and that's at the heart of our architecture. It allowed us to yield and deliver whole wafers.
Nobody had ever been able to do that in the 70-year history of our industry. Really, really smart people struggled. Likely Gene Amdahl, one of the fathers of our industry, had a company called Trilogy that crashed and burned trying to do this. We figured it out.
Harry Stebbings
When you speak about being the fastest, and across all benchmarks being the fastest, what matters the most? Is it being the fastest, being the most efficient, or being the least costly? How do you think about the stack of prioritization for your customers?
Andrew Feldman
I think it varies. If you go to get a cancer diagnosis—for God forbid, your mother or your wife—I think 93% accuracy is just plain not as good as 94% accuracy. You pay a lot and wait another week to understand what the accuracy is. You pay a lot.
On the other hand, if you want Llama 405B to generate data to help you tune Llama 70B, maybe you can wait a few days, 3 days, or a week more. There's no urgency there.
If you want an answer from Perplexity, you don't want to wait 45 seconds for a search answer. You don't want to wait in a chat. You don't want to wait 3 minutes for R1 on GPUs to give you an answer.
In interactive mode, milliseconds matter. In interactive mode, what Google showed years ago was that you can destroy your user's attention with milliseconds of delay. Being the fastest matters in that domain.
You have to be thoughtful and say that in some cases being the fastest doesn't matter. We'll call those batch. Maybe cheapest matters there. In other domains, there is no search if you have to wait 8 minutes to get an answer. That's not a product.
When you go fast, a whole set of new opportunities open up. Netflix used to mail DVDs. That's what happened when the internet was slow: they'd mail DVDs. I look young, Andrew, but I'm not that young. I remember Blockbuster.
First, we used to drive to Blockbuster to get a DVD or a video. Then Netflix was mailing them to us. Then we got broadband, and suddenly Amazon is a studio. It changed everything, and speed in inference does the same thing.
Harry Stebbings
When we chatted before, you gave this great equation for inference. What was the equation? It was really helpful for me in understanding it.
Andrew Feldman
It begins with the following: training makes AI. That's how we make AI. Inference is how we use or consume AI.
Understanding how big the inference market is means understanding the number of people who are going to use it, how often they're going to use it, and how much compute each use takes.
Right now, we're in this rare time where the number of people using AI is growing, the frequency with which they use it is growing, and the amount of compute used in each instance of use is growing. That's why you're getting this extraordinary growth, and that's why it's off the charts right now.
Harry Stebbings
When we think about the distribution of resources between training and inference, what will that look like in 5 years? We've seen a lot of focus go to training and not as much go to inference. What does that look like?
Andrew Feldman
What we made in AI until the middle of 2024 was a novelty. What we made in AI late in 2024 began to be useful.
Harry Stebbings
What was the turning point?
Andrew Feldman
If you look at the models, they became useful. ChatGPT wasn't really a technical innovation; it was a user-interface invention. It gave more people access, but we didn't really know what to do with it right away. It was cool. That's what I mean by novelty: “Whoa, this is cool.”
Now, if your marketing team isn't on an LLM each person several times a day, they're not doing their jobs. The difference between “It's cool” and “This is part of everyday workflow” is what changed, starting sometime in Q4 last year and running into this year.
AI became useful not just to a select group in Silicon Valley, but to my dad, my brothers, doctors, and ordinary people who aren't buried in the Silicon Valley discussion. When you get them, the market is ripping.
Harry Stebbings
Do you not still think we're incredibly early? Going back to your point, how many times bigger are we in 5 years? Are we 100 times bigger? Are we 1,000 times bigger?
Andrew Feldman
I think we're way over 100 times bigger.
Harry Stebbings
What does that mean in terms of what we need to equip ourselves to deliver these? They're incredibly energy-intensive, it's incredibly difficult, and our industry consumes a lot of power.
6. Why is it Total BS That We Have Hit Scaling Laws?
We're seeing some water usage come down, but are we equipped from an energy and data-center standpoint to deliver the inference requirements for a population that is as AI-hungry as we are?
Andrew Feldman
The first thing is to admit that this is a power-intensive problem. Our industry consumes an enormous amount of power.
The second thing to say is that the burden is on us to deliver exceptional value as an industry. You take both the good and the bad. In order to make it worthwhile from a societal perspective to expand all this power, we better deliver the goods.
We better use AI to find cures for diseases. We better use AI to solve a bunch of different societal problems. That's the macro view.
Do I think we're equipped? I think we're in a very unusual situation in the US, where we have plenty of power, but it's in all the wrong places. We have power in Niagara. What we don't have is power where you want to build data centers, where we have good fiber.
What we also don't have is a national way to relax the local regulations that make getting power difficult. When you go to Silicon Valley, if you want to build a data center, you're dealing with local government and vested interests. That's not an efficient way to decide if you want to build a power plant or put a new data center in, especially if it's large.
I think those places that have removed some of that burden—for example, through taxes—are getting a huge number of data centers built.
Harry Stebbings
When I spoke to Jonathan at Gro, he said there were a huge number of data centers being built that weren't really equipped properly. We've seen this massive supply of data centers that are done by tourists, so to speak, and that is a massive problem. The provisioning of these data centers isn't there. Do you agree?
Andrew Feldman
A data center is, to begin with, a construction project. It's access to power, a construction project, and a design and engineering component.
I think there's been a huge push for new-construction data centers. We don't know if they're going to be good enough. Many of them will be fine.
The guys who were there early were some of the Bitcoin-mining companies, like TeraWulf, the guys at Crusoe, and others. There were guys in Europe who were early in building buildings near low-cost power in order to run compute that used a lot of power. They are some of the leaders now in some of the largest projects.
Those are certainly not tourists. They're extremely sophisticated data-center builders. There are some tourists, but there are a lot of very knowledgeable data-center builders building huge facilities right now—gigawatt-scale facilities, both domestically and internationally.
Harry Stebbings
How do you think about how the cost of inference goes down with the surge of demand that we mentioned—over 100 times? Does the price reduce 100 times? Does it follow Moore's law continuously? How do we think about the ever-reducing price of inference?
Andrew Feldman
The cost of inference is built up of several pieces. There's the power and space consumed to generate the response. That's a data-center cost and an OPEX item.
Second, there's the cost of the computer. We can drive down the cost of the computers with each generation by driving up their performance.
The other thing we can do is develop more efficient algorithms. Our AI algorithms today aren't particularly efficient. In a GPU, most of the time it's doing inference, it's 5% or 7% utilized. That means it's 95% or 93% wasted.
Over time, I think that as an industry we get better at things. We can drive the cost of compute down, build more efficient data centers with lower PUEs, and make our algorithms more efficient, so that utilization on our now-cheaper computers is higher.
You get a higher percentage of the maximum number of FLOPS. You get more tokens per unit time for the same power.
Harry Stebbings
When you look at the inefficiency of the algorithms, and what that means for the utilization of the chips, why are people suggesting that we're at scaling laws already? That seems to suggest there is so much room for improvement.
How do you think about what you just said in conjunction with the idea that we're hitting this asymptote point?
Andrew Feldman
I don't think there's a lot of debate among senior ML thinkers that we have tremendous room for algorithmic improvement. I don't think there's a lot of debate there.
There's even debate about whether the scaling laws are over, whether we've run out of mojo to keep making or gathering data to fill these ever-bigger models. But OpenAI's work on o1 shows me that the scaling law, certainly for inference, is fully functional. The more compute you put on inference, the better answer you get.
Many of the leading models are now MoEs, so they're not presenting all of the weights to each token. That's 1 way to do it: present the important stuff, not the unimportant stuff.
There are other ways to do it that we will invent and learn over time. We have human models that aren't all-to-all connected. Many of our models today are all-to-all connected. That's a lot of unnecessary connections, connections that don't produce anything but that we still end up doing math over.
Harry Stebbings
What does “all-to-all connected” mean?
Andrew Feldman
In many of the layers in a neural network, every element is connected to every other one. That's not actually the way the learning happens. Some are more valuable, and some are not valuable at all.
Imagine you're going to read 50 books because you want to learn something. You can read all 50 books, or you could read the 3 books that are really important, or you could read summaries of the 3 books that are the most important.
The problem is that we don't know which they are at the beginning. There's a process that you could learn. There are things called dropout and all these other techniques to use sparsity to help solve these problems.
We are early in the evolution of AI, and that plays right into this point that we'll get better at these algorithms. Transformers aren't the end of the world. We'll get better. Better will mean faster, more accurate, and more efficient.
That's what's exciting about an ever-changing industry. That's why I'm not in all these other industries that don't change quickly. They're the same 9 years ago as they are today.
Harry Stebbings
This show is kind of strange to me because I speak to a lot of people, and they think about the 3 pillars—compute, algorithms, and data—and the common refrain is that we're actually very far along in all of them.
When I hear you, it's actually very exciting. I think they're wrong.
Andrew Feldman
I think they're wrong. I don't think we're very far along, and it's very difficult to say that we're early in an industry but far along on all of its underpinnings. I think we are early in all of them.
Harry Stebbings
If we take them 1 by 1, in 5 years, how much synthetic versus human data will be used to train models? If you had to put a percentage on it?
Andrew Feldman
Almost all synthetic.
Harry Stebbings
And is the utility value of synthetic data the same as human data?
Andrew Feldman
I think so. When you teach a pilot to fly in a simulator, there is a lot of potential data that isn't very useful in teaching a pilot to fly. They spend a lot of time going straight and doing nothing.
Takeoffs and landings are where you want to spend your time, and that's why, when we put pilots in simulators, that's what we have them doing. In simulators, we can create data where engines blow, where there are a whole set of problems, and where learning can take place. That's simulated data.
In the same way, when we think about creating data—whether it's for self-driving or other forms of AI—we want the data that's hard to gather. Otherwise, we just have a bunch of data of people driving straight on a freeway. That's not difficult. We've been able to do that for a decade.
What we want is an unprotected left turn in the snow. It's snowing, it's hard to see, and you've got an unprotected left turn. That's a difficult thing, and you want that thousands or millions of different ways. That's where the synthetic data comes along: to fill in the empty parts where it's really expensive or painful to get that type of data.
Think of the pilot. You want them spending a huge amount of time on things that are rare in their training. It's the same with a surgeon: a huge amount of time on things that are rare. Most of the time it's carpentry, but their expertise is only needed when something rare happens.
That's when their mettle is shown, when the unexpected occurs. I think we will get better synthetic data by a great deal.
Harry Stebbings
I get it from a consumer perspective and from an expectations perspective. If we move the needle on compute, algorithms, and data, what does that mean for the experience of AI?
Andrew Feldman
It gets faster and cheaper. Faster and cheaper is the first answer.
The second is that when things become faster and cheaper, new applications emerge. It's used everywhere.
When computers became faster and cheaper, suddenly they were in cars, then they were in your pocket, then they were in your dishwasher and your TV. We were saying 30 years ago, “I need a computer in my TV? Are you kidding me? I need one in my pocket?”
Now you've got powerful computers in your pocket, in your TV, in your kids' toys, and in the car. That's what happens. Diffusion of innovation accelerates when you make things faster and cheaper.
Harry Stebbings
This is Jevons's paradox and Satya's belief, isn't it?
Andrew Feldman
I know that in the VC community you have to cite 19th-century English economists.
Harry Stebbings
I'm English. I'm English. Come on. If I'm not allowed to cite an English philosopher, what am I here for?
Andrew Feldman
I think there are very few examples in our industry—actually none in compute in 50 years—in which, by making things cheaper and faster, the market got smaller. The market always gets bigger. It always does.
Harry Stebbings
From an architectural standpoint, you mentioned transformers. Is there a world where we move past transformers?
Andrew Feldman
Transformers, 100%. We won't be as dependent on transformers in 3 years or 5 years as we are now—100%. They're not the end-all and be-all.
Harry Stebbings
Why? What will replace them, and what does that look like?
Andrew Feldman
I don't know whether they're going to be state-space models or other types of models. What I know for sure is that innovation doesn't stop, and the transformer has some weaknesses that people are desperate to overcome.
There's a quadratic effect in the attention head. There are all sorts of things that could be improved. But it's pretty darn good now. It's the best we have, and that's what you run with. You run with the best you have, and the minute it's not the best you have, you drop it in favor of the best you have.
I think that's what we're seeing. We're seeing a large number of innovative companies designing models.
Harry Stebbings
What DeepSeek showed us is that you don't need 5,000 people and billions of dollars a year. You can do it with 200 smart people and more hardware than DeepSeek said they had, but less hardware than others had.
Were you very impressed with DeepSeek, and what impressed you most?
Andrew Feldman
I think it was the result of focused engineering, and that impressed me. It was designed to be better. They weren't confused about being model intellectuals, and they weren't confused about whether it was important to break new ground. They were interested in being better.
From an invention standpoint, that's a little boring. From an engineering standpoint, that was a sweet effort. They really built a model that was just plain better at many, many things, and that's cool. I like good engineering projects.
They chose to announce it right around Trump's inauguration, and the politics of it are a separate matter. We can talk about that later.
Harry Stebbings
Did distillation rile people up?
Andrew Feldman
I don't think distillation is wrong. Is summarization wrong? I'm a VC. Are you kidding me? That's what we do. If you didn't summarize, you wouldn't know anything.
7. What Specifically Was So Impressive About DeepSeek?
Harry Stebbings
Exactly. That's exactly right. I don't think distillation is wrong. If distillation is wrong, then certainly using people's copyrighted data is wrong. That's the problem. You've got to be a little bit consistent.
Andrew Feldman
I think neither is wrong, actually, but you have to be consistent.
Harry Stebbings
The thing with it, bluntly, is that DeepSeek is open. Everything that they innovated on, OpenAI can learn from and take too.
8. Why is Distillation Not Wrong and OpenAI Need to Look in the Mirror?
Andrew Feldman
I think there are few examples of an open-source anything having the sort of immediate impact that model had. That model had a giant impact in a technical community of really smart people.
There are very few examples of other open-source software projects that had that type of impact in that amount of time. Usually, you're in the business of betting on these guys: they ramp up, and they go from 10,000 users to 100,000 users to 1 million users. Then you better start a company around that and get those graduate students.
This had a loud boom in the industry immediately. It was, “Whoo, the thing!”
9. Where Will Value Accrue in a World of AI?
Harry Stebbings
The thing I have to think about as a venture investor is where enduring and defensible value is, and how I get in early and build that over time. In hardware, that's well understood.
But on the model side, do you think there is value when you look at the sheer number of players with relatively comparable models?
Andrew Feldman
To demonstrate enduring value, you need both immediate value and a trajectory for more.
The problem in some industries is that you're capable of demonstrating a leadership position for a short period, and then someone else—maybe the next generation—generates the next one, and the next generation generates the next one.
I think that, in the software world, you end up competing against other people's release cadences. You're 4 months ahead; they're 6 months ahead. If that's really where you are, there's not a lot of value.
But if you can stay at the top over years—even if you're not the best, even if you're in the top decile over years, while the people above you are changing constantly—I think there's a lot of value.
Very large Silicon Valley companies have been built with technology that was not the most compelling. It might have started as the most compelling technology, and then it got to a point where it was good enough, easy enough to use, and well distributed. That's when you're at the mature market.
10. How Will NVIDIA’s Market Position Change Over the Next Five Years?
We're a long way from there right now. Right now, we're in the early phases. You characterized my position exactly right: data, compute, and algorithms. I think we have a ton of room for improvement on all of them.
Harry Stebbings
You said that computing hardware is where the value is. How does that value distribution shake out? We've obviously got the 800-pound gorilla that is NVIDIA. How do you think about how the distribution of value shakes out in hardware and compute over the next 5 years?
Andrew Feldman
Historically, 1 of the barriers to entry was the capital intensity of a project. In the world of building chips, there are both scarce resources in expertise and high expense.
Historically, it hasn't fit comfortably in a software company. The things that modern software companies value aren't entirely conducive to chip-making.
When I look down the road, I think that people who build systems endure. Cisco and Juniper endure. Chipmakers have endured. There's a reason Apple and NVIDIA are among the most valuable companies on Earth. What they do is hard, and I think that's why it's worth challenging.
If it weren't hard, enormous, and difficult, why spend time being the underdog and challenging it?
Harry Stebbings
A lot of people place defensibility around NVIDIA's CUDA lock-in. To what extent is that real versus hype in inference?
Andrew Feldman
In inference, it's not real at all. There's no CUDA lock-in in inference. You can move from OpenAI on an NVIDIA GPU to Cerebras, to the Fireworks service on something else, to Together, to Perplexity with 10 keystrokes.
Anybody who actually uses AI knows there's no CUDA lock-in in inference.
There was a fundamental effort to disintermediate CUDA, first by Google with TensorFlow and by some graduate students with Caffe, and later by Google with TensorFlow and Facebook, or Meta, with PyTorch.
Today, most AI is written in PyTorch. You ought to be able to compile it and run it on your hardware.
NVIDIA has many moats. When you're a dominant market-share leader, that in itself is a moat. Being the default solution is a moat. Everybody learns to think about AI in your structures. Those are moats.
The software—compilers are hard, but they're tractable.
Harry Stebbings
I completely agree with you that being the leader is a moat in itself. It's never talked about that way.
Look at Intel. Intel has made nearly a decade of catastrophic decisions until hiring Lip-Bu Tan, and they still own 80% of the x86 market. AMD has worked up to perhaps 25% or 30%. After a decade of screwing up, Intel only lost 20% of its share. That's a moat.
Andrew Feldman
That moat is just unbelievable. You can make a bunch of bad decisions for a decade and only lose 20% share. That's extraordinary.
I'm a huge fan of Lip-Bu Tan. He's an investor in our company, and I wish him well. I think if anybody can change that company, he can.
I think we rarely talk about what being the market-share leader means in terms of a moat. As a challenger, we have to think about it exactly, because it's exactly those characteristics of the moat that we need to get over.
Harry Stebbings
In 5 years, is it Uber, or is it like AWS and cloud? Cloud is an interesting market where a couple of players have relative segments—25% or 30%—and it's shared relatively evenly between them. Or is it 1 like Uber, where Uber has 90%, Lyft has 5%, and alternative providers have the other 5%?
Andrew Feldman
I think it's going to be between those 2. In 5 years, NVIDIA is going to have 60%, somewhere between 50% and 60% of the market. Right now, they have approximately all of it.
Harry Stebbings
Of NVIDIA's usage, what percentage will be training versus inference?
Andrew Feldman
They'll continue to have a meaningful business on both sides. They're exceptional at training. They will not roll over and play dead in inference.
They're a world-class company. They've had 1 of the great decades of any company in history. From 2014, when they were worth $10 billion, to where they are right now, it's 1 of the great decades in corporate history.
I don't think they're going to roll over and say, “We're not going to be in the inference market.” That's not going to happen. They're going to have a meaningful share, but the market is growing, and we'll have a piece. Others will have a piece. I think there'll be some very big companies made in this 100× growth.
Harry Stebbings
Do you think chip providers will be far larger than model providers in terms of enterprise value in the 5-year timeframe?
Andrew Feldman
Yes.
Harry Stebbings
How does that prediction change on a different timeline?
Andrew Feldman
In a shorter timeline, when you price an option, variance and uncertainty increase the option's value. If you look at the way Black-Scholes works, or at any option-pricing model, uncertainty and variability are friends of the value of the option.
When people are paying these extraordinarily high prices for model companies right now, I think part of that is this extraordinary uncertainty and wild variance. In the shorter run, it might not be the case.
In the longer run, as markets mature and we begin to understand the value of these models, their businesses, and their long-term net profitability, we'll have a better understanding.
11. Why is the CUDA Locking for NVIDIA BS? What is Their Weakness?
What did Warren Buffett say about markets? In the short term, they're a voting mechanism, and in the long term, they're a weighing mechanism. At some point, the weighing kicks in. Usually, it's in the public markets, and then investors say, “Which is likely to give me better growth in the future?”
Harry Stebbings
You mentioned the word “public.” I do want to hone in on your business. You're cash-flow positive in a world where everyone else literally bleeds cash.
Help me understand: how did you become cash-flow positive when everyone else is bleeding or hemorrhaging cash?
Andrew Feldman
Traditionally, gross margins were a measure of technical differentiation. If you're running a negative-gross-margin business, I think it speaks for itself. You're selling a commodity. Your value creation isn't being recognized in the market.
12. Why is Trump Better for Business than Biden?
I think our technology is creating an opportunity for us to maintain margins where some others can't.
Harry Stebbings
A lot of your revenue is concentrated in the G42 deal. To what extent is that a strength or a weakness?
Andrew Feldman
It's both. The way you catch 3 large customers is to catch 1 first. The way you build 3 large strategic partners is to learn to be a strategic partner. That's a learned skill.
We didn't arrive knowing how to be a strategic partner at G42. Now that we've worked at it, it's a muscle we can replicate. We could be a better partner to any of a dozen different companies in the world.
Harry Stebbings
What have you learned in the G42 relationship-building process that makes Cerebras a good partner in a way that you weren't before?
Andrew Feldman
We've deployed tens of exaflops of compute, vastly more than anybody else that isn't AMD or NVIDIA. That's a huge amount of compute.
Our software has been hardened on some of the largest AI clusters in the world. We've gone through the growing pains of increasing manufacturing 2×, 5×, and 2× again through unbelievable growth in manufacturing.
We've worked with our supply-chain partners to be sure that they're ready for this extraordinary growth.
When you work with a strategic partner of this size, your organization comes out different on the other side. There are things you've learned and mistakes you've made.
I hadn't done a big relationship in the Middle East. There was a huge amount to learn. I think you come out a much better company and much better prepared to do business with a hyperscaler, another massive partner, or another sovereign.
But it takes real work, and your team has to learn.
Harry Stebbings
You said you come out better. Why go public when you did? When it happened, I thought it seemed preemptive, respectfully.
My question now to companies is: why go public at all? There is so much private capital. The decisions have shown very clearly that you can stay private for a lot longer than you planned to. Databricks has certainly shown that.
Those were historically public-market valuations, and the valuations that Anthropic, OpenAI, and some of the others are getting are historically public-market-only valuations. Your S-1 is live; anyone can read it. I wouldn't want people reading mine.
Andrew Feldman
We have nothing to hide.
Harry Stebbings
No, but your competitors have asymmetric information.
Andrew Feldman
Yes, we've got asymmetric technology. I think you have to be pretty transparent to be public.
You have to be ready organizationally. You have to be ready with your processes. You need to be ready to forecast and predict, and to be held accountable in a way that private companies historically haven't been.
We think there's tremendous value. We think we'll be among the first in the category. We think some of our largest targets would have a stated preference for doing business with public companies. Large enterprises in the US have done that historically.
Those were some of the reasons that led us to it.
Harry Stebbings
How many G42 relationships shall you have in the next 24 months? How fast can you ramp them?
Andrew Feldman
That's a good question. Several.
Harry Stebbings
Remind me, how big is the G42 deal?
Andrew Feldman
It was 87% of revenue.
Harry Stebbings
I know it was big.
Andrew Feldman
When we announced it, some estimated it was north of $1 billion.
Harry Stebbings
Well done. That must be a bit of a high five.
Andrew Feldman
There's tremendous excitement, and then there's every entrepreneur's reality: I have to make a lot more gear.
You make a list of your top 10 vendors and fly it to them all, saying, “Big orders are coming. Be ready.” You work with all your partners to get ready because you need to make a great deal more stuff.
That's 1 of the real differences between hardware and software. When we grow fast, the number of people you need to work with in your supply chain, and the amount of collaboration that needs to happen, is truly extraordinary.
Harry Stebbings
Are NVIDIA going to have a cluster of unhappy customers who, bluntly, have waited so long for chips that by the time they get them, the chips are outdated? Are they going to say, “What happened?”
Andrew Feldman
All of that is an opportunity for us and others. Being a market-share leader isn't easy either.
When the bully falls, everybody wants to give him a kick. A lot of that happened at Intel. They'd been the dominant player, and when they fell, everybody was happy to jump in and kick them when they were down.
I think there's a real opportunity in the potential for NVIDIA customer unhappiness. For those of us competing with them, if you can't get your gear, you may as well test somebody else. That's a huge opening.
Harry Stebbings
In hardware, you mentioned the complexity. Are export controls being implemented properly? Do you think they're a good idea?
13. Quick-Fire Round
Everyone was looking at DeepSeek and saying, “How did this happen? They must have stolen chips. How could this be?” What do you think about that?
Andrew Feldman
It turns out that they probably did use chips in Singapore.
I think managing software compliance and managing hardware compliance are extremely different things because their vectors of diffusion are different. There's a different weight.
If you sell a server that weighs 500 or 600 pounds and arrives on a pallet, you can go visit it. If you want to deploy it in Kazakhstan, you can put it in a data center and have somebody from the embassy visit it and take photos once a month. It's not going anywhere.
You can keep track of who uses it and provide logs. That's much harder with software, and open source is a whole other level.
That's the first observation. The second is that I got to know the leadership in Commerce in the previous administration. I didn't always agree with their policies, but it is a world of unintended consequences.
You sought to limit Chinese access to EDA tools to delay the growth of a Chinese chip market, and US venture capitalists backed tons of Chinese companies in Shenzhen to build EDA tools. This is an unbelievably slippery, dynamic, challenging problem.
I don't know if it's a tractable problem. Delaying another nation's progress on a technical trajectory is an enormously challenging thing.
I came to appreciate just how difficult it was for well-meaning people to predict the impact of policy during the last 2 years.
Harry Stebbings
Do you think this administration is better for AI than the prior administration?
Andrew Feldman
I don't think there's any doubt that's the case.
Harry Stebbings
What makes you say that?
Andrew Feldman
The past administration lined itself up against Big Tech, and that was a mistake.
AI is also in a different place, so it's easier to be for it. It's less scary now than it was in 2021. We have a better picture of the trajectory, both the risks and the benefits.
I think this administration had the foresight to put an AI czar or leader in place as a focal point for discussions. It's probably net a fair bit better.
Harry Stebbings
You said it's very challenging to hinder a nation's development, adoption, and progression of a technology. Respectfully, you chose not to sell to China. Why was that, and does that not go against the difficulty of hindering progression?
Andrew Feldman
I have a very simple rule, and I encourage your team to use it. You don't need a big handbook to help you make good decisions in a company. Just ask yourself, “Would my mother be proud?”
Would she be proud if I did this? Would she be proud if I explained the exact situation? Would she look at me and say, “I'm proud you're doing this, son?”
I asked myself that, and I came to believe that the deal on the table wouldn't be used for good. I wasn't comfortable with that. I wouldn't have been able to explain it to my mother.
That's a moral compass. It wouldn't have been used for good.
Harry Stebbings
I'm naive. What would they use it for?
Andrew Feldman
They could use it to power drones. Some use it for facial recognition to identify minorities for persecution, or to build military equipment—to do things that I either couldn't see or that, from what I saw, didn't feel right.
14. Do We Underestimate China in a World of AI?
It's more important than money.
Harry Stebbings
Do you think we fundamentally underestimate Chinese capabilities?
Andrew Feldman
100%. It is 1 of the most obvious and frequent errors in judgment: you underestimate the other side.
You have to look carefully at what they're doing. Their investment in infrastructure has been extraordinary. The rate at which they generate engineering talent is exceptional.
The government's ability to have a policy and implement it is extraordinary. They're not a democracy; they weren't designed to have checks and balances there.
The funding that flowed into the development of AI technology was significant. Their venture capitalists were backed by their government. They have national champion companies. They've developed a belt-and-suspenders strategy to make much of the developing world dependent on them and their technologies.
They absolutely should not be underestimated. They have a lot of people, and we see a tiny fraction of it.
I think they have produced industrial policy that has moved their nation forward.
Harry Stebbings
What was the most significant part of that, do you think?
Andrew Feldman
The creation of economic zones like Shenzhen was clearly a visionary move. They knew that their own system was in the way, so they created zones that relaxed their own system.
Harry Stebbings
Could the US learn from them in that way?
Andrew Feldman
We did some of the same things in the 1st Trump administration. What did we do? We relaxed our own rules in the development of vaccines. We knew that, in that time, it would be very difficult to go through the steps that we always go through, and we tried to implement thoughtful shortcuts, or workarounds.
Why are they committed to trains as a mode of transportation, and we can't build a decent train system in the US or in California? Why can't we build infrastructure when the rest of the world can build extraordinary high-speed trains linking important cities?
Why do we have 3 different standards for train rails? Why are our bridges and freeways in disarray?
I think those are questions we have to ask ourselves when we see other people doing it differently.
If you watch a good football team, you say, “That's interesting offense.” You're not thinking, “How could our team learn? What could we do? Why did that work? What was it about the people they had, the talent, or the structure that made that a successful series of plays?”
What can I take away from that? How can that inspire me to do better?
I'm always looking for inspiration in others, competitors, and partners. Some of our partners at G42 have an unbelievable work ethic. It inspires me. The scope of the challenge they've undertaken inspires me.
Harry Stebbings
What do you believe that most around you disbelieve?
Andrew Feldman
I think we're closer to peace in the Middle East than people believe.
There is a rise of a moderate, business-focused Arab state that wasn't there 25 or 30 years ago. If you visit the UAE, Qatar, or even Saudi Arabia, what you see is amazing transformation.
I think there's a desire to be included in the West in their own way, but also to enjoy the benefits of it. We are closer than people may think.
Harry Stebbings
What's the most underrated threat to NVIDIA's market-share dominance?
Andrew Feldman
The fundamental architecture of the GPU with off-chip memory is not great for inference. They will continue to do well in inference, but they can be beaten, and I think they know it.
Harry Stebbings
What's a crazy AI prediction you have that most people would call science fiction?
Andrew Feldman
Dario at Anthropic says that we'll live to 150. I don't think we're going to live to 150.
I don't think that 90% of our code will be written by machines this year. But I do think that within a year or 2, AI's penetration will be approximately the same as telephones—cell phones.
Harry Stebbings
What have you changed your mind on in the last 12 months?
Andrew Feldman
There are lots of things. Many decisions I made turned out to be wrong.
There are 2 ways you can be wrong. You can actively be wrong, or you can fight against what was right.
In 2016, JP, 1 of our co-founders and chief system architect, laid out a plan that would have us doing water cooling for our systems. Nobody else was doing it, and I fought so hard. I was so wrong.
JP was right. About a year or 2 later, Google announced that the TPUs were going to be water-cooled. We were 1st, and now NVIDIA is only selling water-cooled parts.
I was dead wrong, and JP was right.
When you make a lot of decisions every day, there are many instances where you're wrong. I've been wrong about people. People I thought were pretty good turned out to be extraordinary. People I thought would be extraordinary were really smart but couldn't finish projects or get things done.
If you're not prepared to be wrong a fair bit, you ought not to be making a lot of decisions, because it comes with the territory.
Harry Stebbings
As a venture capitalist, I'm never wrong, so—
Andrew Feldman
As a venture capitalist, you're wrong 9 times out of 10, and everybody forgets as long as you're really right.
Harry Stebbings
I get a picture of you signing the term sheet with me.
Andrew Feldman
Good. Yours is a perfect industry in which nobody cares about the average. On average, you're wrong all the time. What they care about is the occasional time you're really right, and that moves a fund.
That's different from being a CEO. I think we've got to be mostly right most of the time, but if you're making a lot of decisions, you're still making a ton of mistakes.
Harry Stebbings
This is your 5th startup. You are a sucker for punishment, aren't you? Really—5 times? Did you not get beaten alive enough?
My question is about the value of serial entrepreneurship. I've spoken to many people who don't believe in it, respectfully. How do you think about the inherent benefits of having done it 4 times before?
Andrew Feldman
If you're in a business in which running a business is a benefit, then experience matters a great deal.
If you're in a business in which you look like your customer, there was a reason social networks were started by people right out of college or in college. Dating is top of their mind. They look like their customers, and that was more important than knowing anything about running a business.
In that environment, it will select for people who are of the demographic their customers are. They know that backwards and forwards.
But if you want to have a business that has manufacturing, a supply chain, and hundreds or thousands of engineers managed to a timeline and a schedule, I don't think anybody would turn around your statement and say with a straight face, “What I'm looking for is an engineering leader with no experience.”
You wouldn't want somebody who had only led a team of 400 or 500 people. You'd want somebody who had experienced the challenges of growth.
Harry Stebbings
No, I don't want somebody who's led a team of 400 or 500. What I'm looking for is somebody with no experience. Naivety is a bonus here.
Andrew Feldman
The people who sell that are sometimes consultants. “My guys have no experience in your industry, so they aren't biased.” Maybe a little bit of experience in the industry would help.
Harry Stebbings
Where are people investing today in AI across the stack where you're thinking, “Why is so much cash going to that part?” I'm not asking for a company; I mean a category.
Andrew Feldman
Sometimes money needs to find a home. Some people have raised really big funds, and they have to find a home for that money. Some people don't like to be left out. They're willing to make investments for status purposes or other reasons that don't seem to make sense.
I haven't thought about it in detail. I think there are some underappreciated places of investment.
In the chip world, the sub-milliwatt, tiny little chips that live next to sensors and do inference are an extremely interesting market. These are tiny things that will only send back useful data, and they'll sell in enormous volume.
It's not a part of the market I love to play in. I like to build bigger things and sell them to the data center. But I think that part is extremely interesting. I think it will be fundamental for robotics.
That's an area where I think the opportunity is extremely underappreciated.
Harry Stebbings
If we think about Cerebras in 10 years, where do you envision the business? If everything goes well, where are we in business having that conversation?
Andrew Feldman
10 years ago, NVIDIA was worth $10 billion, so that's a long run in our world right now.
In 3 to 5 years, I would like our technology to have been used to solve 2 important societal problems. I would like it to have been used to find a therapeutic for an affliction that impacts more than 1 million people a year.
I would like our inference to be powering a collection of apps that don't exist today. I would like a meaningful portion of the population in the US and Europe to inadvertently use our technology—to use something that we power without even knowing it.
Harry Stebbings
Andrew, I've wanted to make this show happen for a long time. As I said, I've heard so many good things from Marc for many years.
There have been so many requests to have you on the show. My team was saying, “Just get Andrew on the show, Harry.” I was like, “Okay, okay.”
I tweeted it, obviously, which is how we got this conversation. You tweeted it, and around 40 people sent me messages saying, “How come you're avoiding Harry? How come he has to tweet it?” I was just like, “All right, just call me.”
Andrew Feldman
It's good. Send me a note. I'm happy to come on.
Really thoughtful questions, Harry. Really thoughtful and interesting. It was a really good conversation.