[BidClub_]
20VC · · 63 min

Andrew Feldman, Cerebras Co-Founder and CEO: The AI Chip Wars & The Plan to Break Nvidia's Dominance

Harry StebbingsAndrew Feldman

YouTube
TL;DR
  • Feldman's core attack on NVIDIA: the GPU's off-chip HBM memory is a fundamental architectural limitation for generative inference — serving one word from a 70B-parameter model means moving ~140GB of weights from memory to compute, again for every next word. Wafer scale let Cerebras use fast SRAM at capacity, and "what used to be their advantage is now weakness... it can be beaten and I think they know it." His market-structure call: NVIDIA goes from "approximately all" of the market today to 50-60% share in five years — between Uber's 90/5 and cloud's even split.
  • The CUDA moat is "not real at all" in inference — "you can move from OpenAI on an Nvidia GPU to Cerebras... with 10 keystrokes." The real, rarely discussed moat is market-share leadership itself: Intel made "nearly a decade of catastrophic decisions" and still holds ~75-80% of x86. "I can make a bunch of bad decisions for a decade and only lose 20% share — the moat was just unbelievable."
  • Scaling-law gains continue — Feldman flatly rejects the "far along on compute, algorithms and data" consensus: "I think they're wrong. I think we are early in all of them." A GPU doing inference runs at 5-7% utilization — "95 or 93% wasted" — OpenAI's o1 shows inference scaling laws "fully functional," and we won't be as dependent on Transformers in 3-5 years ("100%"). In five years training data is "almost all synthetic."
  • The inference market equation: users × frequency × compute-per-use — all three growing simultaneously, a rare condition. AI flipped from "novelty" to useful in Q4 2024, the market in five years is "way over 100 times bigger," and faster/cheaper only expands it: "no examples in compute in 50 years in which by making things cheaper, faster, the market got smaller."
  • Chip providers will be worth more than model providers on a 5-year view. Today's model valuations are option pricing — "uncertainty is a friend of the value of the option" — but Buffett's weighing machine eventually kicks in. On models generally: competing on release cadence four months ahead of rivals carries "not a lot of value"; staying top-decile for years does.
  • Cerebras is cash-flow positive where peers hemorrhage cash — "gross margins were a measure of your technical differentiation... if you're running a negative gross margin business, you're selling commodity." The flip side: G42 is 87% of revenue (a deal estimated north of $1bn), which Feldman defends as a learned muscle: "the way you catch three large customers is to catch one first," with several relationships targeted in 24 months.
  • On China: Cerebras refused to sell — the deal "wouldn't be used for good" (facial recognition of minorities, military) and failed his mother test — yet he thinks the US underestimates China "100%," export controls may not be "a tractable problem," and the attempt to cut off EDA tools just spawned US-VC-backed EDA startups in Shenzhen.
Digest · the substance, structured for research

1. The bet in 2015: AI's hard problem is moving data, not computing

  • Feldman's framing of chip design: a chip "does calculations and it moves data" — and AI inverted the difficulty. The math is trivial ("a matrix multiplication and an FMAC can be developed by any second-year electrical engineering student"); the hard part is that results and intermediate results must be moved constantly — to memory, from memory, among GPUs. Cerebras bet that solving data movement would yield a faster, lower-power AI computer.
  • His one admitted miss, notable from a fifth-time founder: "the first time I underestimated the size of the market by a lot." What the team got right was that AI would pressure memory bandwidth and communication structure — the dimensions GPUs weren't built for.
  • Training and fine-tuning are computationally "approximately the same"; generative inference is the outlier. To generate one word from a 70B-parameter model at 16-bit weights, you move ~140 gigabytes from memory to compute — then again for the next word. "That's called memory bandwidth, and if you have an architecture like the GPU, that is your fundamental limitation."

2. Wafer scale and the yield problem nobody solved in 70 years

  • The memory trade: HBM is "phenomenal... but it's slow" — high capacity, built for graphics where you rarely go back to memory. SRAM is "unbelievably fast but has low capacity" — a normal-sized SRAM chip serving a 400B-parameter model needs ~4,000 chips; a DeepSeek 671 needs "six or 8,000... what an administrative nightmare." Wafer scale gets SRAM's speed with enough capacity on one, two, or ten wafers — and less power, since off-chip I/O is among the most power-hungry operations on a chip.
  • Why nobody had done it: yield. His analogy — a wafer is cookie dough, flaws are M&Ms your mother throws blindfolded; "the bigger the cookie, the higher probability you hit an M&M," and traditionally you binned or trashed flawed chips. Cerebras's answer, borrowed from memory-making: build the processor from hundreds of thousands of identical tiles with redundant rows or columns, shut down a flawed tile and route around it. "Nobody had ever been able to do that in the 70-year history of our industry" — likely Gene Amdahl's company, Trilogy, "crashed and burned trying."

3. Speed is not a spec — it's what creates product categories

  • Feldman refuses a single ranking of fast/cheap/accurate: for a cancer diagnosis, "93% accuracy is just plain not as good as 94%" and you'll pay and wait; for batch work like Llama 405B generating tuning data for 70B, cheapest may matter. But in interactive use "milliseconds matter" — Google showed years ago "you can destroy your user's attention with milliseconds of delay." "There is no search if you've got to wait eight minutes."
  • The analogy he leans on: when the internet was slow, Netflix mailed DVDs; broadband arrived and "suddenly Amazon's a studio. It changed everything. Speed in inference does the same thing" — new applications open up that couldn't exist at GPU latencies. Cerebras claims the fastest inference "across a whole set of models" since its August 26 launch, per Artificial Analysis.

4. The inference equation: all three multipliers growing at once

  • His sizing formula: inference market = people using AI × frequency of use × compute per use. "We are in this rare time" where all three are growing simultaneously — hence off-the-charts growth. Five years out the market is "way over 100 times bigger."
  • The turning point wasn't technical: "ChatGPT was not really a technical innovation, it was a user interface invention." Until mid-2024 AI was "a novelty... whoa, this is cool." Starting Q4 2024 it became useful — "if your marketing team isn't on an LLM, each person several times a day, they're not doing their jobs" — reaching "my dad, my brothers, the doctors," and "when you get them, the market is ripping."
  • On power: he concedes the industry "consumes an enormous amount of power... and some water," so "the burden is on us to deliver exceptional value" — cures, societal problems. The US problem isn't supply but geography and process: "we have power in Niagara... what we don't have is power where you want to build data centers" and no national mechanism to override local regulation and installed interests.
  • Pushback on Jonathan at Groq's "tourist data centers" claim — Feldman mostly disagrees: the early movers were Bitcoin miners like TeraWolf and Crusoe, "certainly not tourists... extremely sophisticated data center builders" now leading gigawatt-scale projects. "Sure there's some tourists," but many facilities will be fine.

5. Scaling-law gains continue — "I think they're wrong"

  • Against the refrain that we're far along on compute, algorithms and data: "I think they're wrong. I don't think we're very far along... I think we are early in all of them." Exhibit A: a GPU doing inference is "5 or 7% utilized — that means it's 95 or 93% wasted." Costs fall via cheaper computers, lower-PUE data centers, and better algorithms compounding together.
  • On the scaling-laws debate itself: there's genuine argument about whether we've "run out of mojo" on data, but "OpenAI's work on o1 shows me that the scaling laws certainly for inference are fully functional — the more compute you put on inference, the better answer you get."
  • The algorithmic headroom, as told: many models are still all-to-all connected — "connections that don't produce anything that we still end up doing math over." His analogy: to learn something you can read 50 books, or the three that matter, or summaries of those three — "the problem is we don't know which they are at the beginning." MoE, dropout, sparsity are early steps. And on architecture: "we won't be as dependent on Transformers in three years or five years as we are now — 100%. They're not the end-all be-all" — he won't guess the successor ("I don't know whether they're going to be state-based models"), but the attention head's quadratic effect is a known weakness people are "desperate to overcome."
  • Synthetic data in five years: "almost all synthetic," with utility he thinks is equal to human data. His pilot analogy, worth keeping whole: real driving data is "people driving straight on a freeway — not difficult"; what you want is "an unprotected left turn in the snow... thousands of different ways, millions of different ways." Like simulators for pilots and rare cases for surgeons, synthetic data fills in exactly what's expensive or painful to gather.

6. DeepSeek was "focused engineering" — and the distillation complaint fails a consistency test

  • What impressed him: "they weren't confused about being model intellectuals... they were interested in being better. From an invention standpoint that's a little boring, but from an engineering standpoint that was sweet effort." DeepSeek proved "you don't need 5,000 people and billions of dollars of gear — you can do it with 200 smart people and more gear than DeepSeek said they had, but less gear than others had." The inauguration-timed announcement he files under politics.
  • On distillation: "is summarization wrong?... I don't think distillation is wrong, and if distillation is wrong then certainly using people's copyrighted data is wrong. You've got to be a little bit consistent." And on impact: "there are very few examples of an open-source anything having the sort of immediate impact that model had... this had a loud boom."

7. Where value accrues: chips over models, and the moat nobody names

  • CUDA lock-in "in inference, it's not real at all... none — you can move from OpenAI on an Nvidia GPU to Cerebras to Fireworks to Together with 10 keystrokes." Most AI is written in PyTorch; compilers are "hard but tractable." NVIDIA's real moats are being the default — and Feldman thinks challengers under-study this: Intel made "nearly a decade of catastrophic decisions" pre-Lip-Bu and still holds ~80% of x86 while AMD clawed to 25-30%. "That's a moat... as a challenger we have to think about it exactly, because it's exactly that we need a bridge for."
  • The five-year market structure: between Uber (90/5/5) and cloud's shared oligopoly — "Nvidia is going to have somewhere between 50 and 60% of market... right now they have approximately all of it." He's emphatic NVIDIA won't "roll over and play dead in inference" — "one of the great decades of any company in history," from ~$10bn in 2014. But NVIDIA's long wait times are "a huge opening": "when the bully falls, everybody wants to give him a kick."
  • Chip providers larger than model providers in enterprise value at five years: yes. His mechanism for why models look expensive now: "when you price an option, variance and uncertainty increases the option's value... part of these extraordinarily high prices is this wild variance." Long-run, Buffett applies: markets are a voting machine short-term, "a weighing mechanism" long-term — "at some point the weighing kicks in, and usually it's in the public markets."
  • On model-company defensibility: "you're competing against other people's release cadences — you're four months ahead, they're six months. If that's really where you are, there's not a lot of value. But if you can stay top-decile over years while the people above you are changing constantly, I think there's a lot of value." Hardware endures — Cisco, Juniper, Apple, NVIDIA — because "what they do is hard. That's why it's worth challenging."

8. The business: cash-flow positive, one giant customer, and why go public

  • On being cash-flow positive while rivals bleed: "traditionally your gross margins were a measure of your technical differentiation... if you're running a negative gross margin business, you're selling commodity — your value creation isn't being recognized."
  • G42 at 87% of revenue (estimated north of $1bn when announced) is "both" strength and weakness: "the way you catch three large customers is to catch one first... being a strategic partner is a learned skill." The proof points: tens of exaflops deployed — "vastly more than anybody that isn't AMD or Nvidia" — software hardened on some of the largest AI clusters, manufacturing scaled "2x and 5x and 2x." Target: "several" relationships in the next 24 months.
  • On Harry's challenge that the IPO filing seemed preemptive and hands competitors asymmetric information: "we have nothing to hide... we've got asymmetric technology." The affirmative case: first in category, and "some of our largest targets would have a stated preference for doing business with public companies."

9. China: refuse the sale, respect the rival, doubt the controls

  • On export controls, a hardware-native's distinction: a 500-600lb server "arrives on a pallet — you can have somebody from the embassy visit it, take photos once a month. It's not going anywhere." Software and open source are "a whole other level." He notes DeepSeek "probably did use chips in Singapore." His deeper doubt: "I don't know if it's a tractable problem to delay another nation's progress on a technical trajectory" — the EDA restrictions just meant "US venture capitalists backed tons of Chinese companies in Shenzhen to build EDA tools." This administration is "probably net a fair bit better" for AI than the last, which "lined itself up against big tech."
  • Yet Cerebras itself refused a China deal — his test: "just ask yourself, would my mother be proud?" What he saw or couldn't see — facial recognition "to identify minorities for persecution," military equipment — failed it. "It's more important than money." He holds both positions without smoothing the tension: controls may be futile, but this sale was his choice.
  • Do we underestimate China? "100%... one of the most obvious and frequent errors in judgment." The evidence as he lists it: extraordinary infrastructure investment, exceptional engineering-talent generation, state-backed VCs, national champions, "a belt-and-suspenders strategy to make much of the third world dependent on them." Shenzhen's economic zones were "clearly a visionary move" — and America has done the same when it wanted to (Trump-1 vaccine rule relaxation). His uncomfortable questions: why can't the US build trains, why are "our bridges and our freeways in disarray"?

10. Quick fire: being wrong, tiny chips, and what experience buys

  • His best documented error: fighting co-founder JP's 2016 water-cooling plan — "I fought so hard and I was so wrong." Google announced water-cooled TPUs a year or two later; "now Nvidia's only selling water-cooled parts. I was dead wrong and JP was right." The generalization: "if you're not prepared to be wrong a fair bit, you ought not to be making a lot of decisions." A CEO must be "mostly right most of the time" — unlike VCs, where "on average you're wrong all the time and what they care about is the occasional time you're really right."
  • Underappreciated corner of the chip market: sub-milliwatt inference chips living next to sensors that "only send back useful data" — enormous volume, "fundamental for robotics" — though not his market ("I like to build bigger things and sell them to the data center").
  • On Dario's predictions: he rejects living to 150 and 90% of code machine-written this year, but expects AI penetration approximating cell phones within a year or two. Contrarian belief: Middle East peace is closer than believed, on the "rise of a moderate, business-focused Arab state" across UAE, Qatar, even KSA. And on fifth-time founding: where a business has manufacturing, supply chain, and hundreds or thousands of engineers on a schedule, experience compounds — "nobody would say with a straight face: what I'm looking for is an engineering leader with no experience."

1. What Will Be the Ratio of Synthetic to Human Data Used in 5 Years?

Andrew Feldman

Our AI algorithms today are not particularly efficient. In a GPU, most of the time it's doing inference, it's 5% or 7% utilized. That means it's 95% or 93% wasted.

We won't be as dependent on transformers in 3 years or 5 years as we are now—100%. The fundamental architecture of the GPU with off-chip memory is not great for inference. Now, they will continue to do well in inference, but they can be beaten, and I think they know it.

Harry Stebbings

Andrew, it is such a pleasure to meet you. I've wanted to do this one for a while, and I've heard so many good things from Eric for a long time, so thank you so much for joining me.

2. Where Was AI Landscape in 2015 When Cerebras Founded

I have my pen ready. I feel like this is going to be a learning experience for me. I want to go back to 2015. What did you and the team see in the AI landscape in 2015 that led to the founding of Cerebras?

Andrew Feldman

We saw the rise of a new workload, and this is every computer architect's dream. We saw a new problem to solve, and what that means is maybe you can build a new machine better suited to that problem.

In 2015—and the credit goes to Gary, Shan, JP, and Michael, my co-founders—they saw on the horizon the rise of AI. What that meant was there'd be a new problem for computers, and what the AI software would ask from the underlying chip or processor would be different. We came to believe that we could build a better machine for that problem.

That's what we saw. Obviously, we didn't see it exactly right. I underestimated it. This is my 5th startup, and the first time I underestimated the size of the market by a lot. What we did get right was that this was going to be big, that it would put a different type of pressure on a processor, that it would put pressure on the memory bandwidth, and that it would put pressure on the communication structure.

3. NVIDIA’s Biggest Strength Has Become Their Biggest Weakness

That's what we saw. We dove in, and it's been an extraordinary 9 years.

Harry Stebbings

Can you help me understand how the movement into an age of AI changes the requirements from a chip perspective of what is needed for a provider, and how that resulted in how you built Cerebras?

Andrew Feldman

The way to think about a chip is that it does 2 things: it does calculations and it moves data. Sometimes, along the way, it stores data. That's what a chip does.

What AI presented was a very unusual combination of challenges. First, the underlying calculation is trivial. It's a matrix multiplication, and an FMAC can be developed by any 2nd-year electrical engineering student. You say to yourself, “Holy cow, this has a huge number of very, very simple calculations.”

The hard part with AI work is that results and intermediate results have to be moved a lot. They have to be moved to memory and from memory, and they have to be broken up and moved among GPUs. What we saw was that this was going to be the hard problem, and that if we could solve for that problem, we would build an AI computer that was faster and used less power.

4. What Happens to the Cost of Inference?

Harry Stebbings

When we think about what we're going to build and what we're building for, there are a couple of core elements. Where are you going to focus? Are you focusing on fine-tuning, training, or inference?

You chose all 3. Why? I'm sorry for my basic questions, but I thought GPUs were specialized toward training and weren't specialized toward inference. Can you have a mono-architecture that does all 3 best?

Andrew Feldman

The first step in computer architecture is deciding what you're not going to do. What are we not going to be good at? That's really the first important question to answer.

To answer your question, is the computational work for training from scratch different from fine-tuning? The answer is that it's not different. It's approximately the same.

Inference and training have some different requirements, and generative inference in particular has some very challenging requirements on exactly the communication dimension that I mentioned. In generative inference, you have to move all the weights from memory to compute to generate a single word. You have to move them again to generate the next word, and again and again.

If you have a 70-billion-parameter model—not a giant model—and each weight is 16 bits, you're moving 140 gigabytes of data to generate 1 word. This is an enormous amount of data movement across memory, and that needs memory bandwidth.

If you have an architecture like we saw in the GPU, that's your fundamental limitation. It's a fundamental architectural limitation. That was what we went to wafer scale to solve.

5. Why Are AI Algorithms So Inefficient?

They use a memory called HBM, a type of DRAM, and it's phenomenal memory, but it's slow. It's slow and high-capacity. When they set the architecture for graphics, that's what you wanted. You didn't have to go back and forth to memory very often.

SRAM, on the other hand, is unbelievably fast but has low capacity. We wanted to use SRAM, but if you build a normal-sized chip, you can't hold a model. By going to wafer scale, we were able to put down a huge amount of SRAM and get the benefits of speed and enough capacity.

If you build a normal-sized chip with SRAM and you want to do a 400-billion-parameter model in inference, you might need 4,000 chips. If you want to do a DeepSeek 671, you might need 6,000 or 8,000 chips. What an administrative nightmare.

You can keep it on 1 wafer, 2 wafers, 4 wafers, or 10 wafers. You get all the benefit of the SRAM, and because you've been able to use the wafer, you get this tremendous capacity as well.

Harry Stebbings

I totally get you on HBM and the slowness of it. Why is it, then, that so much of the market just continues to use it? Forty percent of NVIDIA's revenue is using those chips for inference.

Andrew Feldman

There wasn't really, unless you went to wafer scale, a credible other choice. This is the way GPUs had always been made. It's called a graphics processing unit. That's the way they were built, and it was part of their advantage against a CPU: they were built this way.

But now there are dedicated chips like ours, and what used to be their advantage is now a weakness. That's a fun market to be in, when over a very short period of time what you're good at becomes your weakness.

Harry Stebbings

With a market cap like theirs, and with Jensen as good as he is—which I'm sure we both agree with—they must know this.

Andrew Feldman

They do know this. There aren't a lot of choices. They don't make memory, so they're a consumer of other people's memory. That's SK hynix, Samsung, Micron. There are only 3, 4, or 5 companies that make huge amounts of memory. There aren't many choices.

It's part of a complex architectural tradeoff. The flip side is that it's worked really well for them. Look at where it's taken them.

In comparison to those of us who are wafer scale, it's a small set. It's a set of 1: us. We have a real advantage against them on inference.

Harry Stebbings

How do LPUs fit into this? We've got HBM, we've got SRAM with you, and, bluntly, we have many more of them to make it work and scale. Where do LPUs fit into this mix?

Andrew Feldman

There are a lot of ways to skin a cat. Our way is different from NVIDIA's way. It's different from the TPUs, and it's different from Trainium. They're different.

Right now, and every day since August 26, when we launched inference, our way has been the fastest way across a whole set of models tested by Artificial Analysis and others.

Harry Stebbings

When we think about that speed, you said that you're 1 of 1 with wafer scale and the associated architecture. What does that mean in terms of cost? With such efficiency, is it inherently more expensive, and what does that look like from a cost profile?

Andrew Feldman

This isn't our first dance. We've been building computers for a long time, and when you make a choice like wafer scale, you have to weigh the tradeoffs.

We use less power. We use less power because 1 of the most power-hungry things on a chip is the I/O, moving data off-chip. If you're moving data off-chip frequently, you're using more power than if you can keep it in the silicon domain, on-chip.

We knew we would use less power. We knew that if you went to wafer scale, you had to solve some problems that people said were impossible to solve, like yield. We had to invent techniques that allowed us to yield wafers. In fact, we invented techniques that allow us to yield as well as, or better than, others who are building much smaller chips.

Harry Stebbings

Can I interrupt and ask what yield is, and why is it impossible to solve?

Andrew Feldman

A wafer begins as a 12-inch-diameter circle, a slice of silicon, and your chip is punched out of this. It's the way your mother might take a cookie cutter and cut out cookie dough.

During the process, at some point, just like your mom might have done, she lifts up the edges and all the little bits are removed. What's left are just the cookies. Those are your chips.

What happens is there are a set of naturally occurring flaws. It's like your mother closing her eyes and throwing up a handful of M&M's. The bigger the cookie, the higher the probability you hit an M&M. The bigger the chip, the higher the possibility that you have a flaw.

Traditionally, when you had a flaw, you threw away the chip or sold it as a less valuable part. You shut down part of the chip and sold it as a less valuable part, something called binning.

Every wafer is going to have flaws. The bigger your chip, the higher the probability you hit a flaw, and the more of the silicon is wasted when you throw it away. This is what everybody thought was known truth.

One of the things our team realized was that there are other ways to handle flaws. What if, instead, you built your computer—your processor—out of hundreds of thousands of identical tiles? If there was a flaw, you could shut down that tile and work around it. You could have a row or a column of redundant tiles that, when you needed them, you could pull in.

That had traditionally been the technique used in memory-making, and memory yields are extraordinary. It occurred to us that if we could build a processor out of hundreds of thousands of identical tiles, we could use redundancy. When there was a flaw, we could leave it there, shut it down, work around it, and pull in 1 of the redundant tiles.

That had never been done in a computer before, and that's at the heart of our architecture. It allowed us to yield and deliver whole wafers.

Nobody had ever been able to do that in the 70-year history of our industry. Really, really smart people struggled. Likely Gene Amdahl, one of the fathers of our industry, had a company called Trilogy that crashed and burned trying to do this. We figured it out.

Harry Stebbings

When you speak about being the fastest, and across all benchmarks being the fastest, what matters the most? Is it being the fastest, being the most efficient, or being the least costly? How do you think about the stack of prioritization for your customers?

Andrew Feldman

I think it varies. If you go to get a cancer diagnosis—for God forbid, your mother or your wife—I think 93% accuracy is just plain not as good as 94% accuracy. You pay a lot and wait another week to understand what the accuracy is. You pay a lot.

On the other hand, if you want Llama 405B to generate data to help you tune Llama 70B, maybe you can wait a few days, 3 days, or a week more. There's no urgency there.

If you want an answer from Perplexity, you don't want to wait 45 seconds for a search answer. You don't want to wait in a chat. You don't want to wait 3 minutes for R1 on GPUs to give you an answer.

In interactive mode, milliseconds matter. In interactive mode, what Google showed years ago was that you can destroy your user's attention with milliseconds of delay. Being the fastest matters in that domain.

You have to be thoughtful and say that in some cases being the fastest doesn't matter. We'll call those batch. Maybe cheapest matters there. In other domains, there is no search if you have to wait 8 minutes to get an answer. That's not a product.

When you go fast, a whole set of new opportunities open up. Netflix used to mail DVDs. That's what happened when the internet was slow: they'd mail DVDs. I look young, Andrew, but I'm not that young. I remember Blockbuster.

First, we used to drive to Blockbuster to get a DVD or a video. Then Netflix was mailing them to us. Then we got broadband, and suddenly Amazon is a studio. It changed everything, and speed in inference does the same thing.

Harry Stebbings

When we chatted before, you gave this great equation for inference. What was the equation? It was really helpful for me in understanding it.

Andrew Feldman

It begins with the following: training makes AI. That's how we make AI. Inference is how we use or consume AI.

Understanding how big the inference market is means understanding the number of people who are going to use it, how often they're going to use it, and how much compute each use takes.

Right now, we're in this rare time where the number of people using AI is growing, the frequency with which they use it is growing, and the amount of compute used in each instance of use is growing. That's why you're getting this extraordinary growth, and that's why it's off the charts right now.

Harry Stebbings

When we think about the distribution of resources between training and inference, what will that look like in 5 years? We've seen a lot of focus go to training and not as much go to inference. What does that look like?

Andrew Feldman

What we made in AI until the middle of 2024 was a novelty. What we made in AI late in 2024 began to be useful.

Harry Stebbings

What was the turning point?

Andrew Feldman

If you look at the models, they became useful. ChatGPT wasn't really a technical innovation; it was a user-interface invention. It gave more people access, but we didn't really know what to do with it right away. It was cool. That's what I mean by novelty: “Whoa, this is cool.”

Now, if your marketing team isn't on an LLM each person several times a day, they're not doing their jobs. The difference between “It's cool” and “This is part of everyday workflow” is what changed, starting sometime in Q4 last year and running into this year.

AI became useful not just to a select group in Silicon Valley, but to my dad, my brothers, doctors, and ordinary people who aren't buried in the Silicon Valley discussion. When you get them, the market is ripping.

Harry Stebbings

Do you not still think we're incredibly early? Going back to your point, how many times bigger are we in 5 years? Are we 100 times bigger? Are we 1,000 times bigger?

Andrew Feldman

I think we're way over 100 times bigger.

Harry Stebbings

What does that mean in terms of what we need to equip ourselves to deliver these? They're incredibly energy-intensive, it's incredibly difficult, and our industry consumes a lot of power.

6. Why is it Total BS That We Have Hit Scaling Laws?

We're seeing some water usage come down, but are we equipped from an energy and data-center standpoint to deliver the inference requirements for a population that is as AI-hungry as we are?

Andrew Feldman

The first thing is to admit that this is a power-intensive problem. Our industry consumes an enormous amount of power.

The second thing to say is that the burden is on us to deliver exceptional value as an industry. You take both the good and the bad. In order to make it worthwhile from a societal perspective to expand all this power, we better deliver the goods.

We better use AI to find cures for diseases. We better use AI to solve a bunch of different societal problems. That's the macro view.

Do I think we're equipped? I think we're in a very unusual situation in the US, where we have plenty of power, but it's in all the wrong places. We have power in Niagara. What we don't have is power where you want to build data centers, where we have good fiber.

What we also don't have is a national way to relax the local regulations that make getting power difficult. When you go to Silicon Valley, if you want to build a data center, you're dealing with local government and vested interests. That's not an efficient way to decide if you want to build a power plant or put a new data center in, especially if it's large.

I think those places that have removed some of that burden—for example, through taxes—are getting a huge number of data centers built.

Harry Stebbings

When I spoke to Jonathan at Gro, he said there were a huge number of data centers being built that weren't really equipped properly. We've seen this massive supply of data centers that are done by tourists, so to speak, and that is a massive problem. The provisioning of these data centers isn't there. Do you agree?

Andrew Feldman

A data center is, to begin with, a construction project. It's access to power, a construction project, and a design and engineering component.

I think there's been a huge push for new-construction data centers. We don't know if they're going to be good enough. Many of them will be fine.

The guys who were there early were some of the Bitcoin-mining companies, like TeraWulf, the guys at Crusoe, and others. There were guys in Europe who were early in building buildings near low-cost power in order to run compute that used a lot of power. They are some of the leaders now in some of the largest projects.

Those are certainly not tourists. They're extremely sophisticated data-center builders. There are some tourists, but there are a lot of very knowledgeable data-center builders building huge facilities right now—gigawatt-scale facilities, both domestically and internationally.

Harry Stebbings

How do you think about how the cost of inference goes down with the surge of demand that we mentioned—over 100 times? Does the price reduce 100 times? Does it follow Moore's law continuously? How do we think about the ever-reducing price of inference?

Andrew Feldman

The cost of inference is built up of several pieces. There's the power and space consumed to generate the response. That's a data-center cost and an OPEX item.

Second, there's the cost of the computer. We can drive down the cost of the computers with each generation by driving up their performance.

The other thing we can do is develop more efficient algorithms. Our AI algorithms today aren't particularly efficient. In a GPU, most of the time it's doing inference, it's 5% or 7% utilized. That means it's 95% or 93% wasted.

Over time, I think that as an industry we get better at things. We can drive the cost of compute down, build more efficient data centers with lower PUEs, and make our algorithms more efficient, so that utilization on our now-cheaper computers is higher.

You get a higher percentage of the maximum number of FLOPS. You get more tokens per unit time for the same power.

Harry Stebbings

When you look at the inefficiency of the algorithms, and what that means for the utilization of the chips, why are people suggesting that we're at scaling laws already? That seems to suggest there is so much room for improvement.

How do you think about what you just said in conjunction with the idea that we're hitting this asymptote point?

Andrew Feldman

I don't think there's a lot of debate among senior ML thinkers that we have tremendous room for algorithmic improvement. I don't think there's a lot of debate there.

There's even debate about whether the scaling laws are over, whether we've run out of mojo to keep making or gathering data to fill these ever-bigger models. But OpenAI's work on o1 shows me that the scaling law, certainly for inference, is fully functional. The more compute you put on inference, the better answer you get.

Many of the leading models are now MoEs, so they're not presenting all of the weights to each token. That's 1 way to do it: present the important stuff, not the unimportant stuff.

There are other ways to do it that we will invent and learn over time. We have human models that aren't all-to-all connected. Many of our models today are all-to-all connected. That's a lot of unnecessary connections, connections that don't produce anything but that we still end up doing math over.

Harry Stebbings

What does “all-to-all connected” mean?

Andrew Feldman

In many of the layers in a neural network, every element is connected to every other one. That's not actually the way the learning happens. Some are more valuable, and some are not valuable at all.

Imagine you're going to read 50 books because you want to learn something. You can read all 50 books, or you could read the 3 books that are really important, or you could read summaries of the 3 books that are the most important.

The problem is that we don't know which they are at the beginning. There's a process that you could learn. There are things called dropout and all these other techniques to use sparsity to help solve these problems.

We are early in the evolution of AI, and that plays right into this point that we'll get better at these algorithms. Transformers aren't the end of the world. We'll get better. Better will mean faster, more accurate, and more efficient.

That's what's exciting about an ever-changing industry. That's why I'm not in all these other industries that don't change quickly. They're the same 9 years ago as they are today.

Harry Stebbings

This show is kind of strange to me because I speak to a lot of people, and they think about the 3 pillars—compute, algorithms, and data—and the common refrain is that we're actually very far along in all of them.

When I hear you, it's actually very exciting. I think they're wrong.

Andrew Feldman

I think they're wrong. I don't think we're very far along, and it's very difficult to say that we're early in an industry but far along on all of its underpinnings. I think we are early in all of them.

Harry Stebbings

If we take them 1 by 1, in 5 years, how much synthetic versus human data will be used to train models? If you had to put a percentage on it?

Andrew Feldman

Almost all synthetic.

Harry Stebbings

And is the utility value of synthetic data the same as human data?

Andrew Feldman

I think so. When you teach a pilot to fly in a simulator, there is a lot of potential data that isn't very useful in teaching a pilot to fly. They spend a lot of time going straight and doing nothing.

Takeoffs and landings are where you want to spend your time, and that's why, when we put pilots in simulators, that's what we have them doing. In simulators, we can create data where engines blow, where there are a whole set of problems, and where learning can take place. That's simulated data.

In the same way, when we think about creating data—whether it's for self-driving or other forms of AI—we want the data that's hard to gather. Otherwise, we just have a bunch of data of people driving straight on a freeway. That's not difficult. We've been able to do that for a decade.

What we want is an unprotected left turn in the snow. It's snowing, it's hard to see, and you've got an unprotected left turn. That's a difficult thing, and you want that thousands or millions of different ways. That's where the synthetic data comes along: to fill in the empty parts where it's really expensive or painful to get that type of data.

Think of the pilot. You want them spending a huge amount of time on things that are rare in their training. It's the same with a surgeon: a huge amount of time on things that are rare. Most of the time it's carpentry, but their expertise is only needed when something rare happens.

That's when their mettle is shown, when the unexpected occurs. I think we will get better synthetic data by a great deal.

Harry Stebbings

I get it from a consumer perspective and from an expectations perspective. If we move the needle on compute, algorithms, and data, what does that mean for the experience of AI?

Andrew Feldman

It gets faster and cheaper. Faster and cheaper is the first answer.

The second is that when things become faster and cheaper, new applications emerge. It's used everywhere.

When computers became faster and cheaper, suddenly they were in cars, then they were in your pocket, then they were in your dishwasher and your TV. We were saying 30 years ago, “I need a computer in my TV? Are you kidding me? I need one in my pocket?”

Now you've got powerful computers in your pocket, in your TV, in your kids' toys, and in the car. That's what happens. Diffusion of innovation accelerates when you make things faster and cheaper.

Harry Stebbings

This is Jevons's paradox and Satya's belief, isn't it?

Andrew Feldman

I know that in the VC community you have to cite 19th-century English economists.

Harry Stebbings

I'm English. I'm English. Come on. If I'm not allowed to cite an English philosopher, what am I here for?

Andrew Feldman

I think there are very few examples in our industry—actually none in compute in 50 years—in which, by making things cheaper and faster, the market got smaller. The market always gets bigger. It always does.

Harry Stebbings

From an architectural standpoint, you mentioned transformers. Is there a world where we move past transformers?

Andrew Feldman

Transformers, 100%. We won't be as dependent on transformers in 3 years or 5 years as we are now—100%. They're not the end-all and be-all.

Harry Stebbings

Why? What will replace them, and what does that look like?

Andrew Feldman

I don't know whether they're going to be state-space models or other types of models. What I know for sure is that innovation doesn't stop, and the transformer has some weaknesses that people are desperate to overcome.

There's a quadratic effect in the attention head. There are all sorts of things that could be improved. But it's pretty darn good now. It's the best we have, and that's what you run with. You run with the best you have, and the minute it's not the best you have, you drop it in favor of the best you have.

I think that's what we're seeing. We're seeing a large number of innovative companies designing models.

Harry Stebbings

What DeepSeek showed us is that you don't need 5,000 people and billions of dollars a year. You can do it with 200 smart people and more hardware than DeepSeek said they had, but less hardware than others had.

Were you very impressed with DeepSeek, and what impressed you most?

Andrew Feldman

I think it was the result of focused engineering, and that impressed me. It was designed to be better. They weren't confused about being model intellectuals, and they weren't confused about whether it was important to break new ground. They were interested in being better.

From an invention standpoint, that's a little boring. From an engineering standpoint, that was a sweet effort. They really built a model that was just plain better at many, many things, and that's cool. I like good engineering projects.

They chose to announce it right around Trump's inauguration, and the politics of it are a separate matter. We can talk about that later.

Harry Stebbings

Did distillation rile people up?

Andrew Feldman

I don't think distillation is wrong. Is summarization wrong? I'm a VC. Are you kidding me? That's what we do. If you didn't summarize, you wouldn't know anything.

7. What Specifically Was So Impressive About DeepSeek?

Harry Stebbings

Exactly. That's exactly right. I don't think distillation is wrong. If distillation is wrong, then certainly using people's copyrighted data is wrong. That's the problem. You've got to be a little bit consistent.

Andrew Feldman

I think neither is wrong, actually, but you have to be consistent.

Harry Stebbings

The thing with it, bluntly, is that DeepSeek is open. Everything that they innovated on, OpenAI can learn from and take too.

8. Why is Distillation Not Wrong and OpenAI Need to Look in the Mirror?

Andrew Feldman

I think there are few examples of an open-source anything having the sort of immediate impact that model had. That model had a giant impact in a technical community of really smart people.

There are very few examples of other open-source software projects that had that type of impact in that amount of time. Usually, you're in the business of betting on these guys: they ramp up, and they go from 10,000 users to 100,000 users to 1 million users. Then you better start a company around that and get those graduate students.

This had a loud boom in the industry immediately. It was, “Whoo, the thing!”

9. Where Will Value Accrue in a World of AI?

Harry Stebbings

The thing I have to think about as a venture investor is where enduring and defensible value is, and how I get in early and build that over time. In hardware, that's well understood.

But on the model side, do you think there is value when you look at the sheer number of players with relatively comparable models?

Andrew Feldman

To demonstrate enduring value, you need both immediate value and a trajectory for more.

The problem in some industries is that you're capable of demonstrating a leadership position for a short period, and then someone else—maybe the next generation—generates the next one, and the next generation generates the next one.

I think that, in the software world, you end up competing against other people's release cadences. You're 4 months ahead; they're 6 months ahead. If that's really where you are, there's not a lot of value.

But if you can stay at the top over years—even if you're not the best, even if you're in the top decile over years, while the people above you are changing constantly—I think there's a lot of value.

Very large Silicon Valley companies have been built with technology that was not the most compelling. It might have started as the most compelling technology, and then it got to a point where it was good enough, easy enough to use, and well distributed. That's when you're at the mature market.

10. How Will NVIDIA’s Market Position Change Over the Next Five Years?

We're a long way from there right now. Right now, we're in the early phases. You characterized my position exactly right: data, compute, and algorithms. I think we have a ton of room for improvement on all of them.

Harry Stebbings

You said that computing hardware is where the value is. How does that value distribution shake out? We've obviously got the 800-pound gorilla that is NVIDIA. How do you think about how the distribution of value shakes out in hardware and compute over the next 5 years?

Andrew Feldman

Historically, 1 of the barriers to entry was the capital intensity of a project. In the world of building chips, there are both scarce resources in expertise and high expense.

Historically, it hasn't fit comfortably in a software company. The things that modern software companies value aren't entirely conducive to chip-making.

When I look down the road, I think that people who build systems endure. Cisco and Juniper endure. Chipmakers have endured. There's a reason Apple and NVIDIA are among the most valuable companies on Earth. What they do is hard, and I think that's why it's worth challenging.

If it weren't hard, enormous, and difficult, why spend time being the underdog and challenging it?

Harry Stebbings

A lot of people place defensibility around NVIDIA's CUDA lock-in. To what extent is that real versus hype in inference?

Andrew Feldman

In inference, it's not real at all. There's no CUDA lock-in in inference. You can move from OpenAI on an NVIDIA GPU to Cerebras, to the Fireworks service on something else, to Together, to Perplexity with 10 keystrokes.

Anybody who actually uses AI knows there's no CUDA lock-in in inference.

There was a fundamental effort to disintermediate CUDA, first by Google with TensorFlow and by some graduate students with Caffe, and later by Google with TensorFlow and Facebook, or Meta, with PyTorch.

Today, most AI is written in PyTorch. You ought to be able to compile it and run it on your hardware.

NVIDIA has many moats. When you're a dominant market-share leader, that in itself is a moat. Being the default solution is a moat. Everybody learns to think about AI in your structures. Those are moats.

The software—compilers are hard, but they're tractable.

Harry Stebbings

I completely agree with you that being the leader is a moat in itself. It's never talked about that way.

Look at Intel. Intel has made nearly a decade of catastrophic decisions until hiring Lip-Bu Tan, and they still own 80% of the x86 market. AMD has worked up to perhaps 25% or 30%. After a decade of screwing up, Intel only lost 20% of its share. That's a moat.

Andrew Feldman

That moat is just unbelievable. You can make a bunch of bad decisions for a decade and only lose 20% share. That's extraordinary.

I'm a huge fan of Lip-Bu Tan. He's an investor in our company, and I wish him well. I think if anybody can change that company, he can.

I think we rarely talk about what being the market-share leader means in terms of a moat. As a challenger, we have to think about it exactly, because it's exactly those characteristics of the moat that we need to get over.

Harry Stebbings

In 5 years, is it Uber, or is it like AWS and cloud? Cloud is an interesting market where a couple of players have relative segments—25% or 30%—and it's shared relatively evenly between them. Or is it 1 like Uber, where Uber has 90%, Lyft has 5%, and alternative providers have the other 5%?

Andrew Feldman

I think it's going to be between those 2. In 5 years, NVIDIA is going to have 60%, somewhere between 50% and 60% of the market. Right now, they have approximately all of it.

Harry Stebbings

Of NVIDIA's usage, what percentage will be training versus inference?

Andrew Feldman

They'll continue to have a meaningful business on both sides. They're exceptional at training. They will not roll over and play dead in inference.

They're a world-class company. They've had 1 of the great decades of any company in history. From 2014, when they were worth $10 billion, to where they are right now, it's 1 of the great decades in corporate history.

I don't think they're going to roll over and say, “We're not going to be in the inference market.” That's not going to happen. They're going to have a meaningful share, but the market is growing, and we'll have a piece. Others will have a piece. I think there'll be some very big companies made in this 100× growth.

Harry Stebbings

Do you think chip providers will be far larger than model providers in terms of enterprise value in the 5-year timeframe?

Andrew Feldman

Yes.

Harry Stebbings

How does that prediction change on a different timeline?

Andrew Feldman

In a shorter timeline, when you price an option, variance and uncertainty increase the option's value. If you look at the way Black-Scholes works, or at any option-pricing model, uncertainty and variability are friends of the value of the option.

When people are paying these extraordinarily high prices for model companies right now, I think part of that is this extraordinary uncertainty and wild variance. In the shorter run, it might not be the case.

In the longer run, as markets mature and we begin to understand the value of these models, their businesses, and their long-term net profitability, we'll have a better understanding.

11. Why is the CUDA Locking for NVIDIA BS? What is Their Weakness?

What did Warren Buffett say about markets? In the short term, they're a voting mechanism, and in the long term, they're a weighing mechanism. At some point, the weighing kicks in. Usually, it's in the public markets, and then investors say, “Which is likely to give me better growth in the future?”

Harry Stebbings

You mentioned the word “public.” I do want to hone in on your business. You're cash-flow positive in a world where everyone else literally bleeds cash.

Help me understand: how did you become cash-flow positive when everyone else is bleeding or hemorrhaging cash?

Andrew Feldman

Traditionally, gross margins were a measure of technical differentiation. If you're running a negative-gross-margin business, I think it speaks for itself. You're selling a commodity. Your value creation isn't being recognized in the market.

12. Why is Trump Better for Business than Biden?

I think our technology is creating an opportunity for us to maintain margins where some others can't.

Harry Stebbings

A lot of your revenue is concentrated in the G42 deal. To what extent is that a strength or a weakness?

Andrew Feldman

It's both. The way you catch 3 large customers is to catch 1 first. The way you build 3 large strategic partners is to learn to be a strategic partner. That's a learned skill.

We didn't arrive knowing how to be a strategic partner at G42. Now that we've worked at it, it's a muscle we can replicate. We could be a better partner to any of a dozen different companies in the world.

Harry Stebbings

What have you learned in the G42 relationship-building process that makes Cerebras a good partner in a way that you weren't before?

Andrew Feldman

We've deployed tens of exaflops of compute, vastly more than anybody else that isn't AMD or NVIDIA. That's a huge amount of compute.

Our software has been hardened on some of the largest AI clusters in the world. We've gone through the growing pains of increasing manufacturing 2×, 5×, and 2× again through unbelievable growth in manufacturing.

We've worked with our supply-chain partners to be sure that they're ready for this extraordinary growth.

When you work with a strategic partner of this size, your organization comes out different on the other side. There are things you've learned and mistakes you've made.

I hadn't done a big relationship in the Middle East. There was a huge amount to learn. I think you come out a much better company and much better prepared to do business with a hyperscaler, another massive partner, or another sovereign.

But it takes real work, and your team has to learn.

Harry Stebbings

You said you come out better. Why go public when you did? When it happened, I thought it seemed preemptive, respectfully.

My question now to companies is: why go public at all? There is so much private capital. The decisions have shown very clearly that you can stay private for a lot longer than you planned to. Databricks has certainly shown that.

Those were historically public-market valuations, and the valuations that Anthropic, OpenAI, and some of the others are getting are historically public-market-only valuations. Your S-1 is live; anyone can read it. I wouldn't want people reading mine.

Andrew Feldman

We have nothing to hide.

Harry Stebbings

No, but your competitors have asymmetric information.

Andrew Feldman

Yes, we've got asymmetric technology. I think you have to be pretty transparent to be public.

You have to be ready organizationally. You have to be ready with your processes. You need to be ready to forecast and predict, and to be held accountable in a way that private companies historically haven't been.

We think there's tremendous value. We think we'll be among the first in the category. We think some of our largest targets would have a stated preference for doing business with public companies. Large enterprises in the US have done that historically.

Those were some of the reasons that led us to it.

Harry Stebbings

How many G42 relationships shall you have in the next 24 months? How fast can you ramp them?

Andrew Feldman

That's a good question. Several.

Harry Stebbings

Remind me, how big is the G42 deal?

Andrew Feldman

It was 87% of revenue.

Harry Stebbings

I know it was big.

Andrew Feldman

When we announced it, some estimated it was north of $1 billion.

Harry Stebbings

Well done. That must be a bit of a high five.

Andrew Feldman

There's tremendous excitement, and then there's every entrepreneur's reality: I have to make a lot more gear.

You make a list of your top 10 vendors and fly it to them all, saying, “Big orders are coming. Be ready.” You work with all your partners to get ready because you need to make a great deal more stuff.

That's 1 of the real differences between hardware and software. When we grow fast, the number of people you need to work with in your supply chain, and the amount of collaboration that needs to happen, is truly extraordinary.

Harry Stebbings

Are NVIDIA going to have a cluster of unhappy customers who, bluntly, have waited so long for chips that by the time they get them, the chips are outdated? Are they going to say, “What happened?”

Andrew Feldman

All of that is an opportunity for us and others. Being a market-share leader isn't easy either.

When the bully falls, everybody wants to give him a kick. A lot of that happened at Intel. They'd been the dominant player, and when they fell, everybody was happy to jump in and kick them when they were down.

I think there's a real opportunity in the potential for NVIDIA customer unhappiness. For those of us competing with them, if you can't get your gear, you may as well test somebody else. That's a huge opening.

Harry Stebbings

In hardware, you mentioned the complexity. Are export controls being implemented properly? Do you think they're a good idea?

13. Quick-Fire Round

Everyone was looking at DeepSeek and saying, “How did this happen? They must have stolen chips. How could this be?” What do you think about that?

Andrew Feldman

It turns out that they probably did use chips in Singapore.

I think managing software compliance and managing hardware compliance are extremely different things because their vectors of diffusion are different. There's a different weight.

If you sell a server that weighs 500 or 600 pounds and arrives on a pallet, you can go visit it. If you want to deploy it in Kazakhstan, you can put it in a data center and have somebody from the embassy visit it and take photos once a month. It's not going anywhere.

You can keep track of who uses it and provide logs. That's much harder with software, and open source is a whole other level.

That's the first observation. The second is that I got to know the leadership in Commerce in the previous administration. I didn't always agree with their policies, but it is a world of unintended consequences.

You sought to limit Chinese access to EDA tools to delay the growth of a Chinese chip market, and US venture capitalists backed tons of Chinese companies in Shenzhen to build EDA tools. This is an unbelievably slippery, dynamic, challenging problem.

I don't know if it's a tractable problem. Delaying another nation's progress on a technical trajectory is an enormously challenging thing.

I came to appreciate just how difficult it was for well-meaning people to predict the impact of policy during the last 2 years.

Harry Stebbings

Do you think this administration is better for AI than the prior administration?

Andrew Feldman

I don't think there's any doubt that's the case.

Harry Stebbings

What makes you say that?

Andrew Feldman

The past administration lined itself up against Big Tech, and that was a mistake.

AI is also in a different place, so it's easier to be for it. It's less scary now than it was in 2021. We have a better picture of the trajectory, both the risks and the benefits.

I think this administration had the foresight to put an AI czar or leader in place as a focal point for discussions. It's probably net a fair bit better.

Harry Stebbings

You said it's very challenging to hinder a nation's development, adoption, and progression of a technology. Respectfully, you chose not to sell to China. Why was that, and does that not go against the difficulty of hindering progression?

Andrew Feldman

I have a very simple rule, and I encourage your team to use it. You don't need a big handbook to help you make good decisions in a company. Just ask yourself, “Would my mother be proud?”

Would she be proud if I did this? Would she be proud if I explained the exact situation? Would she look at me and say, “I'm proud you're doing this, son?”

I asked myself that, and I came to believe that the deal on the table wouldn't be used for good. I wasn't comfortable with that. I wouldn't have been able to explain it to my mother.

That's a moral compass. It wouldn't have been used for good.

Harry Stebbings

I'm naive. What would they use it for?

Andrew Feldman

They could use it to power drones. Some use it for facial recognition to identify minorities for persecution, or to build military equipment—to do things that I either couldn't see or that, from what I saw, didn't feel right.

14. Do We Underestimate China in a World of AI?

It's more important than money.

Harry Stebbings

Do you think we fundamentally underestimate Chinese capabilities?

Andrew Feldman

100%. It is 1 of the most obvious and frequent errors in judgment: you underestimate the other side.

You have to look carefully at what they're doing. Their investment in infrastructure has been extraordinary. The rate at which they generate engineering talent is exceptional.

The government's ability to have a policy and implement it is extraordinary. They're not a democracy; they weren't designed to have checks and balances there.

The funding that flowed into the development of AI technology was significant. Their venture capitalists were backed by their government. They have national champion companies. They've developed a belt-and-suspenders strategy to make much of the developing world dependent on them and their technologies.

They absolutely should not be underestimated. They have a lot of people, and we see a tiny fraction of it.

I think they have produced industrial policy that has moved their nation forward.

Harry Stebbings

What was the most significant part of that, do you think?

Andrew Feldman

The creation of economic zones like Shenzhen was clearly a visionary move. They knew that their own system was in the way, so they created zones that relaxed their own system.

Harry Stebbings

Could the US learn from them in that way?

Andrew Feldman

We did some of the same things in the 1st Trump administration. What did we do? We relaxed our own rules in the development of vaccines. We knew that, in that time, it would be very difficult to go through the steps that we always go through, and we tried to implement thoughtful shortcuts, or workarounds.

Why are they committed to trains as a mode of transportation, and we can't build a decent train system in the US or in California? Why can't we build infrastructure when the rest of the world can build extraordinary high-speed trains linking important cities?

Why do we have 3 different standards for train rails? Why are our bridges and freeways in disarray?

I think those are questions we have to ask ourselves when we see other people doing it differently.

If you watch a good football team, you say, “That's interesting offense.” You're not thinking, “How could our team learn? What could we do? Why did that work? What was it about the people they had, the talent, or the structure that made that a successful series of plays?”

What can I take away from that? How can that inspire me to do better?

I'm always looking for inspiration in others, competitors, and partners. Some of our partners at G42 have an unbelievable work ethic. It inspires me. The scope of the challenge they've undertaken inspires me.

Harry Stebbings

What do you believe that most around you disbelieve?

Andrew Feldman

I think we're closer to peace in the Middle East than people believe.

There is a rise of a moderate, business-focused Arab state that wasn't there 25 or 30 years ago. If you visit the UAE, Qatar, or even Saudi Arabia, what you see is amazing transformation.

I think there's a desire to be included in the West in their own way, but also to enjoy the benefits of it. We are closer than people may think.

Harry Stebbings

What's the most underrated threat to NVIDIA's market-share dominance?

Andrew Feldman

The fundamental architecture of the GPU with off-chip memory is not great for inference. They will continue to do well in inference, but they can be beaten, and I think they know it.

Harry Stebbings

What's a crazy AI prediction you have that most people would call science fiction?

Andrew Feldman

Dario at Anthropic says that we'll live to 150. I don't think we're going to live to 150.

I don't think that 90% of our code will be written by machines this year. But I do think that within a year or 2, AI's penetration will be approximately the same as telephones—cell phones.

Harry Stebbings

What have you changed your mind on in the last 12 months?

Andrew Feldman

There are lots of things. Many decisions I made turned out to be wrong.

There are 2 ways you can be wrong. You can actively be wrong, or you can fight against what was right.

In 2016, JP, 1 of our co-founders and chief system architect, laid out a plan that would have us doing water cooling for our systems. Nobody else was doing it, and I fought so hard. I was so wrong.

JP was right. About a year or 2 later, Google announced that the TPUs were going to be water-cooled. We were 1st, and now NVIDIA is only selling water-cooled parts.

I was dead wrong, and JP was right.

When you make a lot of decisions every day, there are many instances where you're wrong. I've been wrong about people. People I thought were pretty good turned out to be extraordinary. People I thought would be extraordinary were really smart but couldn't finish projects or get things done.

If you're not prepared to be wrong a fair bit, you ought not to be making a lot of decisions, because it comes with the territory.

Harry Stebbings

As a venture capitalist, I'm never wrong, so—

Andrew Feldman

As a venture capitalist, you're wrong 9 times out of 10, and everybody forgets as long as you're really right.

Harry Stebbings

I get a picture of you signing the term sheet with me.

Andrew Feldman

Good. Yours is a perfect industry in which nobody cares about the average. On average, you're wrong all the time. What they care about is the occasional time you're really right, and that moves a fund.

That's different from being a CEO. I think we've got to be mostly right most of the time, but if you're making a lot of decisions, you're still making a ton of mistakes.

Harry Stebbings

This is your 5th startup. You are a sucker for punishment, aren't you? Really—5 times? Did you not get beaten alive enough?

My question is about the value of serial entrepreneurship. I've spoken to many people who don't believe in it, respectfully. How do you think about the inherent benefits of having done it 4 times before?

Andrew Feldman

If you're in a business in which running a business is a benefit, then experience matters a great deal.

If you're in a business in which you look like your customer, there was a reason social networks were started by people right out of college or in college. Dating is top of their mind. They look like their customers, and that was more important than knowing anything about running a business.

In that environment, it will select for people who are of the demographic their customers are. They know that backwards and forwards.

But if you want to have a business that has manufacturing, a supply chain, and hundreds or thousands of engineers managed to a timeline and a schedule, I don't think anybody would turn around your statement and say with a straight face, “What I'm looking for is an engineering leader with no experience.”

You wouldn't want somebody who had only led a team of 400 or 500 people. You'd want somebody who had experienced the challenges of growth.

Harry Stebbings

No, I don't want somebody who's led a team of 400 or 500. What I'm looking for is somebody with no experience. Naivety is a bonus here.

Andrew Feldman

The people who sell that are sometimes consultants. “My guys have no experience in your industry, so they aren't biased.” Maybe a little bit of experience in the industry would help.

Harry Stebbings

Where are people investing today in AI across the stack where you're thinking, “Why is so much cash going to that part?” I'm not asking for a company; I mean a category.

Andrew Feldman

Sometimes money needs to find a home. Some people have raised really big funds, and they have to find a home for that money. Some people don't like to be left out. They're willing to make investments for status purposes or other reasons that don't seem to make sense.

I haven't thought about it in detail. I think there are some underappreciated places of investment.

In the chip world, the sub-milliwatt, tiny little chips that live next to sensors and do inference are an extremely interesting market. These are tiny things that will only send back useful data, and they'll sell in enormous volume.

It's not a part of the market I love to play in. I like to build bigger things and sell them to the data center. But I think that part is extremely interesting. I think it will be fundamental for robotics.

That's an area where I think the opportunity is extremely underappreciated.

Harry Stebbings

If we think about Cerebras in 10 years, where do you envision the business? If everything goes well, where are we in business having that conversation?

Andrew Feldman

10 years ago, NVIDIA was worth $10 billion, so that's a long run in our world right now.

In 3 to 5 years, I would like our technology to have been used to solve 2 important societal problems. I would like it to have been used to find a therapeutic for an affliction that impacts more than 1 million people a year.

I would like our inference to be powering a collection of apps that don't exist today. I would like a meaningful portion of the population in the US and Europe to inadvertently use our technology—to use something that we power without even knowing it.

Harry Stebbings

Andrew, I've wanted to make this show happen for a long time. As I said, I've heard so many good things from Marc for many years.

There have been so many requests to have you on the show. My team was saying, “Just get Andrew on the show, Harry.” I was like, “Okay, okay.”

I tweeted it, obviously, which is how we got this conversation. You tweeted it, and around 40 people sent me messages saying, “How come you're avoiding Harry? How come he has to tweet it?” I was just like, “All right, just call me.”

Andrew Feldman

It's good. Send me a note. I'm happy to come on.

Really thoughtful questions, Harry. Really thoughtful and interesting. It was a really good conversation.

Andrew Feldman, Cerebras Co-Founder and CEO: The AI Chip Wars & The Plan to Break Nvidia's Dominance | BidClub