[BidClub_]
Sequoia Capital · · 70 min

Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis

Dylan PatelShaun MaguireSonya Huang

YouTube
TL;DR
  • The core call: co-design is the 100x. Optimize hardware, systems software, and model separately and three 2x's multiply to 8x; co-optimize across all layers and "it's actually 100x." The consequence is model-silicon lock-in: DeepSeek V3's expert shapes were built for Hopper (V4's for Blackwell and Huawei's chip), so "TPUs suck at running DeepSeek" — and "the way OpenAI's models are headed, it would be a terrible decision for them to use TPUs," with the mirror image true for Anthropic and Google training on GPUs.
  • The CUDA moat isn't CUDA anymore. Models write kernels now — "all software gets commoditized" — so the moat has migrated to the ecosystem: Chinese open models (DeepSeek, Kimi, Alibaba) are co-designed for Nvidia GPUs, dragging every downstream inference API and RL company onto them. If Google open-sourced really good models, the same gravity would pull toward TPUs — hence Gemma.
  • The compute crunch is demand-led and durable. 20 GW comes online this year, 30+ GW next even counting delays, but Fable/Mythos 5's TAM is "not just 2x that of Opus" while world compute didn't double in the seven-odd months since Opus 4.5. Anthropic was net-income profitable in Q2 ex-stock comp, with Opus 4.8 API token margins "north of 80%" — so it can rent GPUs above market and still print: "whatever price I want to pay, I can pay."
  • Gigawatts aren't fungible. Trainium rents at sub-$10B per GW, GPUs at $12–13B, the SpaceX–Google deal at ~$25B per GW; colocation has gone from $60/kW/month to $120–160, up to $200. Google stuffs ~1.5 GW of hardware into a 1 GW shell by sloshing power, and "a gigawatt given to Anthropic is objectively worth more revenue than a gigawatt given to OpenAI."
  • Everyone builds ASICs, but general-purpose survives — labs "literally don't know what architecture they're going to be doing in a year," so specialized silicon risks racing to a local minimum. Google runs three different-architecture TPU programs yet still pays xAI $11 per GPU-hour. Cerebras's risk: dollars concentrate in the best model, and if frontier models hit 10T+ parameters with million-token context, SRAM-based chips can't fit them.
  • Jensen is engineering a multipolar world: a world where hyperscalers own all compute or three closed labs own all models "is one in which he's screwed," so he backstops neoclouds and neolabs and loves Chinese labs — "you throw a bunch of bait into the water and the best fish will figure out and survive."
  • Space compute is real but not yet: sub-1% of compute by 2030, but "more than half of incremental compute" by 2040. Scale forecast: OpenAI plus Anthropic alone exceed 100 GW combined by 2030, terawatts by 2040, with inference "many percentage points of GDP."
  • The bear case he respects vs the one that infuriates him: Sonya's leverage worry (a Crusoe customer just asked to halt a buildout) is live if models' useful work stops expanding faster than compute — "that tide turns." But "AI has no ROI" infuriates Dylan, while Shaun's model-progress denial is captured by: "the line has been up and to the right in terms of capabilities this entire time."
Digest · the substance, structured for research

1. From motel step-stool to Starcraft grandmaster — the making of an obsessive

  • The origin story as told: his parents ran a motel and a gas station, and he was too short to reach the cigarettes — so he learned to pre-position the step stool by who walked in. "The first neural network I trained" was racially and visually profiling people based on when they enter the gas station, which cigarette to pick. The small-business economics tinge never left.
  • His Xbox 360 hit the red ring of death; he opened it up and shorted the temperature sensor — "opened Pandora's box." By 12 he was moderating Reddit hardware forums, and even as a teen argued Nvidia over the crowd-favorite AMD on margins: "they use a smaller chip to get better performance at better power efficiency."
  • The pattern is obsession: at one point grandmaster on the North American Starcraft 2 ladder. "Obsession is good." Grades? "Fine enough for Asian parents."

2. The 2020 crash-out behind SemiAnalysis

  • The founding was a pile-up: he got "screwed out of a bonus" after making his quant firm well over $10M of risk-free revenue, his grandmother developed dementia and died after a fall, COVID hit. Then an internet argument got him doxxed — he stopped posting for three weeks, asked "what am I doing? Why do I care?", and launched the SemiAnalysis blog on his 24th birthday.
  • He was then effectively homeless from mid-2020 through 2024: a truck-and-tent tour of national parks reading semiconductor textbooks, then Latam, then 40+ conferences a year anywhere in the supply chain. SPIE lithography humbled him — understood ~10% the first time, half the second, 75% the third: "some parts of the supply chain are so arcane and so deep... even now I don't understand everything."
  • The lore that stuck, from a Japanese guy in broken English: in the 1980s the world's only factory for one chemical burned down and memory prices doubled or tripled. "Not too different from today."
  • Today SemiAnalysis is ~90 people — a big chunk are supply-chain technologists and engineers, and another big chunk are people formerly at hedge funds — and the internal fights ("that doesn't matter" / "but cost" / "this technology is the coolest") happen organically. On the $100M revenue rumor: "It's as accurate as the information is."

3. InferenceX tracks 60x model-cost declines

  • Point-in-time benchmarks die on arrival: new models weekly, vLLM/SGLang/PyTorch updating twice a week — "relentless breakthrough after breakthrough" that has driven model cost down ~60x a year for equivalent quality. So InferenceX runs daily, automated, on ~15 chip types across the best Chinese and US open models, on over $50M of donated hardware (over $100M once TPUs and Trainium land) from CoreWeave, Crusoe, Nebius, Oracle, Microsoft, Amazon, Google, and OpenAI.
  • The point is the Pareto-optimal frontier: vendors compare their optimal point against a rival's suboptimal one — "if I drove a Porsche versus some race car driver, obviously I'd drive it slower." InferenceX open-sources containers for the optimal configuration at every point on the curve.
  • His master framework: the throughput-vs-interactivity curve — "most things in hardware, infrastructure, model, application layer — everything is downstream of that curve." One-size-fits-all inference ends: batch workloads take the 4x cost saving, latency-sensitive users pay 4x up — Claude Code fast mode and OpenAI's priority queue are examples of price-segmented points.

4. Space compute and the intelligence-per-watt ledger

  • On what share of inference happens in space: sub-1% by 2030, but "in 20 years I think the vast majority of compute will be going in space" — by 2040, "probably more than half of the incremental compute." He'd buy the SpaceX IPO if he could buy stocks — "not investment advice," stated twice.
  • The scale behind that: by 2030 "just OpenAI and Anthropic will have over 100 gigawatts combined," terawatts by 2040, with inference bigger than oil — "many percentage points of GDP."
  • Intelligence per watt has improved roughly 40x (versus 60x on cost). Versus the human brain "we're many orders of magnitude away — thankfully, doesn't really matter": it's much easier to power computers than humans, who come with "sickness, disease, food preferences... sleep."

5. The core thesis: co-design turns 8x into 100x

  • Shaun's framing — three layers (hardware, kernel-level systems, model), with recent gains mostly from hardware — gets flatly rejected: "Shaun, I completely disagree with you." Hopper to Blackwell gave ~30x on DeepSeek at the most optimized deployment, but the model layer gave more: three years ago the frontier was GPT-4; today a small Qwen — "27B parameters total and like 2 billion active" — is way better.
  • The mechanism: DeepSeek V3's expert shapes were optimized for Hopper; V4's for Blackwell and Huawei's chip. Result: "TPUs are objectively an amazing chip... [but] TPUs suck at running DeepSeek." Shapes, network IO, collectives, and attention arithmetic intensity are co-optimized through the stack — "it's hard to say you can disentangle the gains." Same at Google: each Gemini generation is built for that generation's TPU, and pulled onto old hardware "it's really not that great."
  • On the idea that China does this better: no — "the West doesn't tell people what they do." GPT-4o was roughly the same size as DeepSeek V3, slightly smaller, and came out earlier.
  • The money line: layer-by-layer optimization gives 2x·2x·2x = 8x; co-design across all three and "instead of being multiplicative to 8x, it's actually 100x." A company likely associated with Naveen Rao (a Sequoia investment) is the long-horizon version — silicon, software abstraction, and model simultaneously, potentially analog compute with energy-based models: "Probably won't work, but that's exciting... definitely won't work quickly."

6. The CUDA moat isn't CUDA anymore

  • Shaun's observation: model companies no longer fear other chips — "Claude and Codex are actually quite good at doing a lot of that optimization work," and there are only tens of model companies, not the thousands the CUDA-moat thesis assumed. Dylan concedes: "the CUDA moat and software moat is at least partially disentangled... models are just great at coding and all software gets commoditized."
  • What remains is ecosystem gravity: DeepSeek, Kimi, Zhipu (likely), Alibaba, Tencent (likely), Xiaomi co-design their open models for GPUs, so everyone downstream — inference API providers, RL customization shops — concludes "I guess I need to use Nvidia because the ecosystem uses Nvidia." If Google open-sourced really good models, the pull would reverse toward TPUs — that's the point of Gemma.
  • Big labs don't need the open stack at all: OpenAI forked PyTorch long ago. The frontier game is "I'll choose the best hardware and I'll co-design my model and infrastructure software through and through... and I'll have AI help me write all that software."

7. Nvidia vs TPU: both win — the real risk is local minima

  • He refuses the binary: two years out, Google makes 10M+ TPUs (~$100B+ of TPU a year) while Nvidia ships tens of millions of GPUs ($500B+, "not a specific estimate"). "I could with a straight face argue GPUs are way better than TPUs or TPUs are way better than GPUs" — the resolution is co-design: "the way OpenAI's models are headed, it would be a terrible decision for them to use TPUs," and potentially terrible for Anthropic and Google to train on GPUs.
  • The concrete divergences: OpenAI's models are much sparser, Anthropic's sparser-but-more-dense; matrix-multiply unit sizes differ; NVLink switches connect 72 GPUs while Google's switchless ICI connects 8,000 chips at high bandwidth (passing through other chips). Google runs three separate TPU design programs — Broadcom, MediaTek, and a third undisclosed one — with genuinely different architectures.
  • End state: everyone deploys billions to tens of billions of their own ASICs (Google: hundreds of billions a year), but general-purpose compute survives because labs "literally don't know what architecture they're going to be doing in a year" — a specialized chip can race to a local minimum and be stranded by one attention-mechanism breakthrough. Even Google pays xAI $11/hour per GPU, and some non-Gemini Google bets (drug discovery or Waymo — "I won't say which one") run primarily on GPUs.
  • On Cerebras: "really innovative" — SemiAnalysis uses fast mode "almost exclusively" — but the risk is that fast mode matters most on the best models, revenue concentrates there (large numbers of users switched to Fable and Mythos on launch day despite the price), and a 10T+ parameter model at million-token context won't fit on SRAM chips like Cerebras and Groq. On tokens versus dollars: "who cares about volume by tokens? It's about the dollars" — his F-150s-vs-Camrys point.

8. The physical bottlenecks: ancient memory cells, the 1W/mm² wall, diesel engines

  • Memory, from the technology angle rather than supply chain: the NAND cell is ~25 years old, DRAM ~40, and HBM progress has just been more, faster stacks. Coming in the next few years: stacking memory directly on the compute die, which "makes your bandwidth explode."
  • Power density has been pinned at ~1 watt per mm² for two decades — chips hit 1,400W today, 2,000W with Rubin, ~4,000W with Rubin Ultra, all by adding silicon. Breaking past that wall (in development now) means less silicon per chip, at the cost of hard thermal and electrical-interference engineering. Co-packaged optics is a when-not-if: "the debate is like '27, '28, '29, 2030."
  • Energy has unglamorous fixes: convert the millions of diesel truck engines the US can build to gas, back-drive electric motors as generators, and staff the sites by pulling mechanics out of car shops. The deeper Western problem, in Shaun's framing that Dylan seconds: "why would you want to go work in hardware when you can make ads to ads?"

9. The crunch is real — and Anthropic can pay above market

  • Supply is compounding — 20 GW this year, more than 30 GW next, delays included — but demand compounds faster: Mythos 5/Fable 5's TAM is "not just 2x that of Opus," yet world compute didn't double in the seven-to-eight months since Opus 4.5.
  • The economics beneath it: Anthropic was net-income profitable in Q2 excluding stock-based comp, maybe including it by Q3, with Opus 4.8 API token margins "north of 80%." At 75% gross margin, doubling compute cost still leaves 50% — so "every GPU I rent... I can immediately turn around and sell tokens on it at a positive margin... whatever price I want to pay, I can pay." They even bought GPUs above market from SpaceX (still below what Google later paid, having signed earlier).
  • Sonya's pushback — worth keeping: everyone is levered to "we got to build," and Crusoe publicly said a customer asked to halt construction on a buildout; "high leverage high growth makes me very nervous as an investor." Dylan teases ("small amount of equity has huge upside — you're not a debt investor... go to the school of private equity") but concedes the crux: if models' economically valuable work stops expanding faster than compute capacity, "that tide turns."
  • His base case is it doesn't turn: models are improving faster than six months ago via a "pseudo recursive self-improvement loop" — models writing the infrastructure that launches the next model sooner. Capital is the binding constraint: Google raised despite ~$100B of SpaceX stock unlockable in nine months (Larry Page's $1B at a $10B valuation — "one of the greatest investments of all time. Good job, Larry."); Meta announced a raise and the stock tanked. And the take that "infuriates me": "AI has no ROI" and model-progress denial — "bro, the line has been up and to the right in terms of capabilities this entire time... look at the new benchmark, they're skyrocketing."

10. Gigawatts aren't fungible — Jensen wants a multipolar world

  • Pricing already discriminates: Trainium rents at sub-$10B per gigawatt, GPUs historically $12–13B, and the SpaceX–Google deal "was like 25 or something crazy billion dollars per gigawatt — $25 million per megawatt a year." Colocation went from $60/kW/month to $120–160, as high as $200 (weak-credit tenant, good facility), as low as $80 in India. And on the compute layer: "a gigawatt given to Anthropic is objectively worth more revenue than a gigawatt given to OpenAI" — both could sell every gigawatt they have, especially since Codex 5.5.
  • Operational skill compounds it: Google puts ~1.5 GW of hardware into a 1 GW datacenter by sloshing power around from workload knowledge, and cuts utility deals for 2 GW "except three days of the year." Shaun's addition: SpaceX's Starlink networking and Tesla power-management experience "might be missing from the analysis a lot of people are doing" — and SpaceX sells compute that's "running now, buy it," versus Google selling six months forward on paper to finance its own purchase orders.
  • Why neoclouds exist at all — his 2023 "Amazon Cloud Crisis" report: hyperscaler advantages (Nitro NICs, tenant isolation, custom SSDs, custom networks) were built for time-sliced CPU clouds and turned neutral-to-detrimental when AI customers rent whole racks on long-term contracts; Microsoft's datacenter team "fell on their face" when forecasts doubled. Plus incentives: nobody at a hyperscaler gets rich building faster, while Crusoe's "hyperlevered equity owners" do.
  • Jensen's chess, in Dylan's telling: "Jensen absolutely hates a world where all the hyperscalers have all the power" — a world of only OpenAI/Anthropic/Google models, or only hyperscaler compute, "is one in which he's screwed." So he backstops neoclouds, funds neolabs, and champions Chinese labs — because Crusoe and CoreWeave existing in five years makes TPU and Trainium weaker. "You throw a bunch of bait into the water and the best fish will figure out and survive." Early returns: Thinking Machines' Tinker at a few hundred million dollars of ARR under about six months from launch.
Dylan Patel

I think it's really fun inside of SemiAnalysis because we have 90 people, and a big chunk of them are technologists and engineers across the whole supply chain. Another big chunk is people who are formerly at hedge funds. You see these arguments happen organically: people say, “Oh, well, that doesn’t matter,” and then someone says, “Well, but cost,” and one of the engineers says, “No, no, no, but this technology is the coolest.” You see them organically fight it out.

We’re pretty informal, and given the fact that I was a former moderator, you can imagine why they’re enjoying it. You don’t wrestle with a pig because a pig enjoys it, right?

Shaun Maguire

Exactly. We’re here in the SemiAnalysis office with Dylan Patel. I’m Shaun from Sequoia, and my partner is Sonya Huang. It’s pretty insane what you’ve done. Semis 5 years ago were not very sexy in the West. They were sexy in the East, but people here in the West had kind of forgotten about them.

You did not forget about them, though. You went very long, and you created probably the premier research company in the space. You’ve been educating the world on the state of the art, from very technical details to the supply chain to the bigger picture. There are rumors that SemiAnalysis recently passed $100 million in revenue. I don’t know how accurate those are. Whatever the numbers are, you guys are crushing it.

Dylan Patel

It’s as accurate as the information is.

Shaun Maguire

Yeah, cool. You never know. There are also rumors that you might start a venture fund. I hear all the time in the ecosystem about people wanting affiliation with SemiAnalysis. You’ve built this trusted brand, and whatever you do, it’s working. This is clearly just the beginning of the journey for you. Congratulations on all of that.

But how did this happen? My first question is: What’s your background? How did you get to where you are now?

1. Motel Kid Origins

Dylan Patel

When I was a young boy, coming out of the womb—no. Okay, I grew up in a small business. My parents had a motel; we lived in the motel and ran a gas station. I joke that the first neural network I trained was racially and visually profiling people based on when they entered the gas station and which cigarette they would pick.

The cigarettes were all displayed across the top, and I was too short to reach them. Technically, it wasn’t legal to sell cigarettes at that age, but whatever. I had to move the step stool over to the right area.

Shaun Maguire

I started working my first job before it was legal, too. It’s good experience.

Dylan Patel

Well, I didn’t get paid. It’s a family business.

Shaun Maguire

Same.

Dylan Patel

We had our motel, and across the street was our gas station. Sometimes someone would walk in, and if an old white lady with curly hair walked in, I’d move the ladder or step stool over to where the Camels were. If someone of a different age, demographic, profession, race, or whatever walked in, I’d move the step stool over accordingly.

2. Xbox Repair Spark

I joke that this was the first neural network I trained because if I waited for them to tell me what they wanted, I’d have to move it over and then step up, versus just being ready. Menthols versus 100s, slims, and all these things—I joke that was the first neural network I trained.

It all really goes back to when I was 8 years old. My birthday’s in May, and it was April when the Xbox 360 was announced. For my birthday, I didn’t ask for the Xbox—or I didn’t ask for a birthday gift. My parents asked what I wanted, and I asked for it for Christmas.

We celebrated Christmas, but there was no way—at least at the time, I thought there was no way—they would give me the Xbox 360 for Christmas. Anyway, Christmas comes around, and I get it. Fast-forward a couple of months, and my cousin, who lived in Alabama—we also lived in a motel—was going to come over for spring break.

He’s between me and my older brother in age. My older brother was a bit more of a jock, so he didn’t really care too much about the Xbox. He played sometimes, but he didn’t really care. I wanted my cousin to think I was cool, so I bragged many times on the phone: “Yeah, I got an Xbox.”

Then the Xbox broke. There was a hardware defect called the Red Ring of Death. Long story short, I had to open it up and short the temperature sensor, and it fixed it. There were many other tricks I tried first, and none of them worked. That’s how I got into hardware. It was like opening Pandora’s box.

3. Internet Forums to Semis

By the time I was 12, I was on these forums a lot, reading and posting. This was around the time Reddit ate all the other forums, and I became a moderator of Android, Apple, and Google, as well as hardware. I was watching Intel, Nvidia, and AMD and all these other forums. I was building PCs, watching, reading, and posting a lot, and moderating some of them.

I watched smartphones develop from very simple to much faster and eventually become technologically more advanced than PCs in many ways, architecturally. The same was true of all the new GPUs—just tracking and watching that, and reading every comment.

I always had an economic tinge because I grew up in a small business. I was always looking at the economics. There was a time when all the neckbeards on the internet loved AMD GPUs, and I personally had bought an AMD GPU, too, because of price-performance. But when it came down to what was technically better, I’d always be like, “No, no, no, Nvidia is better because they use a smaller chip to get better performance at better power efficiency, and their margins are better.”

I would always talk about how Nvidia’s margins were better than AMD’s in the GPU landscape, and it was very fun.

Shaun Maguire

And you were 12 at the time.

Dylan Patel

I started moderating when I was 12, but this was all through my teenage, tween, and high school years.

Shaun Maguire

Do you have any other weird hobbies, or was it just semis?

Dylan Patel

I played a ton of StarCraft. At one point, I was Grandmaster on the North American ladder in StarCraft II.

Shaun Maguire

Very serious. You’ve gotten obsessively good at multiple things.

Dylan Patel

Obsession is good.

Shaun Maguire

How were your grades?

Dylan Patel

They were decent. I had mostly A’s, but there were classes that I thought were really boring or that I just didn’t enjoy. Spanish—I got not the greatest grades. I speak fluent Spanish, by the way, so it’s really dumb, but it’s just sort of—

Shaun Maguire

Maybe that’s why you didn’t get a good grade.

Dylan Patel

I didn’t learn Spanish until later, to be fair. But, yeah, my grades were fine—fine enough for Asian parents. I was better than most of school, but it wasn’t like try-hard maxing for all A’s.

4. From Quant to Founder

Shaun Maguire

Okay. So you’re very much a student of the internet. This is how you developed this expertise. At what point did you decide to start SemiAnalysis, and what’s been the biggest surprise since starting the company?

Dylan Patel

I went to school and got a few degrees in stuff that wasn’t related to semiconductors. I was a quant for 2 years at a small quant risk firm. Basically, there was a culmination of events that happened.

One was that I got screwed out of a bonus. I had made my company many millions in risk-free revenue because I exploited a risk-related thing in the market—well over $10 million, I think—and then someone else took credit for my work and all this sort of stuff. Eventually, I did get rightsized, but I lost the social contract with the company I was working with.

My grandparents lived in the motel with us, so I was very close with them. My grandmother got dementia and forgot who I was. She fell down some stairs, had a tragic accident, and passed away. All of that happened in early 2020. Additionally, there were some girl things, so there were a few things that happened that made me very sad.

All of those things culminated, and then COVID happened. My brother was like, “Dude, just come stay with me.” He lived in Nashville, so I went and stayed with him in Nashville. We thought lockdowns would last a few weeks: I could stay with him while they happened, then go back home. Famous last words.

Lockdowns lasted much longer. Living with my brother for a few months was like, okay, I didn’t know what I was doing. I was now at my brother’s home, and everything was his rules. He and his fiancée at the time, now wife, were there, so I basically had to tiptoe around. But I didn’t care about my job, so I was posting even more than normal.

I’d always been posting a lot on the internet, and I’d always been trading stocks a lot. I made a lot of money shorting COVID and going long on COVID, and all this stuff.

Semiconductor shortages happened around then, too. Anyway, I was very much obsessed with posting and things like that. Eventually, around that time, I got into an argument with someone on the internet, and they doxed me. They publicly revealed my identity for my anonymous account. At the time, I was like, “Oh no, I’m scared.” I stopped posting for 3 weeks, and I was like, “What am I doing? Why do I care?”

So then I started posting under my real name. I had blogs and stuff as well. I made a real blog, SemiAnalysis, and on my 24th birthday, I posted 2 blogs. From there, it took off. It wasn’t a newsletter, but I got so much traction because now, instead of posting under an anonymous name, it was a real name, and I put a lot more effort into those 2 posts than I usually did. Instead of just posting on the internet, I put real effort into the blog.

5. Homeless Research Roadtrip

You can actually go back and read those if you want. They’re not that great, but they were good for the time. They were the best stuff you could find on the internet about semis. I just kept posting, posting, posting, and I started getting a lot of consulting business.

6. Cerebras Speed and Limits

In 2020, I was again crashing out. I didn’t know what I wanted to do, so I packed everything up. I took my truck, bought a tent that fit on the back of the truck, bought an air mattress, and drove around all these national parks all around America. For 2, 3, or 4 days of the week, I’d stay in a random motel where I negotiated the price down to around $30 a night for a room, and I would work on other stuff. On the weekends, I’d read books and often textbooks while in some random national park or hiking, and listen to audiobooks about semiconductors, AI, and all the things I cared about.

I got way more educated over those 6 months of going to every national park. The whole time, I was alone and posting blogs. Everyone was like, “Dylan, what the fuck are you doing?”

Shaun Maguire

Pre-Starlink or the very early days of Starlink?

Dylan Patel

Pre-Starlink. Yeah, it was very much like, “What are you doing?” I traveled around LatAm again for about a year, initially with my friend and then with my ex. Then, from the end of 2021 through 2024, I was completely—I’m still completely homeless since mid-2020—but I was traveling around to every conference in the world.

I go to 40-plus conferences a year, no matter where in the supply chain they are. I’m like, “Oh, that looks interesting. I guess I’ll go to that.” I went to 1 conference and thought, “Wow, this is amazing. You get to talk to the experts, and they’ll talk to you because you’re so excited.” In the case of semiconductors, everyone’s a boomer, so it’s great to talk to them. They don’t see young people who are excited about it, so they’re really happy to tell you stuff, and you just have to ask.

Shaun Maguire

Was there a part of the supply chain, or one of these conferences, that particularly changed your view of the semiconductor world, or that you felt then—or feel now—is particularly underrated?

Dylan Patel

I think trade shows and conferences range really widely. Obviously, some of the ones I have the most fun at include NeurIPS.

Shaun Maguire

Why is that?

Dylan Patel

Because it’s 20,000 AI researchers, and they’re generally in my age range. It’s a lot of fun, but they’re also leading AI researchers, so you learn a lot. There are also a lot of parties.

Then it ranges all the way to a random chemical conference in Japan, where it’s 300 Japanese dudes. It’s 20 guys from ASML, 20 guys from TSMC, 20 guys from Intel, and those are the only people who speak English. Everyone else speaks only Japanese, and you’re like, “Huh, I guess they’re still pretty interesting and fun.”

I think one skill set I have is that I’m able to bond with anyone, regardless of their background or who they are. I’m able to talk to them and find something interesting to talk about. Oftentimes, it’s the tech stuff, but it’s not always. I think the most interesting conferences are oftentimes the really big ones, because that’s where the biggest stuff is happening.

But I think the niches that are really exciting are SPIE. There’s IEEE, which is the Institute of Electrical and Electronics Engineers, and there’s SPIE, which is another ecosystem. SPIE conferences are super, super deep in detail. Every single one I went to—especially SPIE Advanced Lithography and SPIE Photomask—I didn’t even understand 90% of what I heard the first time.

Then I read, read, read, and built some context, of course, and the next time I went, I understood half of what I heard. The third time I went, I understood 75% of what I heard. Even now, I go and think, “I still don’t understand everything that’s going on.”

Whereas, when you go to NeurIPS a couple of times, you can understand, “Okay, what’s neurosymbolic reasoning? What’s this? What’s that?” You can kind of get a map of what everything is pretty quickly. But some parts of the supply chain are so arcane, so deep, and so technical that it takes many visits for you to even understand what’s happening in everything.

You go to a conference for a few reasons. You understand the research being published, but what you really care about is understanding how that research intersects with technology. Also, how does that research differ from what’s there today? None of these research papers tell you what’s happening today, but then you just ask people, build contacts, and learn.

Then you learn about the supply chain: this company supplies that company, even though it’s not publicly stated anywhere. You learn that this chemical costs about this much, a tool uses about this much, and you hear the horror stories of how a chemical shortage totally threw off this part of the supply chain. Then it turns out there are only 3 companies in the world that make that chemical.

My favorite one is that I learned from a Japanese guy at that specific Japanese conference I went to, where almost no one spoke English. In very broken English, he told me about how his father worked in this industry in the 1980s, that the only factory in the world that built this chemical burned down, and that caused memory prices to double or triple.

Shaun Maguire

Not too different from today.

Dylan Patel

Not at all.

7. InferenceX and Benchmarking

Shaun Maguire

Crazy. Is inference going to be the biggest market on Earth—the biggest market beyond oil? Agree or disagree?

Dylan Patel

Obviously, the use of tokens is going to be the biggest market, and the value that’s created from tokens is going to be the biggest market. But I think tokenomics—the use of tokens and adoption of AI—is the most important thing that’s happening. Inference, whether it’s open models or closed models, will be one of the biggest markets in the world, much bigger than oil. I think inference of AI will be many percentage points of GDP.

Shaun Maguire

Yeah, right. What you’ve done with InferenceX, I think, is industry standard. Maybe say a word on why you started it, what it does, and what people misunderstand about performance benchmarking for inference.

Dylan Patel

Yeah. To zoom back, SemiAnalysis does a lot of stuff. A lot of it is research for institutional clients and our subscription products, but a lot of it is also, “Hey, this would just be cool to figure out. Let’s figure out how to figure it out and just post it publicly.” That gets more and more scale.

We’ve done this with a lot of GPU benchmarking and testing, training performance, and inference performance. Ultimately, we saw that inference benchmarking was point-in-time. You test it, take some time, release it, and it’s slow, arcane, and outdated because models change all the time. Every week, I feel like there’s a new model, whether it’s a Chinese model or, you know, Mythos 5, Fable [?] dropped today.

On the software layer, PyTorch, vLLM, SGLang, new drivers—something new drops all the time. In fact, the update cycle for most of these libraries is twice a week. You basically have the software updating all the time and, therefore, performance changing. New inference optimizations are coming out and getting updated.

I feel like it’s a relentless breakthrough after breakthrough after breakthrough that keeps driving efficiency and cost down, which is why we’ve seen model costs drop for equivalent quality by 60× a year. It’s incredible. But to stay on top of that, you can’t have point-in-time benchmarking. You need benchmarks to be living and breathing—constantly running on the latest hardware and the latest models.

We embarked on a project and got a lot of buy-in from the ecosystem. This was only possible because we had enough aura with some of the ecosystem to get CoreWeave, Crusoe, Nebius, Oracle, Microsoft, Amazon, Google, and OpenAI to contribute compute to us. We were able to work with SGLang and vLLM, and now Red Hat and Inferact, which are the private companies leading those open-source efforts, to collaborate with us.

We're able to get NVIDIA, AMD, Google, and Amazon involved now because we're adding TPUs and Trainium to collaborate. Now we've got all these people collaborating. We've got over $50 million of hardware donated to us. Once we launch TPUs and Trainium, it should actually be over $100 million of hardware.

Maybe about 15 different chip types are all running these benchmarks every single day on all the latest models, right? The best model from Moonshot, the best model from Alibaba—there are about 5 different Chinese models—the best open-source models from the best Chinese labs. We run benchmarks on their models every day, and also on the best U.S. open-source models: GPT-OSS, Nemotron, et cetera. So we're running these benchmarks every day in an automated fashion, and they run on servers dedicated to us for inference benchmarking. We sweep across so many different configurations and optimization types.

What it creates—and all the results are public, and all the configurations are public—is the Pareto-optimal curve, because a lot of times when people are comparing inference performance, they're taking a suboptimal curve or point for someone else and comparing it to their optimal one. It's like, well, yeah, I can make—I can stick—if I drove a Porsche versus some race car driver, obviously I'd drive it slower. The same thing applies to inference benchmarking.

What we did is create open-source containers for the optimal points across every point on the interactivity—that is, how fast it's responding to me—versus batch size—that is, how many users I'm simultaneously serving—curve. Now anyone who wants the optimal point can just go to InferenceX, download it, and run that as the optimal point. They can check every day if they want, or they can even auto-download the most optimal point for that model, and their inference performance will be near peak.

Shaun Maguire

Is that curve the most important curve, in your opinion?

Dylan Patel

The throughput-interactivity curve is the most important one. Yeah, I think most things in hardware infrastructure, the model application layer—everything is downstream of that curve, right? Is it something that needs to be super, super fast, super-low latency, and I don't really care about the cost, so I make batch size very low and use techniques like speculative decoding or multi-token prediction heavily? There are so many possible techniques there.

Or is it something where I'm actually batch-processing a ton of documents and I don't really care about all these things? I don't use these techniques that are actually worse for cost efficiency but help with speed for an individual user, because I just want to pack a bunch of users together. I don't care if the document takes all night to process, right?

Right now, the way we treat AI infrastructure, it's like one-size-fits-all. But over time, we're going to get to the point where there are workloads where you have batch processing, or you need an instant response, and the whole curve is going to matter for users. We see this with Anthropic, right? Claude Code's fast mode costs way more than regular mode, and the same is true with OpenAI's priority processing.

Shaun Maguire

Sorry, dumb question. How does cost factor into the chart? So, in an imaginary example, I have a batch size of 100—

Dylan Patel

Okay?

Shaun Maguire

—and I can do 10 tokens per second per user. In total, I'm doing 1,000 tokens per second off that one piece of compute. That's one side of the curve: super slow, 10 tokens per second. The other side is that I have 500 tokens per second, but I only have 1 user. So maybe 250 tokens per second, 1 user. Then there are points in the middle that are Pareto-optimal, right? The average person actually wants 50 or 100 tokens a second, and maybe the number of users I can batch together.

So the curve is, okay, 1,000 tokens total per second or 250 tokens total per second, depending on how many users I batch, and there's a curve in the middle. If you had to guess, and you choose the time frame—10 years or 15 years—what percentage of inference compute do you think will happen in space?

Dylan Patel

Ultimately, some workloads will actually want the 4× cost decrease because the same unit of hardware can do 1,000 versus 250, and some users will pay 4× more because they don't care about the price; they care about time because the person using the tokens is expensive, or the feedback loop I have here is expensive.

It can be 0%, 50%—Shaun, 99%? This is a tough one. You choose the time frame, like 10—whatever time frame—and you're—

So I think the nonconsensus, or at least the view against the SpaceX thing—I love SpaceX, by the way, and I would totally buy the IPO if I could buy stocks.

Shaun Maguire

Not investment advice.

Dylan Patel

Not investment advice. Thank you. Thank you. Not investment advice from either. I don't think space data centers will really matter in the next 3 to 5 years. With that said, I think in 20 years the vast majority of compute will be going into space.

The real factor there is the cost: the time frame, the cost of building power on terrestrial land, and how much power you're going to be able to have on terrestrial land. Obviously, my view of where inference is going—how many gigawatts or terawatts are devoted to inference—is a crazy curve for me personally.

Shaun Maguire

What's your forecast? How many gigawatts?

Dylan Patel

By 2030, just OpenAI and Anthropic will have over 100 gigawatts combined. Then you'll add Meta and Google, and so on and so forth. It's a humongous amount of compute that will be dedicated to inference. By around 2040, it'll be terawatts—the curve of productivity that we're going to get.

Inference deployments are going to be huge. If you look at 2040, I think probably more than half of the incremental compute will be going into space. But if you look at 2030, I think it's sub-1%.

Shaun Maguire

Do you think intelligence per watt has been increasing? It seems like there's still a giant gap between where we are in intelligence per watt versus human biology. Do you think we're going to close that gap? And if so, where is that gain going to come from?

Dylan Patel

Yeah, I think it often depends on what you're doing, too, right? A TI-84 is way more intelligent per watt in terms of doing math than us, and it's 30 years old, right? Obviously, this is a dumb, dumb—

Shaun Maguire

General intelligence.

Dylan Patel

Yeah. But general-intelligence-wise, one of the things InferenceX does is measure the power and cost of all this hardware. So we offer not just throughput versus interactivity; we offer cost versus interactivity and power versus interactivity.

As far as whether intelligence per watt has been increasing, I mentioned that it's been a 60× cost decrease for the same benchmark level. We've also seen the same thing with intelligence per watt. It hasn't been exactly 60×; it's been closer to 40×. Some of the efficiencies are non-power-related, but there's been a humongous improvement in intelligence per watt on an annual basis—at least so far this year, last year, the year before, and the year before that. I expect that to continue.

As far as where we are compared with the human brain, we're many orders of magnitude away. Thankfully, it doesn't really matter. We can devote a lot of power to computers. It's much easier to power computers than human brains. We have sickness, disease, food preferences, and sleep.

Shaun Maguire

Yeah, exactly.

Let me ask one more question on the general theme. In terms of intelligence per watt or intelligence per dollar—any of these metrics—I think there are 3 levels of input. You can get hardware improvements, where the hardware is more efficient. You can get low-level systems optimizations, like kernel-level improvements, matrix multiplication libraries, things like that. Or you can get high-level, model-level algorithmic improvements.

It seems to me that in the last 3 years, most of the gains have come from the hardware level, with some from the model level. Do you agree with that? Do you think that's what it will look like in the future? Do you think there's a bunch of juice to squeeze in, say, the kernel level, like—

Dylan Patel

Yeah, Shaun, I completely disagree with you, by the way.

Shaun Maguire

Great. That's why I'm asking this question.

Dylan Patel

Okay, so I think one way is to look at it as these 3 different layers. In that sense, from Hopper to Blackwell, which is all we've had over the last 3 years, there's roughly a 30× improvement on DeepSeek on the most optimized deployment, which you can see on InferenceX. There's about a 30× improvement.

But over the last 3 years, we've had way more improvement in intelligence per watt, a lot of that coming from the model layer, right? If you look back 3 years, it's GPT-4. Now it's maybe Qwen, one of the smaller Qwen models, that's about 27B parameters total and 2B active, and is way better.

You've got this huge improvement on the model layer, and you've got this pretty sizable improvement on hardware, but it's that co-design layer. I think that's what's important, right? If you look at the architecture of any of these models, DeepSeek is the most famous one, at least among the public models that people have seen.

Shaun Maguire

Yeah, DeepSeek got huge efficiency gains from co-optimization, or kernel-level optimization of memory.

Dylan Patel

Yes, I think it’s kernels, of course, but it’s actually that you build the hardware architecture for the chip. If you look at the shapes of all the experts in DeepSeek V3, they were all optimized for Hopper. If you look at V4, they’re optimized for Blackwell and Huawei’s chip.

What’s interesting is that despite the fact that TPUs are objectively an amazing chip—they run all of DeepMind, and they do all the training for Anthropic as well, at least on the pretraining side—TPUs are terrible at running DeepSeek. But they’re really, really great at running other kinds of models that don’t run well on NVIDIA.

There’s some level of such deep optimization that’s been done, whether it’s the shapes, network I/O patterns, how you do the collectives, or things around the arithmetic intensity of the attention mechanism. All these different things are co-optimized between the model, the hardware, and the infrastructure software in between. It’s hard to say you can disentangle the gains.

Shaun Maguire

Do you think that my understanding is that China has done this a lot better than the West over the last few years? DeepSeek was one of the first models to really do this.

Dylan Patel

I don’t necessarily think so. I think it’s more that the West doesn’t tell people what they do, right? OpenAI didn’t tell people how sparse GPT-4o was, what the shape size was, or any of these things. But GPT-4o is roughly the same size—slightly smaller than DeepSeek V3—and 4o came out a little bit earlier, if I recall correctly.

Shaun Maguire

So is your view that all 3 of these things have been happening simultaneously at roughly the same rate, and the biggest gains come when you just co-optimize?

Dylan Patel

I would say there have been more gains on the model layer than on the co-optimization layer, the software infrastructure layer, and the hardware layer. But there have been innovations on every layer, and really, the biggest gain—and the beauty of the best labs—is when they co-optimize all 3.

When Anthropic has used many different kinds of hardware, they don’t really do much inference on TPUs. They mostly train on TPUs, and they run a lot of inference on Trainium and GPUs. GPUs are more of a jack of all trades, but they’ve optimized their hardware, they’ve optimized their model, and they’ve optimized everything so they can do that.

OpenAI’s prior models were optimized for Hopper; now they’re more optimized for Blackwell. You step forward through time with these labs, and it’s the same with Google. They’ve optimized Gemini 2 for the TPU v6e, Gemini 3 was, and then the next Gemini that’s coming out is really optimized for TPU v7.

A lot of these things are being co-optimized, and when you pull that model and run it on the old hardware, it’s really not that great. I think a lot of this co-optimization is the most important thing. It’s called software-hardware co-design, and that’s what’s really exciting about what I think my day-to-day is like: You get to look at one layer, but there are all these innovations happening here, and there are all these innovations happening on every layer.

The real breakthrough innovation is when you leapfrog a few layers, co-optimize and co-design them, and all of a sudden you’ve taken what could have been a 2x gain here, a 2x gain here, and a 2x gain here. Instead of being multiplicative to 8x, it’s actually 100x because you’ve optimized across all 3 layers.

That’s what’s really exciting about what you see at the labs, and what you see at a company like NVIDIA, which isn’t co-optimizing on the model layer per se, but is co-optimizing a little bit from the model layer all the way downstream to silicon. Or you look at a company like TSMC: They’re co-optimizing not just fabrication, but everything from the components, consumables, and tools all the way upstream to what the designs of their customers’ chips are telling them.

This is co-optimization across many layers of the abstraction stack.

Shaun Maguire

There will always be bottlenecks somewhere in that optimization, though, that are lagging behind and then need to get pulled forward. If you had to predict at any level of the stack—it can be literally anywhere—what are some of the bottlenecks you’re tracking most acutely over the next year?

Not necessarily in the supply chain, not in scale, but in terms of the actual... It can be in the supply chain too, but is it memory improvements? Is it just scaling?

Dylan Patel

Memory is an easy one that everyone’s talked about, but I’m not going to talk about it from a supply-chain angle. I’m talking about it from a technology angle. Memory capacity and bandwidth have been improving very slowly.

The NAND cell was invented about 25 years ago. The DRAM cell was invented about 40 years ago, and there’s been no major breakthrough in how a NAND cell or a DRAM cell works. Obviously, NAND is a very simple gate. There is stuff that could come down the pipeline that could be hugely innovative.

But even over the last 5 years, all we’ve really done is add more stacks to HBM and make it faster. Actually, there are new innovations coming in the next few years where, instead of stacking the HBM separately from the chip, you stack the memory directly on the chip, and that makes your bandwidth explode. There are interesting companies in that space and interesting proofs of concept that companies are trying to do there.

I think memory bandwidth is one of the biggest bottlenecks. Another one is that, for the history of silicon—basically, for the last 2 decades at least—how many watts a chip is can be easily predicted just by looking at it. For a data-center or desktop chip, it peaks out at 1 watt per square millimeter. So if a chip is 100 square millimeters, generally the power consumption is around 100 watts or a little bit less.

If you look at the newest NVIDIA silicon and the newest TPU silicon, they’re still in that range of 1 watt per square millimeter. Chips are now getting to 1,400 watts. The next generation is 2,000 watts for NVIDIA, with Rubin and such. As you move forward to Rubin Ultra, it’s going to be 4,000 watts or something like that.

But really, that’s just increasing the amount of silicon. What’s exciting is that we’re now finally doing things—and it’s in development right now—where you can actually pump the amount of power into the silicon to be way more than 1 watt per square millimeter. All of a sudden, that means you need less silicon. Obviously, it’s running at higher power and is less efficient in some cases, but you reduce the amount of silicon and you’re able to—

Shaun Maguire

Like, over thermal issues?

Dylan Patel

Thermal issues. There are electrical interference issues. There are all sorts of different issues that crop up, and that’s why it’s a hard engineering problem. That’s why we’ve stuck at about 1 watt. But what’s exciting is that the world is trying to change these things.

In a different part of the supply chain, people will talk about how energy is hard and how we have energy bottlenecks. But there are actually very simple solutions one could think of. Take the millions of diesel engines for trucks that the U.S. has the capacity to make. You can very trivially convert them to use natural gas on the assembly line and then connect them to an electric motor, back-driving it so the electric motor generates electricity rather than causing the rotation of the wheel, for example—doing it in the opposite direction.

Now you’ve generated electricity by pumping gas into something that the U.S. can make millions of. That sounds like a pain to service, because now you have to have hundreds of these on a data-center site. But you can just pull people out of car mechanic shops and have them run around and repair truck engines.

Actually, it’s pretty trivial to do. I don’t want to say it’s trivial; I couldn’t do it.

Shaun Maguire

I think you’re making a really good point, which is that because the West wasn’t really thinking about semiconductors, or even hardware more broadly, over the last 20 or 30 years, we didn’t have much innovation. We’d have the best minds thinking about how to improve these—

Dylan Patel

Why would you want to go work in hardware when you can make ads to ads?

Shaun Maguire

Yeah, exactly. Okay, I’m dying to ask: NVIDIA versus TPU. What are your thoughts?

Dylan Patel

I think everyone wants to pick one or the other for this, but it’s really a function of the fact that, if you look 2 years from now, Google is going to make 10+ million TPUs through their supply chain, and NVIDIA is going to make many more—tens of millions of GPUs.

Both are going to be $100+ billion. Well, Google is going to be $100+ billion of TPUs created a year, and NVIDIA will be $500+ billion, or whatever. I’m not making a specific estimate.

Shaun Maguire

This is not a revenue forecast.

This is just a thought experiment.

Dylan Patel

Yeah. Or research.

Shaun Maguire

You've been media trained.

Dylan Patel

Absolutely. You know, getting ready for the SpaceX IPO.

Are you guys big in SpaceX?

Shaun Maguire

Okay, so that makes sense. We're very lucky to be very large investors.

Dylan Patel

Awesome. So I would say the case of Google TPUs versus NVIDIA GPUs is that they both have points that are really in their favor, right? NVIDIA will be like, “Oh, well, we have switches and we're general purpose.” And TPUs will be like, “Well, we're more optimized, actually more energy efficient, and our network is more optimized for certain types of network architectures.”

8. Sparse vs Dense Models

So you have these counterpoints that both would really get into. I could, with a straight face, argue with you that GPUs are way better than TPUs or that TPUs are way better than GPUs, but it comes down to hardware-software co-design. Actually, given the way OpenAI's models are headed, it would potentially be a terrible decision for them to use TPUs. And given the way Anthropic's and Google's models are headed, it could potentially be a terrible decision for them to train with GPUs.

Shaun Maguire

What's the fundamental difference there?

Dylan Patel

There are various things, right? The size of the matrix-multiply unit is different, as a very simple example. Therefore, the shape of the matrix multiply you do, the attention mechanism you use, the way that attention mechanism is structured, and the way the experts are structured are all different.

Shaun Maguire

So you think OpenAI and Anthropic are converging on very different model architectures?

Dylan Patel

I think they have quite different model architectures. In fact, OpenAI's models are much more sparse, and that has benefits. Anthropic's models are still sparse, but more dense in general, and that has different benefits.

9. Interconnect Shapes Architecture

There are many other things, right? The network topology is another one. NVIDIA's chips are all connected to NVLink switches. Google has no switch. What they've done is enable their interconnect to connect 8,000 chips at super-high bandwidth, but you have to pass through other chips to get there because there's no switch.

There are trade-offs there—positives and negatives—and that influences the model architecture. It's not necessarily that you should claim one is better than the other because, at the end of the day, how do you say that this is better than that when you can't measure them in isolation? It also extends up to the model layer, right?

10. CUDA Moat Is Shifting

Shaun Maguire

I remember, for a long time, thinking that the programmability of NVIDIA and CUDA itself was such a big moat. It seems to me that narrative has changed, at least in my mind, over the last 3–6 months. Model companies no longer care whether they have to write custom kernels for another chip; so be it. They'll work with 4 or 5 chips if they have to.

Claude and Codex are actually quite good at doing a lot of that optimization work. It also isn't as though there are 10,000 model companies that each need programmability. There are on the order of tens of model companies, maybe. So it seems to me that the fundamental premise of tens of thousands of big customers that need CUDA compatibility—that thesis seems to be changing over the last 3–6 months.

Dylan Patel

Certainly, the CUDA moat and the software moat are at least partially disentangled because models are just great at coding, and all software gets commoditized in that case.

11. Ecosystems and Co-Design

I do think there is some level of open source, and what people call the CUDA moat is not actually anything to do with CUDA. It's the fact that DeepSeek, Kimi, Zhipu, Alibaba, Tencent, and Xiaomi—all these companies—have models that are co-designed for GPUs. Therefore, if I want to run them on TPUs, in some cases they don't run really well on TPUs.

Google just has to create its own open-source model ecosystem or open-source models themselves. They have the Gemma models. So you end up with, well, that's not really CUDA as a moat; it's that the downstream product is more optimized for NVIDIA. In these cases, these companies are just open-sourcing them. Nemotron is just open-sourcing it.

The users of those models—for example, the open inference API providers, the RL companies that are trying to take open models and customize them for companies' business use cases—all these different companies are downstream of the fact that, “Okay, I guess I need to use NVIDIA because the ecosystem uses NVIDIA,” even though I don't particularly care about writing CUDA kernels because the models are great at that.

It's the shape of the model: this expert, the d_model is this, and the hidden dimension is this. Therefore, it's better to run on NVIDIA GPUs than it is on TPUs, and vice versa. If Google were to open-source really good models, people would take those models and say, “Oh, wow, these don't run that well on NVIDIA GPUs. I should actually just rent or buy TPUs and do it on there.”

For small teams, you're going to want to use all the open-source software, like vLLM, SGLang, PyTorch, and all that stuff. But the big labs don't necessarily need to use all of that. OpenAI forked PyTorch long ago, and Anthropic and all these other companies don't necessarily rely heavily on the open-source implementations of these things. They forked things or built them on their own already, so they don't need to rely on the open source.

Therefore, now it's more like, “I'll choose the best hardware, and I'll co-design my model and infrastructure software through and through for the hardware that is the best and most cost-efficient.”

Shaun Maguire

And I'll have AI help me write all that software.

What do you think of Cerebras?

Dylan Patel

I think Cerebras is a really innovative company. In some parts of the market, they're really, really good. They have very fast inference, and I think that's a big market. We use fast mode almost exclusively at SemiAnalysis.

Shaun Maguire

By the way, I love how disciplined you've been about accounting for the dollars spent and the ROI on each task. I don't know if that was one exhibit you did or if you do it consistently, but it's awesome analysis.

Dylan Patel

Yeah, we do it pretty diligently, so thank you. That was the Dark GDP article that we wrote. We also track everyone's token spend by day, and if someone's spend has spiked, I'm like, “What did you do?” It's like, “Okay, thank you for telling me. That seems worth it.” Then I continue with my day.

I think fast mode is obviously worth a lot for high-end tasks, right? I could see so many different use cases where super-fast tokens are worth it. I can also see the flip side, where there are a lot of use cases where super-fast tokens aren't needed, and therefore the market won't pay for them. They'll use GPUs and TPUs instead.

I think the big risk for Cerebras is that I mostly think the best models are the ones that you want to use fast mode on. With small models, you might not necessarily use fast mode. I could see that being wrong with financial markets, maybe—like Jane Street high-frequency trading or something like that, or medium-frequency trading.

Ultimately, running really large models at really long context lengths is very difficult on SRAM-based chips like Cerebras and Groq. So now it all of a sudden becomes, what happens if the models get too big? If OpenAI's model is not on the order of hundreds of billions of parameters or low trillions of parameters, but is actually 10+ trillion parameters, I don't think that will fit on Cerebras.

And if that doesn't fit with a long context length—if you have a 1-million-token context length—that makes it really difficult to justify. So far, we've seen the bulk of revenue and usage at the labs be on their best model, even when the model price has gone up.

There's some data that shows that even though Fable just released today, they've had incredible numbers of people switch to Fable and Mythos, sort of that next-tier model, even though it's way more expensive.

Shaun Maguire

And that's volume by dollars, totally? But was that volume by tokens?

Dylan Patel

Well, I guess who cares about volume by tokens? It's about the dollars.

Shaun Maguire

Fair enough.

Dylan Patel

Right. If I don't care that there are 200,000 Mini Coopers or Toyota Camrys sold if Ford F-150s have 5× the ASP and sell only half as many, then the most lucrative market is pickup trucks in America. I'm mostly being facetious, but—

Shaun Maguire

I do think this is one of the things you've done so well, and it differentiates you from almost everyone else: You care so much about the economics in addition to the technology. Very few people bridge those 2 things well.

Dylan Patel

I think it's really fun inside SemiAnalysis because we have 90 people, and a big chunk of them are technologists and engineers across the whole supply chain. Then another big chunk are people who were formerly at hedge funds.

You see these arguments. People are like, “Oh, well, that doesn't matter.” Then someone's like, “Well, but cost.” And then the engineers are like, “No, no, but this technology is the coolest.”

You see this organically fight it out. We’re pretty informal, and given the fact that I was a forum moderator, you can imagine how much we enjoy it.

Shaun Maguire

You don’t wrestle with a pig because a pig enjoys it.

Dylan Patel

Exactly.

12. ROI Debates and Hot Takes

Shaun Maguire

Just on this topic, before going to the next question: are there trigger topics in semis for you? Like, if someone says, “Memory is the bottleneck,” which is such a meme, does that trigger you?

Dylan Patel

I mean, it’s true, but the one that really gets me is when people say, “AI has no ROI.”

Shaun Maguire

It infuriates me, right? There’s the “What’s the ROI?” argument, or people denying model progress. There are these people who say, “Models aren’t getting better. They’re not reasoning. They can’t think. They’re going to dead-end and plateau.” It’s like, bro, the line has been up and to the right in terms of capabilities this entire time. They’re like, “Look, this benchmark didn’t improve because it was at 90 percent. Look at the new benchmark—you saturated it, and now the scores are skyrocketing.” Right?

Dylan Patel

I think that’s more so the issue and challenge. Semis are really complex, and I don’t fault people for lacking an understanding of them. I learn stuff every day about the semiconductor supply chain from people, and I’ve been studying it for arguably 18 years, since I started moderating the forums when I was 12. I’ve been studying it for that long, and it’s all I care about, but there are so many layers of the abstraction stack.

I learned about a new chemical that does $100 million in sales yesterday, and I was like, “Whoa, I didn’t know this one existed or what process it was used in.” You learn about things all the time. It’s like, okay, $100 million in sales in a couple-hundred-billion-dollar industry is whatever, but—

Shaun Maguire

But it’s essential.

Dylan Patel

It’s essential, and actually every chip requires it. It’s like, wow, I guess there are 1,000 process steps. You like semiconductors? Name every process step. Come on.

What I think is the funniest is when people have all the facts in front of them and then get the conclusion completely wrong.

Shaun Maguire

That happens in our job all the time, too.

Dylan Patel

Yeah.

Shaun Maguire

Yeah. I mean, I can’t—I get it. I think my attitude is not to be mad that you do that; it’s to do it as fast as possible.

13. Ten Year Tech Bets

I think the industry is so important. AI is the most important thing in the world right now, and there are so many near-term bottlenecks. We talk a lot about the near term, but are there longer-term things that you’re really excited about, say, on a 10-year time frame? We talked about orbital data centers, but what about silicon photonics? Do you think it’s underrated or overrated on a 10-year time frame? Are there other things that excite you on that time frame?

Dylan Patel

On space, I think space is super crazy and awesome on a 10-year time frame. I’m super excited about space data centers, mining asteroids, and all these things. I’m super excited about the vision of SpaceX. Again, not investment advice before you hop in.

On the semiconductor side, tremendous market movements and tremendous things can happen just when things happen 1 year later or sooner. That’s all technology that, in terms of co-packaged optics, everyone knows is going to happen by the end of the decade. The debate is whether it happens in 2027, 2028, 2029, or 2030, but at some point along there, it’s going to happen.

I think the more interesting thing is that there are companies like Naveen Rao’s company. Did you guys invest in Naveen Rao’s company?

Shaun Maguire

We did.

Dylan Patel

Okay, yeah. He’s trying to innovate on the silicon layer, the software abstraction layer, and the model layer simultaneously, and he fully understands that it’s not a “we’re going to do this in a few years” kind of thing.

Shaun Maguire

It’s not a 2-year time frame.

Dylan Patel

Yeah, it’s not a few-year time frame. It’s a long-term bet. Stuff like that is exciting. It’s like, okay, we’re going to bring analog compute, energy-based models, and all this crazy stuff together at once. That’s exciting. It probably won’t work, but that’s exciting, and I really look forward to—

Shaun Maguire

It definitely won’t work quickly.

Dylan Patel

Yeah, definitely won’t work quickly is what I should say. I believe in Naveen. I met him very early, and I think he’s one of the first people I met in the industry, funnily enough, in 2020 or 2021. Actually, 2020. That says something about him.

In my experience, he’s always trying to help the younger generation and identify talent.

Shaun Maguire

I baited him on the internet. I baited him on the internet.

Dylan Patel

He’s always trying to help the younger generation. He’s trying to identify talent.

Shaun Maguire

He was also so ahead of his time with Mosaic. I remember getting pitched.

Dylan Patel

No, it was 2019. I was still anonymous then, actually. I baited him on the internet, and he started replying. Then I took it to DMs, then took it to a call, and that was the first person who was really important that I talked to in the entire semiconductor industry. Funny.

Shaun Maguire

But yeah, sorry to interrupt. That’s funny. What do you think is the end state of the ecosystem? Do you think every lab and every hyperscaler just has its own chips? Trainium seems like it’s working now, right? Do we end up with every lab and every hyperscaler having its own chips, at least for inference, and then maybe going to NVIDIA or whoever for training? What do you think is the end state?

Dylan Patel

I think everyone will try and then stop trying. Ultimately, supply chains matter, and what technology you can bring in matters. More and more, as the industry gets bigger, supply-chain diversification happens.

Right now, everyone’s chip more or less looks the same. It’s a big logic-compute die in the center, with some HBM on the right and left and on the top and bottom. The top side is networking, and the bottom side is PCIe and other I/O. That’s the exact same structure for Trainium, TPU, and NVIDIA chips. Most of the startups—not Groq and Furiosa, which are doing something weird, but that’s cool—are doing the same thing.

As you step forward, we’re going to get more bifurcation of hardware architecture and model architecture, and therefore people are going to co-optimize them. Some of them will end up in local minima. If this is like gradient descent, people are trying to get to the most optimized solution. Some people will race to a local minimum, and then the question is: how do you leap, how do you scoot back over to the absolute minimum?

To some extent, a general-purpose NVIDIA chip will always be more general-purpose than anyone else’s chip, at least on a parallel AI-compute basis, because NVIDIA has so many customers who care about different things and who will always give them feedback in the design. The minimum will always be better for them, but is that minimum a local minimum? Is the TPU, Trainium, Groq, Cerebras, or whoever’s design optimized awesomely for here, but in the end state, you actually have to go over here, and so they’re the wrong choice?

Shaun Maguire

Maybe they make a great chip. They’re great for a little bit of time, but then they end up being wrong. That’s the real question.

Dylan Patel

I think there will be a big market for general-purpose AI compute. You talk to people at labs, and they don’t even know what architecture they’re going to be doing in a year. They literally don’t know what architecture they’re going to be doing in a year. They have bets, many research bets, and that’s the exciting thing, but they don’t know where it’s going.

Generally, they know what hardware they have, and they’re trying to co-optimize. Ultimately, if a new breakthrough happens in model architecture, it’s like, just replace the attention mechanism with something else. Who knows? All of a sudden, if something happens, the best hardware will change.

Therefore, are people going to make 5-year investments in hardware solely on an ASIC that is more specialized? They’re going to have some bucket of more general-purpose compute. You see this with Google paying $11 an hour per GPU to xAI for H100 GPUs, right? That’s insane. Obviously, compute is limited and so on and so forth, but it’s a very high amount. At the same time, despite the fact that Google has TPUs, there are some questions there: Why do they do that?

Google actually has 3 different design programs for TPUs. They’re making a TPU with Broadcom. That’s a different architecture from the TPU with MediaTek, which is a different architecture from the third one. I won’t disclose the details of that one, which is being developed by Google Research. They’re making different architectures. It’s not just, “Oh, they’re making TPUs with a couple of vendors using the same architecture.” They’re making different architectures, and the third one is very different from the first 2.

I think people recognize that local minima can happen. Therefore, I think everyone will have their own ASIC program. Everyone will deploy billions of dollars of their own ASICs—tens of billions of dollars, and in Google’s case, hundreds of billions of dollars a year on their own ASICs. Ultimately, they’re also going to have workloads that don’t use TPUs. Some of Google’s bets that are not Gemini or DeepMind actually primarily use GPUs.

They don't use TPUs. Some of them also primarily use TPUs, right? It's a broad thing, but maybe for drug discovery or for Waymo, you might not want to use TPUs. I won't say which one it is, but there are different architecture bets and different paths for AI.

AI for science may have different algorithmic patterns than general-intelligence AGI models. I think we'll see diversity continue to proliferate.

Sonya Huang

Yeah, and because the market has gotten so big, niches will be carved out. That makes it possible for companies to have their niche and actually make money, even if the majority of the pie goes to NVIDIA, TPU, and Trainium.

14. Compute Crunch and NeoClouds

Okay, love that. Can we talk about the data center buildout? First, it seems like, by all accounts, if you look at the charts—dollars per compute hour—we're in the middle of a crazy compute crunch. It seems like it's both a demand- and supply-side crunch, right? Demand for long-running agents is skyrocketing, while supply— all these data center buildouts—are delayed. Do you think we're in a compute crunch for the foreseeable future, or do you think it alleviates at some point?

Dylan Patel

Yes. Every quarter, we're deploying vastly more compute than the prior quarter, and there are more data centers built than in the prior quarter. This year, there's going to be 20 gigawatts, even accounting for the delays, and next year there's going to be more than 30 gigawatts, accounting for the delays. Of course, delays happen on everything, right? Anything involving hardware can have a delay. That's just the reality of life.

Are we going to have a compute crunch for the rest of our lives? It depends on what happens with models. But the TAM for Mythos 5 and Fable 5 is not just 2x that of Opus, right? The model is so much better, and it can do so many more tasks, that the TAM for it is way larger than that.

And yet compute in the world did not double in the last 6 months, right? From Opus—or maybe 7 or 8 months since Opus 4.5 launched—to now, 4.6, 4.7, and 4.8 were improvements, but Fable and Mythos were a huge step-function improvement. The world's compute did not double, quadruple, or whatever in that same time frame. But the demand for useful tasks that can be done by AI—the number of useful tasks and the value of them that can be done by AI—has.

So now the question is, what happens? Obviously, Anthropic in Q2 is net-income profitable, excluding stock-based compensation. I think by Q3 they may even be profitable including stock-based compensation. That's how profitable they're getting. Their margins on an Opus token—at least an Opus 4.8 token—are north of 80% at the API price.

They've got a lot of deals where their total corporate gross margin gets clawed down a little bit because of how they do Bedrock deals, Vertex deals, and things like that. But ultimately, their per-token margin is so high. If you have the capability to pay, you can ultimately buy every GPU above the market rate. They also bought GPUs at above-market rates from SpaceX, which is below the rate of Google, but that's because they signed earlier.

That's something that other companies—maybe a venture-backed company or a company that doesn't really have positive margins—can't necessarily do, right? What is the cost-benefit ratio? Every GPU I rent because I'm out of compute capacity, I can immediately turn around and sell tokens on it. Every TPU or Trainium I rent, I can immediately sell tokens on it at a positive margin.

If I'm running a 75% gross margin and I double the cost of the compute, it's fine. I'm still running a 50% gross margin, and spinning up more compute nodes isn't necessarily a human-requiring task for them if they're renting them. Ultimately, it's like, well, my NOI still goes up, right? I'm going to rent GPUs at whatever price. At some level, whatever price I want to pay, I can pay.

Sonya Huang

I have almost the reverse question. At some point, does this compute buildout go bump in the night? Earlier today, I think there was a tweet that Crusoe publicly said one of its customers had asked to halt construction on one of its data center buildouts. It seems like everybody in the ecosystem is so leveraged right now to, “We have to build. We have to go build. We have to build.”

High leverage, high growth makes me very, very nervous as an investor.

Dylan Patel

Wait, hold on. High leverage, high growth means a small amount of equity has huge upside. You're not a debt investor.

Sonya Huang

You're a credit investor—you're an equity investor, right?

Dylan Patel

Let's go.

Look, you've got to go to the school of private equity. Levered buyouts only.

Sonya Huang

I actually come from the school of private equity.

Dylan Patel

Oh, awesome.

Shaun Maguire

She forgot the school. She's been a VC for too long.

Sonya Huang

Yeah. No, I just do revenue multiples. But are you seeing any signs of that? Are you worried about that?

Dylan Patel

I see what you mean, right? That sort of goes back to the model point. Obviously, if the models expand the total economic value of the work—the “Dark GDP” report that we did, which you mentioned earlier—if the work that these models can do does not expand faster than the compute capacity, then that tide turns, right?

Over the last 6 months, that tide has been very much levered in this direction: The models can do more work, or are expanding the TAM of work they can do, faster than compute is increasing. So prices go up.

It's very possible that all of a sudden model progress stops. You talk to anyone at Anthropic or OpenAI—maybe they're drinking the Kool-Aid—but you talk to basically all of them and they're like, “No, no, no, no. Model progress still goes up.”

Ultimately, current methods could stall somewhere. I'm not sure where that would be. It seems like we have a line of sight to model improvement—rapid model improvement. In fact, models are improving faster than they were 6 months ago or a year ago because there's—I wouldn't call it recursive self-improvement, but basically, the models are helping the engineers write all the infrastructure and launch the next model sooner and sooner and sooner.

You've got this pseudo-recursive-self-improvement loop going, and so the models are getting better and better and better faster.

Ultimately, capital is a big problem, which is why Google raised capital. They've got an ungodly amount of cash, right? They own about 5% of the company.

Sonya Huang

I think a little more, but yeah. Maybe. I think at one point they had around 10%.

Dylan Patel

Larry Page invested $1 billion at a $10 billion valuation and got 10% of the company. It got diluted, and all of that. But that was one of the greatest investments of all time. Good job, Larry—the guy.

So they know they have something like $100 billion in the bank that they can sell in 9 months or whatever, once the lockup expires. They have all the gross profit they do, and yet they still modeled that and were like, “We need to raise capital.” So they did an offering, and it's like, that's insane. That tells you how much they think they need to spend.

Capital is a real problem. Meta did announce that they're going to do a raise. The stock tanked. People don't like it, but all these companies are going to raise capital, whether it be debt or equity. At some point, money spigots will have to slow down.

But right now, every GPU that Amazon adds is making higher revenue, or every TPU or Trainium that anyone adds is making gross profit. I do a little bit of a tee-up on this to turn it into a question for you, but—

Sonya Huang

As we talk about this, the thing going through my head is almost an alternative hypothesis for the Crusoe example. I'll use an analogy in oil. In oil, Saudi Arabia has a way lower cost per barrel to produce oil than a lot of other countries. There's also the purity of the oil. Saudi Arabia generally has very low contaminants in its oil, which makes refining easier, and all of this.

The question for me is, when you look at every gigawatt that's being put in the ground—call it the 20 gigawatts coming online today—how much homogeneity do you see in those gigawatts? I don't know what metric you think is right, but are Google's gigawatts 2x more valuable than, say, most neoclouds because they have optical switches, they've been doing it for a long time, and they know how to do power smoothing?

I think this could be the alternative hypothesis: Some of the people who are good at building data centers should just do it to the max because there's so much demand and they're so much better at it. But maybe we're starting to see the early signs of the people who aren't as good at it getting hit a little. I don't know the reality here. I'm just curious how you think about this.

Dylan Patel

So far, there are metrics for this, right? Trainium sells at a sub-$10 billion-per-gigawatt rental rate to Anthropic and OpenAI. GPUs, at least before the craziness of the last 6 months, usually went around $12 billion to $13 billion per gigawatt.

So the rental rate—and this is from a neocloud versus Amazon, even—and now, when Amazon sells GPUs, they'd also be $13 billion or so—

Sonya Huang

My understanding of that is also that Amazon subsidized those numbers a little bit. I actually think the numbers were even lower—I think the disparity was even more.

Dylan Patel

It's less than 10.

It's less than 10, but there's some weirdness, basically.

Shaun Maguire

And look, my understanding, obviously, is that Anthropic played a big role in making Trainium useful in terms of writing all the libraries, et cetera.

Dylan Patel

Everything I hear is that Trainium's really freaking good hardware, and it's getting way better. Obviously, Anthropic is now using it a lot, so hopefully we would see that price go up. Per the deal they did, there was actually a floor mechanism in it: if it didn't do well, it would be cheaper, to the point where it was cancellable, and if it did really well, the price would be higher. But effectively, less than 10 is where Trainium shakes out, whereas the xAI GPU deal, again, was $25 billion or something crazy per gigawatt, or $25 million per megawatt as a yearly rental rate with Google. I was like, that's a crazy divergence.

Now, obviously, if Amazon was selling Trainium today, it'd probably be more expensive than 10 because of the compute shortages. But you do see this already in the sense that, with data centers, oftentimes a rental price of a data center—if you're doing colocation, not compute in there, but just power—gets priced generally in dollars per kilowatt per month.

They used to be $60 per kilowatt per month, and now you see things transacting anywhere from $120 to $160. But with different-quality data centers, I've seen data centers go as high as $200 when the customer doesn't have such a great credit rating and the data center is a pretty good one. I've seen stuff go as low as $100 still, or in India as low as $80, because the grid isn't reliable, the internet connection isn't great, and it's a pretty mid-level data center—but at least it's a data center.

In the case of data center construction, usually the pitfalls are that they just fail. There are a lot of people who claim they're going to build a data center—there are like 4 guys, and they're like, “Yeah, we bought some turbines. We put the money down for them. We're going to build a data center.” Then they get delayed, delayed, delayed, and fail.

You have to probability-weight and time-weight the time lag, and account for the teams that suck versus the teams that don't. Our data center model does that. We track every data center and try to do this for every single one based on the equipment they're using and all these things.

One of the things you mentioned about Google is that, in a gigawatt data center, they'll actually put 1.5 gigawatts of hardware. Because they understand the workload so well, they're able to slosh the power around. Instead of constantly having a gigawatt of compute—which typically runs at 60% or 70% utilization in terms of power consumption, not utilization of the hardware, because someone's always renting it—they're now running it so that 60% to 70% means it's at a gigawatt and they're using the full gigawatt.

You see people doing deals, including Google with utilities, where they're like, “Oh, well, I know this grid can sustainably take a gigawatt, but except for 3 days of the year, you can actually do 2 gigawatts. So give me 2 gigawatts and then just tell me to turn off.” And so they'll do that.

These sorts of tricks require supreme management of workload, backup power, and all these things, plus generators on-site, to figure out how to actually keep 2 gigawatts sustainably. When people do this, they're able to charge more. Whether it be, “I'm actually selling 2 gigawatts despite only having 1 gigawatt because those 3 days I'm able to deal with via battery, gas, et cetera,” or, “I figured out how to build power on-site. Now I have a gigawatt where no one else does, and so I'm able to do it quickly.”

It's not necessarily transacting for a higher price; it's that I'm selling more gigawatts. Sometimes there are levers where selling more gigawatts means that each gigawatt is selling at a different price. I think, on the data center and energy layer, it's more about just having it versus not, and then whether that is delayed or not. It's more binary.

But on the compute side, I do think there's a lot more interesting work there, right? A gigawatt given to Anthropic is objectively worth more revenue than a gigawatt given to OpenAI. It seems that both of them could sell every gigawatt that they have right now, given rate-limit problems, token max limits, and all these sorts of things at OpenAI and Anthropic. Especially since Codex 5.5 came out, it's much better.

Shaun Maguire

My guess—my suspicion—is that they probably make better use of the hardware than most people. I think people underestimate how much networking experience they have from Starlink in particular, and also how much power-management experience they have via Tesla.

Dylan Patel

Yeah, people like Brett Mayo are incredible.

Shaun Maguire

Pretty good.

Dylan Patel

Yeah.

Shaun Maguire

And so I think that, for me, that's probably the thing that might be missing from a lot of the analysis people are doing. I don't actually know the answer, but I think that might be what's missing.

Dylan Patel

I think it's also the fact that when CoreWeave builds a gigawatt, even though its GPU compute is objectively better than Amazon's, Google's, or Microsoft's in terms of performance—we've tested the performance and reliability—the problem is that CoreWeave sells it 6 months before they have it up. They need to turn around and take the paper they signed to get debt, with that credit backing, so they can actually pay for the purchase order they've already issued—for the order they've already issued.

Whereas xAI was like, “No, no, no. This is running now. Buy it.” It's a big discrepancy when you have a balance sheet to do that versus not, and that also helps your revenue per megawatt be much higher.

Shaun Maguire

Why does the neocloud opportunity even exist? If you had asked me 5 years ago, I would have said the hyperscalers were going to own this. You mentioned just now that CoreWeave has better performance than the hyperscalers. Why does this opportunity exist, maybe at the macro level and then at the execution level?

Dylan Patel

Yeah. So in 2023, I wrote a report that had Amazon really hating me. It was called “Amazon Cloud Crisis.” I talked about how Amazon was the best cloud because they had their Nitro NICs, which offered tenant isolation. The hypervisor ran on the NIC, and then you could sell all the cores.

They had custom SSDs that they made. They'd buy the raw NAND and build their own SSDs, so they had lower costs, and they had their custom Graviton CPUs that drove down the cost per core. They had all these things that enabled them to sell more cores, have better security, and have good networking and storage, but this was all for the traditional CPU and traditional cloud world.

In the AI cloud, a lot of this stuff hurt performance. These Nitro NICs were bad for performance and still have worse performance, although they've caught up a lot because they've had a couple of iterations to improve them. They're still worse for performance.

A lot of the security stuff doesn't matter because it's not like I'm time-slicing users or splitting a socket among many users. No one rents a single GPU in an 8-GPU server. No one rents a single GPU in a 72-GPU rack; they rent the whole rack, and in fact, they rent many of the racks.

There's no, “I rent for 6 hours and give it back.” Everyone has these long-term contracts. The mechanics of the GPU rental market meant that a lot of the expertise of the hyperscalers fell away, and a lot of the expertise that they did have was actually detrimental.

For Google and Amazon, they had custom networks that were better for traditional CPUs and for the stuff they were doing, but actually worked worse for AI. In other cases, Microsoft would save money by building its own data centers, but its data center teams actually weren't that great. When building was predictable, it was fine; when it came time to double the forecast for the year, they fell on their face and had to go get a bunch of neocloud capacity.

I think performance—I think time to market's another one. These massive organizations—no one's getting rich from building this data center faster. But you look at Crusoe, for example: Chase and all the other people on the team. I was going to name some people on the team, but I'd rather not. These people are getting rich if they deliver this compute faster. They're hyper-levered equity owners.

Shaun Maguire

Hey, look, they're also all coming from Bitcoin. And you know, you're not supposed to say that.

Dylan Patel

Their main data center guy came from Microsoft.

Shaun Maguire

I don't know. I'm just teasing. But you learn a lot when you're in a high-fluctuation market.

What do you think Jensen was playing for in this chess game?

Dylan Patel

Jensen absolutely hates a world where all the hyperscalers have all the power.

There's a reason he's blowing money on random AI labs that I don't even know if it makes sense to, but he's blowing money and pumping them up and going to everyone around the world and saying, “You should invest in this company,” because he wants to create a multipolar world. That's why he loves Chinese labs: He wants to create a multipolar world. A world where OpenAI, Anthropic, and Google models are the only models is one in which he's screwed.

Shaun Maguire

Yep.

Dylan Patel

Right. A world in which the hyperscalers are the only ones building compute is one he's screwed in.

Shaun Maguire

Yeah.

Dylan Patel

And so, of course, he needs to point the allocation gun at neoclouds, help backstop their clusters, and do anything and everything because, while today a GPU sold to Crusoe, a GPU sold to CoreWeave, and a GPU sold to Google and Amazon are all the same price for him, 5 years from now, Crusoe and CoreWeave existing means Google TPU will be weaker, and it means Amazon Trainium will be weaker. More inference being done with non-closed-model labs is better for the firm.

So I think the neocloud ecosystem is these people that are Wild West. These neoclouds as well—a lot of them have investments from NVIDIA. It's the Wild West. Some will fail, many will fail, but some will emerge as really great teams, whether it be, oddly, Crusoe, who's a bunch of crypto guys who then started building data centers and doing flared-gas stuff, or CoreWeave, who initially was a bunch of New York hedge fund guys and crypto guys, but then they built. There were a lot of people who didn't bubble up like them, started around the same time, and just failed. Right. And so I think—

Shaun Maguire

I gotta say, both those teams are phenomenal. They deserve a lot of credit, and that's your point.

Dylan Patel

Yeah, my point is you throw a bunch of bait into the water, and the best fish will figure it out and survive, right? In the same way with the neoclouds, he hopes the neolabs will do the same as well. We'll see if any of the neolabs really bubble up, but Thinking Machines has a few hundred million dollars of ARR, right? That's pretty impressive, even though in the media it's, “Oh, they've lost all this talent.” It's like, well, but Tinker is doing a few hundred million dollars of ARR. That's pretty impressive for a product that's less than 6 months old or whatever.

We hope the same happens to other neolabs, and so he wants a multipolar world.

Shaun Maguire

Truly, congratulations on the success. Thank you. Just the last thing I'll say is, I've seen a little bit of this. I think the public can probably tell from listening to you how hard you work, but it's clear you've just been working your ass off for more than a decade, and it led to the last few years of being in the right place at the right time. It's unbelievable what you've accomplished, and I know it's just the beginning.

Dylan Patel

Thank you so much.

Shaun Maguire

Thank you for doing this.

Awesome.