[BidClub_]
The Next Big Thing · · 67 min

Dylan Patel on the infrastructure powering the AI revolution | The Next Big Thing

Dylan PatelChristopher Gannatti

YouTube
TL;DR
  • The memory call is the episode's spine: SemiAnalysis flipped from "memory is the biggest loser from AI" (early '23) to biggest winner when o1 led Patel to predict KV cache would explode, and Patel now says "it's a shortage that's going to last years" — capacity growing only 20-30%/yr for the next three years while demand doubles. Pricing is up ~4x with "another 2x, 3x" coming; the inelastic buyers get forced out (mid/low-end Chinese smartphone shipments already down 40%), iPhone and MacBook prices "have to go up" next year — "a few hundred bucks," not $100 — until "AI gets its fill." Gross margins head to 85-90% before halving back to the 70s or lower.
  • CPUs: real inflection, but a mini-cycle, not a supercycle. Reinforcement-learning environments and agentic tool calls genuinely inflected demand — SemiAnalysis called it in November institutional research and ARM, Intel, and AMD have "ripped" since — but the sell-side "who doesn't really understand technology at all is just making up stuff." The math: at ~$5k per CPU versus "$50 something thousand" per Blackwell, $300-500B of Blackwell implies only $30-50B of CPU sales; today's frenzy is a one-time catch-up on ~10M GPUs or other AI chips shipped over three years with no CPUs attached. "This market was underpriced and it's more fairly priced now."
  • CPO is pushed out — favor copper instead. "People are a little bit too excited on CPO": not 2027, "tail end of 28, but 29 is the real ramp" for scale-up co-packaged optics. Reuben is all copper and Fineman on the GPU is still copper, so likely Amphenol and non-CPO optics are favored medium term while CPO manufacturing and yields catch up.
  • Power is the binding constraint on a buildout going 20 GW this year → 30 GW → 50 GW, and the answer is increasingly on-site: within a couple years half of incremental datacenter power will be generated behind the meter — from GE Vernova combined-cycle turbines down to truck engines converted to gas and serviced by car mechanics ("it's a pain in the ass, but it will work"). Solar-plus-battery undercuts gas in ~2 years; Nvidia dropping 800V from likely Reuben Ultra Kyber pushes that conversion supply chain out.
  • The ROI debate, answered with receipts: Anthropic was free-cash-flow positive and profitable in April and May with June trending the same, revenue past $50B ARR at 70%+ gross margins — "Anthropic is printing." SemiAnalysis's own AI spend went from under $100k annualized in November to ~$11M today (peak week annualizing $14M) for a 90-person firm once Claude Code inflected; companies clamping down on AI spend "are going to get left in the dust."
  • Cost optimization means the newest model, not the cheapest: a task that took Claude 4.6 Opus 100,000 tokens over several turns takes 4.8 Opus 25,000 in one shot. Token efficiency, not benchmark edge, is why "Anthropic has been beating OpenAI" — OpenAI's models win frontier edge cases but burn 3-4x the tokens with a slower human feedback loop. For AI baked into fixed processes, the opposite logic: freeze quality and ride the ~60x/yr cost decline (DeepSeek was 600x cheaper than likely GPT-4 roughly two years later, versus 3,600x implied by two 60x steps).
  • Blackwell measured 30x faster than Hopper somewhere on the continuum on DeepSeek V3 in SemiAnalysis's open-source InferenceX benchmarks — above even Jensen's ridiculed 25x launch claim ("Jensen, I was wrong. You were sandbagging"), and Jensen spent five minutes at GTC on Patel's charts and "Inference King" belt as proof he doesn't sandbag numbers.
  • The reusable framework for every "next shortage" (MLCCs, PCB drill bits, copper foil): flow-through × elasticity × market structure — how many cents of each AI dollar reach the product, whether pricing is spot-commodity (memory) or partner-stable (TSMC takes "5, 10%"), and whether the market is an oligopoly or hundreds of names traded across Taiwan, Japan, and Korea.
Digest · the substance, structured for research

1. From tween shitposter to 90-person research firm

  • Patel's own origin story: "the origin of SemiAnalysis really comes from shitposting" — posting about smartphone SoCs and display specs "before I ever even had a smartphone," moderating Android/Intel/Nvidia forums by age 12, then two years as a quant: "yes, you make money, but it's not as amazing as it seems." He quit in 2020 for a WordPress blog under his real name, shaped by growing up living in his parents' motel in rural Georgia — "I kind of just know business."
  • The first post set the template: when Huawei lost access to TSMC, the U.S. market thought Qualcomm would win; Patel called MediaTek the biggest winner because "geopolitically China would rather buy from a Taiwanese firm than from a US firm." Technology, supply chain, finance, and geopolitics melded — fed by 40 conferences a year up and down the stack, some 300-person and Japanese-only: "I didn't live anywhere."
  • The pivot to an institutional firm came via hire #3, Myin — a hire with a hedge-fund background who answered a hiring note buried in the paid section of the early-'23 post arguing memory was the biggest loser from AI. Models and data services followed, and "the ball started tumbling down the hill": headcount 2→7 (2023-24), 7→20, 20→60, now 90, with 30 added this year.
  • The moat, as Patel tells it, is talent density no one else has: ex-ASML, Applied Materials, and Lam Research engineers upstream; ex-Intel, TSMC, Nvidia, OpenAI, Tesla FSD, and someone likely from Cohere downstream; "someone at my company who built a power plant in Kazakhstan." The other half: ex-hedge-fund, or "random people from the internet who are super passionate."

2. Jensen on stage: "Dylan said I was sandbagging — but I wasn't"

  • InferenceX is SemiAnalysis's open-source benchmarking suite running every single night — because any night a CUDA, PyTorch, driver, or inference-engine version can drop — across eight GPU types plus Google TPUs and Amazon Trainium, on $50M+ of donated hardware from OpenAI, Microsoft, Amazon, Google, CoreWeave, Nebius, Crusoe, and Oracle.
  • The call it produced: Jensen claimed 25x Blackwell-over-Hopper at launch and "a lot of people were like, no, no, no, it's like 3x"; SemiAnalysis's simulator said 15-20x. InferenceX then measured 30x faster than Hopper somewhere on the continuum on DeepSeek V3. Patel emailed: "Jensen, I was wrong. You were sandbagging."
  • At GTC — 20,000 people in the stadium — Jensen put Patel's charts and the WWE-style "Inference King" belt on stage for five minutes as proof he doesn't sandbag numbers: "He talked about us longer than anyone else in the entire presentation." The only thing that got comparable airtime was OpenClaw.

3. Anthropic is printing — and SemiAnalysis's own AI bill went over 100x

  • Patel's first answer to the ROI skeptics: Anthropic was free-cash-flow positive and profitable in April and in May, with June trending the same; recurring revenue "soared past $50 billion ARR" at gross margins above 70%. "Anthropic is printing." OpenAI's revenue is inflecting too as Codex adoption grows.
  • The demand side, from his own books — he calls it ARS, "annual reoccurring spend": under $100k in November (a $200 ChatGPT seat per employee), $4M by end of January once Claude Code hit its inflection with Claude Opus 4.5 and 4.6, ~$11M today with a peak week annualizing at $14M. AI is already more than a third of employee spend, "probably half by the end of the year depending on how Methos and other models" land.
  • For good developers on ~$300k salaries, AI spend "is starting to approach one to one" — and at SemiAnalysis, "a lot of our biggest spenders are people who don't know how to code," who just "iterate, iterate, iterate."
  • The corporate fork in the road: companies blew through full-year AI budgets by Q2. Some cut legacy SaaS, some cut employees instead of AI, some clamp down on AI — "but those companies are going to get left in the dust in terms of productivity gains."

4. Cost optimization means the newest model, not the cheapest

  • Patel splits AI workloads in two. Process-integrated AI (check every inbound document for XYZ): hit a quality bar, freeze it, then ride the cost curve down — models get ~60x cheaper per year at fixed quality. "People freaked out about DeepSeek because it was 600 times cheaper than likely GPT-4." About 2 years after that comparison, the 60x annual curve would imply 3,600x; it actually ended up 600x.
  • AI-as-assistant inverts the logic: "cost optimization is often times taking the newest model." A task that took Claude 4.6 Opus 100,000 tokens and several back-and-forths takes 4.8 Opus 25,000 tokens in one shot — fewer tokens, less human time, lower true cost.
  • Token efficiency is "the main reason why Anthropic has been beating OpenAI": OpenAI's models can crack edge cases in leading science, math, and code that Anthropic's cannot, "but they take 3x as long and 4x as many tokens," and the human-in-the-loop feedback cycle degrades. So SemiAnalysis remains "a majority Anthropic shop," reserving overnight tasks for OpenAI Codex.
  • After both the 4.6→4.7 and 4.7→4.8 Opus releases, his cost fell for about a week — then soared past prior highs as people adjusted: "okay, the work I was doing is done, let me do more."

5. Memory: a shortage measured in years, paid for by your next iPhone

  • The intellectual arc matters: SemiAnalysis called memory the biggest loser in early '23 (AI servers carried far less memory content than regular servers), then flipped in December 2024 when o1 launched — reasoning led Patel to predict KV-cache usage would explode, and while weights are read identically at 1,000 or 100,000 tokens of context, KV cache memory reads scale with context while compute barely moves. Conclusion then: "memory was going to be the biggest winner."
  • The January 2026 note doubled down when people asked if +50% was the top: "no, no, no — I don't think you guys get it." Capacity grows 20-30% a year for the next three years; demand is doubling. "This is not a short-term shortage. It's a shortage that's going to last years."
  • The clearing mechanism is brutal: prices soar until inelastic buyers drop out. Chinese mid/low-end smartphone makers like Xiaomi report shipments down 40%; the high end is untouched so far, which is why "next year iPhone prices have to go up, next year MacBook prices have to go up" — and since $100 won't move that market, "they're going to have to go up a few hundred bucks" until "AI gets its fill."
  • He is explicit this stays cyclical: pricing already ~4x with 2-3x more coming, memory margins headed toward 85-90% gross — "memory necessarily doesn't deserve a margin of 85%" — before halving back to the 70s or lower. Cycles survive; the trough just keeps rising.

6. The framework: flow-through × elasticity × market structure

  • The episode's reusable tool, laid out before every sector answer: for each dollar of AI spend, how many cents flow to this product — "it might be one cent on this product, but it might be 5 cents on this product"? Is the end market up 50%, doubling, or quadrupling? Then, who can reprice: TSMC is non-elastic — "pretty fair with their customers… we'll take up price 5, 10%" — while memory lets spot and contract markets clear, so it 4x's. ASML barely oscillates at all.
  • Apply it to the tail: MLCCs, PCB drill bits, copper foil — "you'll go online and see this is the next shortage, this is the next shortage." The local bumps are real but "very small and many," and the names trade in Taiwan, Japan, Korea — "not just easily accessible to investors."

7. CPUs: agents and RL made the forgotten chip more in demand

  • SemiAnalysis flagged it in November institutional research: OpenAI and Anthropic were striking deals to rent essentially all the CPUs in Amazon's, Google's, and Microsoft's fleets. The host's observation: three years of AI without hearing "CPU," and now it's everywhere.
  • The mechanism: pre-training barely touched CPUs, but reinforcement learning checks every generation against an environment — unit tests, compilers, sandboxed websites and shopping flows — and agentic inference makes constant tool calls into the regular world (searches, databases, Python interpreters). Both are CPU-hungry in ways chat never was.
  • A third demand leg: deployed output. GitHub commits are up "multiple X" versus last year — "a lot of the code is sloppy, but a lot of code is being deployed" — and web scrapers and business-process automations land on standard, cost-effective CPU cores.
  • The winners map: Intel and AMD both raised prices; ARM entered and its stock "has gone gangbusters"; Amazon extracts "incredible margins" renting Graviton rather than selling it; Nvidia's standalone Vera carries $20B of CPU revenue guidance.

8. But CPUs are right-sizing, not the next supercycle

  • Patel's pushback on his own call's momentum: "the sell side, who doesn't really understand technology at all, is just making up stuff" — ratios now imply more CPU than AI compute, which is "false." A full Blackwell runs "$50 something thousand per chip" versus ~$5,000 per CPU; even a 1:1 ratio on $300-500B of Blackwell yields only $30-50B of CPU sales.
  • What's actually happening is a backlog catch-up: roughly 10 million GPUs or other AI chips shipped over three years "that don't have any CPU attached really." Once that fleet is matched, demand drops to the incremental attach rate — "we're in sort of a mini cycle of CPU." His verdict: "this market was underpriced and it's more fairly priced now."
  • The design space splits by workload, per the core-count law — double a core's size and per-core performance rises only ~50%. Nvidia's Vera bets on fewer than 100 fast cores for workloads where AI compute stalls waiting on a CPU response; AMD's 256 cores and Graviton offer the more-core option when workloads are highly batched and no one is waiting on any single core. "For some workloads you do want Vera and for some workloads you want the Graviton or the AMD CPU."

9. CPO slips to 2029 — favor copper in the meantime

  • Networking content is growing faster than any other category — from sub-10% to above 10% of AI-chip-associated spend, and 20-30% once CPO arrives; telecom optics—likely including Ciena—have been ripping.
  • But on co-packaged optics itself: "people are a little bit too excited." "It's not coming in 27 in my view — really the tail end of 28, but 29 is the real ramp for scale-up co-packaged optics." Manufacturing volumes, yields, and chip designs simply aren't there; switch-level CPO comes earlier than GPU-level.
  • The Monday institutional note: medium-term bullish copper and non-CPO optics, "kind of bearish on CPO" — Reuben is all copper, Fineman on the GPU is still copper, and Reuben is likely only just starting to ship. So backplane makers—likely including Amphenol—"are actually going to do way better over the next few years than previously expected."
  • Why copper keeps winning locally: "at the end of the day, integrating optics is so much more expensive than sending something electrically" — until distance forces repeaters or optics. Close your eyes for five years and optics is "way bigger"; some of that is priced in, some isn't.

10. Power: from truck engines to space, the constraint that yields to money

  • The scale: 20 GW of datacenters deployed this year, 30 GW next year, 50 GW the year after — gated by energy first, politics second, construction third. Of the three power legs, transmission is "the hardest to be bullish on" (utility monopolies, cost amortization rules); generation and conversion are where the action is.
  • Behind-the-meter is the release valve: SemiAnalysis predicts half of incremental new datacenter power will be generated on-site within a couple of years. The supply chain runs from GE Vernova/Mitsubishi/Siemens combined-cycle gas down to train, boat, and truck engines converted into power generation; truck engines can be converted to gas, back-driving electric motors, buffered by batteries, serviced by "a bunch of people from car mechanic shops" — 10 GW+ of datacenters planned on such tech. "It's a pain in the ass, but it will work" — the continuum runs from "going full dirty" to "fully into space," where solar panels need no battery at all.
  • The next crossover: "in about 2 years, solar plus battery will be cheaper than gas" thanks to China's manufacturing scale — with the caveat of reliability nines (enough battery for one night is cheap; three rainy days is not).
  • Conversion is its own supply chain — IGBTs, silicon carbide, GaN MOSFETs, the 12V→54V→800V DC transition, solid-state transformers, supercapacitors — and it just took a hit: likely Reuben Ultra Kyber no longer has 800 volt, pushing that chain out. Fittingly, SemiAnalysis's biggest research vertical isn't semiconductors — it's the "DEI team" (datacenters, energy, industrials, "it's a pun internally"), tracking every datacenter and power plant, because "Google's interested in what Meta is able to deploy."

Verification Notes

  • Raw captions contain an internal contradiction: “memory isn't a shortage” is followed by “It's a shortage that's going to last years”; the digest follows the latter.
Christopher Gannatti

I'm joined by my colleague, Klay Hyman, and Dylan Patel, our new partner and the founder of SemiAnalysis, the research group. It's an exciting day, because we're going to go through the current state of the AI infrastructure landscape with Dylan.

You've probably heard Dylan on many different podcasts. I follow his content all the time, as well as the newsletter on the SemiAnalysis site. Most recently, one of the pieces I saw was about data centers in space. If anyone is interested in that topic, they have a great piece with a lot of detail on it.

Dylan, I would love to hear a bit about where the idea for SemiAnalysis came from. I know that in the Substack community these days, people are talking a lot about the firm's revenue and the success it has achieved, but I feel like a lot of the time you see the current level of success and forget about the journey, the origins, and how much work probably went into all of it.

1. SemiAnalysis Grows From Scratch

Dylan Patel

Yeah. I think the origin of SemiAnalysis really comes from shitposting—posting online in a less serious manner. When I go back to my first posts about semiconductors, they were from when I was a tween. I was just posting on the internet about chips, smartphones, smartphone displays, and smartphone SoCs, before I ever even had a smartphone.

It was the same with gaming hardware: PC hardware and console hardware. I was posting about this stuff all the time on the internet, on various forums. By the time I was 12, I was moderating and creating a lot of forums related to Android, Apple, Google, Intel, NVIDIA, and AMD—basically, all the hardware topics of the world. There were various forums on Reddit related to these things.

That's sort of where the origin of it all comes from. I've always been a poster. I've always posted my opinion, replied, fought, and taken comments. One of the things the people on my team say now—we're a 90-person organization, so I actually have people in marketing—is, "Dylan, stop replying to random bozos on the internet. You're making us look bad."

I just have that burning desire to respond to anyone on the internet when they try to criticize me. Maybe it's a bad thing, but throughout my teenage years I was moderating these forums. I started investing as soon as I started making money in my later teenage years. I was a quant for 2 years, and then I started my firm, but the whole time I was posting, posting, posting.

I had anonymous blogs and anonymous posts. Then, in 2020, I got fed up with my job. The disillusionment of being a quant is that it's not as amazing as it seems. Yes, you make money, but it's not as amazing as it seems.

I quit my job and started my company. I wasn't exactly sure how it was going to go, but I was posting on a WordPress website that I made. I was posting under my real name and writing about a mix of technology, business, finance, and supply chain—the things I was most interested in.

I grew up in a small business. I grew up in a motel in rural Georgia. My parents owned the motel, and we lived in it. We later had gas stations, too, so I sort of lived in business and grew up in it. I always loved business.

Supply chain was always interesting from the perspective of investing and from the perspective of how things are made. I think that's always just been a knack of mine: How are things made? The technology aspect, of course, is super exciting, and the finance aspect is super exciting.

Taking the combination of those things, my first posts were about things like China banning—or the U.S. banning—Huawei from access to TSMC. My first post was actually about how MediaTek was the biggest winner. Huawei had the number-one market share in China for smartphone chips and smartphones in general, and obviously that was going to tank because Huawei no longer had access to TSMC.

The U.S. market thought Qualcomm was going to win. But MediaTek, the Taiwanese firm, was actually going to win a lot more share because, geopolitically, China would rather buy from a Taiwanese firm than from a U.S. firm, given that we had just banned Huawei. Both firms benefited, but MediaTek benefited a lot more.

It was that mix of technology, supply chain, finance, and geopolitics, all melded together. Over the coming years, I converted my WordPress into a Substack. At one point, I started charging for it, and I was writing about all these topics across the entire semiconductor and AI supply chains.

I followed AI a lot. I did a little bit of AI when I was a quant, in addition to following semiconductors out of a passion. The Substack just grew and grew and grew.

For 4 years, I traveled all around the world and went to every conference I could. I was going to 40 conferences a year. I went to AI conferences like NeurIPS, ICML, and ICLR. Those are mostly researcher conferences, although some companies go there and present their research.

I also went all the way downstream to random conferences for, let's say, chemicals that are inputs into the semiconductor supply chain. I went up and down the stack, whether it was servers, networking, fabrication, or AI. I went to 40 conferences a year.

Some were super niche, with 300 people there, and they only spoke Japanese except for 5 people. I'd think, "Well, whatever. That is what it is." Others had 10,000 or 20,000 people and were huge. It was the whole spectrum and continuum.

I was able to go across the whole ecosystem. When you go to a conference 3 times, you actually know the language, you know people there, and you build all these contacts so you can ask them questions. I developed this whole ecosystem of knowledge, and I was covering the inflections at each of these points.

I was very curious technologically, but once something stood out—whether it was technology or supply chain—I learned from a conference where it would lead on a supply-chain or finance basis. I was writing about the whole mix.

Sometimes reports would be centered around technology and no one in finance would care. Other times, people would say, "Wow, this is the bottleneck," or, "This is the inflection that's happening," or, "This company is going to gain a ton of market share because they have next-generation technology." I would call it before anyone else on the Street, before any hedge fund, before anyone.

That was the start. As the Substack grew, in 2022 I started hiring people. My first 2 hires were people I had known from Discord for years. After that, my third hire was Mihir. He had worked at a hedge fund before and was moving to Japan to live with his wife, so he was sort of a free agent.

I had put up a post about how, at the time, it was interesting that in early 2023 memory was the biggest loser from AI. The reason was that the amount of memory AI chips and AI servers used, versus regular servers, was a lot less.

In regular servers, about half the bill of materials was memory; in AI servers, it was much less. Part of that was that NVIDIA's margins were much higher, and part of it was that there were a couple of different factors. Of course, NVIDIA's next-generation chips have increased the memory content dramatically, and now it's way more. But at the time, the argument was that memory was the biggest loser.

In the paid section, I said, "Hey, I'm hiring." Myin reached out, and he was the first person from a hedge-fund background. The other 2 people had technical backgrounds. As soon as he joined the firm, we started making all these models, and we really converted the business from what it had been.

We still do the newsletter, and we still post a lot of amazing content there—more than ever before. But we converted the business into one centered around selling information services: selling these reports and selling these data sets.

As that started to happen, the ball started tumbling down the hill. From 2023 to 2024, I went from 2 people to 7. Then, from the end of 2024 to the beginning of 2025, I went from 7 to 20. From 2025 to 2026, I went from 20 to 60. Now we're at 90, so we've added 30 people this year.

It's just been a ball tumbling down the hill, and we've added sectors. I've always been interested in everything, but now I've been able to add people who are experts. I think the most exciting thing about SemiAnalysis is that I don't know another firm with the level and concentration of expertise that we have.

I have people who have worked at ASML, Applied Materials, and Lam Research—the equipment companies that build wafers—all the way upstream to people who have worked at Intel, TSMC, NVIDIA, Microsoft, and Amazon. We also have people who have worked at OpenAI on models, someone who has worked at Tesla on FSD, and someone who has worked at Cohere.

We've got people who have worked on the model layer, and then, on another vertical, we've had people who have worked on data centers.

There’s someone at my company who built a power plant in Kazakhstan. We just have this insane talent density, which is awesome. Half the firm is people who have done engineering across the industry, and the other half is either ex-hedge-fund people or just random people from the internet who are super passionate. I found them on Twitter or Discord, and I’m like, “You’re smart. Come work for me.” And it works.

This is what’s built up, and now SemiAnalysis has many different lines of business: obviously data services, consulting, and information services. We have the newsletter. We’re doing all this media. We’re having a conference soon that’s going to be big, and all sorts of different stuff that we do. It’s just a hell of a ride.

Christopher Gannatti

Speaking of the ride, I had a moment, Dylan, where WisdomTree and SemiAnalysis have been on this journey, working together for a number of months. NVIDIA GTC happens in March, as it does every year, and I’m sitting here watching the livestream in Charlotte, North Carolina. I think there were 55,000 other people watching the livestream, and suddenly—

Dylan Patel

There are 20,000 people in a stadium. There are 20,000 people in the stadium, dude.

Christopher Gannatti

And he’s referring to you directly. You said he was sandbagging a certain number, and your charts were right up on the stage. Admittedly, I had a moment, probably on your behalf, watching the CEO of the biggest company in the world refer to your research and how you were criticizing his take on some of the numbers. I’d love for you to tell us about it. I guess you were in the stadium, from what it sounds like.

Dylan Patel

Yeah. That moment was quite surreal. Basically, one of the things that SemiAnalysis does is we have a number of engineers, and we do open-source benchmarking of all the AI models that are open source, as well as all the hardware. It’s a pretty awesome effort. There are a number of engineers on my side, but we collaborate with the industry heavily.

We get hardware, so we have over $50 million of hardware donated to us from companies like OpenAI, Microsoft, Amazon, Google, CoreWeave, Nebius, and Crusoe—all of the major clouds you can think of have donated to us. Oracle has donated hardware to us that we run these benchmarks on. We have 8 different kinds of GPUs: H100s, H200s, Blackwell, and AMD’s various GPUs. In addition, we have TPUs from Google and Trainium from Amazon.

What we do is run benchmarks on the latest version of software every single day. The reason is that every night a CUDA version could be released, a PyTorch version could be released, a driver update could be released, or an inference-engine version could be released. We run these benchmarks every single night on the entire curve—from how fast you want the tokens to how cost-effective you want them in the optimal scenarios. We run all of this every single day. It’s an automated benchmarking suite that runs.

When Jensen originally launched Blackwell, he claimed it would be a 25x improvement. At the time, no one believed him. It’s Jensen, right? He’s marketing. Even we at the time were like, “Oh, okay.” We were more bullish than ever. We thought it could be a 15x to 20x improvement based on some of the simulations we were running, because we have a simulator for performance.

As we built out this inference benchmarking called InferenceX, we got to the point where we realized, “Oh, wow. In DeepSeek V3, Blackwell is 30x faster than Hopper somewhere on the continuum.” I emailed him as soon as we had the results, and they were automatically published to the open-source GitHub. It’s an open-source collaboration; NVIDIA people helped, and they knew. But I highlighted it to him. I was like, “Hey, Jensen. Back in 2024—or back when you launched Blackwell—you said 25x, and everyone gave you crap. Here are all the people that gave you crap. Even I gave you crap. I was like, ‘There’s no way it’s 25x. It’s maybe 15x to 20x.’”

A lot of people were like, “No, no, no, it’s like 3x.” We were quite bullish, but, Jensen, I was wrong. You were sandbagging it. It was 30x. He took that, and I didn’t know that he was doing anything with it. I’d heard from a couple of customers. Someone at Meta told me there was a meeting they had, and Jensen was using that as proof that he doesn’t sandbag numbers. He was talking about the next-generation chip, and anyway, this all happened. I didn’t expect it to happen onstage.

In addition, in InferenceX, we created this belt. It looks like a WWE belt, and it says “Inference King.” We sent it to all of our collaborators. We sent it to NVIDIA, AMD, and some of the other folks—SGLang, vLLM—all these different people who helped us with the benchmark, as well as people who donated hardware.

It’s an open-source effort where I’m spending a couple million dollars a year on engineer salaries, and other people are spending millions on hardware and donating it, or millions on engineer salaries and donating them toward this open-source effort that we run. I sent Jensen this belt, and he had it on the slide. He held it up, and then there were our charts. He was on the slide for 5 minutes talking about how, “Dylan said I was sandbagging, but I wasn’t. Our performance is the best.” It was such a surreal moment.

Christopher Gannatti

He talked about us longer than anyone else in the entire presentation. The only other thing he talked about as much was OpenClaw, which is obviously taking the world by storm. It was an incredible moment.

Dylan, you mentioned a couple of things there. I think it’s quite interesting that you mentioned open source, and now that we’re transitioning a bit toward some of the recent developments and the markets, there have been discussions around the actual inference efficiency of some of the open-source models versus the closed-source models. There’s also, even to this day, a lot of investors who are questioning the ROI on all of this.

In just the last week or 2, we had a Bloomberg economist talk through how a lot of these AI initiatives may be potentially failing at some of the firms out there. I know you’ve highlighted how your firm is using AI extensively and really leaning into giving your employees lots of access to tokens. You just highlighted that you’re hiring. I’m curious: What’s your take on this end demand, and on the fact that this end demand really does drive the big buildout we’re seeing, which is full of all these constraints that have, at least over the last month—besides the most recent couple of days in the markets—been driving up a lot of the different stocks tied to some of these themes?

2. AI Spending Pays Off

Dylan Patel

Yes. I would say a few things. When you look at the overarching question here—ROI, are companies making enough money from AI, is this going to continue, and are the companies using AI actually getting value out of it?—it’s an overarching question that a lot of people have. When I look at it and think about it, there are a few ways to dissect it.

First and foremost, Anthropic is free-cash-flow positive, and they are profitable in Q2. Even in April, when they closed April’s books, they were profitable. In May, they were free-cash-flow positive and profitable. June looks like it’s going to be the same way. It’s not fully closed yet, but at least for 2 of the 3 months, they’ve been free-cash-flow positive and profitable.

Their recurring revenue has soared past $50 billion ARR, and they’re doing fantastic. That’s one side of the coin: Anthropic is printing. Obviously, there are a lot of companies that aren’t printing, but they’re getting there. OpenAI’s revenue has started to inflect as Codex’s adoption has grown, and others as well. These companies are all getting much more profitable. Anthropic’s gross margins are really, really high; they’re above 70%.

Ultimately, that’s one side of the coin. The other side of the coin, which you were alluding to, is: What about the companies’ spending on AI? At least at SemiAnalysis, we went from our annual recurring spend—I like to call it not ARR, annual recurring revenue; it’s ARS, annual recurring spend. Our annual recurring spend in November, before Claude Code really started to take off for us in December last year, was less than $100K.

What we had was a subscription to every model, or we had a subscription to the $200 tier for ChatGPT for every user, and that’s about it.

So we were spending less than $100,000 on this. If people wanted xAI or Claude, we’d give it to them as well, but our standard was giving everyone the $200 OpenAI subscription. That was the state in November, and I think we were on the bleeding edge even then.

But then Claude Code really started to hit its inflection point with Claude Opus 4.5 and 4.6, and so on and so forth. By the end of January, our ARS—our annual recurring spend—had hit $4 million. That’s because people were using Claude Code.

Today, it’s about $11 million. The highest, if we take a week of spend and multiply it by 52, we’ve been at $11 million. The highest we’ve ever had was $14 million. We’ve oscillated a lot based on what work people are doing, but right now the average looks to be about $1 million of spend a year for a 90-person firm.

That’s freaking insane, right? I just want to be clear. We’re spending more than a third of employee spend on AI, and we’ll probably get to half by the end of the year, depending on how Methos and other models start coming out and getting better and better.

That’s a huge amount of spend. Now the question is, what’s the ROI? I think there’s been huge ROI because we’ve been able to build products, sell more, and increase the efficiency of everyone in the company.

A lot of companies are questioning, “Hey, if I’m spending hundreds of thousands of dollars—if I take a really good developer and they make, call it, $300,000 a year or more, right? There are a lot of devs who make a lot more, of course—their spend on AI is starting to approach one-to-one for good developers.”

For non-developers, the spend ranges and can be lower. But even at SemiAnalysis, a lot of our biggest spenders are people who don’t know how to code. They just tell the model what they want, iterate, iterate, iterate, and get what they want.

You see this soaring spend per employee, and a lot of companies are now rightfully asking, “Hey, we blew through our entire AI budget for the year in Q1 or Q2. We blew through it already. Now what do we do?” The question is, do we cut spend, or do we cut elsewhere?

A lot of companies are saying, “Oh, maybe we need to slow down on AI spend.” But a lot of companies I’ve seen are starting to cut elsewhere. They’re cutting other SaaS products that they’ve used historically. They’re saying, “Hey, we can grow faster, so we’ll just do it.” They’re saying, “Hey, it’s okay to spend on AI. We’ll take the hit temporarily. AI keeps getting cheaper.”

As adoption soars, what I used to do 6 months ago is much cheaper with AI today. Of course, what I’m doing with AI today is much more extensive than what I did 6 months ago. There are a variety of different approaches people are taking.

Some people are even cutting employees instead of cutting AI. Some people are clamping down on AI, but those companies are going to get left in the dust in terms of productivity gains and what they’re able to build.

Christopher Gannatti

Gotcha. One of the ways to, I’ll say, mitigate some of the incremental cost is choosing cheaper, maybe sometimes less intelligent models—maybe not always being at the leading edge there. I’ve said it’s very early, I think, in terms of some of the rumblings there, but I’m curious: is there a point where firms like yours decide there are some use cases that are more optimal for using a DeepSeek V4-type model for a certain type of work, and then obviously you might need to rely on things that cost a bit more for things that require a lot more intelligence? Is that part of the calculus here?

3. Models Differ By Workload

Dylan Patel

I think that’s absolutely part of the calculus for some folks. You have to break out AI workloads into 2 types. One is, “Hey, this is AI integrated into a process that I have.” In those cases, it’s like, “Oh, when a customer sends me a document, I check it for XYZ. I put it into the model, the model checks it, and it’s done.”

There, I just need to hit some level of quality, and then from there I can stop improving the model and start decreasing cost by waiting for newer models, cheaper models, or cost efficiency. We’ve seen AI models improve at a rate of about 60x per year in cost. You take a quality level, and a year later it’s 60x cheaper.

People freaked out about DeepSeek because it was 600 times cheaper than GPT-4. That was actually about 2 years after GPT-4, so 3,600x—60x times 60x, 3,600x—and it actually ended up being 600x cheaper. Somewhere on the curve is how much cheaper it’s getting each year.

DeepSeek V3 versus GPT-4 was 600x cheaper in 2 years. As you step forward, it’s somewhere in that range. If you have a workflow and integrate AI into that workflow, then you get to a quality level, and then you go cheaper.

The other range of work is an AI assistant. That’s where I think there’s actually a bit of a misnomer. If I’m doing my day-to-day work and asking the model to help me with this, help me find this, or help me figure out that, cost optimization isn’t actually going to a cheaper model.

Cost optimization is oftentimes taking the newest model, because the newest model can be much more efficient. Claude Opus 4.6 would take 100,000 tokens to do a task, and it might take a couple of turns—me talking to it back and forth—so it might take 100,000 tokens and 10 minutes of my time. Claude Opus 4.8 can do it in a quarter of the tokens, 25,000 tokens, and it might only take 1 back-and-forth.

The cost is actually less because the number of tokens being generated is less, and the amount of time I’m using is less. When I look at a developer or someone doing intelligence work, how do I reduce the cost? It’s actually not by using a cheaper model. It’s by taking an existing task that could sometimes be done with the model after fighting with it, going to newer and newer models, and now it’s able to do it in just 1 iteration or 1-shot the entire workflow. It’s able to do it in fewer tokens.

What we saw from Claude Opus 4.6 to when Claude Opus 4.7 came out was that my cost actually fell for a week before it soared back up, because people were using it more and more. Why did it soar back up? People had to adjust to the new workflow: “Okay, the work I was doing is done. Let me do more.”

Likewise, when Claude Opus 4.8 came out after 4.7, the cost fell for about a week or a week and a half, and then it soared back up because people were like, “Oh, yeah, now I can do more work.” You have to measure productivity alongside cost.

When it’s an AI assistant, token efficiency is really important. This is why Anthropic has been beating OpenAI: its models are more token-efficient than OpenAI’s. Actually, OpenAI’s models, on the edge cases—in terms of leading science, leading math, and leading code—can oftentimes do a task that Anthropic’s models cannot, but they take 3 times as long and 4 times as many tokens.

Therefore, it costs a lot more, and the feedback loop between human and AI is not as rapid. It ends up being worse on a customer-perception basis. It’s one thing to say, “Hey, model, do this task,” and then come back and check whether the task is done. It’s another thing to say, “Hey, I have 4 hours to do this task,” and whether it’s 1 call to the model and it does work for 4 hours, or 4 calls to the model and it goes back and forth—which one does it better?

It turns out Anthropic, when you have this human-in-the-loop feedback loop, is actually way faster and better because it’s more token-efficient. That’s the main reason why we still remain a majority-Anthropic shop.

For some tasks, people do use OpenAI. Oftentimes, the tasks they let run overnight are the ones they give to OpenAI Codex. But most tasks they keep with Claude Code.

This is one of the interesting factors of what’s going on with the models and token efficiency: cost is a bit hard to parse out. For some tasks, you freeze the model quality and wait for the models to get cheaper, and for others, you actually just want the smartest model because it is cheaper.

Christopher Gannatti

Dylan, I was curious about your thoughts, shifting it a bit to the hardware side. I know earlier this year, in one of the newsletters—I’ve been a big newsletter fan for multiple years, if the audience hasn’t realized it—there was an article talking about memory.

Memory has usually been a cycle. Meaning, maybe it’s 18 to 24 months: you go up, and then 18 to 24 months, you go down. We know that it feels like almost everything is in shortage. If you’re involved in a component that goes into a data center, it feels like it’s not a question of whether you can even get the component.

It's more, okay, how long are you going to have to wait? Because it feels like in the world today, you can barely get any component. With your experience having looked across the hardware side, what do you expect is going to change with something like memory, which used to always be this commoditized product? You ride the upswing, you go through the downswing, and it just repeats, going back the last 40 years.

4. Memory Supply Tightens

Dylan Patel

Yeah. I'm not saying there aren't going to be cycles anymore. I think cycles will happen. Obviously, we're in a supercycle where the upswing is crazy, and there will be some downswing, and it'll be brutal as well. But the downswing—trough to trough—there's still a lot of growth, right?

So I think what's relevant now about memory and other components is the shifting phases of what's happening. Historically, upcycles would be up 50% for the end market, and therefore for commodity markets like memory, where pricing is more elastic, you'd end up with those stocks 2–3x. What we've gotten today is that, instead of being up 50%, spend has already doubled over just the last few years, and it's going to double again.

When you look at the elasticity of different end markets, memory pricing has gone up like 4x, and it's going to go up another 2–3x again, in addition to capacity growth. So you've got the stocks just ripping like crazy before going back down. What's really exciting about memory is that it's not just an end-market thing. It's not just that the market is ripping and it's a very elastic good, and memory is a commodity, and therefore its pricing is very elastic with end-market demand.

What's actually interesting—and this is something we wrote in 2024 when o1 came out—was that OpenAI released o1, the first reasoning model. It created a new boom of reasoning models that OpenAI, Anthropic, DeepSeek, and many others have been exploiting to get models to go after long-horizon agent tasks.

When we look at that, what's interesting is that when o1 came out, the immediate thing we noticed was that the workload changed dramatically. When we were doing chat, when you're talking to ChatGPT, you may send a prompt that might be 50 words or 500 words, but you're going to send a prompt and it's going to give you a response back. That ratio—the context length—is a few thousand. You might have a context length of, let's call it, 2,000.

When you're running inference, every time you generate a token, you read all the weights into the chip, you read all the context into the chip, you process a token, and then you iterate again. You read all the tokens, the context, and the weights. The context is called the KV cache, right? It creates this relationship between all these tokens.

What's interesting is that when you're running model inference on the weights side, whether the context length is 1,000 or 100,000, you still have to read all the weights. So memory intensity on inference is the same on the weights side. But on the side of the KV cache—the context—the memory intensity when you have 1,000 tokens that you're reading in versus 100,000 tokens is a humongous difference, even though the compute amount is roughly the same.

The compute amount is roughly the same because of KV-cache caching in memory and things like that, so you can sort of get away with it: your compute costs don't soar, but your memory costs soar. What we highlighted in our o1 note was that, in December 2024, we talked about the scaling laws and how pretraining scaling laws were giving way to reasoning scaling laws. o1 was a big step-function change. We talked about how the KV cache was going to explode because of reasoning and, therefore, memory was going to be the biggest winner.

We did that in December 2024, and multiple times in 2025, we were really excited about memory. But in January 2026, I think, is when we wrote the note saying that, at the time, people were like, "Okay, memory has gone up 50%. Is it the top of the cycle? Do we need to keep going?" And we wrote a note that was basically like, "No, no, no. I don't think you guys get it." Memory capacity is only growing 20–30% a year for the next 3 years, and yet demand is doubling.

What's going to end up happening is that memory prices are going to keep soaring. Users of memory who are less elastic, or less capable of adapting to the elasticity of pricing, will drop out of the market. Smartphones and laptops, because their costs are going to soar so much, will drop out of the market, and that's going to all give way to AI. The price is just going to have to soar until that happens because capacity is not going up enough.

Ultimately, our point there was that memory isn't a shortage, and this is not a short-term shortage. It's a shortage that's going to last years. What we've seen so far over the rest of Q1 and now Q2 is that memory has just been gangbusters. It's been soaring. There have been days where it's gone down 7–8% for some random reason, but ultimately the chart has been up and to the right.

That's not investment advice, but we see it continuing to soar because pricing continues to go up. We still haven't seen the high-end market get impacted yet, although we've had some Chinese smartphone makers in the mid-range and low-end, like Xiaomi, say their shipments are down 40%. Next year, iPhone prices have to go up. Next year, MacBook prices have to go up.

Right now, if MacBook prices or iPhone prices go up $100, that market's not going to adjust too much. But memory is going to keep getting more and more expensive until AI gets its fill. That means smartphone prices aren't just going to go up $100; they're going to have to go up a few hundred bucks.

At some point, there's going to be an equilibrium where AI gets the demand it needs and mobile and consumer hardware gets pushed down enough. Obviously, at some point people still need new phones and new laptops, so they'll still buy. We're going to have to reach a new equilibrium because memory capacity for this end market doesn't grow fast enough.

As we extend across the ecosystem, what really matters is that a lot of different components are in shortage. Who is taking elasticity? Who has an elastic price and who doesn't? An example is TSMC, which is not elastic on pricing. They're a pretty good company, pretty fair with their customers, and they partner long-term. They're like, "We'll take up prices 5–10%."

Memory companies are in a commodity market. They let the spot market and contract market, with supply and demand balancing, really adjust pricing. So you see pricing increase 2–3x, and someday you'll see pricing halve, because memory doesn't necessarily deserve an 85% margin, which is where it's headed, though. We're still not at 85–90% gross margins for memory, but we'll get there. Then at some point from there, it'll also halve back down to the 70s or maybe even lower.

We'll see this oscillation in memory. In TSMC, you don't see so much oscillation. In other areas, like ASML, we don't see much oscillation in pricing. They make equipment, but different parts of the ecosystem will oscillate differently based on, first, how much of the end AI demand flows through to them.

Different parts of the supply chain are going to have different exposure. For every dollar spent on AI, it might be $0.01 on this product, but it might be $0.05 on this product. So, obviously, there's a difference in this end market—memory versus something else. Ultimately, different end markets in the infrastructure supply chain will benefit differently in terms of demand.

In addition, what are the market dynamics there? Is it one where there's a monopoly or an oligopoly? Is it one where there's a very competitive, large market? Is it one where pricing is pretty stable and there are a lot of long-term agreements? Or is it quite a commodity market where pricing is based on supply and demand?

All of these factors determine what happens in a specific end market, whether it be memory or the shortages people are now talking about—MLCCs, PCB drill bits, PCB foil, copper foil, and all these random components. You'll go online and see, "This is the next shortage. This is the next shortage." What matters is how much flow-through of demand there actually is. Is this end market doubling? Is it going up 50%? Is it quadrupling? How much is pricing going to go up based on the market structure? These are what really determine what happens in the infrastructure supply chain.

Christopher Gannatti

And if you take that framework, it feels like year by year the market wakes up to exactly what you said: a new, quote-unquote, potential shortage. Earlier this year, we had the OpenClaw virality on various sites, which awakened people to the world of AI agents and all the possibilities. Taking the framework you just described, I'd be curious to hear your take on the CPU market, which, for the first 3 years of AI, I don't think I heard the word CPU, and this year I'm hearing CPU everywhere.

5. CPU Demand Inflects

Dylan Patel

Yeah. On the side of CPUs, what's interesting is that in some of our institutional research for our clients, in November last year, we started talking a lot about it. That's because OpenAI and Anthropic had started striking deals with Amazon, Google, Microsoft, and others to buy all the CPUs they had in their fleets and rent them out. Over the course of late last year and now this year, CPU demand has just been inflecting.

Let's talk about the reason first. Initially, when AI was training and doing inference—and inference was mostly a short-context thing—it was mostly just predicated on compute and networking, right? But as pre-training shifted to reinforcement learning, and as chat-style inference turned into agentic inference, we had this big inflection where CPUs became more in demand.

Why is that the case? In pre-training, you're training the entire web dataset into your model. Whereas in reinforcement learning, the model generates some synthetic data or a reasoning trace, and then it checks it against an environment. That environment may involve running unit tests on code. It may be a sandbox that looks like a website, or a sandbox that looks like an engineering system or some other platform that you would use, whether it be a website, a shopping site, or what have you.

It might involve compiling the code. Those environments require a lot of CPUs, whereas before, in pre-training, the actual processing of tokens didn't require much CPU; it was all the environment checking. I've generated these tokens—are they valid? What do they look like inside an environment, whether it be Python or a C compiler, or within a website if I'm trying to buy something through e-commerce? Whatever it is, as an agentic workflow, I'm testing these things constantly, and that requires a lot of CPU.

The flip side is live inference. When you're doing chat, it's, "Okay, I tell it something, it gives me an answer back, and we're done." I might ask it a few more questions, but that's it. But now, when I talk about agentic workflows where the model is making tool calls, it's, "Okay, I'm going to go search for this. I'm going to look this up in a database. I'm going to ask the Python interpreter, and I'm going to write a little bit of code to check my work. I'm going to write some code, compile it, and deploy it."

These agentic flows end up requiring more and more CPU because they have to actually interact with the regular world, right? It was one thing when the human was interacting with the model: I'm telling the model something, the model gives me a response, I read it, and I'm like, "Okay, copy and paste it into whatever it is." It's a different thing when the model is interacting with the internet, right?

There ends up being a lot more compute in the loop, a lot more AI in the loop—or, sorry, a lot more CPUs in the loop—that are bouncing the answers back and forth. Both reinforcement learning and agentic workflows need a lot of CPU.

Now what's ended up happening is that, okay, we need a lot of CPU, but let's evaluate the prior things in our framework. What is the market structure? There are a few people in the market. There's Intel and AMD. Arm is now releasing a CPU, and Arm's stock has gone gangbusters because of that, because they're a new entrant into the market that looks pretty competitive.

Then you've got Amazon, which is the leader in this, as well as Microsoft and Google, releasing their own CPUs that they've developed internally. You've also got NVIDIA releasing its own CPU. So you've got a lot of different competitors in the market, but up until 2 years ago, all of the market was Intel and AMD. Now Amazon has gotten a good amount of share, and NVIDIA and Arm are starting to get more share.

Ultimately, what happens in the end market is that Intel is actually able to increase its price. AMD is also able to increase its price, so they've both increased their pricing. They've obviously gotten demand to go up a lot. Amazon is able to extract incredible margins out of CPUs because they don't make them and sell them; they make them and rent them. Their Graviton CPUs are renting like crazy, and they've increased their orders massively.

NVIDIA, which was previously only selling CPUs attached to its GPUs, is now selling CPUs standalone with Vera. They've given guidance of $20 billion of CPU revenue. For NVIDIA, that doesn't really scratch the surface. It's like, okay, that's a few percentage points of growth. No, I'm just kidding. But when you look at other companies like Intel, AMD, Arm, and Amazon, which gets the revenue instead of just the sales revenue, there are huge things happening there.

Christopher Gannatti

Dylan, maybe on the back of that, with CPUs now, some of the discussion that I've heard has been that CPUs for agents are different than historical CPUs in some regards. The cores are more optimized for agentic activity, is what I remember hearing Jensen saying or implying around the Vera CPU.

There's also a lot of discussion around this GPU-to-CPU ratio, which obviously highlights maybe the direction of the demand and need for CPUs. Can you give us a little bit more color on each of those topics? The concept makes a lot of sense to people at a high level, but there are some technical things that are probably pushed under the rug, if there are any. I'm not sure if it's just marketing or if there's a reality to this.

Dylan Patel

When it comes to agentic workflows, the use of CPUs varies a lot. You have some agentic workflows where the model is running, and then I send a response—all the tokens—to some CPU workflow. I'm waiting on the CPU to do something, and then I send it back to the model and the model works some more.

The question is: Did the compute that the model is running on stall while you're waiting for the CPU? In some cases it does; in some cases it doesn't. In the cases where it does stall, the compute that's running the model just stalls while waiting for the CPU's response. Then the CPU needs to be architected very differently.

The basic concept is: Do I want more cores, or do I want faster cores? There's sort of a law within CPU architecture, which is basically that if you make the CPU core twice as big—which means I have half as many CPU cores on the chip—my performance doesn't go up 2 times per CPU core. My per-CPU-core performance may only go up 50%. Obviously, there's a lot of engineering involved, and the trade-off isn't that simple, but to simplify it, that's a simple way to think about it.

If I look at an NVIDIA Vera CPU, it has fewer than 100 cores, but those cores are faster than the AMD cores. AMD's leading CPU has 256 cores. So you've got this big delta in the number of CPU cores, but the NVIDIA core is faster. It's not twice as fast as an AMD CPU core. There's this trade-off that people are making in the design space.

For some workloads where the AI compute has to stall while waiting for the CPU, then you need to ask: Who cares if I have half as many cores and they're only 50% faster? In total, the performance of the CPUs is lower, but the per-core performance is higher. Therefore, I'm not waiting on the CPU cores as often. I don't need a super-parallel workload. What I really need is this one workload done now.

In that case, where the AI compute is stalling, I want to have the fastest core possible, and I'm willing to sacrifice multicore performance. That's true for some types of agentic workflows.

Other types of agentic workflows are different. If I talk about how I use Claude day to day, or how the team uses Claude—how we spend $11 million a year on Claude on an ARS basis—what is that? I'm calling Claude, and Claude is processing a bunch of tokens, but it's not just using me. It's batching hundreds of thousands of users together across all of its compute.

If I get the response back and now it's waiting on me to implement it, whether it's waiting on me or a CPU core to implement it somewhere, that's okay because the computer is still running, just not for me. It's running for other people. So if the CPU is slower but I get way more of them, it's a different sort of task.

Another question is: Is it the active use of AI, or is it what's AI-generated and then taking what AI generated and deploying it? The beauty is that, if we look at GitHub commits globally, they're up multiple times versus last year. It's not just that total GitHub commits are up 10% or 50%. They're up multiple times.

What that means is that all this code is being generated for the world, and people are deploying a lot of the code. A lot of the code is sloppy, but a lot of code is being deployed.

And when it gets deployed, it’s being put on CPUs. It might just be a web scraper, an analytical engine, or some business-process automation. That doesn’t necessarily need to be on a super-fast CPU core; it can be on a cost-effective CPU core.

When you look at the continuum, NVIDIA has built the highest-performance CPU core, but it’s not necessarily giving you the maximum number of CPU cores times the performance per core if you have a CPU chip. They’re actually not so great at that. Whereas if I look at AMD and Amazon, they have a lot more cores—hundreds—but they have less per-core performance.

You end up with this trade-off. ARM is on that end, too. Where in that continuum do you want to go? For some workloads, you do want Vera, and for some workloads, you want the Graviton or the AMD CPU. I wouldn’t say it’s as simple as that.

As far as the other question you mentioned, which is the ratio, it is indisputable that CPU demand is going up. We were the first to call it out late last year in our institutional research and in January of this year in our newsletter. Since we published that, some of these CPU stocks have ripped: ARM has gone up multiple times, Intel has gone up multiple times, NVIDIA has gone up multiple times, and AMD has ripped. These stocks have ripped.

But now the sell-side, which doesn’t really understand technology at all, is just making things up. It’s getting to the point where the ratio of CPUs to GPUs, or the ratio of CPUs to AI compute, is getting lopsided to the point where it’s more in favor of CPUs than AI compute. That’s false.

Just to reiterate, if you look at a full, all-out Blackwell, it’s like $50,000-something per chip. If you had a 1:1 ratio, CPUs cost something like $5,000. For $300 billion of Blackwell to sell, or $500 billion of Blackwell to sell, you would only get $30 billion or $50 billion of CPU sales. That’s another thing that people are starting to miss.

Yes, this end market is ripping. Ultimately, the majority of the dollars are still going to AI compute and memory. This market was underpriced, and it’s more fairly priced now. I think that’s something that people need to recognize: it’s not like CPUs are going to keep growing and growing in demand beyond that of AI ASICs.

It’s a bit of a rightsizing. In 2023 and 2024, there were years of selling millions of AI chips and very few CPUs. Now, all of a sudden, CPU demand has inflected, and the ratio shouldn’t be here; it should be here. People are in catch-up mode, so now they need to buy a bunch of CPUs to catch up with all the compute they’ve historically bought, in addition to the compute they’re currently buying.

Once I catch up with that backlog of all these AI chips that I had bought previously, and I catch up on CPUs, that demand isn’t there anymore. I’ve already caught it up, and now it’s only the incremental demand. If there’s a ratio of, let’s say, 1 CPU to 2 GPUs, and each of those GPUs costs $50,000 while each of those CPUs costs $5,000, then for every $100,000 I’m spending on GPUs, I’m spending $5,000 on CPUs.

That’s still a great market dynamic in terms of CPU growth. It’s way better than it used to be. But if you think about having 10 million GPUs and AI ASICs that I shipped over the last 3 years without any CPU really attached, then that $5,000 has a huge catch-up component. That’s what we’re experiencing right now: a huge catch-up, as well as a shift in the ratio. You’re seeing demand be ridiculous, but it will eventually subside and reach a steady state. We’re in sort of a mini CPU cycle.

Christopher Gannatti

No, that’s excellent context. Really, really helpful. Then maybe just moving to networking to move into another area of the stack. I think this is one that has come to a lot of investors’ attention, particularly as they dive down into the optics supply chain and some of the constraints there.

We’re seeing some estimates that co-packaged optics is something that’s talked about a lot but is really probably going to be deployed around 2028—2027 or 2028. As you think about it, there’s obviously this concept of “using copper when you can, optics when you must,” and this kind of transition from optics to copper. We also had Jensen talking a lot about it at Computex, along with a lot of other discussions popping up, or at least bringing a lot of attention to firms like Marvell.

Is there any additional thought you have around optics and how you see the architecture of the data center within the networking domain evolving over the next 2 years?

6. CPO Ramp Gets Delayed

Dylan Patel

As models get bigger, how do we run them across nodes? How do we train models? There are a lot of different domains within the optical stack. There are telecom optics—companies like Ciena have been ripping, along with many of the constituent supply chains around them.

Then you have datacom: chip-to-chip communications. That has a copper domain and an optics domain today, and those are all ripping because the growth of networking content is faster than the growth of any other content, in percentage terms. Networking is going from below 10% to above 10% of the spend associated with AI chips. When we get to CPO, networking grows even further—it’s like 20% to 30%.

We’ve got this huge uplift in networking content. On the flip side, CPO is such a huge step-function change in the industry, and everyone recognizes CPO now. Currently, I think people are a little too excited about CPO. It’s not coming in 2027, in my view. It’s really coming in the tail end of 2028, but 2029 is the real ramp for scale-up co-packaged optics.

There have been a lot of problems. It’s a manufacturing thing: if we could deploy it today at a good cost, everyone would do it. But it’s really hard. The manufacturing volumes aren’t there, the yields aren’t there, and the chips aren’t really designed for it yet. It’s a very complex, difficult thing to ramp.

People are going to stay in copper as long as they can. That means Reuben is all copper. Fineman on the GPU is still copper, which is the next-generation NVIDIA GPU. After Reuben, there’s Reuben Ultra, then Fineman, and we’re not even at Reuben shipping yet. Reuben is just starting to ship, so we’ve got a few generations of chips before we get to co-packaged optics on the GPU.

There’s co-packaged optics on switches, which is coming earlier than on the GPU or the AI ASICs. Ultimately, even without that, as the cluster size gets bigger, you need more optics per GPU, or more active electrical cables and things like that.

We’ve seen this big dynamic shift. On Monday, we released a note at SemiAnalysis for our institutional research subscribers, which was on a localized timeline—not saying anything about the end market, right? Obviously, our view is that CPO is going to happen; we’ve been pushing that for a long time. Our view is that copper will be subsumed over time, but on a medium-term basis, we’re actually very bullish on copper, very bullish on optics that aren’t CPO, and actually kind of bearish on CPO because of certain delays on chips that we see downstream.

Fineman isn’t going to be full CPO, along with other things that we’ve seen. Copper names like Amphenol, which makes all the backplane connectors and cables, are actually going to do way better over the next few years than previously expected because we previously thought CPO would ramp sooner, but now it’s delayed.

These things happen in the supply chain. Ultimately, optics is an area that, if you close your eyes today and open them 5 years from now, is going to be way bigger. A lot of that is priced into stocks, a lot of it isn’t, and I’d say there are some localized dislocations.

That’s part of the research that we do, and the work that we’ve been doing with you folks as well: how do we weight that? How do we weight what amount is CPO-favored optics versus non-CPO optics, typical optical transceivers versus copper? Copper has actually got a long way to go. There are a lot of things happening in the copper industry that are innovating and pushing back CPO.

Why would I do CPO? At the end of the day, integrating optics is so much more expensive than sending something electrically. Except if I have to send something electrically, I can’t go that far unless I add repeaters or optics. There’s this trade-off and continuum, and CPO will happen, but it looks like it’s getting pushed out a little bit.

Christopher Gannatti

And Dylan, as we go into what is probably our last overall topic, we’ve done models, GPUs, CPUs, memory, and networking.

We'd probably be remiss not to mention the elephant in the room at any data center: How are you getting the electricity, and how are you getting that electricity into the right form factor? I know you've written about certain things in the newsletter, at least, about direct current versus alternating current.

When you see the hyperscalers spending all this money building these data centers, potentially even putting the power plants on-site, behind the meter, how are we to think about the electrical demand—the grid versus non-grid? I know it's a big topic, but I feel like we'd be remiss not to at least mention it here.

7. Data Centers Need Power

Dylan Patel

Yeah. I would say data center growth is massive. This year, we're deploying 20 gigawatts of data centers. Next year, that number goes up 50% to 30 gigawatts, and then it'll be 50 gigawatts the year after that. The growth in data center capacity is massive.

There are a lot of local dislocations that people are having to deal with, and energy is one of the biggest ones. The other one is political, and the third is construction. Building data centers and getting permits and filings is politically difficult. People are trying to stop it, but the primary factor gating it is really the energy at the end of the day.

What's happening there is that energy can be broken down into a few things. There's generation: Where do I generate the electrons from? There's transmission: How do I transmit the electrons from where they were generated to the data center? And then there's conversion, because the power that gets transmitted is in a form factor that the chips cannot consume. The chips need to consume it in a different form factor. What does that conversion pipeline look like?

I think in all 3 of these areas, there are very bullish aspects. The transmission side is the hardest to be bullish on because of the regulatory and political difficulties with building more transmission capacity, the way local utility monopolies work, and how, if they build a utility line, they have to amortize it across all users, not just the individual user. There are all these various weird dislocations with transmitting power, and so building more grid capacity is difficult on a transmission basis.

But generation-wise and conversion-wise, there are 2 interesting things. Generation-wise, obviously, there's more generation happening on the grid. There's also this big shift toward generating power for the data center. We predict that in a couple of years, half of the power for data centers—the incremental new power for data centers—will be generated on-site, not off-site. Behind-the-meter generation is soaring.

We see this with the behind-the-meter tracker we have in our data center and energy models. I mentioned that someone on our team built a power plant in Kazakhstan. She's leading our energy model—Ellie. She has been tracking this, and we've been building a model of the entire grid: every generation asset, every transmission asset, all the load assets, as well as all the behind-the-meter work.

What's interesting is that we've seen this huge boom in behind-the-meter generation. There's been a lot of fighting on the permitting and regulatory side, whether it be people not wanting to allow air permits or people not wanting to allow the gas pipeline to be built to the site, or things like that. We've seen that with an Oracle data center. There are a lot of different aspects of this that are happening, but ultimately, the end state is that behind-the-meter generation is soaring.

A lot of it is gas. A lot of it was combined-cycle gas turbines from GE Vernova, Mitsubishi, or Siemens. But beyond that, there have also been a lot of different types of energy sources. There's reciprocating engines, industrial gas turbines, and various types of diesel engines. People have taken train engines, boat engines, and truck engines and converted them into power generation for data centers.

We see a sea of innovation happening there. It's not like we don't have the industrial capacity. The U.S. can make millions of reciprocating engines a year. These are just engines that burn fuel and spin, and it's pretty trivial to retool those to run on gas rather than diesel. But even if it's diesel, that's fine.

Then you stick an electric motor on it and basically back-drive it, and that generates electricity. You can do this in large volumes to generate power. We see 10-plus gigawatts of data centers that are going to be built with technologies like this—taking diesel truck engines, converting them to gas, which can be done at the time of production very simply, putting an electric motor on them, back-driving the motor, and then sticking them on-site at a data center.

You have hundreds of these powering a data center, and then you hire a bunch of people from car mechanic shops. These things need to be serviced, so they just run around servicing these diesel engines all day. You have some buffer, so that when they go down, you can service them and keep them going and have maximum power. Obviously, you need some batteries in between because you don't want the ups and downs of the data center to mess with or blow up the engines that you've got.

You've got this entire supply chain of behind-the-meter generation, which is exciting. In addition, in about 2 years, solar plus battery will be cheaper than gas. The supply chains for solar plus battery are difficult, and it depends on what level of reliability you want.

If you have just enough battery to get through the night, it's cheaper. But what if you need enough batteries to get through the night—3 days, right? Because it might rain for 2 days. How many nines of reliability do you want? Solar plus battery is getting cheaper and cheaper at an incredible pace because of China's manufacturing excellence and some of the subsidies, too.

It's going to get cheaper at some point to do solar plus battery. Then you've got space data centers, where you don't even need a battery. You just stick it in space, you've got a solar panel, and that's it. You've got this whole continuum of ways to generate power, whether it be taking diesel engines and making them gas engines, using combined-cycle engines, or, all the way downstream, saying, "Let's just ship the chips into space instead."

There's the whole continuum, and there's a lot of money to be made there. There's a lot of interesting, dynamic things to do there. That's why, actually, the largest data set and research vertical for SemiAnalysis—which you think is semiconductors—is actually data centers, energy, and industrials. We call it the DEI team: the data center, energy, and industrials team. It's a pun internally. The tag is @DEI team.

Jeremy leads that team. He came up with the name. Data centers, energy, and industrials is actually our biggest research vertical because we're tracking every data center and every power plant. When we identify a delay, or that something is happening, or that there's going to be a build, or that a company is going to have a certain number of data centers go online in a particular quarter, it's something that no one else in the industry can do. That's why it's one of our biggest verticals.

Everyone's interested in that. Google is interested in what Meta is able to deploy. Meta is interested in what OpenAI is able to deploy. All of these companies are also looking at what the supply chain is able to do and who has capacity, and all the investors are looking as well. That's our biggest data set.

But it's a market where it's very decentralized. In the case of memory, there are 3 names, so it's pretty simple. In the case of accelerators, there are just a few names. In the case of semiconductor wafer fabrication equipment, there are just a few names.

In this case, there are hundreds of names in the supply chain making all these random little widgets. There are dozens of companies building data centers, and there are dozens of companies trying to do different things, whether you're an independent power producer, doing it behind the meter, offering some sort of battery service, or doing all these other things. It's a very complex supply chain, but one that has a lot of dynamism. Ultimately, I think there's a lot of innovation happening.

While data centers will continue to be a constraint of sorts, they will also not be a constraint because it depends on how crazy you're willing to go. As I said, you can just take truck engines, convert them, hire a bunch of mechanics, and run a site like that.

It's not going to be the best. A lot of people say, "That's disgusting. How reliable is that going to be?" or, "That's going to be really annoying to do." But people are doing it, and it will work. It's a pain in the ass, but it will work.

You know, all the way to, “I’m going to shoot it into space.” It’s going to be a pain in the ass. It’s going to be really hard to make it work, but it will work. And so you’ve got solutions to the data center problem, whether it be going full dirty or going fully into space, whereas other parts of the supply chain, you don’t. I think that’s what makes this market so dynamic: you’re going to see people go up and down a lot.

Then I guess the other part—that’s on the generation and transmission side—but on the conversion side, the other thing is: How do you get the power from where it is generated or transmitted to what the chips want? There’s an entire supply chain of stuff going on there, whether it be IGBTs, silicon carbide, various types of MOSFETs, GaN, gallium nitride MOSFETs—all the names there. What happens when we go from 12-volt to 54-volt to 800-volt DC in the conversion supply chain? What happens with solid-state transformers as those get innovated?

All these things are happening in the space. What happens with UPSs—uninterruptible power supplies—battery backups, supercapacitors, and all these other different ways to smooth out the power, make it from the dirty, variable power that gets created on the left side to the super-clean power, but also the variable usage of it on the right side? How do you match that? That entire conversion pipeline is super, super exciting.

We had a blog on that and 800-volt just last week. We’ve talked more recently to our institutional subscribers about some delays that are happening there on the NVIDIA side as they delay it out of Kyber. Reuben Ultra—Kyber doesn’t have 800 volt anymore. So what does that mean for the supply chain? Well, it gets pushed out a little bit.

Christopher Gannatti

So, Dylan, I want to thank you profusely on our side. This is the first episode, if we think in terms of chapters. This will be the first time we’ve had Dylan on the podcast, but certainly not the last, because there’s a lot more information. As he said numerous times, everything’s changing all the time across the entire stack. It’s a bear to keep track.

Dylan Patel

The other thing I would say is this supply chain is so freaking crazy. A lot of times we talk about the big ones—memory, CPUs, data centers—but actually, when you drill down to the supply chain, the local bumps are very small and many. For a couple of months, we were talking about PCB drill bits—the drill bits that drill into PCBs for the holes for copper foil that goes on PCBs.

There are all these random small things in the supply chain that also have these dislocations. Also, the companies that exist in them are all over the world. They could be trading in Taiwan. They could be trading in Japan. They could be trading in Korea. They could be trading in all parts of the world. It’s not just easily accessible to investors.

I think that’s what’s really exciting about our partnership and the way we’re working together: We get to influence what’s going on. We get to talk a lot about these supply chain disruptions, but also what’s really interesting in the framework that I laid out earlier and the entire landscape that we’re trying to cover. I’m looking forward to coming back on the show more and to our other collaborations.