[BidClub_]
The a16z Show · · 99 min

Dylan Patel on the AI Chip Race - NVIDIA, Intel & the US Government vs. China

Erik TorenbergDylan PatelSarah WangGuido Appenzeller

YouTube
TL;DR
  • The Nvidia–Intel tie-up is a strategic endorsement that could lower Intel’s cost of capital while redrawing the PC and data-center map. Nvidia committed $5 billion after SoftBank’s $2 billion and the U.S. government’s $10 billion, but Dylan Patel still thinks Intel needs roughly $50 billion; Jensen Huang supplies a semiconductor “Buffett effect” before a larger capital raise. Guido Appenzeller called the integrated x86-plus-Nvidia product compelling and warned that when “two arch nemeses suddenly team up,” AMD and Arm face the worst possible news.

  • Huawei is technologically credible, but manufacturing volume—especially high-bandwidth memory—remains the load-bearing constraint. Huawei reached market with a 7 nm Ascend AI chip in 2020, later obtained roughly 2.9 million TSMC-made chips through intermediaries, and now proposes separate prefill and decode products with custom HBM. China can probably produce substantial 7 nm logic and perhaps reach 5 nm with existing equipment, yet Patel stressed that design is not production: HBM3 yields, etch capacity and the transition from stockpiles to domestic scale remain unresolved.

  • China’s rejection of Nvidia chips is simultaneously industrial policy, a risky capacity bet and potentially negotiating leverage. ByteDance and other model builders still prefer Nvidia because it is “way better,” but Beijing can compel domestic adoption while Huawei publicizes an ambitious roadmap—a maneuver Patel called “10,000 IQ,” with Washington “playing checkers while they’re playing chess.” China may temporarily backtrack if domestic supply cannot ramp fast enough, forcing a choice between sovereignty and deploying “super powerful AI” at U.S.-competitive volume.

  • The near-term Nvidia bull case rests on AI capex running materially above Wall Street’s model, not on further market-share gains. Bank consensus puts six hyperscalers at roughly $360 billion of capex next year; Patel’s data-center and supply-chain work points to $450–500 billion, with Nvidia largely growing alongside the market while defending share. OpenAI alone signed more than $300 billion with Oracle, and the industry bull case becomes “multiple trillions a year on AI infrastructure”—though Patel refuses to forecast beyond five years because “the fifth year is sort of YOLO.”

  • Nvidia’s moat is a repeated willingness to risk inventory, redesign late and ship working silicon before competitors finish revising theirs. Jensen allegedly ordered Xbox volume before Microsoft formally awarded the business, placed non-cancellable capacity bets above customers’ own plans, and added Volta’s tensor cores only months before fabrication. Nvidia commonly ships A0 silicon while one Intel data-center processor reached E2—roughly 15 revisions—capturing the cultural difference between “I hate spreadsheets. I just know” and quarter-by-quarter caution.

  • Amazon and Oracle can both gain AI-cloud share without having the best accelerator, because powered capacity and balance-sheet willingness are scarce. Patel expects AWS revenue growth to trough and reaccelerate above 20% as Anthropic, Trainium and GPU deployments fill Amazon’s spare capacity; Trainium remains “very hard to use,” but a lab serving a few high-volume models can hand-optimize it. Oracle’s hardware-neutral engineering and willingness to underwrite OpenAI’s demand make it the bolder counterparty, although whether OpenAI can pay more than $80 billion annually in 2028–29 remains the central risk.

  • GB200 economics are workload-dependent, and reliability can erase headline performance for teams lacking sophisticated infrastructure. Against roughly 1.6x H100 total cost, GB200 may offer only about 2x performance in some training cases but north of 6–7x per GPU for DeepSeek inference; the problem is that one failure now sits inside a 72-GPU NVLink domain. Some clouds consequently promise about 99% availability for 64 GPUs but only 95% for all 72, making “the blast radius of a failure” as important as benchmark speed.

  • The GPU market is tightening again while Nvidia’s next strategic problem becomes what to do with potentially $250 billion of annual free cash flow. Patel also used an ambiguous “$200 million” figure in the same sentence. Hopper capacity at several major neoclouds sold out as reasoning-model inference surged, Blackwell deployment took longer than Hopper, and prices bottomed months ago before creeping upward—small allocations remain easy, large immediate clusters do not. Patel’s preferred outlet is data centers and power, the bottlenecks to GPU growth, but even that may not absorb the cash without turning Nvidia into a culturally different company.

Digest · the substance, structured for research

1. Nvidia’s Intel investment converts rivalry into strategic dependency

  • Patel’s opening reaction was financial theater: Nvidia announced a $5 billion Intel investment, Intel’s stock jumped roughly 30%, and the stake was already “a billion-dollar profit” by his telling. More importantly, a major prospective customer is committing capital and product roadmaps, giving Intel validation before it returns to public debt or equity markets.

  • The product logic is unusually strong. Intel would package one of its chiplets beside an Nvidia chiplet, reversing a history in which Intel faced antitrust litigation over chipsets and paid Nvidia a settlement. Patel called the turn “poetic”: Intel is “sort of crawling to Nvidia,” yet an x86 laptop with fully integrated Nvidia graphics might be the market’s best device.

  • The announced checks remain small against Intel’s needs: $5 billion from Nvidia, $2 billion from SoftBank and $10 billion from the U.S. government, versus Patel’s prior estimate that Intel needed roughly $50 billion immediately. More strategic investors—he floated Apple as one possibility—could create a “Warren Buffett coming into a stock” effect before Intel raises the balance.

  • Appenzeller’s customer view was enthusiastic but brutal for competitors. Intel could reset its uncompetitive internal graphics and AI programs, including the “Gaudi F4” mentioned in the discussion; AMD now faces two historic enemies acting together, while Arm loses the pitch that it is the natural partner for anyone avoiding Intel. “It remixes the cards.”

2. Huawei entered sanctions as a near-frontier chip company

  • Patel asked listeners to begin in 2020, not with the current roadmap. Huawei submitted an Ascend accelerator to impartial public benchmarks, became the first to bring a 7 nm AI chip to market and had only a narrow technological gap with Nvidia before the ban cut off full foreign-supply-chain access. It had also surpassed Apple in TSMC orders.

  • The Trump administration’s restrictions severed that access just as Huawei had the ingredients to challenge the market. Nvidia accelerated while Huawei rebuilt around SMIC, pursued Korean memory and simultaneously used shell companies to place TSMC orders. Huawei had already trained significant models on its limited original Ascend inventory.

  • By late 2024, Patel said that intermediary channel had produced roughly 2.9 million TSMC-made chips from about $500 million of orders before authorities stopped it. He referenced reporting about a possible $1 billion U.S. fine for TSMC but explicitly was unsure whether it had been issued; importantly, visible Ascend deployment had not yet consumed all that inventory.

  • The 2025 H20 ban then removed what SemiAnalysis estimated as more than $20 billion of Nvidia China revenue. Nvidia wrote off inventory, later received permission to resell it and faced a harder decision: restart a supply chain that might again be interrupted, or let Huawei and Cambricon absorb a market whose present “domestic” capacity still uses foreign wafers and memory.

3. China can design frontier architecture sooner than it can manufacture it

  • On logic, Patel’s reading of the equipment rules was more permissive than their public label. Although Washington describes restrictions reaching 14 nm, he argued that the equipment actually blocked is primarily needed below 7 nm. China should therefore make substantial volumes of 7 nm AI chips and might stretch existing tools to 5 nm, albeit with harder economics.

  • Huawei’s roadmap splits inference into specialized products: one chip for recommendation systems and prefill, another for decode. Nvidia and multiple startups are making the same architectural separation, so Huawei’s surprising claim was not the split itself but custom HBM for decode—an approach Nvidia and AMD were also only preparing to adopt the following year.

  • The import trail still signals a bottleneck. China previously devoted 30–40% of equipment imports to stockpiling lithography, versus a historical fab mix around 17–18% and roughly 25% in the EUV era; now etch imports are surging. HBM needs through-silicon vias etched through each layer before stacks are assembled 12-high or 16-high.

  • Huawei had sampled HBM2 but, by Patel’s account, had not begun volume production of HBM3, a technology introduced years earlier. Equipment availability and yield learning both matter: China can catch up faster than it took to invent the technology because the process already exists, but a few months of imports cannot reproduce years of Korean capacity. Torenberg framed the eventual outcome as “a matter of when, not if”; Patel’s own emphasis was that manufacturing capacity and yields remain the bottlenecks.

4. Beijing’s Nvidia ban creates a dangerous transition gap

  • China can initially reject Nvidia because it is converting the 2024 stockpile into accelerators. The difficult interval comes when those parts run down before domestic logic and HBM reach volume. Patel expects China eventually to ramp, but “it’ll take a little bit longer,” potentially creating a period when Beijing backtracks and relaxes policy.

  • Commercial incentives point the other way from industrial policy. ByteDance was described as “begging for Nvidia chips”: it uses some Huawei and Cambricon hardware but wants Nvidia to build the best models and deploy inference efficiently. The government can mandate domestic purchases, yet that does not mean Nvidia has ceased to be competitive.

  • Smuggling and re-exportation continue at low to lower-medium volume, but cannot close a national deployment gap. China ultimately has to choose how heavily it weights an internal supply chain against keeping pace in powerful AI; otherwise it deploys “so many fewer AI chips” than the United States.

5. Huawei may be hyping strength precisely because it still wants Nvidia

  • Patel’s negotiating interpretation was delightfully adversarial: if China wants better U.S. chips, it should advertise domestic self-sufficiency, unveil “the most crazy” multi-year Huawei roadmap possible and announce that Nvidia is banned. American suppliers then warn officials that an irreplaceable market is disappearing and lobby for looser export limits.

  • He called that play “10,000 IQ”—“we’re here playing checkers while they’re playing chess”—while conceding that much of Huawei’s architecture is real. The exaggeration lies principally in manufacturing certainty: the roadmap treats hoped-for capacity and yields as if they already exist.

  • Export policy therefore cannot be reduced to selling everything or nothing. Patel proposed comparing China’s achievable volume at each performance tier, then deciding what U.S. products near or somewhat above that level can be sold. AI’s potential end market is far larger than semiconductor equipment, so maximizing current chip revenue alone is an inadequate objective.

6. Huawei is the competitor Jensen fears beyond China

  • Patel took Jensen’s description of Huawei as “formidable” literally. Before sanctions, Huawei surpassed Apple in TSMC orders and phone share across multiple markets, then began recovering without Western supply chains. Against that history, fearing Huawei more than AMD is not lobbying theater alone.

  • Jensen’s best argument is to make Huawei’s aspirational roadmap become geopolitical reality: persuade policymakers that manufacturing capacity is no constraint and that Huawei will capture China, the Middle East, Southeast and South Asia, Europe and Latin America. Patel’s objection is narrow but crucial—capacity and yield are genuine constraints, even if temporary.

  • The alternative strategy is to “Japanize” China: isolate it until hardware and software become hyper-optimized for its domestic market, like the unusual Japanese PCs Patel invoked, leaving the global platform to U.S. firms. But isolation can also force China onto a superior branch while Western hardware–software co-design settles into a local optimum. “I don’t know if it’s accurate, but it’s an interesting one.”

7. Nvidia’s measurable bull case is $450–500 billion of hyperscaler capex

  • Bank consensus had Microsoft, CoreWeave, Amazon, Google, Oracle and Meta spending about $360 billion next year. Patel’s site-by-site data-center, component and supply-chain model produced roughly $450–500 billion. Nvidia cannot meaningfully add share from its dominant base; the earnings lever is defending share while total infrastructure expands faster than consensus.

  • Oracle’s contract illustrates the magnitude. OpenAI committed more than $300 billion over several years, rising beyond $80–90 billion annually, despite lacking the present cash to pay it. OpenAI had reached roughly $20 billion ARR; estimates for the following year’s exit ranged from $35 billion to $45 billion, while projected annual cash burn ran around $15–25 billion before profitability targeted for 2029.

  • Stack similar revenue and fundraising across OpenAI, Anthropic and other labs, and $500 billion of hyperscaler capex becomes plausible. Nvidia’s maximal thesis is “multiple trillions a year on AI infrastructure,” with GPUs mediating business agents, coding and consumer companionship alike.

  • When pressed for Nvidia’s ultimate ceiling, Patel refused false precision. If AI repeatedly improves AI, value creation could reach hundreds of trillions, but supply chains are only visible three or four years out: “the fifth year is sort of YOLO.” Whether white-collar workers become twice as productive, replaced, or dependent on a constant token stream is beyond an investable five-year forecast.

8. Nvidia built its moat by repeatedly risking the company

  • Nvidia failed early and repeatedly made existential commitments. An industry veteran told Patel that Jensen ordered Xbox production before Microsoft formally awarded the contract—likely with more nuance than the legend preserves—but the order came first. Nvidia’s first successful chip similarly had to work from its only affordable mask set or the company would run out of money.

  • During crypto booms, Nvidia persuaded suppliers that demand was durable gaming, visualization and data-center growth, prompting them to build capacity. When crypto collapsed, Nvidia absorbed inventory write-downs; suppliers retained empty lines. AMD had more efficient mining silicon but declined to raise production aggressively, a sensible risk policy that surrendered the upside.

  • The behavior continued at hyperscaler scale. Nvidia booked non-cancellable, non-returnable capacity above Microsoft’s internal plan, then Microsoft raised its plan toward Nvidia’s number. Patel summarized the founder’s decision system with Jensen’s own line to his CFO: “I hate spreadsheets. I don’t look at them. I just know.”

  • Jensen’s pinball metaphor explains the time horizon: “The reason you win is so you can play again.” Winning funds the next generation, not a fixed 15-year plan. Founder memory matters here—he remembers nearly going bankrupt, yet concludes that Nvidia must keep taking comparable risks rather than optimize for predictable Wall Street quarters.

9. First-pass silicon and late redesigns turn vision into market share

  • Nvidia’s long-tenured engineering organization contains both visionary architects and operators willing to say, “We need to get this silicon out now.” Patel described one private engineering leader as almost mythical and another fellow as famous for cutting features that technologists love, preserving them for the next chip rather than delaying shipment.

  • The measurable edge is stepping. Nvidia often ships A0 silicon and occasionally A1; one Intel data-center processor reached E2, roughly its fifteenth revision. Each new stepping can cost about a quarter, so verification quality becomes a go-to-market weapon rather than engineering hygiene. Patel recalled that Intel was openly jealous that Nvidia “consistently delivered on the first revision.”

  • Nvidia also begins transistor-layer production, pauses before final metal wiring if necessary, then releases volume once validation arrives. That combination of simulation, verification and calculated work-in-process lets it respond before competitors finish revising their masks.

  • Volta is the canonical bet: after observing AI workloads on P100 Pascal, Nvidia added tensor cores only months before sending the design to fabrication. Without that late change, somebody else might have captured the AI accelerator market. Equally important, its software organization delivered drivers and infrastructure quickly enough for first-pass hardware to be useful immediately.

10. Nvidia’s balance sheet is becoming a strategy problem

  • Patel cited roughly $250 billion of annual free cash flow; the transcript also contains an ambiguous “$200 million” figure in the same sentence. Regulators did not allow an Arm acquisition, even the $5 billion Intel investment requires review, and no obvious acquisition can absorb hundreds of billions without creating antitrust or integration problems.

  • Small investments in CoreWeave, model labs and other neoclouds help diversify buyers, but remain “small fries.” Nvidia could finance an entire Anthropic, xAI or OpenAI round, yet selecting winners would alarm every unchosen customer and strengthen their motivation to adopt AMD, TPUs, startups or internal silicon.

  • The cleaner use is to expand the complement: data centers and power. Nvidia could finance data-center and energy infrastructure without operating the cloud layer itself, removing the physical bottleneck to GPU growth while letting increasingly capable cloud competitors handle rentals.

  • The unresolved risk is cultural. Pouring concrete and building power infrastructure require different people from designing accelerated computing, while endless buybacks invite comparison with Apple under Tim Cook—excellent supply-chain execution, but in Patel’s view little transformative investment for nearly a decade. “Nothing requires $300 billion of capital” while also offering an obvious return.

11. Amazon’s spare power can reverse its AI-cloud slowdown

  • Patel’s Q1 2023 “Amazon’s Cloud Crisis” argued that AWS was optimized for the previous era: elastic scale-out networking, custom CPUs and cost reduction, not tightly coupled AI systems maximizing performance per dollar even when absolute cost rises. Neoclouds would commoditize that advantage, and AWS subsequently became the weakest-performing hyperscaler.

  • His new call is a turn, not a retraction of those structural criticisms. AWS year-over-year growth should trough in the current quarter and reaccelerate above 20% as Anthropic, Trainium and Nvidia capacity starts generating revenue. In today’s shortage, having powered space to fill can matter more than offering the industry’s cleanest software or accelerator.

  • Amazon historically ran unusually dense facilities—about 40 kW racks when peers used 12 kW—and its obsessive efficiency made data halls feel “like a swamp,” hot and humid. Retrofitting networking and cooling is less elegant than purpose-built infrastructure, but inexpensive relative to GPUs; Patel cited 2 GW sites with power, transformers, wet chillers and dry chillers already secured.

  • That retrofit creates component winners. SemiAnalysis highlighted Astera Labs around $90 based partly on Amazon connectivity orders; Patel noted the stock reached about $250 the following month. His point was not that Amazon has superior architecture, but that additional networking and cooling hardware is immaterial if the resulting GPU racks can be rented.

12. Anthropic can tolerate Trainium because it serves a narrow workload

  • Trainium remains “very hard to use” and was described by one host’s portfolio company as nearly impossible in summer 2023. Patel did not claim the developer experience had become good. His narrower claim was that production inference already requires hand-written kernels, low-level optimization and sometimes assembly-like work even on Nvidia.

  • TPUs and Trainium use larger, simpler cores with less general functionality, which some Anthropic engineers reportedly prefer once operating at that low level. A general customer still struggles, but Anthropic can optimize a handful of models rather than support the world’s architectures.

  • Patel’s intentionally rough example was serving mostly Sonnet—“Sonnet 3.5, or sorry, 4.5, whatever it is”—while leaving other models on GPUs or TPUs. If Anthropic reaches tens of billions of ARR and one model carries most traffic, spending heavily to optimize perhaps $15 billion of Trainium capacity becomes economically rational despite poor usability.

13. Oracle won by underwriting demand Microsoft would not

  • Oracle combines a large balance sheet with little hardware dogma. It will deploy Arista Ethernet, white-box networking, Nvidia InfiniBand or Spectrum-X, backed by strong network and software teams. That flexibility makes it a natural counterparty for OpenAI’s extreme demand.

  • Microsoft’s exclusivity became a right of first refusal: OpenAI can present an $80 billion annual or $300 billion multi-year requirement, and Microsoft can decline. Oracle accepted the bet because OpenAI lacks a balance sheet and Oracle has one; Microsoft’s caution is defensible, but it created the opening.

  • SemiAnalysis mapped Oracle’s signed and prospective sites through permits, satellite imagery, power equipment, chillers, transformers and generators. Using a simplified GB200-era assumption—a GPU at roughly 1,000 watts, a whole system at roughly 2,000 watts, and about $50,000 of all-in accelerator capex per GPU—it estimated approximately $12 million of annual rental revenue per megawatt, then rolled each site’s quarterly energization into Oracle forecasts.

  • That method closely matched Oracle’s announced 2025–27 revenue path and most of 2028; unseen 2028–29 capacity created the miss. OpenAI’s ability to pay more than $80 billion annually remains uncertain, but Oracle buys GPUs only one or two quarters before rental. Its earlier commitment is mainly data-center capacity, limiting stranded-asset risk, with debt available for later GPU purchases.

14. xAI’s advantage is treating regulation as another engineering constraint

  • AI infrastructure has moved from percentage growth to order-of-magnitude growth. A 100,000-GPU cluster exceeding 100 MW was once extraordinary; Patel’s team now tracks roughly ten of them and reacts to another 200 MW site with a yawning emoji. “It’s only exciting if you do gigawatt scale.”

  • Elon Musk’s first Memphis build still stands out: xAI bought a factory around February 2024 and trained models within six months on roughly 100,000 GPUs. It deployed large-scale liquid cooling, mobile substations, generators and CAT turbines, and tapped a nearby natural-gas line—galvanizing resources while conventional operators would search for another site.

  • Colossus 2 repeats the feat near gigawatt scale. After political and environmental resistance limited expansion in Memphis, xAI bought another facility roughly ten miles away, placed it near the Mississippi border and acquired a power plant in Mississippi, where regulation differed. Patel’s first-principles summary: others say power cannot be built there; Musk says, “Just go across the border.”

15. GB200 rewards the right workload and punishes weak operations

  • SemiAnalysis estimated GB200 total cost of ownership at roughly 1.6x H100. If a workload sees only 2x performance, the upgrade is worthwhile but marginal; for DeepSeek inference, Patel cited more than 6–7x performance per GPU with continuing optimization, turning a 60% cost premium into roughly 3–4x performance per dollar.

  • B200 is operationally simpler: eight GPUs in a conventional server, with less upside but familiar reliability. GB200 NVL72 creates a coherent 72-GPU domain and much larger inference gains, yet heat, novelty and system complexity make it finicky. A single GPU failure has a far larger “blast radius” than one failure in an eight-GPU box.

  • Sophisticated labs run high-priority work on 64 GPUs and low-priority jobs on eight, borrowing a healthy GPU when another fails and postponing physical service. Clouds cannot casually rent those spares to another customer because the NVLink domain must remain coherent.

  • SLAs now encode the compromise: Patel’s stylized example was roughly 99% uptime for 64 GPUs but 95% for all 72, varying by provider. Large labs can schedule around that; small companies may lose the theoretical gain through idle GPUs, downtime and an inability to mix high- and low-priority workloads.

16. Rubin CPX turns prefill into a distinct silicon market

  • Modern training increasingly consists of inference-like reinforcement-learning generation, while inference itself divides into prefill and decode. Prefill processes the prompt and constructs the KV cache; decode autoregressively emits tokens. They stress hardware differently enough that leading labs already place them on separate GPU pools.

  • Chunked prefill can fill unused batch capacity, but it slows concurrent decode. Disaggregation lets operators autoscale long-input and long-output traffic separately while guaranteeing time to first token—the latency users notice most. “I can’t read that fast anyways,” one host observed, but users still abandon products that hesitate before streaming.

  • Decode repeatedly moves model parameters and user-specific KV caches, making memory bandwidth central. A 64,000-token prefill request instead contains enough computation to occupy an accelerator by itself, making raw FLOPS more valuable than loading parameters rapidly.

  • Nvidia’s Rubin CPX specializes for that compute-heavy prefill phase and strips out expensive HBM, which Patel said represents more than half a GPU’s cost. If Nvidia passes through comparable margins, CPX can make long-context inference materially cheaper without forcing data-center operators to redesign the surrounding facility.

17. Large GPU buyers are back in a tightening spot market

  • Patel’s procurement rule was deliberately memorable: “How you buy GPUs, it’s like buying cocaine.” Buyers call or message several suppliers—“Yo, how much you got? What’s the price?”—rather than conduct a pristine enterprise RFP. His team maintains Slack connections with roughly 30 neoclouds and circulates live customer requirements.

  • Several major neoclouds had sold out of Hopper while their Blackwell capacity was still months away. Reasoning-model inference and revenue accelerated faster than supply, while Blackwell’s reliability and deployment learning curve delayed usable capacity relative to Hopper, which could be installed and producing within a month or two.

  • Hopper pricing bottomed roughly three to six months earlier and had begun creeping upward. Patel stopped short of declaring a return to the 2023–24 shortage: a small number of GPUs is easy to obtain, but a large cluster available immediately is difficult. Capacity, once again, matters before the formal price sheet.

Dylan

How you buy GPUs is like buying cocaine. You call up a couple of people, text a couple of people, and ask, “Yo, how much do you have? What’s the price?”

Guido Appenzeller

If your 2 arch nemeses suddenly team up, that’s the worst possible news you can have. I did not see this coming. I think it’s an amazing development.

Dylan

Like Warren Buffett coming into a stock, Jensen is like the Buffett effect for the semiconductor world. It’s kind of poetic that everything’s gone full circle and Intel is sort of crawling to NVIDIA.

Erik Torenberg

Dylan, welcome back to the podcast.

Dylan

Thanks for having me. It just so happens that there’s some big news as we’re having you on: NVIDIA announced a $5 billion investment in Intel, and they’re teaming up to jointly develop custom data centers and PC products. What do you think about the collaboration?

I think it’s hilarious that NVIDIA can invest, have it get announced, and their investment is already up 30%. A $5 billion investment, a billion-dollar profit already, right? I think it’s fun because they need their customers to really have buy-in. So when their potential customers buy in and commit to certain types of products, it makes a lot of sense.

It’s kind of funny because, in the past, there was this whole thing around how Intel was sued for being anticompetitive with its chipsets, and NVIDIA actually got a settlement from Intel way back when, when the graphics were separate from the GPU and were really put on the chipset, which had all this other I/O, like USB and all this stuff.

So it’s kind of a funny turn of events that now Intel is going to make a chiplet and package it alongside a chiplet from NVIDIA, and then that’s a PC product, right? It’s kind of poetic that everything’s gone full circle and Intel is sort of crawling to NVIDIA. But actually, it might just be the best device, right?

I don’t want an ARM laptop because it can’t do a lot of things, so an x86 laptop with NVIDIA graphics fully integrated would probably be the best product on the market. Are you optimistic? How do you think this will go?

I mean, sure. I hope so. I’m a perpetual optimist on Intel because I have to be. I was thinking that the structure of the deal that a lot of the government folks and Intel were trying to pursue was that big customers and the biggest suppliers would directly give capital to Intel.

But this is sort of the other way around, where they’re buying some of the stock and having some ownership, but they’re not really diluting the other shareholders. The other shareholders will get diluted—everyone will get diluted—when Intel finally does raise capital from the capital markets.

But because they’ve announced these deals, and they’re pretty small—$5 billion from NVIDIA, $2 billion from SoftBank, and $10 billion from the US government—these are still relatively small.

Dylan

Pretty small.

Yeah, in the grand scheme of things. I mean, last time I think I said Intel needs something like $50 billion right now. When they go to the capital markets, it’s better, and hopefully they get another couple of these announcements.

There’s all sorts of speculation that Trump is involved in getting these companies to invest in Intel. Now the government is involved as well, of course. Is Apple going to come invest and also do something with Intel? Who else will come in?

That will really boost investor confidence, and they can dilute or go get debt, like Warren Buffett coming into a stock. Jensen is like the Buffett effect for the semiconductor world.

Guido, you were the CTO of Intel’s Data Center and AI Group. What are your thoughts?

Guido Appenzeller

I think it’s really good for customers and consumers in the short term. Specifically for the laptop market, having Intel and NVIDIA collaborate is amazing.

I wonder what’s going to happen with any of the integrated graphics or AI products at Intel. They might just push a reset and give up on that for now. They currently don’t have anything competitive. There was the Gaudi F4, which is more or less done, and there were the Intel graphics chips, which never really competed at the high end.

From that perspective, it makes a lot of sense for both sides. For Intel, they needed a breath of fresh air. They were sort of desperate, so I think it’s a very good thing.

I think AMD is screwed, right? If your 2 arch nemeses suddenly team up, that’s the worst possible news you can have. They were already struggling. Their cards are good; their software stack is not. They were getting very limited traction, and now they have a bigger problem on that side.

I think Arm is a little bit screwed as well, because its biggest selling point was, “Look, we can partner with everybody that doesn’t want to partner with Intel.” In a sense, NVIDIA is probably the most dangerous of the future CPU competitors. Now NVIDIA suddenly has access to Intel technologies and might go in that direction.

It reshuffles the cards. I did not see this coming. I think it’s an amazing development.

Erik Torenberg

Yeah, it will be very interesting to see this play out. To my point, it’s been a packed news week.

The other thing we wanted to pick your brain on, since we have you here, Dylan, is the other news about Huawei unveiling its AI roadmap. Obviously, they’re hyping up the capabilities. I think you guys have been ahead of the curve in trying to gauge what the Atlas 950 SuperCluster can actually do.

I’d love your thoughts on everything that’s going on on the China front. This is coupled with DeepSeek saying its next models are going to be on domestically produced Chinese chips, and the Chinese government banning companies from buying the NVIDIA chips produced specifically for China.

There are just a lot of dominoes falling right now in the semiconductor market in China. I’d love your take overall, and I’d like you to drill into some detail.

Dylan

When you zoom out, let’s walk from 2020, because I think it’s really important to recognize how cracked Huawei is, even historically. They’ve always been really good. Sure, initially they stole Cisco source code and firmware and all this stuff, but then they rapidly passed Cisco up, as well as every other telecom company.

In 2020, they released an Ascend chip and submitted it to impartial public benchmarks. They were the first to bring 7-nanometer AI chips to market. They were the first to do that. You could still say NVIDIA was ahead, but the gap was almost nothing. The market was so nascent then that Huawei could have really taken it over.

Huawei got banned by the first Trump administration from accessing the full foreign supply chain, and that went into effect in 2020. They were only able to make a small volume of these chips, but they had trained significant models on the chips they made.

Over the next couple of years, NVIDIA continued to accelerate. Because Huawei was banned from TSMC, it had to figure out how to manufacture at SMIC, the domestic TSMC. In parallel, Huawei was trying to use shell companies to manufacture at TSMC and acquire memory from Korea, and so on and so forth.

By the end of 2024, this had gotten into full swing, and it was caught. They finally shut it down, but Huawei was able to acquire 3 million chips—2.9 million chips—from TSMC through these other entities, roughly $500 million worth of orders.

That ends up being a $1 billion fine that the US government gave TSMC, if I recall correctly, or at least there was a Reuters article about it. I don’t know if they actually issued it, which is important and interesting to gauge, because the number of Ascends floating out there has not consumed this entire capacity yet.

Now we get to 2025. The H20 got banned at the beginning of the year. NVIDIA had to write off huge amounts of money. Our revenue estimate for NVIDIA in China for just the H20 was north of $20 billion, because that’s what they were booking in capacity and then had to write off.

They cut the supply chain. They just said, “No, we’re not doing this anymore.” Their inventory got reapproved, and they resold the inventory, but now NVIDIA’s question is, “Do we even restart production?”

China is now saying, “We don’t need NVIDIA. We have domestic alternatives,” whether it’s Huawei or Cambricon. These companies have capacity, but most of this capacity is still foreign-produced, whether it’s wafers from TSMC or memory from Korea—Samsung and SK hynix.

So the question is, how much can they do domestically? There are 2 fronts. There’s logic, meaning replacing TSMC, and there’s memory, meaning replacing SK hynix, Samsung, and Micron.

Dylan Patel

On the logic side, they are behind, but they’re really ramping there. I think they can get to the production-capacity estimates needed, and the US is still allowing them to import pretty much all the equipment necessary. The bans are really for beyond the current generation of technology—beyond 7 nanometers. Even though the government says they’re for 14 nanometers, the actual equipment that’s banned is only for below 7 nanometers.

They’ll be able to make a lot of 7-nanometer AI chips and maybe even get to 5 nanometers using existing equipment for 5 nanometers, rather than using the new techniques. There’s the logic side and then there’s the memory side. The aspect of Huawei’s announcement that was surprising was that they’re doing custom memory.

Speaker 2

Right? Yeah.

Speaker 1

That’s the part that’s really exciting. They announced 2 different types of chips for next year: 1 focused on recommendation systems and prefill, and 1 focused on decode.

Speaker 2

There’s the trend these days.

Speaker 1

Yeah. At NVIDIA, it’s the same thing. They just announced a prefill-specific chip recently. Numerous AI hardware startups are really focusing on prefill versus decode, so this split of inference into 2 workloads is becoming common. Huawei is doing the same thing for its chip next year.

What’s interesting is that the decode one has custom HBM. What does that mean? What’s the manufacturing supply chain? That’s the tricky one, right? How much of that custom HBM can they manufacture? NVIDIA and others are also adopting custom HBM starting next year, so it’s not like the manufacturing capacity isn’t there. Maybe it’s going to consume a bit more power, and maybe it’s going to have slightly lower bandwidth, but the fact that they’re able to do some of the same things that NVIDIA plans to do and AMD plans to do in their memory is evidence that they’re catching up.

The main question that remains is production capacity. As far as NVIDIA being banned in China—China saying, “Don’t buy NVIDIA chips”—I think that’s fine for a period of time, from China’s perspective. If I’m China, that’s fine because you have all this capacity that you shipped in 2024 that you haven’t turned into AI chips. Now you’re turning them into AI chips and running that stockpile down.

What about the transition from running that stockpile down to ramping your new stuff? That transition is the one that’s really tricky. China is either shooting itself in the foot by not purchasing NVIDIA chips during that period, or it’s able to ramp.

Dylan Patel

I think they’ll be able to ramp. I think it’ll take a little bit longer, and there’ll be a gap in between where China probably backtracks and says it’s fine. ByteDance will be begging for NVIDIA chips. They don’t want to use Cambricon; they use some Cambricon and some Huawei, but they really want to use NVIDIA because it’s way better.

They don’t care about the domestic supply chain. They want to make the best models and deploy their AI as efficiently as possible. The government can mandate them not to do it. So it’s not that NVIDIA isn’t competitive; it’s that the government is trying to instigate it.

The last thing is that there’s always the argument of, “Hey, if banning NVIDIA chips to China is so good for China, why didn’t China do it for itself?” They’re finally doing it for themselves. So again, it’ll be interesting to see.

Smuggling is still happening. Reexportation of chips from other countries to China is still happening at some volume—low volume, lower-medium volume. Direct shipments of NVIDIA chips that are legally allowed to China aren’t necessarily happening today, but they may have to restart at some point because China won’t have the production capacity. It would just have so many fewer AI chips being deployed domestically versus the US, and at some point you have to pick: “Am I all about the internal supply chain, or am I all about chasing super-powerful AI?”

Erik Torenberg

Is there a negotiation angle here as well? There are still discussions ongoing about exactly what the boundaries are and what can be exported to China. These are sort of well-timed announcements if you want to make a point that the US should allow more exports. Do you think that’s a factor or not?

Dylan Patel

Yeah. In the report we did a few weeks ago about Huawei’s production capacity and the supply chain, there was a bit in there about how, honestly, if you’re China and you want NVIDIA, you do want NVIDIA chips. How do you play this? It’s by hyping up your domestic supply chain.

Erik Torenberg

And it’s like, yes, we can do everything. Huawei announced the craziest shit possible—7 years of shit, or 3 years of roadmaps that are so—did they read your report? Basic question.

Dylan Patel

I think they do. I mean, say we’re banning NVIDIA, right? The government official is going to think alongside the lobbying from domestic players, “Of course we want to ship them better AI chips. We’re losing this market. We can’t lose this market.” It’s sort of 10,000 IQ. We’re here playing checkers while they’re playing chess.

Erik Torenberg

Negotiation aside, in that report you talked about HBM, or high-bandwidth memory, being a bottleneck for Huawei. To your point about one of the surprising aspects of the announcement, do you think it’s credible that it’s no longer a bottleneck based on what they’re saying, or is it just hype?

Dylan Patel

Production-capacity-wise, it is still absolutely a bottleneck. Certain types of equipment required for making HBM need to be imported. They’re working on domestic solutions, but as far as we know, they have not imported enough equipment for this.

Although if you look at Chinese import data for different types of equipment, fabs spend—depending on the process technology—roughly different amounts of money on lithography, etch, deposition, and metrology, these different steps. Historically, lithography has hovered around 17% or 18%; with EUV, it grew to 25%. But China wanted to stockpile lithography and was worried about the coming ban, so it was importing lithography at a much higher rate. Thirty to 40% of its equipment imports were lithography, and it was just stockpiling lithography equipment.

That has sort of reversed now. If you look at the monthly import-export data, both into provinces in China and out of countries, you can see that etch—specifically—is skyrocketing. The main thing about stacking HBM is that with each wafer, you have to etch and create something called a through-silicon via so it can connect from the top to the bottom. Then you stack them on top of each other—12 high or 16 high for HBM. That’s how you make super-high-bandwidth memory, and their imports of etch equipment are skyrocketing now.

Erik Torenberg

They don’t have the production capacity yet. How fast can they ramp it as a function of how much equipment they can get? And B, the yields, right?

Dylan Patel

Improving yields is really hard in manufacturing. Intel and Samsung are really good, and TSMC is just amazing. Not that those companies suck; I think that’s a better way to put it. Those are the 2 things: yield and production capacity.

Yield-wise, they haven’t even started production of HBM3. They’ve only done some sampling of HBM2. HBM3 came out a few years ago, so there’s still quite a ways to go up the learning curve. I expect them to catch up faster than it took for the technology to be developed because it exists. We know how to do it in the world; it’s just a matter of actually doing it versus inventing it.

The other issue is production capacity. A couple months of import-export data isn’t enough to set up years’ worth of supply-chain buildup, which is what we have today in Korea for the Korean companies. SK hynix is also investing in the US, in Illinois, and Micron is primarily in Japan. The American memory companies are primarily in Japan and Taiwan, but they’re also expanding in Singapore and the US now.

There’s so much capital that’s been invested. It would take some time for China to build up that production capacity to actually match the West. When I say the West, I mean East Asia—in non-China East Asia—in production capacity. It’ll take some time to get there. I don’t think it’s a question of whether they can design this; it’s always a question of whether they can manufacture it. The thing Jensen would say is that you’re betting on China not being able to manufacture.

Erik Torenberg

That’s a matter of when, not if.

Dylan Patel

And that’s the whole calculus that I think the US government has to be aware of when it’s like, “Hey, what level of AI chips do we sell? Do we sell everything?” Probably not, because AI is far more powerful, and the end market for AI is going to be way larger than the end market for semiconductors and equipment.

Do we sell, you know, what level do we sell at? How much can China make at each specific performance tier? Then analyze that: What’s the volume? Figure out what is okay, which is maybe a little bit above or around the same level.

Erik Torenberg

Yeah. So, to your point on playing chess versus checkers, if you’re Jensen, what would your next move be, given the situation at hand?

Dylan Patel

It’s partially true that he’s more afraid of Huawei than he is of AMD.

Erik Torenberg

Right? He called them formidable.

Dylan Patel

Yeah. I mean, Huawei has beaten Apple, right? They passed Apple in TSMC orders, and they passed Apple in phone market share—not in the United States, but in many parts of the world—before the bans came down. Even now, they’re growing back again in market share without Western supply chains.

They’ve done this to numerous other industries. I would say Huawei is a formidable competitor, right? They’ve beaten a lot of industries, and so it’s reasonable that he’s afraid of them. It’s sort of—you know, he’s not afraid of AMD.

I think the best thing is to try and see whether what Huawei announced is reality rather than their aspirational target.

Erik Torenberg

Yeah.

Dylan Patel

We shouldn’t wave away all doubt about manufacturing capacity, which I think is not fair. Manufacturing capacity is a real bottleneck for them. Yield learning is a real bottleneck, although perhaps temporarily. We’ll see how long that lasts, how fast the rest of NVIDIA’s technology advances past what Huawei is capable of, and how fast Huawei is able to close the gap.

I think his main pitch would be that Huawei is real. They’re a formidable competitor, and they’re going to take over not just the Chinese market, but also foreign markets—whether it’s the Middle East, Southeast Asia, South Asia, Europe, or Latin America. Everywhere besides America.

I think Noah Smith has this analogy. The whole idea is that you should let China go, right? Make them have their own domestic industry that is so different from the rest of the world—kind of what happened with Japan in the 1970s, 1980s, and 1990s. Their PCs were so specific and hyper-optimized to the Japanese market, with the weird scroll wheel on these Japanese PCs. You literally go like this and it scrolls, and the touchpad is a circle with that around it. Things like that are so weird.

Erik Torenberg

Totally.

Dylan Patel

The rest of the world doesn’t care, but the Japanese market likes it, right? His whole idea is: Let them go—keep their technology within China—and then that’s deadweight loss, and they never expand outside of China, versus serving the whole world.

But the whole risk is that the opposite can also happen, right? Our technology is hyper-optimized to run language models at this scale and reinforcement learning. Hardware-software co-design can take you down a branch of the tree that is a dead end. Then China, because they’re not allowed to access this tree, says, “Oh, okay,” and ends up in the optimal spot, right? We hit a local minimum; they had a global maximum. That sort of technological goal-posting is what Noah Smith’s analogy is. I like it a lot. I don’t know if it’s accurate, but it’s an interesting one.

Erik Torenberg

Yeah. I love that. Well, actually, maybe just taking a step back from current events—even though there’s so much to talk about right now—last time you appeared with us, Nvidia came up, obviously, and you talked about a couple of the potential paths forward for Nvidia. Give us maybe the bull case, the bear case.

Dylan Patel

Fair enough. There’s a lot embedded in their numbers now. What’s interesting is that the consensus from the banks across the hyperscalers—Microsoft, CoreWeave, Amazon, Google, Oracle, and Meta—is $360 billion of spend next year across all of them. Those are the 6 hyperscalers I would consider hyperscalers.

My number is closer to $450–$500 billion. That’s based on all the research we do on data centers, tracking each individual data center, the supply chains, and so on.

Erik Torenberg

So this is just Nvidia spend?

Dylan Patel

This is capex for the hyperscalers. Capex gets split up across different companies, but the vast, vast majority still goes to Nvidia.

Nvidia is in a position where they can’t take share. They grow with the market and defend share.

Erik Torenberg

Yeah.

Dylan Patel

The question is: How fast is the growth rate of capex for hyperscalers and other users? The reason I included Oracle and CoreWeave as hyperscalers, even though they’re traditionally not called hyperscalers, is because they are OpenAI’s hyperscalers.

Erik Torenberg

Right. Right.

Dylan Patel

When you look at the Oracle announcement, I don’t understand why people don’t think this is crazier. They did the most unprecedented thing in the history of stocks, public companies, and companies ever: They gave 4-year guidance. It made Larry the richest man in the world, among all these other things.

The question is: How fast does revenue grow? Do you think OpenAI, which signed a $300 billion-plus deal with Oracle, will actually be able to pay $300 billion through raising capital and revenue? It gets to a rate of over $80 billion or $90 billion a year in just a handful of years.

Do you believe the market will grow that fast? It’s very possible. For OpenAI, what is their revenue going to be exiting next year? Some people think $35 billion, some people think $40 billion, and some people think $45 billion in ARR by the end of next year. This year they hit $20 billion.

If that growth rate is maintained, then all of that cost goes to compute, plus all the capital they continue to raise. The financials they gave investors for their last round were, “We’re going to burn $15 billion next year.” It’s probably more likely to be closer to $20 billion. They’re not generating cash flow, and they’re not going to be profitable until 2029.

They’re going to continue to burn $15–$25 billion of cash each year, plus revenue growth. That’s their compute spend. You do this for Anthropic, you do this for OpenAI, and you do this for all the labs. It’s very possible that the pie gets to more than $500 billion—not $360 billion—next year in total capex, and the pie continues to grow for hyperscalers.

Nvidia says it’s actually going to be multiple trillions of dollars a year in AI infrastructure, and Jensen is going to capture a huge portion of it. That’s his bull case: AI is actually so transformative that the world gets covered in data centers, and the majority of your interactions are with AI—whether it’s business productivity and telling an agent to write code, or you’re just talking to your AI girlfriend, Annie. All of this is running on Nvidia, for the most part. The bear case is, you know, even if it does grow a lot.

Erik Torenberg

Yeah, so the bull case for a second: I think fundamentally the value creation is there. I mean, trillions of dollars of value with AI—I can totally see this happening. So assume it’s true: Where will Nvidia top out?

Dylan Patel

I guess it depends on how much you believe in takeoffs. If there is a takeoff scenario where powerful AI builds more powerful AI, which builds more powerful AI, and each level of intelligence enables more for the economy, how many monkeys can you employ in your business versus how many humans? You know, what is the value creation of a human versus a dog? It’s sort of the same with AI.

Erik Torenberg

I mean, in this case, the value creation could be hundreds of trillions, if not more than that.

Do you even need this? If you take every white-collar worker and make them twice as productive with AI, that’s in the hundreds of trillions, isn’t it?

Dylan Patel

Yeah, but what is “twice”? If you talk to people at the labs, what does “twice as productive” even mean? It’s replaced them, right? It’s 10 times better than that. I don’t know how soon, but if white-collar work is essentially useless without a constant stream of LLM tokens that make people productive, at that point you basically can tax every single knowledge worker in the world, which is most workers in the world long term.

Erik Torenberg

Yeah. So I don’t know. What’s your guess? Give us a number. What’s the cap?

Dylan Patel

Cap? I mean, why aren’t we making a Matrioshka brain? I don’t know. At some point, the machine says humans don’t need to live, and we need even more compute.

Speaker 1

One step before that, maybe—

Speaker 2

Are we colonizing Mars yet?

Dylan Patel

TBD. I don't know, man. I find it completely impossible to predict anything beyond 5 years, given how much stuff is changing. Linear time is a large number; I'll leave that to economists, right? Honestly, supply-chain stuff is 3 or 4 years out, and that's it.

Speaker 1

And then the fifth year is sort of YOLO, right?

Dylan Patel

I just try to ground myself in the supply-chain stuff. What is the adoption of AI? What is the value creation? What's the usage like? You can see that in a short horizon. Beyond that, I don't know. Are we all going to be connected to computers through BCIs and stuff? I don't know, dude.

You saw Elon's thing, right? He's like, "Yeah, humanoid robots are why Tesla's worth more than $10 trillion." Okay, great. What is all that being trained on? Great, Nvidia. Okay, awesome. So that's worth also $10 trillion, right? I don't know. It's too out there for me. I don't like the out-there discussions.

Speaker 1

Very fair.

Speaker 2

Read some sci-fi books.

Speaker 1

So, just pulling out the thread where you talked about how market share can't really grow just because it's such a dominant market share. We talked about—or you guys talked about—the moat of Nvidia last time, and obviously this moat is tied to maintaining the very high market share they currently have. I love the historic journey you took us through with Huawei earlier. Can you walk through what Nvidia did throughout history to build its moat?

Dylan Patel

It's super awesome because Nvidia failed multiple times in the beginning, and they bet the whole company multiple times. Jensen is just crazy enough to bet the whole company, right? Whether it was certain chips, ordering volume before he knew they even worked, when it was all the money he had left, or ordering volumes for projects he had not won yet.

I heard a rumor—or, not a rumor, but a story—from someone who's a graybeard in the industry and who I think would know. He said, "No, Nvidia ordered the volume for the Xbox before Microsoft gave them the order." They were literally just like, "Fuck it. YOLO." I don't know how true that is. I'm sure there's more nuance there, like a verbal indication or whatever, but the order was placed before he got the order, is what he said.

With the crypto bubbles, there were a couple of them. Nvidia did its damn best to convince everyone in the supply chain that it wasn't crypto, that it was gaming, and that it was durable, real demand. It was gaming and data center and professional visualization, and therefore everyone should ramp production.

They all ramped production and spent all this capex on increasing production and building out new lines for Nvidia. They paid per item, bought them, sold them, and made shitloads of money. Then, when it all fell apart, they just had to write down a quarter's worth of inventory, whatever.

Speaker 1

Everyone else was like, "Well, crap, I have all these empty production lines," right?

Speaker 2

But what did AMD do then? Its chips were actually better for crypto mining, right? In terms of the amount of silicon cost versus how much you hash. But AMD just didn't. AMD was like, "We're not really going to raise production," right? As a reasonable thing. It wasn't a matter of striking while the iron was hot.

Dylan Patel

Nvidia has done the same thing in recent times. They've ordered capacity that no one believes, multiple times. They see the demand, obviously, but in many cases their number for Microsoft was higher than Microsoft's internal planning. Microsoft's internal planning went up, but Nvidia's number for Microsoft was way higher. It was like, "We just don't think Microsoft is going to need this much, even though they tell us this." Who the heck is like, "No, no, no, customer, you're going to buy more"?

When the orders come through the supply chain, it's like, "I have to put, 'Pay NCNR'—non-cancelable, non-returnable."

Speaker 1

I asked a question in Taiwan once. Colette, the CFO, and Jensen, the CEO, were both there. It was a room full of mostly finance bros, and they were asking stupid finance questions 3 days before earnings, so obviously they couldn't answer anything because of SEC regulations.

My question to them was, "Look, Jensen, you're so vibes-driven, gut-feel-driven, and visionary. Colette's a CFO. She's amazing in her own right, but those personalities clash. How do you work together?"

He's like, "I hate spreadsheets. I don't look at them. I just know." That was his response. Of course, the best innovators in the world have really good gut instinct, right? So the gut instinct to order—

Speaker 2

With non-cancelable orders, when you don't know, they've had to write down over their history multiple times—many, many billions of dollars in cumulative orders, whether it be the H20, which is more regulatory, or other cases where they've ordered and had to cancel.

Speaker 1

Is that many billions?

Dylan Patel

It's many billions.

Speaker 2

Peanuts.

Dylan Patel

Well, it depends. The crypto write-down was multiple billion when their stock was less than $100 billion. It's peanuts compared to the upside, right?

Speaker 1

I think everything Nvidia did was right, and everything AMD did was wrong in that scenario. But it is crazy, especially in a cyclical industry like semiconductors, where companies go bankrupt all the time, which is why we have all this consolidation.

Dylan Patel

These bets were totally worth taking. Yes.

Speaker 2

If you look at it from a risk-return perspective, these bets were totally worth taking. If you look at it from the perspective of, "I'm a CEO and I want to have predictable quarters for Wall Street," it's a very different story. I think that's where part of the tension is now.

Speaker 1

We made one of Jensen recently and put it on social media—Instagram, TikTok, XHS, Redbook, and Twitter, of course. I really liked it because he's like, "The goal of playing is to win, and the reason you win is so you can play again." He compared it to pinball, where you just play all day and keep getting more rounds. His whole thing is, "I want to win so I can play the next game."

Dylan Patel

It's only about the next generation. It's only about now and the next generation. It's not about 15 years from now, because it's a whole new playing field every time—or 5 years from now. I think the risk-reward is correct.

Speaker 2

But few people take these kinds of risks. Nvidia is the only semiconductor company worth north of $10 billion that was founded as late as it was.

Dylan Patel

MediaTek was founded in the early 1990s, and Nvidia and everyone else are mostly from the 1970s.

Speaker 1

The big ones.

Speaker 2

I think you raised this great point about betting the farm, and he's actually been wrong a couple of times, to your point.

Speaker 1

Mobile, right? What the hell happened with mobile?

Dylan Patel

Exactly. He still takes those risks. Mark actually had this great conversation with Eric where he talked about being founder-run, where you have this memory of the risks you took to get to where you are today.

In a lot of cases, if you're a CEO brought on later, you're sort of like, "Okay, continue to steer the ship as is." But in this case, Jensen remembers all the times they almost went belly-up, and he's like, "I've got to keep making bets like that."

How do you think he's changed? He's been one of the longest-running CEOs—over 30 years. He's kind of right up there with Larry Ellison now.

Speaker 1

I mean, obviously, I'm 29. I don't freaking know what he was like.

Speaker 2

I've watched a lot of old interviews.

Dylan Patel

I won't say he wasn't—

Speaker 1

Longer than you've been alive.

Dylan Patel

Yeah, exactly. Nvidia was founded before I was born. I'm '96, right? Maybe anything over the last couple of years is more relevant.

Speaker 1

No, no, probably better. I think even watching old interviews, right? I watched a lot of old interviews and a lot of old presentations he's given.

Dylan Patel

One thing is that he's just sauced up and dripped up. The charisma he's gotten has only gotten stronger.

Erik Torenberg

Right?

Dylan Patel

Which is an interesting point. I don't know if it's quite relevant.

Erik Torenberg

Totally agree with that.

Dylan Patel

But the man has learned to be more of a rock star, even though he was always charismatic. He's a complete rock star now, and he was a rock star a decade ago, too. It's just that people maybe didn't recognize it.

I think the first live presentation of his that I watched was at—what's the conference? CES, like 2014 or 2015 or whatever. He was talking only about AI. He was telling all these gamers about AlexNet and self-driving cars. It's like, know your audience, first of all, but also, it had nothing to do with consumer electronics. It was gaming, you know. At the time, I was a teenager moderating gaming and gaming-hardware subreddits.

At the time, I was also half like, “Holy crap, this is amazing.” But I was also half like, “I want you to announce a new gaming GPU, right?” On the forums, everyone was quickly like, “Screw this. I want to hear about the gaming GPUs.” Nvidia's price gouging.

Of course, Nvidia has always had the attitude that it prices for the value, plus a little bit, because it's just smart enough to know. I'm guessing Jensen just has the gut feel for how to price things.

Erik Torenberg

He'll change the price, at least on gaming launches, right up until right before the presentation. Wow.

Dylan Patel

So it really is a gut-feel thing, probably. Anyway, he had that charisma to know what was right, but I think a lot of people were like, “Oh, no, whatever. Jensen's wrong. He doesn't know what he's talking about.”

But now, when he talks, people are like, “Oh, very, very...” It might just be that he's been right enough.

Erik Torenberg

Yeah. There's a post on X recently that said he had moved up into god mode with a select group of CEOs.

Dylan Patel

Who's god? Who are the other gods?

Erik Torenberg

It was Zuck. Who are the other gods?

Dylan Patel

Elon.

Erik Torenberg

Elon, Zuck, and Jensen.

Dylan Patel

Nice. Nice. Okay.

Erik Torenberg

A cool crew to be in.

Dylan Patel

So we pray to Silicon Valley. It's sort of a cult now, is it?

Erik Torenberg

Exactly. Just on one last thing about people: you mentioned Colette Kress, his CFO. There's a famously loyal crew at Nvidia, even though all of the OGs could retire at this point. Is there anyone akin to Gwynne Shotwell at SpaceX, or previously Tim Cook to Steve Jobs at Apple, who is at Nvidia today?

Dylan Patel

I mean, he had 2 cofounders, right? Let's not overlook that. 1 of them isn't involved and hasn't been for a long time, but the other one was involved until just a few years ago. So it's not just Jensen running the show, right?

Erik Torenberg

Totally.

Dylan Patel

Although he was running the show, there are quite a few people on the hardware side. There's someone at Nvidia who's mythical to me. When you talk to the engineering teams, he leads a lot of the engineering teams. He's a private person, so I don't want to say his name, actually.

Erik Torenberg

Fair enough.

Dylan Patel

But he's effectively the chief engineering officer. That's his role, and people within his organization know who he is. I think there are people like that. He's intensely loyal, and there are a number of these types of people.

There's another fellow who's associated with all these innovative ideas at Nvidia, but he's the guy who literally says, “We need to get this silicon out now. We're cutting features.” That's what he's famously known for, and all the technologists at Nvidia hate him. This is a second guy, also intensely loyal to Nvidia, who's been around for a long time.

When you have such a visionary, forward-looking company, 1 problem is that you get lost in the sauce. You think, “I want to make this. It's got to be perfect and amazing.” You have to have people who say, “Screw it. Cut it. We'll put it in the next 1. Ship now. Ship faster.” That's really hard to do in a space like silicon.

The thing about Nvidia that's always been super impressive, going back to the beginning, is their first successful chip. They were going to run out of money, and Jensen had to get money from other people just to finish the development. Even then, he had barely enough money because Nvidia had already had a failed chip before that.

When the chip came back, it had to work. Otherwise, Nvidia wouldn't make it. They could only pay for what is called a mask set. Basically, you put these stencils into the lithography tool, and that tells it where the patterns are. You put the stencil in, deposit material, etch material away, and repeat the process.

You stack dozens of layers on top of each other to make a chip. These stencils are custom to each chip, and today they cost tens and tens of billions of dollars. Even back then, it was still a lot of money. They could only pay for 1 set.

The typical thing with semiconductor manufacturing is that, no matter how well you simulate or verify everything, you'll send a design in and have to change it. There's always going to be something. It's so hard to simulate everything perfectly. The thing about Nvidia is that they tend to get it right the first time.

Erik Torenberg

Yeah.

Dylan Patel

Even great-executing companies like AMD or Broadcom often have to ship revisions. They're denoted as A followed by a number or B followed by a number. Nvidia almost always ships A0. They sometimes ship A1.

The letter is basically the transistor layer, and the number is the wiring that connects all the transistors together. Nvidia will start production of the A layer and ramp it really high, then hold it right before transitioning to the metal layers, just in case it needs to change them.

The moment Nvidia confirms that it works, it can blast through a lot of production. Everyone else is like, “Let's get the chip back. Okay, A0 doesn't work. We have to make this tweak, make that tweak, and get the chip back.” It's called a stepping.

At Intel, we were very jealous of Nvidia at that time. Nvidia consistently delivered on the 1st revision. We did not. In the data center CPU group, there was 1 product where Intel got to E2.

Erik Torenberg

E2 is like a 15th revision. This is—

Dylan Patel

This was around the peak of AMD's market share, when AMD skyrocketed in market share versus Intel. Intel was at E2—15 steppings.

Erik Torenberg

Because it's quarters of delay, right? I mean, it's catastrophic for a go-to-market.

Dylan Patel

Yeah. Each 1 is a quarter of delay or something like that. It's absurd. I think that's the other thing about Nvidia: “Screw it, let's ship it. Let's get the volume as soon as possible.”

Nvidia has some of the best simulation and verification, which lets it go from design—or from idea—to shipment as fast as possible. It cuts out unnecessary features that could delay the product and makes sure it doesn't have to do revisions, so it can respond to the market as quickly as possible.

There's a story about Volta, which was the first Nvidia chip with tensor cores. Nvidia saw all the AI work on the prior-generation P100 Pascal and decided to go all in on AI. It added the tensor cores to Volta only a handful of months before sending it to the fab. It said, “Screw it. Let's change it.” It's crazy.

If Nvidia hadn't done that, who knows? Maybe someone else would have taken the AI chip market. There are all these times when Nvidia makes major changes, but there are often minor things that have to be tweaked, too, like number formats or some architectural detail.

Nvidia is just so fast. The other crazy thing is that it has a software division that can keep up with that. If you come out with a chip and basically require no stepping, it's immediately in the market. Being ready with the drivers and all the infrastructure on top of that is just super impressive.

Guido Appenzeller

Yeah, I love that point because you think of Nvidia benefiting from tailwind after tailwind, but I think both of you are saying you've got to move fast enough and execute well enough to take advantage of those tailwinds.

Erik Torenberg

And if you think about it—and by the way, I loved your CES story. I'm just envisioning him more than 10 years ago talking about self-driving cars. But if you think about nailing the video game tailwind, VR, Bitcoin mining, and obviously AI now, one of the things that Jensen talks about today is robotics and AI factories.

Maybe my last question on Nvidia: What do you think about the next 10 to 15 years? I know calling beyond 5 is hard, but what does Nvidia's business look like? It's really a question of—and this is, I think, every time I've talked to some executives at Nvidia, I've asked this question because I really want to know, and they won't answer it, obviously—but what are you going to do with your balance sheet? You are the highest-free-cash-flow company; you have so much cash flow.

Sarah Wang

Now the hyperscalers are all taking their cash flow way down, right, because they're spending on GPUs. What are you going to do with all this cash flow? Even before this whole takeoff, he wasn't allowed to buy ARM, right? So what can he do?

Dylan Patel

With all this capital and all this cash, right? Even this $5 billion investment in Intel has regulatory scrutiny there. It's in the announcement: “Yeah, this is subject to review,” right?

Erik Torenberg

Yeah.

Dylan Patel

You know, I imagine that'll get passed, but he can't buy anything big. He's going to have hundreds of billions of dollars of cash on his balance sheet. What do you do? Is it starting to build AI infrastructure and data centers? Maybe. But why would you do that if you can just get other people to do it and take the cash?

Sarah Wang

Well, he's investing those, right?

Dylan Patel

Investing peanuts, right?

Guido Appenzeller

Right. You know, he recently gave CoreWeave a backstop because today it's really hard to find a large number of GPUs for burst capacity. Like, “Hey, I want to train a model for 3 months. I have my base capacity for all my experiments, but I want to train a big model for 3 months, and then I'm done.”

Sarah Wang

We know from our portfolio, yeah.

Dylan Patel

Yeah, so Nvidia sees this issue. They think it's a real problem with startups; it's why the labs have such an advantage. But what if I could—right now, most companies in the Valley spend, what, 75% of their round on GPUs, right? Or at least—yeah, with [inaudible].

Guido Appenzeller

What if you could do 75% in 3 months on 1 model run? You know, really scale and have some sort of competitive product, and then you have the model. Then you raise more capital or start deploying. What do you do with it? Is it—

Erik Torenberg

Start buying a crapload of humanoid robots and deploying them? But they don't really make good software. They don't make really that amazing software for them. In terms of the models, they make the layer below, which is great. Where they deploy their capital is the question.

Sarah Wang

He has been investing up and down the supply chain a little bit, though, right? Investing in the neoclouds, investing in some of the model-training companies.

Dylan Patel

Yeah, but again, it's small fries. He could have done the entire Anthropic round if he wanted to. Of course, he didn't. He could have done the entire OpenAI round, or the entire xAI round.

Erik Torenberg

Do you think these are things he should be doing?

Sarah Wang

Yeah, good question.

Dylan Patel

I don't know, right? I think picking winners is obviously really tough for him because he has customers all across this ecosystem. And if he starts picking winners, then his customers will be even more anxious to leave and give even more effort to whether it's AMD or some startup or their internal efforts, et cetera, et cetera, right? Buying TPUs, whatever it is. He can't just invest in these—

Erik Torenberg

We'll quote you for the next round that we're raising, but anyways—

Dylan Patel

He could make venture a dead industry, take all of the best rounds.

Sarah Wang

Put a lot of business in it. Yeah.

Erik Torenberg

You know, you could do the seeds and then have Jensen mark you up.

Guido Appenzeller

I don't like it.

Erik Torenberg

Is he also reshaping his market? I mean, look, a couple of years ago there were 4 big purchasers of these cards. You just listed 6. To what extent is that a strategy?

Sarah Wang

Him and Nebius—there's a long list there, of course. Yeah.

Erik Torenberg

Is there a strategy?

Dylan Patel

It is. I think it absolutely is. But he didn't have to put much capital down to do this.

No, but if you look at the grand amount of capital that he spent investing in the neoclouds, it's a few billion, but he has a lot of other levers if he wants to.

Erik Torenberg

Right, right. Allocations, as you mentioned. What's nice is that historically you gave volume discounts to hyperscalers, but because he can use the argument of antitrust, he's like, “Everyone gets the same price.” So what should he do with the cash, or what should guide his decisions?

Dylan Patel

I mean, I think there is an argument that he should invest in data centers—and only the data-center layer, not what goes in the data centers—so that more people build data centers. If market demand continues to grow, data centers and power are not the issue. Invest in data centers and power.

I've said that to them: They should invest in data centers and power, not in the cloud layer, because the cloud layer is not commoditized, but it's quite a commoditized complement, right? That's the whole phrase. And I won't say being a cloud is commoditized, but you certainly have a lot of competitors who are decent now. You've educated commercial real estate and other infrastructure investment firms into going into AI infrastructure as well. So I don't think it's the cloud layer that you invest in.

Do you invest in data centers and energy? Yes. Do you invest in it because that's the bottleneck for your growth, really? A, how much people want to spend and can spend; and B, the ability to actually put them in data centers.

And then robotics. I think there are areas he could invest in, but nothing requires $300 billion of capital. So what do you do with the capital? I really don't know, and I feel like Jensen has to have some idea. There's some visionary plan here because that's what shapes the company.

I mentioned $200 million of free cash flow, $250 billion of free cash flow a year. What do they do with it? Do they just buy back stock forever? Do they go the Apple route? The reason Apple hasn't done anything interesting in nearly a decade is they've got a non-visionary at the head. Tim Cook is great at supply chain, and they're just plowing the money into buybacks. They're not really doing anything in automotive; the self-driving-car thing failed. We'll see what happens with AR/VR. We'll see what happens with wearables, but Meta and OpenAI might be even better than them. We'll see, and others, right?

So what does he invest in? I have no clue. But nothing that requires so much capital and actually gets a return is the tough question.

Sarah Wang

Because the easy thing is my cost of equity, right? I just buy back stock—

Guido Appenzeller

And it doesn't completely change the company culture. I think that's another thing, right? There are probably areas you could invest in, but you suddenly end up with the company doing 2 completely different things, which are very difficult to keep aligned.

Dylan Patel

But they do 10 completely different things, right? I mean, one way to look at it is: We build AI infrastructure, and humanoids around the world are AI infrastructure. Data centers and energy are AI infrastructure.

Sarah Wang

So the humanoids would totally work, right? If you suddenly start pouring concrete and building power plants, that's a completely different culture, a completely different set of people, and it gets much, much harder.

Erik Torenberg

There are all these different areas where he could use capital to allow something to happen, right? Not necessarily owning it himself.

Guido Appenzeller

And look, bear in mind, at Intel, one of the biggest problems we had was that our customer base sucked, right? Most of the chips we sold went into the large hyperscalers. They were way too concentrated, and they build their own chips, so they can push down your prices. So, honestly, spending it on diversifying the customer base…

Erik Torenberg

In 2014, you guys should have just charged so much that your margins were 80%. What would the world have done?

Guido Appenzeller

Nothing. The margins were pretty good back then. That wasn't the problem. That wasn't the primary problem. They were 60, 65.

Erik Torenberg

They were 80 still.

Guido Appenzeller

Yeah.

Erik Torenberg

Oh boy. Jensen's PTSD is kicking in here.

Well, wait. I think Guido's comment is actually a really good segue into something else we wanted to talk to you about, which is the hyperscalers. One of the reasons that I love reading SemiAnalysis is that you guys make these out-of-consensus calls that you're often right about.

Dylan Patel

Only often. But you have a Jensen hit rate that's very high.

Erik Torenberg

Where's my billion-dollar, PV-positive bet?

The one that caught my eye was Amazon's AI resurgence. I wanted to talk to you a little bit about that because I think we found it pretty interesting being on the ground and helping our portfolio companies pick who their partners are. We have some microdata on this, so can you walk us through why they are behind?

Dylan Patel

Yeah. So, in Q1 2023, I wrote an article called “Amazon's Cloud Crisis.” It was about how all these neoclouds were going to commoditize Amazon. It was about how Amazon's entire infrastructure was really good for the last era of computing: what they do with their Elastic Network Adapter, ENA, and EFA, their NICs, and the whole protocol and everything behind them; what they do for custom CPUs, et cetera. It was really good for the last era of scale-out computing, and not this era of scale-up AI infrastructure.

It was also about how neoclouds were going to commoditize them, and how their silicon teams were focused on cost optimization, whereas the name of the game today is maximum performance per cost. Often that means you drive performance up like crazy, even if the cost doubles or triples, because then the cost per performance still falls. That's the name of the game today with NVIDIA's hardware.

It ended up being a really good call. Everyone was calling us out, saying, “No, you're wrong,” because Amazon was the best stock at the time, Microsoft really hadn't started taking off yet, and neither had Oracle or all these other companies. Since then, Amazon has been the worst-performing hyperscaler.

The call here is that they still have structural issues. They still use Elastic Fabric Adapter, although that's getting better. They're still behind NVIDIA's networking and behind Broadcom and Arista-type networking NICs. Their internal AI chip is okay, but the main thing is that they're now waking up and are able to actually capture business.

The main call here is that, since that report, AWS revenue growth has been decelerating consistently, and our big call is that it's actually going to start reaccelerating. That's because of Anthropic and because of all the work we do on data centers—tracking every single data center, when it goes online, and what's in there.

If you know how much the chips, networking, and power cost, and you know generally what the margins are for these things, then you can start estimating revenue. When we build all that up, it's very clear to us that AWS revenue growth troughs this quarter. This is the lowest AWS revenue growth will be on a year-over-year basis for at least the next year, and it's reaccelerating to north of 20% again.

That's because of all these massive data centers they have online with Trainium and GPUs. It depends on which one, depending on the customer. The experience isn't as good as, say, CoreWeave or whatever, but the name of the game is capacity today. CoreWeave can only deploy so much. They can only get so much data center capacity, and they're really fast at building.

The company with the most data center capacity in the world—and still today, although they may get passed up in the next 2 years—is actually Amazon. Based on what we see, they will get passed up, but incrementally, Amazon still has the most spare data center capacity that's going to ramp into AI revenue over the next year.

Erik Torenberg

Is that the right type of data center capacity? For the high-density AI buildouts today, you need massively more cooling. You need enough water close by, and you need enough power close by. Is it in the right place, or is it the wrong type?

Dylan Patel

Data center capacity, in this sense, means everything from power being secured, to substations being built, to transformers, to being able to provide power whips to the racks.

Historically, Amazon has had the highest-density data centers in the world. They went to 40-kilowatt racks when everyone else was still at 12. If you've ever stepped foot inside most data centers, they're pretty cool and dry-ish. If you step inside an Amazon data center, they feel like a swamp. It feels like where I grew up. It's humid and hot.

Because they're optimizing every percentage point, your point here is that Amazon's data centers aren't equipped for the new type of infrastructure. But when you compare that to the cost of the GPU, having a complex cooling arrangement is fine.

We made a call on Astera Labs a couple of months ago, when it was at 90, and it went to 250 the month after because of the orders Amazon is placing with them. There are certain things with Amazon's infrastructure—I won't get too much into it—but its rack infrastructure requires it to use a lot more Astera Labs connectivity products. The same applies to cooling. On the networking and cooling side, they just have to use a lot more of this stuff. But again, this stuff is inconsequential in cost compared to the GPU.

Erik Torenberg

You can build, right? My question was more like, look, I may need a major river close by for cooling at this point. In many areas, I just can't get enough water. It's probably power in the same region, too.

Dylan Patel

There are 2 gigawatt-scale sites where they have all the power secured. Wet chillers and dry chillers—all secured. Everything's fine. It's just not as efficient, but that's fine, right? They're going to ramp the revenue. They're going to add the revenue.

Not that I necessarily think Amazon's internal models are going to be great, or that their internal chip is better than NVIDIA's or competitive with TPUs, or that their hardware architecture is the best. I don't necessarily think that's the case. But they could build a lot of data centers and fill them up with stuff that will be rented out. It's a pretty simple thesis.

Erik Torenberg

How important has Anthropic been to the co-design for Trainium? I remember we had a portfolio company—this was summer 2023—and they invited them to AWS. They spent, I think, 8 hours with them over the course of a week trying to figure out Trainium back then. It was just impossible to work through. Obviously, that portfolio company hasn't gone back and tried it now, but how different is it based on what you're hearing?

Dylan Patel

Oh, it's still bad.

Erik Torenberg

Okay. Got it.

Dylan Patel

It's tough to use. This is sort of the argument that every inference company offers, including the AI hardware startups:

Erik Torenberg

Because I'm only running 3 different models at most, I can hand-optimize everything, write kernels for everything, and even go down to an assembly level, right? How hard can it be?

Dylan Patel

Yeah, it is pretty hard. But you tend to do this for production inference anyway. You aren't using cuDNN, which is NVIDIA's library that's super easy to use. You're still not using these ease-of-use libraries. When you're running inference, you're either using CUTLASS, stamping out your own PTX, or, in some cases, people are even going down to the SASS level.

When you look at, say, OpenAI or Anthropic, when they run inference on GPUs, they're doing this. The ecosystem isn't that amazing once you get all the way down to that level. It's not like using NVIDIA GPUs is easy. You have an intuitive understanding of the hardware architecture because you work on it so much, and everyone's worked on it and you can talk to other people, but at the end of the day, it's not easy.

Whereas, on Anthropic's Trainium or on TPUs, the hardware architecture is a little bit simpler than a GPU.

They have larger, simpler cores rather than all this functionality. They’re less general, so they’re a little easier to code on. There are tweets from Anthropic people saying that, when they’re doing that low-level work, they actually prefer working on Trainium and TPUs because of the simplicity.

Erik Torenberg

Interesting.

Dylan Patel

To be clear, TPUs—and Trainium especially—are very hard to use, not for the faint of heart. It’s very difficult, but you can do it if you’re just running—if I’m Anthropic and I must only run Claude 4.1 Opus or Sonnet, and screw it, I won’t even run Haiku. I’ll just run Haiku on GPUs or whatever, right? I’m just going to run 2 models.

Actually, screw it, I’m just going to run Opus on GPUs too and Sonnet on TPUs. Sonnet is the majority of my traffic anyway. I could spend the time. How often am I changing that architecture—every 4 or 6 months?

Erik Torenberg

Right. How much?

Dylan Patel

It’s not even changing that much, honestly, right?

Erik Torenberg

I think from 3 to 4 definitely did change, right?

Dylan Patel

Yeah. I mean, define architectural change. At a high level, the primitives are more or less the same across the last couple of generations.

Erik Torenberg

I don’t know enough about Anthropic’s model architecture, to be honest. But I think, from what I’ve seen at other places, there have been enough changes that it takes time to program this. The main thing is, if I’m Anthropic and I have, what, $7 billion ARR now, or whatever—north of $10 billion by the end of next year, north of $20 billion, right? ARR is maybe even $30 billion—and my margins are 50%, 70%, that’s $15 billion of Trainium that I need, right, that I can run Sonnet on.

Most of that is going to be Sonnet 3.5—or, sorry, 4.5, whatever it is, right? It’s going to be 1 model serving most of the use cases. So I could spend the time, and it’ll work on this hardware.

Yeah, totally. Maybe on the topic of non-consensus calls you’ve made, I’ll move to another cloud. In June, you guys said that Oracle is winning the AI compute market. In this pod, we’ve already referenced the big jump that Oracle had. I think it was the single largest gain that a company with over $500 billion in market cap has ever had.

Was it Nvidia in Q1 2023? Wasn’t that bigger? It might have been smaller. Okay. I think it was maybe close. We’ll fact-check ourselves. That’s amazing. But, obviously, this is the massive commitment that was announced. Can you walk us through why you made that call, and why Oracle is poised to do so well in such a competitive space?

Dylan Patel

Yeah, so Oracle has the largest balance sheet in the industry that isn’t dogmatic to any type of hardware. They’re not dogmatic to any type of networking. They’ll deploy Ethernet with Arista, Ethernet through their own white boxes, and NVIDIA networking—InfiniBand or Spectrum-X. They have really good network engineers. They have really great software across the board.

Again, ClusterMAX—they were ClusterMAX Gold because their software is great. There are a couple of things they needed to add that would take them higher, and they’re adding those to Platinum, which was where CoreWeave was.

When you couple 2 things, OpenAI has insane compute demand. Microsoft is quite pansy. They’re not willing to invest because they don’t believe OpenAI can actually pay the amount of money. I mentioned earlier the $300 billion deal: OpenAI doesn’t have $300 billion, and Oracle’s willing to take the bet.

Of course, there’s a bit more security in the bet in that Oracle really only needs to secure the data-center capacity. That’s how we came across the bet. We’ve been telling our institutional clients, especially in a super-detailed way—whether they be hyperscalers, AI labs, semiconductor companies, or investors—in our data-center model, because we’re tracking every single data center in the world.

Oracle doesn’t build their own data centers either, by the way. They get them from other companies and co-engineer them, but they don’t physically build them themselves. They’re quite nimble in terms of being able to assess and engineer new data centers. So we saw all these different data centers Oracle is snatching up, in deep discussions, signing, et cetera.

We have a gigawatt here, a gigawatt there, a gigawatt there. Abilene, 2 gigawatts. You have all these different sites that they’re signing up for and discussing, and we’re noting them. We have the timeline because we’re tracking the entire supply chain. We’re tracking all the permits and regulatory filings through language models, using satellite photos constantly, and then the supply chain for chillers, transformer equipment, generators, et cetera.

We’re able to make a pretty strong estimate, quarter by quarter, in our data-center model, of how much power there is for each of these sites. Some of these sites that we know of aren’t even ramping until 2027, but we know that Oracle signed them, and we have the ramp path.

Then it’s a question of, let’s say you have 1 megawatt, for simplicity’s sake, which is a ton of power, but now it doesn’t feel like much—we’re in the gigawatt era. If you’re talking about 1 megawatt, you fill it up with GPUs. How much do the GPUs for 1 megawatt cost?

Actually, it’s even simpler to do the math. If I’m talking about a GB200, each individual GPU is 1,000 watts, but when you talk about the whole system, it’s roughly 2,000 watts. All-in, for simplicity’s sake, it’s $50,000 per GPU. The GPU itself doesn’t cost them that; there are all the peripherals. So that’s $50,000 in capex for 2,000 watts, or $25,000 for 1,000 watts.

Then what’s the rental price for a GPU? If you’re on a really long-term deal, volume is $2.70—$2.60 in that range. Then you end up with, oh, it costs like $12 million per megawatt to rent a megawatt.

Erik Torenberg

Yeah.

Dylan Patel

Each chip is different, so we track each chip—what the capex is and what the networking is. We know what each chip is, so you can predict what chips they’re putting in which data centers, when those data centers go online, and how many megawatts by quarter.

Then you end up with, well, Stargate goes online in this time period. They’re going to start renting at this time. It’s this many chips at each Stargate site. Therefore, this is how much OpenAI would have to spend to rent it.

Then you price that out, and we were able to predict Oracle’s revenue with pretty high certainty. We matched pretty dead-on what they announced for 2025, 2026, and 2027, and we were pretty close on 2028. The surprise for us was that they announced some 2028 and 2029 data centers that we haven’t found yet, but we’ll find them, of course.

This methodology lets you see what data centers you’re getting, how much power, what they’re signing, and how much incremental revenue that is when it comes online. That’s the basis of our Oracle bet.

Obviously, in the newsletter we included a lot less detail, but it was that thesis: they have all this capacity, and they’re going to sign these deals. In our newsletter, we talked about 2 main things: the OpenAI business and the ByteDance business.

Presumably, tomorrow—on Friday—there’s going to be an announcement about TikTok and all this. But the ByteDance business involves huge amounts of data-center capacity that Oracle is also going to lease out to ByteDance. We did the same methodology there.

With ByteDance, it’s pretty certain they’ll pay because they’re a profitable company. With OpenAI, it’s not. There have to be some error bars as you go further out in terms of whether OpenAI will exist in 2028, 2029, or 2030, and whether they’ll be able to pay the $80-plus billion a year they’ve signed up to Oracle for. That’s the only risk here.

If that happens, Oracle’s downside is also somewhat protected because they only sign the data center, which is a minority of the cost. The GPUs are everything, and they purchase those 1–2 quarters before they start renting them. The downside risk is pretty low for them. If they don’t get the deal, they don’t get the revenue, but it’s not like they’re stuck with a bunch of assets they bought that are worthless.

Erik Torenberg

Yeah. Yeah. Is there another angle here? I mean, OpenAI and Microsoft were BFFs, and now they’ve filed divorce papers. They just want to diversify, and that’s pushing them toward other providers.

Dylan Patel

Yeah. Microsoft was the exclusive compute provider. It got reorged to a right of first refusal.

Erik Torenberg

Is it not your last choice or something like that?

Dylan Patel

No, it’s still a right of first refusal. Microsoft and Oracle—those 2 are not mutually exclusive.

Erik Torenberg

Well, if OpenAI is like, “We’re going to sign an $80 billion contract or a $300 billion contract for the next 5 years. Do you guys want it?” and they’re like—

Dylan Patel

“No, what? Okay, cool.” Right? OpenAI needs someone with a balance sheet to actually be able to pay for it. They’ll make tons of money off OpenAI on the margins on the compute and the infrastructure and all these things, but someone’s got to have a balance sheet, and OpenAI doesn’t have one. Oracle does.

Although, given the scale of what they signed, we also had another source of information: they were talking to the debt markets, right? Oracle actually just needs to raise debt to pay for this many GPUs over time. They won’t do it immediately; they can pay for everything this year and next year from their own cash, but in 2027, 2028, and 2029, they’ll start to have to use debt to pay for these GPUs. That’s what CoreWeave has done, and most of the neoclouds are debt-financed.

Even Meta went and got debt for its Louisiana megadatacenter. It’s literally better on a financial basis to do buybacks with your cash and get debt because the debt is cheaper than the return on your stock. It’s a financial engineering thing. Who’s out there, right? It could be Amazon, Google, Microsoft—a very short list.

Erik Torenberg

Or it could be Oracle or Meta, right? Meta’s obviously not. Microsoft has chickened out. Amazon, Google, and Oracle—that’s all that’s left.

Dylan Patel

Google would be an awkward fit. So—

Erik Torenberg

Yeah, Google would be an awkward fit. Amazon would be a fine fit, but—

Dylan Patel

Exactly. Right. It’s like—

Erik Torenberg

Yeah. Well, I guess maybe on the topic of these giant data center buildouts, you guys just released a piece on xAI and Colossus 2. Are you getting less impressed by these feats of building something this massive in 6 months, or is it still very impressive to you guys?

You know, this is the thing I’ve said about AI researchers: they’re the first class of humans to think about things on an order-of-magnitude scale, whereas people have always thought about things in terms of percentage growth. Ever since industrialization—and before that—it was just absolute numbers. Humanity is evolving in how we think because things are changing faster. Everything is an upscale.

Dylan Patel

And so it was really impressive when GPT-2 was trained on so many chips, and then GPT-4 was trained on 20,000 H100s. It was like, “Holy crap.” Then it was the era of 100,000-GPU clusters. We did some reports around 100,000-GPU clusters, but now there are 10 100,000-GPU clusters in the world. I was like, “Okay, this is kind of boring.”

But 100,000 GPUs is over 100 megawatts. Now, literally, in our Slack and some of these channels, it’s like, “Oh, we found another 200-megawatt data center.” There’s someone who puts the yawning emoji every time, and I’m like, “Dude, what?” Now it’s only exciting if you do gigawatt scale.

Erik Torenberg

Gigawatt era. Yeah.

Dylan Patel

Yeah. And I’m sure maybe we’ll start yawning at that, too. But the log scale of this is like—

Erik Torenberg

The capital numbers are crazy, right? It was crazy enough that OpenAI did a $100 million training run. Then they did a $1 billion training run; now we’re talking about $10 billion training runs. It’s crazy that we think in log scale, but yes, things are only impressive—

Dylan Patel

Yeah, when they do it like what Elon’s doing. What Elon’s doing in Tennessee, in Memphis, the first time was crazy, right? 100,000 GPUs in 6 months. He bought a factory in February 2024 and had models training within 6 months.

He did liquid cooling—the first large-scale data center at this scale for AI using liquid cooling—all these crazy firsts. He put generators outside, CAT turbines, and all these things to get the power; mobile substations; all these different crazy things. He tapped the natural gas line running alongside the factory. All of these—

Erik Torenberg

You know, 200, 300 megawatts, right? Now he’s doing it at a gigawatt scale, and he’s doing it just as fast. You would think this is obviously way more impressive that he did it again.

Dylan Patel

Yeah.

Erik Torenberg

But—

Dylan Patel

Like—

Erik Torenberg

Maybe I’m desensitized, but it’s like you’ve given the child too much candy, right?

Dylan Patel

Exactly.

Erik Torenberg

And now the child has no—it’s like he doesn’t like apples, right? I don’t know.

Dylan Patel

So, yeah, a gigawatt data center. There were all these protests around his Memphis facility. People were saying, “Oh, you’re destroying the air.” And it’s like, “Have you looked around that area of Memphis?” There’s a gigawatt gas-turbine plant that’s just powering that area generally. There’s a sewage plant servicing the entire city of Memphis. There are open-air pits—there’s open-air mining. There’s all sorts of disgusting shit around there, which is needed. We need that stuff for a country to run, to be clear. People were complaining about a couple hundred megawatts of power—

Erik Torenberg

Of generation. So he got protests from all sorts of people. He got super into the political side of things, and the NAACP even protested him. He really got some local municipalities to be like, “Oh, I don’t like this.”

Dylan Patel

And so he couldn’t do as much as he wanted to in Memphis. But he still needed the data center to be close because he wanted to connect these data centers with super-high bandwidth, super close. He already had a lot of infrastructure set up there, so he bought another distribution center at this time. It’s still in Memphis, but the cool thing about Memphis is that it’s right across the border from Mississippi. Right. So now—

Erik Torenberg

You know, it’s 10 miles away from his original one, but his facility is a mile away from Mississippi, and he bought a power plant in Mississippi. He’s putting turbines there because the regulation is completely different. If the question is really to galvanize resources and build it really fast, maybe Elon is ahead of everyone. He hasn’t made the best model yet, or he doesn’t have the best model, at least today. You could argue Grok 4 was the best for a little period of time, but it’s truly amazing how fast he’s able to build these things.

From first principles, most people are like, “Shit, we can’t build the power. We can’t do power here anymore. I guess we have to find a new site.” It’s like, “No, just go across the border.”

Dylan Patel

Go to Mississippi.

Erik Torenberg

My favorite thing is that Arkansas is right there, so Mississippi gets mad. I don’t know—the regulation—all future data centers built in places where multiple states meet, is that the—

Dylan Patel

Four Corners, yeah.

Erik Torenberg

The optimal regulation, I think. There’s one. There we go. Is there a point in the U.S. with 5? I know there’s a point with 4—4 states intersect. Yeah. Maybe that’s going to be a data center kind of concern. All right.

Dylan Patel

I’m going to buy real estate in that area, right?

Erik Torenberg

Well, I guess on the topic of just maybe new hardware, you had this piece analyzing TCO for GB200s. I’m kind of going to ask this question on behalf of our portfolio companies, which it sounds like you’re helping them already. One of the findings that I thought was really interesting was that TCO was sort of 1.66x H100s for GB200s. Obviously, there’s this point where that’s the benchmark for the performance boost that you’re going to need to at least make the performance-cost ratio benefit from switching over.

Maybe just talk about what you’ve seen from a performance standpoint, and what do you recommend to portfolio companies, maybe on a smaller scale than xAI, who are thinking about new hardware? Try to get it—there are capacity constraints, obviously.

Dylan Patel

Yeah, that’s a challenge, right? With each generation of GPU, it gets so much faster that you want the new one. In some metrics, you could say GB200 is 3 times faster—or 2 times faster—than the prior generation. In other metrics, you can say it’s way more than that. If you’re doing pretraining versus inference, right—

Erik Torenberg

You can run everything for a bit, right?

Dylan Patel

Yeah, if you can run it for a bit, or just inference, and take advantage of the huge NVLink—NVL72—there are ways you can squint and say GB200 is only 2 times faster than H100. In which case, with 1.66x TCO, it’s—

Erik Torenberg

You know, it’s worthwhile, right? It’s worth going to the next generation—

Dylan Patel

But more marginal.

Speaker 1

It’s more marginal. It’s not a big deal. Then there are other cases where, if you’re running DeepSeek inference, the performance difference per GPU is north of 6–7x, and it continues to optimize for DeepSeek inference. Then it’s like, “I’m only paying 60% more for 6x,” so it’s a 4x or 3x performance-per-dollar gain. Absolutely, right? If you’re running inference of DeepSeek, that can also include RL, right?

Then there’s the question of the GPU being new. There’s also B200, GB200, and B2000. B200 is much simpler from a hardware perspective; it’s just 8 GPUs in a box. So it’s not as much of a performance gain, especially in inference, but you have all this stability. It’s an 8-GPU box; it’s not going to be unreliable. The GB200s are still having reliability challenges, though those are being worked through and getting better by the day. When you have an H100 or H200 8-GPU box and one fails, you take the entire server offline and fix it. If it’s GB200 and one GPU fails, what do you do with 72 GPUs? Do you break the whole thing and get a new 72? The blast radius of a failure is huge. GPU failure rates are at best the same and likely worse generation on generation because everything is getting hotter and faster. A lot of people run a high-priority workload on 64 and use the other 8 for low-priority workloads. When a high-priority workload has a failure, instead of taking the whole rack offline, you take some GPUs from the low-priority workload and put them in the high-priority one, then let the dead GPU sit there until you service the rack later. There are all these complicated infrastructure challenges.

Dylan Patel

That 3x or 2x performance increase in pre-training is lower because the downtime is higher, I’m not using all the GPUs all the time, and I’m not able—or I don’t have the infrastructure—to manage low- and high-priority workloads. It’s not impossible; the labs are doing it. It’s just—

Speaker 3

I mean, if I’m running a cloud, it’s actually really hard, because I probably have to rent the spares out as spot instances or something.

Speaker 2

No, no, no, no. Because it’s a coherent domain; it’s NVLink. You don’t want anyone touching that. So the end customer has to leave them as empty spares. That’s even worse.

Speaker 3

The end customer usually will just be like, “I want them, and I’ll use them.” The SLAs and pricing account for that, right?

Speaker 2

Generally, when you have a cloud, you have an SLA. It says uptime is going to be 99%, blah, blah, blah, for this period. With GB200, it’s 99% for 64 GPUs, not 72, and then it’s 95% for all 72. It differs across every cloud; every cloud has a different SLA. But they’ve adjusted for this because they’re like, “Look, this hardware is just finicky. Do you still want it? We will credit you, in that 64 of them will always work.”

Speaker 3

Right, not 72.

Speaker 2

So the end customer has to be capable of dealing with the unreliability.

Speaker 3

The whole reason you want this 72-GPU domain is so you can have some of these gains, right?

Speaker 2

But you have to be smart enough to be able to do it, and that’s challenging for small companies.

Speaker 3

Totally. So NVIDIA just announced the Rubin prefill cards, like CPX—

Speaker 2

CPX. CPX.

Speaker 3

There we go. What’s your take on that? Does it cannibalize?

Speaker 2

Dude, by the way, I don’t know if this is brain rot or what, but I can’t remember what I had for lunch yesterday, yet I know the model number of every fucking chip.

Speaker 3

In your dreams.

Speaker 1

We’re broken. We’re broken. Living the dream.

Speaker 2

No, no, no. Why do you pre-announce a product that’s 5x faster for certain use cases?

Speaker 3

Is that that much?

Speaker 2

I think, historically, AI chips were AI chips. Then we started getting a lot of people saying, “This is a training chip; this is an inference chip.” Actually, training and inference are switching so fast in terms of what they require that now it’s still like one chip.

Speaker 3

Actually, there are still workload-level dynamics that differ.

Speaker 2

The main workload is inference, even in training, because of RL. Most of that is generating stuff in an environment and trying to achieve a reward, so it’s inference still. Training is now becoming mostly dominated by inference as well.

Inference has 2 main operations. There’s calculating the KV cache for prefill: here are all these documents; do the attention between all of them, between all the tokens, however—whatever type of attention you use. Then there’s decode, which is to autoregressively generate each token.

These are very different workloads. Initially, the infrastructure techniques—the ML systems techniques—were, “Okay, I’ll just make the batch size for every single forward pass this big. I’ll make it, let’s call it, 1,000 big, and maybe I’ll run 32 users concurrently. That way, I still have 900-something left—960 left.”

That 960 is actually doing the prefill. If a request comes in, it chunks it. It’s called chunked prefill. You prefill chunks of it, and now you get really good utilization on GPUs. But that ends up impacting the decode workers. The people who are autoregressively generating each token end up having slower TPS, and tokens per second is really important for user experience and all these other things.

Speaker 3

Everyone, everyone, everyone—

Speaker 2

Together AI, Fireworks, all these guys do prefill-decode disaggregated. They run prefill on a set of GPUs and decode on a certain set of GPUs.

Why is this beneficial? Because you can autoscale them. All of a sudden—or not all of a sudden, but over time—if my traffic mix is not long input and short output, but short input and long output, I can have more decode workers. This way, I can guarantee that my prefill time is at a certain level.

What’s really important in search is how fast you get the page to start loading, not when the response is finished. What do people do in games? The loading screen often has some sort of interactive environment, or it blends in over time, or it has tips and tricks—ways to distract you. The same thing is true here. There are studies and papers out there showing that users prefer a faster time to first token, with the first token streamed to them sooner, even if the total time to get all their tokens is a little bit longer.

Guido Appenzeller

I can’t read that fast anyway, right?

Dylan Patel

I mean, I like to give—I like to give—

Guido Appenzeller

Yeah. Most models return above speed-reading speed.

Dylan Patel

But you need that, right? I think the idea is that you want to guarantee time to first token at a certain level for user-experience reasons. Otherwise, people say, “Screw this, I’m not using AI.” Decode speed matters a lot, too, but not as much as time to first token.

By having separate prefill and decode, you can do this. But now you’ve already done this—and this is all in the same infrastructure. What’s the next logical step? These workloads are so different.

For decode, you have to load all the parameters and the KV cache to generate a single token. You batch a couple of users together, but very quickly you run out of memory capacity or memory bandwidth because everyone’s KV cache is different.

Guido Appenzeller

Yeah. The attention of all the tokens, right? Whereas on prefill, I could even just serve 1 or 2 users at a time. If they send me a 64,000-context request, that is a lot of FLOPs, right?

I’ll use Llama 70B because it’s simple to do the math on—70 billion parameters. That’s 140 gigaflops per token. Seventy times 64,000—that’s many, many petaflops. You can use the entire GPU for about a second, potentially, depending on the GPU, just to do the prefill. And that’s just 1 forward pass.

Dylan Patel

So I don't necessarily care about loading all the tokens or all the parameters into the KV cache quickly. All I care about is the FLOPs. That leads us to CPX, but I had to give this long-winded explanation because it's hard for people to understand what CPX is. I've had a lot of—even my own clients—we sent multiple notes explaining it, and they're like, “I still don't understand.” I'm like, “[expletive], okay.” Send them the Attention Is All You Need paper, and—

You can't expect—I mean, think about a networking person. They're like, “I don't know. I don't need to know about this. You know, Attention Is All You Need, right?” Or think about an investor. There are all these data center operators, and they're like, “Oh, there are 2 chips. Why should I build my data center differently?” It's like, “I have to explain everything.” Or just like, “No, you don't have to build differently.”

At Stanford, at least 25% of all students—not CS students, all students—read the paper Attention Is All You Need. That's low. The majors? And you don't like the philosophy? I find this amazing, anyway. Sorry.

Erik Torenberg

The Middle East—I can't remember what country it is—has AI education starting at around 8, and in high school they have to read Attention Is All You Need.

Dylan Patel

Wow.

Erik Torenberg

Someone told me that their son had to read Attention Is All You Need.

Dylan Patel

Which is—I don't know. Look, top-down mandates for education: maybe they work, maybe they don't. Maybe people like homeschooling their kids. I don't know. I went to public school, but—

Erik Torenberg

Back to your question.

Dylan Patel

Yeah. Just on the topic of hardware cycles, I wanted to maybe—I actually explained what CPX is. CPX is a very compute-optimized chip for prefill, whereas decode, statistically speaking, is like the rest: the normal chips with HBM. HBM is more than half the cost of the GPU.

If you strip that out, you end up with a much cheaper chip that's passed on to the customer. Or, if NVIDIA takes the same margin, the cost of this prefill chip is much, much lower. Now the whole process is way cheaper and more efficient, and long context can be adopted.

Erik Torenberg

All right.

Sarah Wang

Yeah. I love that we're actually going into all this detail because I had a more 10,000-foot-view question for you. I haven't been following the semiconductor market as closely as you have. I probably started with the A100. I remember helping Noam at Character.AI in June 2023 chase down GPUs, and the only thing that mattered at that time was the delivery date because there was a huge capacity crunch.

Then to see that evolve over the last 2 years—let's say 6 to 12 months ago, people were doing these RFPs to 20 neoclouds, right? To some degree, the only thing that mattered was price.

Dylan Patel

People actually do RFPs for GPUs?

Sarah Wang

Yes.

Dylan Patel

So, just to be clear, my opinion on how you buy GPUs is that it's like buying cocaine—or any other drug. This was described to me, not by me. I don't buy cocaine.

Erik Torenberg

Someone tells me this. Someone tells me this. I'm like, “Holy [expletive], it's right.” You call up a couple of people, you text a couple of people, and you ask, “Yo, how much do you have? What's the price?”

Dylan Patel

It's like—

Erik Torenberg

Exactly. This is [expletive] like buying drugs. Oh, sorry. Sorry.

Dylan Patel

No, I mean, to this day, it's the same way. You just send—we have Slack Connects with around 30 neoclouds, as well as some of the major ones, and we just send them a message: “Hey, the customer wants this much. This is what they're looking for.” Then they send quotes.

Erik Torenberg

I know this guy.

Dylan Patel

I know a guy.

Sarah Wang

Well, I think that's actually a very accurate description, and I've sent countless people your original ClusterMAX post because I thought it did a really good job breaking them down. Maybe one question to end on for me is: What era are we in now, with Blackwells coming online? Are we sort of back to the summer of 2023 era, where GPUs were tight, or what is your view on where we are?

Dylan Patel

Very good question. For one of your portfolio companies, after their difficulties with Amazon, we tried to actually get you GPUs. The original deals we got you were gone, but here were some other deals. It turned out that multiple major neoclouds had sold out of Hopper capacity, and their Blackwell capacity comes online in a few months.

So it's a bit of a challenge, right?

Sarah Wang

Due to inference?

Dylan Patel

Inference demand has been skyrocketing this year.

Sarah Wang

Reasoning models, yeah.

Dylan Patel

These reasoning models—the revenue has been skyrocketing this year. Also, Blackwell comes online, but it's hard to deploy, so there's a learning curve to deploying it. You could buy Hopper, install it in the data center, and have it running within a month or 2. For Blackwell, it's a longer time frame because of reliability challenges. It's a new GPU; it's just learning curve and growing pains.

There was this gap in how many GPUs were coming onto the market as revenue started to inflect, so a lot of capacity got sucked up. Actually, prices for Hopper bottomed 3 or 4 months ago—or 5 or 6 months ago.

Sarah Wang

Yeah.

Dylan Patel

They've actually crept up a little bit now. They're still not too bad. I don't think we're quite back to the 2023–2024 era of GPUs being tight, but certainly, if you want just a few GPUs, it's easy. If you want a lot, it's hard.

Sarah Wang

You can't get capacity instantly.

Dylan Patel

Yeah.

Erik Torenberg

Wow. What a time. Shall we wrap on that? Dylan, this was another instant classic. Thank you so much for coming to the podcast.

Dylan Patel

It was like 2 hours, bro. What? I missed—thank you. We couldn't stop. Thanks so much. This is great. Thank you so much for having me.

Dylan Patel on the AI Chip Race - NVIDIA, Intel & the US Government vs. China | BidClub