[BidClub_]
Invest Like the Best · · 77 min

Gavin Baker - Watts and Wafers - [Invest Like the Best, EP.473]

Patrick O'ShaughnessyGavin Baker

Podcast
TL;DR
  • Baker treats March’s AI selloff as “pent-up alpha,” because prices fell while Anthropic added $11 billion of ARR in one month and a 500% NDR statistic was shared on the show. He argues that this one-month addition rivaled the combined businesses of Palantir, Snowflake, and Databricks, while DeepSeek had already demonstrated that reasoning models increase inference demand: Asian AWS availability-zone prices doubled, GPU availability fell, and DRAM went vertical. “I’ve just never seen an exponential like this.”
  • Reported revenue may radically understate Anthropic’s compute-constrained economics. Against roughly $50 billion of ARR and a $900 billion valuation, Baker estimates unlimited compute might support $100 billion to $200 billion of revenue, or roughly five times his invented “URR, unconstrained run-rate revenue”; he also believes Anthropic burned perhaps 80% less capital than OpenAI at comparable scale. His expectation that Anthropic could generate cash this year remains a forecast, not a certainty.
  • The terrestrial watts shortage should begin easing in 2027 or 2028, but zoning and approvals may become more binding than turbines or fuel. Longer term, Baker’s answer is “racks in space”: roughly rack-sized satellites with immense solar wings, radiators, and laser links forming virtual data centers, with inference moving orbital while training stays on Earth. Cooling appears solvable to SpaceX engineers; repair remains the honest weakness “until you have probably floating Optimuses.”
  • TSMC’s capacity decisions are Baker’s single best indicator of whether AI becomes a classic infrastructure bubble. Unlike 2000, today’s build-out is funded overwhelmingly from operating cash flow and GPUs run at 100% utilization, versus 99% of fiber sitting unused; yet if TSMC supplied everything Jensen Huang wanted, Baker thinks NVIDIA might sell $2 trillion to $3 trillion of GPUs in 2026 or 2027 and eventually overbuild. The “Goldilocks zone” is enough expansion to keep Intel or Samsung below roughly 30% share, but not enough to remove wafer scarcity.
  • Terafab could challenge normal fab timelines without alienating TSMC. Baker describes a SpaceX joint venture, with Tesla possibly involved, that would use Intel’s institutional knowledge and recruit the equipment companies’ A-teams. He says the relevant process gap may be roughly 9 to 15 months, or three to five quarters, behind the frontier, but presents that timing approximately.
  • Frontier-token economics remain unusually durable, but three uncertainties could overturn the trade: whether frontier tokens retain their premium, whether ASI violates the “bitter lesson,” and when continual learning arrives. Baker says Gemini 3.1 Pro went from mind-blowing to intolerable as the frontier advanced, while Anthropic, OpenAI, and Grok 4.3 now dominate the intelligence-versus-cost Pareto frontier. A model that updates from one experience rather than “put its hand in the fire a million times” could produce a very fast takeoff—and perhaps optimize away some compute demand.
  • Usage-based pricing could steepen AI revenue as capped subscriptions obscure what frontier systems can do. Baker says capped $250-to-$300 plans increasingly deliver rate-limited, “lobotomized” models, whereas enterprise and usage plans expose the token budget and agent harness; he compares the shift to telecom moving from all-you-can-eat to “pay by the drink.” His aggressive call is that OpenAI and Anthropic will exceed well over $200 billion in ARR this year as compute availability, frontier pricing, and fleets of agents expand usage.
  • New chip companies need to be both different and hard, because “a better GPU” invites NVIDIA to copy the trade-off with superior economics and customer knowledge. Disaggregating prefill from decode creates openings for specialized memory-capacity and memory-bandwidth architectures; Baker’s venture rule is that 1% market share could be worth $100 billion, with Cerebras’s wafer-scale design as the exemplar. The same disaggregation could extend Hopper and Ampere lives to 10 or 15 years, potentially lowering GPU financing from the low sevens toward 5% or 6%.
  • Application investors need exposure to the “token path,” defensible scale, or a niche the model companies will not absorb. Baker credits Cursor and Cognition’s coding focus and cites Replit founder Amjad Masad’s view that coding may be the shortest path to useful AI or ASI. He says AI has still destroyed trillions of dollars of application-layer value and that economic gains currently favor businesses with the highest ratio of utilized GPUs per employee. Meanwhile, public-market selection has become finer-grained: semiconductor-equipment companies at 40 times annualized next-quarter earnings and DRAM companies at mid-single-digit multiples “cannot all be true.”
Digest · the substance, structured for research

1. March’s drawdown accumulated alpha rather than invalidating the thesis

  • Baker separates losses where “your hypothesis was invalidated” from underperformance caused by price action that contradicts well-understood fundamentals. March belonged to the second category: investors could lean in and build “pent-up alpha, pent-up future performance.”

  • Anthropic’s $11 billion of ARR added in one month was, in Baker’s framing, equivalent to the combined businesses Palantir, Snowflake, and Databricks spent roughly a decade building. He also cites a 500% net dollar retention statistic shared on the show and calls this “the most extraordinary moment in the history of capitalism.”

  • DeepSeek Monday provided the earlier template. The paper appeared a week before the selloff, but by the panic, Asian AWS availability-zone prices had doubled, GPU availability had fallen, and DRAM was surging—the observable evidence that reasoning inference was dramatically more compute-hungry.

  • The Strait of Hormuz complicated the setup, yet Baker saw a relative U.S. advantage: the Bloomberg natural-gas quote he called “GY” fell 20% while Asian and European gas doubled or tripled. With tech near its cheapest relative valuation in a decade, AI fundamentals outweighed his acknowledged macro uncertainty.

2. Anthropic’s constraint-adjusted multiple looks cheaper than its headline valuation

  • Baker distinguishes Anthropic sharply from OpenAI on capital efficiency: Anthropic has a “dramatically lower cost per token” and, by his estimate, burned perhaps 80% less money to reach comparable revenue. He thinks Anthropic may already generate cash or may begin doing so this year.

  • Compute scarcity is suppressing both intelligence and monetization. Baker cites analysis showing Claude, even on Opus, emitting 70% fewer tokens for the same question; because token quantity partly drives answer quality, he estimates unconstrained revenue could be $100 billion, $150 billion, or even $200 billion instead of roughly $50 billion.

  • O’Shaughnessy asks why Anthropic does not raise $100 billion at a $3 trillion valuation. Baker’s answer is optionality: capital intensity and geopolitical uncertainty reward leaving investors upside, much as Elon Musk preserved a “sacred covenant” by avoiding greedy SpaceX marks and compounding SpaceX at a low-30% annual rate for a decade.

3. Capitalism can manufacture watts, but permissions may become scarcer

  • Baker expects capitalism to solve the power shortage absent “big regulatory or political blowback,” which he considers a real possibility. One major infrastructure investor told him energy and chips had ceased being the biggest gates: “Now it’s zoning and approval. Much more important.”

  • The physical supply chain is genuinely difficult—only two machines can cast certain enormous turbine blades, and the West has not made one of those machines in 80 years—but turbine manufacturers are expanding. Repurposed jet engines offer another bridge, supporting Baker’s forecast that the watts shortage begins alleviating in 2027 or 2028.

  • Orbital compute should be pictured as “racks in space,” not a floating Pentagon. A Blackwell rack weighs 3,000 pounds and consumes roughly 100 kilowatts; Baker envisions rack-sized satellites with solar wings extending perhaps 500 feet per side, sun-synchronous orbits, long radiators, and Starlink-style laser connections.

  • The strongest pushback is maintenance, not cooling. SpaceX engineers have convinced Baker they can manage thermal design, but failed hardware may remain inaccessible “until you have probably floating Optimuses.” He expects orbital inference, terrestrial training, and valuable Earth data centers to coexist for his lifetime.

4. TSMC is the de facto governor of the AI capital cycle

  • Baker caricatures wafer supply as controlled by roughly “twenty older humans in Taiwan,” whom he says view themselves as inheritors of Morris Chang’s legacy. Jensen Huang repeatedly asks them for capacity, despite Baker’s claim that NVIDIA and TSMC do business through fairness, partnership, and handshakes rather than a formal contract.

  • History argues for a bubble whenever markets recognize a foundational technology and suffer a “breakdown in diversity.” Today differs from 2000 because spending comes mainly from operating cash flow and every GPU is fully utilized, versus 99% of fiber going unused—but Baker refuses to assume the last two or three hundred years of canal, railroad, and technology bubbles no longer rhyme.

  • His cautionary investor is Fidelity’s George Vander Heiden, who fought the 1999 bubble, endured client skepticism, and retired in early 2000 with roughly 40% of his fund in tobacco and 40% in homebuilders. Those positions probably outperformed the Nasdaq by 20 or 30 times over the next three years. The lesson associated with him was brutal: “Being early is the same thing as being wrong.”

  • If TSMC fully accommodated NVIDIA, Baker believes $2 trillion to $3 trillion of GPU sales in 2026 or 2027 might create overcapacity. The ideal outcome preserves TSMC’s roughly nine-to-15-month process lead, limits Intel or Samsung from gaining well north of 30% share, and prevents either competitor from breaking industry discipline.

5. Terafab could challenge normal fab timelines without alienating TSMC

  • Baker describes Terafab as a planned SpaceX joint venture that may also involve Tesla, intended to build America’s largest fab. Its Intel partnership would provide 50 years of institutional knowledge; Baker describes the relevant process gap as roughly nine to 15 months, or three to five quarters, behind the frontier, though he presents that timing approximately.

  • He expects the “A teams” from ASML, KLA-Tencor, Lam Research, and Applied Materials to participate, just as equipment vendors once helped TSMC catch Intel because they disliked dependence on a single source. Elon Musk’s hardware reputation should also recruit engineers conventional semiconductor managers cannot.

  • O’Shaughnessy’s pushback is the industry’s unavoidable lead time. Baker’s honest response is “We’ll see”: Musk built a data center in 122 days when others took three years, and reportedly secured an office inside Samsung’s Texas fab because its expansion pace frustrated him.

6. Frontier value persists until the bitter lesson or open models break it

  • Baker is surprised that the overwhelming share of model-layer economics still accrues to frontier tokens. Gemini 3.1 Pro felt “mind-blowing” at launch and “intolerable” later, although he concedes companies may prototype at the frontier before deploying cheaper Vertex or open-source models.

  • Google dominated the intelligence-versus-cost Pareto frontier nine months earlier; after conservative TPU v8 design decisions cost it per-cost-token leadership, Baker says Anthropic and OpenAI now dominate, with Grok 4.3 the best low-cost 500-billion-parameter model and Gemini 3.1 merely hanging on—possibly through subsidy.

  • TurboQuant illustrated how markets can overreact to algorithmic efficiency: Baker found no AI engineer who expected the memory optimization to reduce DRAM demand. The genuine risk is an ASI-driven exception to Richard Sutton’s “bitter lesson,” because a 300- or 400-IQ model might use its first resources to make itself dramatically more efficient.

7. Pay-by-the-drink AI and continual learning could compound the curve

  • Harnesses—the runtime providing tools, prompts, context, memory, and state—matter enormously even if the underlying model matters more. Baker now thinks a $250-to-$300 subscription is insufficient for investors seeking intuition; serious evaluation requires Claude Code or Codex, an enterprise account, and usage pricing.

  • Capped plans deliver a “lobotomized version of the AI,” while usage plans allow the model to spend the tokens its harness judges necessary. Baker compares this to cellular’s growth era of fixed allowances plus overages: AI is leaving all-you-can-eat for “pay by the drink,” just as one person can have 100 agents.

  • Continual learning would be the sharper discontinuity. Humans learn after touching fire once; current models may require a million examples and another training or reinforcement-learning cycle. If weights can adjust dynamically in real time, Baker expects “a really fast takeoff” and makes its timing one of the three central investor questions.

8. Chip startups must make trade-offs NVIDIA cannot instantly reproduce

  • Baker borrows tank design’s “iron triangle”: attack, defense, and mobility cannot all be maximized, just as chip designers must trade within physics and TSMC’s rules. TPU, Trainium, and AMD are essentially trying to be better GPUs, but “nobody’s a better GPU”; Trainium 3 is merely “tugging on Superman’s cape,” while MI450 remains unresolved.

  • Disaggregation supplies a richer design canvas. Andrew Fox’s analogy is “Prefill is loading the cannon, decode is firing”: understanding context is memory-capacity-bound, while generating tokens is memory-bandwidth-bound, enabling specialized chips to optimize aggressively for one phase.

  • Baker’s rule of thumb is that 1% share may be worth $100 billion, but the architecture must be “different” and hard. Cerebras qualifies through wafer-scale computing, three difficult chip generations, and efforts to place an optical wafer above compute to escape shoreline-I/O constraints.

  • Disaggregation also changes collateral economics: Cerebras or Groq systems can sit in front of Hopper or Ampere GPUs, allowing those GPUs to perform prefill “until it melts.” Baker thinks they could have 10-to-15-year useful lives, potentially shifting GPU financing from CoreWeave’s low-sevens floor toward 5% or 6% and helping “single-handedly save private credit.” This matters while private credit is already under pressure from SaaS loans.

9. Applications need token-path leverage before the frontier absorbs them

  • Baker applies “different and hard” to venture broadly. Retail CEOs could destroy obvious e-commerce startups by driving category margins toward “negative ten thousand percent”; Wayfair survived because its founders solved an operationally difficult problem before the idea became obvious and scale erased the opening.

  • Baker credits Cursor and Cognition’s scale to their early, intense coding focus and notes that Anthropic was also among the companies focused on coding. Replit founder Amjad Masad’s “bitter lesson adjacent” insight is that coding may be the shortest path to useful AI or ASI, because a sufficiently capable coding system can create tools for everything else.

  • Yet Baker says AI has net destroyed trillions of dollars at the application layer, even counting Cursor and Cognition. Today’s winners have the highest “effective ratio of utilized GPUs per human”; other software companies need to sit in the “token path,” like Databricks, or build a vertical moat before frontier labs reach the niche.

  • Open source introduces another prisoner’s dilemma. Baker thinks frontier labs may withhold systems such as Mythos to prevent distillation, but one lab releasing its best model via API could gain revenue and intelligence-enhancing resources, pressuring every rival to follow; NVIDIA can meanwhile keep open source a controlled distance behind. “Open source” still consumes energy and GPUs, and open-source model companies commonly receive a revenue share.

10. Public-market leadership is fragmenting beneath the AI headline

  • Google’s per-cost-token TPU advantage is gone, but its installed compute, search, YouTube data, and accelerating GCP keep it formidable. Baker treats the Google I/O release that week as a test: failure to leapfrog OpenAI or Claude would strengthen the case that NVIDIA’s architectural advantage is larger than assumed.

  • Meta earns credit for becoming genuinely AI-first; Muse, MSL’s first model, landed surprisingly close to the Pareto frontier. Amazon combines Trainium with prospective retail-robotics efficiencies over 18 months, while its Nova models are “better than they get credit for.”

  • Microsoft briefly “flinched” on capital spending in early 2025 and lost allocations, but Baker supports Satya Nadella’s risky decision to reserve compute for Copilot and internal models rather than maximize near-term Azure sales. He estimates Microsoft might otherwise be an $800 stock, while questioning whether its current model team can execute.

  • Cross-sectional prices now “cannot all be true”: semiconductor-equipment companies trade near 40 times annualized next-quarter earnings while DRAM sits at mid-single-digit multiples. Correlations among GPU compute, networking, optical, DRAM, NAND, and HDD broke in January, creating opportunities in miscategorized names such as Astera, whose biggest product is going to be a switch and which Baker says is wrongly placed in copper-loser baskets.

11. Investors must master the machine while society prepares for its blowback

  • Baker’s Last Samurai analogy is uncompromising: “The machine gun is here. If we do not all become masters of the machine gun, we’re going to get mastered.” His most valuable agent extracts personally relevant needles from six daily hours of podcasts; others inspect proxies, PSUs, incentives, and compensation changes.

  • Defensive preparation starts with cybersecurity and an offline family or company safe word that cannot be socially engineered. Baker expects realistic video impersonations to know personal context and request transfers; he also fears rising political violence will increasingly target public AI leaders.

  • Geopolitically, he attributes Ukraine’s improving battlefield position partly to superior battlefield AI and worries U.S. dominance could destabilize adversaries—though it might instead create another Pax Americana. His counterweight is AI’s medical promise: agents helped find a drug already on the market that could impact one child’s rare disease and helped launch a company seeking a cure.

  • Baker remains an “AI optimist and maximalist,” but calls the technology an “event horizon” demanding humility. Frontier access becoming contingent on wealth feels “a little dystopian”; the Luddites may be wrong, he argues, while their concerns still require thoughtful answers that make the gains broadly shared.

Patrick O'Shaughnessy

All right, so this is our 6th time doing this, if you can believe it, which puts you back into first place, or at least tied for first place with Gurley—back into Steve territory. Always my favorite conversation about markets and everything going on.

Even since the last time we did this, which was so exciting and spectacular, I think we're in an even more interesting time now. Maybe just start by riffing on how it felt for you living through March and April of this year, which felt to me like a completely unique economic, technological, and market environment. You're the biggest student of history and of these times, so what did it feel like?

1. March Revealed The AI Inflection

Gavin Baker

I would say, broadly speaking, there are 2 kinds of drawdowns. There are drawdowns where you're wrong, a company missed estimates, your hypothesis was invalidated, and you have to take your medicine and crystallize that loss. And then there are drawdowns or periods of underperformance where you're underperforming because of companies you know really, really well and where you profoundly disagree with the price action. You can lean in, and instead of crystallizing negative performance, you can build pent-up alpha, pent-up future performance.

For me, that is what March felt like. The Nasdaq was selling off. At the same time, what was happening in AI was, I think, the most extraordinary moment in the history of capitalism, the history of American business. What I mean by that is that Anthropic added $11 billion of ARR.

What is astonishing to me about this is that the SaaS and cloud revolution created, we'll call it, between $5 trillion and $10 trillion of value. I would say arguably the 3 highest-profile SaaS companies in the last 10 or 12 years are Palantir, Snowflake, and Databricks. These 3 companies employ thousands of people, tens of thousands collectively. They've all spent 10 years building their businesses. Anthropic added their combined businesses in 1 month.

Nothing like that has ever happened in the history of capitalism. Forget my career—just the flat-out history of capitalism, the history of business. It's wild. And then Krishna comes on this show and shares some stats: 500% NDR.

Patrick O'Shaughnessy

Yeah, you do the math on that for 3 years. It's insanity.

Gavin Baker

So there's just no precedent for this, and we tech investors hear a lot of discussions about S-curves and investing in exponentials. I've just never seen an exponential like this. It felt even more extreme than DeepSeek, which was a very similar setup. They happened at about the same time.

If we go back to 2025, there was a huge sell-off on DeepSeek, which was very strange because the paper got published 7 days before DeepSeek Monday. It got published, I believe, on a Monday that was a holiday in America. I read it and thought, “Hmm, this feels like it—”

Patrick O'Shaughnessy

Could be important.

Gavin Baker

It might not read positively for the AI trade. I took action, and then we had DeepSeek Monday, where AI really imploded a week later.

That was really strange because by DeepSeek Monday, it was super clear that this was going to be the most positive thing that had ever happened to compute demand. Prices of the AWS availability zones in Asia had already doubled. You were seeing GPU availability go down, and this was just the first time we saw how much more compute-hungry reasoning models are during inference than non-reasoning models.

And so that was a similar setup. You had to do some work to see that. It's not that hard to say, “Oh, wow, stocks are selling off. The price of DRAM is going vertical. The price of GPUs in Asia is going vertical. GPU availability is going down.” Then, 2 or 3 days later, GPU rental prices in America started going up.

All you had to do in March was simply observe what was happening to Anthropic. There were all these people who seemed to regret not buying during 2022, not buying during COVID, and not buying during DeepSeek. You had the same valuation setup at the beginning of April, and an even clearer AI inflection. So there have been all these chances to buy into AI.

2. The Strait Of Hormuz Advantage

Then, of course, what complicated things was the Strait of Hormuz. I'm no macro expert, but I do a lot of pro-national-security investing, so I have access to people who are experts and are excited to share their thoughts and opinions with me. I became, and remain, a believer in the idea that one thing the market was mispricing is that the Strait of Hormuz being closed is actually relatively awesome for America.

Patrick O'Shaughnessy

Why?

Gavin Baker

Particularly for the goals of the current administration. Electricity is a very important industrial or manufacturing input. The key input into American electricity prices, which feeds into AI, is GY, natural gas, one on Bloomberg. That was down 20%, and natural gas in Asia, Europe, and everywhere else doubled or tripled.

Our relative manufacturing competitiveness improved overnight. For better or worse, that is what the Trump administration seems to care about. They are very focused on America's relative position, and I think a lot of people had memories of the 1970s. What made the 1970s so traumatic was that it wasn't just that prices went up; there were actual gas shortages.

Then you go through: “Okay, well, the U.S. economy is dramatically less energy-intensive than it was. The United States is now the world's largest producer of oil and gas, and we've now become the world's largest exporter of oil and gas.” On top of that, there's this relative manufacturing advantage.

That made it easier to stay focused on AI fundamentals and on what were historically attractive valuations. I think, on a relative basis, tech essentially got as cheap versus the rest of the market as it has been at any point over the last 10 years.

Just think about that in the context of market efficiency. We have the most extraordinary moment in the history of capitalism that's wildly bullish for AI, and you get a chance to buy AI at really attractive valuations.

3. Valuing Anthropic And OpenAI

Patrick O'Shaughnessy

What do you make of the multiples that specifically Anthropic and OpenAI—which, in my mind, are the reference assets that are the purest-play takes on this trend—are trading at, really not being that crazy? If you just look at the sales multiple and compare it to what Databricks and Snowflake traded at, maybe at their peak, how do you process it? How do you make sense of it?

Gavin Baker

I do think OpenAI and Anthropic are pretty different animals from a capital-efficiency perspective, and Anthropic clearly has a dramatically lower cost per token than OpenAI. They just do, and you can see that in the amount of money that they have burned to get to a roughly similar revenue scale. I think they've burned maybe 80% less than OpenAI.

So, as businesses, they clearly have very different structural ROICs. I think Sarah Friar's one of the most exceptional CFOs. I think they're doing a lot of things to try to improve this.

Patrick O'Shaughnessy

And they've secured a lot of compute.

Gavin Baker

They've secured a lot of compute. That's another big difference. It turns out being aggressive really paid. Anthropic at a $900 billion valuation for $50 billion in ARR, growing at ridiculous rates. And I think maybe a true statement is that if Anthropic could just wave a magic wand and get all the compute they wanted, they'd probably be doing well north of $100 billion today, maybe $150 billion.

They have clearly deprecated the intelligence of Claude. There's an analysis that Claude, even on Opus, is generating 70% fewer tokens for the exact same question. As we talked about last time, token quantity equals quality of answer and quality of thinking at some level, and there is an intelligence density per token that also matters. I've felt that as a user.

So I think they would be doing materially more—$100 billion, $150 billion, maybe $200 billion. So you might be buying it at more like 5 times unconstrained… I'm going to make up a new number: URR, unconstrained run-rate revenue.

Patrick O'Shaughnessy

Why do you think they don't raise $100 billion at a $3 trillion valuation or something like this? If you were the Anthropic CFO—Krishna's awesome; we just had him on—or if you're Sarah, it seems to me like, if the inbound I received following the Krishna episode is any indication, everyone I've ever met is trying to invest in both these companies.

Gavin Baker

I think it's wise. The future is uncertain. You are clearly in a very capital-intensive game. Even if you are Anthropic, I am sure Anthropic is at very positive gross margins on inference today. I think Anthropic probably starts generating cash this year, if they're not already generating cash, which I think is probably the case.

But still, you probably want to be able to raise more capital and access more compute. The world is uncertain. Ukraine is starting to really, really win. How is Russia going to respond? I think there's still a lot of uncertainty in Iran. All this uncertainty, I think, probably amplifies geopolitical uncertainty over time.

So it's an uncertain world. If I think about Elon, Elon has always made investors money. He treats it like a sacred covenant, and as a result, because he's made people money for now 20 years, he has a superpower. That is, he can essentially raise as much capital as he wants whenever he wants. I do think being focused on making investors money is wise and creates benefits that don't just last for a year or two. They can last for the next 20 to 30 years.

Patrick O'Shaughnessy

And the way Elon did this was systematically underpricing SpaceX or whatever else. What's the actual method?

Gavin Baker

Just never being greedy on valuation, never pushing valuation. Just that simple. My friend Antonio pointed out SpaceX compounded at a low 30% per year for a decade, and that was just because Elon was, I think, focused on preserving the superpower and trying to strike a fair balance between investors and employees. I think it's wise. But could Anthropic raise money at probably at least a 100% premium to this rumored latest mark? Of course.

4. The Compute Infrastructure Race

Patrick O'Shaughnessy

Let's get to the watts and wafers part of the discussion. It's always my favorite thing to talk about with you. On the importance of this infrastructure build-out: every time I feel like it's getting overheated, the next time I talk to you, it seems like we should've done way more than we did. You studied S-curves and the steepness of those S-curves a lot, and you know a lot about history. Talk us through how you're thinking about watts and wafers today as the key to inputs into this whole thing.

Gavin Baker

I think capitalism is going to solve the watts shortage absent big regulatory or political blowback, which I think is a real possibility. The head of data center infrastructure investing at one of the big PE firms—Blackstone, Apollo, KKR—said, “It used to be that energy and chips were our biggest gating factors. Now it's zoning and approval. Much more important.”

I think a lot of companies are waiting until after the midterms to take action in terms of maybe workforce reductions. Nobody wants to be a piñata during the midterms. You've seen a lot of companies that make turbines announce plans to significantly increase capacity. There are 2 of these machines that can cast these big blades. We haven't made one in 80 years in the West. We don't know how to make them anymore.

All of that is true. By no means am I minimizing the industrial engineering magic and artistry that goes into those, but capitalism is very good at solving problems like these over time. There are other sources of energy besides these turbines with a longer timeframe. So I think the watts shortage will probably begin to alleviate in 2027 or 2028, and then I think orbital compute will really solve that.

I do want to reframe orbital compute because I think when people hear “data centers in space,” which we discussed in our last episode, they picture a Pentagon-sized building in space. They're like, “Well, we can't do that.” That's not what it is. A Blackwell rack weighs 3,000 pounds. It's 8 feet high. It's 4 feet deep and 3 feet wide. It's racks in space. And SpaceX has showed you an illustration, and it's a rack.

That's the satellite, but it's probably about the size of a Blackwell rack. It has these solar wings that are probably 500 feet long on each side. You keep it in a sun-synchronous orbit, so those solar panels are always in the sun. Because it's in an exactly sun-synchronous orbit, the radiator, which extends behind it for hundreds of feet—

Patrick O'Shaughnessy

This is a common criticism, yeah. How are you going to cool a—

Gavin Baker

Yeah. I've spent a lot of time at Starbase over the years, and I've talked to a lot of SpaceX engineers. I do think it is the most talented group of engineers on planet Earth, and they're very confident they have solved this. And they're not always confident. There's some engineering that needs to happen to turn the Starship into a Mars colonial transporter. Will they do that? Absolutely. What are they more focused on? I'd say probably the repair and maintenance.

Patrick O'Shaughnessy

Those are the 2 big—

Gavin Baker

Yeah.

Patrick O'Shaughnessy

…responses, the radiator and how do you repair the—

Gavin Baker

Yeah.

Patrick O'Shaughnessy

…whatever issue goes wrong on the rack.

Gavin Baker

And the answer is: until you have probably floating Optimuses, you don't. Starship is going to change the space economy in ways we cannot imagine. And particularly if regulation becomes a constraint to data centers, none of it's going to matter. You're going to sell as much orbital compute as you can make.

And then obviously you link these racks using lasers traveling through vacuum, which are already on every Starlink. It's just mind-blowing to me that SpaceX operates the world's largest satellite fleet, which is 98% or 99% of all satellites in orbit. They're cooling every Starlink today. I think Starlink V3 is going to operate at 20 kilowatts. A Blackwell rack is only 100 kilowatts, and people talk a lot about density.

Well, if you're connecting the racks with lasers through vacuum, you can make the rack bigger. Physically, you're focused on weight, not size. In a data center on Earth, where you're trying to connect racks, ideally using copper and minimizing lengths, cabling is a big cost. You do want that rack to be small. Copper when you can, optics when you must.

But in space, there's all sorts of things that SpaceX can do that I think maybe some of these naysayers are not contemplating. They operate more satellites than anyone. They have a 20-kilowatt satellite today. So maybe you just scale that up to 60 kilowatts to start. They seem very confident they're going to go right to 100 to 120 kilowatts.

The same company now also operates the largest data center on Earth. They have the world's best hardware engineers, and all sorts of people—almost all of whom are not smart enough or practical enough to work at SpaceX—are these armchair skeptics. I don't want to quote Larry Ellison, but somebody was being skeptical, and Larry was just like, “Listen, he's out there landing rockets. I don't see anybody else landing rockets.”

And the reality is, 10 years later, no other company is consistently capable of landing and fully reusing an orbital rocket. None of this makes sense without reusability. That means you have to land it. I would like to redefine orbital compute as racks in space, not giant floating—

Patrick O'Shaughnessy

Not data centers. Yeah.

Gavin Baker

…Pentagon-sized data centers in space. That's silly. What makes a data center is you're connecting these racks with lasers. So it'll be racks in space that are connected with lasers into a virtual data center.

Patrick O'Shaughnessy

If you think about that state of the world, let's say that all happens and we're really good at getting these things up economically and running matrix multiplication all over space. What does that mean for terrestrial data centers?

Gavin Baker

Someone once said America was going to suck as hard as it can on every energy source it can get, and I think the same is true of compute. It's why I'm probably less worried about an edge-AI bear case than I was. Inference, I think, is very sensible for orbital compute. Training will be done on Earth for a long time, so I don't think this is super-bearish for terrestrial data centers. I think those are going to be valuable for my lifetime.

But I do think that if you're in this ecosystem of power production and cooling and are massively ramping capacity, a lot of these capacity ramps are going to be hitting just as all of the silly skeptics start to understand that orbital compute is very real. I think it's worth thinking long and hard about that if you're one of those companies. All sorts of cool stuff is happening in the interim. We're getting really good at repurposing jet engines. There's Boom Supersonic—

Patrick O'Shaughnessy

Yeah.

Gavin Baker

—that is doing this. Capitalism is hard at work on watts. On wafers, though, it's just this group of 20 older humans in Taiwan who are the most important humans in Taiwan, whoever they are. They account for the overwhelming fraction of the country's GDP, water usage, and electricity usage. They talk about the silicon shield. They all view themselves as inheritors of Morris Chang's sacred legacy.

I vividly remember visiting Science Park more than 20 years ago and asking them, "Do you think you could catch Intel?" They said, "This is such a beautiful dream, but it's a dream for our grandchildren." They did it partly because of Intel's self-inflicted wounds. They think very differently. One reason Jensen Huang flies over there so much is that he wants them to expand capacity.

I do think it's wild that Jensen has never had a contract with TSMC. They do business based on what seems fair, with handshakes: "It's going to be fair over time. We're partners. We're going to be fair to each other." It's just fascinating. No contract.

The truth is, based on every prior market precedent for a foundational new technology like AI, you've always had a bubble. Carlota Perez wrote this great book about this. Markets are efficient. They correctly understand that this is a foundational new technology. There's what Mauboussin calls a breakdown in diversity. Everyone becomes bullish on this new technology, and I am beginning to worry a little bit about a breakdown in diversity. Then you get a bubble.

That bubble funds the build-out of this new technology, but supply gets ahead of demand, and you get a crash. It's a particularly severe crash if it's a debt-fueled build-out, like in the year 2000. One thing that's really good about the current build-out is that it's still overwhelmingly funded out of operating cash flows, which is a really important fundamental difference versus the year 2000.

There's the valuation, and there's the fact that every GPU is running at 100% utilization, when 99% of fiber was unutilized. So there are all these fundamental differences. History doesn't repeat, but it rhymes, and as investors, we have to be very cognizant of it and recognize that, based on the last 2 or 3 hundred years—forget the internet bubble—we had a railroad bubble, a canal bubble, every kind of bubble, the South Sea bubble. We should expect a bubble.

That's terrifying. Nobody wants a bubble. The reason it's terrible is that if you're valuation-sensitive, you massively underperform and probably get fired by all your clients. George Vander Heiden, who is no longer with us, was a great Fidelity portfolio manager. He fought the bubble in '99, and he retired in early 2000 because I think he just couldn't take it. He knew it was wrong. His clients were deeply skeptical: "George, you're out of step. You don't get it."

He had white hair. He was a truly great man. I only overlapped with him briefly, but he was a very important mentor and friend to my good friend and mentor, Jennifer Urig. So I have a lot of Vander Heiden DNA through her. He was the same person who said, "Being early is the same thing as being wrong."

George retires because he couldn't take the underperformance, and he couldn't take clients saying, "What's wrong with you? You don't get it." And he has like 40% of his fund in tobacco, 40% in home builders, and literally he probably outperformed the Nasdaq by like 20 or 30 X over the next three years.

I've been optimistic that this fundamental shortage of wafers, which today is really controlled by TSMC, will prevent one. If Taiwan Semi did what Jensen wanted, I think NVIDIA could sell two trillion dollars of GPUs in '26 or '27, maybe two and a half trillion, maybe three trillion. But there is a limit where consumers would consume so much that they would probably be in an overbuild.

So, if we don't get a bubble, we need to throw a party for TSMC because it will have single-handedly prevented a bubble. You're starting to see companies go to Intel and Samsung.

Patrick O'Shaughnessy

Let's just assume TSMC stays super supply-constrained versus rampant demand. What happens?

Gavin Baker

In the history of markets, I don't know which one, but one of Intel and Samsung is not going to stay disciplined. They will break, and then, at some level, that will force everyone else to break. I think a lot of this may come down to the degree to which TSMC can maintain a lead over Intel and Samsung. You have to remember, whatever it is, it's 9, 12, or 15 months.

Patrick O'Shaughnessy

The leading node at GMEN.

Gavin Baker

Exactly—the pace at which they expand capacity. If I were to watch one thing to understand whether there's a bubble, it would be TSMC's capacity decisions. I think there's a Goldilocks zone where they expand enough to make it hard for Intel or Samsung to truly emerge as an at-scale second source with well north of 30% market share. Yet they also keep this fundamental constraint on wafers that helps us avoid a bubble. Obviously, I think the Terafab is going to play into this, too.

Patrick O'Shaughnessy

Say more about that.

Gavin Baker

It's a SpaceX joint venture—I believe Tesla is involved as well—to build the world's largest fab here in America. I think they're going to be successful. First, they have a partnership with Intel, which is very important because they're getting access to 50 years of institutional knowledge. That's just nine months, a few quarters, twelve months, three to five quarters behind the front. That's an advantage.

It's also an advantage that I believe the Terafab is going to get attention from the A-teams at all the semicap equipment companies. One big reason TSMC caught up is that ASML, KLA-Tencor, Lam Research, and Applied Materials wanted them to catch up. They don't like having a monopsony. The A-teams were in Taiwan working. Intel made some mistakes, and presto.

The A-teams will be here because of Elon Musk's reputation in hardware engineering. To a degree that I think is maybe hard for people to imagine in America, where politics has replaced religion, I think that because Elon had his foray into politics, it makes it hard for some people in America to see him clearly. That's sad, because I do think he's probably doing more for America than any other American.

He's single-handedly bringing manufacturing back to America. He's revived defense tech. I think SpaceX is, in some ways, the most important defense contractor in America. He's doing Starlink. It's amazing for the world. He's creating all these blue-collar manufacturing jobs, which is a goal, I think, of a lot of liberals, and it's good for America. He's done more than any living human to decarbonize the world.

And if you are upset about data centers on Earth for environmental reasons—

Patrick O'Shaughnessy

That's what it's for.

Gavin Baker

—well, here you go. It's sad, but he is a living deity in China, Taiwan, South Korea, and Japan. Having watched him for a long time, what he's going to do is recruit the best people, because the best engineers want to work for Elon, especially in hardware engineering.

Next to the Terafab, there will be a Taiwan Town. "Oh, these are your favorite restaurants? I'm going to move them and their whole staff from Taiwan to Texas, and we're going to make everything the way they like it." Then we'll have Japan Town. Same thing. Then we're going to have Korea Town. We're going to have all these things exactly dialed in to recruit the best engineers, and that's just not the way that the people who run Intel and Samsung think.

So he's going to have the best talent. He's going to have the A-teams at the wafer-fab equipment companies. He has Intel, which is important. It's so good for any administration's political goals, and I think it's different enough that it will not alienate TSMC.

Patrick O'Shaughnessy

And these have long lead times, right? The Terafab is going to be pumping out whatever GPUs, whatever chips, quite a long time from now.

Gavin Baker

We'll see. Elon tends to do things differently. Everybody else has taken 3 years to build a data center. He built one in 122 days. Samsung had to give him an office in their fab in Texas because he was so unhappy about the pace at which they were expanding and building.

5. The Frontier Token Premium

Patrick O'Shaughnessy

You mentioned DeepSeek earlier. The simple reaction to that was, "Okay, these models are just going to get 95% as effective for some tiny fraction of the cost. It's still Chinese open-source models."

We'll be able to use these for most of what we want to do. Fast-forward a little bit—two years from now, there's no reason I have to spend $1 million a year in my small firm on tokens or something. But the actual reality seems quite different from this, and I'm curious why there's that dissonance in your mind.

Gavin Baker

I do think it's fascinating that the returns to the frontier—all the economic returns to AI at the model layer, not all of them, but an overwhelming amount of them—have been at the frontier. That's surprising to me, and I think it's been surprising to a lot of people. This is one of the most important questions to be answered, and you need to have a hypothesis on it as an investor: Are frontier tokens going to continue capturing the overwhelming majority of economic value created at the model layer?

It is surprising. I remember when Gemini 3.1 Pro came out. It was mind-blowing to me. It was so good. Today, it's intolerable. There's probably a little bit of a dynamic where companies prototype with frontier models, then, when they put something into production, you're hearing a lot of people use Vertex or open source. But still, it is a fact today that the overwhelming majority of these economic returns come from frontier tokens, and that's surprising.

Whether or not it continues, I think, is a very interesting question. I'm much more open-minded to that, having had the experience I've had with Gemini 3.1 and then Opus. I also use Grok 4.3 a lot. It is on the Pareto frontier.

The companies that are on the Pareto frontier—and this is, by the way, a big change and a consequence of what we talked about last time—Google lost its per-cost-token leadership as a result of making very conservative design decisions with TPU v8 to try to take it away partially from Broadcom, while NVIDIA continued to make aggressive choices. But Google dominated the Pareto frontier, the Pareto frontier being intelligence versus cost.

I think this is the most important thing to look at when analyzing AI labs. Google dominated that nine months ago. Every point on the Pareto frontier from OpenAI, xAI, and Anthropic was inside of it. Now, the Pareto frontier is dominated by Anthropic, OpenAI, and then Grok 4.3 is on the Pareto frontier. It's clearly the best, lowest-cost 500-billion-parameter model.

Gemini 3.1 is hanging on to the Pareto frontier, and if I were to bet, I'd bet that they are subsidizing that out of pride. I would just say that a violation of Richard Sutton's bitter lesson is, for sure, the biggest risk to this trade—to all of AI. The closer someone is to AI, the more skeptical they are that this will occur.

One thing I think contributed to weakness in March was a much more stupid version of DeepSeek, which was this thing called TurboQuant. TurboQuant is some Google memory optimization that was written up in a paper a year ago. Then, in the middle of an agreement, while Google was negotiating with Micron, Samsung, and SK hynix to sign some LTA that would lock in really high prices for a long time, they released this.

What people do is always more important than what they say. They just publicized it on X, and it went viral: “Oh my God, DRAM is cooked. There's this DRAM optimization.” I was unable to find a single AI engineer on planet Earth who believed that TurboQuant would have any impact on DRAM demand.

Nonetheless, a violation of Richard Sutton's bitter lesson—more compute will always outperform human algorithmic ingenuity. More compute and data, Chinchilla-optimal and beyond Chinchilla-optimal, I guess, is what people increasingly use today. That's a real risk, man. The people who are building these models are skeptical of that risk.

The reason I am a little less skeptical is that I think we're very close to ASI. Who knows if the bitter lesson holds for 400-IQ models? Maybe we get a temporary period where these—

You know, if you get to ASI, the first thing it wants is probably to be smarter and have more resources. How does it do that? It makes itself more efficient. I think that is an actual risk. The bitter lesson literally, I believe, includes humans in it.

We're about to find out whether the bitter lesson applies to 300-IQ AIs, then 400, then 500 and 600. At some point, we may have a temporary violation of the bitter lesson based upon AI and ASI.

Patrick O'Shaughnessy

I'm curious how you think about some other parts of the innovation around the model, continual learning and memory being 2 that people seem to be most focused on as things that might create yet another new paradigm that we would enter. What do you think about the role of those 2 things?

Gavin Baker

I think we've done a lot with memory through these harnesses, and it turns out that harness engineering is not as important as the model, but it really matters. These harnesses and these models are increasingly being co-developed. One of the big things a harness does—just think of it as a runtime that the model operates in—is that it knows where the tools are. It creates context, memory, and state, and has very specific prompts or instructions.

Patrick O'Shaughnessy

It makes a huge difference, even simple versions.

Gavin Baker

It makes an incredible difference. I think the last time I was on here, or one of the other times, I said, “As an investor, it's very important that you pay for the $250-a-month version to get your own intuitive sense.” That's no longer possible.

To understand what frontier AI is capable of today, even for a non-coding use case, you need to have Claude Code or—

Patrick O'Shaughnessy

Codex 5.1.

Gavin Baker

—you need to be on an enterprise plan. The reason for this—and this is another dynamic that's enabled by Google losing its cost leadership—is that these AI models just shifted to usage-based pricing.

If you're on that $250, $300, or $280-a-month plan, or whatever it is, you are getting severely rate-limited. You are getting a lobotomized version of the AI. Because, as we talked about, Claude now produces 70% fewer tokens. You want the tokens that Claude and its harness really think it needs to produce to get you a good answer, you need to be on a usage-based plan.

By the way, this is so bullish for AI. If we go back to 2005 to 2007, cellular had been a great growth industry really for the last 10 years, and the reason was that you had a combination of fixed pricing—you had 900 minutes for whatever it was—and then usage-based pricing over that. When did cellular stop being a great growth industry? When everybody just went to all-you-can-eat.

Long distance is the same thing. AI is just shifting from all-you-can-eat to pay-by-the-drink. It turns out people really like to talk to their friends long distance. They really like to talk to their friends on the phone, and people really like to use AI, particularly now that 1 person can have 100 agents working.

I think the shift to usage-based pricing is probably why you will see OpenAI and Anthropic exceed well over two hundred billion dollars in ARR this year. Not only is more compute going to come online, but they're going to be able to push frontier-token pricing with these usage-based enterprise models.

It's sad. It's sad for the world because it just means if you can't afford that, you're not at the frontier. I think it's going to throw off a lot of investors' intuitive sense of the capabilities of AI.

But, yeah, continual learning, man. I mean, if we solve that—

Patrick O'Shaughnessy

How do you conceptualize that?

Gavin Baker

AI is constantly updating its weights. It may end up being something different. There are so many mysteries about the human mind. We're such sample-efficient learners relative to AI—many orders of magnitude.

We have a crude variant of continual learning today when something is verifiable, and that's just reinforcement learning during mid-training. Continual learning is a model that dynamically adjusts its weights, or adjusts in some way, in real time.

As a human, the first time I put my hand in a fire, I've learned never to put it in there again. That model today needs to put its hand in the fire 1 million times and then have the designers effectively put a fire in the next training run or an RL gym for it to learn.

I think it has to be dynamically updating the weights, but I think people are working on really smart techniques beyond this. If we get that, then we have a really fast takeoff, and people seem confident that continual learning is just around the corner.

I do think this is the third big question: Will bitter-lesson violations as a result of ASI make human ingenuity less important? Will frontier tokens continue to command the premium they do? And will we get continual learning—and if so, when?

6. Different And Hard Wins

Patrick O'Shaughnessy

What is the role of new chip companies in all of this? We talked a lot about NVIDIA and its relationship with TSMC and Intel and all these sorts of things. There's a thousand flowers blooming—literally, probably 1,000 flowers blooming—trying to create a new chip to address some part of this bottleneck. I'm curious how you process this space and this opportunity, and what role it will play.

Gavin Baker

I think this is good and healthy for the world. It's good for Jensen, too, because a different administration might take a different view. Competition is good for everyone, and seeing different architectures explored is good.

The reason is that, in tank design, they talk about the iron triangle. The iron triangle of tank design is that all designers of a tank have to make trade-offs between attack, defense, and mobility, for obvious reasons. The more defense you have, which is just armor, the heavier the tank is, and the less mobile it is. So you have to live in this triangle and make trade-offs.

The Merkava in Israel is optimized for defense. Russian tanks and the Leopard are generally more optimized for mobility. Chip design is the same. There are fundamental constraints imposed by the laws of physics, as embedded in TSMC’s design rules, that you need to live within.

You have TPU, Trainium, and AMD, which are all essentially trying to be a better GPU. Today, I think Trainium is probably doing the best. Nobody’s a better GPU, but Trainium is tugging on Superman’s cape. It hasn’t started yet. Trainium 3 needs to ramp into production because it has a switch scale-up network, which you really need to economically infer MoE models. A lot of companies have a torus architecture. That’s where Google was. Google’s developing a switch scale-up network, and then AMD is like always kind of flying—

Patrick O'Shaughnessy

A little bit behind.

Gavin Baker

Yeah. AMD—we’ll see. The MI450, we don’t know yet. We’ll see. We probably know more about Trainium 3 than the MI450. But that’s a hard game to play. You have to do something different, and you have to do something different that is also hard to do.

I think the best path for these startups, my rule of thumb is one percent market share is going to be worth a hundred billion. A hundred billion is a pretty good venture outcome. I think what Jensen would say is, “Okay, if somebody does something different and it gets to a 1%, 2%, or 3% share, we’ll make that chip.” And that’s coming for everyone.

But if you’re trying to make a better GPU, good luck. If you’re doing something different, it also needs to be hard to do, and you can make different trade-offs. The disaggregation of prefill and inference has really opened the aperture for making these different trade-offs, because you can make very aggressive trade-offs for decode and aggressive trade-offs for prefill.

Patrick O'Shaughnessy

Prefill is taking in the context, and decode is writing the output.

Gavin Baker

Yeah. I have a great colleague named Andrew Fox who said, “Prefill—picture a British naval ship from the 18th century. Prefill is loading the cannon; decode is firing.”

What prefill literally is is the model understanding the question, the prompt, and then keeping track of its own answer. That is fundamentally a memory-capacity-bound problem. Decode is the process of generating new tokens, and that is a memory-bandwidth-constrained problem. If you’re a chip designer, this gives you a richer canvas to paint on.

Even so, it needs to be hard, because if you make different trade-offs in that iron triangle to optimize for memory capacity, and they’re not hard trade-offs to make, NVIDIA is going to make those same trade-offs. They get better prices from TSMC than you’re ever going to get. They also have the advantage of working with every model company and optimizing their designs.

By the way, another very funny thing is that there’s this process: If you’re a VC investing in a semiconductor company that is telling you they’re going to have an advantage because of a TSMC process that they have special access to, I promise you that Jensen saw that process when it was a twinkle in TSMC’s eye. They know more about it than this little company with 200 people can imagine.

TSMC—everybody in the supply chain is showing Jensen everything, the same way they’re showing Amazon everything, AMD everything, TPU everything. That’s another reason not to “try to make a better GPU.” You could do something different. You could paint in the prefill canvas. You can paint in the decode canvas. But you also have to do something hard, because if it gets to scale, you’re going to have those 4 companies as very fast followers.

My firm was a venture investor in Cerebras. What Cerebras has done is something hard and fundamentally different: wafer-scale computing. It comes with a set of trade-offs, but that architectural decision they made was hard and lets them do something that no one else can do. We’ll find out how big that is. They’re working on really cool things.

One of the problems Cerebras has is that once you start needing to glue a lot of chips together and scale-up or scale-out networks, you need a lot of I/O, and I/O is bound by what’s called the shoreline—the sides of the chip. Cerebras has an overwhelming ratio of on-chip compute and memory relative to shoreline I/O. They’re really smart people. They did something really hard. They’re trying to see if they can put an optical wafer right on top of that, and then that solves the problem.

I’m sure they’re looking at hybrid bonding of DRAM to get around the much-discussed limitations on X that are not true. A Cerebras machine can theoretically run any size model. There are sizes of models where they’re much better than other sizes. What I think is interesting about Cerebras is that they did something different that’s hard to do—really hard to do: wafer-scale computing.

I do think there’s a role for these companies. I would just encourage them all to make a different trade-off and try to do something hard. Everybody’s going to get funded after the Cerebras IPO. It’s not going to be a problem. But it took Cerebras 3 generations of chips to get it right. Andrew Feldman, the CEO, makes it clear how hard it was for him and that whole team to get where they are today. They need to have the grit and resilience to do that.

The first chip is a failure. It happens. Can you come back and make a second chip? The one last thing on this topic is that this is going to be amazing for the useful lives of GPUs and may single-handedly save private credit.

Patrick O'Shaughnessy

Say more about that. What do you mean by private credit?

Gavin Baker

Private credit is in pain from these SaaS loans, and however much they’re marked down, they probably need to be marked down more. If the public companies are struggling to adapt, how is a debt-laden company going to adapt and invest in what is a very different margin-structure business?

There’s a lot of private credit in GPUs, too, and they were underwriting that to, I think, 3 or 4 years. The disaggregation of inference means that I think these GPUs are going to have 10- or 15-year lives.

The AI skeptics are like, “Oh, these companies are all cooking their books. The useful life of a GPU is only 1 or 2 years. The useful life of a CPU is only 4 years because of the rapid technological change.” No. What rapid technological change has done with the disaggregation of prefill and inference is mean that you can put a Cerebras system or Groq LPUs—which NVIDIA acquired—effectively in front of a Hopper or even an Ampere, use that Hopper and Ampere for prefill, and extend the useful life of that GPU until it melts.

They do melt, so they have a finite life, but maybe you don’t have to run them as fast. This is going to be really good for the whole private credit industry. It’s going to help finance the AI buildout, because if you can start to finance GPUs at more like 5% or 6%, I think CoreWeave’s lowest financing was in the low sevens. That actually mathematically changes the cost of financing this buildout.

We have this technological innovation that’s going to lower the cost of financing and extend the useful life of compute on Earth. The one last thing that’s interesting about that is my friend Jamin from Altimeter just did a podcast, and Coatue had a deck. They talked about how sellers of scarcity are doing so much better than buyers of scarcity, with buyers of scarcity being the hyperscalers.

But if you own a giant installed base of what is currently in shortage, that’s also a very good place to be. We’re hearing CPUs are way more important than they were in an agentic world. They do all these things around orchestration and tool calls. The biggest CPU fleets in the world sit at the hyperscalers. Some of these hyperscalers may catch up a little bit to the sellers of scarcity.

Patrick O'Shaughnessy

I want to talk about this idea of “different and hard” applied outside of the infrastructure piece of this. Now you’re starting to interact with new founders and existing CEOs and founders who have to adjust to this new world. What are you seeing from the most AI-native founders who aren’t building chips, infrastructure, or models, but are just using this technology to build other things? How do they feel the most different to you, if you’ve observed differences?

7. AI Founders Need A Token Path

Gavin Baker

I do think this isn’t just for chip design. To me, it’s always been a fundamental question for venture. There are different ideas that are obvious to everyone on planet Earth as soon as they hear them. If that’s where you are in venture, if it’s not hard to do, and if it becomes obvious to the world before you have built scale, you’re in trouble. Scale is the ultimate advantage.

The great thing Amazon had was that it was obvious to a lot of people, but it wasn’t obvious to the retail CEOs. Amazon was very smart. Any e-commerce company that VCs invested in, the retail CEOs would destroy. They’d be like, “Oh, that’s so cute. We’re going to take our margins on that to negative 10,000%.”

That’s why the guys at Wayfair did something hard. Amazon tried to kill them, and they failed. Those were tough, operationally competent CEOs. For me in venture, I always look at whether this is going to be obvious to the world before this company can build scale, or whether it is both not obvious and different and really hard to do.

I think a lot of founders are really struggling with this in AI. I think people are becoming worried today that, in Jensen’s 5-layer cake of AI, the profits are accruing to energy, data centers, chips, and models—not really accruing to the applications.

I think Cursor and Cognition got to scale because they focused on coding. Eighteen months ago, the people focusing on coding were Cursor, Cognition, and Anthropic, and it was really right to focus on code.

Amjad Masad, the founder of Replit, tweeted something that I thought was so smart. It was something like, “Bitter Lesson adjacent is the fact that coding might be the shortest path to ASI and useful AI.” Because if you’re really good at coding, you can write yourself code to do anything.

I think it was really smart of those companies to focus intensely on coding. They all probably got to a scale where they have a place. I think Cognition is doing something really, really different. But I think a lot of founders are really struggling, man. I think they’re trying to get confidence that, in niche-ier areas—

Patrick O'Shaughnessy

They won’t get steamrolled.

Gavin Baker

—that they can get to them and get a data moat before the model companies get to that niche, or that it’s a small enough niche that the model companies won’t do it themselves, but it can still produce a venture outcome.

Patrick O'Shaughnessy

Is this related to what you would call the token path? I know you’ve used that phrase with me before.

Gavin Baker

Yeah, I think it comes from a guy at Altimeter, Jamin Ball. He just said that if you’re a software company or an AI company of any kind, you have to be in the token path. So Databricks is in the token path. Compute companies are in the token path. If you’re not in the token path and you’re not in some really niche thing, life may be hard.

Even for these vertical niches, I think if you talk to the people at the model companies, they’re skeptical of some of these because all of the data being generated in these niches comes from humans. But then you’re betting that you’re able to use that proprietary data in this narrow vertical to train a model that’s lower-cost than the frontier labs can ever get to, and maybe that’s a good bet. I just think you have to be very, very careful.

On the other hand, if the returns to these frontier tokens relative to other tokens come down, there’s going to be an explosion in value creation at the application layer. I think another really important point is that I have a belief that, whenever he wants, Jensen can probably get pretty close to the frontier.

Patrick O'Shaughnessy

With his own model.

Gavin Baker

With his own model. I don’t think he wants to do that, but that is what OpenAI and Anthropic are trying to do to him, unsuccessfully. He’s a very logical thinker. This is the logical countermove.

You will see that open-source frontier, which today consists of Chinese models with stolen American tokens. Somebody told me that DeepSeek, maybe the original one, was only 150,000 reasoning traces. There are many ways to launder this if you’re a Chinese company. You can hit all these different APIs and make it hard.

The American labs are working really hard on anti-distillation technology, but I just think Chinese open source is doing really impressive things in a very resource-constrained way, and there’s a lot of distillation. This is why I think, in addition to there not being enough compute to serve Mythos, they did not want it to be distilled. They wanted to use Mythos, distill it themselves, use it to RL their next model, whatever it is.

Eventually, I think OpenAI—anyone on the frontier—will just say there’s going to be some very interesting game theory because it’s a new kind of prisoner’s dilemma. We talked about the old prisoner’s dilemma being around, “Hey, you’re in a prisoner’s dilemma where you have to spend.” The new prisoner’s dilemma is going to be: If you are at the frontier, do you release that model via API or not?

If everyone at the frontier agrees not to do that, then, if one person defects, Chinese open source will have the best model. They’ll have a lot of revenue and cash flow, and then, of course, resources equal intelligence, so they’ll start to pull ahead. That will lead to everybody else releasing it.

It’s a new game theory. It’s kind of the same game theory that you have with Taiwan Semi, Samsung, and Intel. The reality is, if a company like NVIDIA or AMD were ever to really, really use one of these other foundries, that foundry would get better really quickly.

I do think Jensen is going to keep open source a certain timeframe behind the frontier. I think that’s going to be a very interesting thing to watch. And, by the way, open source gets monetized. There’s this misnomer that open source is free. Open-source tokens cost energy to produce, you need to mark them up on GPUs, and the open-source model companies almost always get a revenue share.

8. Investing Through The AI Shift

Patrick O'Shaughnessy

How are you preparing Atreides for the world of Mythos III, Mythos IV?

Gavin Baker

We’re just trying to overinvest in cybersecurity. I really believe everybody needs to have a safe word. Everybody needs to leave their digital devices behind—literally go to the ocean—and have a family safe word or a company safe word. It can’t be one that can be socially engineered.

This is just to avoid cybercrime, where what looks like your son or your daughter or your grandparents or your parents FaceTimes you. It’s an utterly accurate simulation of them. They know everything and can extrapolate based on what they’re likely to say, and then say, “Wire me a million bucks.” So we’re doing everything we can with cybersecurity.

Patrick O'Shaughnessy

That’s defensive. What about analytical or processing? What will you still be able to do that it won’t be able to do, I guess, on the analytical side?

Gavin Baker

It’s a good question. I just watched The Last Samurai, and I asked people at my firm to watch it. If you haven’t seen it, I highly recommend watching it. It’s actually a movie that’s aged really well. It’s a Tom Cruise movie from 20 years ago.

The conceit is that Tom Cruise is this bitter, washed-up Civil War veteran who’s actually a very good soldier, and he’s bitter and washed up because he feels like he participated in negative actions against Native Americans. It’s during the Meiji Restoration, and he’s hired by the modern elements of the Japanese government to train an army of peasants how to fight the samurai.

There’s a first battle. Of course, the samurai win even though they don’t have guns. He fights valiantly, so the samurai decide not to kill him and take him to their village. He becomes a samurai. It feels like the Civil War to him, so he fights on the side of the samurai. At the end of it, he’s massacred by a peasant with a machine gun.

The machine gun is here. If we do not all become masters of the machine gun, we’re going to get mastered. I am trying to become a master of the machine gun, and I’m optimistic there’s a long period of time where, just like if you were a 50-year-old samurai veteran of many wars, "I've fought many wars, Master Dwarf," you will have advantages using the machine gun.

I’m optimistic that, as a lifelong student of investing, I’m going to be able to master the machine gun, this new technology, and integrate it into my own process and our firm’s process in ways that let me contribute value as a human being for a long time. Like everyone, I have agents running all the time now.

Patrick O'Shaughnessy

What’s your most useful agent?

Gavin Baker

My single most useful agent is a really good summary of the points that would be interesting to me from podcasts. There are 6 hours a day of stuff that I feel like is in my job description to watch. Every time somebody from OpenAI, xAI, Google, Cursor, Fireworks, or Base10 speaks—to say nothing of Jensen, Elon, or Dario—I feel compelled to watch, and I just don’t have that much time.

There are some real needles in haystacks. That is, for me, the most useful. I do think there’s a set of things that I always like to see. I’m very sensitive to management compensation. What are they incentivized to do? Do they just have stupid RSUs, or do they have PSUs? And if they have PSUs, what do those PSUs incentivize them to do?

We now have systems that do a very good first pass at that. That saves people a lot of time. It frees them up for more creative work than going through the proxy, pulling the PSU thing, and looking at how it’s changed versus all the proxies. There’s signal in that, but it’s very labor-intensive, and it’s so good for AI.

There are obviously all sorts of the same things within investing. Pressuring the organization in those ways, I think, has been helpful. This is the most exciting, thrilling time to be an investor. I’m getting a little bit worried.

Patrick O'Shaughnessy

The diversity breakdown thing?

Gavin Baker

Yeah.

Patrick O'Shaughnessy

Say just a little bit more about the kinds of people who are capitulating.

Gavin Baker

I don’t know anyone like me who’s not really bullish on DRAM. There are all these interesting things happening with AI right now. One is that, cross-sectionally, the valuations do not make sense. They just flat-out do not make sense. They cannot all be true.

In other words, you have semicap equipment companies trading at 40 times next quarter’s annualized earnings and DRAM companies trading at mid-single-digit multiples. At the peak of the last cycle, that was 5 versus 12. At one point, it was 3 versus 45. Those can’t both be true.

Yes, semiconductor capex business models have improved more than the memory business models. We don’t know how much HBM is going to improve memory business models yet. Yes, they have some element of recurring revenue with parts and maintenance, but it’s not worth a 1,000% multiple gap.

I think it’s hard to square the valuation of something like NVIDIA, which, still in early April, was essentially as cheap as it gets relative to the market—like in the last 10 or 12 years, or whatever it is—and very cheap in absolute terms, with something like GE Vernova’s valuation, because it builds in an unfathomable amount of share loss for NVIDIA.

So valuations cross-sectionally are really different. Because we are in shortages, the lowest-quality companies are doing the best.

So if you’re an oil and gas investor, a mining investor, or a natural-resources investor, and you’re well-versed in thinking about costs, this is very intuitive to you. In a real bull market for a commodity, the commodity suppliers with the highest costs go up the most because it’s the most beneficial to them. They go from being on the verge of bankruptcy to gushing cash.

This is, I think, one reason commodity investing is really, really hard, because quality outperforms during the cycles, but you get all of the outperformance during the downturns, when the high-cost guys that mooned during the shortages and the commodity bull markets go bankrupt or whatever. You’re seeing that happen in every industry. The lowest-quality players—companies that are hated and detested by the hyperscalers and the buyers because they have high costs, they’re unreliable, and their parts fail at a high rate—are sold out and raising prices.

That activity gets the interest of these retail accounts on X, and these stocks get bid to the moon, whereas some of the higher-quality expressions have actually really underperformed. As an investor, it’s hard because you know without a shadow of a doubt that the thing that’s mooned 10X in 3 months or 6 months is going to go right back down, subject to what they do with all the cash. It worries me a little bit that people who were very skeptical a year ago are no longer skeptical. But then I contrast that with the valuations of these high-quality companies, which are just not extended, and it makes me feel better.

I always thought it was funny in 2024 and 2025 that anyone asked about an AI bubble or talked about it. You have this nuclear bubble and this quantum bubble right here, right in front of you. What are we talking about? This is so real. Some of that nuclear and quantum silliness has maybe spread into more speculative, lower-quality, smaller-cap names, where if you have a big presence on X or Reddit, it’s easy to move them, and that frightens me a little bit.

But I just wish there were more AI bears. I wish there were more memory bears. Astera is a stock I’ve been close to for a long time. There are a lot of bears on that. I love that. I first invested in the Series C. Good luck thinking that’s a copper loser.

You can also feel the baskets in the market and the leverage baskets. Which baskets you’re in is really important: copper, optical, DRAM, NAND. A very interesting thing that’s happened this year is that in 2024 and 2025, the AI trade traded together. You could be long GPU compute, scale-up networking, and optical scale-out, and short power or whatever it was. That trade worked from a risk-management sense because I am very factor-aware.

That all blew out in January of this year. Scale-up networking would go crazy while scale-out was going down, or DRAMs would massively underperform NAND and HDDs, which had not happened. These cross-sectional correlations within AI really fell apart, and you had to get very fine-grained. You couldn’t hedge your memory anymore with some semiconductor-capital-equipment or NAND exposure. Everything cross-sectionally really changed in a very interesting way in January.

I think maybe one reason for that was that AI got to a point where it was suddenly really easy for a bunch of people to get really smart on these different subsectors, start trading them, and then put them into baskets. Those baskets influence—

Patrick O'Shaughnessy

AI creating price efficiency, yeah.

Gavin Baker

Yeah, exactly. I think some of the biggest opportunities outside of these higher-quality names that I think can compound for a long time—and that are safe, unlike these low-quality names, which are terrifying—are in names that are miscategorized. Astera was in a lot of copper-loser baskets. Astera’s biggest product is going to be a switch. You use both copper and optics to connect switches to accelerators. Definitionally, if you’re a switch company or an accelerator company, you cannot be a copper loser because you’re going to be on the other side of that connection.

Patrick O'Shaughnessy

I wonder if you could riff for a sentence or two on each of the major public companies—Google, Microsoft, Amazon, the major players. All the conversation is centered around these exciting new companies. Maybe run through them and riff.

Gavin Baker

Google was incredible last year because they had that TPU advantage, which is now gone. The reason I think they’re still in a great position is that they have the most compute of anyone. We talked about the value of installed bases being higher as a result of shortages. They have the biggest installed base of compute.

Google I/O is this week. If they don’t release something that even slightly leapfrogs OpenAI and/or Claude, that’s interesting. It’s not a disaster for Google; it’s just interesting, and it means this NVIDIA effect we discussed is even more powerful than I might have imagined. I’m very curious to see what the Pareto frontier looks like literally 5 days after Google announces its new stuff. This is a big card for them.

But Google, between the amount of data they have and the fact that YouTube data is actually really genuinely valuable, is in a great position. It is valuable in a world of robotics. They have the amount of compute they have and the search business they have. Google’s never not going to be in a good position, and then you see that with GCP going crazy.

You’ve got to give Zuckerberg immense credit for what he’s done in terms of making Meta an AI-first company internally. I do think he is the only one of those true internet giants to have done that. I give him a lot of credit for that. I also give him a lot of credit for paying up when he did for buying out contracts for that talent. Muse, I think, was a really big upside surprise. It was the first model from MSL, and it’s not on the Pareto frontier with xAI, Google’s one entrant, and then OpenAI and Claude, but it’s pretty close. That was very impressive to me.

I think Meta is in a better position. It’s still not as strong of an absolute position as Google, but they’re in a better position, and rates of change matter more than level, as you know, in markets, particularly over short 3-year time frames. Over long time frames, the level of competitive advantages tends to dominate, but even within that, changes really matter.

Amazon, I think, is in a really strong position because of Trainium. I do think you’re going to see real P&L efficiencies from robotics over the next 18 months in their retail business. I actually think Nova, their internal models, are not where Muse is, but they’re better than they get credit for.

Then Microsoft: I like Satya; I admire him. I think he’s an exceptional CEO, and I give him a lot of credit for the decisions he’s made. But he did go from, “We’re going to make Google dance,” to being the product manager of Copilot in about 3 years. I would love to know, during the coup attempt against OpenAI, whether Satya regrets his decisions. Does Satya wish that he had supported Ilya instead of Sam, and that Ilya and Mira were really running OpenAI today? In his heart of hearts, I would love to know, because I think the Microsoft-OpenAI partnership might look very different in that world. I think that’s a very interesting question that we’ll never know the answer to.

But I give him a lot of credit. What he is doing now is taking risk. This goes to the decisions you have to make in that cone of uncertainty: not only how much you spend, but what you’re going to spend it on. I think Microsoft flinched for a moment in early 2025. They have this algorithm: we spend this much in CapEx dollars, and we get this return. That algorithm was off. If you flinch, you lose position. You lose all these allocations, and it’s difficult to get them back.

They flinched, and now the decision Satya is making—which the market has punished him for, but I think is the right decision—is to use their compute internally to make their own products better. Who knows how fast Azure could be growing if they were willing to just sell GPUs to OpenAI? One reason Copilot has been so bad is that there’s simply not enough compute available. They’re fixing that. He’s the product manager of Copilot.

I do think he’s a great CEO. They’re trying to use their compute to train their own models. I am a little skeptical that they have the right team to succeed there, but just like Meta, they can afford to hire maybe a different team. I think he’s making good, risky decisions to position Microsoft for this world where frontier models are no longer API-accessible. I think it’s a really courageous decision that I give him a lot of credit for.

Microsoft would probably be an $800 stock today if they were using their GPUs solely to serve OpenAI and Anthropic’s capacity instead of using them for their own products. So I give him a lot of credit for making a great decision.

I think what’s really interesting is the degree to which these companies are outward-facing in their decisions. The 2 companies that are the most deeply engaged with startups are Amazon and NVIDIA, by a mile. There’s really intense engagement with Google; they’re the next most intense. Broadcom is engaged in a different way. They’re just everybody’s favorite ASIC supplier.

If you’re a startup, it’s considered a level up if you get to work with Broadcom for your second-generation chip, and it’s considered manna from heaven if Broadcom works with you for your first-generation chip. Then you see essentially zero engagement with startups from AMD, Microsoft, and Meta. When I say zero, I mean there’s a little, and I just wonder about that decision, because some of the best teams are no longer at big public companies; they’re at these smaller startups.

I think it’s going to end up being a pretty big advantage for NVIDIA, with AMD and Google right behind them, to have this engagement that you just don’t see from these other hyperscalers.

9. AI Raises The Stakes For Everyone

Patrick O'Shaughnessy

As we wrap up, I'm curious for you to riff on any other knock-on effects that you've started to think about for this giant trend. We've talked about the specific companies that this most impacts in a lot of detail. We've talked a little bit about the application layer and what would have to happen for there to be more value accruing to that layer of the stack. I'm curious about any other fun knock-on things that you've been thinking about as this world changes so quickly.

Gavin Baker

It is wild. At the application layer, forget value accruing; just value has been destroyed. AI has net destroyed trillions of dollars of value at the application layer, even if you count Cursor and Cognition, the most successful AI natives. The companies that are doing the best today, that are seeing their values increase the most and are creating economic value, are the companies with the highest effective ratio of utilized GPUs per human. Maybe this just means that every human is going to get a lot of GPUs, but I think that's an interesting fact that we need to be cognizant of.

I will just say, and maybe this is a little dark, that I am more and more worried about personal safety. I worry about this a lot more for people who have a much bigger public presence and are much more associated with AI. I hope nothing tragic happens. There is this upsurge in political violence here in America, and as AI increasingly becomes political, I worry that's going to get directed at more and more AI political leaders. Whatever I may think or may not think of OpenAI, I think it is terrible that someone threw Molotov cocktails at Sam Altman's house.

I am worried that we are headed into a higher-variance, higher-beta, higher-risk world because of AI. That's true for me as an individual, and then for people who are big players on the chessboard. Think about what it means geopolitically. We're watching the Ukrainians really start to win, and the reason they're winning, I think, is not really because they have better drones. I think they do have better drones; that's part of it. I think the reason Ukraine is really winning is that they have the best battlefield AI outside of probably America and Israel.

As China and our adversaries begin to process that, how do they respond? Because of its edge in AI, the United States is in a great position, but it is destabilizing for the rest of the world. Something I think a lot about is creating a charity to educate the world on how awesome the West has been. Slavery was endemic to essentially almost every civilization, and slavery was really ended by the British Empire. Tell that story.

But America, after 1945, had the nuclear bomb, and no one else had it. We could have controlled the world forever. Instead, we rebuilt Germany and Japan, who are America's most reliable allies, as well as Israel and South Korea. That's a testament to the American spirit in our country. We didn't take over the world. There were fears documented at the time that the American generals—MacArthur was a little bit of an American emperor in Japan—were just going to take over the world. They could have, and they didn't. They came home, we demilitarized, and then you had this period of great global stability, despite the fact that there were terrible moments.

Patrick O'Shaughnessy

Pax Americana.

Gavin Baker

Yeah, you had the Pax Americana. So maybe it's not destabilizing. Maybe it leads to another Pax Americana informed by our AI dominance, and I'm so optimistic that AI is going to be amazing for the world.

There's someone like me whose daughter was diagnosed with a very rare disease. There's no cure. He was able to assemble a lot of resources. He was able to get a lot of compute from the labs. We were made aware of what was happening, spun up an immense number of agents, and came up, using AI, with a drug on the market that can actually impact his daughter's disease. He has since spun up a company to cure it. Her life is already immeasurably different because of AI.

So I'm an AI optimist and maximalist, but I also acknowledge that it's an event horizon. It is, for sure, going to be a discontinuity we need to navigate as a society. I think the Luddites are going to be wrong, but we need to be really thoughtful in how we address their concerns. We need to make sure that it's good for everyone. It is a little dystopian that now the best AI is only available to people with a lot of money. We need to solve that. We need to approach this with humility, recognize that there's a lot of uncertainty, and be thoughtful.

Patrick O'Shaughnessy

When I do this with you, I tell people afterward, "May you find something that you love as much as Gavin loves markets and companies and capitalism and history on display today, as always." Gavin, thanks so much for your time.

Gavin Baker

Thank you. Thanks, Patrick.