Dwarkesh Podcast · · 11 min
Why smarter AI models could drive up compute prices 10x
TL;DR
- Dwarkesh's core call: if Anthropic keeps 10x'ing revenue year over year (ending last year at $9B, with a projected $100–150B this year, implying $1T by the end of next year if the trend continues) while lab compute only 3x's, the surplus must show up somewhere — and he sees rising compute prices as the remaining escape valve, with everyone in the stack below the labs capturing the gains.
- Basically all three escape valves are already happening: Anthropic's inference margins reportedly went from 40% mid-last-year to upwards of 80% for Fable, GPU spot prices are more than 40% higher than the February trough, and OpenAI's inference share of compute was a quarter (2024, per Epoch) and is likely closer to ~50%, if not higher, now.
- The tradeable anchor: a human-level software engineer running on an H100 equivalent should rent for over $250K a year — over 15x the current spot price — before counting nights and weekends. He questions the "more engineers = lower wages" objection, invoking the lump of labor fallacy; if standard economics holds, he says, "the marginal value of compute should stay astonishingly high."
- Frontier-tranche compute already trades rich: Google is paying SpaceX $900M a month for 110,000 GB200/GB300 GPUs — 2x spot — on top of a spot price itself more than 40% higher than in February.
- Supply can't respond: the 3x decomposes into 1.4x Moore's Law ("a miracle if we can just keep it going"), 1.2x new fabs bottlenecked by ASML EUV up to 2030 and potentially beyond, and 1.8x wafer reallocation that probably hits a wall by the end of next year as AI's share of TSMC leading-edge N3 goes from 60% to 86%.
- Second-order implications: efficient frontier models could charge much larger premiums (Alchian-Allen effect), many current popular AI applications will probably get priced out by AI-research demand for tokens, and challengers may struggle to outbid labs that monetize compute better — with a stated caveat that post-singularity, robot-built chips make compute cheap again.
Digest · the substance, structured for research
1. The arithmetic that forces the whole thesis
- The setup: Anthropic's revenue has 10x'd three consecutive years — $9B at the end of last year, $100–150B likely this year, $1T by the end of next year if the trend holds. He flags the hedge himself: "there's no deep reason why this has to be true… it's ultimately a question of AI capabilities."
- Against that, lab compute only 3x's year over year — so margins rise, compute prices rise, or inference share rises. "Basically all three of these things are already happening": Anthropic's inference margins reportedly went from 40%→80%+ for Fable, spot prices are more than 40% higher since February, and OpenAI inference share was a quarter in 2024 per Epoch and is likely closer to ~50%, if not higher, now.
2. Labs don't want the inference-heavy path — and 90%+ margins remain uncertain
- Spending most compute on inference means "you're basically declaring that AI progress has stalled and you're just now in the business of being a cloud provider" — labs believe models within a year make current ones "look extremely shitty," so training must keep the majority.
- He finds it "really wild" that margins on intelligence exceed 90% without being competed away — leaving one escape valve: compute prices rise, and "everybody in the stack below the lab gets the surplus."
3. Smarter models make the same GPU worth 15x more
- The load-bearing example: a true human-level software engineer on an H100 equivalent justifies $250K+/year rent, over 15x spot, before nights and weekends. Case study: Google pays SpaceX $900M/month for 110K GB200/GB300s at 2x spot, reflecting frontier labs' needs for scale, efficiency, flexibility, and weight security.
- His framing of the demand-glut objection: applied to people, this is "the classic lump of labor fallacy" — economists generally believe high-skill immigration doesn't depress wages long-run. Hedged: "maybe this labor supply shock will be so big and so fast that we can't count on this general heuristic anymore."
4. Expensive compute rewards efficiency and prices out slop
- The Alchian-Allen effect: at $20/hour for an H100, "it would be extremely stupid to use a weaker, less efficient model" — a model that gets the same result with less compute has "in some sense created more compute," so efficient frontier models could charge much larger premiums.
- Consequence: a lot of current popular AI applications will probably get priced out. Google or Anthropic or OpenAI will be willing to pay more for tokens to "automate AI research" than you or I will be willing to pay to make more AI slop talk.
5. Why supply can't answer: the 3x is fragile, not expandable
- He pre-empts the Simon–Ehrlich analogy — Ehrlich's Malthusianism famously lost to ingenuity — but "I'm guessing that the analogy to this bet is probably wrong": compute supply is much less elastic, less able to absorb large demand shocks, and less able to use substitutes than metal extraction, and Ehrlich might well have won in other decades.
- The decomposition: 1.4x Moore's Law (keeping it going would be "a miracle"), 1.2x new fabs bottlenecked by ASML EUV machines up to 2030 and potentially beyond (citing Dylan's earlier episode), 1.8x wafer reallocation from phones/PCs — which probably hits a wall by the end of next year as AI's share of TSMC leading-edge N3 goes from 60% to 86%.
6. Caveats that frame the trade
- This is explicitly the "pre-singularity regime" — eventually robots converting silica and copper into chips make compute cheap again.
- The 10x-revenue-on-3x-compute gap itself shows how strong economies of scale in the model business are: one-time training cost shared across all users, unlike human labor "retrained from scratch." His discomfort, verbatim: "I wish we didn't live in a world with such strong economies of scale for intelligence, because I'm worried about power concentration, but it seems we do."