[BidClub_]
Dwarkesh Podcast · · 11 min

Why smarter AI models could drive up compute prices 10x

Dwarkesh Patel

YouTube
TL;DR
  • Dwarkesh's core call: if Anthropic keeps 10x'ing revenue year over year (ending last year at $9B, with a projected $100–150B this year, implying $1T by the end of next year if the trend continues) while lab compute only 3x's, the surplus must show up somewhere — and he sees rising compute prices as the remaining escape valve, with everyone in the stack below the labs capturing the gains.
  • Basically all three escape valves are already happening: Anthropic's inference margins reportedly went from 40% mid-last-year to upwards of 80% for Fable, GPU spot prices are more than 40% higher than the February trough, and OpenAI's inference share of compute was a quarter (2024, per Epoch) and is likely closer to ~50%, if not higher, now.
  • The tradeable anchor: a human-level software engineer running on an H100 equivalent should rent for over $250K a year — over 15x the current spot price — before counting nights and weekends. He questions the "more engineers = lower wages" objection, invoking the lump of labor fallacy; if standard economics holds, he says, "the marginal value of compute should stay astonishingly high."
  • Frontier-tranche compute already trades rich: Google is paying SpaceX $900M a month for 110,000 GB200/GB300 GPUs — 2x spot — on top of a spot price itself more than 40% higher than in February.
  • Supply can't respond: the 3x decomposes into 1.4x Moore's Law ("a miracle if we can just keep it going"), 1.2x new fabs bottlenecked by ASML EUV up to 2030 and potentially beyond, and 1.8x wafer reallocation that probably hits a wall by the end of next year as AI's share of TSMC leading-edge N3 goes from 60% to 86%.
  • Second-order implications: efficient frontier models could charge much larger premiums (Alchian-Allen effect), many current popular AI applications will probably get priced out by AI-research demand for tokens, and challengers may struggle to outbid labs that monetize compute better — with a stated caveat that post-singularity, robot-built chips make compute cheap again.
Digest · the substance, structured for research

1. The arithmetic that forces the whole thesis

  • The setup: Anthropic's revenue has 10x'd three consecutive years — $9B at the end of last year, $100–150B likely this year, $1T by the end of next year if the trend holds. He flags the hedge himself: "there's no deep reason why this has to be true… it's ultimately a question of AI capabilities."
  • Against that, lab compute only 3x's year over year — so margins rise, compute prices rise, or inference share rises. "Basically all three of these things are already happening": Anthropic's inference margins reportedly went from 40%→80%+ for Fable, spot prices are more than 40% higher since February, and OpenAI inference share was a quarter in 2024 per Epoch and is likely closer to ~50%, if not higher, now.

2. Labs don't want the inference-heavy path — and 90%+ margins remain uncertain

  • Spending most compute on inference means "you're basically declaring that AI progress has stalled and you're just now in the business of being a cloud provider" — labs believe models within a year make current ones "look extremely shitty," so training must keep the majority.
  • He finds it "really wild" that margins on intelligence exceed 90% without being competed away — leaving one escape valve: compute prices rise, and "everybody in the stack below the lab gets the surplus."

3. Smarter models make the same GPU worth 15x more

  • The load-bearing example: a true human-level software engineer on an H100 equivalent justifies $250K+/year rent, over 15x spot, before nights and weekends. Case study: Google pays SpaceX $900M/month for 110K GB200/GB300s at 2x spot, reflecting frontier labs' needs for scale, efficiency, flexibility, and weight security.
  • His framing of the demand-glut objection: applied to people, this is "the classic lump of labor fallacy" — economists generally believe high-skill immigration doesn't depress wages long-run. Hedged: "maybe this labor supply shock will be so big and so fast that we can't count on this general heuristic anymore."

4. Expensive compute rewards efficiency and prices out slop

  • The Alchian-Allen effect: at $20/hour for an H100, "it would be extremely stupid to use a weaker, less efficient model" — a model that gets the same result with less compute has "in some sense created more compute," so efficient frontier models could charge much larger premiums.
  • Consequence: a lot of current popular AI applications will probably get priced out. Google or Anthropic or OpenAI will be willing to pay more for tokens to "automate AI research" than you or I will be willing to pay to make more AI slop talk.

5. Why supply can't answer: the 3x is fragile, not expandable

  • He pre-empts the Simon–Ehrlich analogy — Ehrlich's Malthusianism famously lost to ingenuity — but "I'm guessing that the analogy to this bet is probably wrong": compute supply is much less elastic, less able to absorb large demand shocks, and less able to use substitutes than metal extraction, and Ehrlich might well have won in other decades.
  • The decomposition: 1.4x Moore's Law (keeping it going would be "a miracle"), 1.2x new fabs bottlenecked by ASML EUV machines up to 2030 and potentially beyond (citing Dylan's earlier episode), 1.8x wafer reallocation from phones/PCs — which probably hits a wall by the end of next year as AI's share of TSMC leading-edge N3 goes from 60% to 86%.

6. Caveats that frame the trade

  • This is explicitly the "pre-singularity regime" — eventually robots converting silica and copper into chips make compute cheap again.
  • The 10x-revenue-on-3x-compute gap itself shows how strong economies of scale in the model business are: one-time training cost shared across all users, unlike human labor "retrained from scratch." His discomfort, verbatim: "I wish we didn't live in a world with such strong economies of scale for intelligence, because I'm worried about power concentration, but it seems we do."
Dwarkesh Patel

Today I want to talk about what the compute situation for the labs will look like over the next few years. For the last 3 consecutive years, Anthropic's revenue has 10x'd year over year, and it's likely to do so again this year. They ended last year with $9 billion in revenue. I think they'll probably end this year with somewhere between $100 billion and $150 billion in revenue.

Now, for this trend to continue, Anthropic would need to make $1 trillion in revenue by the end of next year. Of course, there's no deep reason why this has to be true. It's a very wild conclusion, and it's ultimately a question of AI capabilities: Does AI get that useful by the end of next year? But suppose the trend does continue. I want to think through what happens in that world.

1. Compute Revenue Creates a Gap

The other big trend in AI is that lab compute only 3x's year over year. For a lab to keep 10x'ing revenue year over year while compute only 3x's, one of the following 3 things needs to happen, or some combination of the 3. One, lab margins have to increase. Two, the price of compute has to increase. Or three, the percentage of compute that labs spend on inference rather than training has to increase.

My understanding is that all 3 of these things are already happening. With regards to the margins, Anthropic's inference margins reportedly went from 40% in the middle of last year to upwards of 80% now for Fable. With regards to compute, the spot prices for compute are more than 40% higher than they were in the February trough that we had earlier this year. And with regards to the share of compute that goes to training versus inference, in 2024, according to Epoch, OpenAI was spending just a quarter of its compute on inference, and that number is likely closer to 50%, if not higher, now.

Now, labs would prefer not to do this final thing of increasing the share of compute they spend on inference. The way the labs see the world, the whole point of inference revenue is to help convince investors to give you more money in order to train the next bigger, better model. And if you're spending most of your compute on inference, then you're basically declaring that AI progress has stalled and you're just now in the business of being a cloud provider.

This is a less compelling business than building AGI, so the labs do not want to be in this business, nor do they think they are in this world. They think that within a year, they'll have built models that make the current ones look extremely shitty. But they need to invest a lot of their compute—the majority of their compute—into doing the training and experiments that are necessary to build the next model.

So that leaves only 2 options for how you can get out of this gap between the fact that lab compute only increases 3x year over year, but revenue increases 10x. Either the lab's margins have to increase so that they get the surplus, or the price of compute has to increase so that everybody in the stack below the lab gets the surplus.

It's not clear to me which world we end up in. Do we end up in a world where we go from 80% for some of the top models to greater than 90% margins if the lab margin effect dominates? That would require the leading model to be so far ahead of the competition, because the nature of margins—why they exist in a market economy—is that the thing you are serving is so much better than what somebody else could go get and replace you with on the market. But it's just really wild for me to consider that the margins for something like intelligence will be greater than 90% and they don't get competed away at that level.

2. Compute Prices Must Rise

So that leaves only one other possibility of this escape valve between these 2 trends, which is that the price of compute has to increase. As I mentioned, this is already starting to happen. And the effect is even stronger when you look at the tranche of compute that the frontier labs actually need to accumulate, because they can't just go out and buy a spot instance.

They need to make sure that they get enough scale to get really good efficiency and flexibility, and also that they have the kind of compute that lends itself to the security they need for their own weights and for their customers' information. I think a relevant case study here is to look at the compute that Google and Anthropic are renting from SpaceX. Google, for example, is paying $900 million a month for 110,000 GPUs that are a blend of GB200s and GB300s. The price that Google is paying here is 2x the spot price per hour for those GPUs. And that spot price itself is more than 40% higher than it would have been in February.

3. Smarter Models Monetize Compute

I want to emphasize a key conclusion here: As AI models get smarter, they will be better able to monetize the same amount of compute. If a true human-level software engineer could run on an H100 equivalent, then at today's prices for software engineers, that H100 should rent for over $250K a year. That's over 15x the current spot price for an H100. And this is not even accounting for the fact that your AI can work nights and weekends.

Of course, you might expect that if we had 10 million extra software engineers suddenly appear in the economy, the marginal value of a software engineer would decrease, and thus the revenue that that H100 would be able to generate would not be 15x higher than it is right now. But I actually don't know if this is true. If we apply this argument to people instead of AIs, then this would be the classic lump of labor fallacy.

For example, economists generally believe that high-skill immigration does not decrease wages in the long run because of how innovation and specialization increase the value of labor. Maybe this labor supply shock will be so big and so fast that we can't count on this general heuristic anymore. But if you believe what standard economics says, then the marginal value of labor, and thus the marginal value of compute, should stay astonishingly high.

4. Scarcity Rewards Efficiency

So let's think about what changes in such a world. One of the things that would happen is that as the top labs get better and better at monetizing compute, and the cost of compute increases, it becomes harder for anybody else to compete against them, because they have to bid for this resource against somebody who is basically able to make better use of it.

Another thing that will happen—and I think this is actually the most interesting implication of this whole thought exercise—is that if you can train the best, most efficient model, then you'll be able to charge much higher margins than you can today. This is the Alchian–Allen effect in economics, and what it's basically saying is that if it costs $20 an hour to rent an H100, then it would be extremely stupid to use a weaker, less efficient model, because it's gonna burn more tokens on your expensive compute to get the exact same result.

So labs will be able to charge a much larger premium if they can train a model that better economizes this scarce input. Basically, if you have a model that can get the same result by using less compute, then you've, in some sense, created more compute, and the value of compute is gonna increase.

Another thing that will happen is that a lot of current popular applications of AI will probably get priced out. The reason AI is relatively cheap right now is that AI just can't do a lot of things that top humans can do. But this, at some point, will no longer be the case. And at that point, Google or Anthropic or OpenAI will be willing to pay more for the tokens to automate AI research than you or I will be willing to pay to make more AI slop talk.

5. Compute Scarcity May Persist

I'm a bit worried that this kind of analysis honestly pattern-matches a lot onto the ways that people in the past have been wrong about scarcity. I'm thinking, for example, of the famous Simon–Ehrlich bet. Paul Ehrlich was this famous doomer about population growth, and he made this bet that a basket of commodities would increase in price rather than decrease in the decade preceding 1990.

This is a very famous bet because it's supposed to illustrate how Ehrlich's Malthusian worldview was wrong, and how he did not anticipate the way in which market signals and human ingenuity can find better ways to economize scarce inputs. I'm guessing that the analogy to this bet is probably wrong. Other analysis has shown that if that bet had been made in a different decade, Ehrlich might well have won.

But more generally, I think the supply of compute is much less elastic, much less capable of absorbing large demand shocks, and much less capable of being accommodated by using different substitutes than the extraction of different metals is.

To illustrate why I think this 3× in compute capacity year over year is hard to budge or potentially even sustain, I don't see how any of the 3 elements that constitute that 3× can be much accelerated. 1.4× of that is coming from Moore's Law. Far from increasing it, I think it'll be a miracle if we can just keep it going for a few more years.

1.2× is coming from building new fabs. This process is ultimately gonna be bottlenecked up to 2030 and potentially even beyond by just building new ASML EUV machines. Dylan, when he was on the podcast a few months ago, talked about this in great detail.

And 1.8× comes from the fact that AI is absorbing a lot of wafer allocation that was previously going to smartphones and PCs. This is probably gonna hit a wall by the end of next year, when at the leading-edge N3 nodes at TSMC, AI will have gone from 60% to 86%. At some point, you have just absorbed all leading-edge wafer capacity for AI, and you can't keep increasing this number.

So I don't know how we even continue to do 3× compute scaling year over year for the next few years, much less go beyond that.

6. Cheap Compute Comes Later

Now, I wanna clarify that at some point in the future, compute will get cheap again. At some point, we'll just have robots that can convert shores of silica sand and mines of copper into new computer chips, and then the price of compute is basically the raw inputs and the tools required to do this processing.

I'm just talking about this current pre-singularity regime where AI compute merely 3×'s year over year, which is not enough to offset how much more valuable AI is becoming over time. By the way, the fact that Anthropic's revenue has been 10×'ing year over year, whereas their compute has only been 3×'ing year over year, I think illustrates how strong the economies of scale are in the model business.

And logically, this makes sense. When you train a model, you just have to spend this one-time cost to learn all these different skills that then get to be shared across all your users. This is very unlike human labor, where each instance has to be retrained from scratch.

I wish we didn't live in a world with such strong economies of scale for intelligence, because I'm worried about power concentration, but it seems we do.

Okay, this was a narration of a blog post that I also released on my website at dwarkesh.com. Check it out for other posts or to be notified when I release a post in the future. Otherwise, I'll see you for the next full episode.

Why smarter AI models could drive up compute prices 10x | BidClub