[BidClub_]
Dwarkesh Podcast · · 11 分钟

为什么更聪明的 AI 模型可能将算力价格推高10倍

Dwarkesh Patel

YouTube
TL;DR
  • Dwarkesh的核心判断是:如果 Anthropic 的收入继续按年增长10倍(去年年底为90亿美元,预计今年达到1000亿至1500亿美元;若趋势延续,明年年底将达到1万亿美元),而实验室算力仅增长3倍,那么这部分盈余必然要在某处体现,而他认为剩下的出口就是算力价格上涨,实验室以下的整个产业链都将捕获这部分收益。
  • 3个出路基本都已出现:Anthropic的推理利润率据报从去年年中40%升至 Fable 的80%以上,GPU现货价格较2月低点高出40%以上,而 OpenAI 用于推理的算力占比按Epoch口径在2024年为四分之一,如今可能更接近约50%,甚至更高。
  • 一个运行在 H100 等效算力上、达到人类水平的软件工程师,在不计夜间和周末的情况下,年租用成本应超过25万美元——超过当前现货价格的15倍;他对“工程师越多、工资越低”这一异议提出质疑,援引劳动总量谬误;如果标准经济学成立,他说,“算力的边际价值应该始终高得惊人”。
  • 前沿档位的算力已经在高价交易:Google每月向SpaceX支付9亿美元,租用110,000块GB200/GB300 GPU,价格为现货价的2倍;而现货价本身也较2月高出40%以上。
  • 供给无法回应这一需求:3倍增长可拆解为摩尔定律贡献的1.4倍、新建晶圆厂贡献的1.2倍(受ASML EUV光刻机制约,瓶颈将持续至2030年甚至更久),以及晶圆再分配贡献的1.8倍;随着AI在TSMC先进制程N3中的占比从60%升至86%,后者大概率在明年年底撞上上限。
  • 二阶影响包括:高效的前沿模型可能收取大得多的溢价(阿尔钦—艾伦效应);AI研究对token的需求可能把许多当下流行的AI应用挤出价格区间;而能更有效实现算力变现的实验室,可能让挑战者难以出更高的价——但需要说明的是,奇点之后,机器人制造的芯片会再次让算力变得便宜。
摘要 · 为研究而整理的核心内容

1. 迫使整套论点成立的算术

  • Anthropic的收入已连续3年每年增长10倍:去年年底达到90亿美元;今年很可能达到1000亿至1500亿美元;如果趋势延续,明年年底将达到1万亿美元。他本人也主动提示了这一判断的风险:“没有什么深层理由说明这一定为真……归根结底,这是AI能力的问题。”
  • 与之相比,实验室算力同比仅增长3倍,因此要么利润率上升,要么算力价格上升,要么推理占比上升。Dwarkesh说:“这3件事基本都已经在发生”:Anthropic的推理利润率据报从去年年中的40%升至 Fable 的80%以上;GPU现货价自2月以来高出40%以上;按Epoch的数据,OpenAI用于推理的算力占比在2024年为四分之一,如今可能更接近约50%,甚至更高。

2. 实验室不想走推理密集型路径——90%以上的利润率仍不确定

  • 如果把大部分算力都花在推理上,就等于“宣布AI进步已经停滞,而你现在只是在做一家云服务商”。实验室相信,1年内的新模型就会让当前模型“烂得离谱”,因此训练仍必须占据算力的大头。
  • 他觉得,智能业务的利润率在没有被竞争压平的情况下超过90%“非常离谱”,这只留下了一条出路:算力价格上涨,实验室下游产业链上的所有人都拿走这部分盈余。

3. 更聪明的模型让同一块GPU价值提升15倍

  • 这个例子是整套论证的支点:一个真正达到人类水平、运行在H100等效算力上的软件工程师,年租用价格应超过25万美元,且还没算夜间和周末——超过现货价格的15倍。一个现实案例是,Google每月向SpaceX支付9亿美元,以现货价2倍的价格租用110,000块GB200/GB300 GPU,反映出前沿实验室对规模、效率、灵活性和模型权重安全的需求。
  • 针对“工程师供给增加会压低工资”这一质疑,他的框架是:如果把同样的逻辑套到人身上,这就是经典的“劳动总量谬误”;经济学家普遍认为,高技能移民长期并不会压低工资。但他也保留了条件:“也许这次劳动力供给冲击会大到、快到让我们再也不能依赖这条一般经验。”

4. 昂贵的算力奖励效率,也会把低质应用挤出市场

  • 这体现了Alchian-Allen效应:当H100算力价格为每小时20美元时,使用更弱、更低效的模型会“蠢到极点”;如果一个模型用更少算力取得同样结果,它在某种意义上就创造了更多算力,因此高效的前沿模型可能收取大得多的溢价。
  • 结果是,许多当前流行的AI应用可能被挤出价格区间。Google、Anthropic和OpenAI为“自动化AI研究”购买token的意愿,会高于你我为让更多AI垃圾内容开口说话所愿支付的价格。

5. 供给为何无法回应:3倍增长脆弱且不可扩张

  • 他先回应 Simon–Ehrlich 类比:Ehrlich的马尔萨斯式判断曾著名地败给人类巧思,但“我猜,这个赌局与之类比大概不对”。与金属开采相比,算力供给的弹性要低得多,吸收大规模需求冲击的能力更弱,也更难找到替代品;换个年代,Ehrlich很可能才是赢家。
  • 这3倍增长的拆解是:摩尔定律贡献1.4倍——能把它维持下去就是“奇迹”;新建晶圆厂贡献1.2倍,但受ASML的EUV光刻机限制,瓶颈将持续到2030年甚至更久(Dylan此前的一期节目已讨论过);从手机和PC转向AI的晶圆再分配贡献1.8倍,而随着AI在TSMC先进制程N3中的占比从60%升至86%,这条路大概率在明年年底撞上上限。

6. 定义这笔交易的限制条件

  • 这明确处于“前奇点阶段”:最终,机器人把二氧化硅和铜加工成芯片后,算力会再次变得便宜。
  • 收入增长10倍而算力仅增长3倍,本身就说明模型业务拥有多强的规模经济:一次性训练成本可以由所有用户共同分摊,而人类劳动力的训练无法在用户间复用,近乎每次都要从头开始。他对此直言:“我真希望我们不用生活在一个智能拥有如此强规模经济的世界里,因为我担心权力集中,但事实似乎就是如此。”
Dwarkesh Patel

Today I want to talk about what the compute situation for the labs will look like over the next few years. For the last 3 consecutive years, Anthropic's revenue has 10x'd year over year, and it's likely to do so again this year. They ended last year with $9 billion in revenue. I think they'll probably end this year with somewhere between $100 billion and $150 billion in revenue.

Now, for this trend to continue, Anthropic would need to make $1 trillion in revenue by the end of next year. Of course, there's no deep reason why this has to be true. It's a very wild conclusion, and it's ultimately a question of AI capabilities: Does AI get that useful by the end of next year? But suppose the trend does continue. I want to think through what happens in that world.

1. Compute Revenue Creates a Gap

The other big trend in AI is that lab compute only 3x's year over year. For a lab to keep 10x'ing revenue year over year while compute only 3x's, one of the following 3 things needs to happen, or some combination of the 3. One, lab margins have to increase. Two, the price of compute has to increase. Or three, the percentage of compute that labs spend on inference rather than training has to increase.

My understanding is that all 3 of these things are already happening. With regards to the margins, Anthropic's inference margins reportedly went from 40% in the middle of last year to upwards of 80% now for Fable. With regards to compute, the spot prices for compute are more than 40% higher than they were in the February trough that we had earlier this year. And with regards to the share of compute that goes to training versus inference, in 2024, according to Epoch, OpenAI was spending just a quarter of its compute on inference, and that number is likely closer to 50%, if not higher, now.

Now, labs would prefer not to do this final thing of increasing the share of compute they spend on inference. The way the labs see the world, the whole point of inference revenue is to help convince investors to give you more money in order to train the next bigger, better model. And if you're spending most of your compute on inference, then you're basically declaring that AI progress has stalled and you're just now in the business of being a cloud provider.

This is a less compelling business than building AGI, so the labs do not want to be in this business, nor do they think they are in this world. They think that within a year, they'll have built models that make the current ones look extremely shitty. But they need to invest a lot of their compute—the majority of their compute—into doing the training and experiments that are necessary to build the next model.

So that leaves only 2 options for how you can get out of this gap between the fact that lab compute only increases 3x year over year, but revenue increases 10x. Either the lab's margins have to increase so that they get the surplus, or the price of compute has to increase so that everybody in the stack below the lab gets the surplus.

It's not clear to me which world we end up in. Do we end up in a world where we go from 80% for some of the top models to greater than 90% margins if the lab margin effect dominates? That would require the leading model to be so far ahead of the competition, because the nature of margins—why they exist in a market economy—is that the thing you are serving is so much better than what somebody else could go get and replace you with on the market. But it's just really wild for me to consider that the margins for something like intelligence will be greater than 90% and they don't get competed away at that level.

2. Compute Prices Must Rise

So that leaves only one other possibility of this escape valve between these 2 trends, which is that the price of compute has to increase. As I mentioned, this is already starting to happen. And the effect is even stronger when you look at the tranche of compute that the frontier labs actually need to accumulate, because they can't just go out and buy a spot instance.

They need to make sure that they get enough scale to get really good efficiency and flexibility, and also that they have the kind of compute that lends itself to the security they need for their own weights and for their customers' information. I think a relevant case study here is to look at the compute that Google and Anthropic are renting from SpaceX. Google, for example, is paying $900 million a month for 110,000 GPUs that are a blend of GB200s and GB300s. The price that Google is paying here is 2x the spot price per hour for those GPUs. And that spot price itself is more than 40% higher than it would have been in February.

3. Smarter Models Monetize Compute

I want to emphasize a key conclusion here: As AI models get smarter, they will be better able to monetize the same amount of compute. If a true human-level software engineer could run on an H100 equivalent, then at today's prices for software engineers, that H100 should rent for over $250K a year. That's over 15x the current spot price for an H100. And this is not even accounting for the fact that your AI can work nights and weekends.

Of course, you might expect that if we had 10 million extra software engineers suddenly appear in the economy, the marginal value of a software engineer would decrease, and thus the revenue that that H100 would be able to generate would not be 15x higher than it is right now. But I actually don't know if this is true. If we apply this argument to people instead of AIs, then this would be the classic lump of labor fallacy.

For example, economists generally believe that high-skill immigration does not decrease wages in the long run because of how innovation and specialization increase the value of labor. Maybe this labor supply shock will be so big and so fast that we can't count on this general heuristic anymore. But if you believe what standard economics says, then the marginal value of labor, and thus the marginal value of compute, should stay astonishingly high.

4. Scarcity Rewards Efficiency

So let's think about what changes in such a world. One of the things that would happen is that as the top labs get better and better at monetizing compute, and the cost of compute increases, it becomes harder for anybody else to compete against them, because they have to bid for this resource against somebody who is basically able to make better use of it.

Another thing that will happen—and I think this is actually the most interesting implication of this whole thought exercise—is that if you can train the best, most efficient model, then you'll be able to charge much higher margins than you can today. This is the Alchian–Allen effect in economics, and what it's basically saying is that if it costs $20 an hour to rent an H100, then it would be extremely stupid to use a weaker, less efficient model, because it's gonna burn more tokens on your expensive compute to get the exact same result.

So labs will be able to charge a much larger premium if they can train a model that better economizes this scarce input. Basically, if you have a model that can get the same result by using less compute, then you've, in some sense, created more compute, and the value of compute is gonna increase.

Another thing that will happen is that a lot of current popular applications of AI will probably get priced out. The reason AI is relatively cheap right now is that AI just can't do a lot of things that top humans can do. But this, at some point, will no longer be the case. And at that point, Google or Anthropic or OpenAI will be willing to pay more for the tokens to automate AI research than you or I will be willing to pay to make more AI slop talk.

5. Compute Scarcity May Persist

I'm a bit worried that this kind of analysis honestly pattern-matches a lot onto the ways that people in the past have been wrong about scarcity. I'm thinking, for example, of the famous Simon–Ehrlich bet. Paul Ehrlich was this famous doomer about population growth, and he made this bet that a basket of commodities would increase in price rather than decrease in the decade preceding 1990.

This is a very famous bet because it's supposed to illustrate how Ehrlich's Malthusian worldview was wrong, and how he did not anticipate the way in which market signals and human ingenuity can find better ways to economize scarce inputs. I'm guessing that the analogy to this bet is probably wrong. Other analysis has shown that if that bet had been made in a different decade, Ehrlich might well have won.

But more generally, I think the supply of compute is much less elastic, much less capable of absorbing large demand shocks, and much less capable of being accommodated by using different substitutes than the extraction of different metals is.

To illustrate why I think this 3× in compute capacity year over year is hard to budge or potentially even sustain, I don't see how any of the 3 elements that constitute that 3× can be much accelerated. 1.4× of that is coming from Moore's Law. Far from increasing it, I think it'll be a miracle if we can just keep it going for a few more years.

1.2× is coming from building new fabs. This process is ultimately gonna be bottlenecked up to 2030 and potentially even beyond by just building new ASML EUV machines. Dylan, when he was on the podcast a few months ago, talked about this in great detail.

And 1.8× comes from the fact that AI is absorbing a lot of wafer allocation that was previously going to smartphones and PCs. This is probably gonna hit a wall by the end of next year, when at the leading-edge N3 nodes at TSMC, AI will have gone from 60% to 86%. At some point, you have just absorbed all leading-edge wafer capacity for AI, and you can't keep increasing this number.

So I don't know how we even continue to do 3× compute scaling year over year for the next few years, much less go beyond that.

6. Cheap Compute Comes Later

Now, I wanna clarify that at some point in the future, compute will get cheap again. At some point, we'll just have robots that can convert shores of silica sand and mines of copper into new computer chips, and then the price of compute is basically the raw inputs and the tools required to do this processing.

I'm just talking about this current pre-singularity regime where AI compute merely 3×'s year over year, which is not enough to offset how much more valuable AI is becoming over time. By the way, the fact that Anthropic's revenue has been 10×'ing year over year, whereas their compute has only been 3×'ing year over year, I think illustrates how strong the economies of scale are in the model business.

And logically, this makes sense. When you train a model, you just have to spend this one-time cost to learn all these different skills that then get to be shared across all your users. This is very unlike human labor, where each instance has to be retrained from scratch.

I wish we didn't live in a world with such strong economies of scale for intelligence, because I'm worried about power concentration, but it seems we do.

Okay, this was a narration of a blog post that I also released on my website at dwarkesh.com. Check it out for other posts or to be notified when I release a post in the future. Otherwise, I'll see you for the next full episode.