[BidClub_]
SemiAnalysis · · 44 分钟

第013期——AWS利润率跳升10%,Azure和GCP停滞不前(Tokenomics)| Jordan Nanos、Jeremie Eliahou Ontiveros、Joey Brookhart、Crystal Huang

Jordan NanosJeremie Eliahou OntiverosJoey BrookhartCrystal Huang

播客
TL;DR
  • AWS相较Azure和GCP改善的利润率,来自通过Bedrock销售Claude这一高毛利Token业务,而不只是出租加速器。 AWS每季度新增超过1 GW算力,利润率却在上升;Azure和GCP的业务结构仍更多暴露于低毛利的基础设施即服务。Joey的框架是:Token销售比5年期take-or-pay合约保留更多上行空间。

  • 在集群达到“稳定”利用率之前,算力扩张通常会先压低云业务短期利润率。 CoreWeave这类供应商要先承担折旧、租赁和人工成本,GB200等复杂系统在等待数月启用期间不产生收入。因此,AWS能在扩张利润率的同时吸收同样成本,对其最终资本回报率而言是“一个相当不错的信号”。

  • Claude以API为主的增长,正在为Anthropic和AWS带来异常强的经营杠杆。 Joey称Anthropic ARR已达470亿美元,3月、4月和5月每月净新增ARR可能约100亿美元,其中约80%来自API。Amazon当时“在正确的时间处于正确的位置”:Claude约占Bedrock使用量的80%-93%。

  • Anthropic以650亿美元完成的H轮融资,对应9650亿美元投后估值,放在其增长和盈利能力背景下看并没有那么极端。 Joey将约20倍ARR估值倍数与2021年软件公司达到的80倍水平相比,并称Anthropic在不计股权激励费用的情况下已经盈利。需要警惕的是,如果企业削减编码Token消费,经营杠杆可能显著反转;但“眼下列车还没有减速”。

  • SpaceX/xAI算力交易,引发了本期节目围绕AI需求最尖锐的分歧。 Jeremie认为,一个原本购买算力的前买家如今成了供应商,可能是在“放弃前沿竞赛”;Jordan则认为,Anthropic需求强得惊人,收回GPU的选择权很有价值,在xAI自己的研究和分发渠道能够消化这批算力前,用它创收是理性的。

  • AI是否会变成赢家通吃,取决于支出是否集中到那些不存在“够好就够了”的开放式任务上。 Crystal认为,排名第三或第五的模型仍可能替代大量劳动力;Jeremie反驳说,法律、科学、医疗和分析工作会奖励持续采购更强的智能。Jordan的表述更宽泛:“编程不是编程,而是使用计算机。”

  • 能够把前沿模型访问、企业分发和定制芯片结合起来的超大规模云厂商,可能才是长期赢家。 Joey预计到年底,Bedrock可能成为AWS AI业务的主体,而Azure和GCP在小组模型中仍有80%-90%属于基础设施即服务。Trainium和TPU在Token买家永远无需知道具体由哪种加速器提供服务时又多了一重优势:“赢家赢,输家输”。

摘要 · 为研究而整理的核心内容

1. Bedrock将云算力转化为高毛利产品

  • Joey将超大规模云厂商的AI业务分成三类:GitHub Copilot等软件;出租加速器的基础设施即服务;以及“Token即服务”,即客户通过现有云协议购买模型访问权。最后一种模式保留企业级安全性、可用区和统一计费,同时比单纯出租芯片让云厂商捕获更多经济价值。

  • Crystal认为,GPU即服务削弱了过去的云护城河:大型用户越来越想要“只要硬件”,并按自己的方式配置,而不是使用曾经让AWS、Azure和GCP极具防御性的托管平台。Jordan称,如今已有超过200家新型云厂商,说明行业进入门槛已经降低。

  • Bedrock重新建立了差异化的经济模型:Claude使用量主要经由AWS流转,而Token销售的利润率高于固定合约下的基础设施出租。这一业务结构解释了为什么AWS营业利润率在改善,Microsoft则在下滑,Google相对持平。

2. AWS跑赢算力爬坡的利润率陷阱

  • Crystal以CoreWeave为例,将稳态经济性与爬坡阶段区分开来:在5年期take-or-pay合约下,一座完全运转、达到稳定状态的集群可能带来约25%的营业利润率和30%-40%的毛利率。关键在于“稳定”——合约收入和稳定成本结构,只有集群真正运转后才会出现。

  • 在启用前,供应商已经建好数据中心,并开始承担折旧、租赁、人工及其他费用。设备安装需要数月,GB200等复杂系统又进一步拉长了这一周期,导致资产持续产生成本却没有收入。

  • AWS面临同样的物理约束,每季度新增超过1 GW算力,但利润率仍在上升。Crystal的推论明确带有条件:如果在前所未有的交付规模下利润率仍能扩张,那么Token即服务最终达到稳定状态后的利润率和资本回报率可能会“极其丰厚”。

3. Claude的API结构让Anthropic更容易量化

  • Joey估计,Claude占Bedrock使用量的80%-93%;Microsoft的业务更偏向OpenAI,Google则更偏向Gemini。他称Anthropic ARR已达470亿美元,3月、4月和5月每月可能新增约100亿美元,其中约80%的增量ARR来自API。

  • Crystal认为,Anthropic比OpenAI更容易预测,因为API工作流会暴露Token消耗量,且定价一直相对稳定。OpenAI第一季度业务约60%来自消费者订阅,用户可以取消或切换;Anthropic的API消费则更容易观察。

  • 关于Opus 4.8,Crystal保留了这一前提:常规定价没有变化,fast-mode定价不同;Anthropic的图表“据称”显示,其幻觉数量显著少于4.7。她尚未亲自测试,但希望它能终结分析中“凭空捏造数字”的问题,并称其接近Mido的预览版。

  • 650亿美元H轮融资对应Anthropic 9650亿美元投后估值,接近2月讨论的3800亿-4000亿美元水平的2倍。Joey认为,约20倍ARR估值倍数并没有2021年软件公司的80倍那么极端,并称Anthropic在不计股权激励费用的情况下已经盈利,不过编码Token消费削减可能造成显著的经营去杠杆。

4. SpaceX/xAI交易让小组在算力稀缺问题上分裂

  • Jordan称,SpaceX提交的文件披露了与SpaceX相关的“数十亿美元”支出,并包含允许xAI收回GPU的条款。对Anthropic而言,如果算力约束正在限制增长,这批容量可以缓解限流;对xAI而言,这能把算力转化为产生收入的资产,同时不必永久放弃这批算力。

  • Jeremie认为,这笔交易对整体算力需求偏利空:“一个原本应该是需求来源的玩家,变成了供给来源”,既会加剧GPU即服务的竞争,也会减少一个承购方。更直接地说,把算力卖给Anthropic,意味着xAI可能正在放弃前沿竞赛,尽管训练回报已经显示出异常强的水平。

  • Jordan反驳称,交易存在的原因恰恰是Anthropic的需求“强得令人无法应付”,以至于必须从竞争对手手中购买算力。收回条款同样重要:xAI可以先将算力变现,同时推进研究;一旦取得突破,或X、Starlink、Tesla的分发能力足以支撑更大规模的运行,就可以重新调配这批算力。

  • 这场分歧最终落到了资本结构上。Meta可以依靠一个巨额现金创造业务,为长期研究提供资金;xAI则依赖有限的风险资本。Joey提出“训练、利息和税项前利润”(earnings before training interest and taxes)的框架,把推理利润视为经营现金流、把训练视为投资,由此引出一个问题:每家实验室的模型投入是否能取得足够回报。

5. 前沿经济学可能让领先者之下几乎没有价值

  • Crystal提出,排名本身是否重要?排名第三、第四或第五的模型,可能仍足以替代大量工作。Jordan则重新定义了这场竞赛——“编程不是编程,而是使用计算机”——并将Anthropic排在第一、OpenAI第二、Cursor第三,因为Composer把分发能力和可用模型结合在了一起。

  • Jeremie认为,在最大的支出池中,“够好”并不存在。翻译最终可能变成有限需求,但法律、科学、医疗和分析工作都是开放式的:用户可以持续为更强的智能付费,以搜集证据、探索替代方案并击败竞争对手。他的结论是绝对的:“如果你排在第三、第四或第五,就拿不到任何收入。”

  • 他以Mythos为例,区分Token价格与完成任务的成本:尽管Token定价“贵了6倍,我想是5倍”,他仍称其比Opus“便宜三分之二”。如果一个模型聪明10倍、所需Token少10倍,它仍可能是更便宜的工具——从而把需求集中到前沿模型。

  • 算力本身无法决定竞赛结果。Jordan认为,实验室需要算力和人才;Jeremie对此表示认同,两人都认为xAI同时失去人才和算力,说明其东山再起将很困难。Jordan仍不确定,最终胜出的会是哪种收入、分发能力与“hero runs”的组合。

6. 分发与芯片让市场继续集中

  • Jordan以Cursor约20亿美元ARR、Fireworks宣称3.15亿美元ARR为例,质疑赢家通吃的判断。Jeremie承认开源模型的采用势头令人鼓舞,但相较于已经超过1000亿美元的AI市场,这些数字仍无足轻重:“在朋友之间,3亿美元算什么?”

  • Joey预计到年底,Bedrock可能成为AWS AI业务的主体,尽管AI在AWS中的业务占比仍低于其在Azure或GCP中的占比。Azure和GCP目前仍被估计有80%-90%属于基础设施即服务,但Azure可以快速加入Claude,因为超大规模云厂商与客户的关系让Token分发相对容易。

  • 新型云厂商面临更难打通的循环:要成功开展Token即服务,既需要与前沿实验室合作,也需要提前配置具有投机性质的算力,还需要足够资本在没有5年期承购协议的情况下部署GPU。CoreWeave、Nebius和AIREN在这些优势中各有缺失,使推理端点高度集中在排名前三的超大规模云厂商手中。

  • 定制芯片进一步拉大了利润率差距。AWS的Trainium和GCP的TPU提供了垂直整合能力,而这正是高度依赖NVIDIA的Azure所欠缺的;当1 GW算力只服务少数几个模型时,加速器的易用性就没那么重要。正如Jordan所说,Claude Code用户无法分辨一个Token来自GPU、TPU还是Trainium,也不需要知道。

Jordan Nanos

Hello, everyone. Welcome back to SemiAnalysis Weekly, episode number 13—lucky number 13. We're here with Joey, Jeremie Eliahou Ontiveros, and Crystal. We're going to talk about an article that we put out recently called “Anthropic Growth and Bedrock Mix Drive AWS Margins Higher While Peers Lag.”

That means we're going to talk about everything Anthropic, including the recent announcement of their Series H and the release of Opus 4.8, but with a focus on the infrastructure—how exactly they serve these tokens, especially with their partnership with AWS. Guys, welcome to the show. Excited to talk through this.

Joey Brookhart

Thanks, Jordan.

Crystal Huang

Thanks, man.

Jordan Nanos

All right, so let's dig into the article itself and talk a little bit about the backstory. I think a lot of people understand the concept of tokens and understand what GPUs are, but not everybody is getting their tokens from the same place. Can one of you guys give me a backstory on Bedrock? What is AWS? What are they doing for Anthropic, and how are they serving tokens with Bedrock? Joey, start with you.

Joey Brookhart

Perfect. I'll go through that, and I'll do it for all the clouds as well. If we look across all the hyperscalers, and especially the big 3—Amazon, Microsoft, and Google—there are 2 big breakouts, maybe even 3.

There's a bit on the software side, so if you look at Microsoft, things like GitHub Copilot, that's an AI software-as-a-service-type product. They have AI infrastructure as a service, where they're just renting out these accelerator chips. Then they have this token-as-a-service business, which is where they'll essentially expose these outside models or their own models to consumers to interact with.

It's a little bit of a different business because instead of just renting the underlying chip, you're renting the underlying model—buying it through your cloud provider, your CSP account, and your enterprise spending agreement. You have all the same benefits around security and availability zones, and you're able to buy third-party models and some of these first-party models through your cloud provider of choice.

At a high level, those are the 3 big buckets right now at the large hyperscalers that we see, and we're seeing some big changes or differences between them in how they've gone about it strategically. The 2 main things were Anthropic's growth and Amazon's strategy around Bedrock. Those have really driven their margins higher recently, and that was the takeaway from the article. This token-as-a-service business is obviously much better for the hyperscalers than infrastructure as a service on the AI side.

Jordan Nanos

Yeah. If you look at the big 3, they're kind of going in opposite directions right now, just in terms of their operating margin, right? That was the biggest chart from the article. I'll put it up on screen right now.

But Crystal, can you explain, when we look at AWS versus Google and Microsoft, and try to break out the cloud business, is this all to blame on Anthropic from our perspective? What's driving AWS to improve its operating margins while Microsoft's are declining and Google's stay flat here?

Crystal Huang

I feel like a lot of it is Claude usage, right? A lot of people are using Claude more, and it's mostly routing through AWS. As Joey was saying, with the token-as-a-service business model, they just have a better margin on that as opposed to the infrastructure-as-a-service model.

Yeah. I think if you take a step back, the world of clouds has really changed a lot in 2023, as you started seeing these neo-clouds, these GPU-as-a-service businesses. In the old days—which was basically 2022 and before—cloud service providers were basically 3: Amazon, Google Cloud, and Microsoft Azure. They had amazing margins and amazing returns on capital. Some new entrants, like Oracle, were trying to get in, but really the market was dominated by 3 players that had an amazing business.

Now you get to GPU as a service, and what folks started to realize is that the barriers to entry are much lower. Jordan, you're probably the best person to talk about this. You probably know the CEOs of 150 neo-clouds or 200 neo-clouds.

Jordan Nanos

Yeah. Over 200.

Crystal Huang

It is pretty insane. Obviously, some are better than others, but the point is that the market has much lower barriers to entry. The moat that cloud used to have doesn't really exist in the AI era because it's really more about infrastructure and, especially, the end users.

The whole point of cloud computing was to make IT much easier. Folks don't need to have such a big IT department internally; they can just rent through the cloud. It's super easy. Everything is well.

Now the big end users want to have much more control. You shift from platform as a service in the old days to bare metal. Folks want just the metal. OpenAI wants things the way they like them; Microsoft and Meta do too.

This token-as-a-service business is basically the first case at scale where you see an AI cloud provider having a business with a different profile. Obviously, as you have less of a moat in the GPU-as-a-service era, margins go down. Oracle is the best example. They have an RPO of half a trillion dollars, which, as of a few quarters ago—that was as of Q1—is still bigger than Amazon's. Their backlog is bigger than Amazon's, but no one gives them credit because people know it's much riskier. The returns are not the same.

We've seen empirically that every single GPU-as-a-service cloud has faced struggles when it started to ramp up its business. Our view as a firm is that we actually think the GPU-as-a-service business model is sound. Companies like CoreWeave have a sound business.

But there's this lag effect where, as you ramp up and bring more capacity online, there are lags that make your margins go down temporarily because your asset base depreciates, you have to pay data center leases, and so on and so forth. It's really interesting to see that Amazon is basically the first cloud provider that, in a time of unprecedented capacity expansion—over a gigawatt per quarter now—expands margins.

That tells you that if, during a period of accelerated capacity delivery, they can expand margins, you start to think, “Okay, what's the stabilized margin of this business?” One of the points that we make is that the stabilized margins of token as a service for Amazon are actually extremely rich. The return on capital is fundamentally much more so.

Jordan Nanos

Yeah. Let me throw this chart up on screen, actually, from the article. It's the percentage of revenue that is going to Bedrock, Bedrock being the token as a service at Amazon. You can obviously see it ramp up quite a bit at roughly the time that we're in right now—Q4 of last year and the first quarter of this year.

If we compare that to the chart I had up previously, where we're seeing their operating margins improve in the first quarter of this year, they're saying that's due to this. Can you explain in more detail why it's unprecedented to say that Amazon can bring on a gigawatt per quarter and still improve operating margins?

Crystal Huang

To understand this, you basically have to go back to why the pure bare-metal providers are seeing their margins go down. You look at CoreWeave, which is the pure play, so they're the cleanest example. Oracle is kind of the same.

When you look at this chart, essentially what this tells you is that their stabilized business does something like 25% operating margins and 30% to 40% gross margins. But the whole issue is stabilization. Stabilized means that your GPU cluster is fully functional. You're getting the monthly rent, or whatever rent, from your customer on a take-or-pay basis, so it's a flat fee. You know exactly how much revenue you're going to make. Oftentimes, it's a 5-year take-or-pay contract, again, at a fixed rate.

You know your revenue and you know your costs; everything is stable. But before getting there, you obviously have to set up the data center, which is a huge capital expense upfront. Then you have this whole process where you have the data center built, but you need to fill it with equipment. That takes a few months.

With some new types of equipment, like GB200, which is super complicated, we've seen that lag get longer and longer. That means this period of time where you depreciate your assets, pay data center rent, pay some labor, and pay a whole bunch of other costs gets longer, and you don't make any revenue because your cluster is not yet turned on. That's the whole dilemma these guys are facing: they know their business model and that stabilized structure, but they've been facing challenges ramping it up.

Some are a bit conjunctural—again, GB200—and some are more structural because there’s this lag. Amazon is the same thing, right? Like everyone else, they’re bringing on a whole lot of data centers and a whole lot of XPUs. These XPUs, in theory, should take time to bring online. Yet despite this, you’re seeing their margins go up, which is a pretty good sign for them.

Jordan Nanos

Yeah, but it’s clearly different from the others in this space, where the percentage of their total AI revenue that they’re reporting as being from token-as-a-service, as opposed to other products, is much higher than at Google and Azure. They’re not just bringing on capacity; they’re successfully selling it into the labs that are using it to serve tokens for these models.

Joey Brookhart

Yeah, and I guess the charts you showed earlier—what I forgot to mention is obviously the margin buffer that you have when you sell tokens with your infrastructure, as opposed to having a 5-year take-or-pay fixed contract with a capped upside, right? So that’s kind of the key—one of the key points of the article.

Jordan Nanos

Yeah. Can you talk a little bit about the workload mix as well? Obviously, any provider could conceptually do this, but not everybody is doing it successfully. It’s not like Azure or Google doesn’t have a token-as-a-service business. In fact, even Crusoe, CoreWeave, and Nebius are all trying to get into this business, too, but they really need a customer, and they need a customer serving the right type of workload for it to really result in a bunch of growth, I would say.

Joey Brookhart

Yeah, for Amazon specifically, they benefit from having the biggest customer base. They’ve been doing this for 20-plus years now at AWS, and people are very comfortable buying through them. People even buy infrastructure software through them, from providers like MongoDB and Snowflake. It’s a pretty large marketplace business, so customers are really comfortable with the security at this point, buying through them and having a single bill for all of this.

I think this comes back to some of the Anthropic news today, also on the ARR number of $47 billion. When we look at Q1 and even into Q2 here, a lot of Anthropic’s business mix is much, much different. Amazon is benefiting from Bedrock being 80% to 93% Claude, versus Microsoft being heavily OpenAI, obviously. Google has a lot of Gemini, which doesn’t benefit as much from a lot of these agentic coding tasks.

When you look at OpenAI and the coding percentage that’s really driving Anthropic, you’re talking probably like $10 billion a month in net new ARR over the last 3 months—March, April, and May together. Probably 80% of Anthropic’s net new ARR is in this API business. If we go to OpenAI, 60% of that business in Q1 was really consumer subscriptions.

There’s kind of a mix of factors, but Amazon was in the right place at the right time with Anthropic. They were also able to give, with Trainium2, I think, a pretty interesting deal structure for both parties. We mentioned this in the article, too: obviously, with their mix of Trainium, there’s an infrastructure-as-a-service fee component that Anthropic pays for this infrastructure, like Bedrock infrastructure. But then there are some interesting hurdles around revenue share and, really, margin share that happen.

Because Anthropic was probably at $25 million of ARR—I’m making that number up off the top of my head—in Q1 versus probably $6 million back in Q4, the numbers really made sense for both parties, and they both benefited.

Jordan Nanos

Makes sense. Maybe, Crystal, can you talk a little bit about those forecasts you guys were making? Going into the end of Q1 and your forecast for Q2, you don’t necessarily get the disclosures the same way from Anthropic, but we’ve at this point kind of been bang on with the disclosures, with the disclosed revenue figures and the margin figures, right?

Crystal Huang

Mm-hmm. I think for us, it’s a little easier to forecast Anthropic than it was OpenAI, just because so much of Anthropic’s ARR comes from the API side. Because we also use Anthropic, and there’s a lot of data out there about how people are using all of the different Claude models, it’s so much easier to predict the workflow and see what token consumption is going to look like. They’ve kept token pricing relatively stable these past 2 releases, more or less, right?

Whereas with OpenAI, a lot of their revenue comes from these subscriptions, and you never know if consumers are going to switch over to another one. So it’s a lot harder to quantify the number of users who are using the subscription when they can just cancel anytime, versus the API.

Jordan Nanos

Yeah. Can you explain a little bit about the release of Opus 4.8 and 4.7? Pricing has remained the same, but fast mode has changed, and maybe there have been some other changes in terms of how they’re doing pricing on the API.

Crystal Huang

Yeah. They said pricing is the same for regular mode, but for fast mode it’s different. Another cool thing that they said was that it doesn’t hallucinate as much, and they had a pretty cool bar chart showing that its rate of hallucination is a lot lower for 4.8 than for 4.7, supposedly. I haven’t tested it out yet, and they said that it’s super close to Mido's preview, so hopefully it’ll stop pulling numbers out of thin air when we’re using it for our analyses.

Jordan Nanos

That’d be good. That’d be good if numbers weren’t pulled out of thin air.

Joey Brookhart

That’d be great.

Jordan Nanos

Yeah.

Crystal Huang

Yeah.

Jordan Nanos

In terms of fast mode and consumer subscriptions, do you have comments there on what we’ve learned over the past few weeks or months from digging into how Anthropic is running its business on AWS? What’s maybe the biggest percentage of their revenue or their margin contribution across those different mixes—the different types of workloads that people could be consuming tokens on the API for?

Joey Brookhart

We have that in the tokenomics model, and we’re doing a big study right now on what percentage of the coding market is currently represented in token spend, especially on the API side. We’ve done a lot of work on the consumer side and some work on the B2B data, where Anthropic has just been taking a ton of share of net-new customers year to date, both in consumer subscriptions and enterprise.

But it’s really difficult, and it’s a big question among a lot of our clients: How big is the coding market currently? I think Anthropic recently said at their Financial Services Day that financial services was the second-biggest vertical. We know there’s a pretty big gap between coding and financial services, especially in terms of API spend, just from how people use this anecdotally.

There have also been a lot of recent comments on token maxing, especially at the Fortune 500s. How do you budget for that? How do you blow through that spend over time? How do you measure ROI? Things will have to change, I think, and then people will put some processes in place.

I know guys like Jeremie get insane ROI in the data center model, and his team gets that, too. Maybe there’s less policing at some other organizations that just let people go crazy. That’s another contributing factor, I think: the success of coding, and Anthropic starting to win a lot of net-new share on the subscription side in both B2B and consumer, which we saw was really interesting in Q1.

Jordan Nanos

Yeah. I guess 3 things have happened since the last time we talked about this topic on the podcast. First of all, they signed that massive deal with SpaceX, xAI, and Cursor.

Crystal Huang

The SpaceX, xAI, Cursor. Beautiful. Beautiful.

Jordan Nanos

Yeah. Cursor isn’t part of it yet because they’re trying not to change their S-1, I think. But anyway, the SpaceX S-1 revealed the many billions that they’re spending with SpaceX, with a clause to let them back out—meaning xAI has the ability to reclaim these GPUs if they want—but that should be some contribution to Anthropic’s total revenue. In other words, if they were constrained by compute for their ability to grow the business on the consumer subscription side and enforce rate limits or things like that, those should go away pretty quickly.

The second thing that happened is obviously that they raised their Series H: $65 billion in funding at a $965 billion valuation.

So they raised $65 billion at a $900 billion valuation, resulting in $965 billion post-money. I don't know why they didn't round that up to a nice, even trillion, but we'll see.

Crystal Huang

Isn't that almost double February's number—their valuation, right?

Jordan Nanos

Yeah. What was the February number?

They were at $400-something?

Crystal Huang

$380 billion or $400 billion, somewhere around there.

Jordan Nanos

I guess they need their valuation to track with their ARR growth.

Joey Brookhart

It's not a crazy multiple: 20× versus the software bubble back in 2021. Some of those were 80×, like Snowflake and Cloudflare. It's not super expensive.

Then, from the recent Wall Street Journal article and in our financials, which we also had in our model, we know that they're profitable now. When you exclude stock-based compensation, Anthropic's a pretty good business model, and they're seeing a ton of operating leverage. We also know from The Information article, year to date, that OpenAI is not seeing that as well.

It's very clear what the better business is right now and what the better business model is. There's obviously a lot of operating deleverage if people start cutting how much they're spending on coding tokens.

Jordan Nanos

Yeah.

Joey Brookhart

We're not. I mean, yeah, they're still growing at $10 billion in net new ARR a month over the last 3 months, which is pretty crazy.

Jordan Nanos

It's good.

Joey Brookhart

They're probably going to do another $10 billion in June. Who knows where they end up at the end of the year? If Mythos [?] is released as well, that probably does help.

There's no train that's slowing right now, and they're in all the right places. They're seeing the benefits of that throughout the entire business model.

Jeremie Eliahou Ontiveros

One big question here: the xAI deal, or SpaceX deal, with Anthropic—is it bullish or bearish for the market overall?

Jordan Nanos

Which market? Neocloud market?

Jeremie Eliahou Ontiveros

Compute demand overall. Everything's related, right? Stocks, compute demand, NVIDIA, all of it.

Jordan Nanos

I think it's super bullish. Neoclouds, man. Everybody wants to be a neocloud. Neocloud is the terminal business. Even AI labs want to be neoclouds, selling their compute to whoever they choose to right now.

Jeremie Eliahou Ontiveros

I disagree with this because this deal is basically one player that was supposed to be a source of demand becoming a source of supply. So now suddenly there's more competition in the supply market, which is GPU as a service, and the offtakers—there's one less. So I don't know, man.

Jordan Nanos

Yeah, I disagree with this. I think Cursor is training plenty of models on Colossus right now. I think they wouldn't have that provision in the contract to take back their GPUs if they really were going to be no source of demand in the future.

I think there's lots of demand to go around, and the reason that deal's happening is because Anthropic's demand is so overwhelming right now that they need to do these crazy things, like buy compute from their competitors, in order to be able to serve that demand. If they had a different way to serve that demand, they would be doing it, I assume.

Jeremie Eliahou Ontiveros

On the other hand, if you're xAI, Meta, or whoever—other labs that are lagging—you see Anthropic and you're like, “Whoa, bro, this is what I could do if I was at the frontier.” Right? If Groq suddenly was at the frontier, they could be a $100 billion ARR business.

So, to some extent, you would argue this should make them more bullish.

Jordan Nanos

I think, yeah.

Jeremie Eliahou Ontiveros

They should be—

Jordan Nanos

I think their S-1 has about $24 trillion of enterprise AI applications carved out as future market TAM, right? So you don't even need to be a $100 billion ARR business. You can just take that forward a few more quarters and take it to $24 trillion, no problem.

Jeremie Eliahou Ontiveros

So their TAM is $24 trillion, but that's for what? Generative AI applications?

Jordan Nanos

I'm going to bring up that chart now from the S-1. Yeah, it's like there's space and telco, and then there's a really big section for enterprise AI applications.

Jeremie Eliahou Ontiveros

Okay, enterprise AI applications. So that would basically be the TAM for Groq, right? And so they're saying, “Actually, we give up on the $24 trillion TAM. I'd rather give my compute to Anthropic. That's a better use case than fighting for a $24 trillion TAM.”

Jordan Nanos

Well, I think in some ways it's just a matter of when you can spend that money, or when you can spend that compute. In other words, there's potentially some serial nature to the development of AI progress, where you have to wait for you or all of your competitors to run a bunch of experiments to figure out the optimal model architecture and dataset mix—or just the creation of data, synthetic data, whatever it is—before it's done.

Because if we were to take a lot of these models that people are running today and try to run them on hardware from 3 years ago, the model architectures would actually run really well. All the innovations in sparsity and attention—these are huge improvements over dense models from 3 years ago.

Jeremie Eliahou Ontiveros

But they still could have—

Jordan Nanos

These are huge improvements over dense models from 3 years ago.

Jeremie Eliahou Ontiveros

They could have used that compute. If you're saying, “Maybe they had a bottleneck because they weren't able to figure out new architectures that would have enabled them to use that compute efficiently,” then they could have used that compute to do research on those specific topics, right? Less training, more research.

Or they could have used it to give a whole lot of tokens to their employees and make them much more productive at doing a bunch of research tasks. Now we know that AI can do pretty complex science problems—the OpenAI math stuff, which I know a few things about, because my dad does math for a living.

So, yeah, I mean, I don't know. That does tell you that it's an odd decision when you're kind of in the fight and giving up, when you're seeing the strongest signals we've ever seen that this is real and accelerating—and actually, $10 billion of ARR per month, right? So I don't know. It's pretty odd.

Because I feel like this is the opposite of Meta. My sense is that, to some extent, xAI is giving up on the frontier race, whereas Meta is, if anything, getting more bullish because of what they're seeing from Anthropic. They're like, “Hell yeah, this is what I bet on, and I'm going to double down after having doubled down so many times already.” So Meta is the one that's seeing Anthropic and being like, “Hell yeah. I want that.”

Jordan Nanos

Well, I think there are 2 dynamics that you're overlooking here a little bit, potentially. One is that Meta has a huge cash-generating business that they can use to fund all of this, and SpaceX really just doesn't have a business that generates hundreds of billions of dollars of free cash flow that they can pour into compute for the research bets.

So they have to do it based on venture capital, which is unfortunately finite when you're talking about the scale of tens or hundreds of billions of dollars and needs returns on some timeline, whereas Meta can do it on a longer timeline.

The second thing is that I think the optionality of having access to compute that you can then take back and pour into something is actually quite powerful. If they're a cash-generating neocloud business that can do some research on the side and then, in the future, have some breakthrough or have some distribution moat with X, or something in the Starlink relationship, or something in the Tesla relationship, or something that just means they can take advantage of it, they should, in theory, be able to then pour that compute into that thing that's just not ready yet.

And I guess what I'm saying is that I'd really actually love Joey to cover a little bit about the earnings-before-training concept, which is to say that other labs are spending a whole bunch of money training models right now that they need some return on. Right now, they're getting returns on them—namely, Anthropic and OpenAI are getting a return on these models.

But Meta is getting no return on its models outside of the Rexus stuff. There's no return on Muse Spark, for example. xAI has very limited returns. If they had 1 million subscribers to SuperGrok or something, it's quite different from approaching 1 billion MAUs for some of these consumer applications.

And so I think the play to say, “Well, Cursor's doing pretty well training on Kimi. Why don't we just let the open-source guys build us a model for the next year, and then we'll take our compute back and go run a bunch with it, instead of spending a bunch right now just to keep being in fourth or fifth place?” plays into the earnings-before-training argument, right?

Jeremie Eliahou Ontiveros

Huge disagreement.

Jordan Nanos

You disagree with that?

Jeremie Eliahou Ontiveros

Massive disagreement.

Crystal Huang

But we've been monopolizing the speech for a bit, so I'll let Joey take it.

Jeremie Eliahou Ontiveros

But I hugely disagree here.

Jordan Nanos

Well, you’ve got to explain why you disagree now, after he says something.

Jeremie Eliahou Ontiveros

Yeah, sure. It’s pretty simple. What we’re seeing right now is that there are tremendous returns to training compute. I think it’s pretty clear, and you basically want to make sure you have more than others if you want to stay in the race.

Open source versus frontier: I think it’s pretty clear that the gap is expanding, not closing, which everyone was saying last year. Open source is going to catch up, or the gap is going to close. The gap is closing. China is getting closer. No, that’s not happening. The frontier is beating the open-source models to a massive extent, as demonstrated by Anthropic’s ALR trend.

I think anyone who does production workloads sees the difference between Claude and Kimi or DeepSeek V4. If you were to take a guess, would you imagine that the gap is going to expand or is going to narrow? I would assume that it’s going to expand because one has much more compute than the other.

That goes back to the fundamental point, which is that training compute has tremendous returns. Not having compute means that you’re disadvantaged relative to competitors.

Jordan Nanos

Yeah, I think we’re agreeing about one thing, which is that training compute has massive returns if you’re in first place, but not necessarily if you’re in fifth place.

Jeremie Eliahou Ontiveros

No, not necessarily.

Jordan Nanos

Right?

Jeremie Eliahou Ontiveros

Probably? No?

Jordan Nanos

No.

Jeremie Eliahou Ontiveros

Obviously, they need talent as well, but please go ahead.

Jordan Nanos

I think they need the talent. I do think there’s potential for—I think you need both compute and talent, basically. Maybe these go hand in hand: when xAI gives up all, or close to all, of their talent, with all the co-founders leaving, and then they give up all their compute, it kind of goes hand in hand there.

Jeremie Eliahou Ontiveros

Yeah, no, I 100% agree. But I think the point is that these 2 things tell you that they’re basically out of the race. It’s going to be incredibly tough for them to come back and extract value out of the $24 trillion enterprise AI applications market.

Jordan Nanos

Okay, that came up again, so I’m going to pull that up in the S-1. Joey, maybe you can get us away from this argument and talk about this stuff here. Here’s their TAM. It wasn’t $24 trillion; it was $22.7 trillion dedicated enterprise applications. Look, all this down here is Starlink.

Jeremie Eliahou Ontiveros

This is insane.

Jordan Nanos

Anyway.

Crystal Huang

Does the race even matter, though? I feel like, at a certain point, if Meta gets so much more compute and their model gets so much better, even if they’re number 3, number 4, number 5, or whatever, it’s good enough for most people to use, right? It’s probably good enough to be replacing a lot of jobs already, so you don’t have to be number 1 or number 2 to be winning.

Jeremie Eliahou Ontiveros

No.

Jordan Nanos

No, okay. I think this kind of comes down to your perspective on how they actually use the models. We’re saying numbers 3, 4, and 5 here just to define it. From my perspective, which others may disagree with, Anthropic’s in first place right now because I believe coding is the only thing that matters. I think they’ve been proven correct.

Coding is not coding; it’s computer use. Everything a human can do with a computer, an AI can do with a computer, and therefore this is an interface to the computer, not coding. Anthropic’s in first, OpenAI is in second, and I put Cursor in third.

My experience using Composer is significantly better than using Gemini, Muse Spark, or the Grok models from xAI because, with those, it’s like you can’t use them. In some ways, I think that Cursor has both the distribution and the model to be in third place right now. The question is just how much compute Cursor needs to stay in the race.

I think having the optionality to feed them more compute in the future is compelling. They can’t use it right now, so why have it on your balance sheet if you can’t actually use it? Why not turn it into a revenue-generating asset and use it later, once you have more distribution or once you build out the training stack to improve it to the point where you can actually do these hero runs?

We’ll see, because it’s really interesting that there are 4 or 5 labs in the US testing this theory from different angles. Some are stacking compute and have a bunch of revenue. Some are stacking compute and have no revenue. Some actually have quite a bit of revenue, if you look at Cursor, and don’t have that much compute right now on a relative basis.

I’d like to see all 3 pursue it that way because I’m not sure what the right playbook is or who the winner will be. It’s going to be interesting for Anthropic to attempt to defend their number 1 position, because that’s not a position they’ve been in before. They’ve only had to play catch-up.

I think that’s actually quite hard. I think it’s quite hard to retain talent. I think it’s quite hard to keep pressing a compute advantage. I think it’s quite hard to motivate users and consumers to keep consuming more instead of getting distracted by the grass always being greener with some new feature from some competitor. It’s up to them to maintain a trillion-dollar market cap. We’ll see.

Joey Brookhart

Yeah. Jordan, I think it—Jeremie, you want to go?

Jeremie Eliahou Ontiveros

I’ve been talking a lot, man. I want other people to talk, but I had a response for Crystal.

Joey Brookhart

No.

Jeremie Eliahou Ontiveros

But go first, and—

Joey Brookhart

I think it goes into earnings before training, interest, and taxes. I think it’s really interesting. We think of EBTIT—earnings before training, interest, and taxes—as the cash operating profits that you generate from running inference.

If you want to think of training and research as CapEx, back to this conversation, I think this is a big investor question and corporate strategy question: is that return on invested capital?

Right now, we know Anthropic is obviously having massive, massive returns on the invested capital they put not only into these models, but also into coding applications—these computer applications specifically. We’re seeing more and more news of other labs trying to get into this coding market. I think there was some Microsoft news this morning on that.

There are mixed opinions here on how successful that might be, but they’re training more and more models for this because obviously people do want to use frontier models. We even saw the Meta token-maxing article. All that spend is external. Token spend is external because, to Jordan’s compute and coding point, that’s where there’s a ton of product-market fit, and they’re seeing their own ROI when they use the product.

That’s heavy. I’m guessing Jordan’s still on the call, so Jeremie, I’ll send it back over—

Jeremie Eliahou Ontiveros

Yeah, yeah.

Joey Brookhart

To you.

Jeremie Eliahou Ontiveros

I just wanted to respond to Crystal’s point. It obviously depends on how you think the market evolves, but I really like the macro framework that Malcolm has, which is: think about 2030 or 2035. What types of tasks are going to drive the bulk of the total addressable market—the dollars that people actually spend on AI?

There are tasks that have a certain amount where they’re good enough and it’s kind of finite, like translation, maybe. There’s only so much time; it doesn’t make any sense to spend more on translation.

But then there are these very open-ended tasks where the spending is pretty much infinite. In legal, for example, you could assume that if AI is really good at legal, and you want to make sure you beat your competitor, you probably want to spend more on AI than they do and have more intelligence than they do, because you want to gather more evidence and think through many different ways of coming up with a defense and whatnot.

Scientific research is typically the very open-ended use case. Healthcare is another one. For us as analysts, we try to get insights from gathering a lot of data; it’s very open-ended.

I would assume these use cases are going to drive much more spending than the finite, “It’s good enough” use cases. Really, when you think of these open-ended use cases, what matters is being able to do what you want to do in the cheapest way. If you want to do it in the cheapest way, you’re going to have to use frontier models.

That’s the whole point we made: Mythos [?] is actually two-thirds cheaper than Opus. The model, in terms of token pricing, is 6 times—I think 5 times—more expensive than Opus.

But if it's 10 times smarter, if it requires 10X fewer tokens to answer a given task, then it's actually way cheaper to complete that task with the model, right? And so I think at least that's kind of the way I view it. I think the bulk of the market is going to concentrate on the frontier models. I think if you're 3, 4, or 5, you're not going to get any dollars.

And I think if you take a step back and think, what are the signals that we've seen in 2026? Has the AI market beaten or missed versus the expectations we had in 2025? Massive beat. But then what I think is super interesting is the composition of this beat. Did everyone beat, or is it just a few companies, right?

On the frontier side, it's basically 1 company. It's just Anthropic—a monster beat. Google is probably tracking behind to some extent when you look at just Gemini adoption and how much people are spending on Gemini. OpenAI is tracking behind. Obviously, xAI and Meta are nowhere to be seen. You could have hoped that they would've had something, but they don't really.

To be fair, on the open-source side, I think it's also been a beat. I think there's been some good adoption, but the dollars spent are still pretty small. So I don't know. I think it's an interesting composition of a massive beat where it's basically all driven by 1 company, which kind of gives you a sign that it's pretty much winner-takes-all. If you're state of the art, you get the bulk of the value. And if you're not, people don't spend on you, right? I know, Joey, start praying. What do you think?

Joey Brookhart

Can you repeat the last part of that?

Jeremie Eliahou Ontiveros

Holy shit, you didn't listen to my beautiful prose. I'll convert.

Jordan Nanos

No, I'm kidding.

Joey Brookhart

I heard the most recent part.

Jordan Nanos

I got something to respond to. So, Jeremie, I think this totally makes sense, but we're also seeing massive beats, or massive reported ARR numbers, from startups serving open-source models right now. It's not just Anthropic growing. There's huge growth for Fireworks that's tied to Cursor.

Jeremie Eliahou Ontiveros

They're not that big.

Jordan Nanos

Cursor's growing really big. They're not that big.

Jeremie Eliahou Ontiveros

They're not that big, man. That is the thing. They're not that big. Some encouraging signals: I think Cursor is at $2 billion now, and they were at maybe $1-point-something billion at the end of 2025. So it's still really good growth. You could definitely say Cursor is a beat. I don't know—Windsurf, I guess, is in Google, but nowhere to be seen.

Joey Brookhart

Yeah.

Jeremie Eliahou Ontiveros

No, but I agree with you. Some of the open-source guys have had a beat, but in terms of dollar amount, it's still not very meaningful.

Jordan Nanos

Yeah, I mean, the claim from Fireworks is $315 million of ARR, right? That's—

Jeremie Eliahou Ontiveros

Eh, what is $300 million between friends?

Jordan Nanos

Guys, in the past, a startup unicorn was interesting when it had a $1 billion valuation. Now they start to approach $1 billion in ARR, and you go, "Ah, whatever. Fly on the wall."

Jeremie Eliahou Ontiveros

Relative to the size of the market, which is already above $100 billion, it's, in the grand scheme of things, not that big. Think of the positioning of the different hyperscalers. We've talked about business models and whatnot, but the beauty of Tokenomics 2.0—this magic model is just so accurate, man—is that it covers all of the bases.

Joey Brookhart

And what we'll—yeah, I mean, in 2 weeks we've gotten some pretty good feedback so far from some of these hyperscalers' customers and things like that. It's been solid. But I think right now, given that Amazon continues to win, the things that benefit them around their customer base benefit Azure pretty similarly, especially as you look at more and more token-as-a-service-type models that come on a foundry.

And then, obviously, I think Google, if they get a coding model right now, when we look at what was formerly Vertex and is now the Gemini Agent Enterprise platform, we still think the Gemini API is a pretty decent percentage of that, so they're not benefiting a ton from Claude in some of those things. But that's probably the biggest thing right now.

We still see Amazon's Bedrock token-as-a-service as a pretty significant business. I think by the end of the year it could be the majority of the AI business at Amazon. Even though Amazon lags Google, or GCP, and Azure in terms of their AI mix, at AWS, AI is a much smaller percentage of the business. But with Bedrock going to the majority of the AI business, and infrastructure-as-a-service being 80% or 90% at Azure and GCP, it's really, really advantageous.

But to that point, it's very easy for Azure to come in, add Claude as a model, and implement this token-as-a-service business. It's not that hard for people with massive customer bases to implement. It's much harder as you go down to Oracle, and then to the neoclouds like CoreWeave, to implement this at the same level, given that they don't have the massive inertia and customer bases that really benefit from the old-school software distribution—enterprise software distribution moats.

Jordan Nanos

Yeah, let me throw this chart on screen just to make the point Jeremie was making, and you're making right now, which is such a rounding error for anything but the top 3 hyperscalers when it comes to the inference endpoint business. It's a very small portion of the market that's going to the "everybody else" bucket.

Jeremie Eliahou Ontiveros

And that also goes to the same point I was mentioning earlier with regard to the composition of the market. One important point to make in that article is that the key to being a successful token-as-a-service business is actually just to have partnerships with the big labs and have access to frontier models. It's pretty simple, right?

So I guess the big disadvantage that Nebius, CoreWeave, and AIREN currently have is that they don't yet have the partnerships, and they also may not have the capital to be able to deploy GPUs without a 5-year contract. So it's kind of a function of the way the market works that a company like CoreWeave, or AIREN, has the bulk of its business contracted over multiple years. Maybe not Nebius, actually, but IREN has the bulk of its business contracted over multiple years.

Whereas Amazon is free to be more speculative, a bit more on demand. They don't have that 5-year offtake, take-or-pay arrangement locked in.

Joey Brookhart

Another thing, too, Jordan, on who wins is definitely the custom-silicon portion around Trainium and TPUs at GCP. I think that's a big thing. When we look at bringing accelerators in and some of the data-center model numbers, Azure is mostly NVIDIA. Then you look at that vertical integration at AWS and GCP, and that's another big advantage for them, especially on the margin side.

As we've seen a lot of inference get more efficient, and gross margins on inference drastically improve at the 2 major frontier labs over the last 2 years, that's definitely been another key consideration when you think about who's going to win in this market.

Jordan Nanos

Yeah. I mean, from the technical perspective, everything that we've criticized all the chip startups about—and TPU and Trainium—is always about usability. But if the entire chip, a gigawatt of Trainium at Rainier, is all just serving tokens from 1 to 3 models, then the end-user customer doesn't even necessarily need to know that much about what chip is running if they're only buying tokens, and certainly not the actual terminal user of the tokens.

When I'm using Claude Code, I have no understanding of whether my token is coming from a TPU, a Trainium accelerator, or a GPU. It doesn't make a difference, right? Anything left unsaid on the topic of Bedrock, token-as-a-service, or Anthropic's growth?

Joey Brookhart

I think we got it all, Jordan. Winners win, losers lose, and it was a clear trend.

Jordan Nanos

Winners win.

第013期——AWS利润率跳升10%,Azure和GCP停滞不前(Tokenomics)| Jordan Nanos、Jeremie Eliahou Ontiveros、Joey Brookhart、Crystal Huang — 文字稿与摘要 | BidClub