[BidClub_]
SemiAnalysis · · 44 min

Ep. 013 - AWS Margins Jump 10% While Azure and GCP Flatline (Tokenomics) | Jordan Nanos, Jeremie Eliahou Ontiveros, Joey Brookhart, Crystal Huang

Jordan NanosJeremie Eliahou OntiverosJoey BrookhartCrystal Huang

Podcast
TL;DR
  • AWS’s improving margins versus Azure and GCP come from selling Claude through Bedrock as higher-margin tokens, not simply renting accelerators. AWS is adding more than a gigawatt of capacity per quarter while margins rise; Azure and GCP remain far more exposed to lower-margin infrastructure-as-a-service. Joey’s framing: token sales retain more upside than five-year take-or-pay contracts.

  • The capacity ramp normally crushes near-term cloud margins before clusters reach “stabilized” utilization. CoreWeave-style providers pay depreciation, leases, and labor while complex systems such as GB200 wait months for activation and produce no revenue. AWS’s ability to absorb the same costs while expanding margins is therefore “a pretty good sign” for its eventual return on capital.

  • Claude’s API-heavy growth is giving both Anthropic and AWS unusually strong operating leverage. Joey cites Anthropic at $47 billion of ARR, with probably $10 billion of net-new ARR per month across March, April, and May and roughly 80% of the increase coming from APIs. Amazon was “in the right place at the right time”: Claude represents an estimated 80%-93% of Bedrock usage.

  • Anthropic’s $65 billion Series H at a $965 billion post-money valuation looks less extreme against its growth and profitability. Joey compares the roughly 20x ARR multiple with the 80x levels reached by software names in 2021 and says Anthropic is profitable excluding stock-based compensation. The caveat is significant operating deleverage if enterprises curb coding-token consumption, but “there’s no train that’s slowing right now.”

  • The SpaceX/xAI compute deal produced the episode’s sharpest disagreement over AI demand. Jeremie sees a former compute buyer becoming a supplier and possibly “giving up on the frontier race”; Jordan sees overwhelming Anthropic demand, valuable GPU-recall optionality, and a rational way for xAI to earn revenue until its own research and distribution can use the capacity.

  • Whether AI becomes winner-take-all depends on whether spending concentrates in open-ended tasks where “good enough” never arrives. Crystal argues a third- or fifth-ranked model could still replace substantial labor; Jeremie counters that legal work, science, healthcare, and analysis reward continually buying more intelligence. Jordan’s formulation is even broader: “coding is not coding, it’s computer use.”

  • The durable winners may be hyperscalers combining frontier-model access, enterprise distribution, and custom silicon. Bedrock could become the majority of AWS’s AI business by year-end, while Azure and GCP remain 80%-90% infrastructure-as-a-service in the panel’s model. Trainium and TPUs gain another advantage when token buyers never need to know which accelerator served them: “Winners win, losers lose.”

Digest · the substance, structured for research

1. Bedrock turns cloud capacity into a higher-margin product

  • Joey divides hyperscaler AI into three businesses: software such as GitHub Copilot, infrastructure-as-a-service that rents accelerators, and “token as a service,” where customers buy model access through existing cloud agreements. The last retains enterprise security, availability zones, and consolidated billing while letting the cloud capture more economics than bare chip rental.

  • Crystal argues that GPU-as-a-service lowered the old cloud moat: major users increasingly want “just the metal,” configured their way, rather than the managed platform that once made AWS, Azure, and GCP exceptionally defensible. Jordan says there are now more than 200 neo-clouds, illustrating the lower barriers to entry.

  • Bedrock restores a differentiated economic profile: Claude usage routes predominantly through AWS, and token sales carry better margins than renting infrastructure under a fixed contract. That mix explains why AWS operating margins are improving while Microsoft’s decline and Google’s remain comparatively flat.

2. AWS is outrunning the capacity-ramp margin trap

  • Crystal’s CoreWeave example separates steady-state economics from the ramp: a fully functional, stabilized cluster might produce roughly 25% operating margins and 30%-40% gross margins under a five-year take-or-pay contract. The operative word is “stabilized”—contracted revenue and the stable cost profile apply only after the cluster is functioning.

  • Before activation, the provider has already built the data center and begun paying depreciation, leases, labor, and other expenses. Equipment installation takes months, and complex systems such as GB200 have lengthened that interval, leaving assets generating costs but no revenue.

  • AWS faces the same physical constraints while bringing on more than a gigawatt per quarter, yet its margins are rising. Crystal’s inference is explicitly conditional: if margins expand during unprecedented delivery, the eventual stabilized margin and return on capital for token-as-a-service could be “extremely rich.”

3. Claude’s API mix makes Anthropic unusually measurable

  • Joey estimates Claude accounts for 80%-93% of Bedrock, while Microsoft skews toward OpenAI and Google toward Gemini. He cites Anthropic at $47 billion of ARR and probably $10 billion of net-new ARR per month across March-May, with about 80% of that incremental ARR coming from APIs.

  • Crystal finds Anthropic easier to forecast than OpenAI because API workflows expose token consumption and pricing has stayed relatively stable. OpenAI’s Q1 business was roughly 60% consumer subscriptions, where users can cancel or switch; Anthropic’s API consumption is more observable.

  • On Opus 4.8, Crystal preserves the caveat: regular pricing is unchanged, fast-mode pricing differs, and Anthropic’s chart “supposedly” shows substantially fewer hallucinations than 4.7. She had not tested it, but hoped it would stop “pulling numbers out of thin air” in their analyses and said it was close to Mido’s preview.

  • The $65 billion Series H values Anthropic at $965 billion post-money, nearly double the roughly $380-$400 billion figure discussed for February. Joey calls the approximately 20x ARR multiple less extreme than 2021 software valuations and says Anthropic is profitable excluding stock-based compensation, though coding-token cuts could cause significant operating deleverage.

4. The SpaceX/xAI deal divides the panel on compute scarcity

  • Jordan says the SpaceX filing revealed “many billions” in spending with SpaceX, with a provision allowing xAI to reclaim the GPUs. For Anthropic, the capacity could relax rate limits if compute constraints were limiting growth; for xAI, it converts capacity into a revenue-producing asset without permanently surrendering it.

  • Jeremie reads the deal as bearish for aggregate compute demand: “one player that was supposed to be a source of demand becomes a source of supply,” increasing GPU-as-a-service competition while removing an offtaker. More starkly, selling capacity to Anthropic suggests xAI may be abandoning the frontier despite evidence of extraordinary returns to training.

  • Jordan’s rebuttal is that the deal exists because Anthropic’s demand is “so overwhelming” that it must buy capacity from competitors. The recall clause also matters: xAI can monetize compute now, do research on the side, and redirect the fleet after a breakthrough or once distribution through X, Starlink, or Tesla can support a larger run.

  • The disagreement becomes capital-structure specific. Meta can fund long-duration research from a huge cash-generating business; xAI relies on finite venture capital. Joey’s “earnings before training interest and taxes” framework treats inference profit as operating cash and training as investment, raising the question of whether each lab’s model spending earns an adequate return.

5. Frontier economics may leave little value below the leaders

  • Crystal asks whether ranking even matters: a third-, fourth-, or fifth-best model might still be good enough to replace many jobs. Jordan defines the race differently—“coding is not coding, it’s computer use”—and ranks Anthropic first, OpenAI second, and Cursor third because Composer combines distribution with a usable model.

  • Jeremie rejects “good enough” for the largest pools of spending. Translation may eventually become finite, but legal work, science, healthcare, and analysis are open-ended: users can spend more intelligence to gather evidence, explore alternatives, and beat competitors. His conclusion is categorical: “If you’re three or four or five, you’re not gonna get any dollars.”

  • His Mythos example distinguishes token price from completed-task cost: he calls it “two-thirds cheaper than Opus” despite token pricing being “six times, I think, five times more expensive.” If a model is 10x smarter and needs 10x fewer tokens, it might still be the cheaper tool—concentrating demand at the frontier.

  • Compute alone does not settle the race. Jordan argues that labs need both compute and talent; Jeremie agrees, and both see xAI’s loss of talent alongside its compute as evidence that a comeback will be difficult. Jordan remains unsure which combination of revenue, distribution, and “hero runs” ultimately wins.

6. Distribution and silicon keep the market concentrated

  • Jordan challenges the winner-take-all thesis with Cursor at roughly $2 billion of ARR and Fireworks claiming $315 million. Jeremie concedes encouraging open-source adoption but calls those figures immaterial beside an AI market already above $100 billion: “What is $300 million between friends?”

  • Joey expects Bedrock could become the majority of AWS’s AI business by year-end, even though AI remains a smaller share of AWS than of Azure or GCP. Azure and GCP are still estimated at 80%-90% infrastructure-as-a-service, but Azure could add Claude quickly because hyperscaler customer relationships make token distribution relatively easy.

  • Neo clouds face a tougher loop: successful token-as-a-service requires frontier-lab partnerships, speculative capacity, and enough capital to deploy GPUs without a five-year offtake. CoreWeave, Nebius, and AIREN lack some combination of those advantages, leaving inference endpoints overwhelmingly concentrated among the top three hyperscalers.

  • Custom silicon widens the margin gap. Trainium at AWS and TPUs at GCP provide vertical integration that NVIDIA-heavy Azure lacks; usability matters less when a gigawatt serves only a few models. As Jordan puts it, a Claude Code user cannot tell whether a token came from a GPU, TPU, or Trainium—and does not need to.

Jordan Nanos

Hello, everyone. Welcome back to SemiAnalysis Weekly, episode number 13—lucky number 13. We're here with Joey, Jeremie Eliahou Ontiveros, and Crystal. We're going to talk about an article that we put out recently called “Anthropic Growth and Bedrock Mix Drive AWS Margins Higher While Peers Lag.”

That means we're going to talk about everything Anthropic, including the recent announcement of their Series H and the release of Opus 4.8, but with a focus on the infrastructure—how exactly they serve these tokens, especially with their partnership with AWS. Guys, welcome to the show. Excited to talk through this.

Joey Brookhart

Thanks, Jordan.

Crystal Huang

Thanks, man.

Jordan Nanos

All right, so let's dig into the article itself and talk a little bit about the backstory. I think a lot of people understand the concept of tokens and understand what GPUs are, but not everybody is getting their tokens from the same place. Can one of you guys give me a backstory on Bedrock? What is AWS? What are they doing for Anthropic, and how are they serving tokens with Bedrock? Joey, start with you.

Joey Brookhart

Perfect. I'll go through that, and I'll do it for all the clouds as well. If we look across all the hyperscalers, and especially the big 3—Amazon, Microsoft, and Google—there are 2 big breakouts, maybe even 3.

There's a bit on the software side, so if you look at Microsoft, things like GitHub Copilot, that's an AI software-as-a-service-type product. They have AI infrastructure as a service, where they're just renting out these accelerator chips. Then they have this token-as-a-service business, which is where they'll essentially expose these outside models or their own models to consumers to interact with.

It's a little bit of a different business because instead of just renting the underlying chip, you're renting the underlying model—buying it through your cloud provider, your CSP account, and your enterprise spending agreement. You have all the same benefits around security and availability zones, and you're able to buy third-party models and some of these first-party models through your cloud provider of choice.

At a high level, those are the 3 big buckets right now at the large hyperscalers that we see, and we're seeing some big changes or differences between them in how they've gone about it strategically. The 2 main things were Anthropic's growth and Amazon's strategy around Bedrock. Those have really driven their margins higher recently, and that was the takeaway from the article. This token-as-a-service business is obviously much better for the hyperscalers than infrastructure as a service on the AI side.

Jordan Nanos

Yeah. If you look at the big 3, they're kind of going in opposite directions right now, just in terms of their operating margin, right? That was the biggest chart from the article. I'll put it up on screen right now.

But Crystal, can you explain, when we look at AWS versus Google and Microsoft, and try to break out the cloud business, is this all to blame on Anthropic from our perspective? What's driving AWS to improve its operating margins while Microsoft's are declining and Google's stay flat here?

Crystal Huang

I feel like a lot of it is Claude usage, right? A lot of people are using Claude more, and it's mostly routing through AWS. As Joey was saying, with the token-as-a-service business model, they just have a better margin on that as opposed to the infrastructure-as-a-service model.

Yeah. I think if you take a step back, the world of clouds has really changed a lot in 2023, as you started seeing these neo-clouds, these GPU-as-a-service businesses. In the old days—which was basically 2022 and before—cloud service providers were basically 3: Amazon, Google Cloud, and Microsoft Azure. They had amazing margins and amazing returns on capital. Some new entrants, like Oracle, were trying to get in, but really the market was dominated by 3 players that had an amazing business.

Now you get to GPU as a service, and what folks started to realize is that the barriers to entry are much lower. Jordan, you're probably the best person to talk about this. You probably know the CEOs of 150 neo-clouds or 200 neo-clouds.

Jordan Nanos

Yeah. Over 200.

Crystal Huang

It is pretty insane. Obviously, some are better than others, but the point is that the market has much lower barriers to entry. The moat that cloud used to have doesn't really exist in the AI era because it's really more about infrastructure and, especially, the end users.

The whole point of cloud computing was to make IT much easier. Folks don't need to have such a big IT department internally; they can just rent through the cloud. It's super easy. Everything is well.

Now the big end users want to have much more control. You shift from platform as a service in the old days to bare metal. Folks want just the metal. OpenAI wants things the way they like them; Microsoft and Meta do too.

This token-as-a-service business is basically the first case at scale where you see an AI cloud provider having a business with a different profile. Obviously, as you have less of a moat in the GPU-as-a-service era, margins go down. Oracle is the best example. They have an RPO of half a trillion dollars, which, as of a few quarters ago—that was as of Q1—is still bigger than Amazon's. Their backlog is bigger than Amazon's, but no one gives them credit because people know it's much riskier. The returns are not the same.

We've seen empirically that every single GPU-as-a-service cloud has faced struggles when it started to ramp up its business. Our view as a firm is that we actually think the GPU-as-a-service business model is sound. Companies like CoreWeave have a sound business.

But there's this lag effect where, as you ramp up and bring more capacity online, there are lags that make your margins go down temporarily because your asset base depreciates, you have to pay data center leases, and so on and so forth. It's really interesting to see that Amazon is basically the first cloud provider that, in a time of unprecedented capacity expansion—over a gigawatt per quarter now—expands margins.

That tells you that if, during a period of accelerated capacity delivery, they can expand margins, you start to think, “Okay, what's the stabilized margin of this business?” One of the points that we make is that the stabilized margins of token as a service for Amazon are actually extremely rich. The return on capital is fundamentally much more so.

Jordan Nanos

Yeah. Let me throw this chart up on screen, actually, from the article. It's the percentage of revenue that is going to Bedrock, Bedrock being the token as a service at Amazon. You can obviously see it ramp up quite a bit at roughly the time that we're in right now—Q4 of last year and the first quarter of this year.

If we compare that to the chart I had up previously, where we're seeing their operating margins improve in the first quarter of this year, they're saying that's due to this. Can you explain in more detail why it's unprecedented to say that Amazon can bring on a gigawatt per quarter and still improve operating margins?

Crystal Huang

To understand this, you basically have to go back to why the pure bare-metal providers are seeing their margins go down. You look at CoreWeave, which is the pure play, so they're the cleanest example. Oracle is kind of the same.

When you look at this chart, essentially what this tells you is that their stabilized business does something like 25% operating margins and 30% to 40% gross margins. But the whole issue is stabilization. Stabilized means that your GPU cluster is fully functional. You're getting the monthly rent, or whatever rent, from your customer on a take-or-pay basis, so it's a flat fee. You know exactly how much revenue you're going to make. Oftentimes, it's a 5-year take-or-pay contract, again, at a fixed rate.

You know your revenue and you know your costs; everything is stable. But before getting there, you obviously have to set up the data center, which is a huge capital expense upfront. Then you have this whole process where you have the data center built, but you need to fill it with equipment. That takes a few months.

With some new types of equipment, like GB200, which is super complicated, we've seen that lag get longer and longer. That means this period of time where you depreciate your assets, pay data center rent, pay some labor, and pay a whole bunch of other costs gets longer, and you don't make any revenue because your cluster is not yet turned on. That's the whole dilemma these guys are facing: they know their business model and that stabilized structure, but they've been facing challenges ramping it up.

Some are a bit conjunctural—again, GB200—and some are more structural because there’s this lag. Amazon is the same thing, right? Like everyone else, they’re bringing on a whole lot of data centers and a whole lot of XPUs. These XPUs, in theory, should take time to bring online. Yet despite this, you’re seeing their margins go up, which is a pretty good sign for them.

Jordan Nanos

Yeah, but it’s clearly different from the others in this space, where the percentage of their total AI revenue that they’re reporting as being from token-as-a-service, as opposed to other products, is much higher than at Google and Azure. They’re not just bringing on capacity; they’re successfully selling it into the labs that are using it to serve tokens for these models.

Joey Brookhart

Yeah, and I guess the charts you showed earlier—what I forgot to mention is obviously the margin buffer that you have when you sell tokens with your infrastructure, as opposed to having a 5-year take-or-pay fixed contract with a capped upside, right? So that’s kind of the key—one of the key points of the article.

Jordan Nanos

Yeah. Can you talk a little bit about the workload mix as well? Obviously, any provider could conceptually do this, but not everybody is doing it successfully. It’s not like Azure or Google doesn’t have a token-as-a-service business. In fact, even Crusoe, CoreWeave, and Nebius are all trying to get into this business, too, but they really need a customer, and they need a customer serving the right type of workload for it to really result in a bunch of growth, I would say.

Joey Brookhart

Yeah, for Amazon specifically, they benefit from having the biggest customer base. They’ve been doing this for 20-plus years now at AWS, and people are very comfortable buying through them. People even buy infrastructure software through them, from providers like MongoDB and Snowflake. It’s a pretty large marketplace business, so customers are really comfortable with the security at this point, buying through them and having a single bill for all of this.

I think this comes back to some of the Anthropic news today, also on the ARR number of $47 billion. When we look at Q1 and even into Q2 here, a lot of Anthropic’s business mix is much, much different. Amazon is benefiting from Bedrock being 80% to 93% Claude, versus Microsoft being heavily OpenAI, obviously. Google has a lot of Gemini, which doesn’t benefit as much from a lot of these agentic coding tasks.

When you look at OpenAI and the coding percentage that’s really driving Anthropic, you’re talking probably like $10 billion a month in net new ARR over the last 3 months—March, April, and May together. Probably 80% of Anthropic’s net new ARR is in this API business. If we go to OpenAI, 60% of that business in Q1 was really consumer subscriptions.

There’s kind of a mix of factors, but Amazon was in the right place at the right time with Anthropic. They were also able to give, with Trainium2, I think, a pretty interesting deal structure for both parties. We mentioned this in the article, too: obviously, with their mix of Trainium, there’s an infrastructure-as-a-service fee component that Anthropic pays for this infrastructure, like Bedrock infrastructure. But then there are some interesting hurdles around revenue share and, really, margin share that happen.

Because Anthropic was probably at $25 million of ARR—I’m making that number up off the top of my head—in Q1 versus probably $6 million back in Q4, the numbers really made sense for both parties, and they both benefited.

Jordan Nanos

Makes sense. Maybe, Crystal, can you talk a little bit about those forecasts you guys were making? Going into the end of Q1 and your forecast for Q2, you don’t necessarily get the disclosures the same way from Anthropic, but we’ve at this point kind of been bang on with the disclosures, with the disclosed revenue figures and the margin figures, right?

Crystal Huang

Mm-hmm. I think for us, it’s a little easier to forecast Anthropic than it was OpenAI, just because so much of Anthropic’s ARR comes from the API side. Because we also use Anthropic, and there’s a lot of data out there about how people are using all of the different Claude models, it’s so much easier to predict the workflow and see what token consumption is going to look like. They’ve kept token pricing relatively stable these past 2 releases, more or less, right?

Whereas with OpenAI, a lot of their revenue comes from these subscriptions, and you never know if consumers are going to switch over to another one. So it’s a lot harder to quantify the number of users who are using the subscription when they can just cancel anytime, versus the API.

Jordan Nanos

Yeah. Can you explain a little bit about the release of Opus 4.8 and 4.7? Pricing has remained the same, but fast mode has changed, and maybe there have been some other changes in terms of how they’re doing pricing on the API.

Crystal Huang

Yeah. They said pricing is the same for regular mode, but for fast mode it’s different. Another cool thing that they said was that it doesn’t hallucinate as much, and they had a pretty cool bar chart showing that its rate of hallucination is a lot lower for 4.8 than for 4.7, supposedly. I haven’t tested it out yet, and they said that it’s super close to Mido's preview, so hopefully it’ll stop pulling numbers out of thin air when we’re using it for our analyses.

Jordan Nanos

That’d be good. That’d be good if numbers weren’t pulled out of thin air.

Joey Brookhart

That’d be great.

Jordan Nanos

Yeah.

Crystal Huang

Yeah.

Jordan Nanos

In terms of fast mode and consumer subscriptions, do you have comments there on what we’ve learned over the past few weeks or months from digging into how Anthropic is running its business on AWS? What’s maybe the biggest percentage of their revenue or their margin contribution across those different mixes—the different types of workloads that people could be consuming tokens on the API for?

Joey Brookhart

We have that in the tokenomics model, and we’re doing a big study right now on what percentage of the coding market is currently represented in token spend, especially on the API side. We’ve done a lot of work on the consumer side and some work on the B2B data, where Anthropic has just been taking a ton of share of net-new customers year to date, both in consumer subscriptions and enterprise.

But it’s really difficult, and it’s a big question among a lot of our clients: How big is the coding market currently? I think Anthropic recently said at their Financial Services Day that financial services was the second-biggest vertical. We know there’s a pretty big gap between coding and financial services, especially in terms of API spend, just from how people use this anecdotally.

There have also been a lot of recent comments on token maxing, especially at the Fortune 500s. How do you budget for that? How do you blow through that spend over time? How do you measure ROI? Things will have to change, I think, and then people will put some processes in place.

I know guys like Jeremie get insane ROI in the data center model, and his team gets that, too. Maybe there’s less policing at some other organizations that just let people go crazy. That’s another contributing factor, I think: the success of coding, and Anthropic starting to win a lot of net-new share on the subscription side in both B2B and consumer, which we saw was really interesting in Q1.

Jordan Nanos

Yeah. I guess 3 things have happened since the last time we talked about this topic on the podcast. First of all, they signed that massive deal with SpaceX, xAI, and Cursor.

Crystal Huang

The SpaceX, xAI, Cursor. Beautiful. Beautiful.

Jordan Nanos

Yeah. Cursor isn’t part of it yet because they’re trying not to change their S-1, I think. But anyway, the SpaceX S-1 revealed the many billions that they’re spending with SpaceX, with a clause to let them back out—meaning xAI has the ability to reclaim these GPUs if they want—but that should be some contribution to Anthropic’s total revenue. In other words, if they were constrained by compute for their ability to grow the business on the consumer subscription side and enforce rate limits or things like that, those should go away pretty quickly.

The second thing that happened is obviously that they raised their Series H: $65 billion in funding at a $965 billion valuation.

So they raised $65 billion at a $900 billion valuation, resulting in $965 billion post-money. I don't know why they didn't round that up to a nice, even trillion, but we'll see.

Crystal Huang

Isn't that almost double February's number—their valuation, right?

Jordan Nanos

Yeah. What was the February number?

They were at $400-something?

Crystal Huang

$380 billion or $400 billion, somewhere around there.

Jordan Nanos

I guess they need their valuation to track with their ARR growth.

Joey Brookhart

It's not a crazy multiple: 20× versus the software bubble back in 2021. Some of those were 80×, like Snowflake and Cloudflare. It's not super expensive.

Then, from the recent Wall Street Journal article and in our financials, which we also had in our model, we know that they're profitable now. When you exclude stock-based compensation, Anthropic's a pretty good business model, and they're seeing a ton of operating leverage. We also know from The Information article, year to date, that OpenAI is not seeing that as well.

It's very clear what the better business is right now and what the better business model is. There's obviously a lot of operating deleverage if people start cutting how much they're spending on coding tokens.

Jordan Nanos

Yeah.

Joey Brookhart

We're not. I mean, yeah, they're still growing at $10 billion in net new ARR a month over the last 3 months, which is pretty crazy.

Jordan Nanos

It's good.

Joey Brookhart

They're probably going to do another $10 billion in June. Who knows where they end up at the end of the year? If Mythos [?] is released as well, that probably does help.

There's no train that's slowing right now, and they're in all the right places. They're seeing the benefits of that throughout the entire business model.

Jeremie Eliahou Ontiveros

One big question here: the xAI deal, or SpaceX deal, with Anthropic—is it bullish or bearish for the market overall?

Jordan Nanos

Which market? Neocloud market?

Jeremie Eliahou Ontiveros

Compute demand overall. Everything's related, right? Stocks, compute demand, NVIDIA, all of it.

Jordan Nanos

I think it's super bullish. Neoclouds, man. Everybody wants to be a neocloud. Neocloud is the terminal business. Even AI labs want to be neoclouds, selling their compute to whoever they choose to right now.

Jeremie Eliahou Ontiveros

I disagree with this because this deal is basically one player that was supposed to be a source of demand becoming a source of supply. So now suddenly there's more competition in the supply market, which is GPU as a service, and the offtakers—there's one less. So I don't know, man.

Jordan Nanos

Yeah, I disagree with this. I think Cursor is training plenty of models on Colossus right now. I think they wouldn't have that provision in the contract to take back their GPUs if they really were going to be no source of demand in the future.

I think there's lots of demand to go around, and the reason that deal's happening is because Anthropic's demand is so overwhelming right now that they need to do these crazy things, like buy compute from their competitors, in order to be able to serve that demand. If they had a different way to serve that demand, they would be doing it, I assume.

Jeremie Eliahou Ontiveros

On the other hand, if you're xAI, Meta, or whoever—other labs that are lagging—you see Anthropic and you're like, “Whoa, bro, this is what I could do if I was at the frontier.” Right? If Groq suddenly was at the frontier, they could be a $100 billion ARR business.

So, to some extent, you would argue this should make them more bullish.

Jordan Nanos

I think, yeah.

Jeremie Eliahou Ontiveros

They should be—

Jordan Nanos

I think their S-1 has about $24 trillion of enterprise AI applications carved out as future market TAM, right? So you don't even need to be a $100 billion ARR business. You can just take that forward a few more quarters and take it to $24 trillion, no problem.

Jeremie Eliahou Ontiveros

So their TAM is $24 trillion, but that's for what? Generative AI applications?

Jordan Nanos

I'm going to bring up that chart now from the S-1. Yeah, it's like there's space and telco, and then there's a really big section for enterprise AI applications.

Jeremie Eliahou Ontiveros

Okay, enterprise AI applications. So that would basically be the TAM for Groq, right? And so they're saying, “Actually, we give up on the $24 trillion TAM. I'd rather give my compute to Anthropic. That's a better use case than fighting for a $24 trillion TAM.”

Jordan Nanos

Well, I think in some ways it's just a matter of when you can spend that money, or when you can spend that compute. In other words, there's potentially some serial nature to the development of AI progress, where you have to wait for you or all of your competitors to run a bunch of experiments to figure out the optimal model architecture and dataset mix—or just the creation of data, synthetic data, whatever it is—before it's done.

Because if we were to take a lot of these models that people are running today and try to run them on hardware from 3 years ago, the model architectures would actually run really well. All the innovations in sparsity and attention—these are huge improvements over dense models from 3 years ago.

Jeremie Eliahou Ontiveros

But they still could have—

Jordan Nanos

These are huge improvements over dense models from 3 years ago.

Jeremie Eliahou Ontiveros

They could have used that compute. If you're saying, “Maybe they had a bottleneck because they weren't able to figure out new architectures that would have enabled them to use that compute efficiently,” then they could have used that compute to do research on those specific topics, right? Less training, more research.

Or they could have used it to give a whole lot of tokens to their employees and make them much more productive at doing a bunch of research tasks. Now we know that AI can do pretty complex science problems—the OpenAI math stuff, which I know a few things about, because my dad does math for a living.

So, yeah, I mean, I don't know. That does tell you that it's an odd decision when you're kind of in the fight and giving up, when you're seeing the strongest signals we've ever seen that this is real and accelerating—and actually, $10 billion of ARR per month, right? So I don't know. It's pretty odd.

Because I feel like this is the opposite of Meta. My sense is that, to some extent, xAI is giving up on the frontier race, whereas Meta is, if anything, getting more bullish because of what they're seeing from Anthropic. They're like, “Hell yeah, this is what I bet on, and I'm going to double down after having doubled down so many times already.” So Meta is the one that's seeing Anthropic and being like, “Hell yeah. I want that.”

Jordan Nanos

Well, I think there are 2 dynamics that you're overlooking here a little bit, potentially. One is that Meta has a huge cash-generating business that they can use to fund all of this, and SpaceX really just doesn't have a business that generates hundreds of billions of dollars of free cash flow that they can pour into compute for the research bets.

So they have to do it based on venture capital, which is unfortunately finite when you're talking about the scale of tens or hundreds of billions of dollars and needs returns on some timeline, whereas Meta can do it on a longer timeline.

The second thing is that I think the optionality of having access to compute that you can then take back and pour into something is actually quite powerful. If they're a cash-generating neocloud business that can do some research on the side and then, in the future, have some breakthrough or have some distribution moat with X, or something in the Starlink relationship, or something in the Tesla relationship, or something that just means they can take advantage of it, they should, in theory, be able to then pour that compute into that thing that's just not ready yet.

And I guess what I'm saying is that I'd really actually love Joey to cover a little bit about the earnings-before-training concept, which is to say that other labs are spending a whole bunch of money training models right now that they need some return on. Right now, they're getting returns on them—namely, Anthropic and OpenAI are getting a return on these models.

But Meta is getting no return on its models outside of the Rexus stuff. There's no return on Muse Spark, for example. xAI has very limited returns. If they had 1 million subscribers to SuperGrok or something, it's quite different from approaching 1 billion MAUs for some of these consumer applications.

And so I think the play to say, “Well, Cursor's doing pretty well training on Kimi. Why don't we just let the open-source guys build us a model for the next year, and then we'll take our compute back and go run a bunch with it, instead of spending a bunch right now just to keep being in fourth or fifth place?” plays into the earnings-before-training argument, right?

Jeremie Eliahou Ontiveros

Huge disagreement.

Jordan Nanos

You disagree with that?

Jeremie Eliahou Ontiveros

Massive disagreement.

Crystal Huang

But we've been monopolizing the speech for a bit, so I'll let Joey take it.

Jeremie Eliahou Ontiveros

But I hugely disagree here.

Jordan Nanos

Well, you’ve got to explain why you disagree now, after he says something.

Jeremie Eliahou Ontiveros

Yeah, sure. It’s pretty simple. What we’re seeing right now is that there are tremendous returns to training compute. I think it’s pretty clear, and you basically want to make sure you have more than others if you want to stay in the race.

Open source versus frontier: I think it’s pretty clear that the gap is expanding, not closing, which everyone was saying last year. Open source is going to catch up, or the gap is going to close. The gap is closing. China is getting closer. No, that’s not happening. The frontier is beating the open-source models to a massive extent, as demonstrated by Anthropic’s ALR trend.

I think anyone who does production workloads sees the difference between Claude and Kimi or DeepSeek V4. If you were to take a guess, would you imagine that the gap is going to expand or is going to narrow? I would assume that it’s going to expand because one has much more compute than the other.

That goes back to the fundamental point, which is that training compute has tremendous returns. Not having compute means that you’re disadvantaged relative to competitors.

Jordan Nanos

Yeah, I think we’re agreeing about one thing, which is that training compute has massive returns if you’re in first place, but not necessarily if you’re in fifth place.

Jeremie Eliahou Ontiveros

No, not necessarily.

Jordan Nanos

Right?

Jeremie Eliahou Ontiveros

Probably? No?

Jordan Nanos

No.

Jeremie Eliahou Ontiveros

Obviously, they need talent as well, but please go ahead.

Jordan Nanos

I think they need the talent. I do think there’s potential for—I think you need both compute and talent, basically. Maybe these go hand in hand: when xAI gives up all, or close to all, of their talent, with all the co-founders leaving, and then they give up all their compute, it kind of goes hand in hand there.

Jeremie Eliahou Ontiveros

Yeah, no, I 100% agree. But I think the point is that these 2 things tell you that they’re basically out of the race. It’s going to be incredibly tough for them to come back and extract value out of the $24 trillion enterprise AI applications market.

Jordan Nanos

Okay, that came up again, so I’m going to pull that up in the S-1. Joey, maybe you can get us away from this argument and talk about this stuff here. Here’s their TAM. It wasn’t $24 trillion; it was $22.7 trillion dedicated enterprise applications. Look, all this down here is Starlink.

Jeremie Eliahou Ontiveros

This is insane.

Jordan Nanos

Anyway.

Crystal Huang

Does the race even matter, though? I feel like, at a certain point, if Meta gets so much more compute and their model gets so much better, even if they’re number 3, number 4, number 5, or whatever, it’s good enough for most people to use, right? It’s probably good enough to be replacing a lot of jobs already, so you don’t have to be number 1 or number 2 to be winning.

Jeremie Eliahou Ontiveros

No.

Jordan Nanos

No, okay. I think this kind of comes down to your perspective on how they actually use the models. We’re saying numbers 3, 4, and 5 here just to define it. From my perspective, which others may disagree with, Anthropic’s in first place right now because I believe coding is the only thing that matters. I think they’ve been proven correct.

Coding is not coding; it’s computer use. Everything a human can do with a computer, an AI can do with a computer, and therefore this is an interface to the computer, not coding. Anthropic’s in first, OpenAI is in second, and I put Cursor in third.

My experience using Composer is significantly better than using Gemini, Muse Spark, or the Grok models from xAI because, with those, it’s like you can’t use them. In some ways, I think that Cursor has both the distribution and the model to be in third place right now. The question is just how much compute Cursor needs to stay in the race.

I think having the optionality to feed them more compute in the future is compelling. They can’t use it right now, so why have it on your balance sheet if you can’t actually use it? Why not turn it into a revenue-generating asset and use it later, once you have more distribution or once you build out the training stack to improve it to the point where you can actually do these hero runs?

We’ll see, because it’s really interesting that there are 4 or 5 labs in the US testing this theory from different angles. Some are stacking compute and have a bunch of revenue. Some are stacking compute and have no revenue. Some actually have quite a bit of revenue, if you look at Cursor, and don’t have that much compute right now on a relative basis.

I’d like to see all 3 pursue it that way because I’m not sure what the right playbook is or who the winner will be. It’s going to be interesting for Anthropic to attempt to defend their number 1 position, because that’s not a position they’ve been in before. They’ve only had to play catch-up.

I think that’s actually quite hard. I think it’s quite hard to retain talent. I think it’s quite hard to keep pressing a compute advantage. I think it’s quite hard to motivate users and consumers to keep consuming more instead of getting distracted by the grass always being greener with some new feature from some competitor. It’s up to them to maintain a trillion-dollar market cap. We’ll see.

Joey Brookhart

Yeah. Jordan, I think it—Jeremie, you want to go?

Jeremie Eliahou Ontiveros

I’ve been talking a lot, man. I want other people to talk, but I had a response for Crystal.

Joey Brookhart

No.

Jeremie Eliahou Ontiveros

But go first, and—

Joey Brookhart

I think it goes into earnings before training, interest, and taxes. I think it’s really interesting. We think of EBTIT—earnings before training, interest, and taxes—as the cash operating profits that you generate from running inference.

If you want to think of training and research as CapEx, back to this conversation, I think this is a big investor question and corporate strategy question: is that return on invested capital?

Right now, we know Anthropic is obviously having massive, massive returns on the invested capital they put not only into these models, but also into coding applications—these computer applications specifically. We’re seeing more and more news of other labs trying to get into this coding market. I think there was some Microsoft news this morning on that.

There are mixed opinions here on how successful that might be, but they’re training more and more models for this because obviously people do want to use frontier models. We even saw the Meta token-maxing article. All that spend is external. Token spend is external because, to Jordan’s compute and coding point, that’s where there’s a ton of product-market fit, and they’re seeing their own ROI when they use the product.

That’s heavy. I’m guessing Jordan’s still on the call, so Jeremie, I’ll send it back over—

Jeremie Eliahou Ontiveros

Yeah, yeah.

Joey Brookhart

To you.

Jeremie Eliahou Ontiveros

I just wanted to respond to Crystal’s point. It obviously depends on how you think the market evolves, but I really like the macro framework that Malcolm has, which is: think about 2030 or 2035. What types of tasks are going to drive the bulk of the total addressable market—the dollars that people actually spend on AI?

There are tasks that have a certain amount where they’re good enough and it’s kind of finite, like translation, maybe. There’s only so much time; it doesn’t make any sense to spend more on translation.

But then there are these very open-ended tasks where the spending is pretty much infinite. In legal, for example, you could assume that if AI is really good at legal, and you want to make sure you beat your competitor, you probably want to spend more on AI than they do and have more intelligence than they do, because you want to gather more evidence and think through many different ways of coming up with a defense and whatnot.

Scientific research is typically the very open-ended use case. Healthcare is another one. For us as analysts, we try to get insights from gathering a lot of data; it’s very open-ended.

I would assume these use cases are going to drive much more spending than the finite, “It’s good enough” use cases. Really, when you think of these open-ended use cases, what matters is being able to do what you want to do in the cheapest way. If you want to do it in the cheapest way, you’re going to have to use frontier models.

That’s the whole point we made: Mythos [?] is actually two-thirds cheaper than Opus. The model, in terms of token pricing, is 6 times—I think 5 times—more expensive than Opus.

But if it's 10 times smarter, if it requires 10X fewer tokens to answer a given task, then it's actually way cheaper to complete that task with the model, right? And so I think at least that's kind of the way I view it. I think the bulk of the market is going to concentrate on the frontier models. I think if you're 3, 4, or 5, you're not going to get any dollars.

And I think if you take a step back and think, what are the signals that we've seen in 2026? Has the AI market beaten or missed versus the expectations we had in 2025? Massive beat. But then what I think is super interesting is the composition of this beat. Did everyone beat, or is it just a few companies, right?

On the frontier side, it's basically 1 company. It's just Anthropic—a monster beat. Google is probably tracking behind to some extent when you look at just Gemini adoption and how much people are spending on Gemini. OpenAI is tracking behind. Obviously, xAI and Meta are nowhere to be seen. You could have hoped that they would've had something, but they don't really.

To be fair, on the open-source side, I think it's also been a beat. I think there's been some good adoption, but the dollars spent are still pretty small. So I don't know. I think it's an interesting composition of a massive beat where it's basically all driven by 1 company, which kind of gives you a sign that it's pretty much winner-takes-all. If you're state of the art, you get the bulk of the value. And if you're not, people don't spend on you, right? I know, Joey, start praying. What do you think?

Joey Brookhart

Can you repeat the last part of that?

Jeremie Eliahou Ontiveros

Holy shit, you didn't listen to my beautiful prose. I'll convert.

Jordan Nanos

No, I'm kidding.

Joey Brookhart

I heard the most recent part.

Jordan Nanos

I got something to respond to. So, Jeremie, I think this totally makes sense, but we're also seeing massive beats, or massive reported ARR numbers, from startups serving open-source models right now. It's not just Anthropic growing. There's huge growth for Fireworks that's tied to Cursor.

Jeremie Eliahou Ontiveros

They're not that big.

Jordan Nanos

Cursor's growing really big. They're not that big.

Jeremie Eliahou Ontiveros

They're not that big, man. That is the thing. They're not that big. Some encouraging signals: I think Cursor is at $2 billion now, and they were at maybe $1-point-something billion at the end of 2025. So it's still really good growth. You could definitely say Cursor is a beat. I don't know—Windsurf, I guess, is in Google, but nowhere to be seen.

Joey Brookhart

Yeah.

Jeremie Eliahou Ontiveros

No, but I agree with you. Some of the open-source guys have had a beat, but in terms of dollar amount, it's still not very meaningful.

Jordan Nanos

Yeah, I mean, the claim from Fireworks is $315 million of ARR, right? That's—

Jeremie Eliahou Ontiveros

Eh, what is $300 million between friends?

Jordan Nanos

Guys, in the past, a startup unicorn was interesting when it had a $1 billion valuation. Now they start to approach $1 billion in ARR, and you go, "Ah, whatever. Fly on the wall."

Jeremie Eliahou Ontiveros

Relative to the size of the market, which is already above $100 billion, it's, in the grand scheme of things, not that big. Think of the positioning of the different hyperscalers. We've talked about business models and whatnot, but the beauty of Tokenomics 2.0—this magic model is just so accurate, man—is that it covers all of the bases.

Joey Brookhart

And what we'll—yeah, I mean, in 2 weeks we've gotten some pretty good feedback so far from some of these hyperscalers' customers and things like that. It's been solid. But I think right now, given that Amazon continues to win, the things that benefit them around their customer base benefit Azure pretty similarly, especially as you look at more and more token-as-a-service-type models that come on a foundry.

And then, obviously, I think Google, if they get a coding model right now, when we look at what was formerly Vertex and is now the Gemini Agent Enterprise platform, we still think the Gemini API is a pretty decent percentage of that, so they're not benefiting a ton from Claude in some of those things. But that's probably the biggest thing right now.

We still see Amazon's Bedrock token-as-a-service as a pretty significant business. I think by the end of the year it could be the majority of the AI business at Amazon. Even though Amazon lags Google, or GCP, and Azure in terms of their AI mix, at AWS, AI is a much smaller percentage of the business. But with Bedrock going to the majority of the AI business, and infrastructure-as-a-service being 80% or 90% at Azure and GCP, it's really, really advantageous.

But to that point, it's very easy for Azure to come in, add Claude as a model, and implement this token-as-a-service business. It's not that hard for people with massive customer bases to implement. It's much harder as you go down to Oracle, and then to the neoclouds like CoreWeave, to implement this at the same level, given that they don't have the massive inertia and customer bases that really benefit from the old-school software distribution—enterprise software distribution moats.

Jordan Nanos

Yeah, let me throw this chart on screen just to make the point Jeremie was making, and you're making right now, which is such a rounding error for anything but the top 3 hyperscalers when it comes to the inference endpoint business. It's a very small portion of the market that's going to the "everybody else" bucket.

Jeremie Eliahou Ontiveros

And that also goes to the same point I was mentioning earlier with regard to the composition of the market. One important point to make in that article is that the key to being a successful token-as-a-service business is actually just to have partnerships with the big labs and have access to frontier models. It's pretty simple, right?

So I guess the big disadvantage that Nebius, CoreWeave, and AIREN currently have is that they don't yet have the partnerships, and they also may not have the capital to be able to deploy GPUs without a 5-year contract. So it's kind of a function of the way the market works that a company like CoreWeave, or AIREN, has the bulk of its business contracted over multiple years. Maybe not Nebius, actually, but IREN has the bulk of its business contracted over multiple years.

Whereas Amazon is free to be more speculative, a bit more on demand. They don't have that 5-year offtake, take-or-pay arrangement locked in.

Joey Brookhart

Another thing, too, Jordan, on who wins is definitely the custom-silicon portion around Trainium and TPUs at GCP. I think that's a big thing. When we look at bringing accelerators in and some of the data-center model numbers, Azure is mostly NVIDIA. Then you look at that vertical integration at AWS and GCP, and that's another big advantage for them, especially on the margin side.

As we've seen a lot of inference get more efficient, and gross margins on inference drastically improve at the 2 major frontier labs over the last 2 years, that's definitely been another key consideration when you think about who's going to win in this market.

Jordan Nanos

Yeah. I mean, from the technical perspective, everything that we've criticized all the chip startups about—and TPU and Trainium—is always about usability. But if the entire chip, a gigawatt of Trainium at Rainier, is all just serving tokens from 1 to 3 models, then the end-user customer doesn't even necessarily need to know that much about what chip is running if they're only buying tokens, and certainly not the actual terminal user of the tokens.

When I'm using Claude Code, I have no understanding of whether my token is coming from a TPU, a Trainium accelerator, or a GPU. It doesn't make a difference, right? Anything left unsaid on the topic of Bedrock, token-as-a-service, or Anthropic's growth?

Joey Brookhart

I think we got it all, Jordan. Winners win, losers lose, and it was a clear trend.

Jordan Nanos

Winners win.

Ep. 013 - AWS Margins Jump 10% While Azure and GCP Flatline (Tokenomics) | Jordan Nanos, Jeremie Eliahou Ontiveros, Joey Brookhart, Crystal Huang | BidClub