[BidClub_]
SemiAnalysis · · 51 min

Ep. 020 - Anthropic vs OpenAI Usage, Margins, Meta Compute, Future of MSL (Tokenomics)

CrystalMax KanJoey BrookhartJordan Nanos

YouTube
TL;DR
  • Enterprise token austerity is aimed at the wrong workloads: coding consumes the budget, while premium-model emails are effectively free. Crystal says most employees never approach their caps, so pooled company budgets make more sense than per-person ceilings. Max calls engineering-only access a “caste system,” while Joey says the top 1% of companies already spend roughly $100,000 per employee annually on AI and continue when the ROI survives scrutiny.

  • The $200 coding subscriptions are heavily subsidized and may be loss-making at power-user utilization. The team estimates Codex can provide about $12,000 of API-equivalent usage and Anthropic about $8,000; break-even utilization falls to 10% for Claude Max 20x and 5.7% for OpenAI Pro 20x. At likely utilization levels, Max says these plans “might just be negative margin.”

  • Anthropic’s enterprise/API-heavy mix is already producing operating leverage that OpenAI’s free consumer base cannot match. More than 80% of Anthropic ARR is API-based and historically more than 90% enterprise; Joey says some profitability is on a non-GAAP basis excluding stock compensation, while operating profit was positive in Q2 and could reach $1 billion-plus in Q3. OpenAI’s roughly 950 million weekly users convert only about 6% to paid plans, and servicing the free population lowers blended gross margin by around 20 points.

  • OpenAI has nevertheless returned to a genuine two-horse race because model quality, not incumbency, remains decisive. After likely flat ARR growth in March and April, 5.5 became the predicted inflection point; some people now judge 5.6 “as good as Opus 4.8” at half the price and likely with a smaller model. Net-new monthly ARR is reportedly comparable with Anthropic, prompting Dylan’s reversal: “I am nothing if not flexible.”

  • Compute optionality may matter as much as compute ownership for xAI and Meta. Max’s revised view is that surplus capacity can be rented at 3x-4x market rates while teams retain enough to prove they can reach the frontier, protected by 90-day clawback clauses. Meta compute could similarly justify heavier 2027-28 capex and become a profitable fallback if MSL disappoints; by contrast, Max reads Google’s non-recallable, long-term TPU commitments as evidence of weak conviction in building RSI.

  • Hyperscalers win token-as-a-service distribution through existing enterprise relationships, while independent inference providers can prosper without winning the market. Anthropic’s indirect token volume may have risen from roughly 5% to 20% of its business in six months, and AWS may collect a 20%-30% revenue share largely because customers already buy through Bedrock. Together, Fireworks and Baseten can still grow rapidly because inference could be “the largest market ever,” even if open-source volume grows more slowly than frontier-model volume.

  • RL environments are becoming both the next capability bottleneck and an unusually lucrative engineering market. Frontier-lab data budgets could exceed $10 billion this year — about 10x last year — and might increase another 10x in 2027; top coding tasks command well over five figures each, while people who are very good at creating RL tasks can earn seven figures or more annually. The episode makes its exponential thesis falsifiable with a $400 billion Anthropic ARR over-under for end-2027: Max and Dylan take the over, Crystal, Joey and Jeremy the under, and the loser must explain publicly why they were wrong.

Digest · the substance, structured for research

1. Token caps are targeting the wrong workloads

  • Crystal’s field read is that social media exaggerates token-budget pain: only a handful of power users at most companies approach the limit. Because usage is highly uneven, one company-wide monthly pool is more rational than identical per-person allowances.

  • Smaller organizations are already routing around limits: a cheap model processes and condenses material before an expensive Anthropic or OpenAI call. Where internal models do not count toward the budget, employees “really push those to the limit” and reserve metered models for the final work.

  • Max’s objection is that many policies optimize trivia. Whether a salesperson uses Opus or Sonnet for an email “basically does not matter” from a token perspective; coding is “by far the most token-hungry use case.”

  • Restricting Claude Code or Codex to engineers creates what Max calls a “caste system,” denying other functions the chance to discover valuable workflows. Joey’s supporting evidence: he estimates that 70% or more of Anthropic’s API-side ARR is in coding, while Meta contributes only about 3%-5% of total Anthropic ARR — usage is heavily driven by power users.

2. Subscription arbitrage exposes the real margin pool

  • Max’s practical rule is blunt: “Anybody who can use a subscription definitely should,” because it is so heavily subsidized. Some startups simply charge five or however many individual accounts they need to the company card and add another whenever they hit a limit; Max says buying 10 subscriptions per person can be more cost-effective than API credits.

  • At $200 monthly, the team estimates Codex supplies roughly $12,000 of API-equivalent credits and Anthropic around $8,000. The corresponding break-even utilization is about 20% for the Claude Pro and 5x plans, 10% for Claude Max 20x, 11.4% for OpenAI Plus and Pro 5x, and 5.7% for Pro 20x.

  • Max believes average 20x-plan utilization exceeds those thresholds, making the products potentially loss-making rather than merely lower-margin. Anthropic API gross margin is estimated at 85%+, and Anthropic enterprise subscriptions include no usage: customers pay the fee, then consume at API prices.

3. Anthropic’s API mix converts model demand into profit

  • More than 80% of Anthropic ARR is API-based and historically more than 90% comes from enterprise. Joey says some of the company’s profitability is on a non-GAAP basis excluding stock-based compensation, while operating profit was positive in Q2 and could reach $1 billion-plus in Q3.

  • That profitability partly reflects an inability to hire or reinvest in training as quickly as revenue is growing. Joey’s prescription is not to race toward 30% EBITDA margins, but to spend the gross-profit advantage extending model capability: “continue to invest as much as you can into training.”

  • OpenAI had the inverse mix in Q1: roughly 60% consumer ARR, around 950 million weekly active users, only 6% paying, and most subscribers on $20 plans. Free users lower blended gross margin by about 20 points; Dylan notes reported daily active-user counts rising from 7 million to 8 million against more than 900 million free weekly users.

  • The team believes Codex, 5.5 and 5.6 may already have flipped OpenAI to roughly 40% consumer and 60% enterprise in Q2. Its closing wager captures the stakes: Anthropic ARR above or below $400 billion at end-2027, with Max and Dylan over, Crystal, Joey and Jeremy under.

4. Better models have pulled OpenAI back into contention

  • Crystal says Claude Code initially retained users through habit — “that’s just what I’ve been using” — but ChatGPT’s consumer traction is now trickling into work. Former Claude Code power users who ignored Codex are beginning to experiment with it.

  • Max believes OpenAI ARR growth was likely flat in March and April, setting off “alarm bells,” before 5.5 delivered the expected inflection at roughly Opus 4.8 quality. Some people now rate 5.6 level with Opus 4.8 despite half the price; the team also believes it is likely much smaller.

  • Dylan’s change of mind is model-led: he now defaults to OpenAI for his work, but rejects the Codex app’s thread model, weak remote-SSH experience and missing tab workflow. He remains a CLI user because the app appears aimed at global “vibe coders” making more B2B SaaS.

5. Frontier compute is most valuable when it can be recalled

  • Max dismisses Grok 4.5 and Muse Spark 1.1 as adequate scaling proofs rather than frontier products: “unless you can deliver a true frontier model… who cares.” He plans to shift his token usage to either, though their existence proves a nonzero chance of catching up.

  • His harsher ranking puts Gemini “clearly in fifth place,” potentially forever unless Gemini 3.5 Pro beats current industry chatter. The distinction is not merely installed compute, but whether management preserves the option to redirect that compute toward frontier training.

  • Max revised his interpretation of Elon renting nearly all of Colossus 1 to Anthropic. The strategy is to leave the internal team enough capacity to prove scaling works, rent the rest at perhaps 3x-4x market rates, and retain a 90-day clawback if the model reaches the frontier.

  • Dylan’s pushback — worth keeping — is that a provider may only claw capacity back once before losing customer trust. Max agrees that one recall is ideal, but argues Anthropic might return anyway if Elon later relents and Anthropic remains desperate for capacity: “Why would they say no?”

6. Meta compute is a capex backstop, while distribution powers token resale

  • Meta can apply the same structure through recallable rentals, token-as-a-service, or capacity assigned to MSL, preserving the option to return everything to MSL. Joey sees this as a backstop permitting more aggressive capex in 2027 and potentially 2028 if MSL is deemed unsuccessful.

  • Selling one’s own frontier-model tokens remains the best business, but premium compute rental is hardly a consolation prize. The discussed SpaceX-plus-Google transaction priced near 4x market, versus neocloud economics around $12 billion per gigawatt and potentially approaching $50 billion.

  • Anthropic’s revenue density illustrates why buyers can tolerate those prices: about $16 million per megawatt last year, more than doubled over three quarters, with some expecting another doubling within two or three quarters. “Token throughput’s been pretty impressive,” and new models monetize at higher prices.

  • Dylan initially read an older chart as evidence that Microsoft Foundry was doing badly; Joey corrected him that it was roughly 90% OpenAI until recently and should be revised upward after recent API gains. Bedrock and Azure remain attractive because Fortune 500 and Global 2000 customers already have large cloud relationships, credits or ELAs, plus security and compliance.

7. RL environments are becoming the next capability scaling constraint

  • Max calls RL “probably the most important scaling law” for current capability gains. Many people believe sufficient environments could teach models nearly any computer-based human task through repeated attempts; aggregate frontier-lab data budgets may exceed $10 billion this year, about 10x last year, and potentially 10x again in 2027.

  • This market has almost no demand constraint: labs will not reject excellent data merely because it is expensive. Max attributes Anthropic’s coding lead before 5.6 partly to buying coding environments far more aggressively than rivals, which are now trying to catch up.

  • Modern environment creation bears little resemblance to bounding boxes or content labels. A strong software task may require a really good human engineer’s full day, integration-test verifiers and a rubric, plus a prompt that is simultaneously unambiguous enough for fair training and natural enough to resemble a real request.

  • Labs will pay well over five figures for one top coding task, and people who are very good at creating RL tasks can make seven figures or more annually. Max’s own experience is that creating data is much harder now than it was eight months ago: identify a precise failure mode, let the model attempt it ten times, then iterate. “Your first thought is almost certainly too easy — try making it 10 times harder.”

Dylan Patel

Hello, everyone. Welcome back to SemiAnalysis Weekly. We're here for Episode 20 with the Tokconomics team. Things are moving fast. We are recording this on Wednesday, July 15. I assume that by the time this episode comes out, there are going to be a lot more model releases and disclosures from these companies about their profitability, revenue, and how many active users they've got.

Regardless, we're going to record a point in time right now, talking about token budgeting, Meta's compute strategy, MSL futures, the release of Fable, Soul, and Anthropic's profit margins. I'm excited to dig in.

With me today, we've got Max. How are you doing, buddy?

Max

Doing great. Thanks for having me.

Dylan Patel

Welcome. We've got Crystal. How's it going, Crystal?

Crystal

Hey, Jordan.

Dylan Patel

And we've got Joey.

Joey

Hey, Jordan.

Dylan Patel

Cool. So we're going to start with token budgeting. Crystal, over to you. We've had lots of conversations with enterprises over time, just asking them what's going on. At SemiAnalysis, we're still token-maxing, but others are moving into a time of austerity. Can you take me through a little bit of what you found with this article?

Crystal

There are a lot of enterprises right now that are cracking down on their budgeting because they think their employees are spending too much. A lot of people are getting scared because they're thinking, "What if I don't have enough tokens to do my work?" I don't think that's necessarily true, though, and it applies to most people, right? A lot of people aren't even getting close to that limit. There are probably a handful of power users at each of these companies who will really be affected.

Other than that, I don't think it has had as big of an impact on most people's workflows as it has been portrayed on social media. We've seen a lot of different strategies that companies have been imposing, whether it's on a per-person basis or a monthly budget for the entire company. I think the per-company basis obviously makes a lot more sense than a per-person basis, given that some users will get more use out of it than others.

Dylan Patel

How do you see people actually enforcing this? Clearly, you can burn through a budget faster if you're going with the super-ultra-premium Max thinking model, but in some cases, when you're actually enforcing a budget, you take different approaches, right?

Crystal

Yeah. I've seen some people do this. I was talking to people at a smaller company, and they have a much smaller budget. The way they do it is by trying to optimize how they use the more expensive tokens from Anthropic or OpenAI. They use a cheaper model to process a lot of the information and condense it down into a smaller prompt, and then they use the more expensive models.

They also use something that their company isn't counting. There are a lot of discrepancies in terms of which models actually count. Some enterprises are counting their own models in the budget, while others aren't. If they aren't counting them, a lot of people will really push those models to the limit, use them a lot, and then use the ones that actually cost money.

You're seeing people—maybe, Max, I can bring you in on this one—using more tokens for coding rather than other use cases, right? That seems to be the market that's growing the fastest.

Max

Yeah, I think coding and software engineering in general is by far the most token-hungry use case. I actually think this is why a lot of the token-budgeting discourse is pretty uninformed. I think a lot of the budgets themselves are pretty uninformed.

I hear people saying that they want to make sure their sales guys aren't using Opus or Fable to write emails, so they should definitely be using Sonnet instead. It's like, dude, generating your email is basically free from a token perspective. It does not matter if you use Opus or Sonnet to write an email. I have no idea why this is the policy you're enforcing to try to reduce your token spend.

I've also heard lots of stories about certain companies where only the engineering team is allowed to use Claude Code or Codex, and everyone else only gets something like Codework or something like that. I think that's pretty shortsighted. You shouldn't have a caste system where your engineers are at the top and everyone else is forced to use something like Demoris [?]. You should really be giving everyone the opportunity to try these tools and figure out how they can become more productive with them.

Dylan Patel

Joey, are you using these models right now? How would you feel if I got access to them and you didn't?

Joey

I would probably be pretty pissed. I do have some access, though not as much as Max and other people at SemiAnalysis. That's probably just the nature of our work, or my work right now, and that's changing a little bit.

But, yeah, I'd be pretty mad because I still use Opus 4.8 a lot for AI research tasks that pull in a lot of different data sources and do more research-type work.

Dylan Patel

How do you think about the ROI question? A lot of the motivation behind token budgeting is clearly that somebody isn't seeing a return on this investment, or they're just not seeing it yet. But in a lot of the research you're doing, you're seeing that there are a lot of benefits to using these models, right?

Joey

Yeah, I think with the market being so coding-focused today, probably 70% or more of Anthropic's ARR on the API side is in coding. Second, from what we've heard, it's very power-user- and power-company-driven. Coding is pretty widespread, but there isn't one major big company for Anthropic. Meta is probably 3% to 5% of total Anthropic ARR, but it's heavily driven by power users.

If you're at a point where you're a power user at one of these companies, I think the Ramp data shows that the top 1%—the 99th percentile of companies—spend $100,000 per year on AI per employee. If you've gotten to the point where you're spending that much, you're obviously getting pretty good ROI.

Our conversations with a lot of enterprises have shown that these top engineers could go well above their limited budget, whatever their budget limit was. But if you were someone who was spending an insane amount of money, you were getting budgeted pretty quickly if you weren't seeing ROI. There are always projects that you'll try that don't get ROI. You'll scrap those and move on.

It was pretty obvious that, because the market is so coding-focused, people are getting ROI and continuing to spend. I think we continue to hear that net-new ARR for both Anthropic and now OpenAI—with Codex and 5.5 and 5.6—is broadening. Since late March, spend has been broadening, and there's clearly a lot of API token spend today.

There are a lot of organizations that aren't traditionally doing a ton of development-type work. They're only allowing their middle- and back-office workers to use Copilot 365 or SAT. But I think the market will continue to broaden out.

Dylan Patel

Do you guys have a take on the balance between API usage in applications and coding plans when it comes to people's individual usage? When you're talking to people, do you ask them whether they're using a Claude Code subscription versus paying per token through the API?

Max

Anybody who can use a subscription definitely should use a subscription because it's so subsidized. OpenAI and Anthropic are very aware of this, so when they have their enterprise plans—SemiAnalysis is on an Anthropic enterprise plan—they don't let you use a subscription. They force you to pay per token because they know that's where the margin is.

In general, anyone who can use a subscription does use a subscription. It's just the people who need more limits than that who pay for the API.

Joey

I've even talked to a few people at startups, and they're not on the team or enterprise plan at all. They just buy 5—or however many—Claude or OpenAI subscriptions they need per month, charge it to the company card, and keep doing that. If they hit the limit on however many accounts they already have, they just buy another one for the month because it's more cost-effective for them to operate that way than to use API billing.

If you're a startup that's young enough that you don't care about security or any of the other enterprise guarantees, just absolutely buy 10 subscriptions per person rather than paying for API credits.

Max

I think the recent study we did on this had a decently viral tweet. For Codex, I think $200 a month gets you around $12,000 worth of API credits, and for Anthropic, it was $200 a month for around $8,000. Obviously, that's a no-brainer for someone who is able to pick between those two.

Dylan Patel

Can you dig into that a little bit more, Max? When you're talking about the margin profile of these businesses and considering the subscription plan versus the API, what does that look like? I'm going to bring up the chart on screen that I think is the key one from that tweet thread.

Max

Yeah.

Max

Obviously, when you hear that a $200-a-month plan can generate $8,000 worth of API credits, the subscription is going to be much lower margin if the average user has anywhere near 100% utilization. The key question is: What is the actual average utilization across our entire user base?

This chart that Jordan showed us here is the break-even percentage for the various plans. We can see that for the Claude Max 20x plan, it’s just 10%. For the Pro and 5x plans, it’s a little better at 20%. OpenAI, because they are even more generous with their subscription limits, is slightly lower at 11.4% for the Plus and Pro 5x, and just 5.7% for the Pro 20x.

We’re pretty confident that the average utilization, especially for these 20x plans, is much higher than 10% and 5.7%, respectively, for Anthropic and OpenAI. That would mean that, forget having lower margins than API—which we think is likely 85%+ for Anthropic at least—it might just be negative margin in general for these 20x plans. Of course, these businesses want to move as much volume as possible to API.

I think maybe, Joey, since she wrote the newsletter on Anthropic’s business, this would be a good segue to talk about why we think Anthropic is potentially in a stronger position than OpenAI.

Joey

Yeah. And I think, too, once you move on to enterprise subscription plans at Anthropic, there’s no usage included in your subscription fee. It’s all done at API pricing.

At Anthropic, what we found was really interesting is that the business today is so much more—and this is changing—but more than 80% of ARR is on the API side. As Max showed, that’s really high margin, so they’ve gotten to a point where they’re profitable today. Some of that is on a non-GAAP basis, so excluding stock-based compensation, but operating profit was positive in Q2 and could be profitable to the tune of $1 billion-plus in Q3.

Some of that is a function of the fact that they’ve just grown so much. I don’t think they can hire enough or plow enough back into training as they’d like. But because of that margin, that gross profit advantage that they have today, they can plow as much back into training as they want and can extend their model advantage and model capability advantage.

We think that’s something they should obviously try to do. Just because it’s very profitable for them today doesn’t mean they should run at 30% EBITDA margins right away or as quickly as possible and show that this is a very profitable business model. They should continue to invest as much as they can into training.

Dylan Patel

Yeah. Can you talk about the mix between consumer and enterprise, assuming that consumer is those paid plans and enterprise is pay-per-use on the API?

Max

Yeah. When you look at these things, clearly the bulk of Twitter users are using the subscription plans. But then you get one really big customer who’s spending so much more money on API tokens that it kind of drowns out everything else, and that, I guess, makes it a good business.

Joey

Yeah. And because Anthropic has historically been so focused on enterprise, it was 90%+ enterprise. ChatGPT, by contrast, started heavy in consumer. Back in Q1, 60% of ARR was in consumer, and only 6% of the 950 million weekly active users end up actually paying for a plan. Most people pay for the $20-a-month plan.

It’s a much different business mix because OpenAI has to subsidize—or not subsidize, but there’s a cost to serving—the other 900 million weekly active users who are free. That lowers their gross margins on a blended basis by about 20 points.

That’s changing. We think that, with some of the rise in OpenAI’s API business since late March with Codex, and then 5.5 and 5.6, the 60/40 consumer-enterprise split has flipped now to 40/60 consumer versus enterprise here in Q2. From what we’re hearing in the channel, especially out of token-as-a-service businesses like Bedrock and Foundry, OpenAI has recognized this and is starting to shift. Even net new ARR is starting to look a little more even between Anthropic and OpenAI.

Dylan Patel

Yeah. And it’s kind of shocking. With the release of GPT-5.5 and GPT-5.6, people like Tibo on X who are tweeting about the daily active user counts are saying that they’re celebrating going to 7 million and then 8 million active users. But to be clear, they’ve got over 900 million weekly active users on the free tier of ChatGPT.

They’ve got a lot of room to grow just to get people using the number-one paid product. Maybe they’ve got a few extra users using other paid products, but I would assume the bulk of people using the paid products are using Codex, right? So that’s not a great conversion rate in terms of who’s actually using the full paid product from the base of 900 million.

Joey

Yeah, it’s funny. You actually see it, too, in the paying rates for even the consumer version of Claude. Claude is like 5.9% of free users end up paying, so it’s a much more focused user base.

If we think of those 900 million weekly active users, a lot of it is Google Search-type replacement, things like that. And then, obviously, because of that consumer focus, they’ve been working on advertising a lot. Consumers traditionally monetize poorly: people aren’t willing to pay, and the retention curves are typically pretty poor. How many people are paying 3 months later, 6 months later, or 12 months later?

Ads are a tough business model. It’s tough to scale, but maybe that’s the eventual path. They’ve put out some pretty crazy targets for 2030 advertising revenue. But yeah, consumer is a tough business. Enterprise, given the margins on tokens for API, is obviously the place to be.

I think Anthropic had a pretty good lead because they were so levered to coding early in the first 6 months of this year, and it seems like it’s shifting to more of an equal two-horse race here in mid-July.

Dylan Patel

Yeah. Crystal, what are you seeing in that two-horse race? Are people that you’re talking to using Anthropic’s Claude Code, Codex, or both?

Crystal

I feel like a few weeks ago, when I was talking to people, they were pretty heavy on using Anthropic’s Claude Code specifically. If I asked them why, they would say, “That’s just what I’ve been using, right? And I’m used to using it now.”

But now, as people in the consumer segment start using ChatGPT a lot more than Claude, that effect in the consumer sector starts trickling into enterprise a little bit. These people who were previously Claude Code power users and didn’t touch Codex at all are kind of like, “Okay, what if I consider tinkering with Codex?”

Even though consumer is a really small portion of the pie and the margins obviously aren’t as strong as enterprise, and historically speaking OpenAI hasn’t done as well in the consumer market as Anthropic has, they have that traction in the consumer space that’s eventually going to trickle in. We’re already starting to see that just this past week.

Dylan Patel

Max, what’s your take? OpenAI versus Anthropic. Where do things stand for you?

Max

I think I remember that the last time I was on this podcast, we were talking about how it was right when 5.5 came out, and we were talking about how things were looking really dire for OpenAI at the start of the year. I think ARR growth was likely flat in March and April, which meant the alarm bells were definitely ringing in Sam’s head when that happened.

But then we said that 5.5 would be an inflection point. This was actually a really good model, probably Opus 4.8-level. I think that prediction has played out even more quickly and more optimistically than we would have guessed.

Some people here are saying that 5.6 is as good as Opus 4.8, sort of just straight up, even though it’s only half the price. We also think it’s a much smaller model, which would be a testament to OpenAI’s training abilities.

As Joey mentioned earlier, net new ARR month over month is comparable between OpenAI and Anthropic now, which is pretty shocking. People were counting OpenAI out a little bit, and I think Jordan might have been one of those people. But they’re back, guys. They’re back.

Dylan Patel

I am nothing if not flexible. It’s always been about the model, guys. Everybody’s like, “When are you going to stop hating on OpenAI?” As soon as they ship a model that’s good—and it’s good.

For the things that I want to do, I’m basically default OpenAI at this point, despite the application not doing what I want it to do with multiple tabs and stuff like that.

Dylan Patel

Wait, are you using the Codex app or the CLI?

Dylan Patel

CLI.

Dylan Patel

What’s your form factor—CLI? I feel like they don’t want you to use the CLI, dude. They very actively want you to use the app. What’s pulling you back?

Dylan Patel

Tabs. I need multiple tabs going.

I think the way you’re supposed to do this—because I agree, the first time I tried using Codex, I was like, every single time I open a new chat, it’s just lost in the ether in the left-hand sidebar of all those tabs.

Dylan Patel

I think how you're supposed to do it is you have 1 pinned thread per project you're working on, and then you ask that thread to spin up subthreads anytime you have a discrete ask. You actually never even interact with the vast majority of threads in the left sidebar. It's just your couple of pinned threads that you use.

Dylan Patel

Yeah, that sucks. Second of all, I'm in remote SSH sessions all the time, and the remote support in the thing sucks. So, yeah, I'm very against both of those.

Dylan Patel

Makes sense.

Dylan Patel

I think they are working on this, but, yeah, I'm just not reading the code, to be clear.

Dylan Patel

Thanks for clarifying. The last time you were on, we had that discussion, but you still have to run some things from the terminal and see some output in the logs at some point. This is just not part of the standard Codex experience, but I'm not the primary market. I understand that they are targeting vibe coders around the world.

We need to help people make more B2B SaaS. That's the core business model of this, and I'm not the target market. That's cool. I'll keep using the CLI. It's all good.

Actually, I wonder if, internally, most of their software engineers are using the app with the CLI. At SemiAnalysis, for sure, I think all of our real engineers are on the CLI. They're diehards, right? They'll never give it up.

Dylan Patel

That would be a good question. You're going to ask them that.

Dylan Patel

Yeah. Well, next time I see some OpenAI people, I'll ask.

Dylan Patel

Sounds good. Yeah.

Dylan Patel

Okay, let's change gears and talk a little bit about who's going to come in third. There have been some releases that we've covered a little bit. SpaceX AI is back. They've got Grok with Cursor bolted on, and things are going well. Meta has Muse Spark out in the world. They've also stacked up a whole bunch of compute, and they've teased launching a neocloud, which, of course, SpaceX was the first to do.

What do you guys want to talk about first: the models or the compute for these guys?

Dylan Patel

Let's do compute. I think it's more interesting. Max is going to walk you through the models.

Max

I mean, to be clear, the models aren't bad, but at this point I don't think it matters unless you can deliver a true frontier model. They're obviously not true frontier models, so who cares?

Dylan Patel

Do you think anyone will use their API product?

Max

I don't think anyone's going to use the API until it's a frontier model.

Dylan Patel

Okay.

Max

Yeah. So when they have an API, their token-as-a-service business will have fire margins, too. I mean, that's only if it's better than whatever the best Anthropic and OpenAI model is at that time, which I think is easier said than done.

I think Grok 4.5 and Muse Spark 1.1 are only useful as proving points along the scaling curve. It's like, okay, these guys, unlike Gemini, are still kind of in the game right now. They can train adequate models. It's not quite frontier yet, but there is a nonzero chance they can catch up. I think that's the way to view them.

I will personally be shifting my token usage to either Grok 4.5 or Muse Spark 1.1.

Dylan Patel

Okay. But you said the other name, which is shockingly, in my view, in fifth place right now: Gemini. Do you still have that view?

Max

Oh, 100%. I think they're clearly in fifth place. Unless Gemini 3.5 Pro is better than the industry chatter we're hearing, I think they're going to stay in fifth place, maybe forever.

Dylan Patel

Maybe forever. Okay. Talk to me about compute. Staying in fifth place forever would be related to the quality of the model, but the quality of the model is tied to how much compute these guys have.

Max

Yeah. I have to say that when I first saw the SpaceX-Anthropic deal, where Elon rented basically all of Colossus 1 to Anthropic, I think I incorrectly viewed that as Elon renting out his compute, when in reality I think the really important question is: do you have the ability to claw back all of your compute, assuming that you have proven you can truly reach the frontier by scaling up your compute?

I think that's the view Elon is taking with Cursor and xAI. He's essentially saying, "I'm going to leave you guys with enough compute to prove that you're capable of reaching the frontier. In the meantime, I'm going to monetize all the additional compute by renting it out on these extremely high-margin, 3x, maybe 4x market-rate rental deals."

I'm going to make sure there's a 90-day clawback clause included in every single one of them, such that if you do really well, I'm just going to claw back all the compute and give it to you so you can make digital costs. I think that is a valid strategy for someone who still truly thinks they can build RSI.

You can think of it as OpenAI and Anthropic monetizing their compute by selling inference, while Elon is doing these short-term compute deals instead. I think this is fine.

On the topic of Meta's compute, there are rumors that they're going to become a neocloud. As long as they do these SpaceX-style deals with a clawback, or they just give it all to MSL, or they do token as a service—any monetization method that allows them, in theory, to claw it all back if MSL proves they're doing well—I think that's an excellent, smart decision on Meta's part.

It probably allows them to be more aggressive with buying compute in the short term, while still giving them a chance to build RSI in the long term. I think the issue with Google, and why they're in fifth place aside from their model, like Gemini 35 Spy Flash being bad, is that they have signed all these long-term compute deals where they cannot claw back the TPUs that they're giving away. That's just committed for the long term, and I think it indicates a lack of conviction on their part that they can build RSI.

I think Noam Shazeer, John Jumper, and all these other high-profile guys at DeepMind leaving is partly because of this fact. They are high-pilled, their leadership is not high-pilled, and they think the only thing they can do in this situation is go work somewhere else.

Dylan Patel

Wow. Okay. Hot takes. The 2 things I want to dig in on there: the first is related to the clawback. This is a valid strategy, but you can basically only do the clawback once, and then nobody's ever going to buy from you or trust you again. Or do you think you can actually do it more than once?

Max

In an ideal world, you only need to do it once, because you only do it if you're confident that you're going to be able to breach the frontier by continuing to scale, right? But I also think it's possible you could do it more than once.

In principle, Anthropic is viewing this as a short-term deal. If Elon said, "Hey, actually, I want it all back 6 months from now," and then a year from now he's like, "Just kidding, we actually still can't train a frontier model," if Anthropic is still ripping at that point and is in desperate need of compute capacity, why would they say no?

Dylan Patel

Yeah, I guess financial motivations can make friends out of enemies. The second thing is just on the topic of the business. Maybe, Joey, I can bring you in here. What's a better business: selling tokens of your frontier model or close-to-frontier model, or selling compute access in the short term at a premium over the market?

Joey

Selling your own tokens is probably the best model. Even selling someone else's tokens, if you're AWS or Azure, is a pretty good business model as well, if you can participate in some of the upside revenue share.

But if you can run above-market rates, it's not a bad business. From the Meta compute perspective, this has been floated for a while. Meta compute has been floated since Q3 of last year.

It gives Zuck a backstop to say, "Okay, if MSL isn't successful, we have the means to essentially make decent ROI on our capex spend." It allows them now to spend a ton more on capex in 2027, potentially 2028, because there's a market. You've gotten these market signals that Anthropic can pay 3x market rates, and you can do a small deal here.

In the background, investors can feel comfortable that if MSL is deemed unsuccessful, Meta could have a pretty nice single-client or multi-client neocloud business just serving the labs.

Dylan Patel

Yeah, but I thought this chart that you guys had in the Meta compute article was striking, where you can see that the SpaceX-plus-Google deal is monetizing this compute at 4x the market rate. I think in the past everything you're saying makes perfect sense, but I would challenge a little bit that these near-term SpaceX GB300 rack deals aren't a really good business. They're clearly high-margin, right?

Joey

Yeah, that's extremely high-margin. When you think about the ARR per gigawatt, I think a lot of people would be happy to get $48 billion of actual token ARR from inference there. So I don't know. That deal doesn't seem very economical for Google.

Dylan Patel

Do you think Meta Compute is them telling a story to the market that they're going to be able to replicate this SpaceX AI thing, which is probably a one-time thing unless demand keeps growing so exponentially that there's such a compute crunch that everybody's fighting for access to GPUs?

While we have them, we can monetize them at 4x the market rate from neoclouds. I think our house view is that a neocloud is a pretty good business on its own at around $12 billion per gigawatt, approaching $50 billion per gigawatt.

Joey

Yeah, and the labs have gotten so good at the ARR-per-gigawatt number. It has gone up a ton.

Dylan Patel

Token throughput’s been pretty impressive—the models as well, and how they’ve monetized them. The new models price higher. If you look at Anthropic, revenue per megawatt was about $16 million per megawatt last year, or $16 billion per gigawatt, and now that’s more than doubled in 3 quarters. I think some people believe that can double again in 2–3 quarters.

Let’s look at that trend. I’ve got this chart on screen right now. Total token-as-a-service market—2 things jumped out at me. One, how fast this thing is growing. I thought the market was pretty big in Q4 of last year, and that looks puny compared to the forecast you guys have on this chart toward the end of this year—way more than doubling, as you said. But the other thing that jumps out to me is just how badly Microsoft Foundry is doing here. Can you—

Joey

Yeah, and that’ll be revised up with the recent OpenAI API success. I think, up until recently, Foundry was 90% OpenAI, and OpenAI was so consumer-focused. This is not the latest version of this chart. Foundry is doing a little better now. Maybe Vertex AI is doing a little worse, given Max’s thoughts on Gemini, than that chart.

Dylan Patel

I think we’re of the view—and when we talk to enterprises, token as a service is a very popular choice for consuming tokens. AWS and Azure have massive businesses with Fortune 500 and Global 2000 companies that spend 9 figures or more on cloud a year. If I can go to an existing vendor and have more model optionality, increase my ELA or my credits, and then burn that down, it’s pretty attractive for both AWS and Azure. If token as a service for Anthropic was 5% of indirect sources 6 months ago, it’s probably 20% of its business now.

Crystal, what’s your take on people using the big 3 hyperscalers’ token-as-a-service offerings versus the smaller startups, like Together, Baseten, and Fireworks—anything they can get on OpenRouter?

Crystal

I feel like, as we see more of the bigger enterprises, a lot of the financial services industry still hasn’t unlocked and used AI to the level that it probably can and should. That’s where Azure and Bedrock are going to benefit, because most of them probably already buy their services, right? Once they get internal approval or whatever to use these AI tools, they’ll just buy them through whatever existing channel they already have. That’s probably where the bulk of the market is.

Dylan Patel

Makes sense. Yeah. Max, how about you? What do you think of token as a service from startups like Together, Fireworks, and Baseten? How fast are they growing compared to the hyperscalers that are tied to the big labs, but maybe aren’t actually selling a lot of tokens to the startups of tomorrow?

Max

Together, Baseten, and Fireworks are all super impressive businesses. I do actually expect open-source token volumes to grow slower than frontier token volumes, but that’s only because I think frontier token volumes are going to absolutely explode. I think open-source token volumes are also going to explode, just to a slightly smaller degree.

It’s kind of funny. If you talk to the VCs investing in Together, Fireworks, and Baseten, and ask them for their explanation of why, oftentimes they lead with, “Inference is going to be the largest market ever.” It doesn’t even matter if these companies can only capture a super, super tiny slice of this extremely large market. That’s good enough for us.

I think that thesis is honestly more or less right. I don’t expect these guys to ever be doing more volume than OpenAI or Anthropic, or anywhere near that. In terms of global token volume, I think this might be the highest percentage they’ll ever be, rather than today. But they’ll still be good businesses in the future, I think.

Dylan Patel

How about the rumors that they can hitch their wagon to some of the bigger labs? SpaceX AI starts to win a little bit, and then Fireworks grows because they’re exposed to Cursor and actually have some exposure to that growth.

Max

I don’t think that’s going to last long term. I think the only reason Fireworks got that exposure is because originally Composer was post-trained on them, right? If you’re a frontier lab that has made a new model from scratch, there’s no reason to believe Fireworks’ engineers would be better at optimizing that model than your own engineers. I don’t think they can get any share of that margin in the future.

Dylan Patel

Okay. So why do the hyperscalers get a share of that margin—the big 3? I mean, they have the enterprise distribution that the neoclouds don’t. I’m sure Joey can speak more on this, but Anthropic isn’t giving up this margin to Bedrock so that Amazon’s engineers can make Claude run faster on Trainium and B300s or whatever. It’s so all the existing customers that rely on Bedrock can also use Claude models.

Max

Makes sense, yeah.

Dylan Patel

Yeah, I think the deal—I mean, we hear more. We don’t hear anything, but people tell us they think the deal could get reworked, because obviously it was very beneficial for AWS. I think we’ve written about AWS and its margins and how it monetizes versus just selling the bare compute for Anthropic to run inference on. To get that 30%, or 20% to 30%, revenue share just falls down to the bottom line. That’s pretty attractive.

AWS, Azure to a lesser extent, and GCP obviously have massive customer bases. People’s cloud estates are there, and enterprises are very comfortable with security and compliance. With everything now in the cloud, even more regulated industries are comfortable, so it’s a natural place for them to want to buy. But to see that much of the economics for Anthropic is obviously a lot. I think a lot of that was because, when the deal was struck and there was an incentive to make Trainium work, they were able to run with those terms and make it pretty favorable. So, yeah, we’ll see what happens if those revenue-share deals can continue.

Okay, let me make one more attempt at the bull case for these token-as-a-service companies that aren’t the hyperscalers, like AWS and Google. Obviously, a big portion of the benefit of getting token as a service from them is that their engineers are working on Trainium and TPU, and that might be different from the experience that the labs have with GPUs, where they have more experience running this themselves.

There’s a whole class of chip startups that are coming to market right now, and an obvious way in which they come to market is by partnering with these token-as-a-service companies. The biggest frontier labs aren’t going to spend a whole bunch of time optimizing for the 7th-best chip startup that’s coming to market right now. But if they do, and that chip startup strikes something really nice for a given model that makes it a lot more cost-effective or a lot higher-performance to run instead of Nvidia, then they’ve got a shot at doing something super unique.

Do you think that’s a potential future in terms of where these kernel engineers who have learned a lot end up going and spending their time over the next few months or years?

Max

This is a good point, and it’s reasonable in the short term. But if any of these new chip startups actually reach sufficient scale, I think the labs will just dedicate teams to making their model run really well on that chip. I think there’s no world in which you have a new accelerator that’s meaningfully better than Nvidia and is also being sold at large volume, where OpenAI is relying on Together to serve its model on that chip instead of just working directly with that company to develop the first-party capabilities to run its model on that chip.

Dylan Patel

Yeah. You think it’s going to go the way of Cerebras, where the chip companies, to be successful, are effectively going to have to become a neocloud themselves and have an existing relationship with the frontier labs?

Okay, makes sense. Maybe we can finish by talking a little bit about MSL at Meta, particularly, Max. I just loved the crash course on RL that this article turned into—not necessarily how it works from a technical perspective, but how it works from a business perspective: where people buy data, how they build these environments, and what the market looks like.

Can you give an overview? Previously, in a world where everything’s pre-training, whoever has a frontier-class team and the most compute can just train the biggest model and win. They follow the scaling laws and go there. But now there’s a scaling law related to RL. How is this playing out in your mind at a high level?

Max

Yeah. I think it’s really important for everyone to understand that reinforcement learning, or RL, is probably the most important scaling law for improving model capabilities today.

And there are a lot of people who believe that the only thing stopping models from being able to do literally anything a human can do on a computer is having sufficient RL environments. This is sort of like data that lets the model try to complete white-collar tasks itself until it can repeatedly try completing the task and fully learn how to solve it. There’s an entire new industry and supply chain of RL-environment startups whose entire job is to convert real-world, economically viable tasks into these RL environments, which they can then sell to the labs and allow them to use to improve their models.

You see most of the main players on this chart Jordan has pulled up here. I actually think some of these ARR numbers might be slightly understated. I think it’s pretty much consensus that the total data budgets at the frontier labs this year—primarily OpenAI, Anthropic, Google, Meta, and xAI, but also the long tail of companies like Amazon, Microsoft, and Thinking Machines—are going to sum to well over $10 billion. That’s roughly a 10x increase relative to last year. It’s very possible we see 10x again in 2027.

This is a hugely important market for improving AI capabilities. This is probably the only market in the world where customer demand isn’t even a question for all these startups. They will never have a contract turned down by the labs because it’s too expensive. It’s just a question of whether they can scale up creating high-quality data fast enough, and if the answer is yes, the labs will pay any price for it. One other side note: one of the reasons Anthropic’s models are the best at coding today—or at least they definitely were before 5.6 came out—is that they were by far the most aggressive about buying coding data from all these RL-environment startups. I think some of the other labs are starting to realize this and catch on, but it is definitely a super important industry that everyone should be aware of.

Dylan Patel

Can you dig into the process of creating some of these tasks? You ran through this in the article by describing how Meta has moved thousands of engineers into doing this work, and you also dispelled the notion that this work is meaningless, soul-crushing stuff. It’s actually pretty economically valuable and intellectually stimulating. I don’t want to steal your thunder, but you had a nice line on that one.

Max

I think I said it was both potentially more economically valuable and intellectually stimulating than your average big tech job. A lot of people hear the phrase “AI data” and still think, “We have some random people in the Philippines drawing bounding boxes or labeling text as NSFW.” The models have already fully solved that. Your data is only valuable if it’s doing something the models don’t already know how to do.

In the case of software engineering, in order to make a good software-engineering task today, it typically needs to be something that would take a really good human engineer maybe a full day of work to solve. It has to be something the model can’t already solve on its own today. To create this data, you essentially need a really good human engineer to sit down and think of an example problem they would actually want to do. You then need them to create a verifier—usually a set of integration tests, maybe along with a rubric—that can check whether the model successfully completed this day-long task.

Obviously, this is easier said than done. You also need the engineer to write a prompt for the model that asks it to do the task, but the prompt has to fulfill two competing factors. It needs to be 100% clear and unambiguous what you want the model to do, because you can’t incorrectly fail the model during training if it successfully did what your prompt asked for but the prompt wasn’t specific enough and the model didn’t know it had to do some extra thing. At the same time, the prompt needs to be realistic, natural-sounding, and representative of something a human would actually ask an AI to do in the real world. It’s very difficult to get that balance right.

We’ve heard that for the highest-quality coding tasks, the labs are willing to pay well over 5 figures for a single task. That’s already entering the realm of how much you would pay a decent engineer for a full week of work. I think that should dispel any myths about this being easy, mind-numbing work. To all the listeners out there: if you’re looking for a new job and you’re really good at creating RL tasks, you can make 7 figures or more annually at this point. Maybe consider that as a new job option.

Dylan Patel

Yeah, that may have some listeners excited. Can you actually go one click lower? Do you have personal experience with striking that balance between making something easy enough for the AI to do but not impossible for the AI to do? What sort of intuition could you give a listener about what that means?

Max

Honestly, it’s always changing. Back in the day, I actually did sell some environment data to the labs myself. I don’t do it anymore, but it was much easier to create data even 8 months ago than it is today. Today, it often looks like you need to identify a specific failure mode that you’re aware of in the model, and then you need to create RL tasks that specifically target that failure mode.

The only way to know whether it’s at the right difficulty for the model is to have the model try solving it 10 times and see how many times it’s successful. You just repeat that iteration. I don’t know if I can provide any blanket advice on how to find the right difficulty, other than this: if it’s your first time trying to do this, your first thought is almost certainly too easy. Try making it 10 times harder, and maybe you’ll be at the right level.

Dylan Patel

Interesting. Cool stuff.

Max

Yeah.

Dylan Patel

Okay, guys, we got a whirlwind tour. Is there anything you think I’ve missed as we’ve gone through token budgeting, Anthropic’s profit margins, Meta compute, MSL? What have we missed, Crystal?

Crystal

I don’t know. Nothing I can think of.

Dylan Patel

We’ve just been too busy watching the World Cup. Here we are on Wednesday, July 15, right when we’re recording this. We just got to watch England get knocked out as Argentina stormed back for a nice 2–1 victory. That was crazy. There are a bunch of British people in the office right now, and they’re all depressed downstairs. It was crazy.

Joey, how about you? What’s on your mind as we wrap up here?

Joey

Nothing. Token spend is up and to the right right now, so it’s good. I’m excited for hyperscaler earnings starting next week. We’ve got Google, Amazon, and Microsoft the week after. It should be pretty good on the top line.

Dylan Patel

Okay, let me go around the horn. We’ll close by getting a vibe check from everybody. Joey, what’s your vibe on the market?

Joey

On the market? I don’t know. As long as lab ARR is going up at a good pace and doesn’t decelerate, I think things are fine. Right now, it’s going up. OpenAI is catching up, so vibes should be good. That’s not what you’re seeing in semis over the last few days, but that’s just summer momentum. It doesn’t work, so it’ll come back. People will come back to the office from vacation, and semis will rip in a year. Things will be good.

Dylan Patel

Yeah, we’re going to see an acceleration after the summer pop. Joey’s going to shoot 84, see lab ARR go up, and be happy. Crystal, how’s your vibe?

Crystal

It’s going good, I have to say. Same thing Joey said. I need to leave San Francisco before the AI bubble pops, so I’m moving away from San Francisco very quickly before it’s too late.

Dylan Patel

Crystal, you think it’s a bubble? Next, we should publicly talk about our bet here.

Crystal

Yeah.

Dylan Patel

Right now. We have a bet. The loser—Max chose the over-under of Anthropic ARR at $400 billion by the end of 2027.

Joey

2027.

Dylan Patel

Yes, 2027. Does somebody know who took the over and who took the under?

Crystal

I took the under.

Dylan Patel

I took the over. Jeremy and Joey actually both took the under, but we haven’t decided what the actual bet is going to be yet.

Joey

Not this year. We’ll figure it out.

Did you guys see our pot at Rays last week? I pulled out a Canadian $20. Rake pulled out 200 Singapore dollars, as well as some rupees. We had Dylan pull out some euros. We had 5 different currencies going on at the table.

I will throw in Canadian currency and take the over with Max because we are—

Max

Let’s go.

Dylan Patel

We’re exponential extrapolators here.

Max

Yes.

Dylan Patel

Exactly.

Max

The loser has to write a newsletter post about why they were wrong.

Dylan Patel

Okay, that’s good.

Max

No, that’s what it was. I didn’t realize we agreed to this. I must have missed that in the Slack thread.

Dylan Patel

Are we doing Frontier Lab fantasy here? We need to pick a model for a given week and set up our team.

Max

We’ll bet on Jeremy’s spend.

Dylan Patel

Jeremy’s on Jeremy.

Max

We’ll bet on Jeremy’s weekly spend, and we’ll have an over-under.

Dylan Patel

Over-under.

Max

Get everyone involved. Isn’t Meta building an internal Polymarket or Kalshi? We’ll do that for SemiAnalysis. Someone can code something like Polymarket.

Dylan Patel

Dude, Meta needs to shut down that effort right now, dude. What are they doing?

Max

We’ll have a market for betting on Jeremy’s tokens.

Dylan Patel

There we go.

Max

I like it, guys.

Dylan Patel

Yeah.

Max

No way. Totally fair.

Dylan Patel

Hey, Jeremy, I’ve got 6,000 rupees riding on you. I need you to hammer the data center model dashboard this week, buddy. Okay, well, guys, I appreciate you taking the time. Hopefully the listeners enjoyed it. It devolved a little bit at the end. Let’s all get back to work. Keep tracking those tokens.

Max

Yeah.

Ep. 020 - Anthropic vs OpenAI Usage, Margins, Meta Compute, Future of MSL (Tokenomics) | BidClub