Dylan Patel
Hello, everyone. Welcome back to SemiAnalysis Weekly. We're here for Episode 20 with the Tokconomics team. Things are moving fast. We are recording this on Wednesday, July 15. I assume that by the time this episode comes out, there are going to be a lot more model releases and disclosures from these companies about their profitability, revenue, and how many active users they've got.
Regardless, we're going to record a point in time right now, talking about token budgeting, Meta's compute strategy, MSL futures, the release of Fable, Soul, and Anthropic's profit margins. I'm excited to dig in.
With me today, we've got Max. How are you doing, buddy?
Max
Doing great. Thanks for having me.
Dylan Patel
Welcome. We've got Crystal. How's it going, Crystal?
Crystal
Hey, Jordan.
Dylan Patel
And we've got Joey.
Joey
Hey, Jordan.
Dylan Patel
Cool. So we're going to start with token budgeting. Crystal, over to you. We've had lots of conversations with enterprises over time, just asking them what's going on. At SemiAnalysis, we're still token-maxing, but others are moving into a time of austerity. Can you take me through a little bit of what you found with this article?
Crystal
There are a lot of enterprises right now that are cracking down on their budgeting because they think their employees are spending too much. A lot of people are getting scared because they're thinking, "What if I don't have enough tokens to do my work?" I don't think that's necessarily true, though, and it applies to most people, right? A lot of people aren't even getting close to that limit. There are probably a handful of power users at each of these companies who will really be affected.
Other than that, I don't think it has had as big of an impact on most people's workflows as it has been portrayed on social media. We've seen a lot of different strategies that companies have been imposing, whether it's on a per-person basis or a monthly budget for the entire company. I think the per-company basis obviously makes a lot more sense than a per-person basis, given that some users will get more use out of it than others.
Dylan Patel
How do you see people actually enforcing this? Clearly, you can burn through a budget faster if you're going with the super-ultra-premium Max thinking model, but in some cases, when you're actually enforcing a budget, you take different approaches, right?
Crystal
Yeah. I've seen some people do this. I was talking to people at a smaller company, and they have a much smaller budget. The way they do it is by trying to optimize how they use the more expensive tokens from Anthropic or OpenAI. They use a cheaper model to process a lot of the information and condense it down into a smaller prompt, and then they use the more expensive models.
They also use something that their company isn't counting. There are a lot of discrepancies in terms of which models actually count. Some enterprises are counting their own models in the budget, while others aren't. If they aren't counting them, a lot of people will really push those models to the limit, use them a lot, and then use the ones that actually cost money.
You're seeing people—maybe, Max, I can bring you in on this one—using more tokens for coding rather than other use cases, right? That seems to be the market that's growing the fastest.
Max
Yeah, I think coding and software engineering in general is by far the most token-hungry use case. I actually think this is why a lot of the token-budgeting discourse is pretty uninformed. I think a lot of the budgets themselves are pretty uninformed.
I hear people saying that they want to make sure their sales guys aren't using Opus or Fable to write emails, so they should definitely be using Sonnet instead. It's like, dude, generating your email is basically free from a token perspective. It does not matter if you use Opus or Sonnet to write an email. I have no idea why this is the policy you're enforcing to try to reduce your token spend.
I've also heard lots of stories about certain companies where only the engineering team is allowed to use Claude Code or Codex, and everyone else only gets something like Codework or something like that. I think that's pretty shortsighted. You shouldn't have a caste system where your engineers are at the top and everyone else is forced to use something like Demoris [?]. You should really be giving everyone the opportunity to try these tools and figure out how they can become more productive with them.
Dylan Patel
Joey, are you using these models right now? How would you feel if I got access to them and you didn't?
Joey
I would probably be pretty pissed. I do have some access, though not as much as Max and other people at SemiAnalysis. That's probably just the nature of our work, or my work right now, and that's changing a little bit.
But, yeah, I'd be pretty mad because I still use Opus 4.8 a lot for AI research tasks that pull in a lot of different data sources and do more research-type work.
Dylan Patel
How do you think about the ROI question? A lot of the motivation behind token budgeting is clearly that somebody isn't seeing a return on this investment, or they're just not seeing it yet. But in a lot of the research you're doing, you're seeing that there are a lot of benefits to using these models, right?
Joey
Yeah, I think with the market being so coding-focused today, probably 70% or more of Anthropic's ARR on the API side is in coding. Second, from what we've heard, it's very power-user- and power-company-driven. Coding is pretty widespread, but there isn't one major big company for Anthropic. Meta is probably 3% to 5% of total Anthropic ARR, but it's heavily driven by power users.
If you're at a point where you're a power user at one of these companies, I think the Ramp data shows that the top 1%—the 99th percentile of companies—spend $100,000 per year on AI per employee. If you've gotten to the point where you're spending that much, you're obviously getting pretty good ROI.
Our conversations with a lot of enterprises have shown that these top engineers could go well above their limited budget, whatever their budget limit was. But if you were someone who was spending an insane amount of money, you were getting budgeted pretty quickly if you weren't seeing ROI. There are always projects that you'll try that don't get ROI. You'll scrap those and move on.
It was pretty obvious that, because the market is so coding-focused, people are getting ROI and continuing to spend. I think we continue to hear that net-new ARR for both Anthropic and now OpenAI—with Codex and 5.5 and 5.6—is broadening. Since late March, spend has been broadening, and there's clearly a lot of API token spend today.
There are a lot of organizations that aren't traditionally doing a ton of development-type work. They're only allowing their middle- and back-office workers to use Copilot 365 or SAT. But I think the market will continue to broaden out.
Dylan Patel
Do you guys have a take on the balance between API usage in applications and coding plans when it comes to people's individual usage? When you're talking to people, do you ask them whether they're using a Claude Code subscription versus paying per token through the API?
Max
Anybody who can use a subscription definitely should use a subscription because it's so subsidized. OpenAI and Anthropic are very aware of this, so when they have their enterprise plans—SemiAnalysis is on an Anthropic enterprise plan—they don't let you use a subscription. They force you to pay per token because they know that's where the margin is.
In general, anyone who can use a subscription does use a subscription. It's just the people who need more limits than that who pay for the API.
Joey
I've even talked to a few people at startups, and they're not on the team or enterprise plan at all. They just buy 5—or however many—Claude or OpenAI subscriptions they need per month, charge it to the company card, and keep doing that. If they hit the limit on however many accounts they already have, they just buy another one for the month because it's more cost-effective for them to operate that way than to use API billing.
If you're a startup that's young enough that you don't care about security or any of the other enterprise guarantees, just absolutely buy 10 subscriptions per person rather than paying for API credits.
Max
I think the recent study we did on this had a decently viral tweet. For Codex, I think $200 a month gets you around $12,000 worth of API credits, and for Anthropic, it was $200 a month for around $8,000. Obviously, that's a no-brainer for someone who is able to pick between those two.
Dylan Patel
Can you dig into that a little bit more, Max? When you're talking about the margin profile of these businesses and considering the subscription plan versus the API, what does that look like? I'm going to bring up the chart on screen that I think is the key one from that tweet thread.
Max
Yeah.
Max
Obviously, when you hear that a $200-a-month plan can generate $8,000 worth of API credits, the subscription is going to be much lower margin if the average user has anywhere near 100% utilization. The key question is: What is the actual average utilization across our entire user base?
This chart that Jordan showed us here is the break-even percentage for the various plans. We can see that for the Claude Max 20x plan, it’s just 10%. For the Pro and 5x plans, it’s a little better at 20%. OpenAI, because they are even more generous with their subscription limits, is slightly lower at 11.4% for the Plus and Pro 5x, and just 5.7% for the Pro 20x.
We’re pretty confident that the average utilization, especially for these 20x plans, is much higher than 10% and 5.7%, respectively, for Anthropic and OpenAI. That would mean that, forget having lower margins than API—which we think is likely 85%+ for Anthropic at least—it might just be negative margin in general for these 20x plans. Of course, these businesses want to move as much volume as possible to API.
I think maybe, Joey, since she wrote the newsletter on Anthropic’s business, this would be a good segue to talk about why we think Anthropic is potentially in a stronger position than OpenAI.
Joey
Yeah. And I think, too, once you move on to enterprise subscription plans at Anthropic, there’s no usage included in your subscription fee. It’s all done at API pricing.
At Anthropic, what we found was really interesting is that the business today is so much more—and this is changing—but more than 80% of ARR is on the API side. As Max showed, that’s really high margin, so they’ve gotten to a point where they’re profitable today. Some of that is on a non-GAAP basis, so excluding stock-based compensation, but operating profit was positive in Q2 and could be profitable to the tune of $1 billion-plus in Q3.
Some of that is a function of the fact that they’ve just grown so much. I don’t think they can hire enough or plow enough back into training as they’d like. But because of that margin, that gross profit advantage that they have today, they can plow as much back into training as they want and can extend their model advantage and model capability advantage.
We think that’s something they should obviously try to do. Just because it’s very profitable for them today doesn’t mean they should run at 30% EBITDA margins right away or as quickly as possible and show that this is a very profitable business model. They should continue to invest as much as they can into training.
Dylan Patel
Yeah. Can you talk about the mix between consumer and enterprise, assuming that consumer is those paid plans and enterprise is pay-per-use on the API?
Max
Yeah. When you look at these things, clearly the bulk of Twitter users are using the subscription plans. But then you get one really big customer who’s spending so much more money on API tokens that it kind of drowns out everything else, and that, I guess, makes it a good business.
Joey
Yeah. And because Anthropic has historically been so focused on enterprise, it was 90%+ enterprise. ChatGPT, by contrast, started heavy in consumer. Back in Q1, 60% of ARR was in consumer, and only 6% of the 950 million weekly active users end up actually paying for a plan. Most people pay for the $20-a-month plan.
It’s a much different business mix because OpenAI has to subsidize—or not subsidize, but there’s a cost to serving—the other 900 million weekly active users who are free. That lowers their gross margins on a blended basis by about 20 points.
That’s changing. We think that, with some of the rise in OpenAI’s API business since late March with Codex, and then 5.5 and 5.6, the 60/40 consumer-enterprise split has flipped now to 40/60 consumer versus enterprise here in Q2. From what we’re hearing in the channel, especially out of token-as-a-service businesses like Bedrock and Foundry, OpenAI has recognized this and is starting to shift. Even net new ARR is starting to look a little more even between Anthropic and OpenAI.
Dylan Patel
Yeah. And it’s kind of shocking. With the release of GPT-5.5 and GPT-5.6, people like Tibo on X who are tweeting about the daily active user counts are saying that they’re celebrating going to 7 million and then 8 million active users. But to be clear, they’ve got over 900 million weekly active users on the free tier of ChatGPT.
They’ve got a lot of room to grow just to get people using the number-one paid product. Maybe they’ve got a few extra users using other paid products, but I would assume the bulk of people using the paid products are using Codex, right? So that’s not a great conversion rate in terms of who’s actually using the full paid product from the base of 900 million.
Joey
Yeah, it’s funny. You actually see it, too, in the paying rates for even the consumer version of Claude. Claude is like 5.9% of free users end up paying, so it’s a much more focused user base.
If we think of those 900 million weekly active users, a lot of it is Google Search-type replacement, things like that. And then, obviously, because of that consumer focus, they’ve been working on advertising a lot. Consumers traditionally monetize poorly: people aren’t willing to pay, and the retention curves are typically pretty poor. How many people are paying 3 months later, 6 months later, or 12 months later?
Ads are a tough business model. It’s tough to scale, but maybe that’s the eventual path. They’ve put out some pretty crazy targets for 2030 advertising revenue. But yeah, consumer is a tough business. Enterprise, given the margins on tokens for API, is obviously the place to be.
I think Anthropic had a pretty good lead because they were so levered to coding early in the first 6 months of this year, and it seems like it’s shifting to more of an equal two-horse race here in mid-July.
Dylan Patel
Yeah. Crystal, what are you seeing in that two-horse race? Are people that you’re talking to using Anthropic’s Claude Code, Codex, or both?
Crystal
I feel like a few weeks ago, when I was talking to people, they were pretty heavy on using Anthropic’s Claude Code specifically. If I asked them why, they would say, “That’s just what I’ve been using, right? And I’m used to using it now.”
But now, as people in the consumer segment start using ChatGPT a lot more than Claude, that effect in the consumer sector starts trickling into enterprise a little bit. These people who were previously Claude Code power users and didn’t touch Codex at all are kind of like, “Okay, what if I consider tinkering with Codex?”
Even though consumer is a really small portion of the pie and the margins obviously aren’t as strong as enterprise, and historically speaking OpenAI hasn’t done as well in the consumer market as Anthropic has, they have that traction in the consumer space that’s eventually going to trickle in. We’re already starting to see that just this past week.
Dylan Patel
Max, what’s your take? OpenAI versus Anthropic. Where do things stand for you?
Max
I think I remember that the last time I was on this podcast, we were talking about how it was right when 5.5 came out, and we were talking about how things were looking really dire for OpenAI at the start of the year. I think ARR growth was likely flat in March and April, which meant the alarm bells were definitely ringing in Sam’s head when that happened.
But then we said that 5.5 would be an inflection point. This was actually a really good model, probably Opus 4.8-level. I think that prediction has played out even more quickly and more optimistically than we would have guessed.
Some people here are saying that 5.6 is as good as Opus 4.8, sort of just straight up, even though it’s only half the price. We also think it’s a much smaller model, which would be a testament to OpenAI’s training abilities.
As Joey mentioned earlier, net new ARR month over month is comparable between OpenAI and Anthropic now, which is pretty shocking. People were counting OpenAI out a little bit, and I think Jordan might have been one of those people. But they’re back, guys. They’re back.
Dylan Patel
I am nothing if not flexible. It’s always been about the model, guys. Everybody’s like, “When are you going to stop hating on OpenAI?” As soon as they ship a model that’s good—and it’s good.
For the things that I want to do, I’m basically default OpenAI at this point, despite the application not doing what I want it to do with multiple tabs and stuff like that.
Dylan Patel
Wait, are you using the Codex app or the CLI?
Dylan Patel
CLI.
Dylan Patel
What’s your form factor—CLI? I feel like they don’t want you to use the CLI, dude. They very actively want you to use the app. What’s pulling you back?
Dylan Patel
Tabs. I need multiple tabs going.
I think the way you’re supposed to do this—because I agree, the first time I tried using Codex, I was like, every single time I open a new chat, it’s just lost in the ether in the left-hand sidebar of all those tabs.
Dylan Patel
I think how you're supposed to do it is you have 1 pinned thread per project you're working on, and then you ask that thread to spin up subthreads anytime you have a discrete ask. You actually never even interact with the vast majority of threads in the left sidebar. It's just your couple of pinned threads that you use.
Dylan Patel
Yeah, that sucks. Second of all, I'm in remote SSH sessions all the time, and the remote support in the thing sucks. So, yeah, I'm very against both of those.
Dylan Patel
Makes sense.
Dylan Patel
I think they are working on this, but, yeah, I'm just not reading the code, to be clear.
Dylan Patel
Thanks for clarifying. The last time you were on, we had that discussion, but you still have to run some things from the terminal and see some output in the logs at some point. This is just not part of the standard Codex experience, but I'm not the primary market. I understand that they are targeting vibe coders around the world.
We need to help people make more B2B SaaS. That's the core business model of this, and I'm not the target market. That's cool. I'll keep using the CLI. It's all good.
Actually, I wonder if, internally, most of their software engineers are using the app with the CLI. At SemiAnalysis, for sure, I think all of our real engineers are on the CLI. They're diehards, right? They'll never give it up.
Dylan Patel
That would be a good question. You're going to ask them that.
Dylan Patel
Yeah. Well, next time I see some OpenAI people, I'll ask.
Dylan Patel
Sounds good. Yeah.
Dylan Patel
Okay, let's change gears and talk a little bit about who's going to come in third. There have been some releases that we've covered a little bit. SpaceX AI is back. They've got Grok with Cursor bolted on, and things are going well. Meta has Muse Spark out in the world. They've also stacked up a whole bunch of compute, and they've teased launching a neocloud, which, of course, SpaceX was the first to do.
What do you guys want to talk about first: the models or the compute for these guys?
Dylan Patel
Let's do compute. I think it's more interesting. Max is going to walk you through the models.
Max
I mean, to be clear, the models aren't bad, but at this point I don't think it matters unless you can deliver a true frontier model. They're obviously not true frontier models, so who cares?
Dylan Patel
Do you think anyone will use their API product?
Max
I don't think anyone's going to use the API until it's a frontier model.
Dylan Patel
Okay.
Max
Yeah. So when they have an API, their token-as-a-service business will have fire margins, too. I mean, that's only if it's better than whatever the best Anthropic and OpenAI model is at that time, which I think is easier said than done.
I think Grok 4.5 and Muse Spark 1.1 are only useful as proving points along the scaling curve. It's like, okay, these guys, unlike Gemini, are still kind of in the game right now. They can train adequate models. It's not quite frontier yet, but there is a nonzero chance they can catch up. I think that's the way to view them.
I will personally be shifting my token usage to either Grok 4.5 or Muse Spark 1.1.
Dylan Patel
Okay. But you said the other name, which is shockingly, in my view, in fifth place right now: Gemini. Do you still have that view?
Max
Oh, 100%. I think they're clearly in fifth place. Unless Gemini 3.5 Pro is better than the industry chatter we're hearing, I think they're going to stay in fifth place, maybe forever.
Dylan Patel
Maybe forever. Okay. Talk to me about compute. Staying in fifth place forever would be related to the quality of the model, but the quality of the model is tied to how much compute these guys have.
Max
Yeah. I have to say that when I first saw the SpaceX-Anthropic deal, where Elon rented basically all of Colossus 1 to Anthropic, I think I incorrectly viewed that as Elon renting out his compute, when in reality I think the really important question is: do you have the ability to claw back all of your compute, assuming that you have proven you can truly reach the frontier by scaling up your compute?
I think that's the view Elon is taking with Cursor and xAI. He's essentially saying, "I'm going to leave you guys with enough compute to prove that you're capable of reaching the frontier. In the meantime, I'm going to monetize all the additional compute by renting it out on these extremely high-margin, 3x, maybe 4x market-rate rental deals."
I'm going to make sure there's a 90-day clawback clause included in every single one of them, such that if you do really well, I'm just going to claw back all the compute and give it to you so you can make digital costs. I think that is a valid strategy for someone who still truly thinks they can build RSI.
You can think of it as OpenAI and Anthropic monetizing their compute by selling inference, while Elon is doing these short-term compute deals instead. I think this is fine.
On the topic of Meta's compute, there are rumors that they're going to become a neocloud. As long as they do these SpaceX-style deals with a clawback, or they just give it all to MSL, or they do token as a service—any monetization method that allows them, in theory, to claw it all back if MSL proves they're doing well—I think that's an excellent, smart decision on Meta's part.
It probably allows them to be more aggressive with buying compute in the short term, while still giving them a chance to build RSI in the long term. I think the issue with Google, and why they're in fifth place aside from their model, like Gemini 35 Spy Flash being bad, is that they have signed all these long-term compute deals where they cannot claw back the TPUs that they're giving away. That's just committed for the long term, and I think it indicates a lack of conviction on their part that they can build RSI.
I think Noam Shazeer, John Jumper, and all these other high-profile guys at DeepMind leaving is partly because of this fact. They are high-pilled, their leadership is not high-pilled, and they think the only thing they can do in this situation is go work somewhere else.
Dylan Patel
Wow. Okay. Hot takes. The 2 things I want to dig in on there: the first is related to the clawback. This is a valid strategy, but you can basically only do the clawback once, and then nobody's ever going to buy from you or trust you again. Or do you think you can actually do it more than once?
Max
In an ideal world, you only need to do it once, because you only do it if you're confident that you're going to be able to breach the frontier by continuing to scale, right? But I also think it's possible you could do it more than once.
In principle, Anthropic is viewing this as a short-term deal. If Elon said, "Hey, actually, I want it all back 6 months from now," and then a year from now he's like, "Just kidding, we actually still can't train a frontier model," if Anthropic is still ripping at that point and is in desperate need of compute capacity, why would they say no?
Dylan Patel
Yeah, I guess financial motivations can make friends out of enemies. The second thing is just on the topic of the business. Maybe, Joey, I can bring you in here. What's a better business: selling tokens of your frontier model or close-to-frontier model, or selling compute access in the short term at a premium over the market?
Joey
Selling your own tokens is probably the best model. Even selling someone else's tokens, if you're AWS or Azure, is a pretty good business model as well, if you can participate in some of the upside revenue share.
But if you can run above-market rates, it's not a bad business. From the Meta compute perspective, this has been floated for a while. Meta compute has been floated since Q3 of last year.
It gives Zuck a backstop to say, "Okay, if MSL isn't successful, we have the means to essentially make decent ROI on our capex spend." It allows them now to spend a ton more on capex in 2027, potentially 2028, because there's a market. You've gotten these market signals that Anthropic can pay 3x market rates, and you can do a small deal here.
In the background, investors can feel comfortable that if MSL is deemed unsuccessful, Meta could have a pretty nice single-client or multi-client neocloud business just serving the labs.
Dylan Patel
Yeah, but I thought this chart that you guys had in the Meta compute article was striking, where you can see that the SpaceX-plus-Google deal is monetizing this compute at 4x the market rate. I think in the past everything you're saying makes perfect sense, but I would challenge a little bit that these near-term SpaceX GB300 rack deals aren't a really good business. They're clearly high-margin, right?
Joey
Yeah, that's extremely high-margin. When you think about the ARR per gigawatt, I think a lot of people would be happy to get $48 billion of actual token ARR from inference there. So I don't know. That deal doesn't seem very economical for Google.
Dylan Patel
Do you think Meta Compute is them telling a story to the market that they're going to be able to replicate this SpaceX AI thing, which is probably a one-time thing unless demand keeps growing so exponentially that there's such a compute crunch that everybody's fighting for access to GPUs?
While we have them, we can monetize them at 4x the market rate from neoclouds. I think our house view is that a neocloud is a pretty good business on its own at around $12 billion per gigawatt, approaching $50 billion per gigawatt.
Joey
Yeah, and the labs have gotten so good at the ARR-per-gigawatt number. It has gone up a ton.
Dylan Patel
Token throughput’s been pretty impressive—the models as well, and how they’ve monetized them. The new models price higher. If you look at Anthropic, revenue per megawatt was about $16 million per megawatt last year, or $16 billion per gigawatt, and now that’s more than doubled in 3 quarters. I think some people believe that can double again in 2–3 quarters.
Let’s look at that trend. I’ve got this chart on screen right now. Total token-as-a-service market—2 things jumped out at me. One, how fast this thing is growing. I thought the market was pretty big in Q4 of last year, and that looks puny compared to the forecast you guys have on this chart toward the end of this year—way more than doubling, as you said. But the other thing that jumps out to me is just how badly Microsoft Foundry is doing here. Can you—
Joey
Yeah, and that’ll be revised up with the recent OpenAI API success. I think, up until recently, Foundry was 90% OpenAI, and OpenAI was so consumer-focused. This is not the latest version of this chart. Foundry is doing a little better now. Maybe Vertex AI is doing a little worse, given Max’s thoughts on Gemini, than that chart.
Dylan Patel
I think we’re of the view—and when we talk to enterprises, token as a service is a very popular choice for consuming tokens. AWS and Azure have massive businesses with Fortune 500 and Global 2000 companies that spend 9 figures or more on cloud a year. If I can go to an existing vendor and have more model optionality, increase my ELA or my credits, and then burn that down, it’s pretty attractive for both AWS and Azure. If token as a service for Anthropic was 5% of indirect sources 6 months ago, it’s probably 20% of its business now.
Crystal, what’s your take on people using the big 3 hyperscalers’ token-as-a-service offerings versus the smaller startups, like Together, Baseten, and Fireworks—anything they can get on OpenRouter?
Crystal
I feel like, as we see more of the bigger enterprises, a lot of the financial services industry still hasn’t unlocked and used AI to the level that it probably can and should. That’s where Azure and Bedrock are going to benefit, because most of them probably already buy their services, right? Once they get internal approval or whatever to use these AI tools, they’ll just buy them through whatever existing channel they already have. That’s probably where the bulk of the market is.
Dylan Patel
Makes sense. Yeah. Max, how about you? What do you think of token as a service from startups like Together, Fireworks, and Baseten? How fast are they growing compared to the hyperscalers that are tied to the big labs, but maybe aren’t actually selling a lot of tokens to the startups of tomorrow?
Max
Together, Baseten, and Fireworks are all super impressive businesses. I do actually expect open-source token volumes to grow slower than frontier token volumes, but that’s only because I think frontier token volumes are going to absolutely explode. I think open-source token volumes are also going to explode, just to a slightly smaller degree.
It’s kind of funny. If you talk to the VCs investing in Together, Fireworks, and Baseten, and ask them for their explanation of why, oftentimes they lead with, “Inference is going to be the largest market ever.” It doesn’t even matter if these companies can only capture a super, super tiny slice of this extremely large market. That’s good enough for us.
I think that thesis is honestly more or less right. I don’t expect these guys to ever be doing more volume than OpenAI or Anthropic, or anywhere near that. In terms of global token volume, I think this might be the highest percentage they’ll ever be, rather than today. But they’ll still be good businesses in the future, I think.
Dylan Patel
How about the rumors that they can hitch their wagon to some of the bigger labs? SpaceX AI starts to win a little bit, and then Fireworks grows because they’re exposed to Cursor and actually have some exposure to that growth.
Max
I don’t think that’s going to last long term. I think the only reason Fireworks got that exposure is because originally Composer was post-trained on them, right? If you’re a frontier lab that has made a new model from scratch, there’s no reason to believe Fireworks’ engineers would be better at optimizing that model than your own engineers. I don’t think they can get any share of that margin in the future.
Dylan Patel
Okay. So why do the hyperscalers get a share of that margin—the big 3? I mean, they have the enterprise distribution that the neoclouds don’t. I’m sure Joey can speak more on this, but Anthropic isn’t giving up this margin to Bedrock so that Amazon’s engineers can make Claude run faster on Trainium and B300s or whatever. It’s so all the existing customers that rely on Bedrock can also use Claude models.
Max
Makes sense, yeah.
Dylan Patel
Yeah, I think the deal—I mean, we hear more. We don’t hear anything, but people tell us they think the deal could get reworked, because obviously it was very beneficial for AWS. I think we’ve written about AWS and its margins and how it monetizes versus just selling the bare compute for Anthropic to run inference on. To get that 30%, or 20% to 30%, revenue share just falls down to the bottom line. That’s pretty attractive.
AWS, Azure to a lesser extent, and GCP obviously have massive customer bases. People’s cloud estates are there, and enterprises are very comfortable with security and compliance. With everything now in the cloud, even more regulated industries are comfortable, so it’s a natural place for them to want to buy. But to see that much of the economics for Anthropic is obviously a lot. I think a lot of that was because, when the deal was struck and there was an incentive to make Trainium work, they were able to run with those terms and make it pretty favorable. So, yeah, we’ll see what happens if those revenue-share deals can continue.
Okay, let me make one more attempt at the bull case for these token-as-a-service companies that aren’t the hyperscalers, like AWS and Google. Obviously, a big portion of the benefit of getting token as a service from them is that their engineers are working on Trainium and TPU, and that might be different from the experience that the labs have with GPUs, where they have more experience running this themselves.
There’s a whole class of chip startups that are coming to market right now, and an obvious way in which they come to market is by partnering with these token-as-a-service companies. The biggest frontier labs aren’t going to spend a whole bunch of time optimizing for the 7th-best chip startup that’s coming to market right now. But if they do, and that chip startup strikes something really nice for a given model that makes it a lot more cost-effective or a lot higher-performance to run instead of Nvidia, then they’ve got a shot at doing something super unique.
Do you think that’s a potential future in terms of where these kernel engineers who have learned a lot end up going and spending their time over the next few months or years?
Max
This is a good point, and it’s reasonable in the short term. But if any of these new chip startups actually reach sufficient scale, I think the labs will just dedicate teams to making their model run really well on that chip. I think there’s no world in which you have a new accelerator that’s meaningfully better than Nvidia and is also being sold at large volume, where OpenAI is relying on Together to serve its model on that chip instead of just working directly with that company to develop the first-party capabilities to run its model on that chip.
Dylan Patel
Yeah. You think it’s going to go the way of Cerebras, where the chip companies, to be successful, are effectively going to have to become a neocloud themselves and have an existing relationship with the frontier labs?
Okay, makes sense. Maybe we can finish by talking a little bit about MSL at Meta, particularly, Max. I just loved the crash course on RL that this article turned into—not necessarily how it works from a technical perspective, but how it works from a business perspective: where people buy data, how they build these environments, and what the market looks like.
Can you give an overview? Previously, in a world where everything’s pre-training, whoever has a frontier-class team and the most compute can just train the biggest model and win. They follow the scaling laws and go there. But now there’s a scaling law related to RL. How is this playing out in your mind at a high level?
Max
Yeah. I think it’s really important for everyone to understand that reinforcement learning, or RL, is probably the most important scaling law for improving model capabilities today.
And there are a lot of people who believe that the only thing stopping models from being able to do literally anything a human can do on a computer is having sufficient RL environments. This is sort of like data that lets the model try to complete white-collar tasks itself until it can repeatedly try completing the task and fully learn how to solve it. There’s an entire new industry and supply chain of RL-environment startups whose entire job is to convert real-world, economically viable tasks into these RL environments, which they can then sell to the labs and allow them to use to improve their models.
You see most of the main players on this chart Jordan has pulled up here. I actually think some of these ARR numbers might be slightly understated. I think it’s pretty much consensus that the total data budgets at the frontier labs this year—primarily OpenAI, Anthropic, Google, Meta, and xAI, but also the long tail of companies like Amazon, Microsoft, and Thinking Machines—are going to sum to well over $10 billion. That’s roughly a 10x increase relative to last year. It’s very possible we see 10x again in 2027.
This is a hugely important market for improving AI capabilities. This is probably the only market in the world where customer demand isn’t even a question for all these startups. They will never have a contract turned down by the labs because it’s too expensive. It’s just a question of whether they can scale up creating high-quality data fast enough, and if the answer is yes, the labs will pay any price for it. One other side note: one of the reasons Anthropic’s models are the best at coding today—or at least they definitely were before 5.6 came out—is that they were by far the most aggressive about buying coding data from all these RL-environment startups. I think some of the other labs are starting to realize this and catch on, but it is definitely a super important industry that everyone should be aware of.
Dylan Patel
Can you dig into the process of creating some of these tasks? You ran through this in the article by describing how Meta has moved thousands of engineers into doing this work, and you also dispelled the notion that this work is meaningless, soul-crushing stuff. It’s actually pretty economically valuable and intellectually stimulating. I don’t want to steal your thunder, but you had a nice line on that one.
Max
I think I said it was both potentially more economically valuable and intellectually stimulating than your average big tech job. A lot of people hear the phrase “AI data” and still think, “We have some random people in the Philippines drawing bounding boxes or labeling text as NSFW.” The models have already fully solved that. Your data is only valuable if it’s doing something the models don’t already know how to do.
In the case of software engineering, in order to make a good software-engineering task today, it typically needs to be something that would take a really good human engineer maybe a full day of work to solve. It has to be something the model can’t already solve on its own today. To create this data, you essentially need a really good human engineer to sit down and think of an example problem they would actually want to do. You then need them to create a verifier—usually a set of integration tests, maybe along with a rubric—that can check whether the model successfully completed this day-long task.
Obviously, this is easier said than done. You also need the engineer to write a prompt for the model that asks it to do the task, but the prompt has to fulfill two competing factors. It needs to be 100% clear and unambiguous what you want the model to do, because you can’t incorrectly fail the model during training if it successfully did what your prompt asked for but the prompt wasn’t specific enough and the model didn’t know it had to do some extra thing. At the same time, the prompt needs to be realistic, natural-sounding, and representative of something a human would actually ask an AI to do in the real world. It’s very difficult to get that balance right.
We’ve heard that for the highest-quality coding tasks, the labs are willing to pay well over 5 figures for a single task. That’s already entering the realm of how much you would pay a decent engineer for a full week of work. I think that should dispel any myths about this being easy, mind-numbing work. To all the listeners out there: if you’re looking for a new job and you’re really good at creating RL tasks, you can make 7 figures or more annually at this point. Maybe consider that as a new job option.
Dylan Patel
Yeah, that may have some listeners excited. Can you actually go one click lower? Do you have personal experience with striking that balance between making something easy enough for the AI to do but not impossible for the AI to do? What sort of intuition could you give a listener about what that means?
Max
Honestly, it’s always changing. Back in the day, I actually did sell some environment data to the labs myself. I don’t do it anymore, but it was much easier to create data even 8 months ago than it is today. Today, it often looks like you need to identify a specific failure mode that you’re aware of in the model, and then you need to create RL tasks that specifically target that failure mode.
The only way to know whether it’s at the right difficulty for the model is to have the model try solving it 10 times and see how many times it’s successful. You just repeat that iteration. I don’t know if I can provide any blanket advice on how to find the right difficulty, other than this: if it’s your first time trying to do this, your first thought is almost certainly too easy. Try making it 10 times harder, and maybe you’ll be at the right level.
Dylan Patel
Interesting. Cool stuff.
Max
Yeah.
Dylan Patel
Okay, guys, we got a whirlwind tour. Is there anything you think I’ve missed as we’ve gone through token budgeting, Anthropic’s profit margins, Meta compute, MSL? What have we missed, Crystal?
Crystal
I don’t know. Nothing I can think of.
Dylan Patel
We’ve just been too busy watching the World Cup. Here we are on Wednesday, July 15, right when we’re recording this. We just got to watch England get knocked out as Argentina stormed back for a nice 2–1 victory. That was crazy. There are a bunch of British people in the office right now, and they’re all depressed downstairs. It was crazy.
Joey, how about you? What’s on your mind as we wrap up here?
Joey
Nothing. Token spend is up and to the right right now, so it’s good. I’m excited for hyperscaler earnings starting next week. We’ve got Google, Amazon, and Microsoft the week after. It should be pretty good on the top line.
Dylan Patel
Okay, let me go around the horn. We’ll close by getting a vibe check from everybody. Joey, what’s your vibe on the market?
Joey
On the market? I don’t know. As long as lab ARR is going up at a good pace and doesn’t decelerate, I think things are fine. Right now, it’s going up. OpenAI is catching up, so vibes should be good. That’s not what you’re seeing in semis over the last few days, but that’s just summer momentum. It doesn’t work, so it’ll come back. People will come back to the office from vacation, and semis will rip in a year. Things will be good.
Dylan Patel
Yeah, we’re going to see an acceleration after the summer pop. Joey’s going to shoot 84, see lab ARR go up, and be happy. Crystal, how’s your vibe?
Crystal
It’s going good, I have to say. Same thing Joey said. I need to leave San Francisco before the AI bubble pops, so I’m moving away from San Francisco very quickly before it’s too late.
Dylan Patel
Crystal, you think it’s a bubble? Next, we should publicly talk about our bet here.
Crystal
Yeah.
Dylan Patel
Right now. We have a bet. The loser—Max chose the over-under of Anthropic ARR at $400 billion by the end of 2027.
Joey
2027.
Dylan Patel
Yes, 2027. Does somebody know who took the over and who took the under?
Crystal
I took the under.
Dylan Patel
I took the over. Jeremy and Joey actually both took the under, but we haven’t decided what the actual bet is going to be yet.
Joey
Not this year. We’ll figure it out.
Did you guys see our pot at Rays last week? I pulled out a Canadian $20. Rake pulled out 200 Singapore dollars, as well as some rupees. We had Dylan pull out some euros. We had 5 different currencies going on at the table.
I will throw in Canadian currency and take the over with Max because we are—
Max
Let’s go.
Dylan Patel
We’re exponential extrapolators here.
Max
Yes.
Dylan Patel
Exactly.
Max
The loser has to write a newsletter post about why they were wrong.
Dylan Patel
Okay, that’s good.
Max
No, that’s what it was. I didn’t realize we agreed to this. I must have missed that in the Slack thread.
Dylan Patel
Are we doing Frontier Lab fantasy here? We need to pick a model for a given week and set up our team.
Max
We’ll bet on Jeremy’s spend.
Dylan Patel
Jeremy’s on Jeremy.
Max
We’ll bet on Jeremy’s weekly spend, and we’ll have an over-under.
Dylan Patel
Over-under.
Max
Get everyone involved. Isn’t Meta building an internal Polymarket or Kalshi? We’ll do that for SemiAnalysis. Someone can code something like Polymarket.
Dylan Patel
Dude, Meta needs to shut down that effort right now, dude. What are they doing?
Max
We’ll have a market for betting on Jeremy’s tokens.
Dylan Patel
There we go.
Max
I like it, guys.
Dylan Patel
Yeah.
Max
No way. Totally fair.
Dylan Patel
Hey, Jeremy, I’ve got 6,000 rupees riding on you. I need you to hammer the data center model dashboard this week, buddy. Okay, well, guys, I appreciate you taking the time. Hopefully the listeners enjoyed it. It devolved a little bit at the end. Let’s all get back to work. Keep tracking those tokens.
Max
Yeah.