第020期 - Anthropic 对 OpenAI:使用量、利润率、Meta 算力、MSL 的未来(Tokenomics)
企业端的 token 紧缩针对的是错误的工作负载:编程吃掉预算,而高阶模型写邮件几乎不花钱。 Crystal认为,多数员工根本触及不了用量上限,因此公司共享预算池比人均设限更合理。Max称只允许工程师使用相关工具是“种姓制度”,Joey则表示,头部1%的公司已经在 AI 上按每名员工每年约10万美元的水平投入,只要 ROI 经得起审视,就会继续花钱。
每月200美元的编程订阅被大幅补贴,在重度用户的使用强度下可能亏损。 团队估算,Codex可提供约1.2万美元的 API 等值用量,Anthropic约为8000美元;Claude Max 20x的盈亏平衡使用率为10%,OpenAI Pro 20x为5.7%。按可能的使用水平,Max认为这些套餐“可能就是负利润率”(might just be negative margin)。
Anthropic 以企业/API为主的收入结构已经形成 OpenAI 免费消费者基盘无法匹敌的经营杠杆。 Anthropic超过80%的 ARR 来自 API,历史上超过90%来自企业端;Joey称,公司部分盈利按剔除股权薪酬的非 GAAP 口径计算,但Q2经营利润已经转正,Q3可能达到10亿美元以上。OpenAI约9.5亿周活用户中,只有约6%转化为付费用户,服务免费用户使整体毛利率下降约20个百分点。
尽管如此,OpenAI仍重新回到了真正的双雄竞争格局,因为决定胜负的依然是模型质量,而不是先发地位。 在3月和4月 ARR 增长可能持平后,5.5被视为预期中的拐点;如今有些人认为5.6以一半价格达到了“与 Opus 4.8 一样好”的水平,而且很可能是更小的模型。据称,OpenAI月度净新增 ARR 已与 Anthropic 相当,促使 Dylan 改口:“我向来都很灵活”(I am nothing if not flexible)。
对 xAI 和 Meta 来说,算力调度的可选性可能与算力所有权同等重要。 Max修正后的观点是,团队可以把闲置算力按市场价3x-4x的价格出租,同时保留足够算力证明自己能够触达前沿水平,并以90天回收条款作为保障。Meta算力也可能为2027-28年更重的资本开支提供支撑;如果 MSL表现不佳,算力出租还可以成为一项有利可图的后备方案。相比之下,Max认为 Google 不可召回的长期 TPU 承诺,说明其对构建 RSI 信心不足。
云巨头凭借既有企业关系赢下 token-as-a-service 的分发,而独立推理服务商无需赢得整个市场也能做大。 Anthropic通过间接渠道获得的 token 量,6个月内可能已从业务的约5%升至20%;AWS很可能仅因客户本来就通过 Bedrock 采购,就能拿走20%-30%的收入分成。即使开源模型的调用量增速低于前沿模型,Fireworks和Baseten仍可快速增长,因为推理可能成为“史上最大的市场”。
RL 环境正同时成为下一阶段的能力瓶颈和一个异常高利润的工程市场。 前沿实验室的数据预算今年可能超过100亿美元,约为去年的10倍,2027年可能再增10倍;顶级编程任务的单价远超5位数美元,而特别擅长设计 RL 任务的人每年可以赚到7位数美元以上。本期还用2027年底 Anthropic ARR 4000亿美元的高低押注,让其指数级增长论点变得可证伪:Max和Dylan押高,Crystal、Joey和Jeremy押低,输家必须公开解释自己为何判断错误。
1. Token 上限瞄准了错误的工作负载
Crystal的一线观察是,社交媒体夸大了 token 预算压力:多数公司最多只有少数重度用户会触及上限。由于使用量分布极不均衡,全公司共享一个月度 token 池,比所有人拿到相同的人均配额更合理。
小型组织已经在绕开限制:先用便宜模型处理并压缩材料,再调用 Anthropic 或 OpenAI 的高价模型。在内部模型不计入预算的情况下,员工会“把这些额度压榨到极限”,把计量模型留给最后一步工作。
Max反对的是许多政策在优化无关紧要的细节。从 token 消耗看,销售人员用 Opus 还是 Sonnet 写邮件“基本无关紧要”;编程“远远是最吃 token 的用例”。
将 Claude Code 或 Codex 限定给工程师使用,在 Max看来就是一种“种姓制度”,让其他职能失去了发现高价值工作流的机会。Joey的佐证是,他估计 Anthropic API 端 ARR 的70%以上来自编程,而 Meta 仅贡献 Anthropic 总 ARR 的约3%-5%——使用量主要由重度用户驱动。
2. 订阅套利暴露真实利润池
Max的实用规则很直接:“凡是能用订阅的,肯定都应该用”,因为补贴力度太大。有些初创公司直接把5个、或者实际需要数量的个人账户挂到公司卡上,触及上限就再加一个;Max称,每人购买10个订阅可能比购买 API credits 更划算。
按每月200美元计算,团队估算 Codex提供约1.2万美元的 API 等值额度,Anthropic约为8000美元。相应的盈亏平衡使用率,Claude Pro和5x套餐约为20%,Claude Max 20x为10%,OpenAI Plus和Pro 5x为11.4%,Pro 20x为5.7%。
Max认为,20x套餐的平均使用率高于上述阈值,因此这些产品可能不是低毛利,而是直接亏损。Anthropic API 毛利率估计超过85%;Anthropic企业订阅不包含用量,客户先付订阅费,之后按 API 价格消耗。
3. Anthropic 的 API 结构将模型需求转化为利润
Anthropic超过80%的 ARR 来自 API,历史上超过90%来自企业端。Joey称,公司部分盈利按剔除股权薪酬的非 GAAP 口径计算,但Q2经营利润已经转正,Q3可能达到10亿美元以上。
这部分盈利也反映出,公司的招聘和训练再投资速度跟不上收入增长。Joey的建议不是急着把 EBITDA 利润率推向30%,而是把毛利优势投入到模型能力扩展上:“尽可能继续把能投的钱都投进训练。”
OpenAI在Q1则是相反的收入结构:约60%的 ARR 来自消费者端,周活用户约9.5亿,付费率只有6%,且大多数订阅用户购买20美元套餐。免费用户使整体毛利率下降约20个百分点;Dylan指出,日活用户数从700万升至800万,但免费周活用户超过9亿。
团队认为,Codex、5.5和5.6可能已经在Q2把 OpenAI 的收入结构扭转为约40%消费者、60%企业。收尾的赌注体现了其中的 stakes:2027年底 Anthropic ARR 高于还是低于4000亿美元;Max和Dylan押高,Crystal、Joey和Jeremy押低。
4. 更强的模型让 OpenAI 重回竞争
Crystal表示,Claude Code最初主要靠使用习惯留住用户——“我一直用的就是它”——但 ChatGPT 在消费者端的势头现在正逐步渗透到工作场景。此前无视 Codex 的 Claude Code 重度用户,开始尝试使用 Codex。
Max认为,OpenAI的 ARR 增长在3月和4月可能持平,触发了“警报”,随后5.5以大致达到 Opus 4.8 的质量带来了预期中的拐点。如今有些人认为5.6已经达到 Opus 4.8 的水平,价格却只有一半;团队还认为,5.6很可能小得多。
Dylan改变看法的原因在模型本身:他现在工作时默认使用 OpenAI,但不认可 Codex app 的线程模型,认为其远程 SSH 体验糟糕,也缺少 tab 工作流。他仍然是 CLI 用户,因为这款 app 显然是面向全球“vibe coder”群体、服务于更多 B2B SaaS 开发。
5. 前沿算力在可召回时最有价值
Max认为 Grok 4.5 和 Muse Spark 1.1 只是足以证明扩展有效的产品,而不是前沿模型:“除非你能交付真正的前沿模型……否则谁在乎。”他计划将 token 用量转向其中任一产品,但这两者的存在至少证明了追赶并非零概率。
Max更苛刻的排名是,Gemini“明确排在第5”,除非 Gemini 3.5 Pro 超过当前行业传闻的水平,否则可能永远如此。关键区别不只是手里有多少算力,而是管理层是否保留将这些算力转去训练前沿模型的选择权。
Max修正了此前对 Elon 将 Colossus 1 几乎全部租给 Anthropic 的解读:内部团队保留足够算力,以证明扩展路径有效;其余算力则按市场价约3x-4x出租,同时保留90天回收权,一旦模型达到前沿水平就可以收回。
Dylan的反驳值得保留:服务商可能只能收回一次算力,随后就会失去客户信任。Max同意收回一次最理想,但认为如果 Elon 后来改变主意,而 Anthropic仍然迫切需要算力,Anthropic可能还是会回来:“他们为什么要拒绝?”
6. Meta 算力是资本开支的后盾,分发能力驱动 token 转售
Meta可以通过可召回的算力租赁、token-as-a-service,或将算力分配给 MSL,采用同样的结构,同时保留将全部算力归还 MSL 的选择权。Joey认为,这能为2027年、甚至在 MSL被认定为未达预期时的2028年更激进资本开支提供后盾。
自己销售前沿模型 token 仍然是最好的生意,但高价算力租赁也绝非退而求其次。讨论中的 SpaceX+Google 交易定价接近市场价4倍;相比之下,neocloud 的经济性约为每吉瓦120亿美元,且可能逼近500亿美元。
Anthropic的收入密度解释了买方为何能接受这种价格:去年每兆瓦约1600万美元,3个季度内翻了一倍以上,部分人预计未来2-3个季度还会再翻一倍。“Token 吞吐量一直相当可观”,而且新模型能以更高价格变现。
Dylan一度把旧图表理解为 Microsoft Foundry 表现不佳的证据;Joey纠正称,该业务直到最近约90%仍来自 OpenAI,在近期 API 增长后应向上修订。Bedrock和 Azure 之所以仍具吸引力,是因为《财富》500强和 Global 2000 客户本来就拥有大型云关系、云额度或 ELA,以及安全和合规能力。
7. RL 环境正成为下一项能力扩展约束
Max认为,RL“可能是当前能力提升最重要的 Scaling Law”。不少人相信,只要拥有足够的环境,就能通过反复尝试教会模型几乎所有基于计算机的人类任务;前沿实验室合计数据预算今年可能超过100亿美元,约为去年的10倍,2027年还可能再翻10倍。
这个市场几乎不存在需求约束:实验室不会仅仅因为价格高,就拒绝优质数据。Max认为,Anthropic在5.6之前的编程领先,部分源于其购买编程环境的力度远超竞争对手,而竞争对手现在正试图追赶。
如今创建环境,和画边界框或添加内容标签几乎不是一回事。一个高质量软件任务可能需要一名非常优秀的软件工程师投入整整一天,配合集成测试验证器和评分标准;同时,提示词既要足够明确,保证训练公平,又要足够自然,像真实用户提出的请求。
实验室为单个顶级编程任务支付的价格远超5位数美元,而特别擅长设计 RL 任务的人每年可以赚到7位数美元以上。Max的亲身体验是,创建数据现在比8个月前难得多:先找出精确的失败模式,让模型尝试10次,再不断迭代。“你的第一个想法几乎肯定太简单——试着把难度提高10倍。”
Dylan Patel
Hello, everyone. Welcome back to SemiAnalysis Weekly. We're here for Episode 20 with the Tokconomics team. Things are moving fast. We are recording this on Wednesday, July 15. I assume that by the time this episode comes out, there are going to be a lot more model releases and disclosures from these companies about their profitability, revenue, and how many active users they've got.
Regardless, we're going to record a point in time right now, talking about token budgeting, Meta's compute strategy, MSL futures, the release of Fable, Soul, and Anthropic's profit margins. I'm excited to dig in.
With me today, we've got Max. How are you doing, buddy?
Max
Doing great. Thanks for having me.
Dylan Patel
Welcome. We've got Crystal. How's it going, Crystal?
Hey, Jordan.
Dylan Patel
And we've got Joey.
Joey
Hey, Jordan.
Dylan Patel
Cool. So we're going to start with token budgeting. Crystal, over to you. We've had lots of conversations with enterprises over time, just asking them what's going on. At SemiAnalysis, we're still token-maxing, but others are moving into a time of austerity. Can you take me through a little bit of what you found with this article?
There are a lot of enterprises right now that are cracking down on their budgeting because they think their employees are spending too much. A lot of people are getting scared because they're thinking, "What if I don't have enough tokens to do my work?" I don't think that's necessarily true, though, and it applies to most people, right? A lot of people aren't even getting close to that limit. There are probably a handful of power users at each of these companies who will really be affected.
Other than that, I don't think it has had as big of an impact on most people's workflows as it has been portrayed on social media. We've seen a lot of different strategies that companies have been imposing, whether it's on a per-person basis or a monthly budget for the entire company. I think the per-company basis obviously makes a lot more sense than a per-person basis, given that some users will get more use out of it than others.
Dylan Patel
How do you see people actually enforcing this? Clearly, you can burn through a budget faster if you're going with the super-ultra-premium Max thinking model, but in some cases, when you're actually enforcing a budget, you take different approaches, right?
Yeah. I've seen some people do this. I was talking to people at a smaller company, and they have a much smaller budget. The way they do it is by trying to optimize how they use the more expensive tokens from Anthropic or OpenAI. They use a cheaper model to process a lot of the information and condense it down into a smaller prompt, and then they use the more expensive models.
They also use something that their company isn't counting. There are a lot of discrepancies in terms of which models actually count. Some enterprises are counting their own models in the budget, while others aren't. If they aren't counting them, a lot of people will really push those models to the limit, use them a lot, and then use the ones that actually cost money.
You're seeing people—maybe, Max, I can bring you in on this one—using more tokens for coding rather than other use cases, right? That seems to be the market that's growing the fastest.
Max
Yeah, I think coding and software engineering in general is by far the most token-hungry use case. I actually think this is why a lot of the token-budgeting discourse is pretty uninformed. I think a lot of the budgets themselves are pretty uninformed.
I hear people saying that they want to make sure their sales guys aren't using Opus or Fable to write emails, so they should definitely be using Sonnet instead. It's like, dude, generating your email is basically free from a token perspective. It does not matter if you use Opus or Sonnet to write an email. I have no idea why this is the policy you're enforcing to try to reduce your token spend.
I've also heard lots of stories about certain companies where only the engineering team is allowed to use Claude Code or Codex, and everyone else only gets something like Codework or something like that. I think that's pretty shortsighted. You shouldn't have a caste system where your engineers are at the top and everyone else is forced to use something like Demoris [?]. You should really be giving everyone the opportunity to try these tools and figure out how they can become more productive with them.
Dylan Patel
Joey, are you using these models right now? How would you feel if I got access to them and you didn't?
Joey
I would probably be pretty pissed. I do have some access, though not as much as Max and other people at SemiAnalysis. That's probably just the nature of our work, or my work right now, and that's changing a little bit.
But, yeah, I'd be pretty mad because I still use Opus 4.8 a lot for AI research tasks that pull in a lot of different data sources and do more research-type work.
Dylan Patel
How do you think about the ROI question? A lot of the motivation behind token budgeting is clearly that somebody isn't seeing a return on this investment, or they're just not seeing it yet. But in a lot of the research you're doing, you're seeing that there are a lot of benefits to using these models, right?
Joey
Yeah, I think with the market being so coding-focused today, probably 70% or more of Anthropic's ARR on the API side is in coding. Second, from what we've heard, it's very power-user- and power-company-driven. Coding is pretty widespread, but there isn't one major big company for Anthropic. Meta is probably 3% to 5% of total Anthropic ARR, but it's heavily driven by power users.
If you're at a point where you're a power user at one of these companies, I think the Ramp data shows that the top 1%—the 99th percentile of companies—spend $100,000 per year on AI per employee. If you've gotten to the point where you're spending that much, you're obviously getting pretty good ROI.
Our conversations with a lot of enterprises have shown that these top engineers could go well above their limited budget, whatever their budget limit was. But if you were someone who was spending an insane amount of money, you were getting budgeted pretty quickly if you weren't seeing ROI. There are always projects that you'll try that don't get ROI. You'll scrap those and move on.
It was pretty obvious that, because the market is so coding-focused, people are getting ROI and continuing to spend. I think we continue to hear that net-new ARR for both Anthropic and now OpenAI—with Codex and 5.5 and 5.6—is broadening. Since late March, spend has been broadening, and there's clearly a lot of API token spend today.
There are a lot of organizations that aren't traditionally doing a ton of development-type work. They're only allowing their middle- and back-office workers to use Copilot 365 or SAT. But I think the market will continue to broaden out.
Dylan Patel
Do you guys have a take on the balance between API usage in applications and coding plans when it comes to people's individual usage? When you're talking to people, do you ask them whether they're using a Claude Code subscription versus paying per token through the API?
Max
Anybody who can use a subscription definitely should use a subscription because it's so subsidized. OpenAI and Anthropic are very aware of this, so when they have their enterprise plans—SemiAnalysis is on an Anthropic enterprise plan—they don't let you use a subscription. They force you to pay per token because they know that's where the margin is.
In general, anyone who can use a subscription does use a subscription. It's just the people who need more limits than that who pay for the API.
Joey
I've even talked to a few people at startups, and they're not on the team or enterprise plan at all. They just buy 5—or however many—Claude or OpenAI subscriptions they need per month, charge it to the company card, and keep doing that. If they hit the limit on however many accounts they already have, they just buy another one for the month because it's more cost-effective for them to operate that way than to use API billing.
If you're a startup that's young enough that you don't care about security or any of the other enterprise guarantees, just absolutely buy 10 subscriptions per person rather than paying for API credits.
Max
I think the recent study we did on this had a decently viral tweet. For Codex, I think $200 a month gets you around $12,000 worth of API credits, and for Anthropic, it was $200 a month for around $8,000. Obviously, that's a no-brainer for someone who is able to pick between those two.
Dylan Patel
Can you dig into that a little bit more, Max? When you're talking about the margin profile of these businesses and considering the subscription plan versus the API, what does that look like? I'm going to bring up the chart on screen that I think is the key one from that tweet thread.
Max
Yeah.
Max
Obviously, when you hear that a $200-a-month plan can generate $8,000 worth of API credits, the subscription is going to be much lower margin if the average user has anywhere near 100% utilization. The key question is: What is the actual average utilization across our entire user base?
This chart that Jordan showed us here is the break-even percentage for the various plans. We can see that for the Claude Max 20x plan, it’s just 10%. For the Pro and 5x plans, it’s a little better at 20%. OpenAI, because they are even more generous with their subscription limits, is slightly lower at 11.4% for the Plus and Pro 5x, and just 5.7% for the Pro 20x.
We’re pretty confident that the average utilization, especially for these 20x plans, is much higher than 10% and 5.7%, respectively, for Anthropic and OpenAI. That would mean that, forget having lower margins than API—which we think is likely 85%+ for Anthropic at least—it might just be negative margin in general for these 20x plans. Of course, these businesses want to move as much volume as possible to API.
I think maybe, Joey, since she wrote the newsletter on Anthropic’s business, this would be a good segue to talk about why we think Anthropic is potentially in a stronger position than OpenAI.
Joey
Yeah. And I think, too, once you move on to enterprise subscription plans at Anthropic, there’s no usage included in your subscription fee. It’s all done at API pricing.
At Anthropic, what we found was really interesting is that the business today is so much more—and this is changing—but more than 80% of ARR is on the API side. As Max showed, that’s really high margin, so they’ve gotten to a point where they’re profitable today. Some of that is on a non-GAAP basis, so excluding stock-based compensation, but operating profit was positive in Q2 and could be profitable to the tune of $1 billion-plus in Q3.
Some of that is a function of the fact that they’ve just grown so much. I don’t think they can hire enough or plow enough back into training as they’d like. But because of that margin, that gross profit advantage that they have today, they can plow as much back into training as they want and can extend their model advantage and model capability advantage.
We think that’s something they should obviously try to do. Just because it’s very profitable for them today doesn’t mean they should run at 30% EBITDA margins right away or as quickly as possible and show that this is a very profitable business model. They should continue to invest as much as they can into training.
Dylan Patel
Yeah. Can you talk about the mix between consumer and enterprise, assuming that consumer is those paid plans and enterprise is pay-per-use on the API?
Max
Yeah. When you look at these things, clearly the bulk of Twitter users are using the subscription plans. But then you get one really big customer who’s spending so much more money on API tokens that it kind of drowns out everything else, and that, I guess, makes it a good business.
Joey
Yeah. And because Anthropic has historically been so focused on enterprise, it was 90%+ enterprise. ChatGPT, by contrast, started heavy in consumer. Back in Q1, 60% of ARR was in consumer, and only 6% of the 950 million weekly active users end up actually paying for a plan. Most people pay for the $20-a-month plan.
It’s a much different business mix because OpenAI has to subsidize—or not subsidize, but there’s a cost to serving—the other 900 million weekly active users who are free. That lowers their gross margins on a blended basis by about 20 points.
That’s changing. We think that, with some of the rise in OpenAI’s API business since late March with Codex, and then 5.5 and 5.6, the 60/40 consumer-enterprise split has flipped now to 40/60 consumer versus enterprise here in Q2. From what we’re hearing in the channel, especially out of token-as-a-service businesses like Bedrock and Foundry, OpenAI has recognized this and is starting to shift. Even net new ARR is starting to look a little more even between Anthropic and OpenAI.
Dylan Patel
Yeah. And it’s kind of shocking. With the release of GPT-5.5 and GPT-5.6, people like Tibo on X who are tweeting about the daily active user counts are saying that they’re celebrating going to 7 million and then 8 million active users. But to be clear, they’ve got over 900 million weekly active users on the free tier of ChatGPT.
They’ve got a lot of room to grow just to get people using the number-one paid product. Maybe they’ve got a few extra users using other paid products, but I would assume the bulk of people using the paid products are using Codex, right? So that’s not a great conversion rate in terms of who’s actually using the full paid product from the base of 900 million.
Joey
Yeah, it’s funny. You actually see it, too, in the paying rates for even the consumer version of Claude. Claude is like 5.9% of free users end up paying, so it’s a much more focused user base.
If we think of those 900 million weekly active users, a lot of it is Google Search-type replacement, things like that. And then, obviously, because of that consumer focus, they’ve been working on advertising a lot. Consumers traditionally monetize poorly: people aren’t willing to pay, and the retention curves are typically pretty poor. How many people are paying 3 months later, 6 months later, or 12 months later?
Ads are a tough business model. It’s tough to scale, but maybe that’s the eventual path. They’ve put out some pretty crazy targets for 2030 advertising revenue. But yeah, consumer is a tough business. Enterprise, given the margins on tokens for API, is obviously the place to be.
I think Anthropic had a pretty good lead because they were so levered to coding early in the first 6 months of this year, and it seems like it’s shifting to more of an equal two-horse race here in mid-July.
Dylan Patel
Yeah. Crystal, what are you seeing in that two-horse race? Are people that you’re talking to using Anthropic’s Claude Code, Codex, or both?
I feel like a few weeks ago, when I was talking to people, they were pretty heavy on using Anthropic’s Claude Code specifically. If I asked them why, they would say, “That’s just what I’ve been using, right? And I’m used to using it now.”
But now, as people in the consumer segment start using ChatGPT a lot more than Claude, that effect in the consumer sector starts trickling into enterprise a little bit. These people who were previously Claude Code power users and didn’t touch Codex at all are kind of like, “Okay, what if I consider tinkering with Codex?”
Even though consumer is a really small portion of the pie and the margins obviously aren’t as strong as enterprise, and historically speaking OpenAI hasn’t done as well in the consumer market as Anthropic has, they have that traction in the consumer space that’s eventually going to trickle in. We’re already starting to see that just this past week.
Dylan Patel
Max, what’s your take? OpenAI versus Anthropic. Where do things stand for you?
Max
I think I remember that the last time I was on this podcast, we were talking about how it was right when 5.5 came out, and we were talking about how things were looking really dire for OpenAI at the start of the year. I think ARR growth was likely flat in March and April, which meant the alarm bells were definitely ringing in Sam’s head when that happened.
But then we said that 5.5 would be an inflection point. This was actually a really good model, probably Opus 4.8-level. I think that prediction has played out even more quickly and more optimistically than we would have guessed.
Some people here are saying that 5.6 is as good as Opus 4.8, sort of just straight up, even though it’s only half the price. We also think it’s a much smaller model, which would be a testament to OpenAI’s training abilities.
As Joey mentioned earlier, net new ARR month over month is comparable between OpenAI and Anthropic now, which is pretty shocking. People were counting OpenAI out a little bit, and I think Jordan might have been one of those people. But they’re back, guys. They’re back.
Dylan Patel
I am nothing if not flexible. It’s always been about the model, guys. Everybody’s like, “When are you going to stop hating on OpenAI?” As soon as they ship a model that’s good—and it’s good.
For the things that I want to do, I’m basically default OpenAI at this point, despite the application not doing what I want it to do with multiple tabs and stuff like that.
Dylan Patel
Wait, are you using the Codex app or the CLI?
Dylan Patel
CLI.
Dylan Patel
What’s your form factor—CLI? I feel like they don’t want you to use the CLI, dude. They very actively want you to use the app. What’s pulling you back?
Dylan Patel
Tabs. I need multiple tabs going.
I think the way you’re supposed to do this—because I agree, the first time I tried using Codex, I was like, every single time I open a new chat, it’s just lost in the ether in the left-hand sidebar of all those tabs.
Dylan Patel
I think how you're supposed to do it is you have 1 pinned thread per project you're working on, and then you ask that thread to spin up subthreads anytime you have a discrete ask. You actually never even interact with the vast majority of threads in the left sidebar. It's just your couple of pinned threads that you use.
Dylan Patel
Yeah, that sucks. Second of all, I'm in remote SSH sessions all the time, and the remote support in the thing sucks. So, yeah, I'm very against both of those.
Dylan Patel
Makes sense.
Dylan Patel
I think they are working on this, but, yeah, I'm just not reading the code, to be clear.
Dylan Patel
Thanks for clarifying. The last time you were on, we had that discussion, but you still have to run some things from the terminal and see some output in the logs at some point. This is just not part of the standard Codex experience, but I'm not the primary market. I understand that they are targeting vibe coders around the world.
We need to help people make more B2B SaaS. That's the core business model of this, and I'm not the target market. That's cool. I'll keep using the CLI. It's all good.
Actually, I wonder if, internally, most of their software engineers are using the app with the CLI. At SemiAnalysis, for sure, I think all of our real engineers are on the CLI. They're diehards, right? They'll never give it up.
Dylan Patel
That would be a good question. You're going to ask them that.
Dylan Patel
Yeah. Well, next time I see some OpenAI people, I'll ask.
Dylan Patel
Sounds good. Yeah.
Dylan Patel
Okay, let's change gears and talk a little bit about who's going to come in third. There have been some releases that we've covered a little bit. SpaceX AI is back. They've got Grok with Cursor bolted on, and things are going well. Meta has Muse Spark out in the world. They've also stacked up a whole bunch of compute, and they've teased launching a neocloud, which, of course, SpaceX was the first to do.
What do you guys want to talk about first: the models or the compute for these guys?
Dylan Patel
Let's do compute. I think it's more interesting. Max is going to walk you through the models.
Max
I mean, to be clear, the models aren't bad, but at this point I don't think it matters unless you can deliver a true frontier model. They're obviously not true frontier models, so who cares?
Dylan Patel
Do you think anyone will use their API product?
Max
I don't think anyone's going to use the API until it's a frontier model.
Dylan Patel
Okay.
Max
Yeah. So when they have an API, their token-as-a-service business will have fire margins, too. I mean, that's only if it's better than whatever the best Anthropic and OpenAI model is at that time, which I think is easier said than done.
I think Grok 4.5 and Muse Spark 1.1 are only useful as proving points along the scaling curve. It's like, okay, these guys, unlike Gemini, are still kind of in the game right now. They can train adequate models. It's not quite frontier yet, but there is a nonzero chance they can catch up. I think that's the way to view them.
I will personally be shifting my token usage to either Grok 4.5 or Muse Spark 1.1.
Dylan Patel
Okay. But you said the other name, which is shockingly, in my view, in fifth place right now: Gemini. Do you still have that view?
Max
Oh, 100%. I think they're clearly in fifth place. Unless Gemini 3.5 Pro is better than the industry chatter we're hearing, I think they're going to stay in fifth place, maybe forever.
Dylan Patel
Maybe forever. Okay. Talk to me about compute. Staying in fifth place forever would be related to the quality of the model, but the quality of the model is tied to how much compute these guys have.
Max
Yeah. I have to say that when I first saw the SpaceX-Anthropic deal, where Elon rented basically all of Colossus 1 to Anthropic, I think I incorrectly viewed that as Elon renting out his compute, when in reality I think the really important question is: do you have the ability to claw back all of your compute, assuming that you have proven you can truly reach the frontier by scaling up your compute?
I think that's the view Elon is taking with Cursor and xAI. He's essentially saying, "I'm going to leave you guys with enough compute to prove that you're capable of reaching the frontier. In the meantime, I'm going to monetize all the additional compute by renting it out on these extremely high-margin, 3x, maybe 4x market-rate rental deals."
I'm going to make sure there's a 90-day clawback clause included in every single one of them, such that if you do really well, I'm just going to claw back all the compute and give it to you so you can make digital costs. I think that is a valid strategy for someone who still truly thinks they can build RSI.
You can think of it as OpenAI and Anthropic monetizing their compute by selling inference, while Elon is doing these short-term compute deals instead. I think this is fine.
On the topic of Meta's compute, there are rumors that they're going to become a neocloud. As long as they do these SpaceX-style deals with a clawback, or they just give it all to MSL, or they do token as a service—any monetization method that allows them, in theory, to claw it all back if MSL proves they're doing well—I think that's an excellent, smart decision on Meta's part.
It probably allows them to be more aggressive with buying compute in the short term, while still giving them a chance to build RSI in the long term. I think the issue with Google, and why they're in fifth place aside from their model, like Gemini 35 Spy Flash being bad, is that they have signed all these long-term compute deals where they cannot claw back the TPUs that they're giving away. That's just committed for the long term, and I think it indicates a lack of conviction on their part that they can build RSI.
I think Noam Shazeer, John Jumper, and all these other high-profile guys at DeepMind leaving is partly because of this fact. They are high-pilled, their leadership is not high-pilled, and they think the only thing they can do in this situation is go work somewhere else.
Dylan Patel
Wow. Okay. Hot takes. The 2 things I want to dig in on there: the first is related to the clawback. This is a valid strategy, but you can basically only do the clawback once, and then nobody's ever going to buy from you or trust you again. Or do you think you can actually do it more than once?
Max
In an ideal world, you only need to do it once, because you only do it if you're confident that you're going to be able to breach the frontier by continuing to scale, right? But I also think it's possible you could do it more than once.
In principle, Anthropic is viewing this as a short-term deal. If Elon said, "Hey, actually, I want it all back 6 months from now," and then a year from now he's like, "Just kidding, we actually still can't train a frontier model," if Anthropic is still ripping at that point and is in desperate need of compute capacity, why would they say no?
Dylan Patel
Yeah, I guess financial motivations can make friends out of enemies. The second thing is just on the topic of the business. Maybe, Joey, I can bring you in here. What's a better business: selling tokens of your frontier model or close-to-frontier model, or selling compute access in the short term at a premium over the market?
Joey
Selling your own tokens is probably the best model. Even selling someone else's tokens, if you're AWS or Azure, is a pretty good business model as well, if you can participate in some of the upside revenue share.
But if you can run above-market rates, it's not a bad business. From the Meta compute perspective, this has been floated for a while. Meta compute has been floated since Q3 of last year.
It gives Zuck a backstop to say, "Okay, if MSL isn't successful, we have the means to essentially make decent ROI on our capex spend." It allows them now to spend a ton more on capex in 2027, potentially 2028, because there's a market. You've gotten these market signals that Anthropic can pay 3x market rates, and you can do a small deal here.
In the background, investors can feel comfortable that if MSL is deemed unsuccessful, Meta could have a pretty nice single-client or multi-client neocloud business just serving the labs.
Dylan Patel
Yeah, but I thought this chart that you guys had in the Meta compute article was striking, where you can see that the SpaceX-plus-Google deal is monetizing this compute at 4x the market rate. I think in the past everything you're saying makes perfect sense, but I would challenge a little bit that these near-term SpaceX GB300 rack deals aren't a really good business. They're clearly high-margin, right?
Joey
Yeah, that's extremely high-margin. When you think about the ARR per gigawatt, I think a lot of people would be happy to get $48 billion of actual token ARR from inference there. So I don't know. That deal doesn't seem very economical for Google.
Dylan Patel
Do you think Meta Compute is them telling a story to the market that they're going to be able to replicate this SpaceX AI thing, which is probably a one-time thing unless demand keeps growing so exponentially that there's such a compute crunch that everybody's fighting for access to GPUs?
While we have them, we can monetize them at 4x the market rate from neoclouds. I think our house view is that a neocloud is a pretty good business on its own at around $12 billion per gigawatt, approaching $50 billion per gigawatt.
Joey
Yeah, and the labs have gotten so good at the ARR-per-gigawatt number. It has gone up a ton.
Dylan Patel
Token throughput’s been pretty impressive—the models as well, and how they’ve monetized them. The new models price higher. If you look at Anthropic, revenue per megawatt was about $16 million per megawatt last year, or $16 billion per gigawatt, and now that’s more than doubled in 3 quarters. I think some people believe that can double again in 2–3 quarters.
Let’s look at that trend. I’ve got this chart on screen right now. Total token-as-a-service market—2 things jumped out at me. One, how fast this thing is growing. I thought the market was pretty big in Q4 of last year, and that looks puny compared to the forecast you guys have on this chart toward the end of this year—way more than doubling, as you said. But the other thing that jumps out to me is just how badly Microsoft Foundry is doing here. Can you—
Joey
Yeah, and that’ll be revised up with the recent OpenAI API success. I think, up until recently, Foundry was 90% OpenAI, and OpenAI was so consumer-focused. This is not the latest version of this chart. Foundry is doing a little better now. Maybe Vertex AI is doing a little worse, given Max’s thoughts on Gemini, than that chart.
Dylan Patel
I think we’re of the view—and when we talk to enterprises, token as a service is a very popular choice for consuming tokens. AWS and Azure have massive businesses with Fortune 500 and Global 2000 companies that spend 9 figures or more on cloud a year. If I can go to an existing vendor and have more model optionality, increase my ELA or my credits, and then burn that down, it’s pretty attractive for both AWS and Azure. If token as a service for Anthropic was 5% of indirect sources 6 months ago, it’s probably 20% of its business now.
Crystal, what’s your take on people using the big 3 hyperscalers’ token-as-a-service offerings versus the smaller startups, like Together, Baseten, and Fireworks—anything they can get on OpenRouter?
I feel like, as we see more of the bigger enterprises, a lot of the financial services industry still hasn’t unlocked and used AI to the level that it probably can and should. That’s where Azure and Bedrock are going to benefit, because most of them probably already buy their services, right? Once they get internal approval or whatever to use these AI tools, they’ll just buy them through whatever existing channel they already have. That’s probably where the bulk of the market is.
Dylan Patel
Makes sense. Yeah. Max, how about you? What do you think of token as a service from startups like Together, Fireworks, and Baseten? How fast are they growing compared to the hyperscalers that are tied to the big labs, but maybe aren’t actually selling a lot of tokens to the startups of tomorrow?
Max
Together, Baseten, and Fireworks are all super impressive businesses. I do actually expect open-source token volumes to grow slower than frontier token volumes, but that’s only because I think frontier token volumes are going to absolutely explode. I think open-source token volumes are also going to explode, just to a slightly smaller degree.
It’s kind of funny. If you talk to the VCs investing in Together, Fireworks, and Baseten, and ask them for their explanation of why, oftentimes they lead with, “Inference is going to be the largest market ever.” It doesn’t even matter if these companies can only capture a super, super tiny slice of this extremely large market. That’s good enough for us.
I think that thesis is honestly more or less right. I don’t expect these guys to ever be doing more volume than OpenAI or Anthropic, or anywhere near that. In terms of global token volume, I think this might be the highest percentage they’ll ever be, rather than today. But they’ll still be good businesses in the future, I think.
Dylan Patel
How about the rumors that they can hitch their wagon to some of the bigger labs? SpaceX AI starts to win a little bit, and then Fireworks grows because they’re exposed to Cursor and actually have some exposure to that growth.
Max
I don’t think that’s going to last long term. I think the only reason Fireworks got that exposure is because originally Composer was post-trained on them, right? If you’re a frontier lab that has made a new model from scratch, there’s no reason to believe Fireworks’ engineers would be better at optimizing that model than your own engineers. I don’t think they can get any share of that margin in the future.
Dylan Patel
Okay. So why do the hyperscalers get a share of that margin—the big 3? I mean, they have the enterprise distribution that the neoclouds don’t. I’m sure Joey can speak more on this, but Anthropic isn’t giving up this margin to Bedrock so that Amazon’s engineers can make Claude run faster on Trainium and B300s or whatever. It’s so all the existing customers that rely on Bedrock can also use Claude models.
Max
Makes sense, yeah.
Dylan Patel
Yeah, I think the deal—I mean, we hear more. We don’t hear anything, but people tell us they think the deal could get reworked, because obviously it was very beneficial for AWS. I think we’ve written about AWS and its margins and how it monetizes versus just selling the bare compute for Anthropic to run inference on. To get that 30%, or 20% to 30%, revenue share just falls down to the bottom line. That’s pretty attractive.
AWS, Azure to a lesser extent, and GCP obviously have massive customer bases. People’s cloud estates are there, and enterprises are very comfortable with security and compliance. With everything now in the cloud, even more regulated industries are comfortable, so it’s a natural place for them to want to buy. But to see that much of the economics for Anthropic is obviously a lot. I think a lot of that was because, when the deal was struck and there was an incentive to make Trainium work, they were able to run with those terms and make it pretty favorable. So, yeah, we’ll see what happens if those revenue-share deals can continue.
Okay, let me make one more attempt at the bull case for these token-as-a-service companies that aren’t the hyperscalers, like AWS and Google. Obviously, a big portion of the benefit of getting token as a service from them is that their engineers are working on Trainium and TPU, and that might be different from the experience that the labs have with GPUs, where they have more experience running this themselves.
There’s a whole class of chip startups that are coming to market right now, and an obvious way in which they come to market is by partnering with these token-as-a-service companies. The biggest frontier labs aren’t going to spend a whole bunch of time optimizing for the 7th-best chip startup that’s coming to market right now. But if they do, and that chip startup strikes something really nice for a given model that makes it a lot more cost-effective or a lot higher-performance to run instead of Nvidia, then they’ve got a shot at doing something super unique.
Do you think that’s a potential future in terms of where these kernel engineers who have learned a lot end up going and spending their time over the next few months or years?
Max
This is a good point, and it’s reasonable in the short term. But if any of these new chip startups actually reach sufficient scale, I think the labs will just dedicate teams to making their model run really well on that chip. I think there’s no world in which you have a new accelerator that’s meaningfully better than Nvidia and is also being sold at large volume, where OpenAI is relying on Together to serve its model on that chip instead of just working directly with that company to develop the first-party capabilities to run its model on that chip.
Dylan Patel
Yeah. You think it’s going to go the way of Cerebras, where the chip companies, to be successful, are effectively going to have to become a neocloud themselves and have an existing relationship with the frontier labs?
Okay, makes sense. Maybe we can finish by talking a little bit about MSL at Meta, particularly, Max. I just loved the crash course on RL that this article turned into—not necessarily how it works from a technical perspective, but how it works from a business perspective: where people buy data, how they build these environments, and what the market looks like.
Can you give an overview? Previously, in a world where everything’s pre-training, whoever has a frontier-class team and the most compute can just train the biggest model and win. They follow the scaling laws and go there. But now there’s a scaling law related to RL. How is this playing out in your mind at a high level?
Max
Yeah. I think it’s really important for everyone to understand that reinforcement learning, or RL, is probably the most important scaling law for improving model capabilities today.
And there are a lot of people who believe that the only thing stopping models from being able to do literally anything a human can do on a computer is having sufficient RL environments. This is sort of like data that lets the model try to complete white-collar tasks itself until it can repeatedly try completing the task and fully learn how to solve it. There’s an entire new industry and supply chain of RL-environment startups whose entire job is to convert real-world, economically viable tasks into these RL environments, which they can then sell to the labs and allow them to use to improve their models.
You see most of the main players on this chart Jordan has pulled up here. I actually think some of these ARR numbers might be slightly understated. I think it’s pretty much consensus that the total data budgets at the frontier labs this year—primarily OpenAI, Anthropic, Google, Meta, and xAI, but also the long tail of companies like Amazon, Microsoft, and Thinking Machines—are going to sum to well over $10 billion. That’s roughly a 10x increase relative to last year. It’s very possible we see 10x again in 2027.
This is a hugely important market for improving AI capabilities. This is probably the only market in the world where customer demand isn’t even a question for all these startups. They will never have a contract turned down by the labs because it’s too expensive. It’s just a question of whether they can scale up creating high-quality data fast enough, and if the answer is yes, the labs will pay any price for it. One other side note: one of the reasons Anthropic’s models are the best at coding today—or at least they definitely were before 5.6 came out—is that they were by far the most aggressive about buying coding data from all these RL-environment startups. I think some of the other labs are starting to realize this and catch on, but it is definitely a super important industry that everyone should be aware of.
Dylan Patel
Can you dig into the process of creating some of these tasks? You ran through this in the article by describing how Meta has moved thousands of engineers into doing this work, and you also dispelled the notion that this work is meaningless, soul-crushing stuff. It’s actually pretty economically valuable and intellectually stimulating. I don’t want to steal your thunder, but you had a nice line on that one.
Max
I think I said it was both potentially more economically valuable and intellectually stimulating than your average big tech job. A lot of people hear the phrase “AI data” and still think, “We have some random people in the Philippines drawing bounding boxes or labeling text as NSFW.” The models have already fully solved that. Your data is only valuable if it’s doing something the models don’t already know how to do.
In the case of software engineering, in order to make a good software-engineering task today, it typically needs to be something that would take a really good human engineer maybe a full day of work to solve. It has to be something the model can’t already solve on its own today. To create this data, you essentially need a really good human engineer to sit down and think of an example problem they would actually want to do. You then need them to create a verifier—usually a set of integration tests, maybe along with a rubric—that can check whether the model successfully completed this day-long task.
Obviously, this is easier said than done. You also need the engineer to write a prompt for the model that asks it to do the task, but the prompt has to fulfill two competing factors. It needs to be 100% clear and unambiguous what you want the model to do, because you can’t incorrectly fail the model during training if it successfully did what your prompt asked for but the prompt wasn’t specific enough and the model didn’t know it had to do some extra thing. At the same time, the prompt needs to be realistic, natural-sounding, and representative of something a human would actually ask an AI to do in the real world. It’s very difficult to get that balance right.
We’ve heard that for the highest-quality coding tasks, the labs are willing to pay well over 5 figures for a single task. That’s already entering the realm of how much you would pay a decent engineer for a full week of work. I think that should dispel any myths about this being easy, mind-numbing work. To all the listeners out there: if you’re looking for a new job and you’re really good at creating RL tasks, you can make 7 figures or more annually at this point. Maybe consider that as a new job option.
Dylan Patel
Yeah, that may have some listeners excited. Can you actually go one click lower? Do you have personal experience with striking that balance between making something easy enough for the AI to do but not impossible for the AI to do? What sort of intuition could you give a listener about what that means?
Max
Honestly, it’s always changing. Back in the day, I actually did sell some environment data to the labs myself. I don’t do it anymore, but it was much easier to create data even 8 months ago than it is today. Today, it often looks like you need to identify a specific failure mode that you’re aware of in the model, and then you need to create RL tasks that specifically target that failure mode.
The only way to know whether it’s at the right difficulty for the model is to have the model try solving it 10 times and see how many times it’s successful. You just repeat that iteration. I don’t know if I can provide any blanket advice on how to find the right difficulty, other than this: if it’s your first time trying to do this, your first thought is almost certainly too easy. Try making it 10 times harder, and maybe you’ll be at the right level.
Dylan Patel
Interesting. Cool stuff.
Max
Yeah.
Dylan Patel
Okay, guys, we got a whirlwind tour. Is there anything you think I’ve missed as we’ve gone through token budgeting, Anthropic’s profit margins, Meta compute, MSL? What have we missed, Crystal?
I don’t know. Nothing I can think of.
Dylan Patel
We’ve just been too busy watching the World Cup. Here we are on Wednesday, July 15, right when we’re recording this. We just got to watch England get knocked out as Argentina stormed back for a nice 2–1 victory. That was crazy. There are a bunch of British people in the office right now, and they’re all depressed downstairs. It was crazy.
Joey, how about you? What’s on your mind as we wrap up here?
Joey
Nothing. Token spend is up and to the right right now, so it’s good. I’m excited for hyperscaler earnings starting next week. We’ve got Google, Amazon, and Microsoft the week after. It should be pretty good on the top line.
Dylan Patel
Okay, let me go around the horn. We’ll close by getting a vibe check from everybody. Joey, what’s your vibe on the market?
Joey
On the market? I don’t know. As long as lab ARR is going up at a good pace and doesn’t decelerate, I think things are fine. Right now, it’s going up. OpenAI is catching up, so vibes should be good. That’s not what you’re seeing in semis over the last few days, but that’s just summer momentum. It doesn’t work, so it’ll come back. People will come back to the office from vacation, and semis will rip in a year. Things will be good.
Dylan Patel
Yeah, we’re going to see an acceleration after the summer pop. Joey’s going to shoot 84, see lab ARR go up, and be happy. Crystal, how’s your vibe?
It’s going good, I have to say. Same thing Joey said. I need to leave San Francisco before the AI bubble pops, so I’m moving away from San Francisco very quickly before it’s too late.
Dylan Patel
Crystal, you think it’s a bubble? Next, we should publicly talk about our bet here.
Yeah.
Dylan Patel
Right now. We have a bet. The loser—Max chose the over-under of Anthropic ARR at $400 billion by the end of 2027.
Joey
2027.
Dylan Patel
Yes, 2027. Does somebody know who took the over and who took the under?
I took the under.
Dylan Patel
I took the over. Jeremy and Joey actually both took the under, but we haven’t decided what the actual bet is going to be yet.
Joey
Not this year. We’ll figure it out.
Did you guys see our pot at Rays last week? I pulled out a Canadian $20. Rake pulled out 200 Singapore dollars, as well as some rupees. We had Dylan pull out some euros. We had 5 different currencies going on at the table.
I will throw in Canadian currency and take the over with Max because we are—
Max
Let’s go.
Dylan Patel
We’re exponential extrapolators here.
Max
Yes.
Dylan Patel
Exactly.
Max
The loser has to write a newsletter post about why they were wrong.
Dylan Patel
Okay, that’s good.
Max
No, that’s what it was. I didn’t realize we agreed to this. I must have missed that in the Slack thread.
Dylan Patel
Are we doing Frontier Lab fantasy here? We need to pick a model for a given week and set up our team.
Max
We’ll bet on Jeremy’s spend.
Dylan Patel
Jeremy’s on Jeremy.
Max
We’ll bet on Jeremy’s weekly spend, and we’ll have an over-under.
Dylan Patel
Over-under.
Max
Get everyone involved. Isn’t Meta building an internal Polymarket or Kalshi? We’ll do that for SemiAnalysis. Someone can code something like Polymarket.
Dylan Patel
Dude, Meta needs to shut down that effort right now, dude. What are they doing?
Max
We’ll have a market for betting on Jeremy’s tokens.
Dylan Patel
There we go.
Max
I like it, guys.
Dylan Patel
Yeah.
Max
No way. Totally fair.
Dylan Patel
Hey, Jeremy, I’ve got 6,000 rupees riding on you. I need you to hammer the data center model dashboard this week, buddy. Okay, well, guys, I appreciate you taking the time. Hopefully the listeners enjoyed it. It devolved a little bit at the end. Let’s all get back to work. Keep tracking those tokens.
Max
Yeah.