[BidClub_]
SemiAnalysis · · 50 分钟

第036期——200美元买到价值12,000美元的 Opus Tokens:我们买了所有套餐(代币经济学)

Max KanJordan NanosAndrew Megalaa

AI与软件技术企业经营
YouTube ↗
TL;DR
  • Anthropic 的 200美元订阅,按 API 价格计算,能提供约 12,000美元的使用量,而且未必亏钱。 Max 引用了 API 毛利率超过 90% 的市场共识预测;如果平均订阅用户只消耗最高额度的约 10%,订阅业务仍可能实现 50–60% 的毛利率。更尖锐的结论是,API 定价可能属于“惩罚性、剥削性、投机性定价”,而不是每个订阅 token 都在被补贴。

  • 订阅的价值不能被压缩成一个美元数字,因为不同模型和 token 类型对套餐额度的消耗方式不同。 新输入、缓存写入、缓存读取和输出的权重各不相同,因此同一个 200美元套餐,根据工作负载不同,可能对应截然不同的 API 等价价值。最稳妥的表述始终是:“某个模型上的某种工作负载”对应多少价值。

  • 缓存是 Agent 推理经济学中的决定性杠杆。 测试中的 Agent 工作负载约 96% 是缓存读取;由于读取的是已有的 KV cache,而不是重复进行模型计算,其成本可能只有新输入的 0.1x 甚至更低。在不同部署、内存层级和暂停时段中维持高缓存命中率,需要大量系统工程——这“正是区分优秀推理服务商和无用服务商的东西”。

  • 在嘉宾进行的同口径 200美元对比中,Anthropic 提供的可用 token 明显多于 OpenAI。 他们测得约 19亿个 Astra token,对比 31亿个 Fable 5.1 token;但 Fable 只消耗了 Anthropic 额度的 50%,剩余一半仍可用于 Opus。尽管批评者引用 Artificial Analysis 的对比:最高成本的 Opus 5.5 任务约 5美元,而 GPT-5.1-Codex 任务为 0.72美元,Andrew Megalaa 认为,在统一基准设置后,效率差距会收窄至 2–3x,而 Anthropic 提供的 API 等价价值约为 5x。

  • 行业优势最终归属于模型所有者,而套壳服务商在结构上承受更差的经济性。 Anthropic 可以以超过 90% 的毛利率变现 API 流量,而 Perplexity、Cognition 等服务商可能只有相当于 API 成本约 20% 的空间来覆盖自身经济模型,因此只能为 Anthropic 模型提供低得多的额度。节目的投资框架很直接:在 AI 领域,拥有一个“所有人都想用”的模型是巨大优势,而做“套壳”尤其不划算。

  • 订阅的实际毛利率取决于普通用户的消耗远低于公开上限。 Max 认为,如果泄露的财务数据要能自洽,Anthropic 的利用率“不能高于太多 10%”,这意味着订阅毛利率约为 40–70%。OpenAI 宣布将额度上限下调 50% 后,公告前购买套餐的用户尚未受到影响,因为有 1个月的宽限期;这可能把反弹推迟到额度下降真正落地之后。

  • 开放模型可以承接大量“够用”的工作,但这并不意味着前沿模型的经济性失效。 企业完全可以把更简单、对利润率敏感的任务交给 GLM-5.3;当前沿模型在网络安全任务上的拒答阻断工作流时,Jordan Nanos 就这样做了,随后又复用同一上下文,用约十分之一的 token 完成一个简单的 10行 PR。前沿模型论点尚未解决的问题是:经济体能否以快于日常工作迁移到更便宜模型的速度,盈利性地吸收相当于“1亿名超级聪明的博士”的智力供给。

摘要 · 为研究而整理的核心内容

1. 12,000美元的标题数字反映的是 API 定价,并不证明订阅亏损 11,800美元

  • Max 的框架是:花 200美元获得按 API 价格计算价值 12,000美元的 Claude token,毫无疑问是“一笔划算的交易”,但这并不能证明订阅业务的毛利率为负。Anthropic 的 API 毛利率估计超过 90%,而大多数订阅用户会闲置绝大部分额度。

  • 在约 10% 的利用率下,Max 的 tokenomics 模型认为 Anthropic 的订阅毛利率约为 50–60%,而 API 销售的毛利率超过 90%。他的另一种挑衅性解释是,API token 本身带有“惩罚性、剥削性、投机性定价”;Jordan 的回应更简单:愿意购买的用户就在支付市场价格。

  • 团队否定了将一个统一 API 等价价值赋予套餐的病毒式分析。每个套餐本质上都是一篮子内部额度,而每种模型与 token 的组合消耗额度的方式都不同:“Anthropic 的 200美元套餐”只有在指定模型和指定工作负载下,才对应一个具体价值。

2. 4种 token 类型让一个订阅变成多种不同产品

  • Andrew 的分类从新输入开始,成本约为 1x;随后是约 1.25x 的缓存写入,以及约 0.1x 或更低的缓存读取。输出,也就是模型的解码工作,通常最贵:示例中,每百万输出 token 的价格为 50美元,而每百万输入 token 为 10美元。

  • 缓存读取调取的是已有 KV cache,而不是重新处理整段对话。由于自回归聊天会不断携带此前的历史上下文,缓存上下文可能占据总量的大头;当工作负载大部分处于缓存状态时,缓存读取价格几乎可以决定总成本。

  • 这个物理类比很关键:读取缓存主要是搬运内存,而生成输出则是在加速器上进行连续解码。因此,工作负载构成才是比较中的承重变量,而不只是 token 总数。

3. 额度测试通过实验拆出了每一种 token 的价格

  • Andrew 的团队买下每一种套餐,并模仿主流编程 CLI 发出的调用。进行输入和缓存测试时,他们在《战争与和平》一段 token 密集型文本前放入一个小标签,然后要求模型给出约 4个 token 的简单答案,避免输出干扰结果。

  • 随机化标签并省略缓存头,可以隔离新输入;设置缓存头则形成缓存写入;重复完全相同的缓存请求,则可以单独测出缓存读取。对于输出,他们要求模型生成一篇很长的技术文章,最多 16,000个生成 token,再测量每周额度消失了多少。

  • 重复测试形成了多条曲线,由此可以反推出内部权重。他们观察到的 Agent 工作负载约 96% 是缓存读取;作为示例的聊天构成则按四舍五入口径表述为:约 75% 的缓存输入、10% 的未缓存输入和 13% 的输出 token。

4. 缓存命中率是藏在 token 价格里的基础设施护城河

  • 生产环境中的服务商必须在数千、数万乃至数百万次间隔请求之间维持对话状态:把 token 写入某处,将热状态保存在 HBM 中,暂停后降级到更低层级,并让下一次请求回到正确的模型部署,同时避免处理单元过载或闲置。

  • TTL 策略会直接转化为用户成本。缓存可能在最高层级停留 5分钟,也可能停留 1小时;一旦旧对话过期,通过完整输入或缓存写入恢复上下文,成本可能达到读取在线缓存的 10x 甚至更高。

  • Max 表示,Anthropic 早期的比例大约是:缓存读取为 1,输入为 10,输出为 50。这意味着 Anthropic 可能拿走了缓存读取约 99% 的利润。因此,竞争性降价瞄准的是毛利率最高、影响也最大的环节:Astra 的缓存读取仍收费每百万 1美元,而 Fable 为 0.25美元。

  • OpenRouter 的“有效价格”之所以有用,正是因为它纳入了对缓存命中构成的假设。节目的运营结论是:高命中率不是会计层面的细节,而是困难的路由和内存工程问题。

5. Anthropic 在实测套餐中胜出,但模型质量仍有争议

  • 在 200美元档位上,测试得到约 19亿个 Astra token,对比 31亿个 Fable 5.1 token。Fable 的 API 定价在这组比较中也便宜约 40%,原因是其缓存读取价格低 75%。不过 Fable 只用了 Anthropic 额度的一半,剩余 50% 还能兑换数十亿个 Opus token,因此按标题数字绘制的美元图表反而低估了 Anthropic 的优势。

  • 批评者援引 Artificial Analysis 的数据:最高成本的 Opus 5.5 任务约 5美元,而 GPT-5.1-Codex 任务约 0.72美元。Andrew 的反驳是,基准测试并不完美;中等设置下的 Opus 5.5 大致相当于 X high 设置下的 6.1 Soul,统一质量后的成本差距更接近 2–3x。与约 5x 的订阅价值相比,Opus 仍能提供约 2倍的价值。

  • 分歧依然明显。Andrew 认为 Opus 5.5 高出 6.1 Soul 一个层级,而 Max 和 Jordan 强调,模型偏好取决于任务,也带有个人因素。Andrew 的 SemiAnalysis 代码库还会自动拒绝 Pull Request,除非指定模型是 Claude Opus 4.5 或 Claude Sonnet 4.5。

6. 模型所有权、利用率和任务质量决定可持续的经济性

  • 在节目提及的中国套餐中,Moonshot 的表现最弱,API 等价价值约为 7x;一些同业约为 12–15x,接近 OpenAI。没有哪个套餐提供了压倒性的价值,足以让用户为了更弱的模型放弃 OpenAI 或 Anthropic;Meta 的 50美元套餐则被指出可提供约 2,500美元的 Spark 使用量,潜在对应每月 130亿个 token。

  • Max 认为,如果泄露的财务数据要能自洽,Anthropic 的利用率不可能明显高于 10%,这意味着订阅毛利率约为 40–70%。拥有多个账户的重度用户主导了线上讨论,但如果企业席位只消耗其表面 API 价值的一小部分,买方和供应商仍然都能获得满意结果。

  • OpenAI 在 DevDay 前夕约晚上 9点宣布将套餐额度上限削减 50%;公告被一系列积极发布包围,并以 1个月宽限期缓和冲击。嘉宾预计,等到继承原有条款的用户真正感受到更低上限后,市场反应可能发生变化;他们也强调,所有估算都只是某一时点的结果,新模型或新定价都可能让数字迅速改变。

  • token 效率本身就难以基准测试。Andrew 表示,Astra 可能更快完成眼前任务,而 Opus 或 Fable 可能会立即完成后续润色,最终形成更干净的长期代码库;排行榜无法清晰区分模型质量与操作者水平,也无法量化“你做的任何事情中究竟有多少是模型完成的”。

  • 开放模型与前沿模型之争,最终取决于任务需求。更简单的工作可以理性地迁移到更便宜的开放模型,而前沿模型的经济性要求新的高智力机会增长得更快;Andrew 的思想实验是,经济体能否以足够高的回报吸收“1亿名超级聪明的博士”。

  • Jordan 的网络安全经历展示了两面性:模型拒答把他推向 GLM-5.3,而 GLM-5.3 利用已有上下文,以远少得多的 token 完成了一个普通的仪表盘 PR。Jordan 与知名黑客的交流表明,解除限制后的 5.6 Soul、Astra 6 Cyber 或 Mythos 5.1 可能强得多;Andrew 则另行表示,前沿实验室的研究人员可能确实担心网络安全能力。整个交流也提出了“监管抢夺”这一解释,但并未使其更清晰。

完整逐字稿
Jordan Nanos

Hello everyone, welcome back to SemiAnalysis Weekly. This week, I'm joined by Max and Andrew, authors of a very controversial but interesting article, “Anthropic Subscriptions Are Five Times More Profitable Than OpenAI.”

We published this material a few days ago, where the guys conducted limit testing of each AI tariff plan. It's not just OpenAI with Codex and Anthropic with Claude. This also includes Meta, SpaceX AI, MiniMax, Moonshot, Z.ai, Cursor, and Cognition—any service where you can purchase a comprehensive subscription plan.

We'll talk about the conclusions they drew. Guys, welcome to the show.

Max Kan

Thanks for the invitation, Jordan. First of all, a TV or a monitor or something like that—this studio is too professional now. I'm not used to this.

1. Subsidies and Credits

Jordan Nanos

Well, yes. Other people need me to guide them with questions, but for you, I think I can let you lead yourselves. Max, these plans are heavily subsidized, aren't they?

Max Kan

It depends on how you want to define subsidization. Compared to API prices, these plans are obviously more cost-effective. We can put a graphic on the screen, or people can read the news if they haven't seen it yet, but some of the results show that you can pay $200 for the Claude plan and get $12,000 worth of tokens at API prices. So this is definitely a good deal.

However, are these subscriptions subsidized in absolute terms, or does Anthropic have a negative gross margin on subscription sales? I think the answer is no. The reason is that, firstly, their margin on the API is simply incredible. The consensus forecast is over 90% margin at this point.

Secondly, most people don't use the service 100%. If you combine a realistic usage of, say, 10%, with that 90% margin on the API, it turns out that subscriptions are about 50%–60% of the gross margin for Anthropic. So it's still a pretty good business, but compared to their API business, with a margin of over 90%, subscriptions look so-so.

Jordan Nanos

Maybe the conclusion is not that subscription plans are subsidized, but that API pricing for current models is punitive, exploitative, speculative pricing—whatever you want to call it. It's pretty expensive.

Max Kan

They sell to willing buyers at a fair market price. This is capitalism.

Jordan Nanos

Yes, of course. Good. So tell me about 100% utilization and how you did the testing to get those numbers and results. This is a question for Andrew.

Andrew Megalaa

Yes. To figure out the results for subscription plans, we purchased each individual plan and measured each token type. Because of that, we can tell what type of workload we're running. In this case, it's generally agentless. That's exactly what I used these subscription plans for.

That's usually about 96% cache reads, a little bit of input—about 2% or so—and a little bit of output. Mix it together, and you get those big numbers we found.

Max Kan

Before you get into the details, I want to point out that there were already people on Twitter doing similar analyses. They were essentially just presenting a single dollar value as the equivalent of the API cost for a subscription. This is, in fact, an overly simplified and incorrect analysis.

The way to think about your subscription plan is that, by paying the company, say, $200 per month, you get a certain number of credits. It's just some other currency. You should think of each combination of model and token type as costing a different number of credits and consuming your limit differently.

Because the cost ratio between different token-type models may differ from their API cost ratio, the actual API cost equivalent of your subscription plan can vary dramatically depending on which model you use and what workload you are running. This means you can't just make a blanket statement like, “The $200 Anthropic plan gives you $10 of value.”

You should say, “The $200 Anthropic plan, when you do a certain workload on a certain model, gives you such-and-such a dollar amount.” This is the nuance that we highlight in our newsletter, and I really want to make sure people understand that.

2. Token Types and Cache

Jordan Nanos

Yes, that makes sense. Can you go back to basics and simply describe the different types of tokens? Since you mentioned token types, I'm not sure everyone understands what that means.

Andrew Megalaa

When you use the Claude Code CLI or the Codex application, there are different types of tokens used when communicating with the model. The first type is input tokens. Input tokens are not cached. These are simply new tokens that you send to the model.

When you have an agent-based, multipass dialogue, most of your input tokens come not from direct chat, but from, for example, web search, where you get something and then analyze it with a small subagent in one step. In this case, caching is not required, so the cache is not used, and this is considered the cost of the new input token.

The second type is a cache write. Caching means that during a multipass dialogue, instead of reprocessing the entire query each time you access the LLM API server, you can save previous results in a cache and reuse them for faster output and lower costs.

Cache-write tokens are previous tokens that you write to the cache. They're charged once on write, and then when you read them, they're charged as a cache read. The nice thing about reading from the cache is that it's extremely cheap, as you're effectively just getting the data from the KV cache on the LLM server.

So now you have input, cache write, and cache read. Inputs usually have a base price of 1×. A cache write typically costs about 1.25× the cost of the input data. Reading from cache costs about 0.1× or less of the cost of the input data.

Finally, you have the output tokens. This is the output data that the model directly generates. This is usually the most expensive type of token—for example, $50 per million tokens compared with $10 for input tokens.

Max Kan

Yes, it's quite simple. You can draw a parallel between the token type and the actual processing that happens on the accelerator that creates those tokens. In other words, reading the KV cache from memory is a cache read. This is significantly cheaper in terms of performance than creating new tokens, which is the decoding step, right?

Andrew Megalaa

Yes, that's right. And read caching is probably the most important thing in how you configure caching with an LLM, because these models work like autoregressive models. If you don't use caching, you process the entire request again each time, which is extremely time-consuming and expensive.

Most of your chat is actually cached reading because you're processing everything from scratch every time. If everything is cached, it's essentially a continuous cache read. You want the cache-read price to be as low as possible, because it almost completely determines the cost.

Jordan Nanos

Got it. Can you tell us about the technical side?

Max Kan

I was going to say that this was a bit of a tease for our internal X-endpoint test. We really have a high percentage of hits in the cache. This is exactly what distinguishes good inference providers from worthless ones. It's actually a surprisingly difficult problem, but it makes a huge difference in the cost of running the same workload, considering everything I just said about different token prices.

Jordan Nanos

Yes, and OpenRouter recently added this too. If you go to OpenRouter for any model and scroll down past the main provider prices, they have a great “effective price” that shows the real value of the mixed tokens, assuming a certain percentage of cache hits from different providers. That's pretty cool.

So explain why cache-hit rate, or CHR, is such an important metric for those doing inference, and why it's so difficult.

Andrew Megalaa

As we said, cache is a conversational state, right? Why is it difficult to direct a new request to where that history is stored, or to move it to a slower but cheaper storage tier?

If you're running a real production workload where there are many different racks and deployments of the same model, and each of them is processing a batch with a bunch of user requests, when I send, say, the first part of my request, you have to be smart enough to write those tokens to a cache somewhere.

Most likely, they'll stay in HBM at first if you're sure I will write the next prompt immediately. But if there's a big pause, you'll probably offload them to a cheaper memory tier. There's a circle of life there.

When I come back with the next step of my query, you've obviously already processed thousands, tens of thousands, or millions of other queries in the meantime. You have to take my new query, probably send it to the same model deployment, find my previous tokens—which you hopefully have cached somewhere—and then route them to the same deployment and process the next step.

This process of looking up the previous cache, routing the next steps of the same request to the same deployment, and balancing everything properly so that none of your PPUs are overloaded, while avoiding a bunch of PPUs sitting idle and burning money, is all very complex systems design.

Max Kan

Yes, and there are TTL dynamics here too, right? Each new request is assigned a lifetime. It might stay in HBM, at the highest level, for maybe 5 minutes, and another one might stay for an hour or something.

But it's a penalty for the user if you let your old conversation die. When it's no longer cached and you want to continue, you have to pay the full cost of the input, or a cache write that's 10 times more expensive, to restore it. Often, at this point, it's even more than 10 times.

One of the things we pointed out to our All Access subscribers a few months ago is that, if you look at the initial pricing ratios of the Anthropic API, if the cache read was 1, the input tokens were about 10 and the output tokens were about 50. It was a ratio of 1 to 10 to 50.

We noticed recently that, with the price cuts from OpenAI and Anthropic, they were mostly focused on making cache reads cheaper. Why is this happening? First of all, as we have already found out, this is the most important lever for a real reduction in the total effective price. Secondly, as we pointed out to our All Access subscribers a few months ago, cache reads were initially by far the highest-margin token type of the three. It is likely that, at the initial price ratio, Anthropic received about 99% of the profit from cache reads, which is simply incredible.

Now we are finally seeing these prices come down again under competitive pressure. This is one of the points I want to draw attention to. Unfortunately, I don't think this has reached OpenAI yet. Astra still costs $1 per cache read, compared to 25 cents for Fable. Yes, yes, yes. This is actually a good reason why Astra may have a higher dollar value in the Pro 200 plan compared to Fable 5.1 in the pricing plans. Although in reality Fable 5.1 has more tokens and more usage.

Because the cost of reading from the cache is 25 cents per million tokens versus $1 in Astra. And that's not to mention that Fable 5.1 only takes up 50% of your plan, and you have another 50% to use.

3. Model Value Charts

Jordan Nanos

So, explain to me in detail how OpenAI and Anthropic compare in terms of their leading models in these plans.

Max Kan

This is the main graph that many of those reading the article have seen. It shows how Astra directly compare to Fable 5.1 across the different plans they offer, in terms of how much value you get per token when using these models.

Choosing which model you prefer now seems like a very personal matter. I don't think anyone could say, at least for my tasks, that Astra is better than Fable or Fable is better than Astra. They just seem to offer different benefits depending on the tasks I'm trying to accomplish.

Jordan Nanos

But tell me about the actual usage.

Max Kan

Yeah. Strangely enough, I think this is not the graph that gained popularity on Twitter. We can discuss that next. People didn't criticize this one because it shows that OpenAI and Anthropic are roughly comparable for the same $200 or $100 a month, or whatever it is.

It's worth emphasizing one point that people missed and that Andrew talked about earlier. Initially, it was claimed that our different schedule was supposedly unfair to GPT 6.1 Soul because the API price is lower than Opus.

Regarding this main graph, where we compare Astra and Fable: Fable is actually 40% cheaper in API pricing than Astra because their cache reads are 75% cheaper, which is over 9% of your total tokens. If you look at it from a token-count perspective, on the $200 per month plan you get 1.9 billion Astra tokens, but on the same $200 plan you get about 3.1 billion Fable tokens.

Not only that, but those 3.1 billion tokens represent only 50% of your limit, not the full 100%, as Fable usage was limited to only 50% of your token limit. So you can actually use billions more Opus tokens on top of that for the same $200. I think even this graph clearly shows that Anthropic offers better value.

If we were to go to a Soul vs. Opus chart, I think it would show an even more crushing dominance by Anthropic. Maybe Andrew can discuss this.

Andrew Megalaa

Yeah, that was a very controversial graph on Twitter, or X, where everyone was saying, “This is so unfair because the 6.1 Soul is so much more efficient per token than the Opus 5.5.”

Everyone kept citing the Artificial Analysis Intelligence Index analysis of the maximum cost of Opus 5.5 per task and the maximum cost of 6.1 Soul per task, which came out to something like $5 versus 72 cents for GPT-5.1-Codex. People used that to justify what they called an unfair comparison.

There are several reasons why this is problematic. Even assuming that GPT-5.1-Codex is indeed that much more efficient in terms of tokens, mathematically it still works out that Claude Opus 4.5 provides double or triple the value in terms of the number of tokens.

The exact math is this: first, we disagree that the Artificial Analysis Intelligence Index is necessarily a good way to measure the effectiveness of token usage. We will return to this issue later. Even assuming that using these artificial, often imperfect and oversaturated benchmarks is a good way to measure token performance, the Opus 5.5 at max is the number one model in the index, while the 6.1 Soul is not.

And if you just use the Opus 5.5 on medium settings, it beats all of the 6.1 Soul's results, except perhaps max. So, I think specifically the 6.1 Soul and X high, as well as the Opus 5.5 on the medium, have the same scores on the AI Intelligence Index. If you look at the cost per task for those 2 specific runs, the difference is about 2 to 3 times. Then you get 5 times the equivalent API value according to this table.

In reality, you still get twice as much value from Claude Opus 4.5 as from GPT-5.1-Codex, even if you assume that this imperfect benchmark index is a good indicator of token performance.

Max Kan

Yes, it's a shame that the Opus 5.5 doesn't really belong in the 6.1 Soul class. It actually competes directly with Astra.

Andrew Megalaa

I don't know if I completely agree with that.

Max Kan

A lot of people are replacing, for example, GPT-5.1-Codex-Max with Claude Sonnet 4.5. This is Andrew's point of view as an Anthropic fan, okay?

Andrew Megalaa

Maybe that's my point of view. This could be my point of view. But I would say that Opus 5.5 even on imperfect benchmarks still remains number one in a good intelligence index.

A lot of people are replacing, for example, GPT-5.1-Codex-Max with Claude Sonnet 4.5. Anthropic says it is as good as, if not better than, GPT-5.1-Codex-Max. So, in that scenario, you have Claude Sonnet 4.5 as a competitor to GPT-5.1-Codex-Max at half the cost of your plan. That's essentially equivalent to the $200 plan you already get on ChatGPT.

But you also have Claude Opus 4.5, which is a beast for the rest of your plan, costing around 5 or 6 times as much. When you compare that the question of who provides more value compared to the H 2 BT doesn't even stand up. H 2 BT vs. Anthropic subscriptions. Because Opus 5.5 is simply a level above 6.1 in most categories.

Max Kan

For the audience, Andrew has a repository on SemiAnalysis. he has an agent.md where you have to specify which model was used to create a pull request, and if it's not Claude Opus 4.5 or Claude Sonnet 4.5, it automatically rejects the pull request.

Andrew Megalaa

Yes. He's actually not even allowed to contribute to this repository.

Max Kan

Yes, okay. That's pretty much what I was saying earlier. At this point, choosing a model is a very personal decision, and it's not necessarily for everyone. Andrew is forcing all GPT-5.1-Codex-Max fans to upgrade to Claude Opus 4.5 if they want to contribute to this repository, and they saw the light. Everyone who started doing this converted.

Andrew Megalaa

This is true. This is true. It works. They saw the light.

4. Test Methodology

Jordan Nanos

Can you tell me a little more about the methodology? How did you actually measure the number of tokens of each type and how the limits are filled? I think that's what confuses people, too. They run one really big request, and their hourly limit, 5-hour limit, or whatever it is just runs out.

You mentioned at the beginning, Max, that no one uses these things 100% of the time. What scenario did you actually use to figure out how many tokens could be consumed in these time periods?

Max Kan

Instead of running some task or test and seeing how many tokens it uses, we wanted to conduct a kind of scientific experiment: extract the exact values for input tokens, cache writes, cache reads, and output token prices. We do this by mimicking how all popular command-line interfaces call the model API. We're simulating how you would access the inference server as a regular user with a subscription.

We use 2 types of queries that allow us to maximize the type of tokens we're interested in and minimize the types we're not interested in. For input, cache writes, and cache reads, we use a small tag at the beginning of the request, followed by a whole block of text from War and Peace that takes up a large number of tokens. Then we ask the LLM a very simple question to minimize the number of output tokens it generates in response. It's usually about 4 tokens that just point to history, fiction, or something like that.

The reason this works is that, for input tokens, we can set a cache header to indicate whether something needs to be cached or not. We can also randomize the tag before the initial prompt, which essentially ensures that nothing gets cached, and that the API doesn't try to cache it since there is no cache header. This is present in the Anthropic API, OpenAI, and other models.

Then we read the response to check what types of tokens were actually used, so we know what was consumed. For caching, the situation is similar to input, except that we mark it for caching, so we can see that the entire prompt is cached, except perhaps for the last model-generated responses or output tokens.

For reading from the cache, we run a similar experiment and simply set the cache header. One call will be a cache write, which we drop from our count. We then continue sending the same request, and that gives us a cache read for each subsequent request.

Finally, for output tokens, we write a very, very long technical essay, and then the model generates 16,000 output tokens. This consumes our entire budget, and we can look at the response and the counters and see that it used about 2% or 5% of our weekly limit. Output tokens are expensive. They probably use 5% to 10% of the limit in a single call.

By repeating this and building graphs, you can do the calculations and determine exactly how much each type of token is worth. We can then use this for any analysis. For example, if you're running a light load, then with this token distribution, this will be the cost of your subscription.

And with a certain chat workflow, it will be the cost of your subscription, and so on. It’s also worth discussing how you get these ratios. We keep a good eye on what the chat workload actually looks like compared with the agent workload, right—in terms of the input-to-output ratio, cache reads, and cache writes.

Jordan Nanos

Mm-hmm.

Max Kan

Yes, we took the data on agents from our economic panel, where we have the internal SemiAnalysis statistics for all months, and we used the data for September. It was one of the freshest months, so this was our own token distribution.

Everything we do in corporate accounts is the work of agents. Chats are a bit more complicated, because the only way to get a real token distribution is to have ChatGPT with a real workload. But this can be roughly estimated based on the fact that there is a system prompt for each message in the chat, or for each chat thread.

That system prompt is essentially cached for all users of Claude, ChatGPT, and other services. So this alone already accounts for a significant portion of cached reads. That’s where we start, and then we assume that users send fairly short queries. Maybe they leave and don’t respond to the thread anymore, or they come back and write once or twice.

Under these assumptions, we get the distribution given in the article: somewhere around 5% of the input data and approximately 20% of the cached reads. Yes, 13% output tokens, 75% cached input tokens, and 10% of the uncached input tokens. Logically.

5. Other Plans

Jordan Nanos

Makes sense. What about other coding plans? For example, one of the subheadings talks about other consumer plans that aren’t from Anthropic or OpenAI. First of all, does this even matter to anyone? But let’s assume that it does.

What do you think of Muse by Meta, MiniMax, Super Grok Heavy, Cursor Composer, GLM, and Kimmy, and how do they rate their value compared to others? I was a bit surprised that some Chinese developers, like Moonshot, Z.ai, and MiniMax, offer pretty good API value in dollar terms.

Max Kan

The reason this surprised me is that they all have low computational costs. I think if you look at their API prices, the margins there are even lower than those of the leading labs, because it’s a more competitive business. It’s more like a commodity market.

Some of them—Moonshot seems to be the worst—offer about 7 times the value for the equivalent API. But some of them were in the same range, at 12 to 15 times the value, as OpenAI. Considering that OpenAI is the complete opposite in terms of compute costs and has cheap API prices, at least for GPT-5, this was a rather unexpected conclusion for me.

Jordan Nanos

But that’s not true. I guess my main takeaway is that when you look at these graphs, it isn’t obvious that any of these pricing tiers are so much better value for money that you’d want to walk away from OpenAI or Claude, thinking, “Okay, I’m getting a much better return on my Manus subscription, so I’m willing to use a slightly worse model because I get a lot more tokens.”

This is probably 2 or 3 times higher than the OpenAI or Anthropic limits. Or maybe it’s a good thing that it starts at $50 instead of $200 or $500 for access to their top-of-the-line model. But as far as we can see, no one is having a price war on monthly subscriptions right now, right?

Max Kan

Yes, that’s true. I think the main reason is that no one can match Anthropic’s margin on API prices. So if they offer similar equivalent value per dollar, that probably means their real margin is much worse than Anthropic’s.

This only highlights that, in the age of artificial intelligence, having a model that everyone wants to use is a huge advantage. Being a wrapper is especially disadvantageous.

We covered this at the end of the article, but if you look at the use of Anthropic models in Perplexity or Cognition, the limits there are much lower than at the original source. This is obvious if you think about it for 2 seconds, because Anthropic has over 90% of the gross profit, while services like Perplexity or Cognition have, at best, about 20% of the API cost or something like that, compared with API prices.

But it was pretty cool to see that this claim was backed up by data.

Jordan Nanos

Yes, that makes sense. How about a little thank you to Meta for one of the best value plans out there. With a $50 Meta subscription, you can get about $2,500 worth of new Spark features. That’s probably somewhere between 140 million and 260 million tokens per dollar, or 13 billion tokens per month. That’s quite a lot.

If they keep the same rate and the new Manus is a little better, it could be a great deal.

Max Kan

Yes, dude. For just this $50 plan, you can get around 250, maybe 300, tennis-court reservations in San Francisco through the new Spark.

Court reservations for a dollar?

6. Utilization and Limits

Jordan Nanos

Yeah, I’m not sure, dude. I think—yes, yes. Okay, let me ask you a question about this. We’ve posted a similar chart before, Max, and I found it extremely interesting then, and it’s even more interesting now.

We don’t see what the real usage rates of these plans are. Of course, there are some loud “pro users” on X who use all 7 plans at 100%, squeezing the most out of them. And there are other people, a little more sensible, like me, who spend half the day recording podcasts, don’t use my coding agent, or let it run in the background until it freezes, and then I have to move on.

I’m dozing, sleeping, you know.

Max Kan

Are you napping, seriously?

Jordan Nanos

I’m not napping, but—never mind. Yes, yes, but wait. Of course, yes, yes. My daughter is napping.

So, if you had to guess, what percentage gross margin do these companies actually have on the coding plans for these models? Where would you place them? Does the average user use 10%, 100%, or something in between?

Max Kan

I think for Anthropic, this figure cannot be much higher than 10%. The reason is that if you look closely at all their financial-data leaks, the numbers don’t add up if the subscription business has extremely negative margins. I think it has to be a business with a decent margin—say, 40% to 70%—for all the financial-statement leaks to converge.

We have broken down all the details in our tokenomics model. That’s why I believe the utilization rate cannot be much higher than 10%. This also makes sense, considering that many subscribers are actually corporate customers or use business plans.

If you’re a business buying a $1,000-a-month plan for your employees, and they're only spending $800 or $1,000 a month on credit, you’re still happy with that deal. That’s a utilization rate of less than 10%.

Jordan Nanos

Right. In your discussions with people who have these plans, what is their first reaction to what you’re saying? It seems that at first, anyone who has tried paying for tokens and then had access to unlimited plans says something like, “I will never pay for tokens voluntarily, because these unlimited plans are a bargain.”

Roughly speaking, what is the reaction when you talk to them about the article? Do they agree or disagree, or do you think people are at 10% to 20% usage or thereabouts?

Max Kan

The guys on Twitter who wrote a lot about it definitely didn’t agree. These are the 100% users, with 5 accounts, creating new ones every 3 days or something like that.

But I think their sense of the plans somewhat coincides with our actual data. Many people have been saying in recent weeks that Opus 5.5 Auto Cloud 1 plan is, in particular, the best value for money you can get.

It’s also important to note that the 50% reduction in limits for the $200-per-month plans that OpenAI announced at DevDay has not actually taken effect yet for those who purchased the plan before the announcement. They received a 1-month grace period from OpenAI.

So we were pretty surprised that Twitter didn’t smash OpenAI harder for this 50% limit cut, and for providing a bunch of code saying, “Oh, since GPT-7 Luna will be smarter than GPT- 6 Astra, that’s a win for you.” It was kind of pointless, but no one criticized them for it.

I think this is largely because they timed it well with a bunch of positive PR on DevDay, and because they provided that month-long grace period. Maybe in a month, when the limits start to really affect most users, we’ll see more buzz around it.

Jordan Nanos

It was good of them to announce the cuts around 9 p.m. on the eve of DevDay, and then cover them the next day with 20 awesome announcements about all sorts of other things.

Max Kan

Dude, just imagine if Anthropic announced this. They would simply be eaten alive. They would become a meme.

7. Token Efficiency

Jordan Nanos

Logically. Um, it's a good thing we haven't discussed what's worth talking about here yet. Virtually everything related to subscription restrictions. It's probably worth emphasizing that this mailing material is very much tied to a point in time. As new models come out, new subscription levels appear, and new things like ultrafast beta become available—or providers decide, “Hey, I actually want an 8% margin on my subscriptions, not 50%”—we expect all of these numbers to change quite dramatically.

We will constantly monitor this and let you know about updates as they come out. We’ll also keep track of model quality and token efficiency, which will definitely come in handy, I’m sure.

Max Kan

Dude, it’s so hard to actually measure token efficiency well. If only there were some way to do it.

Jordan Nanos

Yes. You can’t rely on benchmark tests, man.

Max Kan

But without benchmark tests, how can you actually verify that 2 models have the same quality?

Jordan Nanos

I don’t know, man. Give 2 people the same task, check in a week, evaluate the quality of the work, and see how many tokens they used.

Max Kan

I think there’s a huge hidden variable here, which is the person doing the task.

Jordan Nanos

So what do you think has the greatest impact on token efficiency: 1 person using 2 different models, say Astro and Fable, or 2 different people using the same model to achieve a goal? Andrew has a strong opinion on this matter. He came, looked at the SemiAnalysis token-efficiency leaderboard, and developed his own strong opinion on this.

Andrew Megalaa

So, if you look at the leaderboard, people who are top users on Astra and Fable are spending about the same amount, but Fable users are using more tokens. That doesn't really give us much information about how much of anything you do, which is the hardest part of answering this question. I would say from personal experience trying to work with 3 models, the effectiveness of OpenAI tokens is definitely noticeable.

I think of it as short-term token efficiency. Yes, I'll complete this immediate task, but I feel like there are so many refinements I'd like to add that Opus or Fable would do right away, while Sonnet or Astra could ignore them to complete the task faster. In that sense, in the long run, I feel that better models create a better codebase that is much more readable.

If you read some of these Astra tests, that's something. But that's just my opinion. I think the models themselves have a pretty big impact on the performance of tokens, even more than which user uses them.

Max Kan

What about resellers of these tools, like Cursor, Cognition, and Perplexity? How do you assess their influence on the choice of model that you can use? Have you ever considered using a router where you switch models, or letting Perplexity or someone else choose a model for a specific task because they think it's better, even though their choice is different from yours?

Andrew Megalaa

I don't use Cursor or Devin myself, but if I did, I'd probably still choose my favorite model. I know Devin has Suite 2 and Fusion, which supposedly use a cheaper model along with a bigger and better one to keep prices low. It looks great on benchmarks, but I haven't personally tested it in real-world conditions.

My overall impression is that if I can afford it and my company is willing to pay, I would definitely choose Opus over a cheaper model like the Suite 2, or even the cheaper Sonic 55, simply because I trust these models to always get things right. I don't have the same confidence in other, lesser-known models.

I think when you work with them, you start to understand what each model is good at and where it can let you down. That is actually quite important when designing systems.

8. Open Models

Max Kan

What about the growing popularity of open models? I saw David Friedberg's statement on the All-In podcast that there has been a shift from closed to open models over the past 12 weeks, from an 80/20 to a 20/80 ratio. We wrote a whole article for our Tokenomics subscribers about how that's not entirely true, because they're only looking at API data.

But I realized that there's a pretty significant shift happening, at least in terms of token volume—though maybe not quite in terms of revenue—where people are really using open models more often. Do you have any personal experience or conversations with others that lead you to believe that frontier labs are doing something wrong with the pricing or quality of new models? Are they not sufficiently ahead of the open ecosystem, forcing people to switch to open models and stay with them?

Andrew Megalaa

I guess the question is how far ahead the closed-source models are. Actually, I think it's just a matter of how high-quality a model I need for my task. It's true that many fields of software development, and intellectual work in general, don't require the intelligence of GPT-4 or GPT-5. Therefore, I believe that many businesses, especially those with low margins, are making a rational decision to delegate simpler tasks to these increasingly powerful open models.

I think this will continue to happen in the future. The question is whether the frontier-lab business model is sustainable. It comes down to your belief that the opportunity space for even smarter AI will outweigh the share of tasks moving to cheaper models.

I think this is perhaps the biggest difference in views between those who believe in frontier labs and those who believe in open source. If you magically created 100 million super-smart PhDs—experts in all fields who didn't need to eat or sleep—would the economy be able to absorb that quickly with a high rate of return?

Do you think that, on a planetary scale, the world is simply not smart enough to absorb this? I think sensible people have very different views on this issue. If you believe the answer to that question is yes, you're optimistic about Anthropic. If the answer is no, you probably think that we're going into a recession relatively soon. If you're really optimistic, type together.

9. Cyber and Capture

Jordan Nanos

Yeah, interesting. I have a reason to ask this question because I've been using Perplexity so much in Slack lately, specifically GLM-5.3, because of all the denials from OpenAI and Anthropic on anything cybersecurity-related, which made my job impossible. I had to use GLM-5.3, and I think that led me to discover that it's pretty good for a lot of things.

They have a model with a level of harm, right? I mean, it was just amazing in cybersecurity tests, and I literally can't test it. For me, this is the advanced cybersecurity model because I literally can't use any other model. It doesn't matter—I can only look at the model scorecards. I can't really use other things, even though I'm in the CVP program to get verified for all this cyber work.

I still can't get around these restrictions. I need a more creative approach to prompts to get around the limitations built into the model itself, rather than the classifier outside of it. But anyway, it's a waste of time trying to convince a model that my grandmother is in danger and that I'll bribe her with a peanut butter sandwich if she doesn't do this thing.

It's easier to switch to GLM, and it will do the job. What happened is that I was in the same chat history, with all the data there, and I didn't switch to something else. I just asked it to make a PR for the dashboard repository to post the data we got as a result of that work, using 10 times fewer tokens.

Max Kan

As for the price, she did this PR for the dashboard. Hmm, no problem. I didn't really need Fable 5.1. For, you know, 10 lines of code. This is also music to Dylan Patel's ears. Many things definitely don't require GPT-5.1.

Jordan Nanos

Yes, but I also feel like I use these things day and night. My wife will tell you that I stay up late using them, and I don't make the top 10 on the leaderboard. So I still have no idea what the hell you actually do with all these tokens.

Max Kan

We're constantly out of the top 10, you guys. You really should interrogate guys like Jeremy, Kyle, Andrew Wagner, and the others. Jeremy will be back soon. When I'm done asking him about data centers, I'll ask him what this fast mode actually gives him and how many subagents he uses.

He told me he was just taking it. He once admitted that he chooses arbitrary odd numbers and says that this is exactly how many subagents to run for a task that doesn't require them at all. It just feels like he's saying, "Give me 37 agents to investigate all the permits for this one data center site."

It is a matter of honor for them to do this. Then he's like, "Oh my God. That was $2,000. Oh well."

What? This is research, guys. This is research. Don't touch my husband. Don't question my methods.

Jordan Nanos

You know, I want to add something else about GLM-5.3. After talking to very well-known hackers recently—really world-leading hackers—it seems that even the Soul 5.6, if it has the cyber capabilities unlocked, is a much, much better model than GLM-5.3 in real cyber tasks.

While some of these models may show high results in benchmarks, I think that Soul 5.6 Cyber, Astra 6 Cyber, and Mythos 5.1 with unlocked capabilities can be much more powerful than many people think.

Andrew Megalaa

Yes, there is something a little scary. There is a real possibility that the people working at these frontier labs have access to the models earlier than others and understand their capabilities better. They are not necessarily lying or building schemes for complete regulatory takeover, but are genuinely concerned about the cyber capabilities of these models.

Jordan Nanos

Do you think it's possible that this is exactly the scenario that's unfolding now?

Andrew Megalaa

I think it's quite possible. But the question is whether they are seeking regulatory takeover.

Jordan Nanos

No, I'm just giving a short answer. I mean, the researchers are sincere when they say—

Andrew Megalaa

I'm giving you a short answer to that question. Maybe researchers—

Jordan Nanos

Sorry, sorry. OpenAI just released a new repository called openai/math, where they seem to have made more progress in this field than in the last 10 years combined. I got distracted.

Andrew Megalaa

Sorry about that, but I didn't hear your question.

Jordan Nanos

I think they just solved the Millennium Prize problem.

Andrew Megalaa

I think that's a better answer to my question than what I asked, which was, "Yes, maybe the people who see these breakthroughs months before us and are really worried about what's coming next are being genuine in their concerns, rather than trying to implement some convoluted regulatory takeover scheme."

Jordan Nanos

That's exactly what I said.

Andrew Megalaa

More broadly and, well, this, this, this is a ridiculous idea, Jordan. This is clearly a regulatory grab. Oh my God. Anyway, guys, I think we've gone off topic enough, we need to read an OpenAI blog post about some crazy, uh, mathematical stuff. So it seems the entire repository on GitHub is openai/math. It's like a high level and just a bunch of stuff. We'll see you later. And then, and then, and then about 10 more tasks that I've never heard of, but which are probably very important. I will study math tonight. Well, guys, thanks for joining. Uh, we'll talk about subscription limits and token efficiency another time. Sounds good. Thank you for inviting us. Thank you, Jordan. See you guys. See you later.