第011期 - GPT 5.5 对决 Claude 4.7:OpenAI 绝处反攻(Tokenomics)| Jordan Nanos、Dylan Patel、Doug O'Laughlin、Max Kan
- GPT-5.5 让 OpenAI 重新回到牌桌,此前一段时间被 Max Kan 形容为「实在太惨」。凭借 Opus 4.5 在编码和智能体能力上的跨越式提升,Anthropic 在4月上中旬按同口径收入反超 OpenAI(泄露数据显示,Anthropic ARR 约190亿美元,而 OpenAI 年初约240亿美元);GPT-5.4 则「说实话就是个笑话」,发布说明甚至没有拿它与 Opus 做比较。GPT-5.5 重回前沿,但还不能说「确定优于」Opus 4.6/4.7,尽管 Twitter 舆论机器在发布日极力推动这一说法。
- Dylan Patel 的发布日历是:「两周内大家都会发布」——Google 和 OpenAI 确定在列。「Spud」(5.5)在预训练尚未完成时就出货了,下一次发布预计会补完预训练并加入更多 RL;Google 的版本则预计「基本就是」一次多模态替换。
- Fast 模式的速度溢价正在衰减,但价格仍为 6x:Opus 4.6 fast 的加速倍率已从 2.5x 降至不足 2x(90→70 tok/s,而未变的基准速度为 35–40 tok/s)。这却是 SemiAnalysis 工程师第一次选择更快的 token,而不是质量更高的 token;Jordan Nanos 的判断是,「对现在的用户来说,4.7 的质量相较4.6并没有实质提升」。
- Token 定价开始让连重度专业用户都用不起。Dylan 表示,团队「快要用不起」这种价格;Mythos 的报价含义不明,可能是25美元/150美元,也可能是25美元/125美元,而 Opus 为5美元/25美元;Doug 将其概括为约5倍,fast mode 还要再叠加6x。Doug 花800美元用 AI 抓取数据,而同样任务通过数据增强 API 只需55–100美元,这就是警示案例;成本增长来自新任务(Jevons 效应),而不是旧任务涨价。
- Doug 对前沿模型定价的结构性看空逻辑是:Opus 4.5 可能已经跨过了一道门槛,日常任务多数无需监督即可一次完成。如果在「可能不到一年」的时间里,100–200B 参数规模的模型就能达到 GPT-5.5 级别的智能,可能大多数人都不再需要前沿级智能。
- 基准测试已经退化成「确认模型不是彻底垃圾的体感检查」。Humanity’s Last Exam 是晦涩的选择题;SWE-bench 则从 GitHub 抓取 issue,再用实现范围明确的单元测试评估。同时,4.7 的新 tokenizer 可能让相同输出的成本增加35%;而团队内部的「说真话」版本——「Opus 4.7 其实就是 Sonnet」——则说明了这个模型散发出的体量感有多小。
- Dylan 直截了当地说,中国开源模型与美国的差距又在扩大——背后是算力约束。中国前沿模型的权重恰好能塞进8x H200 pod 的内存域,Ascend 内核只能部分支持 DeepSeek V4 推理;Max 预计 Meta 虽然目前落后、但正在签下巨额算力合同,到 H2’26 或 H1’27 将「甩开所有中国团队」。
- 形态之争正在进行:Dylan 称 CLI 是「注定消亡的遗物」,认为 OpenAI 的 app 才承载真正的智能体编排愿景;Max 则反驳:「一切最终都是 CLI。」纯 maxi 派愿景。Dylan 从 Tri Dao 的群聊里引述了一套 kernel 开发流程:先让 Codex 写,再让 Opus 修正粗糙部分——「顺序不能反过来」,尽管「公司里其他人其实都更喜欢反过来」。
1. GPT-5.5 将 OpenAI 从悬崖边拉回;Anthropic 收入曾反超
- Max 对自己在 SemiAnalysis 发布的首篇文章作出的总结是:The Information 泄露的数据显示,Anthropic ARR 约190亿美元,而 OpenAI 年初约240亿美元;即便不考虑超大规模云厂商采用净额还是总额确认收入的差异,Anthropic「在4月上旬至中旬按同口径计算明显超过了 OpenAI」。核心驱动是 Opus 4.5 在编码和智能体能力上的「真正跨越式提升」:从11月底到4月初,「所有人基本都在用 Opus 4.5、4.6 疯狂跑各种工作负载」。
- GPT-5.4 的结论毫不留情:「说实话就是个笑话」——它的发布说明只与 OpenAI 过去的模型比较,没有拿 Opus 做对照。「这基本已经说明了一切。」到了5.5,Opus 又回到了对比表里,OpenAI 也「重新回到前沿」;但它并不能说确定优于4.6/4.7,尽管 Twitter 舆论机器在发布日期极力推动这一说法,不过「肯定重新回到了牌桌」。
- Dylan 给出的后续发布日历是:「两周内有新版本……Google、OpenAI,也可能 Anthropic。」Spud「有点像预训练根本没做完就直接发布了」,所以接下来大概率是先完成预训练、再进行更多 RL,然后推出新版本;Google 的版本则「基本就是一次多模态替换」。
2. Fast mode:速度溢价正在衰减,且这是首次有人为速度牺牲质量
- Doug 已经没法持续使用 Codex:「使用上限每天都把我硬生生卡住,兄弟。」团队一致认为,OpenAI 的 fast mode「相当于障眼法」:推理深度被削弱,但速度并没有明显提升。Jordan 的分类是,priority mode 收取约2x溢价,买到的是 SLA,而不是更快的交互;5.3 Codex Spark 则「完全是障眼法,那就是另一个模型」。
- Jordan 给出的 Opus 数据显示,4.6 fast 上线时单用户速度约为90 tok/s,普通模式为35–40 tok/s;现在 fast 约70 tok/s,而普通模式仍未变——「花6x的价格,速度甚至不到2x」。
- 用户行为已经说明问题:这是 SemiAnalysis 工程师第一次宁愿选择速度,也不愿购买质量更高的 token。Jordan 的解释是:「对现在的用户来说,4.7 的质量相较4.6并没有实质提升。」
3. 代币经济学:被价格挤出,以及前沿智能是否真的必要
- Doug 通过 OpenAI 与 Cerebras 的交易提出结构性看空逻辑:如果 Opus 4.5 已经跨过一道关键能力门槛,许多日常任务无需监督就能一次完成,那么即便 Cerebras 永远无法运行大于约100–200B 参数的模型,它也可能「不到一年」就获得 GPT-5.5 级别的智能。届时,大多数人可能不再需要前沿级智能。
- Dylan 表示,SemiAnalysis「快要用不起」Mythos fast。Dylan 估算 Mythos 的价格可能是25美元/150美元,也可能是25美元/125美元,而 Opus 为5美元/25美元;Doug 认为大约贵5x,fast mode 还要再叠加6x。Doug 说目前的支出「还算合理」,「但如果再翻倍……我们可能得关掉 fast mode 了,伙计们」——「毛利率很重要」。
- Doug 对成本的关键判断是:「我不认为新模型完成同一个任务会更贵……问题在于,你会拿它去做新的任务。」他还认为,如果 Mythos 的 token 效率更高,那么对很多任务来说,Mythos fast 可能比4.6 fast更便宜。Max 说:「这就叫 Jevons 效应,老兄。」最令人沮丧的未来,是用户不得不先问自己:「这个任务真的值得支付 Mythos fast 的 token 价格吗?」
- 几个样本很能说明问题:Max 用 Opus 4.6 fast 在 DigitalOcean 的 droplet 上搭建基准测试,烧掉了约400美元;Doug 用800美元 token 抓取10,000份员工资料,而同样的数据通过数据增强 API 只需55–100美元——「我只是试图用 AI 把整个海洋煮干」。随后节目用《Rick and Morty》收尾:「你的用途是什么?」「把黄油递给我。」
4. 4.7 与 4.6 打平,基准测试也无法裁决
- Doug 的体感是,指令遵循「客观上变差了」:模型会漏掉 CLAUDE.md 中的指令和 skills,还总是在收工:「今天已经做了很多,去享受你的周末吧。」他明明是在周一使用,于是回了一句:「回去工作。」他的理论是,推理优化叠加用户数增加10–100x,正在拖累体验——「那是4.6的黄金时代,当时还没量化……还没被削弱。那段日子真好。」他的解决方案是:「我需要一款 NIMBY 前沿模型。」
- Doug 认为基准测试无法解决争议:接近前沿是必要条件,但「基准排名第一,不一定意味着你就是最好的模型」;如今它们只是「确认模型不是彻底垃圾的体感检查」。HLE 是「你见过的最晦涩的选择题」;SWE-bench 则从 GitHub 抓取并不具备清晰任务边界的 issue,单元测试还会要求模型输出特定的20词错误信息,而提示词里从未提到这一点。
- Jordan 梳理了4.7的变化:在 high 和 max 之间新增 extra-high 推理档位,支持高分辨率图像,思考过程默认隐藏,加入任务预算,以及词表扩大35%的 tokenizer——这可能使相同输出的成本增加35%。节目中打盹后醒来的 Dylan 反驳说,更大的词表应该压缩输出;Jordan 在概念上承认这一点,但表示实践中4.7的 token 效率反而更低,并追问:如果4.7没有明显变好,为什么要换 tokenizer?
- Doug 说:「我不认为 Anthropic 真会做半成品模型,OpenAI 显然会。」5.3 Codex 就只在代码上做了 RL。真相派的判断是:「Opus 4.7 其实就是 Sonnet,老兄。」Dylan 又补了一句:「而且 Opus 4.7 是 Mitas。」——「这个模型一闻就是小模型。」
5. DeepSeek V4 与中国的算力高墙
- 节目的诱导性问题是:由于算力约束,中国与美国之间的开源模型差距是否又在扩大?Dylan 的完整答案只有一个字:「是。」
- 记录重点提到,V4 支持1M token 上下文,而 Kimi 只有256K,这可能成为长周期智能体任务的优势;此外,曾发布后又撤下的「Reasoning in Visual Space」代码库,也可能暗示多模态版本的存在。但同一位嘉宾也承认,与他们「经常拿来处理各种杂事」的初代 DeepSeek 不同,这一版他们实际上并不怎么用。
- Doug 说,这次没有出现让人惊掉下巴的时刻:「它发布了,只是当前开源水平的代表」,并没有明显优于 Kimi K2.6。真正的新闻在推理优化:如果 Ascend 内核能够部分支持它的推理,「将首次真正为中国解锁更多算力」;同时,它的权重恰好能塞进8x H200 pod 的内存域,而当前 SOTA 没有更大的模型在提供服务。「很明显,他们开始撞上某种算力高墙了。」
- Jordan 认为 DeepSeek 的工程实现很有意思,重点在于注意力机制变体和 KV-cache 压缩;Doug 则质疑更长上下文的实际价值:即使 Opus 把上下文从256K扩到1M,质量也「烂得不行」,而上下文压缩过程非常痛苦。
- Max 对行业斜率的判断是:目前 DeepSeek/Kimi「可能领先」Meta、Grok、Cursor 和 xAI,但算力是关键投入;Meta 正在签下巨额算力合同,也已经消化了先裁员、后重招的包袱。他预计 Meta 会在 H2’26 或 H1’27「甩开所有中国团队」。撇开蒸馏不谈,Dylan 表示 Mistral 蒸馏的是中国实验室的模型,而不是 Anthropic 的模型。
6. CLI 对 app:创新者困境,还是纯 maxi 派愿景
- Dylan 的激进判断是:Anthropic 把 CLI 做成了创新者困境的极致——「CLI 是死路,是 H1’26 和 H2’25 注定留下的遗物」;而 OpenAI 的 app 才承载着「真正的智能体编排平台愿景」,包括语音和多模态能力。Max 另行提到,Codex app 正在加入生成式 UI。
- Max 的极致派反驳是:操作系统甚至不必存在——只要有硬件、终端和 Claude API,就能替你搭出操作系统。「一切最终都是 CLI。纯 maxi 派愿景。」他认为 Claude Code「显然只是 CLI 的封装」,而 Codex CLI「显然只是 app 的封装」。Jordan 则是 VS Code 插件派:「别指控我在读代码。」
- Dylan 在研究高性能 kernel 的群聊里讲了 Tri Dao 的一则轶事:Codex 不聪明,但会直接实现;Claude 则会在小众微架构细节上反复兜圈。所以流程应该是「让 Codex 写,再让 Opus 修。顺序不能反过来。」Max 回应:「其实公司里其他人都更喜欢反过来。」Dylan 说:「我们又不是在写他妈的 Tri Dao kernel。」
7. 长上下文、压缩与假新闻式前沿
- 对上下文压缩,Doug 的态度非常明确:「压缩糟透了……去他的压缩。」还不如清空上下文、重新开始。Jordan 指出,DeepSeek 3.2 论文也得出了完全相同的结论:在他们测试的任务上,超过上下文窗口后,彻底清空的效果甚至优于总结。
- Llama 4 Scout 当初宣布的10M token 上下文为什么还没人真正发布?Dylan 说:「从1M到10M上下文,我到底有什么对下一个 token 生成有用的数据?少得可怜。」同样的数据荒地,也让250K–1M的上下文「反正都是垃圾」,即便是 Opus 也一样。
- 当天走红的创业公司 SubQ 遭到 Max 的评价:「极度、超级、离谱地可疑。」如果真的存在 KV-cache 突破,「内存股今天应该暴跌 10^15%」;而且「我就是不相信这帮人会解决整个 AI 领域最难的问题」。Dylan 的模式匹配是:每次去亚洲,都会冒出一篇美国研究人员从未听说过的 KV-cache 缩减论文;TurboQuant 已经是假新闻。
- 但 Jordan 对市场的判断仍然成立:「资本多得超过了机会。」Dylan 认为 SubQ 可能以10亿美元估值融资5000万美元。
8. 收尾观点
- Jordan 称 Claude Code 是「2026年2月的拐点」,并表示 Doug 在5月的胜利巡礼已经完成。Dylan 希望下一个拐点能更令人兴奋。
Hello, everyone. Welcome back to SemiAnalysis Weekly. We have a big show this week. We have some heavy hitters: Doug, Max—and anybody else out there? Oh, yeah, of course, that's Dylan Patel calling in, nicely in a nice setup office with his headphones on and a SemiAnalysis-logo water bottle. We're going to talk about Claude Code, which is one of the only things we talk about on this podcast now. We're going to talk about GPT-5.5 possibly, DeepSeek, and, yeah, just get into it.
It's crazy. Today's May 5th.
Yeah.
GPT-5.5.
Mm-hmm.
It's Max's birthday.
Pew, pew, pew, pew.
Happy birthday, Max. Max debuted in the SemiAnalysis newsletter last week with his first lead-author article, the coding-assistant breakdown.
He's asking Doug and Michelle, “Hey, am I going to be approved on my work trial?” It's like, bro, obviously. What are you talking about? Also, Max is now a full-time employee at SemiAnalysis, not on a work trial.
Happy to be here. It's really exciting.
Welcome to the team. Okay, so let's run through this article. Basically, when we pushed it out, or when we were preparing to write it, it was effectively just going to be a review of GPT-5.5 Pro—or just 5.5, whatever the new model was. I think we assumed it was going to be called 5.5 when we were testing different checkpoints.
Potato, potato. It's a spud, right?
Yeah.
Potato, potato.
Yeah, we're calling it Spud.
Or is it tomato, tomato?
We tried some different checkpoints. But the basis for the article was just going to be a review of that model, and then Anthropic quickly released Claude Opus 4.7, which was a big change, with some pretty interesting features that we covered in the article. We also got DeepSeek V4, the 4-, 5-, possibly 6-month-delayed release from DeepSeek. So it became this article about all the latest releases. Max, maybe you could give us a quick summary. What was your high-level takeaway? Then we can dig into the details of the changes for these models.
Yeah, happy to. The TL;DR is that things were looking really dire for OpenAI for a while. At the start of the year, Anthropic was quickly encroaching in terms of revenue. The Information leaked $19 billion in ARR for Anthropic versus around $24 billion for OpenAI. Anthropic then quickly surpassed them, even with all the accounting discrepancies related to whether you recognize net or gross revenue from the hyperscalers. I think it pretty clearly surpassed them on a like-for-like basis in early to mid-April.
This was primarily because Opus 4.5 was a real step change in coding and overall agentic abilities. From late November through the end of March and early April, everyone was basically spamming Opus 4.5 and 4.6 for all their workloads. OpenAI tried to fire back with GPT-5.4, but that thing was honestly just an embarrassment. In the model release card, they didn't even compare it to the Opus models—just to past OpenAI models. That tells you all you need to know.
But then, finally, with 5.5, they're back on the frontier. Opus was re-included in the model release card. I wouldn't say it's definitively better than 4.6 or 4.7, despite what the Twitter propaganda machine was trying to push on release day, but I think it's definitely in the conversation. I'm happy to use it, and it allowed OpenAI to come back into the game. Things were looking dire, but now I think they have a shot again.
Doug, what are you using as your daily driver? Are you using both or just one right now?
I've tried so many times to use Codex, but the usage limits raw-dog me every day, bro.
Wait, what do you mean?
I just can't do it.
Use the API, bro.
I use the API. Don't get me started. We already had this—
How do you get usage-limited?
Trust me. I don't know, bro.
Dude, just go bitch in the OpenAI Slack.
I have. I literally—dude. But if it takes me longer than 15 minutes, I don't try to actually fix it.
We're trying to get you started here, Doug, for what it's worth. We're trying to get you started.
If it takes me longer than 15 minutes, I essentially don't try to actually fix it. xhigh seems really expensive and blows out the context window. But, in my impression, it's very neck and neck. They're pretty replaceable to me. The thing that I really want is for someone to give me fast mode and high uptime, and no one has high uptime yet.
Wasn't OpenAI's fast mode fake news?
Yeah, it's pretty fake.
Well, they reduce the reasoning depth, and it's not that fast.
Opus fast mode isn't that fast anymore, either. It's not even 2 times faster.
Sounds like we need more compute.
It started out at 2.5 times faster, and now it's less than 2 times faster. But OpenAI has 3 versions of fast mode: priority mode, fast mode, and GPT-5.3 Codex Spark. GPT-5.3 Codex Spark is definitely fake. That's just a different model. But I feel like fast mode is comparable to Opus fast mode. You guys don't think it's actually faster, or much faster?
I've been using it in Codex, and it feels pretty slow. I don't even notice it really being faster than when fast mode is turned off.
I thought that's what you were showing me, Jordan, the other day: the distribution of what's the peak tokens per second, what's the median, and what's the trough. Or maybe it was someone else, but it seemed like OpenAI's priority mode doesn't actually make it always faster. It can be faster, though.
Yeah. Priority mode is guaranteed execution with an SLA, but it doesn't provide faster interactivity for people who need things done on the API and will pay a premium for that. It's a small premium—around 2 times. I think what I was showing you was the data we have on 4.6 Fast versus 4.6 Base.
Oh.
It started out at, let's say, around 90 tokens per second per user on Fast Mode and around 35 or 40 on Base. It's still 35 or 40 on Base, but it's now around 70 tokens per second on Fast Mode consistently. So it's not even 2 times faster for 6 times the price.
The guys internally are—Max made this point in the article, or I guess we made this point in the article—this is the first time any of the engineers at SemiAnalysis have made the trade-off of wanting fast over higher-quality tokens, and I'm not sure what drives that. People definitely want Fast Mode. It definitely feels faster, but maybe the big thing is that 4.7 is just not meaningfully better quality than 4.6 for people today.
I was actually thinking about this earlier today. If you treat it in the context of the OpenAI-Cerebras deal, and you assume that Opus 4.5 just passed some key threshold in intelligence or model capability such that a lot of your day-to-day tasks are now one-shottable by the models, you don't really have to supervise them at all. You're not even looking at the outputs.
Maybe it's just the case that, even if you can never run anything larger than, let's say, 1 or 200 billion parameters on Cerebras, you're going to have GPT-5.5-level intelligence in that form factor in probably less than a year. It might just be that we've passed the inflection point where the majority of people don't need frontier-level intelligence for their day-to-day workload.
This is probably doubly true if the models keep getting more expensive. Originally, I thought there might be a chance that SemiAnalysis would be able to afford Mythos Fast. I think that's probably not true anymore. A lot of people have already been priced out of the models.
Dylan, can we, though?
I think we're on the cusp of getting priced out.
Was that a definitive no?
I don't think we can, right? Mythos is 6 or 7 times more expensive.
It's $25/$150, I think. $25/$150 or $25/$125, one of the two.
Versus $5/$15, right?
It's $5/$25 for Opus.
$5/$25.
It's 5 times.
So it's 5 times more expensive.
Yeah.
And then Fast Mode is 6 times on top of that. If we were to take our current token spend, I can justify that, but if you were to even double it, I'd be like, “Oh, fuck. Maybe we have to turn off Fast Mode, guys.” At this point, it is—
Yeah.
Margins matter, you know.
Do you think there's a possibility that we could define something so important in SemiAnalysis research in the future that we would want to pay the premium just for one project or one task?
I think the flip side is that Mythos is more token-efficient.
Mythos fast mode is probably cheaper than 4.6 fast mode for most tasks. I just imagine. At least that's the case with Codex versus 5.4 versus 5. Even if they make the model more expensive, that's not a huge jump in price for the model. I don't expect new models to be more expensive to do the same task. The problem with cost is that you're going to do new tasks.
Yeah.
They call that Jevons, dude.
Yeah. We're going to be coming up with new stuff to do with these models based on the work that we do in the next few months with the models. We'll come up with new things, new tasks that are harder and more complex and need to use the bigger models for them. Right?
Yeah. But I feel like right now, at least, I don't even think about whether a task is worthy of spending tokens on. I just spend the tokens. But if the models get much more expensive, you might have to think carefully: Is this task really worth Mythos fast token pricing? I think that'd be really sad, honestly, when it happens.
Question: What is the trade-off? I have a personal anecdote where it's more expensive to burn tokens than to do it. What are your examples of when the cost isn't worth it?
There are times when I was setting up some benchmark on a DigitalOcean droplet, and I was using Opus 4.6 fast, and that was $400 or something. I was like, “I don't know if it's worth $400.”
That wasn't worth 10 minutes of your time?
Yeah.
Jordan, what's yours?
I think stuff where you can very clearly do it from a script or by writing. Writing docs, for example: maybe the first pass with the model is good, but editing stuff is sometimes just annoying to use the model for instead of doing it yourself because it might screw it up or edit the wrong thing.
It's a quality-versus-cost conversation. What's this? It's like, “Wow, this is just a complete waste of tokens.” The tokens in were 2 times more expensive.
I don't think I've had that experience yet, to be honest.
Okay, so I'm going to give mine.
Yeah. What's yours? Yeah.
Okay. Scraping large data sets: at some point, there's a diminishing return. You're like, “Hey, give me 10,000 employees from this company.” And you're like, “Great, I'm going to have it hit the search API.” I'll do a lot of variation to figure out a profile about every person who's worked at this company. Then you actually run the cost, and there's a data enrichment API—and the data enrichment API is literally 1/10 the cost. You're like, “Oh, shit.”
I think I burned $800 in tokens to do what would take maybe $55 to $100 in API calls. One is probably slop, but the other one is, in theory, verified by another slop cannon. I just think there's some value, and that's probably one of the most interesting places where the replacement cost of the tokens versus the actual information—it's still a lot cheaper to essentially serve data via an API. But that cost, the pressure on the top will move that down, if it makes sense.
Yeah.
Just scraping the entire internet is not token-efficient at all. That's a good example where I tried to boil the ocean using AI, and I've been really curious about the trade-off that's going to happen where it's just not worth this much intelligence. Making me coffee—it's the “Rick and Morty” meme. It's like, “What's your purpose?” “Pass me the butter.” This is a waste, man. We've got to find better token efficiency.
Yeah. You can actually hire some people for cheaper to do some menial tasks than you can with tokens. But some of the analysis that we've done recently kind of goes the opposite way so often that you just get used to tokens being the cheaper approach or the faster approach in so many cases that you don't even consider the alternative.
In aggregate, it is, for sure. No way.
So what's the takeaway, Dylan? Is OpenAI so back at this point? Do you think they're going to—
I don't know, man.
—take off on a rocket ship?
New release in 2 weeks. New release in 2 weeks. Everyone's releasing in 2 weeks: Google, OpenAI, maybe Anthropic. I don't know about Anthropic, but Google and OpenAI are definitely releasing in 2 weeks.
What are they releasing?
More everything. More continued pretraining because the Spud—they kind of didn't finish the pretraining and just released it. So, finish the pretraining, do more RL, drop the model. Google is mostly just going to do a multimodal swap.
I guess my question is: the narrative is that, even for the normies who don't use fast mode, 4.7 is worse than 4.6. Do you guys agree with this or not?
I think the instruction following has gotten objectively worse. It keeps missing CLAUDE.md instructions, or you pull a skill and you're like, “Dude, you didn't do exactly what was laid out in the skill.” That seems to be a consistent problem. I don't know; I feel like it's a compute problem more than anything else.
It still has the, “We've done a lot for today. Go enjoy your weekend.” I was like, “It's fucking Monday. Get back to work.” It annoys me so much that it tries to enforce me to stop working, and I'm like, “Ugh, clearly a usage issue.”
I think the 4.6 golden age, when it wasn't quantized in the beginning—those were the days, okay? Fast mode 4.6, pre-nerf. Ugh. But I just think there's this maturity of the models as they become more inference-optimized and more people use them, that it becomes a worse experience as you 10X or 100X the users. And I think that's happened.
I just think 4.7's fine. I think 4.6 and 4.7 are probably the same. It's kind of the same level of experience. Yeah, it's just too many users, bro. I need a NIMBY AI. I need a NIMBY frontier model where no one else uses it except for me, so I can use it. I can get it at a higher rate. That's the appeal of Codex right now, I think—in theory, you should be able to have higher rates.
Can we talk benchmarks? Max, when you saw the 4.7 release, for example, and we reviewed some of the benchmark scores, it was better on most, not better on all, and it led us to talk about where benchmarks are useful and where they're not. Doug's given the vibes, like obviously individual experience can be different across people, and it just drives their preferences, but there should be some objective way to say 4.7 is better or is not better than 4.6.
Honestly, benchmarks obviously try to be that objective measure. I think they just no longer are today. I would say you need to be close to frontier performance in order to have a shot at being the true best model, but being number one on the benchmark ranking does not necessarily imply that you actually are the best model. And so I would say benchmarks today are most useful as a vibe check to make sure that the model's not total trash. We kind of talked about this in the newsletter article, but it's surprising to me how few people actually look into the details of the benchmarks to understand how unrepresentative they are of real LLM use cases.
I think a lot of people just hear a name like Humanity's Last Exam, and they assume, “Oh my God, surely if a model can solve Humanity's Last Exam, that implies that it's smarter than all of humanity, right? We've passed AGI or something.” In reality, you look at the individual questions, and they're just the most esoteric multiple-choice questions you've ever seen that are not at all representative of anything you've ever asked an LLM to do. It's very intentionally multiple-choice to make verification easy, even though obviously when you use an LLM in real life, it's open-ended.
This even applies to benchmarks that, on the surface, you might think would be better, like SWE-bench 2, where it's like, “Oh, coding is this verifiable task. Surely you can just come up with some coding problem for the model and then write some tests that verify if it's successful or not.” But then you dig into the details, and it's actually really hard to write a coding problem that is both naturally worded yet still perfectly unambiguous, with exactly 1 correct solution.
The SWE-bench problems, at least in the original version, don't at all fit those criteria. They just scrape GitHub issues, and all the developers listening know that GitHub issue descriptions are not meant to be well-scoped tasks that you just copy and paste and give to a model.
They often include lots of unit tests that are scoped to particular implementation details, going as far as asking the AI to output a specific 20-word error message that isn't at all mentioned in the task description. Obviously, later versions of SWE-bench tried to solve these issues, but it's still not perfect, and I think it really underscores a lot of the issues with benchmarks.
Makes sense. So let's talk about 4.7 specifically. There were a few things that improved—or changed, let's say—in terms of features, as opposed to just the benchmark itself. I think you made the point in the article that people don't necessarily care as much about the quality of the model anymore. It's the model plus the harness—the product—that should be tested, as opposed to the model on a generic bash-only harness or something like that.
By harness, I mean Claude Code is the thing you're testing, not Opus 4.7. So it's Claude Code versus Codex; it's not Opus versus GPT.
To that end, during the release for 4.7, they announced an extra-high reasoning-effort option that slots between high and max. They announced high-resolution image support, which people can use for screenshots and styling in front-end applications. They're omitting thinking content by default, so people won't see when the model is thinking. They've got this task-budget concept that lets you tell the model how much it should think or work before it can actually run out of context in the context window. And perhaps most importantly, they updated their tokenizer.
It's potentially costing people 35% more for the exact same output from the model than from the previous model, just because they're counting tokens differently and using it. I'm just laughing at Dylan right now. Dude, I literally put him to sleep with that monologue. Holy shit. What the fuck?
Oh my God. Okay, well—
This isn't a joke. No, man, come on. Help me out here.
What are you even talking about, man? The tokenizer is so boring that it puts Dylan to sleep. That's the takeaway.
It's not like the pod was compelling and great at this point, but still—
It is now. I'm compelled. Now we can talk about all the shit we want.
Yeah.
What's the Michelangelo painting or sculpture where the guy is—
The Thinker?
Yeah, The Thinker.
Yo, he's back.
He's back.
Dude, could you actually not hear us through your headphones?
I can hear you.
What do you mean? He's been asleep this entire time.
I was just thinking, man.
You've been gone for a few hours.
Call me GPT-5.5 xhigh.
Just thinking right now. Tokenizer: Opus 4.7, Opus 4.6. We've been dancing around the fact that GPT-5.5 is fine—not goated, definitely good enough, lots of capacity. They'll catch up on the margin. Sounds good? Great.
Yeah.
Anything else? Benchmarks suck. That's a really good take. Anything else?
What's the tokenizer difference?
They changed the tokenizer between 4.6 and 4.7.
The exact same output could have 35% more tokens with the new one.
Oh, they made the tokenizer—They made the vocabulary smaller.
They made the vocabulary bigger for 4.7—
Yeah.
Compared to 4.6.
Oh, sorry.
There are 35% more tokens in the tokenizer—more vocabulary.
So wouldn't that make the average output smaller? Because you can represent longer things with fewer tokens.
No, more tokens.
Well, sorry. If you had only 28 tokens to make English, then you would have to use every letter. But if you wanted to do English with 500 tokens, sure, you'd have every letter, but then you'd also have tokens for “of” and “the.” Wouldn't that make the output smaller, with fewer tokens?
This is a good take.
Oh.
In practice, no, but conceptually, yeah, you could train the model with full words as tokens. But I think what people have seen is that the model is currently less token-efficient with a larger vocabulary.
Oh, interesting.
But that's a good point. Yeah. Man, the guy came back in with a heater here. So, yeah.
Wow. He was thinking all this time. He—Oh.
Okay. Yeah. Well, the whole concept of being more token-efficient is that you can solve tasks with fewer tokens, because the more granular breakup of the tokens—or the larger breakup in the token size, or just having a larger vocabulary—would mean that you would have more information represented in latent space.
The concept of—
The relationship.
SemiAnalysis might be a token instead of “Semi” and “Analysis,” right?
Yeah. But then it creates more context, right? Versus maybe even 4 tokens of SemiAnalysis. I get it, right? There are choices.
But it's richer information. The embedding for SemiAnalysis might be close to the embedding for Dylan in latent space, right? The embedding for Semi or the embedding for Analysis is probably not close to it. So whatever.
I think it's unclear whether this is significantly improving performance, because everybody still thinks it's a toss-up between 4.7 and 4.6 as to what's better. But then I also think it's a toss-up as to whether this is more token-efficient or worse. It seems like people are saying, if it's not a significantly better model, then why are we doing this new tokenizer thing?
Yeah.
So maybe it's just an early checkpoint. They've got to do more RL and improve 4.7, and then have a 4.8 drop that really improves things.
I don't think Anthropic really does half-baked models. OpenAI clearly does. GPT-5.3-Codex was RL'd only on code, and then 5.4 was like—
Is this the most half-baked of any model we've previously seen? We saw Sonnet before Opus, right? 4.5 Sonnet, then 4.5 Opus, then 4.6, and—Right? This is the first new one that's just Opus.
It's because Opus 4.7 is actually Sonnet, dude. Are you not a truther?
Yeah, yeah, and then Opus 4.7 is Mitas.
There you go. Yeah, there you go. This is truther stuff, bro. It's time for truth-truthers, bro. Opus 4.6 was actually Sonnet all along. Does everyone know—Do you not remember that? That was, like, a big model that smelled small.
Yeah, removing the mask. Yeah, yeah.
Exactly. This model smells small. Actually, I wanted to maybe pull this back to the DeepSeek portion of the article, because I don't think we talked about DeepSeek in—
The deep state?
Yeah, the deep-state version of the article. We didn't talk about DeepSeek in depth, but I think we have a lot more internal takes than what we put in the article. And I guess my question is, do you think the gap between open source in China and the United States is now widening again? Because it feels like it is now. And it's because of compute constraints. I'd like to have some takes on that. Just my take is the take.
Yeah.
Yes.
I'll give a quick one. I think one thing that's been overlooked with DeepSeek is the fact that this is a 1-million-token-context-window model, which many of the leading open-source models that perform great on benchmarks people use for coding don't have. And then I think that the “Reasoning in Visual Space” paper that they posted and then took down as a GitHub repo makes me think they're going to release a multimodal version of this as well, or that they're in the process of developing it, and there are going to be new weights for that. Both of those things are fascinating for making China catch up on a product basis.
If we're comparing Claude Code to Codex to DeepSeek in OpenCode, or DeepSeek in some other harness, it'll be able to support all the same features in addition to being pretty good at all the other stuff. With that said, at the time of the first DeepSeek release, I used it all the time for random stuff, and I don't use this one for anything, really. So I—
Isn't Kimi K2.6 better anyway?
Yeah, and I don't use that either.
No, I'm just saying it wasn't like DeepSeek—I mean, the DeepSeek moment was that it was so cooked and it was from the ether, right? Or it cooked so hard, rather, and it was from the ether—not cooked, right? But this round, it comes out, it's just state-of-the-art, or it's state-of-the-open-source art, if that makes sense. It's not exactly better than Kimi K2.6. I think the other stuff in it is clearly just inference optimization, right? They talked about the Ascend kernel being partially able to run inference on it, which would really, really unlock more compute for China for the first time.
Then also, if you look at the weight size, it looks very convenient. I feel like the softmax for China is effectively the size of an H200 8× pod. All the models you're looking at are essentially able to run inference within that memory-domain space, and there's nothing bigger that's served at the state of the art.
Clearly, that seems to be the cap, right? Maybe they can do that, but they won't release it on their B200 pods to the public. It just clearly feels like they're starting to hit some kind of wall. Agree or disagree? Do you think that will keep or cap Chinese progress because they can't run inference on this at all? I'd like to hear some hot takes here.
Okay, I think—
Dude, what the hell was that, Jordan?
He's getting hot and sweaty over DeepSeek.
Yeah, you're fucking deep-panting, bro.
He's getting hot and sweaty over DeepSeek.
The DeepSeek engineering release is fascinating. All the new attention variants and the compression on the KV caches are fascinating stuff. They do so well on the infrastructure stuff. I think, again, the fact that Kimi's 256K context and DeepSeek's 1 million is a significant difference for long-horizon agentic tasks.
But isn't the context from 256K to 1 million dogshit anyway, even on Opus?
Yes.
That's not my—
It's garbage. I mean, okay, it's not true garbage. It's probably a step off of the state of the art, in theory, but the problem is you don't want to just be clearing your context every time. If you're doing a big task, seeing the whole context window is really nice.
Compaction sucks.
Compaction blows.
Not being able to read a 1 million-token context on this stuff sucks. Now—
It's actually better to clear. That's my hot take. I would rather just start over. I would literally be like, “Make a summary of what we've done, copy-paste that, and just start over.” Fuck the compaction.
No, that's what DeepSeek saw in their 3.2 paper. For their benchmarks—which I think are probably bad, but they published this—if you go beyond the context window for a given task, it's better, on the tasks they were testing, to completely clear the context, not even make a summary. There is a compaction—
Yeah, isn't making a summary just what compaction is under the hood? To be clear—
So when I say “make a summary,” I mean literally, okay, you can use the entire context window versus just what I was doing last. I'm trying to remember, because I'm not going to read 1 million tokens of slop. It's like, “Okay, what was I doing here?” Read this, and then I'll Control+C a very small part. So I'm not even—
There are different—
I'm not summarizing the entire thing. I'm just doing the task.
I see. Okay.
I'm passing off tasks. Yeah.
There are different ways to do compaction, but I believe compaction is different from summarization because compaction is removing the thinking traces.
Initially, it's not actually having the model write its own summary of the full context.
Oh, I thought they were literally just taking your entire context and saying, “Yo, please summarize this.”
That is an approach, and they've done multiple of them. But the—
I've got a hot take. Anthropic, with Claude Code, did well with the CLI, so they just kept making the CLI experience amazing. But the CLI experience is not the end-all, be-all of agent orchestration, and therefore they've really cooked themselves into an innovator's dilemma, where they keep making the CLI better. OpenAI has the true vision of what the true agent-orchestration platform of the future is, where you'll be able to integrate voice, multimodality, and all these other things into the app. The app is so much better than the CLI, and the real—
Mm-hmm.
—the point is that you should develop, and users should be using it, in the app, not on the CLI, because the CLI is a dead end, a foregone relic of H1 2026 and H2 2025.
Do you mean the app forever, or do you mean their device? They're talking about releasing a consumer device next year too.
No, no, no. I mean the laptop—
Like that, huh?
—you know, app. The Codex app.
Yeah.
The Codex app.
I have a question then. Why does the Codex app suck?
Yeah. True say.
So, look—
It's long-term planning.
I think—
Long-term planning.
Dylan, in my opinion—
Look—
I think you are thinking too small because, in the perfect, true maxi world—
Yeah.
—the operating system doesn't need to exist. You will just get a piece of hardware. You will plug in your thing. It will pull up the terminal, and you will connect your Claude API, and it will build the OS for you. Thinking—
So, Codex app, they're adding generative UI stuff too, which is pretty interesting.
I'm just saying, I think if you're a coding purist, generative UI is downstream of the CLI. I think I'm a CLI purist. Dude, I don't know. This is just a slop preference thing. I just love the CLI, man. Claude Code usage is clearly just a CLI wrapper, and you can tell, and then Codex CLI is clearly just an app wrapper. I feel like they forced it over. I think there are 2 opinions about the future. Who knows who will win out in the very long run? I'm definitely going to keep it open for competition.
But at this beautiful moment, a true maxi's vision and dream is that it's all downstream from the CLI. It's just tokens. It's the most efficient version of everything, man. Gotta Elon-max, okay? All you need is just an API and then inputs. That's it. Your app and all that stuff, that's all obfuscation. Mythos would know better than OpenAI. Who are we little brains to know what UI we want? No, dude, it's CLI all the way down. Pure maxi vision.
I'm a VS Code plugin guy. I literally tested this yesterday. I was bothering Max, who told me to go away, about using Ghostty for the CLI stuff. It just doesn't work.
Wait, no, I still think having 6 Ghostty terminals open with Claude Code CLI is a superior experience to having 6 different chats going in the Codex app.
Oh, for sure.
I feel more productive.
Yeah, I'm comparing it to the VS Code plugin: 6 different windows in VS Code, also with a file browser on the left side so I can right-click and copy—
Well, I think the difference is that you're still writing real code and you kind of care about the output, whereas I just don't even look at it. I just directly push the—
No, I'm not looking at the code. I'm copying in images or Excel files. Don't accuse me of reading the code. Sorry.
He's fucking Edison.
My favorite is—I mean, I actually agree that the API, CLI, whatever, is the new compiler. No one's reading the compiler. No one cares. They don't need to touch the magic.
This will be true in the future, but it is currently—for any code you actually care about—this is currently not true. It's still producing a bunch of stuff that is bad and should be fixed by coaxing the model to fix it. I'm not saying I type code anymore, but I do read some code.
One of my group chats—
Yeah.
—one of my group chats, I was reading it this morning, and it's a group chat with all the most cracked kernel programmers in the world. It turns out what they do—
Wait, why are you in it?
Because they're my boys, bro. Look, Max, come on. I'm a master networker, okay?
So, anyway, I think it was Tri Dao. TreeDAO's like, “Yeah, dude, Codex is so dumb, but I always just have it create it, and it works and it's smarter, but the code is slop, and then I have Opus rewrite it. But you can't go the other way around. You can't have Opus write the thing and then have Codex fix it. You have to have—
You know what's funny?
Really? I go the other way around.
Codex write it and then have Opus fix it. Everyone else at the firm prefers the other way around.
Yeah.
The entire firm's preference is the other way around, actually.
Yeah, but we're not writing fucking kernels, right? We're not writing fucking Tri Dao kernels.
That's probably fair.
We're doing benchmarks of kernels, but yeah. We wrote some kernels.
Oh, come on. Dude, they're not Tri Dao kernels.
No, they're not Tri Dao kernels. They're just GPU MODE kernel competition kernels.
Because apparently, if you talk about niche and microarchitecture details, Claude will waffle on about shit instead of actually just doing it. Whereas if you describe it to Codex, it'll just try and implement it all, and then it'll be slop. But then you tell Opus to fix it, Opus won't waffle on; it'll just fix it.
Yeah. Doug, this is the context-window stuff, which is like, when you are pumping in so many docs about the ISA of a given GPU to write a kernel or something, you need performance at a million context. You just run out of space on the smaller stuff. So I don't know. Do you want to go back to DeepSeek and any hot takes on DeepSeek, Dylan? Why didn't it crash the market this time if KV cache is reduced by 90%?
Dude, you know, it's been a while since I've been in Asia, but every time I go to Asia, they reference some fucking new paper that reduces KV cache every fucking time for the last 3 years. Some paper, they're reducing KV cache, and no researcher in America has even heard of this paper. It's the fucking best thing ever. DeepSeek and TurboQuant were the most precipitous ones that popped up the most, and TurboQuant was obviously fake news. But yeah, I think it's very funny. I don't know.
I guess they're tired of being robbed.
Okay. Well, look, yeah, I think that's fair. It just doesn't matter. Gemini's working, clearly with the price of the GPU going up. That's all you need to know. Now, if we're gonna talk about real fake news, let's talk about SubQ. Let's do some—I mean, we're not gonna write an article about it. We're not gonna write a post about it. This is free alpha. Did anyone else read the SubQ thing today? It's pretty sus. It's actually extremely ultra-mega sus.
Yeah. It seems like people are launching their startup, right?
Honestly, they should close funding, and then they'd be like, “Wow, it was just Opus with 10 context windows stapled together.” I mean—
Do you think the market is hot enough for them to close, you know, a $200 million at $1 billion round next month or something? If they did, I would be impressed.
I don't know if they could do $200 million, but I think they could do $50 million at a bill. A tril—
A tril? What?
Sorry, bill. No, tril.
There's more capital than there is opportunity.
Wait, wait, wait. Are you familiar with what we're even talking about, Dylan?
No, sorry, I just thought you guys were talking about Anthropic.
No, we're not talking about Anthropic.
Oh.
He's talking about model sparsity. I'm talking about the worst—did you not? It's like this fake-news Twitter thing today called SubQ.
Oh.
Um—
Yeah, yeah, yeah, yeah.
Yeah, yeah.
That's another fake-news one.
Don't worry. We requested API access. We made sure to use our SemiAnalysis email to improve our odds.
Yeah. We're like, “Please, give us this API for this very real model, bro.”
Who knows? Maybe it's a state-space model. Maybe Mamba cooks, or—
No, it's not an SSM. I don't think it's an SSM.
I think it's—I don't know. It's just really funny because, again, we're talking about DeepSeek people freaking out. Dude, if this was real, memory stocks should be down like whatever, a quadrillion percent today. But obviously it's not real because if you look at these guys and you're like, yeah, man, I just don't think these guys are gonna be the guys to crack the single hardest problem in all of AI. No offense—maybe the founder's super legit.
So, okay, maybe one thing this reminded me of was the fact that Llama 4 Scout or Maverick—I think Scout, the smallest one—was released with a 10-million-token context window, or announced with it, but not supporting it officially in the released weights or something. And I'm just really surprised that we haven't seen anybody with effectively an unlimited compute budget give it a go for a more expensive model with a larger context window. Like—
But what—where are you gonna get the data, right? Most people pretrain with 16K context or 4K, you know, something like that, 32K context, and then they post-train it so that they can add and hack in the rest of the context. But it's like, what data do I have? That's why my 250K to 1M context is trash anyways, is because there's no data on this stuff. And so the model doesn't generalize the context really well. And then if you stick it to 10 million, it's like, what fucking data do I have that is useful for the next-token generation that exists from 1 million context to 10 million context? There's so little.
Yeah. I mean, it makes sense. Possibly synthetic stuff, possibly—I mean, why'd they do it in the first place? It seems obvious that people would be working on it, and we haven't even seen anybody announce 2 million. So there's some arbitrary limit—
Wait, but Google serves 2 million.
Google serves 2 million on Gemini 3.1 Pro?
They did on Gemini 2.5—2 million.
Well, maybe that's—
It's 1 million.
That's the answer to me, yeah.
It's 1 million today.
It's 1 million today?
On Gemini 3.1 Pro.
One of their announcements—
No.
They announced 10 million. They started at 1, and then they updated it to 2 at some point in one of the models.
Makes sense, yeah. I mean, they got a big scale-up domain. Why not give it a go with the TPUs? Yeah, maybe another thing that was a little bit missed in the article, and you kind of talked about it when you brought up DeepSeek. Max, I want your take on this. You—because Doug asked the bait question about whether China is catching up or they're still behind. It kind of depends on how you look at it. But I was bugging you the other day: Is DeepSeek or Kimi currently ahead or behind Meta? And are they ahead or behind Grok, Cursor, or SpaceX's xAI?
I would say that today they're probably ahead of all those companies. But the thing that really matters is slope from here. This is a pretty basic take at this point, but I do think the amount of compute you have is actually just one of the key inputs to how good your model's gonna be. And obviously Meta's signing all these monster deals. It seems like they've overcome the overhang of having to fire and then rehire their entire AI team, and they're in the process of making some good models now. So I would expect Meta to pull away from all the Chinese guys, if not in the second half of this year, then in the first half of 2027.
And Meta's not distilling.
Oh, I thought they were.
I thought they were all distilling.
I thought they were distilling from the Chinese guys. They were just running the open-source models.
I mean, that's what Mistral does. They don't distill from Anthropic. They distill from the Chinese guys. That's fair.
Yeah.
Why are we talking about the leading French frontier model company, Mistral?
Dude, you know their revenue's really strong.
Yeah, I do, actually. You know what they've bro'd down on? Also, dude, the bottles.
Because they're a new product.
You keep picking up the bottle on the mic.
Every frontier model—
Well, no, not neocloud. They keep trying to—
…is in a cloud—neocloud.
No, yeah, they're trying to become a neocloud, or at least they're doing fine-tunes.
Every chip company. Cerebras is becoming a neocloud. NVIDIA's launching neoclouds.
The ultimate business model—
AMD—
…for any company in the world is to become a neocloud.
It's starting—
SemiAnalysis will become a neocloud. And then we will be ClusterMAX Platinum.
Diamond. No, dude, we gotta introduce a new tier. Yeah, diamond. Tungsten.
Yeah. SemiAnalysis, lithium.
I don't know, germanium? I don't know, I'm just making up shit. What's the rarest?
What's your favorite semiconductor, Doug? We should make the tiers semiconducting materials only.
Oh, yeah? You like semiconductors?
They already are, but—
Name all of them.
Name them all, yeah. We should be rhodium—the rarest and most expensive precious metal.
Vanadium?
All right, boys, this is getting off track. We’ve got to get out of here.
Yeah, okay. Any other parting takes?
No, I think the hot takes have run out. Claude Code was the inflection point in February 2026. Doug, your victory lap today in May is complete. I appreciate all the hot takes today.
I hope it’s not the next inflection point. I hope it’s more exciting than that.
RLHF is very appreciated. Yeah, better than Dylan.
I made it to the end.
Yeah. He didn’t even make it to the end.
Yeah.