Jordan Nanos
Hello, everyone. Welcome back to SemiAnalysis Weekly. We have a big show this week. We have some heavy hitters: Doug, Max—and anybody else out there? Oh, yeah, of course, that's Dylan Patel calling in, nicely in a nice setup office with his headphones on and a SemiAnalysis-logo water bottle. We're going to talk about Claude Code, which is one of the only things we talk about on this podcast now. We're going to talk about GPT-5.5 possibly, DeepSeek, and, yeah, just get into it.
Dylan Patel
It's crazy. Today's May 5th.
Jordan Nanos
Yeah.
Dylan Patel
GPT-5.5.
Jordan Nanos
Mm-hmm.
Dylan Patel
It's Max's birthday.
Max Kan
Pew, pew, pew, pew.
Jordan Nanos
Happy birthday, Max. Max debuted in the SemiAnalysis newsletter last week with his first lead-author article, the coding-assistant breakdown.
Dylan Patel
He's asking Doug and Michelle, “Hey, am I going to be approved on my work trial?” It's like, bro, obviously. What are you talking about? Also, Max is now a full-time employee at SemiAnalysis, not on a work trial.
Max Kan
Happy to be here. It's really exciting.
Jordan Nanos
Welcome to the team. Okay, so let's run through this article. Basically, when we pushed it out, or when we were preparing to write it, it was effectively just going to be a review of GPT-5.5 Pro—or just 5.5, whatever the new model was. I think we assumed it was going to be called 5.5 when we were testing different checkpoints.
Dylan Patel
Potato, potato. It's a spud, right?
Jordan Nanos
Yeah.
Dylan Patel
Potato, potato.
Jordan Nanos
Yeah, we're calling it Spud.
Dylan Patel
Or is it tomato, tomato?
Jordan Nanos
We tried some different checkpoints. But the basis for the article was just going to be a review of that model, and then Anthropic quickly released Claude Opus 4.7, which was a big change, with some pretty interesting features that we covered in the article. We also got DeepSeek V4, the 4-, 5-, possibly 6-month-delayed release from DeepSeek. So it became this article about all the latest releases. Max, maybe you could give us a quick summary. What was your high-level takeaway? Then we can dig into the details of the changes for these models.
Max Kan
Yeah, happy to. The TL;DR is that things were looking really dire for OpenAI for a while. At the start of the year, Anthropic was quickly encroaching in terms of revenue. The Information leaked $19 billion in ARR for Anthropic versus around $24 billion for OpenAI. Anthropic then quickly surpassed them, even with all the accounting discrepancies related to whether you recognize net or gross revenue from the hyperscalers. I think it pretty clearly surpassed them on a like-for-like basis in early to mid-April.
This was primarily because Opus 4.5 was a real step change in coding and overall agentic abilities. From late November through the end of March and early April, everyone was basically spamming Opus 4.5 and 4.6 for all their workloads. OpenAI tried to fire back with GPT-5.4, but that thing was honestly just an embarrassment. In the model release card, they didn't even compare it to the Opus models—just to past OpenAI models. That tells you all you need to know.
But then, finally, with 5.5, they're back on the frontier. Opus was re-included in the model release card. I wouldn't say it's definitively better than 4.6 or 4.7, despite what the Twitter propaganda machine was trying to push on release day, but I think it's definitely in the conversation. I'm happy to use it, and it allowed OpenAI to come back into the game. Things were looking dire, but now I think they have a shot again.
Jordan Nanos
Doug, what are you using as your daily driver? Are you using both or just one right now?
Doug O'Laughlin
I've tried so many times to use Codex, but the usage limits raw-dog me every day, bro.
Dylan Patel
Wait, what do you mean?
Doug O'Laughlin
I just can't do it.
Dylan Patel
Use the API, bro.
Doug O'Laughlin
I use the API. Don't get me started. We already had this—
Dylan Patel
How do you get usage-limited?
Doug O'Laughlin
Trust me. I don't know, bro.
Dylan Patel
Dude, just go bitch in the OpenAI Slack.
Doug O'Laughlin
I have. I literally—dude. But if it takes me longer than 15 minutes, I don't try to actually fix it.
Jordan Nanos
We're trying to get you started here, Doug, for what it's worth. We're trying to get you started.
Doug O'Laughlin
If it takes me longer than 15 minutes, I essentially don't try to actually fix it. xhigh seems really expensive and blows out the context window. But, in my impression, it's very neck and neck. They're pretty replaceable to me. The thing that I really want is for someone to give me fast mode and high uptime, and no one has high uptime yet.
Dylan Patel
Wasn't OpenAI's fast mode fake news?
Doug O'Laughlin
Yeah, it's pretty fake.
Jordan Nanos
Well, they reduce the reasoning depth, and it's not that fast.
Opus fast mode isn't that fast anymore, either. It's not even 2 times faster.
Doug O'Laughlin
Sounds like we need more compute.
Jordan Nanos
It started out at 2.5 times faster, and now it's less than 2 times faster. But OpenAI has 3 versions of fast mode: priority mode, fast mode, and GPT-5.3 Codex Spark. GPT-5.3 Codex Spark is definitely fake. That's just a different model. But I feel like fast mode is comparable to Opus fast mode. You guys don't think it's actually faster, or much faster?
Doug O'Laughlin
I've been using it in Codex, and it feels pretty slow. I don't even notice it really being faster than when fast mode is turned off.
Dylan Patel
I thought that's what you were showing me, Jordan, the other day: the distribution of what's the peak tokens per second, what's the median, and what's the trough. Or maybe it was someone else, but it seemed like OpenAI's priority mode doesn't actually make it always faster. It can be faster, though.
Jordan Nanos
Yeah. Priority mode is guaranteed execution with an SLA, but it doesn't provide faster interactivity for people who need things done on the API and will pay a premium for that. It's a small premium—around 2 times. I think what I was showing you was the data we have on 4.6 Fast versus 4.6 Base.
Doug O'Laughlin
Oh.
Jordan Nanos
It started out at, let's say, around 90 tokens per second per user on Fast Mode and around 35 or 40 on Base. It's still 35 or 40 on Base, but it's now around 70 tokens per second on Fast Mode consistently. So it's not even 2 times faster for 6 times the price.
The guys internally are—Max made this point in the article, or I guess we made this point in the article—this is the first time any of the engineers at SemiAnalysis have made the trade-off of wanting fast over higher-quality tokens, and I'm not sure what drives that. People definitely want Fast Mode. It definitely feels faster, but maybe the big thing is that 4.7 is just not meaningfully better quality than 4.6 for people today.
Doug O'Laughlin
I was actually thinking about this earlier today. If you treat it in the context of the OpenAI-Cerebras deal, and you assume that Opus 4.5 just passed some key threshold in intelligence or model capability such that a lot of your day-to-day tasks are now one-shottable by the models, you don't really have to supervise them at all. You're not even looking at the outputs.
Maybe it's just the case that, even if you can never run anything larger than, let's say, 1 or 200 billion parameters on Cerebras, you're going to have GPT-5.5-level intelligence in that form factor in probably less than a year. It might just be that we've passed the inflection point where the majority of people don't need frontier-level intelligence for their day-to-day workload.
This is probably doubly true if the models keep getting more expensive. Originally, I thought there might be a chance that SemiAnalysis would be able to afford Mythos Fast. I think that's probably not true anymore. A lot of people have already been priced out of the models.
Jordan Nanos
Dylan, can we, though?
Dylan Patel
I think we're on the cusp of getting priced out.
Jordan Nanos
Was that a definitive no?
Doug O'Laughlin
I don't think we can, right? Mythos is 6 or 7 times more expensive.
Dylan Patel
It's $25/$150, I think. $25/$150 or $25/$125, one of the two.
Doug O'Laughlin
Versus $5/$15, right?
Dylan Patel
It's $5/$25 for Opus.
Doug O'Laughlin
$5/$25.
Dylan Patel
It's 5 times.
Doug O'Laughlin
So it's 5 times more expensive.
Dylan Patel
Yeah.
Doug O'Laughlin
And then Fast Mode is 6 times on top of that. If we were to take our current token spend, I can justify that, but if you were to even double it, I'd be like, “Oh, fuck. Maybe we have to turn off Fast Mode, guys.” At this point, it is—
Jordan Nanos
Yeah.
Doug O'Laughlin
Margins matter, you know.
Jordan Nanos
Do you think there's a possibility that we could define something so important in SemiAnalysis research in the future that we would want to pay the premium just for one project or one task?
Doug O'Laughlin
I think the flip side is that Mythos is more token-efficient.
Mythos fast mode is probably cheaper than 4.6 fast mode for most tasks. I just imagine. At least that's the case with Codex versus 5.4 versus 5. Even if they make the model more expensive, that's not a huge jump in price for the model. I don't expect new models to be more expensive to do the same task. The problem with cost is that you're going to do new tasks.
Jordan Nanos
Yeah.
Max Kan
They call that Jevons, dude.
Jordan Nanos
Yeah. We're going to be coming up with new stuff to do with these models based on the work that we do in the next few months with the models. We'll come up with new things, new tasks that are harder and more complex and need to use the bigger models for them. Right?
Max Kan
Yeah. But I feel like right now, at least, I don't even think about whether a task is worthy of spending tokens on. I just spend the tokens. But if the models get much more expensive, you might have to think carefully: Is this task really worth Mythos fast token pricing? I think that'd be really sad, honestly, when it happens.
Doug O'Laughlin
Question: What is the trade-off? I have a personal anecdote where it's more expensive to burn tokens than to do it. What are your examples of when the cost isn't worth it?
Max Kan
There are times when I was setting up some benchmark on a DigitalOcean droplet, and I was using Opus 4.6 fast, and that was $400 or something. I was like, “I don't know if it's worth $400.”
Jordan Nanos
That wasn't worth 10 minutes of your time?
Max Kan
Yeah.
Doug O'Laughlin
Jordan, what's yours?
Jordan Nanos
I think stuff where you can very clearly do it from a script or by writing. Writing docs, for example: maybe the first pass with the model is good, but editing stuff is sometimes just annoying to use the model for instead of doing it yourself because it might screw it up or edit the wrong thing.
Max Kan
It's a quality-versus-cost conversation. What's this? It's like, “Wow, this is just a complete waste of tokens.” The tokens in were 2 times more expensive.
Jordan Nanos
I don't think I've had that experience yet, to be honest.
Max Kan
Okay, so I'm going to give mine.
Jordan Nanos
Yeah. What's yours? Yeah.
Doug O'Laughlin
Okay. Scraping large data sets: at some point, there's a diminishing return. You're like, “Hey, give me 10,000 employees from this company.” And you're like, “Great, I'm going to have it hit the search API.” I'll do a lot of variation to figure out a profile about every person who's worked at this company. Then you actually run the cost, and there's a data enrichment API—and the data enrichment API is literally 1/10 the cost. You're like, “Oh, shit.”
I think I burned $800 in tokens to do what would take maybe $55 to $100 in API calls. One is probably slop, but the other one is, in theory, verified by another slop cannon. I just think there's some value, and that's probably one of the most interesting places where the replacement cost of the tokens versus the actual information—it's still a lot cheaper to essentially serve data via an API. But that cost, the pressure on the top will move that down, if it makes sense.
Jordan Nanos
Yeah.
Doug O'Laughlin
Just scraping the entire internet is not token-efficient at all. That's a good example where I tried to boil the ocean using AI, and I've been really curious about the trade-off that's going to happen where it's just not worth this much intelligence. Making me coffee—it's the “Rick and Morty” meme. It's like, “What's your purpose?” “Pass me the butter.” This is a waste, man. We've got to find better token efficiency.
Jordan Nanos
Yeah. You can actually hire some people for cheaper to do some menial tasks than you can with tokens. But some of the analysis that we've done recently kind of goes the opposite way so often that you just get used to tokens being the cheaper approach or the faster approach in so many cases that you don't even consider the alternative.
Doug O'Laughlin
In aggregate, it is, for sure. No way.
Jordan Nanos
So what's the takeaway, Dylan? Is OpenAI so back at this point? Do you think they're going to—
Dylan Patel
I don't know, man.
Jordan Nanos
—take off on a rocket ship?
Dylan Patel
New release in 2 weeks. New release in 2 weeks. Everyone's releasing in 2 weeks: Google, OpenAI, maybe Anthropic. I don't know about Anthropic, but Google and OpenAI are definitely releasing in 2 weeks.
Jordan Nanos
What are they releasing?
Dylan Patel
More everything. More continued pretraining because the Spud—they kind of didn't finish the pretraining and just released it. So, finish the pretraining, do more RL, drop the model. Google is mostly just going to do a multimodal swap.
I guess my question is: the narrative is that, even for the normies who don't use fast mode, 4.7 is worse than 4.6. Do you guys agree with this or not?
Doug O'Laughlin
I think the instruction following has gotten objectively worse. It keeps missing CLAUDE.md instructions, or you pull a skill and you're like, “Dude, you didn't do exactly what was laid out in the skill.” That seems to be a consistent problem. I don't know; I feel like it's a compute problem more than anything else.
It still has the, “We've done a lot for today. Go enjoy your weekend.” I was like, “It's fucking Monday. Get back to work.” It annoys me so much that it tries to enforce me to stop working, and I'm like, “Ugh, clearly a usage issue.”
I think the 4.6 golden age, when it wasn't quantized in the beginning—those were the days, okay? Fast mode 4.6, pre-nerf. Ugh. But I just think there's this maturity of the models as they become more inference-optimized and more people use them, that it becomes a worse experience as you 10X or 100X the users. And I think that's happened.
I just think 4.7's fine. I think 4.6 and 4.7 are probably the same. It's kind of the same level of experience. Yeah, it's just too many users, bro. I need a NIMBY AI. I need a NIMBY frontier model where no one else uses it except for me, so I can use it. I can get it at a higher rate. That's the appeal of Codex right now, I think—in theory, you should be able to have higher rates.
Jordan Nanos
Can we talk benchmarks? Max, when you saw the 4.7 release, for example, and we reviewed some of the benchmark scores, it was better on most, not better on all, and it led us to talk about where benchmarks are useful and where they're not. Doug's given the vibes, like obviously individual experience can be different across people, and it just drives their preferences, but there should be some objective way to say 4.7 is better or is not better than 4.6.
Doug O'Laughlin
Honestly, benchmarks obviously try to be that objective measure. I think they just no longer are today. I would say you need to be close to frontier performance in order to have a shot at being the true best model, but being number one on the benchmark ranking does not necessarily imply that you actually are the best model. And so I would say benchmarks today are most useful as a vibe check to make sure that the model's not total trash. We kind of talked about this in the newsletter article, but it's surprising to me how few people actually look into the details of the benchmarks to understand how unrepresentative they are of real LLM use cases.
I think a lot of people just hear a name like Humanity's Last Exam, and they assume, “Oh my God, surely if a model can solve Humanity's Last Exam, that implies that it's smarter than all of humanity, right? We've passed AGI or something.” In reality, you look at the individual questions, and they're just the most esoteric multiple-choice questions you've ever seen that are not at all representative of anything you've ever asked an LLM to do. It's very intentionally multiple-choice to make verification easy, even though obviously when you use an LLM in real life, it's open-ended.
This even applies to benchmarks that, on the surface, you might think would be better, like SWE-bench 2, where it's like, “Oh, coding is this verifiable task. Surely you can just come up with some coding problem for the model and then write some tests that verify if it's successful or not.” But then you dig into the details, and it's actually really hard to write a coding problem that is both naturally worded yet still perfectly unambiguous, with exactly 1 correct solution.
The SWE-bench problems, at least in the original version, don't at all fit those criteria. They just scrape GitHub issues, and all the developers listening know that GitHub issue descriptions are not meant to be well-scoped tasks that you just copy and paste and give to a model.
They often include lots of unit tests that are scoped to particular implementation details, going as far as asking the AI to output a specific 20-word error message that isn't at all mentioned in the task description. Obviously, later versions of SWE-bench tried to solve these issues, but it's still not perfect, and I think it really underscores a lot of the issues with benchmarks.
Jordan Nanos
Makes sense. So let's talk about 4.7 specifically. There were a few things that improved—or changed, let's say—in terms of features, as opposed to just the benchmark itself. I think you made the point in the article that people don't necessarily care as much about the quality of the model anymore. It's the model plus the harness—the product—that should be tested, as opposed to the model on a generic bash-only harness or something like that.
By harness, I mean Claude Code is the thing you're testing, not Opus 4.7. So it's Claude Code versus Codex; it's not Opus versus GPT.
To that end, during the release for 4.7, they announced an extra-high reasoning-effort option that slots between high and max. They announced high-resolution image support, which people can use for screenshots and styling in front-end applications. They're omitting thinking content by default, so people won't see when the model is thinking. They've got this task-budget concept that lets you tell the model how much it should think or work before it can actually run out of context in the context window. And perhaps most importantly, they updated their tokenizer.
It's potentially costing people 35% more for the exact same output from the model than from the previous model, just because they're counting tokens differently and using it. I'm just laughing at Dylan right now. Dude, I literally put him to sleep with that monologue. Holy shit. What the fuck?
Doug O'Laughlin
Oh my God. Okay, well—
Jordan Nanos
This isn't a joke. No, man, come on. Help me out here.
Doug O'Laughlin
What are you even talking about, man? The tokenizer is so boring that it puts Dylan to sleep. That's the takeaway.
Jordan Nanos
It's not like the pod was compelling and great at this point, but still—
Doug O'Laughlin
It is now. I'm compelled. Now we can talk about all the shit we want.
Jordan Nanos
Yeah.
What's the Michelangelo painting or sculpture where the guy is—
Doug O'Laughlin
The Thinker?
Jordan Nanos
Yeah, The Thinker.
Doug O'Laughlin
Yo, he's back.
Jordan Nanos
He's back.
Doug O'Laughlin
Dude, could you actually not hear us through your headphones?
Jordan Nanos
I can hear you.
Doug O'Laughlin
What do you mean? He's been asleep this entire time.
Dylan Patel
I was just thinking, man.
Doug O'Laughlin
You've been gone for a few hours.
Dylan Patel
Call me GPT-5.5 xhigh.
Doug O'Laughlin
Just thinking right now. Tokenizer: Opus 4.7, Opus 4.6. We've been dancing around the fact that GPT-5.5 is fine—not goated, definitely good enough, lots of capacity. They'll catch up on the margin. Sounds good? Great.
Speaker 0
Yeah.
Doug O'Laughlin
Anything else? Benchmarks suck. That's a really good take. Anything else?
Dylan Patel
What's the tokenizer difference?
Doug O'Laughlin
They changed the tokenizer between 4.6 and 4.7.
Jordan Nanos
The exact same output could have 35% more tokens with the new one.
Dylan Patel
Oh, they made the tokenizer—They made the vocabulary smaller.
Speaker 0
They made the vocabulary bigger for 4.7—
Doug O'Laughlin
Yeah.
Jordan Nanos
Compared to 4.6.
Dylan Patel
Oh, sorry.
Jordan Nanos
There are 35% more tokens in the tokenizer—more vocabulary.
Dylan Patel
So wouldn't that make the average output smaller? Because you can represent longer things with fewer tokens.
Speaker 0
No, more tokens.
Dylan Patel
Well, sorry. If you had only 28 tokens to make English, then you would have to use every letter. But if you wanted to do English with 500 tokens, sure, you'd have every letter, but then you'd also have tokens for “of” and “the.” Wouldn't that make the output smaller, with fewer tokens?
Speaker 0
This is a good take.
Doug O'Laughlin
Oh.
Jordan Nanos
In practice, no, but conceptually, yeah, you could train the model with full words as tokens. But I think what people have seen is that the model is currently less token-efficient with a larger vocabulary.
Dylan Patel
Oh, interesting.
Speaker 0
But that's a good point. Yeah. Man, the guy came back in with a heater here. So, yeah.
Doug O'Laughlin
Wow. He was thinking all this time. He—Oh.
Speaker 0
Okay. Yeah. Well, the whole concept of being more token-efficient is that you can solve tasks with fewer tokens, because the more granular breakup of the tokens—or the larger breakup in the token size, or just having a larger vocabulary—would mean that you would have more information represented in latent space.
The concept of—
Doug O'Laughlin
The relationship.
Speaker 0
SemiAnalysis might be a token instead of “Semi” and “Analysis,” right?
Doug O'Laughlin
Yeah. But then it creates more context, right? Versus maybe even 4 tokens of SemiAnalysis. I get it, right? There are choices.
Jordan Nanos
But it's richer information. The embedding for SemiAnalysis might be close to the embedding for Dylan in latent space, right? The embedding for Semi or the embedding for Analysis is probably not close to it. So whatever.
I think it's unclear whether this is significantly improving performance, because everybody still thinks it's a toss-up between 4.7 and 4.6 as to what's better. But then I also think it's a toss-up as to whether this is more token-efficient or worse. It seems like people are saying, if it's not a significantly better model, then why are we doing this new tokenizer thing?
Doug O'Laughlin
Yeah.
Jordan Nanos
So maybe it's just an early checkpoint. They've got to do more RL and improve 4.7, and then have a 4.8 drop that really improves things.
Doug O'Laughlin
I don't think Anthropic really does half-baked models. OpenAI clearly does. GPT-5.3-Codex was RL'd only on code, and then 5.4 was like—
Speaker 0
Is this the most half-baked of any model we've previously seen? We saw Sonnet before Opus, right? 4.5 Sonnet, then 4.5 Opus, then 4.6, and—Right? This is the first new one that's just Opus.
Speaker 3
It's because Opus 4.7 is actually Sonnet, dude. Are you not a truther?
Dylan Patel
Yeah, yeah, and then Opus 4.7 is Mitas.
Speaker 3
There you go. Yeah, there you go. This is truther stuff, bro. It's time for truth-truthers, bro. Opus 4.6 was actually Sonnet all along. Does everyone know—Do you not remember that? That was, like, a big model that smelled small.
Speaker 0
Yeah, removing the mask. Yeah, yeah.
Speaker 3
Exactly. This model smells small. Actually, I wanted to maybe pull this back to the DeepSeek portion of the article, because I don't think we talked about DeepSeek in—
Dylan Patel
The deep state?
Speaker 3
Yeah, the deep-state version of the article. We didn't talk about DeepSeek in depth, but I think we have a lot more internal takes than what we put in the article. And I guess my question is, do you think the gap between open source in China and the United States is now widening again? Because it feels like it is now. And it's because of compute constraints. I'd like to have some takes on that. Just my take is the take.
Speaker 0
Yeah.
Dylan Patel
Yes.
Speaker 0
I'll give a quick one. I think one thing that's been overlooked with DeepSeek is the fact that this is a 1-million-token-context-window model, which many of the leading open-source models that perform great on benchmarks people use for coding don't have. And then I think that the “Reasoning in Visual Space” paper that they posted and then took down as a GitHub repo makes me think they're going to release a multimodal version of this as well, or that they're in the process of developing it, and there are going to be new weights for that. Both of those things are fascinating for making China catch up on a product basis.
If we're comparing Claude Code to Codex to DeepSeek in OpenCode, or DeepSeek in some other harness, it'll be able to support all the same features in addition to being pretty good at all the other stuff. With that said, at the time of the first DeepSeek release, I used it all the time for random stuff, and I don't use this one for anything, really. So I—
Doug O'Laughlin
Isn't Kimi K2.6 better anyway?
Speaker 0
Yeah, and I don't use that either.
Doug O'Laughlin
No, I'm just saying it wasn't like DeepSeek—I mean, the DeepSeek moment was that it was so cooked and it was from the ether, right? Or it cooked so hard, rather, and it was from the ether—not cooked, right? But this round, it comes out, it's just state-of-the-art, or it's state-of-the-open-source art, if that makes sense. It's not exactly better than Kimi K2.6. I think the other stuff in it is clearly just inference optimization, right? They talked about the Ascend kernel being partially able to run inference on it, which would really, really unlock more compute for China for the first time.
Dylan Patel
Then also, if you look at the weight size, it looks very convenient. I feel like the softmax for China is effectively the size of an H200 8× pod. All the models you're looking at are essentially able to run inference within that memory-domain space, and there's nothing bigger that's served at the state of the art.
Clearly, that seems to be the cap, right? Maybe they can do that, but they won't release it on their B200 pods to the public. It just clearly feels like they're starting to hit some kind of wall. Agree or disagree? Do you think that will keep or cap Chinese progress because they can't run inference on this at all? I'd like to hear some hot takes here.
Jordan Nanos
Okay, I think—
Dylan Patel
Dude, what the hell was that, Jordan?
Doug O'Laughlin
He's getting hot and sweaty over DeepSeek.
Dylan Patel
Yeah, you're fucking deep-panting, bro.
Doug O'Laughlin
He's getting hot and sweaty over DeepSeek.
Jordan Nanos
The DeepSeek engineering release is fascinating. All the new attention variants and the compression on the KV caches are fascinating stuff. They do so well on the infrastructure stuff. I think, again, the fact that Kimi's 256K context and DeepSeek's 1 million is a significant difference for long-horizon agentic tasks.
Dylan Patel
But isn't the context from 256K to 1 million dogshit anyway, even on Opus?
Doug O'Laughlin
Yes.
Jordan Nanos
That's not my—
Doug O'Laughlin
It's garbage. I mean, okay, it's not true garbage. It's probably a step off of the state of the art, in theory, but the problem is you don't want to just be clearing your context every time. If you're doing a big task, seeing the whole context window is really nice.
Jordan Nanos
Compaction sucks.
Doug O'Laughlin
Compaction blows.
Jordan Nanos
Not being able to read a 1 million-token context on this stuff sucks. Now—
Doug O'Laughlin
It's actually better to clear. That's my hot take. I would rather just start over. I would literally be like, “Make a summary of what we've done, copy-paste that, and just start over.” Fuck the compaction.
Jordan Nanos
No, that's what DeepSeek saw in their 3.2 paper. For their benchmarks—which I think are probably bad, but they published this—if you go beyond the context window for a given task, it's better, on the tasks they were testing, to completely clear the context, not even make a summary. There is a compaction—
Doug O'Laughlin
Yeah, isn't making a summary just what compaction is under the hood? To be clear—
Jordan Nanos
So when I say “make a summary,” I mean literally, okay, you can use the entire context window versus just what I was doing last. I'm trying to remember, because I'm not going to read 1 million tokens of slop. It's like, “Okay, what was I doing here?” Read this, and then I'll Control+C a very small part. So I'm not even—
Doug O'Laughlin
There are different—
Jordan Nanos
I'm not summarizing the entire thing. I'm just doing the task.
Doug O'Laughlin
I see. Okay.
Jordan Nanos
I'm passing off tasks. Yeah.
Doug O'Laughlin
There are different ways to do compaction, but I believe compaction is different from summarization because compaction is removing the thinking traces.
Jordan Nanos
Initially, it's not actually having the model write its own summary of the full context.
Max Kan
Oh, I thought they were literally just taking your entire context and saying, “Yo, please summarize this.”
Jordan Nanos
That is an approach, and they've done multiple of them. But the—
Dylan Patel
I've got a hot take. Anthropic, with Claude Code, did well with the CLI, so they just kept making the CLI experience amazing. But the CLI experience is not the end-all, be-all of agent orchestration, and therefore they've really cooked themselves into an innovator's dilemma, where they keep making the CLI better. OpenAI has the true vision of what the true agent-orchestration platform of the future is, where you'll be able to integrate voice, multimodality, and all these other things into the app. The app is so much better than the CLI, and the real—
Jordan Nanos
Mm-hmm.
Dylan Patel
—the point is that you should develop, and users should be using it, in the app, not on the CLI, because the CLI is a dead end, a foregone relic of H1 2026 and H2 2025.
Max Kan
Do you mean the app forever, or do you mean their device? They're talking about releasing a consumer device next year too.
Dylan Patel
No, no, no. I mean the laptop—
Jordan Nanos
Like that, huh?
Dylan Patel
—you know, app. The Codex app.
Jordan Nanos
Yeah.
Dylan Patel
The Codex app.
Max Kan
I have a question then. Why does the Codex app suck?
Jordan Nanos
Yeah. True say.
Max Kan
So, look—
Dylan Patel
It's long-term planning.
Max Kan
I think—
Dylan Patel
Long-term planning.
Max Kan
Dylan, in my opinion—
Dylan Patel
Look—
Max Kan
I think you are thinking too small because, in the perfect, true maxi world—
Jordan Nanos
Yeah.
Max Kan
—the operating system doesn't need to exist. You will just get a piece of hardware. You will plug in your thing. It will pull up the terminal, and you will connect your Claude API, and it will build the OS for you. Thinking—
Dylan Patel
So, Codex app, they're adding generative UI stuff too, which is pretty interesting.
Max Kan
I'm just saying, I think if you're a coding purist, generative UI is downstream of the CLI. I think I'm a CLI purist. Dude, I don't know. This is just a slop preference thing. I just love the CLI, man. Claude Code usage is clearly just a CLI wrapper, and you can tell, and then Codex CLI is clearly just an app wrapper. I feel like they forced it over. I think there are 2 opinions about the future. Who knows who will win out in the very long run? I'm definitely going to keep it open for competition.
But at this beautiful moment, a true maxi's vision and dream is that it's all downstream from the CLI. It's just tokens. It's the most efficient version of everything, man. Gotta Elon-max, okay? All you need is just an API and then inputs. That's it. Your app and all that stuff, that's all obfuscation. Mythos would know better than OpenAI. Who are we little brains to know what UI we want? No, dude, it's CLI all the way down. Pure maxi vision.
Jordan Nanos
I'm a VS Code plugin guy. I literally tested this yesterday. I was bothering Max, who told me to go away, about using Ghostty for the CLI stuff. It just doesn't work.
Max Kan
Wait, no, I still think having 6 Ghostty terminals open with Claude Code CLI is a superior experience to having 6 different chats going in the Codex app.
Jordan Nanos
Oh, for sure.
Max Kan
I feel more productive.
Jordan Nanos
Yeah, I'm comparing it to the VS Code plugin: 6 different windows in VS Code, also with a file browser on the left side so I can right-click and copy—
Max Kan
Well, I think the difference is that you're still writing real code and you kind of care about the output, whereas I just don't even look at it. I just directly push the—
Jordan Nanos
No, I'm not looking at the code. I'm copying in images or Excel files. Don't accuse me of reading the code. Sorry.
Dylan Patel
He's fucking Edison.
Max Kan
My favorite is—I mean, I actually agree that the API, CLI, whatever, is the new compiler. No one's reading the compiler. No one cares. They don't need to touch the magic.
Jordan Nanos
This will be true in the future, but it is currently—for any code you actually care about—this is currently not true. It's still producing a bunch of stuff that is bad and should be fixed by coaxing the model to fix it. I'm not saying I type code anymore, but I do read some code.
Dylan Patel
One of my group chats—
Max Kan
Yeah.
Dylan Patel
—one of my group chats, I was reading it this morning, and it's a group chat with all the most cracked kernel programmers in the world. It turns out what they do—
Max Kan
Wait, why are you in it?
Dylan Patel
Because they're my boys, bro. Look, Max, come on. I'm a master networker, okay?
So, anyway, I think it was Tri Dao. TreeDAO's like, “Yeah, dude, Codex is so dumb, but I always just have it create it, and it works and it's smarter, but the code is slop, and then I have Opus rewrite it. But you can't go the other way around. You can't have Opus write the thing and then have Codex fix it. You have to have—
Max Kan
You know what's funny?
Jordan Nanos
Really? I go the other way around.
Max Kan
Codex write it and then have Opus fix it. Everyone else at the firm prefers the other way around.
Jordan Nanos
Yeah.
Max Kan
The entire firm's preference is the other way around, actually.
Dylan Patel
Yeah, but we're not writing fucking kernels, right? We're not writing fucking Tri Dao kernels.
Max Kan
That's probably fair.
Jordan Nanos
We're doing benchmarks of kernels, but yeah. We wrote some kernels.
Dylan Patel
Oh, come on. Dude, they're not Tri Dao kernels.
Jordan Nanos
No, they're not Tri Dao kernels. They're just GPU MODE kernel competition kernels.
Dylan Patel
Because apparently, if you talk about niche and microarchitecture details, Claude will waffle on about shit instead of actually just doing it. Whereas if you describe it to Codex, it'll just try and implement it all, and then it'll be slop. But then you tell Opus to fix it, Opus won't waffle on; it'll just fix it.
Jordan Nanos
Yeah. Doug, this is the context-window stuff, which is like, when you are pumping in so many docs about the ISA of a given GPU to write a kernel or something, you need performance at a million context. You just run out of space on the smaller stuff. So I don't know. Do you want to go back to DeepSeek and any hot takes on DeepSeek, Dylan? Why didn't it crash the market this time if KV cache is reduced by 90%?
Dylan Patel
Dude, you know, it's been a while since I've been in Asia, but every time I go to Asia, they reference some fucking new paper that reduces KV cache every fucking time for the last 3 years. Some paper, they're reducing KV cache, and no researcher in America has even heard of this paper. It's the fucking best thing ever. DeepSeek and TurboQuant were the most precipitous ones that popped up the most, and TurboQuant was obviously fake news. But yeah, I think it's very funny. I don't know.
I guess they're tired of being robbed.
Max Kan
Okay. Well, look, yeah, I think that's fair. It just doesn't matter. Gemini's working, clearly with the price of the GPU going up. That's all you need to know. Now, if we're gonna talk about real fake news, let's talk about SubQ. Let's do some—I mean, we're not gonna write an article about it. We're not gonna write a post about it. This is free alpha. Did anyone else read the SubQ thing today? It's pretty sus. It's actually extremely ultra-mega sus.
Jordan Nanos
Yeah. It seems like people are launching their startup, right?
Max Kan
Honestly, they should close funding, and then they'd be like, “Wow, it was just Opus with 10 context windows stapled together.” I mean—
Jordan Nanos
Do you think the market is hot enough for them to close, you know, a $200 million at $1 billion round next month or something? If they did, I would be impressed.
Dylan Patel
I don't know if they could do $200 million, but I think they could do $50 million at a bill. A tril—
Max Kan
A tril? What?
Dylan Patel
Sorry, bill. No, tril.
Jordan Nanos
There's more capital than there is opportunity.
Max Kan
Wait, wait, wait. Are you familiar with what we're even talking about, Dylan?
Dylan Patel
No, sorry, I just thought you guys were talking about Anthropic.
Max Kan
No, we're not talking about Anthropic.
Dylan Patel
Oh.
Max Kan
He's talking about model sparsity. I'm talking about the worst—did you not? It's like this fake-news Twitter thing today called SubQ.
Dylan Patel
Oh.
Max Kan
Um—
Dylan Patel
Yeah, yeah, yeah, yeah.
Max Kan
Yeah, yeah.
Dylan Patel
That's another fake-news one.
Max Kan
Don't worry. We requested API access. We made sure to use our SemiAnalysis email to improve our odds.
Jordan Nanos
Yeah. We're like, “Please, give us this API for this very real model, bro.”
Max Kan
Who knows? Maybe it's a state-space model. Maybe Mamba cooks, or—
Dylan Patel
No, it's not an SSM. I don't think it's an SSM.
Max Kan
I think it's—I don't know. It's just really funny because, again, we're talking about DeepSeek people freaking out. Dude, if this was real, memory stocks should be down like whatever, a quadrillion percent today. But obviously it's not real because if you look at these guys and you're like, yeah, man, I just don't think these guys are gonna be the guys to crack the single hardest problem in all of AI. No offense—maybe the founder's super legit.
Jordan Nanos
So, okay, maybe one thing this reminded me of was the fact that Llama 4 Scout or Maverick—I think Scout, the smallest one—was released with a 10-million-token context window, or announced with it, but not supporting it officially in the released weights or something. And I'm just really surprised that we haven't seen anybody with effectively an unlimited compute budget give it a go for a more expensive model with a larger context window. Like—
Dylan Patel
But what—where are you gonna get the data, right? Most people pretrain with 16K context or 4K, you know, something like that, 32K context, and then they post-train it so that they can add and hack in the rest of the context. But it's like, what data do I have? That's why my 250K to 1M context is trash anyways, is because there's no data on this stuff. And so the model doesn't generalize the context really well. And then if you stick it to 10 million, it's like, what fucking data do I have that is useful for the next-token generation that exists from 1 million context to 10 million context? There's so little.
Jordan Nanos
Yeah. I mean, it makes sense. Possibly synthetic stuff, possibly—I mean, why'd they do it in the first place? It seems obvious that people would be working on it, and we haven't even seen anybody announce 2 million. So there's some arbitrary limit—
Dylan Patel
Wait, but Google serves 2 million.
Jordan Nanos
Google serves 2 million on Gemini 3.1 Pro?
Dylan Patel
They did on Gemini 2.5—2 million.
Jordan Nanos
Well, maybe that's—
Max Kan
It's 1 million.
Jordan Nanos
That's the answer to me, yeah.
Max Kan
It's 1 million today.
Dylan Patel
It's 1 million today?
Max Kan
On Gemini 3.1 Pro.
Dylan Patel
One of their announcements—
Jordan Nanos
No.
Dylan Patel
They announced 10 million. They started at 1, and then they updated it to 2 at some point in one of the models.
Jordan Nanos
Makes sense, yeah. I mean, they got a big scale-up domain. Why not give it a go with the TPUs? Yeah, maybe another thing that was a little bit missed in the article, and you kind of talked about it when you brought up DeepSeek. Max, I want your take on this. You—because Doug asked the bait question about whether China is catching up or they're still behind. It kind of depends on how you look at it. But I was bugging you the other day: Is DeepSeek or Kimi currently ahead or behind Meta? And are they ahead or behind Grok, Cursor, or SpaceX's xAI?
Max Kan
I would say that today they're probably ahead of all those companies. But the thing that really matters is slope from here. This is a pretty basic take at this point, but I do think the amount of compute you have is actually just one of the key inputs to how good your model's gonna be. And obviously Meta's signing all these monster deals. It seems like they've overcome the overhang of having to fire and then rehire their entire AI team, and they're in the process of making some good models now. So I would expect Meta to pull away from all the Chinese guys, if not in the second half of this year, then in the first half of 2027.
Dylan Patel
And Meta's not distilling.
Jordan Nanos
Oh, I thought they were.
Max Kan
I thought they were all distilling.
Jordan Nanos
I thought they were distilling from the Chinese guys. They were just running the open-source models.
Dylan Patel
I mean, that's what Mistral does. They don't distill from Anthropic. They distill from the Chinese guys. That's fair.
Jordan Nanos
Yeah.
Max Kan
Why are we talking about the leading French frontier model company, Mistral?
Dylan Patel
Dude, you know their revenue's really strong.
Max Kan
Yeah, I do, actually. You know what they've bro'd down on? Also, dude, the bottles.
Jordan Nanos
Because they're a new product.
Dylan Patel
You keep picking up the bottle on the mic.
Jordan Nanos
Every frontier model—
Dylan Patel
Well, no, not neocloud. They keep trying to—
Jordan Nanos
…is in a cloud—neocloud.
Dylan Patel
No, yeah, they're trying to become a neocloud, or at least they're doing fine-tunes.
Jordan Nanos
Every chip company. Cerebras is becoming a neocloud. NVIDIA's launching neoclouds.
Doug O'Laughlin
The ultimate business model—
Jordan Nanos
AMD—
Doug O'Laughlin
…for any company in the world is to become a neocloud.
Jordan Nanos
It's starting—
Doug O'Laughlin
SemiAnalysis will become a neocloud. And then we will be ClusterMAX Platinum.
Dylan Patel
Diamond. No, dude, we gotta introduce a new tier. Yeah, diamond. Tungsten.
Jordan Nanos
Yeah. SemiAnalysis, lithium.
Dylan Patel
I don't know, germanium? I don't know, I'm just making up shit. What's the rarest?
Jordan Nanos
What's your favorite semiconductor, Doug? We should make the tiers semiconducting materials only.
Doug O'Laughlin
Oh, yeah? You like semiconductors?
Jordan Nanos
They already are, but—
Doug O'Laughlin
Name all of them.
Dylan Patel
Name them all, yeah. We should be rhodium—the rarest and most expensive precious metal.
Jordan Nanos
Vanadium?
Doug O'Laughlin
All right, boys, this is getting off track. We’ve got to get out of here.
Dylan Patel
Yeah, okay. Any other parting takes?
Jordan Nanos
No, I think the hot takes have run out. Claude Code was the inflection point in February 2026. Doug, your victory lap today in May is complete. I appreciate all the hot takes today.
Dylan Patel
I hope it’s not the next inflection point. I hope it’s more exciting than that.
RLHF is very appreciated. Yeah, better than Dylan.
Speaker 0
I made it to the end.
Speaker 3
Yeah. He didn’t even make it to the end.
Speaker 0
Yeah.