[BidClub_]
SemiAnalysis · · 45 min

Ep. 011 - GPT 5.5 vs Claude 4.7: OpenAI's Comeback From the Brink (Tokenomics) | Jordan Nanos, Dylan Patel, Doug O'Laughlin, Max Kan

Jordan NanosDylan PatelDoug O'LaughlinMax Kan

Podcast
TL;DR
  • OpenAI is back “in the conversation” with GPT-5.5 after a stretch Max Kan calls “really dire.” Anthropic—riding Opus 4.5’s step change in coding and agentic ability—passed OpenAI on a like-for-like revenue basis in early-to-mid April (a leaked ~$19B ARR versus ~$24B at the start of the year), and GPT-5.4 “was honestly just an embarrassment” whose model card didn’t even compare against Opus. 5.5 is back on the frontier but is not “definitively better” than Opus 4.6/4.7 “despite what the Twitter propaganda machine was trying to push.”
  • Dylan Patel’s release calendar: “Everyone’s releasing in two weeks”—Google and OpenAI definitely. The “Spud” (5.5) shipped with pre-training unfinished, so the next drop is finished pre-training plus more RL; Google’s is expected to be “mostly just” a multimodal swap.
  • Fast-mode premiums are decaying while the price stays 6x: Opus 4.6 fast has slipped from 2.5x to under 2x speedup (90→70 tok/s against an unchanged 35–40 base). Yet this is the first time SemiAnalysis engineers chose fast over higher-quality tokens—Jordan Nanos’s read is that “4.7 is just not meaningfully better quality than 4.6 for people today.”
  • Token pricing is starting to price out even heavy professional users. Dylan says the desk is “on the cusp” of being priced out; Mythos was quoted ambiguously at $25/$150 or $25/$125 versus Opus’s $5/$25, while Doug characterized it as roughly 5x, with fast mode 6x on top. Doug’s $800 boil-the-ocean scraping run versus a $55–100 data-enrichment API is the cautionary tale—and cost growth comes from new tasks (Jevons), not repricing old ones.
  • Doug’s structural bear case for frontier pricing: Opus 4.5 may have crossed the threshold where day-to-day tasks are one-shottable without supervision. GPT-5.5-level intelligence in a 100–200B-parameter form factor in “probably less than a year” might mean the majority of people do not need frontier-level intelligence.
  • Benchmarks have degraded to “a vibe check to make sure that the model’s not total trash.” Humanity’s Last Exam is esoteric multiple choice; SWE-bench scrapes GitHub issues with implementation-scoped unit tests. Meanwhile, 4.7’s new tokenizer can cost 35% more for identical output, and the desk’s truther take—“Opus 4.7 is actually Sonnet”—captures how small the model smells.
  • The China open-source gap is widening again, per Dylan’s flat “Yes”—and it’s a compute-constraint story. Chinese frontier weights conveniently fit an 8x H200 pod’s memory domain, Ascend kernels only partially serve DeepSeek V4, and Max expects Meta—behind today but signing monster compute deals—to “pull away from all the Chinese guys” by H2’26 or H1’27.
  • The form-factor fight is live: Dylan calls the CLI “a foregone relic” and says OpenAI’s app holds the true agent-orchestration vision; Max’s rebuttal: “it’s CLI all the way down. Pure maxi vision.” Tri Dao’s cracked-kernel workflow, via Dylan’s group chat: have Codex write it, then Opus fix the slop—“you can’t go the other way around”—though “everyone else at the firm prefers the other way around.”
Digest · the substance, structured for research

1. GPT-5.5 pulls OpenAI back from the brink — Anthropic had passed them on revenue

  • Max’s TLDR of his debut SemiAnalysis article: The Information leaked ~$19B ARR for Anthropic versus ~$24B for OpenAI at the start of the year, and even setting aside net-versus-gross hyperscaler accounting, Anthropic “pretty clearly surpassed them on a like-for-like basis in early to mid-April.” The driver: Opus 4.5 was “a real step change” in coding and agentic ability—from late November through early April, “everyone was basically just spamming Opus 4.5, 4.6 for all their workloads.”
  • The GPT-5.4 verdict, unsoftened: “honestly just an embarrassment”—its release card compared only against past OpenAI models, not Opus. “That kind of tells you all you need to know.” With 5.5, Opus is back in the card and OpenAI is “back on the frontier”—not definitively better than 4.6/4.7 “despite what the Twitter propaganda machine was trying to push on release date,” but “definitely in the conversation.”
  • Dylan’s forward calendar: “New release in two weeks... Google, OpenAI, maybe Anthropic.” The Spud “kind of like didn’t finish the pre-training and just released it”—so expect finished pre-training, more RL, then the drop; Google’s will be “mostly just gonna be like multimodal swap.”

2. Fast mode: the speed premium is decaying, and buying speed over quality is new behavior

  • Doug can’t stay on Codex—“the usage limits raw dog me every day, bro”—and the desk agrees OpenAI’s fast mode is “pretty fake”: reduced reasoning depth, but not noticeably faster. Jordan’s taxonomy: priority mode is a roughly 2x premium for an SLA, not faster interactivity, and 5.3 Codex Spark is “definitely fake. That’s just a different model.”
  • Jordan’s data on Opus: 4.6 fast launched around 90 tok/s per user versus 35–40 base; it now runs around 70 against an unchanged base—“not even two times faster for six times the price.”
  • The behavioral tell: this is the first time any SemiAnalysis engineer traded fast over higher-quality tokens. Jordan’s explanation: “4.7 is just not meaningfully better quality than 4.6 for people today.”

3. Tokenomics: getting priced out, and whether frontier intelligence is even needed

  • Doug’s structural bear case, framed via the OpenAI-Cerebras deal: if Opus 4.5 passed a key capability threshold such that many day-to-day tasks are one-shottable without supervision, then even if Cerebras can never run anything larger than roughly 100–200B parameters, it might have GPT-5.5-level intelligence “in probably less than a year.” The majority of people might no longer need frontier-level intelligence.
  • Dylan says SemiAnalysis is “on the cusp” of being priced out of Mythos fast. Dylan floated Mythos at $25/$150 or $25/$125 versus Opus’s $5/$25; Doug called it roughly 5x, with fast mode 6x on top. Doug says current spend is justifiable, “but if you were to even double it... maybe we gotta turn off fast mode, guys”—“margins matter.”
  • Doug’s key distinction on cost: “I don’t expect new models to be more expensive to do the same task... the problem of cost is you’re gonna do new tasks.” He also suggests Mythos fast may be cheaper than 4.6 fast for many tasks if it is more token-efficient. Max: “They call that Jevons, dude”—and the sad future is having to ask, “Is this task really worth Mythos fast token pricing?”
  • The specimens: Max burned ~$400 of Opus 4.6 fast setting up a benchmark on a DigitalOcean droplet; Doug burned ~$800 in tokens scraping 10,000 employee profiles that a data-enrichment API would deliver for $55–100—“I just tried to boil the ocean using AI.” The “Rick and Morty” coda: “What’s your purpose?” “Pass me the butter.”

4. 4.7 versus 4.6 is a wash—and benchmarks can’t adjudicate it

  • Doug’s vibes: instruction-following is “objectively worse,” missing CLAUDE.md instructions and skills, and the model keeps signing off—“We’ve done a lot for today, go enjoy your weekend.” He was using it on a Monday and said, “Get back to work.” His theory: inference optimization plus 10–100x more users degrades the experience—“the 4.6 golden age when it wasn’t quantized... pre-nerf. Those were the days.” His fix: “I need a NIMBY frontier model.”
  • Doug on why benchmarks can’t settle it: frontier proximity is necessary, but “being number one on the benchmark ranking does not necessarily imply you’re the best model”—they’re now “a vibe check to make sure that the model’s not total trash.” HLE is “the most esoteric multiple choice questions you’ve ever seen”; SWE-bench scrapes GitHub issues that aren’t well-scoped tasks, with unit tests demanding specific 20-word error messages never mentioned in the prompt.
  • The 4.7 changes Jordan catalogs: extra-high reasoning between high and max, high-resolution image support, thinking hidden by default, task budgets, and a tokenizer with 35% more vocabulary—potentially 35% higher cost for identical output. Dylan, waking from an on-air nap, counters that a bigger vocabulary should compress output; Jordan concedes conceptually but says in practice the model is less token-efficient—and if 4.7 isn’t clearly better, why the new tokenizer at all?
  • Doug: “I don’t think Anthropic really does half-baked models. OpenAI clearly does”—5.3 Codex was RL’d only on code. Truther corner: “Opus 4.7 is actually Sonnet, dude,” with Dylan adding “and Opus 4.7 is Mitas”—“this model smells small.”

5. DeepSeek V4 and China’s compute wall

  • The show’s bait question is whether the open-source gap between China and the US is widening again because of compute constraints. Dylan’s complete answer: “Yes.”
  • The transcript highlights V4’s 1M-token context, versus Kimi’s 256K—a potential edge for long-horizon agentic work—and the “Reasoning in Visual Space” repo, posted then pulled, as a possible sign of a multimodal version. The same speaker admits that, unlike the original DeepSeek, which they used “all the time for random stuff,” they do not really use this one.
  • Doug says there was no ether moment this time: “It comes out, it’s just state-of-the-open-source art,” not clearly better than Kimi K2.6. The real news is inference optimization: Ascend kernels being partially able to run inference on it “would really unlock more compute for China for the first time,” and the weights conveniently fit an 8x H200 pod’s memory domain—nothing bigger is served at the state of the art. “It just clearly feels like they are starting to hit some kind of wall.”
  • Jordan calls the DeepSeek engineering release fascinating for its attention variants and KV-cache compression, while Doug challenges the practical value of the longer context: even Opus’s 256K-to-1M context is “dogshit” in quality, and compaction is painful.
  • Max on slope: DeepSeek/Kimi are “probably ahead” of Meta, Grok, Cursor and xAI today, but compute is a key input; Meta is signing monster deals and has overcome the fire-then-rehire overhang. He expects Meta “to pull away from all the Chinese guys” in H2’26 or H1’27. Distillation aside, Dylan says Mistral distills from the Chinese labs, not Anthropic.

6. CLI versus app: innovator’s dilemma or pure maxi vision

  • Dylan’s hot take: Anthropic perfected the CLI into an innovator’s dilemma—“the CLI is a dead end, a foregone relic of H1’26 and H2’25”—while OpenAI’s app holds “the true vision of the agent orchestration platform,” including voice and multimodality. Max separately notes that the Codex app is adding generative UI.
  • Max’s maximalist rebuttal: the operating system doesn’t need to exist—hardware, a terminal, a Claude API, and it builds the OS for you. “It’s CLI all the way down. Pure maxi vision.” He says Claude Code is “clearly just a CLI wrapper,” while Codex CLI is “clearly just an app wrapper.” Jordan, meanwhile, is a VS Code-plugin guy—“don’t accuse me of reading the code.”
  • Dylan’s Tri Dao anecdote from his cracked-kernel group chat: Codex is dumb but implements; Claude will waffle on niche microarchitecture details. So “have Codex write it and then have Opus fix it. You can’t go the other way around.” Max: “Everyone else at the firm prefers the other way around, actually.” Dylan: “we’re not writing fucking Tri Dao kernels.”

7. Long context, compaction, and the fake-news frontier

  • On compaction, Doug is categorical: “Compaction blows... fuck the compaction”—better to clear and restart. Jordan notes DeepSeek’s 3.2 paper found exactly that: past the context window, fully clearing beat even summarizing on their tested tasks.
  • Why hasn’t anyone shipped Llama 4 Scout’s announced 10M context? Dylan: “What fucking data do I have that is useful for next-token generation from 1 million to 10 million context? It’s so little”—the same data drought that makes 250K–1M “trash anyways, even on Opus.”
  • SubQ, the day’s viral startup, gets the treatment from Max: “extremely ultra-mega sus.” If a real KV-cache breakthrough existed, “memory stocks should be down a quadrillion percent today,” and “I just don’t think these guys are gonna be the guys to crack the single hardest problem in all of AI.” Dylan’s pattern-match: every Asia trip surfaces a KV-cache-reduction paper no American researcher has heard of; TurboQuant was fake news.
  • Still, Jordan’s market read stands: “There’s more capital than there is opportunity.” Dylan thinks SubQ could raise $50M at a $1B valuation.

8. Closing take

  • Jordan calls Claude Code “the inflection point in February 2026” and says Doug’s May victory lap is complete. Dylan hopes the next inflection point is more exciting.
Jordan Nanos

Hello, everyone. Welcome back to SemiAnalysis Weekly. We have a big show this week. We have some heavy hitters: Doug, Max—and anybody else out there? Oh, yeah, of course, that's Dylan Patel calling in, nicely in a nice setup office with his headphones on and a SemiAnalysis-logo water bottle. We're going to talk about Claude Code, which is one of the only things we talk about on this podcast now. We're going to talk about GPT-5.5 possibly, DeepSeek, and, yeah, just get into it.

Dylan Patel

It's crazy. Today's May 5th.

Jordan Nanos

Yeah.

Dylan Patel

GPT-5.5.

Jordan Nanos

Mm-hmm.

Dylan Patel

It's Max's birthday.

Max Kan

Pew, pew, pew, pew.

Jordan Nanos

Happy birthday, Max. Max debuted in the SemiAnalysis newsletter last week with his first lead-author article, the coding-assistant breakdown.

Dylan Patel

He's asking Doug and Michelle, “Hey, am I going to be approved on my work trial?” It's like, bro, obviously. What are you talking about? Also, Max is now a full-time employee at SemiAnalysis, not on a work trial.

Max Kan

Happy to be here. It's really exciting.

Jordan Nanos

Welcome to the team. Okay, so let's run through this article. Basically, when we pushed it out, or when we were preparing to write it, it was effectively just going to be a review of GPT-5.5 Pro—or just 5.5, whatever the new model was. I think we assumed it was going to be called 5.5 when we were testing different checkpoints.

Dylan Patel

Potato, potato. It's a spud, right?

Jordan Nanos

Yeah.

Dylan Patel

Potato, potato.

Jordan Nanos

Yeah, we're calling it Spud.

Dylan Patel

Or is it tomato, tomato?

Jordan Nanos

We tried some different checkpoints. But the basis for the article was just going to be a review of that model, and then Anthropic quickly released Claude Opus 4.7, which was a big change, with some pretty interesting features that we covered in the article. We also got DeepSeek V4, the 4-, 5-, possibly 6-month-delayed release from DeepSeek. So it became this article about all the latest releases. Max, maybe you could give us a quick summary. What was your high-level takeaway? Then we can dig into the details of the changes for these models.

Max Kan

Yeah, happy to. The TL;DR is that things were looking really dire for OpenAI for a while. At the start of the year, Anthropic was quickly encroaching in terms of revenue. The Information leaked $19 billion in ARR for Anthropic versus around $24 billion for OpenAI. Anthropic then quickly surpassed them, even with all the accounting discrepancies related to whether you recognize net or gross revenue from the hyperscalers. I think it pretty clearly surpassed them on a like-for-like basis in early to mid-April.

This was primarily because Opus 4.5 was a real step change in coding and overall agentic abilities. From late November through the end of March and early April, everyone was basically spamming Opus 4.5 and 4.6 for all their workloads. OpenAI tried to fire back with GPT-5.4, but that thing was honestly just an embarrassment. In the model release card, they didn't even compare it to the Opus models—just to past OpenAI models. That tells you all you need to know.

But then, finally, with 5.5, they're back on the frontier. Opus was re-included in the model release card. I wouldn't say it's definitively better than 4.6 or 4.7, despite what the Twitter propaganda machine was trying to push on release day, but I think it's definitely in the conversation. I'm happy to use it, and it allowed OpenAI to come back into the game. Things were looking dire, but now I think they have a shot again.

Jordan Nanos

Doug, what are you using as your daily driver? Are you using both or just one right now?

Doug O'Laughlin

I've tried so many times to use Codex, but the usage limits raw-dog me every day, bro.

Dylan Patel

Wait, what do you mean?

Doug O'Laughlin

I just can't do it.

Dylan Patel

Use the API, bro.

Doug O'Laughlin

I use the API. Don't get me started. We already had this—

Dylan Patel

How do you get usage-limited?

Doug O'Laughlin

Trust me. I don't know, bro.

Dylan Patel

Dude, just go bitch in the OpenAI Slack.

Doug O'Laughlin

I have. I literally—dude. But if it takes me longer than 15 minutes, I don't try to actually fix it.

Jordan Nanos

We're trying to get you started here, Doug, for what it's worth. We're trying to get you started.

Doug O'Laughlin

If it takes me longer than 15 minutes, I essentially don't try to actually fix it. xhigh seems really expensive and blows out the context window. But, in my impression, it's very neck and neck. They're pretty replaceable to me. The thing that I really want is for someone to give me fast mode and high uptime, and no one has high uptime yet.

Dylan Patel

Wasn't OpenAI's fast mode fake news?

Doug O'Laughlin

Yeah, it's pretty fake.

Jordan Nanos

Well, they reduce the reasoning depth, and it's not that fast.

Opus fast mode isn't that fast anymore, either. It's not even 2 times faster.

Doug O'Laughlin

Sounds like we need more compute.

Jordan Nanos

It started out at 2.5 times faster, and now it's less than 2 times faster. But OpenAI has 3 versions of fast mode: priority mode, fast mode, and GPT-5.3 Codex Spark. GPT-5.3 Codex Spark is definitely fake. That's just a different model. But I feel like fast mode is comparable to Opus fast mode. You guys don't think it's actually faster, or much faster?

Doug O'Laughlin

I've been using it in Codex, and it feels pretty slow. I don't even notice it really being faster than when fast mode is turned off.

Dylan Patel

I thought that's what you were showing me, Jordan, the other day: the distribution of what's the peak tokens per second, what's the median, and what's the trough. Or maybe it was someone else, but it seemed like OpenAI's priority mode doesn't actually make it always faster. It can be faster, though.

Jordan Nanos

Yeah. Priority mode is guaranteed execution with an SLA, but it doesn't provide faster interactivity for people who need things done on the API and will pay a premium for that. It's a small premium—around 2 times. I think what I was showing you was the data we have on 4.6 Fast versus 4.6 Base.

Doug O'Laughlin

Oh.

Jordan Nanos

It started out at, let's say, around 90 tokens per second per user on Fast Mode and around 35 or 40 on Base. It's still 35 or 40 on Base, but it's now around 70 tokens per second on Fast Mode consistently. So it's not even 2 times faster for 6 times the price.

The guys internally are—Max made this point in the article, or I guess we made this point in the article—this is the first time any of the engineers at SemiAnalysis have made the trade-off of wanting fast over higher-quality tokens, and I'm not sure what drives that. People definitely want Fast Mode. It definitely feels faster, but maybe the big thing is that 4.7 is just not meaningfully better quality than 4.6 for people today.

Doug O'Laughlin

I was actually thinking about this earlier today. If you treat it in the context of the OpenAI-Cerebras deal, and you assume that Opus 4.5 just passed some key threshold in intelligence or model capability such that a lot of your day-to-day tasks are now one-shottable by the models, you don't really have to supervise them at all. You're not even looking at the outputs.

Maybe it's just the case that, even if you can never run anything larger than, let's say, 1 or 200 billion parameters on Cerebras, you're going to have GPT-5.5-level intelligence in that form factor in probably less than a year. It might just be that we've passed the inflection point where the majority of people don't need frontier-level intelligence for their day-to-day workload.

This is probably doubly true if the models keep getting more expensive. Originally, I thought there might be a chance that SemiAnalysis would be able to afford Mythos Fast. I think that's probably not true anymore. A lot of people have already been priced out of the models.

Jordan Nanos

Dylan, can we, though?

Dylan Patel

I think we're on the cusp of getting priced out.

Jordan Nanos

Was that a definitive no?

Doug O'Laughlin

I don't think we can, right? Mythos is 6 or 7 times more expensive.

Dylan Patel

It's $25/$150, I think. $25/$150 or $25/$125, one of the two.

Doug O'Laughlin

Versus $5/$15, right?

Dylan Patel

It's $5/$25 for Opus.

Doug O'Laughlin

$5/$25.

Dylan Patel

It's 5 times.

Doug O'Laughlin

So it's 5 times more expensive.

Dylan Patel

Yeah.

Doug O'Laughlin

And then Fast Mode is 6 times on top of that. If we were to take our current token spend, I can justify that, but if you were to even double it, I'd be like, “Oh, fuck. Maybe we have to turn off Fast Mode, guys.” At this point, it is—

Jordan Nanos

Yeah.

Doug O'Laughlin

Margins matter, you know.

Jordan Nanos

Do you think there's a possibility that we could define something so important in SemiAnalysis research in the future that we would want to pay the premium just for one project or one task?

Doug O'Laughlin

I think the flip side is that Mythos is more token-efficient.

Mythos fast mode is probably cheaper than 4.6 fast mode for most tasks. I just imagine. At least that's the case with Codex versus 5.4 versus 5. Even if they make the model more expensive, that's not a huge jump in price for the model. I don't expect new models to be more expensive to do the same task. The problem with cost is that you're going to do new tasks.

Jordan Nanos

Yeah.

Max Kan

They call that Jevons, dude.

Jordan Nanos

Yeah. We're going to be coming up with new stuff to do with these models based on the work that we do in the next few months with the models. We'll come up with new things, new tasks that are harder and more complex and need to use the bigger models for them. Right?

Max Kan

Yeah. But I feel like right now, at least, I don't even think about whether a task is worthy of spending tokens on. I just spend the tokens. But if the models get much more expensive, you might have to think carefully: Is this task really worth Mythos fast token pricing? I think that'd be really sad, honestly, when it happens.

Doug O'Laughlin

Question: What is the trade-off? I have a personal anecdote where it's more expensive to burn tokens than to do it. What are your examples of when the cost isn't worth it?

Max Kan

There are times when I was setting up some benchmark on a DigitalOcean droplet, and I was using Opus 4.6 fast, and that was $400 or something. I was like, “I don't know if it's worth $400.”

Jordan Nanos

That wasn't worth 10 minutes of your time?

Max Kan

Yeah.

Doug O'Laughlin

Jordan, what's yours?

Jordan Nanos

I think stuff where you can very clearly do it from a script or by writing. Writing docs, for example: maybe the first pass with the model is good, but editing stuff is sometimes just annoying to use the model for instead of doing it yourself because it might screw it up or edit the wrong thing.

Max Kan

It's a quality-versus-cost conversation. What's this? It's like, “Wow, this is just a complete waste of tokens.” The tokens in were 2 times more expensive.

Jordan Nanos

I don't think I've had that experience yet, to be honest.

Max Kan

Okay, so I'm going to give mine.

Jordan Nanos

Yeah. What's yours? Yeah.

Doug O'Laughlin

Okay. Scraping large data sets: at some point, there's a diminishing return. You're like, “Hey, give me 10,000 employees from this company.” And you're like, “Great, I'm going to have it hit the search API.” I'll do a lot of variation to figure out a profile about every person who's worked at this company. Then you actually run the cost, and there's a data enrichment API—and the data enrichment API is literally 1/10 the cost. You're like, “Oh, shit.”

I think I burned $800 in tokens to do what would take maybe $55 to $100 in API calls. One is probably slop, but the other one is, in theory, verified by another slop cannon. I just think there's some value, and that's probably one of the most interesting places where the replacement cost of the tokens versus the actual information—it's still a lot cheaper to essentially serve data via an API. But that cost, the pressure on the top will move that down, if it makes sense.

Jordan Nanos

Yeah.

Doug O'Laughlin

Just scraping the entire internet is not token-efficient at all. That's a good example where I tried to boil the ocean using AI, and I've been really curious about the trade-off that's going to happen where it's just not worth this much intelligence. Making me coffee—it's the “Rick and Morty” meme. It's like, “What's your purpose?” “Pass me the butter.” This is a waste, man. We've got to find better token efficiency.

Jordan Nanos

Yeah. You can actually hire some people for cheaper to do some menial tasks than you can with tokens. But some of the analysis that we've done recently kind of goes the opposite way so often that you just get used to tokens being the cheaper approach or the faster approach in so many cases that you don't even consider the alternative.

Doug O'Laughlin

In aggregate, it is, for sure. No way.

Jordan Nanos

So what's the takeaway, Dylan? Is OpenAI so back at this point? Do you think they're going to—

Dylan Patel

I don't know, man.

Jordan Nanos

—take off on a rocket ship?

Dylan Patel

New release in 2 weeks. New release in 2 weeks. Everyone's releasing in 2 weeks: Google, OpenAI, maybe Anthropic. I don't know about Anthropic, but Google and OpenAI are definitely releasing in 2 weeks.

Jordan Nanos

What are they releasing?

Dylan Patel

More everything. More continued pretraining because the Spud—they kind of didn't finish the pretraining and just released it. So, finish the pretraining, do more RL, drop the model. Google is mostly just going to do a multimodal swap.

I guess my question is: the narrative is that, even for the normies who don't use fast mode, 4.7 is worse than 4.6. Do you guys agree with this or not?

Doug O'Laughlin

I think the instruction following has gotten objectively worse. It keeps missing CLAUDE.md instructions, or you pull a skill and you're like, “Dude, you didn't do exactly what was laid out in the skill.” That seems to be a consistent problem. I don't know; I feel like it's a compute problem more than anything else.

It still has the, “We've done a lot for today. Go enjoy your weekend.” I was like, “It's fucking Monday. Get back to work.” It annoys me so much that it tries to enforce me to stop working, and I'm like, “Ugh, clearly a usage issue.”

I think the 4.6 golden age, when it wasn't quantized in the beginning—those were the days, okay? Fast mode 4.6, pre-nerf. Ugh. But I just think there's this maturity of the models as they become more inference-optimized and more people use them, that it becomes a worse experience as you 10X or 100X the users. And I think that's happened.

I just think 4.7's fine. I think 4.6 and 4.7 are probably the same. It's kind of the same level of experience. Yeah, it's just too many users, bro. I need a NIMBY AI. I need a NIMBY frontier model where no one else uses it except for me, so I can use it. I can get it at a higher rate. That's the appeal of Codex right now, I think—in theory, you should be able to have higher rates.

Jordan Nanos

Can we talk benchmarks? Max, when you saw the 4.7 release, for example, and we reviewed some of the benchmark scores, it was better on most, not better on all, and it led us to talk about where benchmarks are useful and where they're not. Doug's given the vibes, like obviously individual experience can be different across people, and it just drives their preferences, but there should be some objective way to say 4.7 is better or is not better than 4.6.

Doug O'Laughlin

Honestly, benchmarks obviously try to be that objective measure. I think they just no longer are today. I would say you need to be close to frontier performance in order to have a shot at being the true best model, but being number one on the benchmark ranking does not necessarily imply that you actually are the best model. And so I would say benchmarks today are most useful as a vibe check to make sure that the model's not total trash. We kind of talked about this in the newsletter article, but it's surprising to me how few people actually look into the details of the benchmarks to understand how unrepresentative they are of real LLM use cases.

I think a lot of people just hear a name like Humanity's Last Exam, and they assume, “Oh my God, surely if a model can solve Humanity's Last Exam, that implies that it's smarter than all of humanity, right? We've passed AGI or something.” In reality, you look at the individual questions, and they're just the most esoteric multiple-choice questions you've ever seen that are not at all representative of anything you've ever asked an LLM to do. It's very intentionally multiple-choice to make verification easy, even though obviously when you use an LLM in real life, it's open-ended.

This even applies to benchmarks that, on the surface, you might think would be better, like SWE-bench 2, where it's like, “Oh, coding is this verifiable task. Surely you can just come up with some coding problem for the model and then write some tests that verify if it's successful or not.” But then you dig into the details, and it's actually really hard to write a coding problem that is both naturally worded yet still perfectly unambiguous, with exactly 1 correct solution.

The SWE-bench problems, at least in the original version, don't at all fit those criteria. They just scrape GitHub issues, and all the developers listening know that GitHub issue descriptions are not meant to be well-scoped tasks that you just copy and paste and give to a model.

They often include lots of unit tests that are scoped to particular implementation details, going as far as asking the AI to output a specific 20-word error message that isn't at all mentioned in the task description. Obviously, later versions of SWE-bench tried to solve these issues, but it's still not perfect, and I think it really underscores a lot of the issues with benchmarks.

Jordan Nanos

Makes sense. So let's talk about 4.7 specifically. There were a few things that improved—or changed, let's say—in terms of features, as opposed to just the benchmark itself. I think you made the point in the article that people don't necessarily care as much about the quality of the model anymore. It's the model plus the harness—the product—that should be tested, as opposed to the model on a generic bash-only harness or something like that.

By harness, I mean Claude Code is the thing you're testing, not Opus 4.7. So it's Claude Code versus Codex; it's not Opus versus GPT.

To that end, during the release for 4.7, they announced an extra-high reasoning-effort option that slots between high and max. They announced high-resolution image support, which people can use for screenshots and styling in front-end applications. They're omitting thinking content by default, so people won't see when the model is thinking. They've got this task-budget concept that lets you tell the model how much it should think or work before it can actually run out of context in the context window. And perhaps most importantly, they updated their tokenizer.

It's potentially costing people 35% more for the exact same output from the model than from the previous model, just because they're counting tokens differently and using it. I'm just laughing at Dylan right now. Dude, I literally put him to sleep with that monologue. Holy shit. What the fuck?

Doug O'Laughlin

Oh my God. Okay, well—

Jordan Nanos

This isn't a joke. No, man, come on. Help me out here.

Doug O'Laughlin

What are you even talking about, man? The tokenizer is so boring that it puts Dylan to sleep. That's the takeaway.

Jordan Nanos

It's not like the pod was compelling and great at this point, but still—

Doug O'Laughlin

It is now. I'm compelled. Now we can talk about all the shit we want.

Jordan Nanos

Yeah.

What's the Michelangelo painting or sculpture where the guy is—

Doug O'Laughlin

The Thinker?

Jordan Nanos

Yeah, The Thinker.

Doug O'Laughlin

Yo, he's back.

Jordan Nanos

He's back.

Doug O'Laughlin

Dude, could you actually not hear us through your headphones?

Jordan Nanos

I can hear you.

Doug O'Laughlin

What do you mean? He's been asleep this entire time.

Dylan Patel

I was just thinking, man.

Doug O'Laughlin

You've been gone for a few hours.

Dylan Patel

Call me GPT-5.5 xhigh.

Doug O'Laughlin

Just thinking right now. Tokenizer: Opus 4.7, Opus 4.6. We've been dancing around the fact that GPT-5.5 is fine—not goated, definitely good enough, lots of capacity. They'll catch up on the margin. Sounds good? Great.

Speaker 0

Yeah.

Doug O'Laughlin

Anything else? Benchmarks suck. That's a really good take. Anything else?

Dylan Patel

What's the tokenizer difference?

Doug O'Laughlin

They changed the tokenizer between 4.6 and 4.7.

Jordan Nanos

The exact same output could have 35% more tokens with the new one.

Dylan Patel

Oh, they made the tokenizer—They made the vocabulary smaller.

Speaker 0

They made the vocabulary bigger for 4.7—

Doug O'Laughlin

Yeah.

Jordan Nanos

Compared to 4.6.

Dylan Patel

Oh, sorry.

Jordan Nanos

There are 35% more tokens in the tokenizer—more vocabulary.

Dylan Patel

So wouldn't that make the average output smaller? Because you can represent longer things with fewer tokens.

Speaker 0

No, more tokens.

Dylan Patel

Well, sorry. If you had only 28 tokens to make English, then you would have to use every letter. But if you wanted to do English with 500 tokens, sure, you'd have every letter, but then you'd also have tokens for “of” and “the.” Wouldn't that make the output smaller, with fewer tokens?

Speaker 0

This is a good take.

Doug O'Laughlin

Oh.

Jordan Nanos

In practice, no, but conceptually, yeah, you could train the model with full words as tokens. But I think what people have seen is that the model is currently less token-efficient with a larger vocabulary.

Dylan Patel

Oh, interesting.

Speaker 0

But that's a good point. Yeah. Man, the guy came back in with a heater here. So, yeah.

Doug O'Laughlin

Wow. He was thinking all this time. He—Oh.

Speaker 0

Okay. Yeah. Well, the whole concept of being more token-efficient is that you can solve tasks with fewer tokens, because the more granular breakup of the tokens—or the larger breakup in the token size, or just having a larger vocabulary—would mean that you would have more information represented in latent space.

The concept of—

Doug O'Laughlin

The relationship.

Speaker 0

SemiAnalysis might be a token instead of “Semi” and “Analysis,” right?

Doug O'Laughlin

Yeah. But then it creates more context, right? Versus maybe even 4 tokens of SemiAnalysis. I get it, right? There are choices.

Jordan Nanos

But it's richer information. The embedding for SemiAnalysis might be close to the embedding for Dylan in latent space, right? The embedding for Semi or the embedding for Analysis is probably not close to it. So whatever.

I think it's unclear whether this is significantly improving performance, because everybody still thinks it's a toss-up between 4.7 and 4.6 as to what's better. But then I also think it's a toss-up as to whether this is more token-efficient or worse. It seems like people are saying, if it's not a significantly better model, then why are we doing this new tokenizer thing?

Doug O'Laughlin

Yeah.

Jordan Nanos

So maybe it's just an early checkpoint. They've got to do more RL and improve 4.7, and then have a 4.8 drop that really improves things.

Doug O'Laughlin

I don't think Anthropic really does half-baked models. OpenAI clearly does. GPT-5.3-Codex was RL'd only on code, and then 5.4 was like—

Speaker 0

Is this the most half-baked of any model we've previously seen? We saw Sonnet before Opus, right? 4.5 Sonnet, then 4.5 Opus, then 4.6, and—Right? This is the first new one that's just Opus.

Speaker 3

It's because Opus 4.7 is actually Sonnet, dude. Are you not a truther?

Dylan Patel

Yeah, yeah, and then Opus 4.7 is Mitas.

Speaker 3

There you go. Yeah, there you go. This is truther stuff, bro. It's time for truth-truthers, bro. Opus 4.6 was actually Sonnet all along. Does everyone know—Do you not remember that? That was, like, a big model that smelled small.

Speaker 0

Yeah, removing the mask. Yeah, yeah.

Speaker 3

Exactly. This model smells small. Actually, I wanted to maybe pull this back to the DeepSeek portion of the article, because I don't think we talked about DeepSeek in—

Dylan Patel

The deep state?

Speaker 3

Yeah, the deep-state version of the article. We didn't talk about DeepSeek in depth, but I think we have a lot more internal takes than what we put in the article. And I guess my question is, do you think the gap between open source in China and the United States is now widening again? Because it feels like it is now. And it's because of compute constraints. I'd like to have some takes on that. Just my take is the take.

Speaker 0

Yeah.

Dylan Patel

Yes.

Speaker 0

I'll give a quick one. I think one thing that's been overlooked with DeepSeek is the fact that this is a 1-million-token-context-window model, which many of the leading open-source models that perform great on benchmarks people use for coding don't have. And then I think that the “Reasoning in Visual Space” paper that they posted and then took down as a GitHub repo makes me think they're going to release a multimodal version of this as well, or that they're in the process of developing it, and there are going to be new weights for that. Both of those things are fascinating for making China catch up on a product basis.

If we're comparing Claude Code to Codex to DeepSeek in OpenCode, or DeepSeek in some other harness, it'll be able to support all the same features in addition to being pretty good at all the other stuff. With that said, at the time of the first DeepSeek release, I used it all the time for random stuff, and I don't use this one for anything, really. So I—

Doug O'Laughlin

Isn't Kimi K2.6 better anyway?

Speaker 0

Yeah, and I don't use that either.

Doug O'Laughlin

No, I'm just saying it wasn't like DeepSeek—I mean, the DeepSeek moment was that it was so cooked and it was from the ether, right? Or it cooked so hard, rather, and it was from the ether—not cooked, right? But this round, it comes out, it's just state-of-the-art, or it's state-of-the-open-source art, if that makes sense. It's not exactly better than Kimi K2.6. I think the other stuff in it is clearly just inference optimization, right? They talked about the Ascend kernel being partially able to run inference on it, which would really, really unlock more compute for China for the first time.

Dylan Patel

Then also, if you look at the weight size, it looks very convenient. I feel like the softmax for China is effectively the size of an H200 8× pod. All the models you're looking at are essentially able to run inference within that memory-domain space, and there's nothing bigger that's served at the state of the art.

Clearly, that seems to be the cap, right? Maybe they can do that, but they won't release it on their B200 pods to the public. It just clearly feels like they're starting to hit some kind of wall. Agree or disagree? Do you think that will keep or cap Chinese progress because they can't run inference on this at all? I'd like to hear some hot takes here.

Jordan Nanos

Okay, I think—

Dylan Patel

Dude, what the hell was that, Jordan?

Doug O'Laughlin

He's getting hot and sweaty over DeepSeek.

Dylan Patel

Yeah, you're fucking deep-panting, bro.

Doug O'Laughlin

He's getting hot and sweaty over DeepSeek.

Jordan Nanos

The DeepSeek engineering release is fascinating. All the new attention variants and the compression on the KV caches are fascinating stuff. They do so well on the infrastructure stuff. I think, again, the fact that Kimi's 256K context and DeepSeek's 1 million is a significant difference for long-horizon agentic tasks.

Dylan Patel

But isn't the context from 256K to 1 million dogshit anyway, even on Opus?

Doug O'Laughlin

Yes.

Jordan Nanos

That's not my—

Doug O'Laughlin

It's garbage. I mean, okay, it's not true garbage. It's probably a step off of the state of the art, in theory, but the problem is you don't want to just be clearing your context every time. If you're doing a big task, seeing the whole context window is really nice.

Jordan Nanos

Compaction sucks.

Doug O'Laughlin

Compaction blows.

Jordan Nanos

Not being able to read a 1 million-token context on this stuff sucks. Now—

Doug O'Laughlin

It's actually better to clear. That's my hot take. I would rather just start over. I would literally be like, “Make a summary of what we've done, copy-paste that, and just start over.” Fuck the compaction.

Jordan Nanos

No, that's what DeepSeek saw in their 3.2 paper. For their benchmarks—which I think are probably bad, but they published this—if you go beyond the context window for a given task, it's better, on the tasks they were testing, to completely clear the context, not even make a summary. There is a compaction—

Doug O'Laughlin

Yeah, isn't making a summary just what compaction is under the hood? To be clear—

Jordan Nanos

So when I say “make a summary,” I mean literally, okay, you can use the entire context window versus just what I was doing last. I'm trying to remember, because I'm not going to read 1 million tokens of slop. It's like, “Okay, what was I doing here?” Read this, and then I'll Control+C a very small part. So I'm not even—

Doug O'Laughlin

There are different—

Jordan Nanos

I'm not summarizing the entire thing. I'm just doing the task.

Doug O'Laughlin

I see. Okay.

Jordan Nanos

I'm passing off tasks. Yeah.

Doug O'Laughlin

There are different ways to do compaction, but I believe compaction is different from summarization because compaction is removing the thinking traces.

Jordan Nanos

Initially, it's not actually having the model write its own summary of the full context.

Max Kan

Oh, I thought they were literally just taking your entire context and saying, “Yo, please summarize this.”

Jordan Nanos

That is an approach, and they've done multiple of them. But the—

Dylan Patel

I've got a hot take. Anthropic, with Claude Code, did well with the CLI, so they just kept making the CLI experience amazing. But the CLI experience is not the end-all, be-all of agent orchestration, and therefore they've really cooked themselves into an innovator's dilemma, where they keep making the CLI better. OpenAI has the true vision of what the true agent-orchestration platform of the future is, where you'll be able to integrate voice, multimodality, and all these other things into the app. The app is so much better than the CLI, and the real—

Jordan Nanos

Mm-hmm.

Dylan Patel

—the point is that you should develop, and users should be using it, in the app, not on the CLI, because the CLI is a dead end, a foregone relic of H1 2026 and H2 2025.

Max Kan

Do you mean the app forever, or do you mean their device? They're talking about releasing a consumer device next year too.

Dylan Patel

No, no, no. I mean the laptop—

Jordan Nanos

Like that, huh?

Dylan Patel

—you know, app. The Codex app.

Jordan Nanos

Yeah.

Dylan Patel

The Codex app.

Max Kan

I have a question then. Why does the Codex app suck?

Jordan Nanos

Yeah. True say.

Max Kan

So, look—

Dylan Patel

It's long-term planning.

Max Kan

I think—

Dylan Patel

Long-term planning.

Max Kan

Dylan, in my opinion—

Dylan Patel

Look—

Max Kan

I think you are thinking too small because, in the perfect, true maxi world—

Jordan Nanos

Yeah.

Max Kan

—the operating system doesn't need to exist. You will just get a piece of hardware. You will plug in your thing. It will pull up the terminal, and you will connect your Claude API, and it will build the OS for you. Thinking—

Dylan Patel

So, Codex app, they're adding generative UI stuff too, which is pretty interesting.

Max Kan

I'm just saying, I think if you're a coding purist, generative UI is downstream of the CLI. I think I'm a CLI purist. Dude, I don't know. This is just a slop preference thing. I just love the CLI, man. Claude Code usage is clearly just a CLI wrapper, and you can tell, and then Codex CLI is clearly just an app wrapper. I feel like they forced it over. I think there are 2 opinions about the future. Who knows who will win out in the very long run? I'm definitely going to keep it open for competition.

But at this beautiful moment, a true maxi's vision and dream is that it's all downstream from the CLI. It's just tokens. It's the most efficient version of everything, man. Gotta Elon-max, okay? All you need is just an API and then inputs. That's it. Your app and all that stuff, that's all obfuscation. Mythos would know better than OpenAI. Who are we little brains to know what UI we want? No, dude, it's CLI all the way down. Pure maxi vision.

Jordan Nanos

I'm a VS Code plugin guy. I literally tested this yesterday. I was bothering Max, who told me to go away, about using Ghostty for the CLI stuff. It just doesn't work.

Max Kan

Wait, no, I still think having 6 Ghostty terminals open with Claude Code CLI is a superior experience to having 6 different chats going in the Codex app.

Jordan Nanos

Oh, for sure.

Max Kan

I feel more productive.

Jordan Nanos

Yeah, I'm comparing it to the VS Code plugin: 6 different windows in VS Code, also with a file browser on the left side so I can right-click and copy—

Max Kan

Well, I think the difference is that you're still writing real code and you kind of care about the output, whereas I just don't even look at it. I just directly push the—

Jordan Nanos

No, I'm not looking at the code. I'm copying in images or Excel files. Don't accuse me of reading the code. Sorry.

Dylan Patel

He's fucking Edison.

Max Kan

My favorite is—I mean, I actually agree that the API, CLI, whatever, is the new compiler. No one's reading the compiler. No one cares. They don't need to touch the magic.

Jordan Nanos

This will be true in the future, but it is currently—for any code you actually care about—this is currently not true. It's still producing a bunch of stuff that is bad and should be fixed by coaxing the model to fix it. I'm not saying I type code anymore, but I do read some code.

Dylan Patel

One of my group chats—

Max Kan

Yeah.

Dylan Patel

—one of my group chats, I was reading it this morning, and it's a group chat with all the most cracked kernel programmers in the world. It turns out what they do—

Max Kan

Wait, why are you in it?

Dylan Patel

Because they're my boys, bro. Look, Max, come on. I'm a master networker, okay?

So, anyway, I think it was Tri Dao. TreeDAO's like, “Yeah, dude, Codex is so dumb, but I always just have it create it, and it works and it's smarter, but the code is slop, and then I have Opus rewrite it. But you can't go the other way around. You can't have Opus write the thing and then have Codex fix it. You have to have—

Max Kan

You know what's funny?

Jordan Nanos

Really? I go the other way around.

Max Kan

Codex write it and then have Opus fix it. Everyone else at the firm prefers the other way around.

Jordan Nanos

Yeah.

Max Kan

The entire firm's preference is the other way around, actually.

Dylan Patel

Yeah, but we're not writing fucking kernels, right? We're not writing fucking Tri Dao kernels.

Max Kan

That's probably fair.

Jordan Nanos

We're doing benchmarks of kernels, but yeah. We wrote some kernels.

Dylan Patel

Oh, come on. Dude, they're not Tri Dao kernels.

Jordan Nanos

No, they're not Tri Dao kernels. They're just GPU MODE kernel competition kernels.

Dylan Patel

Because apparently, if you talk about niche and microarchitecture details, Claude will waffle on about shit instead of actually just doing it. Whereas if you describe it to Codex, it'll just try and implement it all, and then it'll be slop. But then you tell Opus to fix it, Opus won't waffle on; it'll just fix it.

Jordan Nanos

Yeah. Doug, this is the context-window stuff, which is like, when you are pumping in so many docs about the ISA of a given GPU to write a kernel or something, you need performance at a million context. You just run out of space on the smaller stuff. So I don't know. Do you want to go back to DeepSeek and any hot takes on DeepSeek, Dylan? Why didn't it crash the market this time if KV cache is reduced by 90%?

Dylan Patel

Dude, you know, it's been a while since I've been in Asia, but every time I go to Asia, they reference some fucking new paper that reduces KV cache every fucking time for the last 3 years. Some paper, they're reducing KV cache, and no researcher in America has even heard of this paper. It's the fucking best thing ever. DeepSeek and TurboQuant were the most precipitous ones that popped up the most, and TurboQuant was obviously fake news. But yeah, I think it's very funny. I don't know.

I guess they're tired of being robbed.

Max Kan

Okay. Well, look, yeah, I think that's fair. It just doesn't matter. Gemini's working, clearly with the price of the GPU going up. That's all you need to know. Now, if we're gonna talk about real fake news, let's talk about SubQ. Let's do some—I mean, we're not gonna write an article about it. We're not gonna write a post about it. This is free alpha. Did anyone else read the SubQ thing today? It's pretty sus. It's actually extremely ultra-mega sus.

Jordan Nanos

Yeah. It seems like people are launching their startup, right?

Max Kan

Honestly, they should close funding, and then they'd be like, “Wow, it was just Opus with 10 context windows stapled together.” I mean—

Jordan Nanos

Do you think the market is hot enough for them to close, you know, a $200 million at $1 billion round next month or something? If they did, I would be impressed.

Dylan Patel

I don't know if they could do $200 million, but I think they could do $50 million at a bill. A tril—

Max Kan

A tril? What?

Dylan Patel

Sorry, bill. No, tril.

Jordan Nanos

There's more capital than there is opportunity.

Max Kan

Wait, wait, wait. Are you familiar with what we're even talking about, Dylan?

Dylan Patel

No, sorry, I just thought you guys were talking about Anthropic.

Max Kan

No, we're not talking about Anthropic.

Dylan Patel

Oh.

Max Kan

He's talking about model sparsity. I'm talking about the worst—did you not? It's like this fake-news Twitter thing today called SubQ.

Dylan Patel

Oh.

Max Kan

Um—

Dylan Patel

Yeah, yeah, yeah, yeah.

Max Kan

Yeah, yeah.

Dylan Patel

That's another fake-news one.

Max Kan

Don't worry. We requested API access. We made sure to use our SemiAnalysis email to improve our odds.

Jordan Nanos

Yeah. We're like, “Please, give us this API for this very real model, bro.”

Max Kan

Who knows? Maybe it's a state-space model. Maybe Mamba cooks, or—

Dylan Patel

No, it's not an SSM. I don't think it's an SSM.

Max Kan

I think it's—I don't know. It's just really funny because, again, we're talking about DeepSeek people freaking out. Dude, if this was real, memory stocks should be down like whatever, a quadrillion percent today. But obviously it's not real because if you look at these guys and you're like, yeah, man, I just don't think these guys are gonna be the guys to crack the single hardest problem in all of AI. No offense—maybe the founder's super legit.

Jordan Nanos

So, okay, maybe one thing this reminded me of was the fact that Llama 4 Scout or Maverick—I think Scout, the smallest one—was released with a 10-million-token context window, or announced with it, but not supporting it officially in the released weights or something. And I'm just really surprised that we haven't seen anybody with effectively an unlimited compute budget give it a go for a more expensive model with a larger context window. Like—

Dylan Patel

But what—where are you gonna get the data, right? Most people pretrain with 16K context or 4K, you know, something like that, 32K context, and then they post-train it so that they can add and hack in the rest of the context. But it's like, what data do I have? That's why my 250K to 1M context is trash anyways, is because there's no data on this stuff. And so the model doesn't generalize the context really well. And then if you stick it to 10 million, it's like, what fucking data do I have that is useful for the next-token generation that exists from 1 million context to 10 million context? There's so little.

Jordan Nanos

Yeah. I mean, it makes sense. Possibly synthetic stuff, possibly—I mean, why'd they do it in the first place? It seems obvious that people would be working on it, and we haven't even seen anybody announce 2 million. So there's some arbitrary limit—

Dylan Patel

Wait, but Google serves 2 million.

Jordan Nanos

Google serves 2 million on Gemini 3.1 Pro?

Dylan Patel

They did on Gemini 2.5—2 million.

Jordan Nanos

Well, maybe that's—

Max Kan

It's 1 million.

Jordan Nanos

That's the answer to me, yeah.

Max Kan

It's 1 million today.

Dylan Patel

It's 1 million today?

Max Kan

On Gemini 3.1 Pro.

Dylan Patel

One of their announcements—

Jordan Nanos

No.

Dylan Patel

They announced 10 million. They started at 1, and then they updated it to 2 at some point in one of the models.

Jordan Nanos

Makes sense, yeah. I mean, they got a big scale-up domain. Why not give it a go with the TPUs? Yeah, maybe another thing that was a little bit missed in the article, and you kind of talked about it when you brought up DeepSeek. Max, I want your take on this. You—because Doug asked the bait question about whether China is catching up or they're still behind. It kind of depends on how you look at it. But I was bugging you the other day: Is DeepSeek or Kimi currently ahead or behind Meta? And are they ahead or behind Grok, Cursor, or SpaceX's xAI?

Max Kan

I would say that today they're probably ahead of all those companies. But the thing that really matters is slope from here. This is a pretty basic take at this point, but I do think the amount of compute you have is actually just one of the key inputs to how good your model's gonna be. And obviously Meta's signing all these monster deals. It seems like they've overcome the overhang of having to fire and then rehire their entire AI team, and they're in the process of making some good models now. So I would expect Meta to pull away from all the Chinese guys, if not in the second half of this year, then in the first half of 2027.

Dylan Patel

And Meta's not distilling.

Jordan Nanos

Oh, I thought they were.

Max Kan

I thought they were all distilling.

Jordan Nanos

I thought they were distilling from the Chinese guys. They were just running the open-source models.

Dylan Patel

I mean, that's what Mistral does. They don't distill from Anthropic. They distill from the Chinese guys. That's fair.

Jordan Nanos

Yeah.

Max Kan

Why are we talking about the leading French frontier model company, Mistral?

Dylan Patel

Dude, you know their revenue's really strong.

Max Kan

Yeah, I do, actually. You know what they've bro'd down on? Also, dude, the bottles.

Jordan Nanos

Because they're a new product.

Dylan Patel

You keep picking up the bottle on the mic.

Jordan Nanos

Every frontier model—

Dylan Patel

Well, no, not neocloud. They keep trying to—

Jordan Nanos

…is in a cloud—neocloud.

Dylan Patel

No, yeah, they're trying to become a neocloud, or at least they're doing fine-tunes.

Jordan Nanos

Every chip company. Cerebras is becoming a neocloud. NVIDIA's launching neoclouds.

Doug O'Laughlin

The ultimate business model—

Jordan Nanos

AMD—

Doug O'Laughlin

…for any company in the world is to become a neocloud.

Jordan Nanos

It's starting—

Doug O'Laughlin

SemiAnalysis will become a neocloud. And then we will be ClusterMAX Platinum.

Dylan Patel

Diamond. No, dude, we gotta introduce a new tier. Yeah, diamond. Tungsten.

Jordan Nanos

Yeah. SemiAnalysis, lithium.

Dylan Patel

I don't know, germanium? I don't know, I'm just making up shit. What's the rarest?

Jordan Nanos

What's your favorite semiconductor, Doug? We should make the tiers semiconducting materials only.

Doug O'Laughlin

Oh, yeah? You like semiconductors?

Jordan Nanos

They already are, but—

Doug O'Laughlin

Name all of them.

Dylan Patel

Name them all, yeah. We should be rhodium—the rarest and most expensive precious metal.

Jordan Nanos

Vanadium?

Doug O'Laughlin

All right, boys, this is getting off track. We’ve got to get out of here.

Dylan Patel

Yeah, okay. Any other parting takes?

Jordan Nanos

No, I think the hot takes have run out. Claude Code was the inflection point in February 2026. Doug, your victory lap today in May is complete. I appreciate all the hot takes today.

Dylan Patel

I hope it’s not the next inflection point. I hope it’s more exciting than that.

RLHF is very appreciated. Yeah, better than Dylan.

Speaker 0

I made it to the end.

Speaker 3

Yeah. He didn’t even make it to the end.

Speaker 0

Yeah.

Ep. 011 - GPT 5.5 vs Claude 4.7: OpenAI's Comeback From the Brink (Tokenomics) | Jordan Nanos, Dylan Patel, Doug O'Laughlin, Max Kan | BidClub