Dylan Patel
All right, Max. We're going to do a podcast. We're going to talk about everything related to Kimi K3 and maybe some other models that just came out. How are you doing?
Max
Doing great. Looking forward to it, and thanks for having me, Jordan.
Dylan Patel
I'm not having you. I'll say thanks for having me. All right, on the docket: Kimi K3 hot takes. Is this the 3rd-best model in the world? What's the impact on OpenAI and Anthropic, architecture changes, and personal usage that we've had so far? What do we think about their open-source strategy, and maybe more? All right, Max, quick hot take: Is this the 3rd-best model in the world right now?
Max
I think the answer is a clear yes. People love shitting on benchmarks. I think benchmarks definitely have their problems, but if you take a composite of all the main benchmarks and look at all their rankings, they've been directionally correct over time.
I think if you look at that composite today, there's a pretty clear top 3 with stable, Soul 5.6, and now Kimmy K3. They're always above everyone else, which includes, of course, other open-source models like Deep Sea, Conjure, and whoever. It also notably includes Google and Meta and SpaceX.
I think it's honestly a truly impressive and very remarkable feat from the Moonshot guys. Google, in particular, should feel incredibly embarrassed right now. At one point, as recently as November or December 2025, everyone thought the clear AI big 3 were Google, Anthropic, and OpenAI. Even when I talk to boomers today, they still seem to think the top 3 are Google, Anthropic, and OpenAI. That's clearly not the case anymore.
I'd say it's definitely the 3rd-best model in the world. I do think it's overall still worse than Fable and Soul 5.6. It's kind of funny that they explicitly said that in their model-release blog post. Maybe it's some old-fashioned Chinese humility. Maybe they don't want to incur scrutiny from the U.S. government or anything, because obviously there were some delays with the Fable 1.56 release. Overall, though, I'm very impressed with the model.
Dylan Patel
Yeah, in the limitations section of the blog post, they said, “Despite being a highly competitive model, overall K3 nonetheless exhibits a noticeable gap in user experience compared with Fable 5 and GPT 5.6 full.”
My experience using this personally is that it is good. It's really slow, which is really annoying. It's motivated me to try open-source harnesses for the first time, and I feel like I'm learning more about the harnesses than I am about the models because, frankly, all these models are good enough to do the basic work that I've been doing so far. I can't really find a lot of complicated stuff that it can't do, which in and of itself is a bit of a feat.
Here's my hot take: For me, this might be the 2nd-best model in the world right now, because every time I try to do something meaningful with Opus, I get rejected and sent down to Opus. Even though I don't know if this is better than Opus, it is less annoying not to get rejected whenever I'm trying to do something.
However, I'm not getting rejected when I use my API key and pay for tokens, but I am hitting limits whenever I try to use the web console, deep research, or the coding plan. I haven't used the coding plan, but some other guys at Anthropic have. That leads me to ask: What's the strategy here? These guys clearly don't have enough GPUs to serve the demand they're seeing for this model.
Previously, that was solved by an open-source strategy where they just dropped the weights, and then other people served the model and served that demand. But they haven't dropped the weights yet, so I think they said the weights would be available in 10 days or something.
Max
Yeah, I think that's what they said.
Dylan Patel
What do you think the strategy is for the delay between the announcement of the model, the API being available, and the weights not being available yet?
Max
To be clear, this is all pure speculation on my part. I think one big reason is that they need to give the LLM and SU lane guys enough time to make sure they can serve this model performantly. If they just dropped it today, you'd have all this hype, but then everyone else serving the model would be giving you 22 seconds or something. That's probably really bad for the brand.
They have this incredible opportunity to get a bunch of huge PR and adoption, and they also need capital a little bit. I think another possibility is that they're actively talking to Together, Fireworks, Nebius, and CoreWeave to figure out how they can sign some sort of licensing deal, have them serve the model, and develop capacity on GB300s or whatever.
In my mind, those are the 2 main reasons why you would wait 10 days to actually drop the model weights.
Dylan Patel
Yeah, that makes sense. Functionally, this is really interesting because, just to talk about the model architecture for a second, it's 2.8 trillion parameters. This does not fit in a B200, so you need to have a B300, GB300, or, I guess, MI355X in order to serve this model on a single system, like a single 8-way HGX server.
Of course, you can do unique things where you have pipeline parallelism across multiple nodes and stuff, but that's going to really impact performance. I think there are a lot of recipes being cooked up, and only people who have the latest and greatest chips are going to be able to serve this model.
Let's go back to your comment on Google for a second. The idea that this model is truly competitive at the frontier with 2.8 trillion parameters gives us some insight into how big the closed-source frontier models are, right? It would be even more embarrassing if they're hitting these levels of performance and being compared to 10-trillion-parameter models with a lot more active parameters, right? We have to assume that this is in the same range as what Sonnet and Opus are, right?
Max
Yeah, I think that's a great point, and you just have to be correct. I still believe in the competence and correctness of all the OpenAI and Anthropic researchers. If there are some people on Twitter who like to claim that current closed-source models have 10 trillion total parameters or something, if that's actually true, guys, it's time to pack up the bags. The stock price probably should crash 50% tomorrow. It's over.
I'm pretty confident that Kimi K3 cannot be much smaller. If anything, it might even be slightly bigger than the leading closed-source models today. If that's true, this further highlights the point we've been harping on for a while at SemiAnalysis: The margins for these closed-source labs have to be absolutely mind-boggling.
If you're telling me Kimi's probably not operating at negative margins and they're serving K3 at $3/$15 per million output tokens, that's sort of the same price as Sonic. If you're telling me that Fable is probably similarly sized and the profit can charge $10/$50 per million output tokens, then this should immediately dispel any remaining concerns people have about the AI labs being unprofitable businesses.
Selling tokens at API prices might be even better than SaaS, honestly. It's an incredible business today.
Dylan Patel
Yeah, that makes sense. No cost of employees, just the GPUs. Can you compare this pricing strategy to the previous stuff? You said it's at $3/$15. The previous version from Moonshot directly was at $0.95 and $4, so we're talking about—
Max
Yeah.
Dylan Patel
—more than a 3× pricing increase from 2.7 code to Kimmy K3. Do they have even more room to increase pricing? What's the curve to get to frontier open-source intelligence, or frontier soon-to-be-open-weight intelligence, with the license yet to be determined?
Max
Honestly, I don't think they have that much more room to push pricing up. I would guess that even at $3/$15, there will be a lot of people who say, “This is a little too expensive for me. My task is easy enough for a GLM-5.2 or a MiniMax M3, and I might just use one of those models instead.”
On one end, you have the SemiAnalysis of the world, right? We don't really care how much money we're costing Dylan when we burn tokens all day. We're very happy using Opus for even a relatively easy task that we're pretty confident one of these other open-source models can do pretty well.
On the other end, you have people who are extremely cost-conscious. You may only get $200 worth of tokens per week, as you've heard some large companies like Tesla and Uber are implementing. Pretty much all the people in that second bucket are going to want to use the GLM-kind-of-pricing-tier models because they're already good enough for most everyday tasks. Everyone in the SemiAnalysis bucket is still using GPT-5.6, too, and Opus.
I think there actually is a pretty interesting question of who the user is that will actually be switching to Kimi K3. It might just be a lot of people who philosophically love open source and are excited to try this new hype model and support it.
Max
But it wouldn't surprise me at all if there isn't serious adoption of this model among, say, large enterprises.
Dylan Patel
Okay. What do you think about where we go from here? This is obviously a new base model and a completely new architecture for these guys: 2.8 trillion parameters. They've got Kimi Delta Attention, attention residuals, and the Stable Latent MoE that they keep using. It's a scaled-up, bigger version of the previous models—clearly about 2 times bigger.
Max
Mhm.
Dylan Patel
Previously, with Kimi K2.5, we saw Cursor train Composer based on
Jordan
Mhm.
Dylan Patel
just continued pre-training, as well as some RL. Then we saw Kimi give us K2.5, K2.6, and K2.7 checkpoints as they continued the RL. This is a new base model, and it seems pretty complete. In my usage, it's working pretty well. It's not screwing up anything basic when it comes to writing a PR description or totally going off the rails, the way we've seen some other models that are raw without a bunch of RL have rough edges at the beginning.
So where do we go from here? When does Kimi K3.1 come out? How does pricing change over time? Do we get a Composer based on Kimi K3? I mean, a Composer based on Kimi K3 is definitely not happening because I think the Cursor guys are pretty set on training their own model from scratch now.
As for when Kimi K3.1, Kimi K3.2, or whatever comes out, I imagine we'll probably see 2 or 3 updates within the next few months, each a month or 2 apart, as they continue post-training this thing. I would guess pricing stays about the same, just because they're not going to be able to run it on new hardware in the next 2 or 3 months. They're not going to get a huge throughput increase there to reduce pricing.
Maybe it's possible that some really crack engineers figure out how to reduce the cost to serve this thing so it's closer to DFC V4 pricing or something. That would be really impressive, but given that it's a 3 trillion parameter model, I'm a little skeptical. I would guess that the current pricing we see for the MiniMaxes and the GLMs is already pushing the limits of what you can charge to serve a 1T-to-1.5T model without having embarrassingly bad margins.
So I think this pricing is probably here to stay for at least the next few months. I think the most interesting question is whether the open-versus-closed gap is going to continue shrinking, and whether open models will ever fully match closed-source models with true frontier-level parity. I'm curious what your thoughts are there, Jordan. I think it has serious implications for our whole industry if it actually happens.
Dylan Patel
Yeah, I mean, my view is that I believe the reason this gap has closed right now is squarely due to the US government imposing restrictions on Anthropic, resulting in us not getting the actual best models that these guys have. They've artificially caught up, basically.
Max
Interesting.
Dylan Patel
Clearly, we see this with Mythos versus Fable. I can't use Opus; I can only use Sonnet sometimes if I ask it nicely. 5.6 Soul, I think our host view is that it's not the biggest model OpenAI has ever trained. It's not the size of GPT-4.5. To me, they have a bigger model somewhere.
I think the result is that we're only going to be able to access frontier intelligence if government entities allow us to. That's a very interesting change to the setup going forward because I think it represents an opportunity for many of the players that are in 4th, 5th, 6th, or 7th place to catch up to a limit, at which point it's okay to release everything and start battling for user share without really being able to find the frontiers and have the frontier dominate.
I think it's possible that we see the frontier take another big step toward the end of the summer. It's possible the politics change a little bit.
Jordan
Yeah.
Dylan Patel
It's possible that we start to find other modalities beyond coding where these guys can really improve, and they start exploring those areas. We didn't intend to talk about this right away, but I loved the release of Inkling by Thinking Machines. I thought the native audio input would be super interesting and super useful in the future, and kind of a sign of what's to come. But anyway, yeah, I—
Max
And on the topic of Inkling, there's definitely huge demand for a Western open-source model that doesn't suck. I'm shocked that markets are still so inefficient and that we haven't had a single American company that's at least on par with the 5th-best Chinese company.
One, it's only a matter of time until the US government bans Chinese open-source models entirely. Maybe that's a can of worms we don't have to go down in this conversation. But even if that doesn't happen, I feel like the average large American enterprise is simply unwilling to put all of its proprietary data through a Chinese open-source model, even though you can make tautological arguments like, “You're just loading their weights in your air-gapped data center. There's no way the CCP is actually going to see any of your data.”
I don't think the executives will actually buy that, and I don't think they really care. There are a lot of people who, A, care about token budgeting and, B, are only interested in running a Western model or a non-Chinese model. It's shocking to me that we're not actually closer to the open-source frontier in America.
Dylan Patel
Yeah, I mean, there was NVIDIA Nemotron, and then there was Inkling. It's really inspiring to see Tinker go for it. I think they have 2 business opportunities there. They've got to be better than the bulk of Chinese open-source models. They have to be in the game there to be considered.
Max
Yeah.
Dylan Patel
But then they also need to be better than Sauna or better than Terra Luna—the tier-2 and tier-3 models from the frontier labs. You can build a bunch of cheap applications using close-to-frontier intelligence with Bedrock or Foundry or whatever, get access to the Anthropic or OpenAI models, and save money by going with their 2nd-best model.
I never really understand the Western open-source angle of saving people money. I think it is real, and getting those models into the ecosystem of companies like Fireworks, Together, and Baseten, and anybody who's serving open source, is a good thing because it is a market. But to me, the bulk of the market is government.
One interesting view on the Chinese models is that Xi has been encouraging the Chinese companies to keep the models open source. That is the view from their party. I think the big reason for that is that a bunch of the Chinese government wants to download the weights and run them on servers that they own, and they want the support of the local ecosystem.
I think the American government should work the exact same way. That's a pretty pragmatic strategy: you need to give the people in your country access and support to run this stuff. Maybe the other thing worth commenting on is that, in the Kimi K3 blog, they mentioned post-training—sorry, quantization during the SFT stage. They were commenting on natively using MXFP4 and MXFP8 weights and activations, respectively, for broad hardware compatibility.
Well, what other hardware do you think Moonshot cares about, Jordan?
Jordan
I've got a list of 11 Chinese accelerators. Huawei Ascend, Baidu, Kunlun Haxen, and the Moore Threads guys—there are all sorts of different chips showing up in papers. We're seeing code. It's a national priority for China to get these frontier models—these are frontier models now—running on their domestic accelerators.
Dylan Patel
Yeah, I mean, if we're calling Google a frontier lab, in 2025 we have to call—
Max
Call Moonshot a frontier lab now.
Dylan Patel
Kind of crazy, dude.
Max
It's like vanity sizing.
Dylan Patel
I'm still a 34 waist.
Max
Yeah, yeah. And so are the 7 other Chinese labs.
Dylan Patel
Yeah. No, no. Funny enough, my dad is actually visiting China right now, and he's telling me that the hotel he's currently staying at is totally booked because Xi Jinping is going to be in the area soon and is going to give a speech about how AI is a top priority for China.
Circling back to what you said earlier about the US government and how, if they keep kneecapping the frontier models OpenAI and Anthropic have—forcing them to delay them, forcing them to only have their 2nd-best model publicly available—and therefore giving all the other players, the Googles, the Space X's, the Metas, whoever, time to catch up, do you think that completely destroys the frontier-lab business model?
If you're OpenAI or Anthropic, you just lose all pricing power at that point, right? I don't see how Anthropic can still accelerate net-new ARR if its model is on par or comparable with the Meta model, the xAI model, the Google model, the Moonshot model, and the DeepSeek model. What happens to our industry at that point, Jordan?
Max
Yeah, I mean, first of all, no, I don't think that's going to happen, and I think I can explain why. But first of all, I don't know for sure, so we'll have to see it play out.
Max
Interesting to think about. I think the biggest thing that I've realized in my personal usage of this stuff is, one, how hard it's getting to differentiate between using the absolute frontier model and the max thinking mode versus high versus medium effort on those models.
Dylan Patel
Yeah.
Max
It's really, really hard for me to find day-to-day tasks that these models can't figure out. My behavior defaults to the biggest and hardest thinking because I don't care about Dylan's budget. But when it comes to actually using this, there is an aspect of the hardest being part of the product. So, testing Kimi K3 requires me to take a serious look at OpenCode, Hermes, and Pi.
The hardest is totally part of the product still. Simple things can cause me to want to use one model over the other. Can I install it on my remote SSH server? How easy are the keystrokes to get stuff in? Can I edit previous commands? These little tiny features in the harness actually impact where I'm going to send my tokens, which results in where I'm going to send my budget, right?
Dylan Patel
That's an interesting point because I think a lot of people talk about the token machine, right? Correct me if I'm wrong, but what I'm hearing from your description of your own workflow is that even for tasks where I'm pretty confident that a GLM could successfully do it, I'm happy routing it to Opus and doing it on max intelligence because the ROI of that task is still worth the Opus price to me.
There's always going to be some risk in the back of your mind where it's like, if I use GLM instead, even on medium thinking mode, it's way cheaper. Maybe it's not actually as high quality as Opus would have been, right? So even if the benchmarks claim that a lot of your tasks can move to GLM, you're still fine keeping them on Anthropic models or OpenAI models for the foreseeable future.
Max
Mostly, yes, but I use a lot of Slack bots right now. I actually don't know what model is running behind the scenes on those Slack bots. Specifically, in Perplexity Computer, I think if it starts routing it to Kimi K3, if it starts routing it to GLM, or if it starts routing it to Sonnet—and I know it's doing it today because I looked at my usage a few weeks ago and found how much of the OpenAI models I was using because it was making that decision—I don't really care which model they're using, right?
For a first cut at a PR before I go in and actually fix some stuff up, I don't really care which model they're using. That is about the quality of the harness there for what I'm using.
Dylan Patel
That might actually be a pitch for outcome-based pricing, if anything. One of these labs could potentially just get 95%-plus margins if they do outcome-based pricing for you because, as you said, all these tasks you're happy to pay even stable pricing for could probably get done at a fraction of the price even today.
Max
Yeah, yeah, 100%. Certainly with dialing in the thinking mode, which is where a ton of the expense ends up going, I can totally imagine them building a router.
The second thing, though, just on the competition thing you said earlier, is I don't think we're out of use cases or ideas for these guys to work on. I think they can continue to train incredible models to try and hit RSI on the coding side without ever releasing it to us, the proletariat, and keep their bourgeois models training each other. They can keep distilling them and giving us little tastes of it while still pursuing a research objective that includes all sorts of other uses of AI.
We're really exploring coding right now, but we do some video generation stuff. We do a lot of audio-to-audio stuff. We do lots of deep research that really doesn't look like coding in some ways. I think there are lots of use cases that they can continue to explore without really encountering the cybersecurity issues.
Robotics and world models are a simple one, right? What if Anthropic sets its sights on automating away a whole bunch of manual labor jobs instead of knowledge-work jobs? The idea that there's no way for them to build a sustainable business with great ROI for their—
Dylan Patel
Mhm.
Max
—greatest technology the world's ever seen, I don't believe that at all.
Dylan Patel
Yeah, it doesn't pass the smell test.
Max
No, not at all. But even beyond that, I use these models so much every day. First of all, I see how much my friends who work in technology, who are software engineers, spend 10 times less than me and use them 10 times less right now.
One person using Opus is like, you know, a person using Sonnet and a person using Opus. They use it both the same amount on the same day, but the person using Opus spends 10 times more at 90% margins. They make up the bulk of the business, right?
As soon as those people—of which I would say there's maybe, maybe I'm in the top 10%, maybe even the 1% of the industry—start using the bigger models, they use them more. That's just more demand for all of this business.
The models don't even need to get any better for them to discover they can use them for the really important tasks or the bigger ideas that they have. Then I need to go and talk to my neighbors who don't work in technology. There, I'm certainly in the 1%, probably in the 0.1%, maybe 0.01%, and maybe we've got a thousand times more to go from here.
So I get back to Masason's golden-goose exponential chart to the right, sort of this point. It's an exponential.
Dylan Patel
Okay, just to take this totally off the rails. But actually, before we go there, I want to say that I think what you said about being early is totally right. This is exactly why I don't think Kimi K3 is going to cause net-new ARR at Anthropic and OpenAI to decelerate.
Even if you want to assume that some non-negligible portion of people who are using Opus and GPT-5.6 today are going to switch to Kimi K3 because it's cheaper and can do their workloads, I think that is completely overwhelmed by the people who still haven't seriously tried this technology.
The people who've kind of tried it a little bit but are every day discovering new use cases—new, cool, high-ROI things they can do with the models—I think all those people are going to be using GPT-5.6 or Opus 5 as a default to unlock these new use cases. You're just not going to see an ARR slowdown or ARR growth slowdown because that isn't going to ramp up so fast.
Max
Yeah, I think we're in agreement on that. Think about how many people there are left to subscribe to this podcast and follow SemiAnalysis then.
Dylan Patel
Dude, it's crazy to me. I went to ICML last week. I went to AI Engineer the week before that. These are normally AI conferences, right? I thought people would be pretty plugged in there.
I would say 80%-plus of people had never heard of SemiAnalysis before. I was like, “Guys, what are we doing, man? You claim to work in AI, but you haven't read and you've never even heard of SemiAnalysis? We're still so early. It's insane.”
Max
That's an ego check, man. Come on, man. You should calm it down a little bit.
Max
Maybe tone our own horn a little bit.
Dylan Patel
Okay, man. I think we could keep talking about this all day, but it's probably good to wrap here. Anything you think was left unsaid? Any burning questions?
Max
If the stock market crashes because all the investors have DeepSeek R1, domain part 2 [?], buy the stocks, guys. Not investment advice, though. Do your own due diligence. Not investment advice.
Dylan Patel
Love it. Let's end it on that. Clip of Max saying anything about stocks, and let's finish with me saying, “Do your own due diligence.”
Good job, man. All right.
Max
Yeah. Cool.