[BidClub_]
SemiAnalysis · · 32 分钟

[紧急节目] Moonshot 的 Kimi K3 来了!中国有了前沿模型

Jordan NanosDoug O'Laughlin

YouTube
TL;DR
  • Max称Kimi K3是全球第3强模型,而Dylan认为就自己的实际使用而言,它可能排第2。 Max表示,基准测试综合结果显示,前3名格局清晰且稳定,其中包括Soul 5.6和Kimi K3,领先于Google、Meta和SpaceX;他称Google应该感到尴尬。K3目前仍更慢,官方口径也落后于Fable 5和GPT 5.6 full,但Dylan发现,在访问被拒或受限时,K3没有Opus那么令人抓狂。
  • K3的2.8万亿参数让推理产能立刻成为瓶颈。 Dylan表示,K3无法装入B200,在单台8路HGX系统上提供服务需要B300、GB300或MI355X级别硬件;跨节点的流水线并行会拖累性能。Max推测,约10天的权重延迟可能是为了给服务团队留出准备时间,也可能让Moonshot有机会与Together、Fireworks、Nebius和CoreWeave安排授权及GB300产能。
  • K3的定价凸显了前沿模型颇具吸引力的单位经济性,但买家基础有限。 价格已从此前2.7 Code模型每100万输入/输出token收费0.95美元/4美元,涨至3美元/15美元。Max认为,如果K3大概率并未亏损运行,同等规模的Fable若收费10美元/50美元,就意味着惊人的利润率,甚至可能优于SaaS。注重成本的用户可能选择GLM-5.2或MiniMax M3,而SemiAnalysis这类用户会继续使用GPT-5.6和Opus;Max怀疑大型企业会认真采用K3。
  • Dylan认为,开源与闭源模型差距收窄,根本原因就是美国限制了用户接触Anthropic最好的模型。 他以Mythos对比Fable,称自己无法使用Opus,只能偶尔使用Sonnet,并认为Soul 5.6可能并非OpenAI训练的最大模型。他预计,政策变化或另一轮前沿突破可能在夏末前后改变局面,同时认为开源模型最终可能达到真正的前沿水平。
  • 西方开源模型在降本之外,还有主权计算机会。 Max认为,市场强烈需要一款“不拉胯”的西方模型,并预计即便是在物理隔离系统中运行,中国模型权重也会被许多美国企业拒绝;他还称,美国可能即将出台中国模型禁令。Dylan表示,西方模型必须同时击败中国绝大多数开源模型,以及前沿实验室的2/3线模型,政府将是更大的市场。K3在SFT期间进行量化,采用MXFP4权重和MXFP8激活值,有助于广泛兼容硬件,也支持中国国产加速器的发展。
  • Harness、路由和尚未释放的用户需求,仍在支撑前沿实验室的收入。 Max表示,在最高、高和中等思考力度下,模型质量在日常使用中越来越难以区分,因此OpenCode、Hermes和Pi等工具的功能成为关键。Dylan认为,按结果定价可能带来95%以上的利润率,Max也认同路由器和可调节的思考力度可以实现这一点。Max认为,编程、音频、视频、深度研究、机器人和世界模型等领域仍有大量新用例空间;Dylan和Max相信,新用户的涌入会压过任何K3替代效应,避免ARR增速放缓。
摘要 · 为研究而整理的核心内容

1. Kimi K3进入前沿模型竞争

  • Max明确认定K3是全球第3强模型。他表示,基准测试综合结果虽然只能提供方向性参考,但仍显示出清晰且稳定的前3名格局,其中包括Soul 5.6和Kimi K3,领先于其他开源模型以及Google、Meta和SpaceX。他称Kimi的成绩非常惊人,并表示Google应该感到尴尬。

  • Max明确给这一排名加了限定:K3整体仍弱于Fable和Soul 5.6。发布博客称,K3与Fable 5及GPT 5.6 full之间仍存在明显的用户体验差距。Max猜测,这种自我降调可能源于中国式谦逊,也可能是为了避免引发美国审查。

  • Dylan表示,K3速度较慢,但能力足够强,自己还没找到多少它处理不了的复杂工作。他称K3是自己实际使用中的第2选择,因为Opus的访问体验可能非常令人沮丧:无论是Web控制台、深度研究还是Coding Plan,他有时会被拒绝或受到限制。不过,使用付费API key后,他表示自己没有再遭遇同样的拒绝。

2. 2.8万亿参数模型让硬件成为发布策略的一部分

  • Dylan将K3描述为全新的基础模型和架构,规模约为此前模型的2倍,并加入Kimi Delta Attention、attention residuals和Stable Latent MoE。其2.8万亿参数无法装入B200;若在单台8路HGX系统上提供服务,需要B300、GB300或MI355X级别硬件。跨节点进行流水线并行虽然可行,但会牺牲性能。

  • Max强调,他对权重延迟约10天的解释纯属推测。一种可能是,服务栈团队需要时间把模型调到足够好的表现,再由Moonshot大规模放出;另一种可能是,Moonshot正在与Together、Fireworks、Nebius和CoreWeave讨论授权和算力容量,其中可能包括部署GB300。

  • 这一规模对比也影响了Max对闭源模型的判断。他表示,如果闭源模型真的拥有约10万亿总参数,却只能达到K3的水平,那就该“收拾包袱走人”,相关公司的股价甚至可能应该下跌50%。但他反而认为,K3大概率并没有比领先闭源模型小多少,甚至可能略大。

3. K3的定价揭示了惊人的利润率,但买家仍不确定

  • Dylan将K3每100万输入/输出token收费3美元/15美元的价格,与此前2.7 Code模型的0.95美元/4美元进行对比,后者价格涨至3倍以上。Max怀疑Moonshot还有多少继续涨价的空间:对普通任务而言,注重成本的用户可能会转向GLM-5.2或MiniMax M3。

  • Max将客户分成两类。SemiAnalysis这类用户愿意在Opus或其他高端模型上大量消耗token;但包括Tesla和Uber在内的一些大型企业,可能把员工每周的token额度限制在约200美元。这些用户更可能选择更便宜的GLM档位。Max表示,即便不考虑部分用户在理念上偏好开源,他也不会对K3在大型企业中几乎没有严肃采用感到意外。

  • 但这套经济账反而强化了闭源实验室的商业逻辑。Max表示,如果Kimi在3美元/15美元的价格下大概率没有亏损,那么同等规模的Fable若收费10美元/50美元,就意味着惊人的利润率。Dylan补充称,主要成本是GPU,而不是员工;Max则认为,token API甚至可能是比SaaS更好的生意。

  • Dylan预计,未来几个月内K3会推出2或3次后训练更新,间隔可能为1或2个月;由于近期不太可能发生能显著提升吞吐量的硬件迁移,价格大致会维持不变。他表示,基于K3打造Composer肯定不会发生,因为Cursor似乎已经决定从头训练自己的模型。

4. 政策正在压缩前沿差距,也在推动开源模型国家化

  • Dylan表示,当前开源与闭源模型的差距之所以收窄,根本就是因为美国对Anthropic的限制让用户拿不到实验室最好的模型。他以Mythos对比Fable,称自己无法使用Opus,只能偶尔使用Sonnet,并认为Soul 5.6不是OpenAI训练的最大模型。他预计,政策变化或另一轮前沿突破会改变这一局面,时间可能就在夏末前后。

  • Max表示,市场对一款不拉胯的西方开源模型有巨大需求。他对美国没有一家公司的能力至少达到中国第5强公司的水平感到意外。他预计,即便有人认为模型权重可以在物理隔离的数据中心内运行,许多美国企业仍会拒绝中国模型;美国全面禁止中国开源模型,可能只是时间问题。

  • Dylan表示,西方模型必须同时击败中国绝大多数开源模型,以及前沿实验室推出的2/3线模型。低成本应用已经可以通过Bedrock或Foundry等服务调用这些接近前沿的模型,因此他认为,西方开源模型更大的市场是政府,而不只是企业的降本需求。

  • Dylan和Max将开放权重视为主权计算战略。Max表示,Xi鼓励中国公司保持模型开放,是因为中国政府部门希望下载权重、部署到自有服务器上,并带动国内生态。Dylan认为,美国政府也应该采取同样务实的做法。

  • Dylan指出,K3的博客介绍了SFT期间的量化方案,使用原生MXFP4权重和MXFP8激活值,以实现广泛的硬件兼容性。Jordan列举了包括Huawei Ascend、Baidu、Kunlun Haxen和Moore Threads在内的一系列中国加速器;随着中国优先推动在国产芯片上运行前沿模型,这种兼容性变得更加重要。

5. 工具编排层、路由与未释放的需求仍支撑前沿实验室收入

  • Max表示,在日常使用中,已经越来越难区分最高思考力度下的绝对前沿模型,与高或中等思考力度模式之间的差异。因此,测试K3时必须考察周边产品——OpenCode、Hermes和Pi——而不只是模型本身。

  • 工具编排层的一些小功能会影响Max把token发给哪个模型:例如能否在远程SSH服务器上安装工具、能否使用顺手的快捷键,或能否编辑此前的命令。他还表示,Slack机器人和Perplexity Computer等产品可以在K3、GLM、Sonnet和OpenAI模型之间路由任务,用户并不在意底层到底是哪一个模型;对于这些首轮任务而言,工具编排层的质量更重要。

  • Dylan由此提出按结果定价的逻辑。路由器可以把简单任务分配给更便宜的模型,把最高思考力度留给更难的工作,从而有可能实现95%以上的利润率。Max认同,动态调节思考力度可以捕获其中大部分机会。

  • 针对Dylan提出的担忧——政策强制形成的模型同等水平可能摧毁Anthropic和OpenAI的定价权——Max的答案是否定的。他认为,这些实验室仍可以继续在内部训练更强模型,在不对外发布的情况下推进coding RSI,并探索视频生成、audio-to-audio、深度研究、机器人和世界模型。他的概括是,所谓“中产阶级模型”(“bourgeois”)可以继续彼此训练,同时只向公众释放少量能力。

  • Max表示,他身边的科技从业者使用AI的频率大约只有自己的1/10,而自己可能处于用户群前10%、前1%甚至更小的分位。随着这些用户转向更大模型,以及非技术用户发现新的应用场景,他认为前方可能仍有1000倍的需求增长空间。Dylan认同,现有Opus或GPT-5.6用户转向K3带来的迁移,都会被那些几乎还没尝试过这项技术的人淹没,因此K3不应导致Anthropic或OpenAI的ARR增速放缓。

  • Max最后开玩笑说,如果市场因为所有人都拥有一个“DeepSeek R1……第2部”而崩盘,大家就应该买入这些股票,同时补充这不是投资建议。Dylan最后说:“请自行尽调”(“Do your own due diligence”)。

Dylan Patel

All right, Max. We're going to do a podcast. We're going to talk about everything related to Kimi K3 and maybe some other models that just came out. How are you doing?

Max

Doing great. Looking forward to it, and thanks for having me, Jordan.

Dylan Patel

I'm not having you. I'll say thanks for having me. All right, on the docket: Kimi K3 hot takes. Is this the 3rd-best model in the world? What's the impact on OpenAI and Anthropic, architecture changes, and personal usage that we've had so far? What do we think about their open-source strategy, and maybe more? All right, Max, quick hot take: Is this the 3rd-best model in the world right now?

Max

I think the answer is a clear yes. People love shitting on benchmarks. I think benchmarks definitely have their problems, but if you take a composite of all the main benchmarks and look at all their rankings, they've been directionally correct over time.

I think if you look at that composite today, there's a pretty clear top 3 with stable, Soul 5.6, and now Kimmy K3. They're always above everyone else, which includes, of course, other open-source models like Deep Sea, Conjure, and whoever. It also notably includes Google and Meta and SpaceX.

I think it's honestly a truly impressive and very remarkable feat from the Moonshot guys. Google, in particular, should feel incredibly embarrassed right now. At one point, as recently as November or December 2025, everyone thought the clear AI big 3 were Google, Anthropic, and OpenAI. Even when I talk to boomers today, they still seem to think the top 3 are Google, Anthropic, and OpenAI. That's clearly not the case anymore.

I'd say it's definitely the 3rd-best model in the world. I do think it's overall still worse than Fable and Soul 5.6. It's kind of funny that they explicitly said that in their model-release blog post. Maybe it's some old-fashioned Chinese humility. Maybe they don't want to incur scrutiny from the U.S. government or anything, because obviously there were some delays with the Fable 1.56 release. Overall, though, I'm very impressed with the model.

Dylan Patel

Yeah, in the limitations section of the blog post, they said, “Despite being a highly competitive model, overall K3 nonetheless exhibits a noticeable gap in user experience compared with Fable 5 and GPT 5.6 full.”

My experience using this personally is that it is good. It's really slow, which is really annoying. It's motivated me to try open-source harnesses for the first time, and I feel like I'm learning more about the harnesses than I am about the models because, frankly, all these models are good enough to do the basic work that I've been doing so far. I can't really find a lot of complicated stuff that it can't do, which in and of itself is a bit of a feat.

Here's my hot take: For me, this might be the 2nd-best model in the world right now, because every time I try to do something meaningful with Opus, I get rejected and sent down to Opus. Even though I don't know if this is better than Opus, it is less annoying not to get rejected whenever I'm trying to do something.

However, I'm not getting rejected when I use my API key and pay for tokens, but I am hitting limits whenever I try to use the web console, deep research, or the coding plan. I haven't used the coding plan, but some other guys at Anthropic have. That leads me to ask: What's the strategy here? These guys clearly don't have enough GPUs to serve the demand they're seeing for this model.

Previously, that was solved by an open-source strategy where they just dropped the weights, and then other people served the model and served that demand. But they haven't dropped the weights yet, so I think they said the weights would be available in 10 days or something.

Max

Yeah, I think that's what they said.

Dylan Patel

What do you think the strategy is for the delay between the announcement of the model, the API being available, and the weights not being available yet?

Max

To be clear, this is all pure speculation on my part. I think one big reason is that they need to give the LLM and SU lane guys enough time to make sure they can serve this model performantly. If they just dropped it today, you'd have all this hype, but then everyone else serving the model would be giving you 22 seconds or something. That's probably really bad for the brand.

They have this incredible opportunity to get a bunch of huge PR and adoption, and they also need capital a little bit. I think another possibility is that they're actively talking to Together, Fireworks, Nebius, and CoreWeave to figure out how they can sign some sort of licensing deal, have them serve the model, and develop capacity on GB300s or whatever.

In my mind, those are the 2 main reasons why you would wait 10 days to actually drop the model weights.

Dylan Patel

Yeah, that makes sense. Functionally, this is really interesting because, just to talk about the model architecture for a second, it's 2.8 trillion parameters. This does not fit in a B200, so you need to have a B300, GB300, or, I guess, MI355X in order to serve this model on a single system, like a single 8-way HGX server.

Of course, you can do unique things where you have pipeline parallelism across multiple nodes and stuff, but that's going to really impact performance. I think there are a lot of recipes being cooked up, and only people who have the latest and greatest chips are going to be able to serve this model.

Let's go back to your comment on Google for a second. The idea that this model is truly competitive at the frontier with 2.8 trillion parameters gives us some insight into how big the closed-source frontier models are, right? It would be even more embarrassing if they're hitting these levels of performance and being compared to 10-trillion-parameter models with a lot more active parameters, right? We have to assume that this is in the same range as what Sonnet and Opus are, right?

Max

Yeah, I think that's a great point, and you just have to be correct. I still believe in the competence and correctness of all the OpenAI and Anthropic researchers. If there are some people on Twitter who like to claim that current closed-source models have 10 trillion total parameters or something, if that's actually true, guys, it's time to pack up the bags. The stock price probably should crash 50% tomorrow. It's over.

I'm pretty confident that Kimi K3 cannot be much smaller. If anything, it might even be slightly bigger than the leading closed-source models today. If that's true, this further highlights the point we've been harping on for a while at SemiAnalysis: The margins for these closed-source labs have to be absolutely mind-boggling.

If you're telling me Kimi's probably not operating at negative margins and they're serving K3 at $3/$15 per million output tokens, that's sort of the same price as Sonic. If you're telling me that Fable is probably similarly sized and the profit can charge $10/$50 per million output tokens, then this should immediately dispel any remaining concerns people have about the AI labs being unprofitable businesses.

Selling tokens at API prices might be even better than SaaS, honestly. It's an incredible business today.

Dylan Patel

Yeah, that makes sense. No cost of employees, just the GPUs. Can you compare this pricing strategy to the previous stuff? You said it's at $3/$15. The previous version from Moonshot directly was at $0.95 and $4, so we're talking about—

Max

Yeah.

Dylan Patel

—more than a 3× pricing increase from 2.7 code to Kimmy K3. Do they have even more room to increase pricing? What's the curve to get to frontier open-source intelligence, or frontier soon-to-be-open-weight intelligence, with the license yet to be determined?

Max

Honestly, I don't think they have that much more room to push pricing up. I would guess that even at $3/$15, there will be a lot of people who say, “This is a little too expensive for me. My task is easy enough for a GLM-5.2 or a MiniMax M3, and I might just use one of those models instead.”

On one end, you have the SemiAnalysis of the world, right? We don't really care how much money we're costing Dylan when we burn tokens all day. We're very happy using Opus for even a relatively easy task that we're pretty confident one of these other open-source models can do pretty well.

On the other end, you have people who are extremely cost-conscious. You may only get $200 worth of tokens per week, as you've heard some large companies like Tesla and Uber are implementing. Pretty much all the people in that second bucket are going to want to use the GLM-kind-of-pricing-tier models because they're already good enough for most everyday tasks. Everyone in the SemiAnalysis bucket is still using GPT-5.6, too, and Opus.

I think there actually is a pretty interesting question of who the user is that will actually be switching to Kimi K3. It might just be a lot of people who philosophically love open source and are excited to try this new hype model and support it.

Max

But it wouldn't surprise me at all if there isn't serious adoption of this model among, say, large enterprises.

Dylan Patel

Okay. What do you think about where we go from here? This is obviously a new base model and a completely new architecture for these guys: 2.8 trillion parameters. They've got Kimi Delta Attention, attention residuals, and the Stable Latent MoE that they keep using. It's a scaled-up, bigger version of the previous models—clearly about 2 times bigger.

Max

Mhm.

Dylan Patel

Previously, with Kimi K2.5, we saw Cursor train Composer based on

Jordan

Mhm.

Dylan Patel

just continued pre-training, as well as some RL. Then we saw Kimi give us K2.5, K2.6, and K2.7 checkpoints as they continued the RL. This is a new base model, and it seems pretty complete. In my usage, it's working pretty well. It's not screwing up anything basic when it comes to writing a PR description or totally going off the rails, the way we've seen some other models that are raw without a bunch of RL have rough edges at the beginning.

So where do we go from here? When does Kimi K3.1 come out? How does pricing change over time? Do we get a Composer based on Kimi K3? I mean, a Composer based on Kimi K3 is definitely not happening because I think the Cursor guys are pretty set on training their own model from scratch now.

As for when Kimi K3.1, Kimi K3.2, or whatever comes out, I imagine we'll probably see 2 or 3 updates within the next few months, each a month or 2 apart, as they continue post-training this thing. I would guess pricing stays about the same, just because they're not going to be able to run it on new hardware in the next 2 or 3 months. They're not going to get a huge throughput increase there to reduce pricing.

Maybe it's possible that some really crack engineers figure out how to reduce the cost to serve this thing so it's closer to DFC V4 pricing or something. That would be really impressive, but given that it's a 3 trillion parameter model, I'm a little skeptical. I would guess that the current pricing we see for the MiniMaxes and the GLMs is already pushing the limits of what you can charge to serve a 1T-to-1.5T model without having embarrassingly bad margins.

So I think this pricing is probably here to stay for at least the next few months. I think the most interesting question is whether the open-versus-closed gap is going to continue shrinking, and whether open models will ever fully match closed-source models with true frontier-level parity. I'm curious what your thoughts are there, Jordan. I think it has serious implications for our whole industry if it actually happens.

Dylan Patel

Yeah, I mean, my view is that I believe the reason this gap has closed right now is squarely due to the US government imposing restrictions on Anthropic, resulting in us not getting the actual best models that these guys have. They've artificially caught up, basically.

Max

Interesting.

Dylan Patel

Clearly, we see this with Mythos versus Fable. I can't use Opus; I can only use Sonnet sometimes if I ask it nicely. 5.6 Soul, I think our host view is that it's not the biggest model OpenAI has ever trained. It's not the size of GPT-4.5. To me, they have a bigger model somewhere.

I think the result is that we're only going to be able to access frontier intelligence if government entities allow us to. That's a very interesting change to the setup going forward because I think it represents an opportunity for many of the players that are in 4th, 5th, 6th, or 7th place to catch up to a limit, at which point it's okay to release everything and start battling for user share without really being able to find the frontiers and have the frontier dominate.

I think it's possible that we see the frontier take another big step toward the end of the summer. It's possible the politics change a little bit.

Jordan

Yeah.

Dylan Patel

It's possible that we start to find other modalities beyond coding where these guys can really improve, and they start exploring those areas. We didn't intend to talk about this right away, but I loved the release of Inkling by Thinking Machines. I thought the native audio input would be super interesting and super useful in the future, and kind of a sign of what's to come. But anyway, yeah, I—

Max

And on the topic of Inkling, there's definitely huge demand for a Western open-source model that doesn't suck. I'm shocked that markets are still so inefficient and that we haven't had a single American company that's at least on par with the 5th-best Chinese company.

One, it's only a matter of time until the US government bans Chinese open-source models entirely. Maybe that's a can of worms we don't have to go down in this conversation. But even if that doesn't happen, I feel like the average large American enterprise is simply unwilling to put all of its proprietary data through a Chinese open-source model, even though you can make tautological arguments like, “You're just loading their weights in your air-gapped data center. There's no way the CCP is actually going to see any of your data.”

I don't think the executives will actually buy that, and I don't think they really care. There are a lot of people who, A, care about token budgeting and, B, are only interested in running a Western model or a non-Chinese model. It's shocking to me that we're not actually closer to the open-source frontier in America.

Dylan Patel

Yeah, I mean, there was NVIDIA Nemotron, and then there was Inkling. It's really inspiring to see Tinker go for it. I think they have 2 business opportunities there. They've got to be better than the bulk of Chinese open-source models. They have to be in the game there to be considered.

Max

Yeah.

Dylan Patel

But then they also need to be better than Sauna or better than Terra Luna—the tier-2 and tier-3 models from the frontier labs. You can build a bunch of cheap applications using close-to-frontier intelligence with Bedrock or Foundry or whatever, get access to the Anthropic or OpenAI models, and save money by going with their 2nd-best model.

I never really understand the Western open-source angle of saving people money. I think it is real, and getting those models into the ecosystem of companies like Fireworks, Together, and Baseten, and anybody who's serving open source, is a good thing because it is a market. But to me, the bulk of the market is government.

One interesting view on the Chinese models is that Xi has been encouraging the Chinese companies to keep the models open source. That is the view from their party. I think the big reason for that is that a bunch of the Chinese government wants to download the weights and run them on servers that they own, and they want the support of the local ecosystem.

I think the American government should work the exact same way. That's a pretty pragmatic strategy: you need to give the people in your country access and support to run this stuff. Maybe the other thing worth commenting on is that, in the Kimi K3 blog, they mentioned post-training—sorry, quantization during the SFT stage. They were commenting on natively using MXFP4 and MXFP8 weights and activations, respectively, for broad hardware compatibility.

Well, what other hardware do you think Moonshot cares about, Jordan?

Jordan

I've got a list of 11 Chinese accelerators. Huawei Ascend, Baidu, Kunlun Haxen, and the Moore Threads guys—there are all sorts of different chips showing up in papers. We're seeing code. It's a national priority for China to get these frontier models—these are frontier models now—running on their domestic accelerators.

Dylan Patel

Yeah, I mean, if we're calling Google a frontier lab, in 2025 we have to call—

Max

Call Moonshot a frontier lab now.

Dylan Patel

Kind of crazy, dude.

Max

It's like vanity sizing.

Dylan Patel

I'm still a 34 waist.

Max

Yeah, yeah. And so are the 7 other Chinese labs.

Dylan Patel

Yeah. No, no. Funny enough, my dad is actually visiting China right now, and he's telling me that the hotel he's currently staying at is totally booked because Xi Jinping is going to be in the area soon and is going to give a speech about how AI is a top priority for China.

Circling back to what you said earlier about the US government and how, if they keep kneecapping the frontier models OpenAI and Anthropic have—forcing them to delay them, forcing them to only have their 2nd-best model publicly available—and therefore giving all the other players, the Googles, the Space X's, the Metas, whoever, time to catch up, do you think that completely destroys the frontier-lab business model?

If you're OpenAI or Anthropic, you just lose all pricing power at that point, right? I don't see how Anthropic can still accelerate net-new ARR if its model is on par or comparable with the Meta model, the xAI model, the Google model, the Moonshot model, and the DeepSeek model. What happens to our industry at that point, Jordan?

Max

Yeah, I mean, first of all, no, I don't think that's going to happen, and I think I can explain why. But first of all, I don't know for sure, so we'll have to see it play out.

Max

Interesting to think about. I think the biggest thing that I've realized in my personal usage of this stuff is, one, how hard it's getting to differentiate between using the absolute frontier model and the max thinking mode versus high versus medium effort on those models.

Dylan Patel

Yeah.

Max

It's really, really hard for me to find day-to-day tasks that these models can't figure out. My behavior defaults to the biggest and hardest thinking because I don't care about Dylan's budget. But when it comes to actually using this, there is an aspect of the hardest being part of the product. So, testing Kimi K3 requires me to take a serious look at OpenCode, Hermes, and Pi.

The hardest is totally part of the product still. Simple things can cause me to want to use one model over the other. Can I install it on my remote SSH server? How easy are the keystrokes to get stuff in? Can I edit previous commands? These little tiny features in the harness actually impact where I'm going to send my tokens, which results in where I'm going to send my budget, right?

Dylan Patel

That's an interesting point because I think a lot of people talk about the token machine, right? Correct me if I'm wrong, but what I'm hearing from your description of your own workflow is that even for tasks where I'm pretty confident that a GLM could successfully do it, I'm happy routing it to Opus and doing it on max intelligence because the ROI of that task is still worth the Opus price to me.

There's always going to be some risk in the back of your mind where it's like, if I use GLM instead, even on medium thinking mode, it's way cheaper. Maybe it's not actually as high quality as Opus would have been, right? So even if the benchmarks claim that a lot of your tasks can move to GLM, you're still fine keeping them on Anthropic models or OpenAI models for the foreseeable future.

Max

Mostly, yes, but I use a lot of Slack bots right now. I actually don't know what model is running behind the scenes on those Slack bots. Specifically, in Perplexity Computer, I think if it starts routing it to Kimi K3, if it starts routing it to GLM, or if it starts routing it to Sonnet—and I know it's doing it today because I looked at my usage a few weeks ago and found how much of the OpenAI models I was using because it was making that decision—I don't really care which model they're using, right?

For a first cut at a PR before I go in and actually fix some stuff up, I don't really care which model they're using. That is about the quality of the harness there for what I'm using.

Dylan Patel

That might actually be a pitch for outcome-based pricing, if anything. One of these labs could potentially just get 95%-plus margins if they do outcome-based pricing for you because, as you said, all these tasks you're happy to pay even stable pricing for could probably get done at a fraction of the price even today.

Max

Yeah, yeah, 100%. Certainly with dialing in the thinking mode, which is where a ton of the expense ends up going, I can totally imagine them building a router.

The second thing, though, just on the competition thing you said earlier, is I don't think we're out of use cases or ideas for these guys to work on. I think they can continue to train incredible models to try and hit RSI on the coding side without ever releasing it to us, the proletariat, and keep their bourgeois models training each other. They can keep distilling them and giving us little tastes of it while still pursuing a research objective that includes all sorts of other uses of AI.

We're really exploring coding right now, but we do some video generation stuff. We do a lot of audio-to-audio stuff. We do lots of deep research that really doesn't look like coding in some ways. I think there are lots of use cases that they can continue to explore without really encountering the cybersecurity issues.

Robotics and world models are a simple one, right? What if Anthropic sets its sights on automating away a whole bunch of manual labor jobs instead of knowledge-work jobs? The idea that there's no way for them to build a sustainable business with great ROI for their—

Dylan Patel

Mhm.

Max

—greatest technology the world's ever seen, I don't believe that at all.

Dylan Patel

Yeah, it doesn't pass the smell test.

Max

No, not at all. But even beyond that, I use these models so much every day. First of all, I see how much my friends who work in technology, who are software engineers, spend 10 times less than me and use them 10 times less right now.

One person using Opus is like, you know, a person using Sonnet and a person using Opus. They use it both the same amount on the same day, but the person using Opus spends 10 times more at 90% margins. They make up the bulk of the business, right?

As soon as those people—of which I would say there's maybe, maybe I'm in the top 10%, maybe even the 1% of the industry—start using the bigger models, they use them more. That's just more demand for all of this business.

The models don't even need to get any better for them to discover they can use them for the really important tasks or the bigger ideas that they have. Then I need to go and talk to my neighbors who don't work in technology. There, I'm certainly in the 1%, probably in the 0.1%, maybe 0.01%, and maybe we've got a thousand times more to go from here.

So I get back to Masason's golden-goose exponential chart to the right, sort of this point. It's an exponential.

Dylan Patel

Okay, just to take this totally off the rails. But actually, before we go there, I want to say that I think what you said about being early is totally right. This is exactly why I don't think Kimi K3 is going to cause net-new ARR at Anthropic and OpenAI to decelerate.

Even if you want to assume that some non-negligible portion of people who are using Opus and GPT-5.6 today are going to switch to Kimi K3 because it's cheaper and can do their workloads, I think that is completely overwhelmed by the people who still haven't seriously tried this technology.

The people who've kind of tried it a little bit but are every day discovering new use cases—new, cool, high-ROI things they can do with the models—I think all those people are going to be using GPT-5.6 or Opus 5 as a default to unlock these new use cases. You're just not going to see an ARR slowdown or ARR growth slowdown because that isn't going to ramp up so fast.

Max

Yeah, I think we're in agreement on that. Think about how many people there are left to subscribe to this podcast and follow SemiAnalysis then.

Dylan Patel

Dude, it's crazy to me. I went to ICML last week. I went to AI Engineer the week before that. These are normally AI conferences, right? I thought people would be pretty plugged in there.

I would say 80%-plus of people had never heard of SemiAnalysis before. I was like, “Guys, what are we doing, man? You claim to work in AI, but you haven't read and you've never even heard of SemiAnalysis? We're still so early. It's insane.”

Max

That's an ego check, man. Come on, man. You should calm it down a little bit.

Max

Maybe tone our own horn a little bit.

Dylan Patel

Okay, man. I think we could keep talking about this all day, but it's probably good to wrap here. Anything you think was left unsaid? Any burning questions?

Max

If the stock market crashes because all the investors have DeepSeek R1, domain part 2 [?], buy the stocks, guys. Not investment advice, though. Do your own due diligence. Not investment advice.

Dylan Patel

Love it. Let's end it on that. Clip of Max saying anything about stocks, and let's finish with me saying, “Do your own due diligence.”

Good job, man. All right.

Max

Yeah. Cool.