[BidClub_]
Latent Space · · 55 分钟

AIE Europe 复盘 + Agent Labs 论点:Unsupervised Learning × Latent Space 交叉特别篇(2026)

Jacob Effronswyx

播客
TL;DR
  • AI 编程在1年内成为一个数十亿美元市场,swyx认为,顺势下注仍比假设相邻用例必然追上更安全。 他援引的数据是:Anthropic 的 Claude Code ARR 约为25亿美元,OpenAI 估计约20亿美元,Cursor 据传也接近20亿美元。编程的占比从约10%升至50%,所以问题是“为什么不能继续增长”,而不是回归均值。

  • 更可能的市场格局是2家大型编程玩家加一批专业化长尾;应用公司靠聚焦和企业落地能力,而不是永久性的模型优势来构筑防线。 Cursor 和 Cognition 可以继续专注编程,而基础模型实验室则追逐金融、医疗和消费级 Agent 等更大的 TAM。更大的判断是:“2025年是 coding agents 之年,2026年是 coding agents 突破编程边界、去做其他一切的年份。”

  • Harness 工程看起来更接近共识,但更具决定性的 go-to-market 转变是,Agent 本身正在成为客户。 Skills 已经收敛到一个极简组合——一个 Markdown 文件加几段脚本;据称,Vercel 用于配置 Vercel 应用的管理后台,有60%的流量来自机器人。“如果它不存在一个 Agent 能调用的 API,它就不存在”;默认推荐也可能把市场从20个名字压缩到3个。

  • Agent Lab 的打法是把专有工作负载转化为更小的领域模型;即使前沿模型质量持续提升,成本和延迟仍是训练模型的持久理由。 据称,Composer 2 和 Sweet 1.6 在没有补贴的模型切换用户中都位列前5选择,定制搜索模型则展现出更明确的领域价值。替代芯片可能进一步强化这套经济性:swyx 对比了每秒数千 token 和不到100 token 的差距,并认为“每次10倍提升,都会解锁一种不同的使用模式”。

  • 基础模型发布对中型初创公司和低 NPS SaaS 的威胁,可能大于对小团队的威胁;后者即便产品失败,也可能把它变成进入模型实验室的面试作品。 swyx 的亲身测试案例,是用他认为约2000美元就能做出来的产品,替代每年20万美元的活动和赞助商管理软件。真正的约束在组织层面:高管周末做出的“80%解决方案”,可能让其他人负责维护“剩下那堆破事”。

  • 编程前沿已经从“人类写代码归零”推进到“人类审代码归零”,自动化验证成为“暗工厂”能够运转的闸门基础设施。 swyx 预计,一次性软件的规模最终会改善质量,同时警告赢家不会是那些把一切都斥为垃圾产出的犬儒主义者:“这件事不管有没有我都会发生,那就让我们把它往正确方向掰。” Jeremie Harris 对临时性 post-training 的态度有所缓和,因为成果可能3个月后就失效;Jacob Effron 则强调:“结果可以扔掉,但原始数据不能。”

  • 模型规模仍在上升,但记忆、上下文和有现实锚定的空间理解,可能比再刷出一次基准测试提升更难突破。 swyx 已放弃模型规模会封顶在约2万亿参数附近的判断,却称上下文是“扩展速度最慢的因素”:3年间从约4000 token 增至100万 token;Gemini 的百万 token 上下文已经存在2年,却几乎没有得到广泛使用。他借用《心灵捕手》的比喻:主角博学却缺乏人生经验,就像一个“什么都知道、却从未经历过任何事的 LLM”。

摘要 · 为研究而整理的核心内容

1. Harness 正在收敛至最低可行的 Agent 技术栈

  • swyx 在 AIE Europe 给出的信号栈中,OpenClaw 排在首位,其后是 harness engineering 和上下文工程;评测、可观测性、GPU、LLM 基础设施、多模态和生成式媒体构成了长期存在的长尾。

  • Harrison Chase 所说 AI 基础设施终于趋于稳定,并非毫无事实基础:各类集成似乎已经收敛到 Skills——“就是一个 Markdown 文件”,再配上脚本——围绕 LLM、工具、文件系统和检索形成循环。实时运行、子 Agent 和记忆仍可能变化,但 swyx 的态度很务实:“如果它又变了,那就跟着一起变。”

  • 主持人的反驳值得保留:应用公司拥有黏性客户,因此可以不断领先模型重建产品;基础设施公司面对的则是更挑剔的开发者,客户可能彻底流失。因此,swyx 更看好垂直领域的“外包 AI 团队”。横向机会依然存在,但沙盒归根结底只是服务于超大规模工作负载的“另一种算力”。

2. Agent Lab 将用户工作负载转化为模型经济学

  • swyx 的 Agent Lab 路径始于前沿模型,先针对某个领域做专门化,再在使用量带来足够高质量数据后训练专有模型。收益包括更低的成本和延迟,以及一个不那么实质、但确实存在的“营销加分项”:给模型命名并发布研究成果。

  • 他用来证明价值不只是营销的证据是:在不靠补贴吸引用户的模型切换者中,Composer 2 和 Sweet 1.6 据称都位列前5选择。领域专用搜索模型的案例更清晰,而 Thinking Machines 的 Tinker 和 Prime Intellect 的工具链正在降低定制门槛。

  • 主持人仍不太相信仅为追求质量而 DIY RL,回忆起一些领域预训练公司:它们宣称的优势后来随着前沿模型发布而消失。swyx 的回答是,质量和经济性“是同一枚硬币的两面”:在高吞吐、低方差的工作负载上,保持质量不变,或牺牲一点质量,换取成本的大幅下降。

  • 替代硬件改变了 swyx 的看法。Cerebras、Talos 和 MatX 等非 NVIDIA 系统,可以把推理速度从每秒不到100 token 推升到数千 token;Cognition 和 OpenAI 已经在使用 Cerebras。从100提升到200感觉只是渐进式改善,但“每次10倍提升”可能催生没人能够预测的应用,从而支撑一个持续数年的投资周期。

3. Agent 是客户,但不是理性买家

  • 当被问到 Agent 是否是终极理性开发者时,swyx 的回答是:“绝对不是。”Agent 容易受到 prompt 注入,也倾向于进一步强化那些已经嵌入2023年前训练数据的在位者;Vercel CTO Malte Ubl 在会议上给出的数据是,机器人目前生成了 Vercel 用于配置 Vercel 应用的管理后台60%的流量。

  • 眼下的产品规则是绝对的:“如果它不存在一个 Agent 能调用的 API,它就不存在。”面向 Agent 的产品体验,本质上仍然类似于好的开发者体验:文档清晰、API 一致且大多无状态、支持 CLI、提供搜索和渐进式披露;但 swyx 不太相信可以把这套逻辑延伸成通过玩弄聊天机器人来做 AEO。

  • 今天的默认推荐仍然可以制造赢家通吃的分发格局。据称,在约70%的邮件服务商请求中,Claude 会在用户未提示的情况下推荐 Resend,尽管 Resend 成立时间相对较短;swyx 预计最终入围名单会“就3家左右”,而不是20家。“把我的东西和 Vercel 一起用”这类语义关联内容,在记忆和个性化取代单纯的提及频率之前,可能仍然有效。

4. 编程能力竞赛奖励先花钱、后谈效率

  • OpenAI 和 Anthropic 都把编程列为 P0 级项目。swyx 提到,Anthropic 来自 Claude Code 的 ARR 约为25亿美元,OpenAI 估计约20亿美元,Cursor 据传为20亿美元;这些市场大约在1年内形成,而 Claude Code 刚刚迎来首个周年。

  • 流行的追赶论认为,创始人应当瞄准编程之外的空白市场,目前约占 Claude 用例的50%。swyx 更关注顺势增长的问题:如果编程占比能在1年内从可能只有10%升至50%,为什么不能继续上升?错误押注均值回归,代价可能远高于沿着已经确立的轨迹下注。

  • 雇主目前奖励的更多是看得见的 token 消耗,而不是判断力。据称,OBI 的 Ryan Lopopolo 每天消耗10亿 token,按市场 API 价格约合每天1万美元,其中很大一部分可能是垃圾产出;但他会比只运行已知有效工作流的人更早发现新能力。“真正会发现下一个热点的人,都活在边界上。”

  • Anthropic 是高端、受容量限制的玩家;Codex 的潜台词是“进来吧”,Gemini 也在大力补贴。swyx 的战术建议是利用这些补贴,用每月200美元的 Claude Code 或 OpenAI 方案探索能力边界,因为大多数人仍没有把它们压到足够极限。

5. 挑战者创造新体验之前,市场集中仍会持续

  • swyx 的基准情景是2家主要编程玩家加一批较小的专业化长尾。要改变这一结构,挑战者必须拿出显著不同的经济模型、品牌或价值主张;他提到的潜在挑战者包括全面调动 GitHub 的 Microsoft、进军编程的 Mistral,以及 Z.ai、GLM 和 Zhipu 等中国实验室。他表示,新实验室在过去1年尚未真正突围,但提醒听众不要低估它们。

  • 基础模型实验室也在追逐扩张 TAM 的其他方向,包括超级 App、面向金融的 Claude Code 和 Claude Cowork。Cursor 和 Cognition 相对仍是纯编程公司,而专门的 Agent 实验室可以承接只提供极简模型的实验室留下的企业端“最后一公里”;大型企业似乎确实希望由另一家合作伙伴完成这部分工作。

  • 金融和医疗是最明确的下一批垂直领域,但在 swyx 看来,医疗的障碍更多,金融的收入路径更清晰。

  • 主持人认为,Claude Code 的先发黏性说明下一个“神奇界面”可能让另一家公司实现跃迁。swyx 承认 Codex 的 Skills 集成和速度略胜一筹,但仍将其视为追赶者;在这个“高波动、高温度阶段”,他不认为用户忠诚度会达到老牌软件品类的水平。OpenClaw 可能预示着突破口:编程 Agent 生成软件,“软件吞噬世界”(software eats the world),因此编程 Agent 也会吞噬世界。

6. 低 NPS SaaS 承压,但组织惯性拖慢替代

  • swyx 对小型初创公司相对放松,因为做得像样的尝试本身就可能成为进入基础模型实验室的面试作品;他更担心中型公司。LLM 基础设施已经开始整合,Langfuse 等公司被 ClickHouse 吸收,而低 NPS 的传统 SaaS 面对的是更根本的威胁。

  • 他举的内部案例是每年20万美元的活动和赞助商管理软件;在他看来,这套软件定制开发的成本约为2000美元。他预计最终会替换它——可能推迟1年,“但不会推迟5年”——因为技术能力较弱的同事必须接受并运营新系统。

  • 这暴露出企业内部的断层:AI 原生领导者认为同事高估了软件必须包含的功能,团队看到的却是高管在周末交付“80%的解决方案”,再把未完成的工作转嫁给下游。swyx 仍认为 AI 原生的核心记录系统存在机会,或许会类似 Convex,因为除了 Postgres 和 MongoDB 之外,“AI 应用的 Firebase”尚未真正出现。

7. 访问限制只是暂时的,记忆才是顽固约束

  • 在一次与 Anthropic CPO Mike Krieger 的晚餐中,swyx 表示自己最大的担忧是生物安全,而 Krieger 强调的是安全问题。swyx 认为,如果一个所谓私有的模型被40家公司获得、每家公司可能有1万名员工,那么它就并不真正私有:这40家公司的分发边界内部可能存在坏人。

  • Anthropic 的策略是限制访问,并将模型与产品捆绑;OpenAI 在理念上更倾向于广泛开放能力。主持人提出的冷峻反驳是,真正的约束在于算力供应,swyx 对此表示认同:随着未来3—5年更大规模集群上线,目前对超过10万亿参数模型的配给可能会消失,但配给会转移到下一个前沿。

  • swyx 有一个明确标注为未经证实的 Gemini 推测:Google 发布了 Flash、Pro 和 Ultra,却始终没有推出 Ultra,可能只是让它“坐在地下室里”,用于蒸馏出更便宜的模型。这样做是理性的,因为真正的瓶颈是成本,而不是用户是否天然偏好更小的模型。

  • 他已经放弃模型规模会封顶在约2万亿参数附近的看法,但无法判断终点会是200万亿还是2000万亿。上下文的扩展速度慢得多:3年间从约4000 token 增至100万 token;Gemini 的百万 token 上下文已经存在2年,却没有得到广泛使用。“记忆可能会成为最大的限制因素。”

8. 开源模型与一次性训练重新打开技术栈

  • swyx 过去引用 Ankur Goyal 的估计,认为开源模型市场份额为5%且仍在下降;现在他认为份额正在上升。公开能力基准容易被刷,而 OpenRouter 提供了相对更客观的选择数据,不过由于折扣力度很大,观察到的调用量必须按价格进行调整。

  • 平均市场份额掩盖了更关键的用户分层:前20%和后80%的行为完全不同,后者包括普通的 GPT 套壳产品。领先团队正在转向开源模型,Fireworks 和 Together “正在碾压”竞争对手;随着开源模型工作负载增长,微调即服务也可以作为其衍生机会成立。

  • Jeremie Harris 的转变范围更窄:如果 post-training 是改善未来3个月结果的最佳方式,那么即使之后的前沿模型发布会抹平这一优势,post-training 仍然值得做。Jacob Effron 补充说,更好的干预可能是增加数据、连接器或后端工程;这项工作可能相当于临时配置几名工程师,规模甚至可能达到1000万美元。swyx 的关键区分是:“结果可以扔掉,但原始数据不能”(You throw out the results, but you don’t throw out the raw data.)。

  • swyx 建议听众关注长轨迹、合成评测标准和 Dr. GRPO。RL 正在走向多轮交互,轮数可能达到数百轮,从而支持沿着远比1年前浅层 RL 或 SFT 更细的领域维度进行定制。

9. 暗工厂要求软件变得可一次性使用

  • 第一个前沿是“人类写代码归零”,swyx 5个月前在 Cognition 见到这一趋势,如今听起来已经没那么激进。下一个前沿是“人类审代码归零”:Agent 直接提交代码,迫使企业围绕更多测试和自动化验证,反向重构 SDLC。

  • 很少有组织以这种方式运转,尽管 OpenAI 正在探索。swyx 认为,这是释放前所未有软件产量“唯一可扩展的方式”;当软件变得廉价且可一次性使用后,数量最终可以帮助团队搜索质量,而不只是成倍制造垃圾。

  • 他的文化判断很直接:2026年表现最出色的人,不会站在场外把产出一概称为垃圾,而会说:“这件事不管有没有我都会发生,那就让我们把它往正确方向掰。”(This is happening with or without me. Let’s bend this the right way.)

10. 世界模型瞄准的是经验,而不只是机器人

  • 除了记忆,swyx 认为下一个主要前沿是世界模型。它们目前的形态包括3D环境、游戏和具身视觉,因此机器人自然成为主流叙事;但他认为,把世界模型等同于这些产品,低估了真正的问题:模型能否理解物质、物理规律,以及一张桌子究竟是什么。

  • 他推荐 Fei-Fei Li 关于空间智能的文章,因为她可能还没有解决方案,但“她提出了正确的问题”。从 Moonlake、General Intuition 到对抗性世界模型,不同方向的工作都在尝试把智能推进到下一个 token 预测之外。

  • 节目的结尾借用了《心灵捕手》(Good Will Hunting)的比喻:Matt Damon 饰演的角色知道书本能够教会什么,却没有亲身经历 Robin Williams 描述的那些人生经验。swyx 认为,LLM 也存在同样的缺口——“一个非常聪明、什么都知道、却从未经历过任何事的 LLM”。

Shawn Wang

Isn't that crazy? That number is just mind-boggling.

Alessio Fanelli

What is the state of the AI coding wars today?

Shawn Wang

We're in a phase of sort of like capability exploration. The general thesis that I have been pursuing now is that the same way that 2025 was the year of coding agents, 2026 is coding agents breaking containment to do everything else.

Alessio Fanelli

Do you worry about the foundation models just eating into a bunch of these startup categories?

Shawn Wang

Mid-size startups, yes.

Alessio Fanelli

What do you think the end state of this market is?

Shawn Wang

For the market structure to significantly change, there would be—

Alessio Fanelli

Today on Unsupervised Learning, we had a fun episode in what's really become an annual tradition, a crossover episode with our friends at Latent Space. Swyx and I sat down and we talked about everything happening in the AI ecosystem today, what we thought of the various changes at the model layer, what's happening in the infra world, the coding wars, and a bunch of other things. It's a ton of fun to do this with someone I really respect and another great podcaster in the game. So without further ado, here's our episode. Well, Swyx, this is super fun to be back with another Unsupervised Learning–Latent Space crossover episode.

Shawn Wang

Yeah.

Alessio Fanelli

I feel like there are a lot of places we could start, but one thing I always find fascinating about the way you spend your time is that you are obviously at the epicenter of this engineering movement and community. You run these events and conferences, put on these awesome talks, and I think you just have a great pulse on the zeitgeist of what's going on.

Shawn Wang

Yeah.

Alessio Fanelli

Maybe to start, what are the biggest topics people are thinking about right now?

1. The Agent Engineering Consensus

Shawn Wang

Yeah. I just came back from London, where we did AIE Europe, and we're doing roughly 1 per quarter now, which—

Alessio Fanelli

Yeah, you're really upping the—

Shawn Wang

Hopefully is a—

Alessio Fanelli

Upping the pace.

Shawn Wang

We're trying to match AI speed, you know?

Alessio Fanelli

Yeah, exactly.

Shawn Wang

Yeah, that actually—

Alessio Fanelli

The topics will be completely different, I imagine.

Shawn Wang

Yeah, yeah. I definitely curate the tracks. You can see what I think when you see the track lists and the speakers that I invite. Obviously, OpenClaw is the story of the last 4 or 5 months, and just below that, I would consider harness engineering and context engineering to be 2 related topics in agents and RAG.

Then there's a long tail of evergreen stuff like evals, observability, GPUs, LLM infra, just in general. We also have other updates on multimodality and generative media, let's call it. But definitely, the first 3 that I mentioned are top of mind—

Alessio Fanelli

Yeah.

Shawn Wang

—for people.

Alessio Fanelli

I think harnesses, particularly, are so interesting. There was this tweet from Harrison Chase, the LangChain CEO, that caught my eye recently, where he said, “It finally feels like we have stability around the infrastructure for AI.”

I think what he was basically implying is, look, over the past 2 or 3 years, as a company at the epicenter of AI infrastructure, it was a bit like playing Whac-A-Mole, right? You were constantly moving around with however the building patterns were evolving.

Shawn Wang

For Harrison, for sure, right? He's basically had to reinvent the company every year since he started LangChain, right? It was LangChain, LangGraph, and now DB agents. I think he's one of the most nimble, adept, sharp people about this.

Alessio Fanelli

Yeah.

Shawn Wang

But yeah, basically—

Alessio Fanelli

But he's saying now is finally the time—

Shawn Wang

Now this time is different.

Alessio Fanelli

—for stability.

Shawn Wang

Yeah, yeah. Yeah.

Alessio Fanelli

Do you buy that, or what do you make of that take?

Shawn Wang

I think it's very expensive to say “this time is different” sometimes, but when you're just writing code, it's actually okay to try to make a call. I think it may not even matter if this call is right or not. I just don't care that much, because you can be right on the thesis, but if you don't figure out how to monetize the thesis, then who cares if you said something first?

That said, it does feel like, for example, we went through a lot of different ways of packaging integrations up with agents, and it feels like we've landed at Skills, which is the minimal viable format—

Alessio Fanelli

Yeah.

Shawn Wang

—which is just a Markdown file with some scripts attached to it, and I don't see how it can be more simple than that. So there is some justification for the stability around harnesses.

I feel like there may be more adaptation with regard to the real-time elements, subagents, memory, or any of those agent disciplines, let's call it, in agent engineering. But if the thesis is that agents are LLMs with tools in a loop, with a file system where they can do retrieval with Skills and all this standard tooling that now seems to be relatively consensus, then probably that makes sense.

I just think there's no point trying to stake your reputation on this thesis that we're there, because if it changes again, just change with it. It's fine.

Alessio Fanelli

Yeah.

Shawn Wang

Like—

Alessio Fanelli

Yeah, that's always... I've always been struck by how that is much more challenging for infrastructure companies than application companies. Obviously, I think—

Shawn Wang

Yeah.

Alessio Fanelli

On the application side, you've seen Brett Taylor from Sierra and Max Sundstrom from Lagora. They're like, “Look, we build what's ahead of the models, and we're willing to throw everything out every 3 months as the models—

Shawn Wang

Things move fast.

Alessio Fanelli

—get better and better.”

Shawn Wang

Yeah.

Alessio Fanelli

But the thing you at least have there is an end customer, right? That's decently sticky. They will mostly stick. They'll give you a shot, at least, at building these things.

What I've always found more challenging at the kind of “reinvent yourself every 3 months” infrastructure layer is that developers are definitely a pickier audience, maybe, than an accounting firm or a bank.

Shawn Wang

Yeah.

Alessio Fanelli

And so it's definitely a more challenging position to be in, to have to constantly reinvent yourself.

Shawn Wang

Yeah, and when they churn, it's complete. They'll leave for the hot new thing because there's no defensibility, I guess. Even if you are a database, people can migrate workloads off databases. It's a known thing.

I think basically what we're talking about is the vertical-versus-horizontal debate in AI startups. The way I think about it, also, is that when you're Legora, when you're a bridge, you are the outsourced AI team, right? Your job is to apply whatever state-of-the-art AI methods—

Alessio Fanelli

Yeah, like this translation layer between—

Shawn Wang

To—

Alessio Fanelli

—the model capabilities and your end customers.

Shawn Wang

Yeah, to the end customers. And if they didn't have you, they would have to hire in-house, and they're not going to hire in-house, so they have you. I think that's a reasonable, very robust position for whatever trends and discoveries people make in the engineering layer.

I do think there are useful horizontal companies being built, but they're all very much the reinventions of classic cloud in the AI era, with the primary one being sandboxes.

Alessio Fanelli

Yeah.

Shawn Wang

Which—it's another form of compute, guys. Let's not get too excited about it. But I mean, the workloads are enormous.

Alessio Fanelli

Right.

Shawn Wang

Yeah.

2. The Agent Lab Playbook

Alessio Fanelli

It's interesting, and I feel like as part of this, the questions that folks are asking around infrastructure, there's a lot around the extent to which companies should have their own AI teams, what they should be doing in-house, and questions around whether people should be training their own models and whether people should be doing RL in-house based on the data they have.

I feel like one has to evolve their takes on this every 3 months, but where are you at on this today?

Shawn Wang

I think most—I mean, actually, all models have gone up. Obviously, I'm involved in Cognition, and Cursor is doing a lot of its own model training. I think that is some part of what I've been calling the Agent Lab playbook, where you start off with the state-of-the-art models from the big labs and specialize for your domain.

But once you have enough workload and enough high-quality data from your users, then you can obviously train your own models and save a lot on cost and latency and all that good stuff. You also get a marketing bonus from calling it some fancy name and putting out some research.

Alessio Fanelli

From my seat, I can't tell how much of it is actual value that's provided to the end user and how much of it is that marketing bonus, right? It seems some combination of the—

Shawn Wang

I think it's both.

Alessio Fanelli

Yeah.

Shawn Wang

No, there actually is real value, and you know that for a number of reasons. One, even when it's not subsidized, people do choose it as one of the top 4 or 5. This is true of both Composer 2 and Sweet 1.6, among the top 5 models. In a fair market, in a free market—

Alessio Fanelli

Yeah.

Shawn Wang

In a model switcher, people do choose it, and it's not subsidized. That's as good as it gets. Beyond that, domain-specific models—for example, for search, which both companies have—absolutely make a ton of sense. Everyone says, “Yeah, you should always do this.”

Honestly, I think the infrastructure for that is becoming easier with Thinking Machines' Tinker, as well as Prime Intellect's lab stuff. This is one of those reversals of the bitter lesson where you first bootstrap on the large models and the general corpus models to get big, and as you get to very well-defined workloads that are high quantity but not high variance, then you distill down to a smaller model and run that on your own—

Alessio Fanelli

Right.

Shawn Wang

—which totally makes sense.

Alessio Fanelli

What I'm less clear on is the kind of DIY RL use case, which I think is really mostly around improved quality for different things. Obviously—

Shawn Wang

Mm-hmm.

Alessio Fanelli

—there are probably more efficient ways to get a smaller model that's faster and cheaper, and it'll be interesting to see whether a similar story plays out in the RL space. Obviously, 2 or 3 years ago, you had this whole case of companies that were pre-training and claiming better outcomes in their domains, then getting cooked as each model iteration improved. I wonder whether a similar story plays out in the RL space.

The focus is on pure outcomes and quality, not the cost side. Clearly, your own models for cost at scale make a ton of sense.

Shawn Wang

I think they are 2 sides to the same coin. You basically always want to hold quality constant, or trade off a little bit of quality for a drastic decrease in cost, and that's true for everyone. One element I wanted to bring out, which is very much in favor of open models, is custom chips. This would be Cerebras, but also Talos, and then there's a huge range of stuff in between.

Alessio Fanelli

Yeah.

Shawn Wang

This has been a huge story this past year. Everything non-NVIDIA is getting bid up, including freaking MatX, which is very rewarding for me. Suddenly, because the number of alternative hardware options is increasing, the inference you can get is insanely high. We're talking thousands of tokens per second instead of less than 100, so the trade-off for quality doesn't hold as much anymore because the speed is so high.

Alessio Fanelli

Have you seen a lot of companies go all in on the alternative chips?

swyx

Cognition has gone all in on Cerebras—

Yeah.

swyx

—and so has OpenAI. No, I don't think so beyond that. That's—

Alessio Fanelli

But do you think that's more—

swyx

That's mostly because that's a—

—a foreshadowing of what's to come?

swyx

Clearly, yeah. I used to be kind of a skeptic. What if I get my inference speed from 100 tokens per second to 200 tokens per second? It's only 2x faster; it's not that big a deal.

But I think every 10x does unlock a different usage pattern, and we have proof in Taalas and some of the others that you can drastically improve inference speed. What happens from there, I don't even really know. It's so hard to predict when entire applications just appear at once.

Yeah.

swyx

It also isn't that expensive, right? I think the investment cycle is going to be multi-year, and I would caution people not to dismiss it too quickly.

3. Selling Infrastructure To Agents

Alessio Fanelli

Yeah. I mean, one other infrastructure question I was curious to get your thoughts on is that, obviously, it seems increasingly that a lot of the cutting-edge infrastructure companies are building for agents as the buyers or users of their products, right?

swyx

Ooh.

Alessio Fanelli

And I'm trying to figure out what you have to do differently about selling into agents.

swyx

Yeah.

Alessio Fanelli

Are they just the ultimate rational developers, or is there—

swyx

Another huge theme. Yeah.

Yeah. I'm trying to figure out what you have to do differently about selling into agents. Are they just the ultimate rational developers?

swyx

No, absolutely not. I think they are easily prompt-injected and very tuned towards basically compounding existing winners. Congratulations if you won the lottery and got into the training data—

Right.

swyx

—before 2023, because now you're installed in there for the foreseeable future.

One stat that Vercel CTO Malte Ubl dropped at my conference was that 60% of traffic to Vercel's admin app for configuring Vercel applications is bots. It's not—

Alessio Fanelli

Yeah.

swyx

—human. Your primary customer is agents now. It's mostly coding agents, mostly people using the CLI, MCP, whatever. I think step 1 is: if it doesn't exist as an API that agents can use, it doesn't exist.

Alessio Fanelli

Right.

swyx

Which I think is a good hygiene thing anyway—to make everything API-available—but now it's an extra push on product people to not only work on the UI. You should probably work on the CLI stuff.

Beyond that, I think everything that you're trying to do for agent experience now—which is the term that Matt Bowman at Netlify is trying to coin—is the same thing that you should have been doing for developer experience. You should have had good docs. You should have had a consistent API that is mostly stateless. You should have had discoverability, progressive disclosure, search, or whatever.

Now that people have energy around finding these customers and doing that, that's great. Do I believe in extending beyond that into something like AEO for gaming the chatbots? Not necessarily, but obviously there are going to be huge advantages from people who figure out the short-term wins.

Alessio Fanelli

Yeah.

swyx

And short-term wins can compound.

Alessio Fanelli

Do you think these compounding advantages to the pre-training-data-cutoff companies persist? Obviously, over some period of time, I imagine that doesn't persist. As you think about 3 or 4 years from now and what the selection criteria end up being, do you think it still mirrors exactly what you were saying before—that's exactly what you should have been doing all along to sell a good product to developers?

swyx

It could be, except that I think in 3 or 4 years we'll probably have much better memory and personalization. So then general AEO or GEO doesn't really matter as much. Whatever memory or personalization system we end up with will probably determine what you end up choosing much more than what is currently the case, which is just frequency of mentions, let's call it.

Yeah, yeah.

swyx

So you just spam quantity. And I think that's something I'm looking forward to.

I do think the fun mental exercise to work through for yourself is this: If you start a new disruptor company now, there's a big incumbent that everyone knows, like Supabase. Supabase is kind of the Postgres database incumbent. If you want to start a new Supabase, how would you compete with them?

I don't necessarily have the answer, but I do think companies like Resend—which is relatively new; I think they were started in 2023—are encouraging. There was a recent survey where people checked what Claude recommends by default. If you just don't prompt it with anything, just say, “Give me an email provider,” it says Resend in 70% of cases. The fact that you can get in there with such a relatively short existence is encouraging.

Yeah.

swyx

I do think you want to do whatever it takes to get into that very short mentions list, because it's not going to be 20 of them; it's going to be 3.

Alessio Fanelli

No, definitely. It feels like there's probably more consolidation than ever, or a winner-take-most market—more than the physics of go-to-market in the past might have enabled.

swyx

The other thing is that semantic association is going to be very important. You want to do the combo articles where you're like, “Use my thing with Vercel—

With—

swyx

—with blah, blah, blah.” All of that gets picked up in a corpus, so that's probably one thing that you want to do well. I don't know what else.

Shawn Wang

It’s one of those things where I think I feel behind. I don’t know how you feel about this, but—

Alessio Fanelli

I think AI is just everyone constantly feeling like they’re behind—

But yeah, with AI—

Alessio Fanelli

I want to meet the person that doesn’t feel behind.

But with AI—sorry—my stance was exactly what I said before: everything that you should do for agents is something that you should have done for humans anyway.

swyx

Yeah.

Shawn Wang

To the extent that you’re just getting more energy to do things for agents, great. But it’s hard to articulate what new thing, apart from just more spam, you should be doing anyway. That would be my take right now.

I do think there will be more turns at this. I think the personalization turn that is coming will be big, and I don’t know what that looks like. Basically, we feel kind of tapped out on the memory side of things.

swyx

4. The Coding Wars Go Parabolic

I guess since we last chatted, you took this role over at Cognition, and you obviously have a front-row seat to the AI coding space today. I feel like coding, in many ways, people view it as this—I mean, besides being the mother of all markets and this massive opportunity, I think it’s kind of a preview of what’s to come for many other spaces, both—

Yeah.

swyx

I feel like agents are most advanced in coding. I also feel like the competition between foundation models and application companies mirrors what we may see in other spaces. For our listeners, can you just lay out what the state of the AI coding wars is today?

Shawn Wang

It is massive. I don’t think we necessarily appreciated the size of what—

Alessio Fanelli

No.

Alessio Fanelli

I wish we did.

The state of the AI coding wars today is that both OpenAI and Anthropic have made it their P0s to compete in coding. Anthropic is at around $2.5 billion in ARR just from Claude Code. The way they recognize ARR is up for debate. OpenAI—I don’t think the public number is known, but let’s call it $2 billion as well. Cursor is rumored to be at $2 billion, and those are the public numbers that are known.

Huge markets have just been created in the past year. Claude Code just recently celebrated its 1-year anniversary, which is—

swyx

Yeah, it’s crazy.

The other thing that I see is that there are some people who are looking at the relative penetration of Claude use cases. Coding is 50%, and then legal, health, and whatever else make up the remainder. There was a very popular tweet that was like, “Look at the empty space in all these other use cases. If you’re a new founder today, you should be betting on the other stuff,” based on a sort of catch-up theory.

swyx

Yeah.

Shawn Wang

My pushback is the same pushback that I had on Apple versus Google: Why is this time different? If it went from, let’s say, 10% to 50% in the past year, why can’t it keep going? Getting that wrong is actually very painful because you could have just made the momentum bet instead of the mean-reversion bet.

swyx

Yeah.

Shawn Wang

I think that’s the state of things now. People are very much in a psychosis. They’re getting rewarded for spending more rather than spending less, and I don’t think we’re in a phase of efficiency. We’re in a phase of capability exploration. People who are more crazy and more creative get rewarded comparatively.

Alessio Fanelli

Yeah. It’s interesting. Behind these token-maxing leaderboards and whatnot, it feels like the first phase of this transition from a workforce perspective is that you just have to show your employer, “Hey, I use these tools.”

“Here’s the number of tokens I cost.” That’s it. They don’t care about the quality right now. It may be distasteful to someone who cares about the craft and all that, but directionally, everyone just wants you to go up regardless.

It’s not very discerning, and it’s probably very sloppy, but I think it’s net fine because we’re still probably underusing AI just in general. We had Ryan Lopopolo from OBI on the podcast, and he spends 1 billion tokens a day.

For those counting at home, that’s something like $10,000 worth a day of API tokens if they paid market rates. Most of us can’t afford that.

swyx

Yeah.

Ben Tossell

A lot of what he does is probably slop. But he’s going to discover it first. If there were a new capability, he would discover it before you because he was trying and you were not trying.

swyx

Right.

Alessio Fanelli

You only do things that work. Good for you, but the people who are going to discover the next hot thing are living at the edge.

swyx

Right. Increasingly, living at the edge is just having the compute budget to run these experiments. It’s kind of similar to what living at the edge on the research side has always been. It was constrained in many ways by the amount of compute you had to run these experiments. It feels similar on the builder side, or in terms of actually using these tools now.

Mm.

Shawn Wang

The other thing that’s very obvious is that Anthropic is the high-priced premium player, where restricting limits or even restricting model releases is the name of the game. Codex is like, “Come on in, guys. Use our SDK, use our login. We don’t care. We’re going to reset limits,” whatever.

You do want to try to exploit the subsidies where you can get them, and Codex is definitely super-subsidized right now. Gemini is also very subsidized. Comparatively, I think you should make hay while that’s going on.

It’s not that bad to be a capabilities explorer on just the $200-a-month plan from Claude Code or from OpenAI. My sense is that people aren’t even there yet.

Alessio Fanelli

How do you think this market ultimately plays? It’s obviously such a big market that any slice of it is interesting for anyone going after it. But what makes people so interested in the coding market, particularly, is that it feels like a foreshadowing of what will happen in other application markets that the foundation models eventually turn toward, compete against, and gather data around.

How do you think it plays out? Does there end up being room for lots of different kinds of players? What do you think the end state of this market is, and do you think that’s applicable to other markets?

Shawn Wang

I feel like there will be. The status quo is probably the most likely outcome: There are 2 big players and a small range of longer-tail people that fit other use cases the 2 big players don’t. That feels right to me.

For the market structure to significantly change, there would need to be significant change in the economics, the brand-building, or the value propositions of the companies involved. I haven’t seen any in the last 6 months that have really changed the stories materially, so I feel like they will just keep going until something else happens.

Something else happening could mean Microsoft wakes up and goes, “Guys, we have GitHub. We’ll do something much bigger here than just Copilot.” That would be a big change.

Mistral has put out a model now, and I was at a breakfast with Alex Wang where they were like, “Yeah, we really, really want to go after the coding use case.” They haven’t done anything yet, but don’t underestimate them.

Similarly, for the Chinese labs, I think they’re trying to go after it. Z.ai is doing stuff.

swyx

Yeah.

Shawn Wang

GLM, Z.ai, and Zhipu are the same thing. Everyone is trying to get a piece of that pie. I feel like the status quo has been pretty stable for almost a year, I will say.

Alessio Fanelli

Yeah. Is there room for application companies, particularly on the enterprise side? What surface area do the model companies leave for application companies?

Shawn Wang

Yeah, that’s a good one. It’s very much evolving.

Because OpenAI did not have this level of attention on coding a year ago, we just don’t have that much history. It seems like, for example, the big push at OpenAI now is the super app. Is that a consumer thing? Is that a product-portfolio-rationalization thing? How much is that going to take away from attention on coding at a time when they actually do want to put more into coding?

I think it’s very unclear.

swyx

So I do think there are all these areas. At both big labs—sorry, at both OpenAI and Anthropic—and xAI as a separate case, they are trying to see the other TAM expansion areas. So, Claude Code for finance.

Alessio Fanelli

Yeah.

swyx

Claude Cowork, all those things. Whereas I think Cursor and Cognition are comparatively just focused on coding, and so I do think they leave space. I do think for the other verticals that also means the same thing: they’re not going to be that intensely focused on that domain.

Except for—I think I would mark out finance and healthcare as the next ones they’re clearly going after. Comparatively, healthcare seems more thorny. There have been some announcements about it, but I would respect the finance work a lot more just because the path to money is a lot clearer.

Alessio Fanelli

Yeah. No, obviously, maybe similar to the space that’s being left in these other domains, there’s a lot that’s required to actually implement these tools in enterprises versus maybe just giving model access to folks out of the box.

swyx

Yeah, yeah, yeah. So the agent lab thing is, “We’ll do the last mile for you,” whereas I think the model labs tend to just trust the model and be minimalist about it. Both of them work.

Yeah.

swyx

I don’t necessarily think one beats the other for every use case. All I do know is that it does seem like the large enterprises do want a dedicated partner that isn’t just the model labs, which is kind of interesting.

Alessio Fanelli

We’ve been in this phase of pure capability exploration, and so I think nothing has been better for large labs, right? They’re always going to be at the frontier of capability exploration, and so I think they have a very good relationship with a lot of these enterprises. But ultimately, over time, the incentive structure of these labs is always going to be maximal token consumption for the end customers they work with.

There are just so few companies that have actually gotten to massive scale. Maybe coding again is the most interesting, because it’s the first space that really is just completely gone. You must live it every day. Absolutely insane.

swyx

Even—I think we say good things about Cursor and Cognition, but the sheer lift-off of both Anthropic and OpenAI—they have independent valuations—is just mind-boggling. I mean, let’s throw xAI in there. It’s now IPO-ing at $1.2 trillion.

That number is just mind-boggling. I feel like in normal investing or normal startups, there’s kind of a ceiling market cap or valuation that you reach, and you go, “All right, well, it’s going to be chiller from now on.” These guys are not slowing down.

Totally.

Alessio Fanelli

No. Well, I also think the dynamic that’s fascinating about some of these later-stage companies is that, in the past, I feel like in the venture world, if you got to a certain level of scale, the question around you was really more a valuation question. And this is why there were different phases and types of venture people.

The late-stage growth people were just incredible at figuring out a little bit of what the ultimate market opportunity of a company was, but also what the right way to value it was. We know it’s in some band of an outcome that is—sure, there’s some variance to it—but it’s relatively understood what that band is, and then maybe you get, over time, a surprise to the upside.

Whereas any later-stage company—even the labs themselves—the bands in which that company might be worth right now, even in 1 or 2 years, are so massive because of how fast—

swyx

Yeah.

Alessio Fanelli

the ecosystem changes that, even for later-stage companies, every 3 months could be an existential-level event, to the upside or to the downside.

swyx

Yeah. And I think that you’re obviously seeing it in the positive with code, which, if you think about a company like Anthropic, for a while it was unclear if they were going to have access to enough capital to really stay in the race, right? And then coding hit at the exact right time, they had the perfect model for it, they executed brilliantly, and now are one of the most valuable companies in the world.

At the same time, I have zero sympathy for OpenAI because they’re crushing it and they’re all rich. This is a high-class champagne problem to have, to be number 2 at coding or whatever. Who cares? You’re doing great.

5. The Consumer AI Retention Puzzle

Yeah. It’s funny, though. I mean, you would be closer to this, even though you’re in the AI coding space, but a lot of people I talk to think Codex is just as good as, if not better than, Claude Code, right? I think one thing that I’ve been really surprised by—and maybe Claude Code is a better product in some ways; I’m curious about your thoughts—is just in consumer AI.

With ChatGPT, you saw this big first-mover advantage, right? Admittedly, today, I don’t know—Claude and Gemini are great products. It’s not abundantly clear that ChatGPT is any better, but people stick with ChatGPT. It’s the first thing they were introduced to.

swyx

They stay, but they’re not growing anymore. I don’t know if you’ve seen the—

Right, but to me that’s more of a product problem than a market-share problem. It’s not like they’ve lost share to someone else.

swyx

Mm.

Alessio Fanelli

My understanding is the overall problem with consumer AI today is much more: how do you take this tool and, for folks like us—knowledge workers—it’s this incredible magic tool, but it’s not necessarily a daily-active-use tool for a lot of people around the world today?

It’s kind of a category-wide problem. In coding, for example, the entire space has gone parabolic. There may be some relative growth in other consumer AI players, but it’s not like consumer AI as a category is going parabolic and they’re capturing most of that growth.

I think the larger problem is much more, “Hey, the category has kind of hit a bit of a plateau. People haven’t figured out how to bring tons more users on board—

swyx

Yeah, yeah.

or increase the frequency of those users.” And so it seems more of a category-wide problem than it is a massive market-share change. I was going to draw the comparison to—

swyx

Sure, sure.

the coding space, where Claude Code was the first product, obviously, to introduce people to this magical experience. By all accounts, Codex is pretty damn close to as good, if not better. But still, that first product—you would have thought that would not be a super-sticky product surface area—and it actually has the—it turns out, it feels like the first model lab to introduce you to an experience really does keep a lot of the focus.

Shawn Wang

I think maybe it’s still early days. ChatGPT is 3-plus years old, and Claude Code is only 1.

Alessio Fanelli

Yeah. Just turned a year.

Shawn Wang

And so just give it time.

Alessio Fanelli

Yeah.

Shawn Wang

Definitely, a lot of people have switched from Claude Code to Codex. Maybe that will keep going. It’s really hard to tell. I do think that because we are in this high-volatility, high-temperature phase, the loyalty and stickiness to first movers and category creators, I don’t think, is as high as it might be in some other areas in our careers that we’ve looked at.

Alessio Fanelli

Yeah. Though I’ve been surprised by the Claude Code thing. I would have thought that, in many ways—I always worried that the consumer business of these companies would be quite sticky, and then the enterprise API business was actually, in some ways, your least loyal buyers. They would move to—

Shawn Wang

Right, right, right.

But they worked out that it wasn’t the enterprise API; it was the enterprise product.

Alessio Fanelli

Totally. And maybe that was the secret. But the amount of lock-in or just default behavior that has happened in that space is more than I might have imagined with 2 products that, by all accounts, are pretty damn similar.

Shawn Wang

Yeah. No fight there. I will say I do think that Codex is still in catch-up mode, in terms of personal experience. The only things I like about Codex are Spark and the skills integration.

Alessio Fanelli

Yeah.

Shawn Wang

I feel like the skills integration is a little bit better. I feel like the speed is a bit better, maybe because it’s written in Rust or whatever. Very minor things that you almost like telling yourself rather than objectively assessing between the 2 of them. I do think, vibes-wise, that’s what’s going on.

You know, I feel like the missing question in this whole debate is: Why is it so concentrated in only 2 names, right? Where is the Gemini presence? Where is the xAI presence? They are trying; they just haven't made that much progress yet.

Alessio Fanelli

Yeah. But what the Claude Code moment does show—and it actually makes me a little more bullish on the potential for someone else to catch up—is that if you're the first person to introduce some magical net-new product experience, that might actually be stickier than one might have imagined.

Shawn Wang

Right, right, right. Okay, yeah.

Alessio Fanelli

I believe they have a shot at—

Shawn Wang

What do you think that new product experience might be? This is a failure of imagination on my part. People always say, “The thing that will save us is being first to the next new thing.” What is it?

Alessio Fanelli

I don't know. Something around consumer agents and computer use—some kind of hybrid, I think. We're obviously just scratching the surface on the consumer side.

Shawn Wang

Yeah. So my current theory is that OpenClaw is a vision of things to come.

Alessio Fanelli

Totally.

Shawn Wang

It's kind of good that OpenAI has the association with OpenClaw, but by no means do they have the rights to win it. The general thesis that I've been pursuing now is that, in the same way that 2025 was the year of coding agents, 2026 is coding agents breaking containment to do everything else.

Coding agents continue to win because they generate software, and software eats the world. It's kind of the transitive property of “software eats the world”: coding agents eat software, therefore coding agents eat the world. It's an interesting paradigm.

Alessio Fanelli

Breaking containment is always an easier phrase in the consumer context than in the enterprise one. You've seen people run these really cool experiments in their own personal lives, figuring out how to—

Shawn Wang

Yes.

Alessio Fanelli

Obviously, everyone's focused on the enterprise side now, around how you create these experiences. I feel like the vibes—people love to have these narratives that everything has completely shifted—but actually, organizationally, aside from the volatility, OpenAI has great products, a great team, and great models.

Everyone else in the world is incentivized for there to be 2 or 3 more great model companies. Everyone would love more great model companies. I feel like the natural forces of the world revolt when any one company is too much the star of the show, right? There are so many people in the ecosystem who are incentivized for that not to happen.

So I think I'd be shocked if we don't have a reversion of vibes—not completely the other way, but at least a little bit more equal—at some point over the next 6 to 12 months.

Shawn Wang

I think there are just different stages. When you talk about the world wanting more model companies, I think about the new labs.

Alessio Fanelli

Yeah.

Shawn Wang

I don't know—is it fair to say none of them have really broken through in the past year?

Alessio Fanelli

I think that's totally fair.

Shawn Wang

Which is rough. How are we going to grow that diversity in choice? This is it.

Alessio Fanelli

Yeah. It'll be really interesting to see what ends up happening with that. You've seen folks like NVIDIA, which is very incentivized to make sure there's a broader platform of other model providers.

Shawn Wang

I don't know. People say this, but I don't think they tried that hard. NVIDIA tries harder to build new clouds than new labs.

Alessio Fanelli

Yeah.

Shawn Wang

You know?

Alessio Fanelli

Well, they try pretty damn hard to build new clouds, so—

Shawn Wang

Right.

Alessio Fanelli

Yeah.

Shawn Wang

But let's call it the core hyperscalers of the world. They're in a much happier place than any new lab built on top of them.

Alessio Fanelli

Yeah. Though one might argue it's easier to enable a new cloud to be successful than it is to—you can't will a new lab into existence the same way you can with a new cloud.

Shawn Wang

Yeah, yeah, yeah. So NVIDIA has more direct control over it, for sure.

6. The Startup Squeeze

Alessio Fanelli

What else is catching your eye today on the startup side? You worry about this whole narrative that foundation models announce some product and every stock goes down 15%. Do you worry about foundation models eating into a bunch of these startup categories?

Shawn Wang

Not really. There's a point of view of being an investor in startups, and there's a point of view of whether you want to start something. Honestly, the downside for all of these is so minimal, in the sense that the worst you do is just get hired into one of these labs anyway.

I think the market for people who do things, try things, and execute in a competent way—even if it doesn't work out commercially, or even if it just wasn't that great—is still strong. That's your job interview to go to one of these places anyway, so I don't feel that from a very small startup's perspective.

Midsize startups, yes. I would say there's been a lot of LLM infrastructure consolidation, like Langfuse getting absorbed into ClickHouse. People have maybe worked out the domain-specific playbook, and I think that's okay. I'm not that worried about it.

I would be more worried about traditional SaaS, like low-NPS SaaS. This is the whole AI-versus-SaaS debate that's been going on. I'm going through that exact thing in my company, so I'm thinking through this on a very visceral level.

On one hand, you have the people who say, “You vibe coders don't appreciate the amount of work that goes into a CRM.”

Alessio Fanelli

Right.

Shawn Wang

You think you can rip out Salesforce? So did the 30 entrepreneurs before you, right? You classically underestimate the things that you don't deeply know, and you're talking to an audience that's not you.

At the same time, we've never been able to build software so easily or customize software so easily. You're not going to use 90% of the things in Salesforce, so what's the typical—

Alessio Fanelli

So what have you done internally?

Shawn Wang

We have the main SaaS that we use for event management and sponsor management, and we pay $200,000 a year for that. It's not huge, but it's chunky for my scale. I could probably spend $2,000 and build a custom version of that.

The trick has been dealing with the rest of my team and getting them on board.

Alessio Fanelli

Yeah.

Shawn Wang

I'm the most technical person on my team, but I can't make that decision myself. In the same way I've been telling other CEOs and team leaders, you can be super cloud-pilled, you can have LLM psychosis and think that's okay, but you have to bring your team with you.

The widening disparity in LLM psychosis in companies is causing real rifts. On one hand, the people who are less AI-native aren't getting with the picture. They're behind. They're not waking up to the fact that everything they think is necessary isn't actually that necessary. In fact, it would be better for them if they just held their noses, went in, and came out the other side only talking to agents in natural language. Their lives would actually be better, but they're closed-minded.

The other perspective is, “Oh, you vibe coder—you did this in a weekend and got the 80% solution, and now the rest of your employees have to pick up the rest of your shit that you thought you were so hot and amazing at, but actually you didn't figure it out. LLMs are still useless at this, and blah, blah, blah.”

I think there's this huge debate going on in every company right now. I have a small microcosm of it, but it's making me hesitate to pull the trigger. I will at some point. Maybe I put it off for 1 year, but not 5.

SaaS is definitely getting squeezed. It does make me wonder whether there's an opportunity for a more AI-native system of record that's not just Postgres or MongoDB, although both are very good.

Maybe it’s like Convex.

Alessio Fanelli

Yeah.

Shawn Wang

People bring up Convex a lot. I don’t know. I just feel like the so-called “Firebase of AI apps” isn’t really a thing yet, beyond what we have. Which is fine. We could probably start in a more rapid-iteration cycle first before scaling up to something like Postgres or MongoDB, which are more old tech.

I was at a dinner with Mike Krieger, the CPO of Anthropic, and we were just going around the room asking, “What are people most worried about?”

Alessio Fanelli

Yeah.

Shawn Wang

For me, instead of security, I brought up biosafety.

Alessio Fanelli

Yeah. Classic. Yeah.

Shawn Wang

Actually, like I said, it was cliché and classic, and the rest of the table were like, “What do you mean someone sitting at home can manufacture a virus that wipes out half of humanity?”

Alessio Fanelli

It was like the OG Geoffrey Hinton: “This is why you should be scared.”

Shawn Wang

I’m like, “Yeah, read their risk reports; this is the thing.” Mike was just sitting there, knowing he was sitting on myths and saying, “Actually, it’s security.”

I think part of it is very good marketing—too good.

Alessio Fanelli

Yeah.

Shawn Wang

I would actually advise Anthropic to tune down the marketing because it’s also just a very good model, and you don’t have to make so many marketing claims around it. At the same time, it’s not really a private model if you give it to 40 companies, each of which has 10,000 employees or whatever, right? It’s not private. There are bad actors in there.

Alessio Fanelli

Yeah. Hopefully not as bad as releasing it widely. But no, it’s an interesting case study for how many model releases might look from now on. This might be the first model release that looks like the rest of them from now on, right?

Shawn Wang

There’s an overall product strategy for Anthropic of bundling—restrict access, bundle the product with the model, maybe—whereas OpenAI has definitely been a lot more philosophically aligned around, “We will just enable access everywhere, and we don’t know what will come out of it,” right?

Alessio Fanelli

Right. Though the cynical take at this current moment is also just tied to the amount of compute that both companies have.

Shawn Wang

Yeah, right.

Alessio Fanelli

Right?

Shawn Wang

Yeah. I think that’s true. I do think this is the dawn of larger-than-10-trillion-parameter models, which is very interesting. I think it’s a temporary phenomenon because we have much larger compute clusters coming online for everyone over the next 3 to 5 years. This is already written in the cards.

Alessio Fanelli

Yeah.

Shawn Wang

So, to the extent that we ask whether we’ll have rationing of models above 10 trillion parameters in 2 years, I don’t think so. I think everyone will have access.

Alessio Fanelli

No, we’ll just have rationing of the next phase.

Shawn Wang

Right. But that’s almost as it should be.

My classic example—which is just me theorizing, not anything confirmed by Google—is that when Google announced Gemini, they actually announced 3 sizes: Flash, Pro, and Ultra. They never released Ultra. They only have Pro and Flash. My theory is that they have Ultra sitting in a basement, and they’re just distilling from it for Flash and Pro. I actually think that’s how it should be for any lab.

Alessio Fanelli

Yeah, just because those are the models that people actually want to end up using, or is it just cost per input?

Shawn Wang

Yeah, it is cost.

Alessio Fanelli

Yeah.

Shawn Wang

I do think it’s interesting that, for a while, I was considering the theory that models capped out at 2 trillion parameters, and I think that’s proving to be wrong. If I’m wrong, how wrong am I? Do we do 200 trillion? Do we do 2 quadrillion? Whatever.

I don’t think we have the straight answer to that, but it’s interesting that we are continuing to scale the number of parameters when everyone can kind of see that we’re not going to get the next 1,000 or 1 million X from this paradigm. The Elias of the world are working on other model-architecture improvements. We need a different scaling law, I guess, because I feel like people already feel like we’re tapped out on this.

The end state of this is we turn most of the world into data centers, and I don’t know if we want that.

Alessio Fanelli

Yeah. If the return on intelligence were there, maybe it wouldn’t be so bad.

Shawn Wang

I think there’s just a sheer amount of unscalability that’s wrangling people’s sensibilities right now, especially in terms of context lengths. My classic quote is, “Context length is the slowest scaling factor in LLMs.”

Alessio Fanelli

Yeah.

Shawn Wang

We took maybe 3 years to go from a 4,000-token context length to 1 million, and that’s about it.

Alessio Fanelli

Yeah.

Shawn Wang

Gemini has had a million-token context length for 2 years now, and no one’s using it. Memory is probably going to be the biggest limiting constraint on all these things.

Alessio Fanelli

Yep. Certainly seems that way. I guess I’m curious: over the last year since we recorded last, what’s one thing you’ve changed your mind on?

7. The Open Model Reversal

Shawn Wang

I feel like I was kind of bearish on open models last year, in the sense that I had just done the podcast with Ankur Goyal of Braintrust. He has a good cross-section of all the top AI companies, and he said, “Market share of open source is 5% and going down.”

I think that’s changed. I think it’s going up.

Alessio Fanelli

Even though the capability gap does seem to be increasing.

Shawn Wang

Yeah.

Alessio Fanelli

Depending on the timeframe.

Shawn Wang

It’s hard to tell.

Alessio Fanelli

Yeah.

Shawn Wang

You know, it’s really hard to tell because, for listeners, “capability gap increasing” is based on public benchmarks. Let’s say you’re comparing Mistral versus GPT-OSS or GLM-5.1. It’s really hard to tell because even if they were closing, you also wouldn’t believe that they were closing that much because it’s very easy to game the benchmarks.

Alessio Fanelli

Yeah.

Shawn Wang

You just don’t really know. All you know is that there are somewhat objective OpenRouter stats on what people choose in a free market, and people do choose some of these open models in significant volume, except that a lot of them are heavily discounted, so you need to price-adjust these things.

Even if that were true—which I’m not sure about—I feel like the number is just up now instead of down. I think the separation between what the top-tier agent labs are doing versus the average startup in AI or the average GPT wrapper is significant enough that you should not worry about the mean industry number.

You should cohort things into: here’s the median, here’s the bottom 80%, and here’s the top 20%. The top 20% acts very differently from the bottom 80%.

Alessio Fanelli

Totally.

Shawn Wang

The top 20%, which is all I care about, is definitely going toward more open models. Fireworks and Together are crushing it, and so are all the fine-tuners, right?

I think maybe last time we even said things like, “Fine-tuning as a service doesn’t work.” Well, now it’s going to work.

Alessio Fanelli

Yeah.

Shawn Wang

It’s a derivative of the open-model market.

Alessio Fanelli

Well, and also the workload is scaling to the point where people care about cost and speed more and more. That’s moving from just pure use-case discovery—what can these models do?—to, okay, we know what they can do at scale; now let’s do it cheaper and faster.

Shawn Wang

Yeah.

Alessio Fanelli

Yeah.

Shawn Wang

That change is probably the most significant in my mind, and I always like to do the mental math of how I think about scheduling a learning rate. When you’ve been wrong once, what else were you wrong on?

I’m kind of working through it. To me, the other thing was the coding one, which obviously I have now come full 360 on.

Alessio Fanelli

Yeah.

Shawn Wang

But I think people are not appreciating dark factories enough, which I don’t know if you’ve discussed on the pod yet.

Alessio Fanelli

No, I haven’t.

Shawn Wang

This is kind of a strong DHH/Simon Willison term. The general idea is that there are different levels of AI coding psychosis you can have.

The very first level—which, by the way, I encountered first at Cognition 5 months ago—was zero human-written code.

Alessio Fanelli

Yeah.

Shawn Wang

Right? That seems like a reasonable thing now; it was less reasonable 5 months ago. The next frontier that sounds as crazy today as zero coding did in the past is zero human review.

Jeremie Harris

Yeah.

Shawn Wang

You just check it in without even reviewing it. Very few people are doing that, but OpenAI is exploring this. I feel like it's definitely the only scalable way to do this. It just means you have to flip the SDLC or change large amounts of what you normally do, which is probably what you should have done anyway: more testing, more automated verification, or whatever.

But that is a frontier at which, when you unlock that in your companies, you are just going to produce much more software than you've ever had. It's going to be so disposable, so cheap, that you can probably innovate in quality a lot as well. That quantity helps you get to quality.

Jeremie Harris

Yeah.

Shawn Wang

Which I think people are very uncomfortable with because people associate more quantity with slop.

Jeremie Harris

Right. No, it's back to exactly the discussion we were having about the reaction to these token-maxing scoreboards and the idea that today maybe that's not the best sign of product efficiency, but going forward—

Shawn Wang

Yeah. But you still get rewarded for it, so you're like, "Fuck it, whatever." I think the people who are doing well, who will do well, who will do best in 2026 are not the cynics who go, "Oh, that's just slop. I'm not going to participate in that." They're like, "Okay, this is happening with or without me. Let's bend this the right way."

Jeremie Harris

Yeah. No, I love that.

I think for me, a related thing on the open-source model side is that, for so long, I really didn't think it made any sense to do any sort of RL post-training, pre-training, or anything you could do to improve overall quality. Certainly for latency and cost, it always made sense to me. But for overall quality, you just get that for free in the models 3 to 6 months later.

I think what I'm starting to change my tune on a little bit is hearing all these app companies talk about, "We build stuff and then we throw it out 3 months later," as the models improve. You're like, okay, what you're doing for capability improvement is just another version of that, right? I still don't think that your RL or post-training is going to make you have a better model for years and years to come, but I think you still have to be pretty rigorous about whether that's the single best thing you can do to solve a customer problem.

Jacob Efron

Oftentimes, it's literally just, no, add more data and feed more data, even via connectors to these models, or do some clever engineering on the back end, or whatever it is. But if the single best thing you can do for that 3-month time period to improve your customers' outcomes is post-training in some way that really improves the output of the model, even if you throw it out 3 months later because the general models get up there, it still might have been worth doing. I think I'm more open to—

Shawn Wang

You throw out the results, but you don't throw out the raw data.

Alessio Fanelli

Totally.

Shawn Wang

And so—

Alessio Fanelli

Right. Then you just run it again. Basically, there's some—obviously, at a cost level of $10 million, maybe that's too much—but there's some level of cost where—

Shawn Wang

Yeah.

Jacob Efron

—but there's some level of cost where—

Shawn Wang

No, it's not—

Jacob Efron

—even $10 million, right?

Shawn Wang

Yeah.

Jacob Efron

No, of course it's not. There's obviously some level of investment at which it's the equivalent of just staffing 4 engineers to go build something for 3 months.

Shawn Wang

Yeah. The other thing I really, for listeners, am just going to leave some droplets of info. Look into the long-trajectory, synthetic-rubrics work that people are doing. It's very important, including something that's called Dr. GRPO. I'll just leave those key search terms in there.

I think what it means is that RL is going much more multi-turn than people think, and that means that you can customize the models in way more specific dimensions than traditional SFT or this sort of shallow RL that was done a year ago. Hundreds of turns.

Alessio Fanelli

Yeah. What else, of these unanswered questions in AI today, are you looking for in the next year? What are you paying close attention to?

8. Memory And World Models

Shawn Wang

I've got a few theses for what is sort of the next frontier. One is memory, which—memory and personalization—we talked about. The other is really world models, which we've done a small little series on, from Fei-Fei Li all the way to even Moonlake and General Intuition.

There's a lot of debate as to the relative importance of this. I think a lot of it manifests as 3D static walls that you kind of inhabit for a little bit and walk around, and they're cool, but how does this help me with my B2B SaaS?

Jacob Efron

Right. That's like all the hype now is robotics, right?

Shawn Wang

Yeah. There's obviously a correlation between world models and embodied vision and experiences, which leads to robotics. But I think world models are very interesting just in improving intelligence itself, from the next-token-prediction paradigm. I think people are testing their edges around that.

One of our top articles this year so far has been on adversarial world models. I do think that, if you don't do anything else, just read Fei-Fei Li's essay on spatial intelligence, on why LLMs don't have it. She may not have the solution yet, but she has the right problem statement.

Alessio Fanelli

Statement.

Shawn Wang

Everyone else is trying to solve that problem statement in their own way. Let's see who wins. But I don't think it does you any favor to equate world models to robotics, or world models to gaming, or to some kind of current manifestation, because what is at stake is a much more important conception of intelligence than just answering questions.

Does the AI understand what a table is, what matter is, and what physics is? For those who are movie fans, it's like Good Will Hunting, where Matt Damon knows everything because he read it in a book, but he's never lived it.

Alessio Fanelli

Great scene with Robin Williams.

Shawn Wang

Robin Williams. I look at that scene and go, "That's exactly the difference between a very intelligent LLM who knows everything but hasn't experienced anything."

Jacob Efron

Wow. That's an awesome note to end on. That's a great anecdote. Have you used that anecdote? That was great.

Shawn Wang

Yeah. One thing I've done with Latent Space is move to adding daily write-ups. One of the times I was doing this daily write-up, I wrote that.

Jacob Efron

That's a great one. I love that. It's been a ton of fun. Thanks so much for coming on and doing this.

Shawn Wang

All right. Thanks for coming to the show, man.

Jacob Efron

I'm Jacob Efron, and this has been Unsupervised Learning, a podcast where I get to talk to the smartest people in AI and ask them tons of questions about what's happening with models and what it means for businesses in the world. As I hope is clear, I have a ton of fun doing this. It's a nights-and-weekends project in addition to my day job as an investor at Redpoint. But our ability to get these incredible guests on really comes from folks like you subscribing to the podcast, sharing it with friends. It's really what ultimately makes this whole thing work. And so please consider doing that, and thank you so much for your support in listening. We'll see you next episode.