[BidClub_]
The a16z Show · · 36 分钟

AI 的现状:模型、护城河与消费领域的复兴

Anish AcharyaJen Kha

YouTube
TL;DR
  • Anish Acharya 明确站在前沿实验室“多赢家”一侧。 过去2周,xAI“从模型端甚至算不上真正的竞争者,变成了3家之一”——双雄竞争变成三足鼎立;与此同时,Anthropic 看似占据的主导地位让位于 OpenAI 表现出色的3个月(新模型、Codex harness,以及 ChatGPT 桌面应用),多家实验室在彼此取得成功的同时都在增长。Jen 补充了可交易的观察框架:X 是不完美的风向标,但能提前反映开发者情绪;Claude 正面临 token 使用量的反弹,而 Anthropic 将于今年晚些时候上市。
  • 市场讨论不足的宏观情景不是泡沫,而是“如果我们的乐观程度还不够呢”。 B200 每小时价格正在上涨,尽管它并非最前沿的 GPU,而算力通常具有通缩属性;这“说明供给极度受限,而需求实际上无限”。SaaS 方面,2月的30–40%回撤属于过度抛售,不少公司已经反弹40%,但随着 SBC 扭曲的经济性开始显现,行业进入“加速,否则出局”。
  • 多数护城河能够穿越低成本智能充裕的时代,但整合护城河正面临风险。 网络效应、规模/分销和品牌效应“处于有史以来最好的状态”(“再多 coding agents 也不会让 Nike 不再是 Nike”);与此同时,coding agents 让 SAP 式整合能力大幅提升,也让 SI 和 GSI 面临一个“生死攸关的问题”。
  • Token 支出会理性地按上行空间是否有边界进行分配。 对销售和产品而言,“为一个哪怕只聪明1个 IQ 点的模型支付几乎任何价格,在经济上都是理性的”——“你的 Fable 5、Gro 或 GPT56”;对财务而言,“你不可能把结账做得比准确好10倍”,因此,开放权重模型加强化学习可能是帕累托高效的选择。模型不是大宗商品:神经质、字面化的 GLM 5.2/5.3,与开放、擅自推断的 Kimi K3 迥异;组织需要这两类头脑。
  • 实验室正在向下垂直整合至推理,而不是向上整合到应用层——这逆转了2025年初的恐慌。 Anthropic 的法律“插件”(其实只是长提示词文件的集合)引发了一轮恐慌,Thomson Reuters 等法律股随之下跌;但推理工作负载同质化且可规模化,应用层则高度特异、运营开支沉重;而在多模型构成帕累托前沿的世界里,实验室更难拿走“你全部的毛利率”。
  • 消费者的季度可能终于到来,但“我们还处在 AI 的 DOS 时代”,正等待属于自己的 Windows。 开放权重模型正在让 AI 软件变得显著更便宜、性能更强(Jen 自己的 X 时间线应用接入1名新用户的成本是250美元);市场仍没有 AI 原生应用商店,消费者却很有兴趣尝试并付费购买新软件。Anish 将这一时刻比作2009年圣诞节前后的 iPhone,只是如今消费者每月可能会支付200美元。
  • 投资姿态已经转向真实上线的产品和技术型创始人。 如今,“无论处于哪个阶段,融资路演中没有展示真实上线产品,都是直接出局的理由”;创始人群体变得“少一些 MBA,多一些研究人员”;而根据 Ben 在 offsite 上的说法,“过去最大的风险是想法太大,现在最大的风险是想法太小”。除去 COVID 期间的一个高峰时刻,新企业成立数量已处于历史最高水平——那个本来会成为 YouTube 创作者的25岁年轻人,如今正在为所在社区开发 SaaS。
摘要 · 为研究而整理的核心内容

1. Grok Bots 买下牛仔裤——宏观数据表明需求实际上无限

  • 节目开场的轶事奠定了整集基调:Anish 平时主要穿 Frame 牛仔裤,他拍下自己当前这条的照片,隔夜对 Grok Bots 说“别花超过500美元,把事情办妥”,醒来时,一条经过调研、已经买下并在运输途中的牛仔裤出现在他面前——版型相同、洗水不同,用他的信用卡支付。他的判断是:“Grok Bots 的定义性特征是资源整合能力”(“the defining characteristic of Grok Bots is resourcefulness”);下一次跃迁来自资源整合能力,加上消费者能够理解的产品架构。
  • 对于谁会成为赢家,他的答案是“多赢家”:xAI 在2周内从毫无竞争力变成3家之一;Anthropic 从看似占据主导,转向 OpenAI 表现出色的一段时期——新模型异常出色、Codex harness,以及一款“做得非常好”的 ChatGPT 桌面应用;各家实验室的专业化方向正在分化,多家实验室同时增长。Jen 的补充是:X 是一个不完美的“风向标”,但能提前反映开发者情绪;Claude 正面临 token 使用量的反弹,Anthropic 将于今年晚些时候上市。
  • 泡沫论已经“被充分讨论”;真正超出常规分布的议题是乐观程度不足。证据在于:B200 每小时价格正在上涨,尽管这款并非最前沿的 GPU 通常属于一个高度通缩的市场——“需求实际上无限,供给高度受限”。
  • SaaS 的剧烈反转体现了市场心理:2月回撤30–40%后,a16z 表示软件遭到过度抛售;不少公司已经反弹40%(“我不太确定我们集体完成了什么”)。软件仅占企业支出的8–12%,因此,用 vibe-coding 自己的薪资系统或 CRM,其上行空间很小,下行空间则“实际上无限”;但随着 SBC 扭曲的经济性暴露,潮水已经退去,这些公司必须“加速,否则出局”。

2. 护城河大多能够存续;Token 支出按有限与无限上行空间分流

  • Anish 从《7 Powers》的框架出发,认为多数护城河不会因智能充裕而受损:“再多 coding agents 也不会让 Nike 不再是 Nike”,Instagram 的力量从来不在于应用本身有多复杂。在他看来,最明显暴露的护城河是整合能力——SAP 复杂到甚至从一个版本迁移到下一个版本都可能带来生死攸关的风险,SI/GSI 作为整合入口,其价值正面临一个生死攸关的问题。
  • 理性的企业架构是:创造 alpha 的职能(产品、销售、工程、研究)使用前沿模型 token,因为上行空间没有上限——为多1个 IQ 点支付几乎任何价格都值得,“你的 Fable 5、Gro 或 GPT56”。支持性职能的上行空间有边界:“结账最好的方式就是准确。你不可能把结账做得比准确好10倍。” 对这些场景而言,开放权重模型配合强化学习,可能处于成本曲线上的帕累托高效点。
  • Jen 提到 Decagon 创始人 Jesse Zhang 的一篇帖子:对很多初创公司而言,开源模型“实际上是唯一选项”——原因不只是成本,也包括本地化、训练和微调。Anish 补充说,在专业问题上进行强化学习,并利用其推理轨迹,可以形成不断复利的领域优势;Harvey 在法律领域已经取得了强劲结果。代价是通用性下降——对法律或客服而言,这完全可以接受。

3. 模型不是大宗商品——实验室向下整合,而不是向上整合

  • 他用“五大模型”的框架来描述:一类是“自闭型模型”——GLM 5.2 和 5.3“非常字面化,只会严格执行你告诉它的事情”;另一类是 Kimi K3,“开放、擅自推断且富有创造力”。模型不可能同时高度开放又高度神经质;会计可能偏好神经质,设计则可能偏好开放,因此组织需要多个头脑。
  • 法律插件恐慌是一个典型案例:Anthropic 的插件只是“一个技能文件的 ZIP……本质上就是长提示词”,却引发了市场恐慌,Thomson Reuters 及其他法律股大幅下跌。但实验室并没有向上整合到应用层,而是向下整合至推理和算力;后者的工作负载同质化,可以实现巨大的规模效应。应用层的定价、打包和采购行为高度特异,对实验室而言则意味着“运营开支沉重”。
  • 专业化已经体现在 harness 中:OpenAI 的新 GPT 模型和桌面应用,为知识工作提供了强大的产品容器;Claude Code 连同终端 UI,则面向软件工程。聚合可以胜过任何单一实验室——以 Expedia 为喻,Cursor 用前沿模型做规划,用较弱的模型执行;创意工具把 ElevenLabs(语音和音乐)与 Black Forest Labs(视频和创意方向)结合起来;研究可以将同一个查询通过训练数据不重叠的模型进行对抗式处理,再用另一个模型收敛答案。实验室在结构上只能提供自研模型,而应用聚合者可以提供各领域的最佳模型。

4. 应用将智能这一基础能力产品化;循环才是企业端的突破口

  • 核心类比是:智能是一种类似云的基础能力;正如 Salesforce 将 AWS 云这一基础能力转化为 CRM,“你确实需要 Harvey 把它转化为法律行业的经济结果”。信用合作社说明了需求的特殊性——多数信用合作社并不想把员工数量砍半,而是希望在保持经济效率的同时扩大到2倍。
  • Agent 去神秘化之后,就是“一个处于循环中的模型,加上工具和记忆”。编码循环是:收到 bug 报告、复现、生成修复方案、验证;低风险情况下还可以自主发布。这一逻辑可以延伸到定价优化和采购,乃至业务循环——模型提出建议:“我认为我们需要在 Tijuana 开一家分行。”
  • Marc 的说法是“行业,而不是市场”:编码这一基础能力,从 Claude Code 向开发者暴露“原始马力”,到 Replit 为不会编码的小企业主进行抽象,都是定价、产品化和打包方式的不同变体。至于实验室是否会挤压应用层,2023年单一模型占主导时,“实验室最终会拿走你全部的毛利率”;而如今,帕累托前沿上有多个选择,实验室更难做到这一点。

5. 消费复兴:DOS 时代、个人 Agent 与 Town 的记忆复利

  • 消费市场此前受阻的原因是:消费者不喜欢为软件付费,而 AI 又带来了真实的分发和互动边际成本——Jen 的 X 时间线应用接入1名用户的成本是250美元,难以支撑面向大众的免费产品;如今开放权重模型正以更低成本和更强性能改变这一点。市场没有 AI 原生应用商店,因此分发更像 Web 2.0,而不是移动互联网;“我们正处于 AI 的 DOS 时代……我们需要 Windows”。
  • 已经奏效的方向包括:面向“数字原生创业者”的 coding agents——过去围绕 YouTube 创作者的道德恐慌,如今被重新定义为年轻人想要创建互联网生意,而他们现在能够做出年收入10万—100万美元的“夫妻店式 SaaS”;这些公司不值得风险投资,但“对这个国家而言非常酷”。个人 Agent 正从 OpenClaw 1月“Homebrew Computer Club”式的氛围,进入 Grok Bots 和 ChatGPT Work 等消费软件。
  • 他对消费者的定义是:如果无法证明销售驱动型获客合理,通常意味着 ACV 约为1.5万美元,那么它就是消费业务——水管工也算。娱乐将成为巨大的市场(Character 可以说是一家娱乐公司;亚洲短剧也开始进入海外):“多数人想花时间,而不是省时间。”
  • 更广义的个人 Agent 模型是一组生活循环——家庭、友谊、金钱和健康;信息变化带来决策、代理、执行,再进入下一轮迭代。OpenAI 正聚焦健康和金融,初创公司则在做购物;Anish 预计生活质量将显著改善,其中80%的剩余价值会交付给大众市场。
  • Town(Alex Rampell 的投资项目)体现了记忆复利模式:第1天它只是一个新员工;第30天,它已经能基于持续吸收的上下文做出出色判断——最终体现为留存率和单客户定价能力。Jen 说她的工作邮箱已经清零,但个人邮箱里约有20,000封邮件;Town 会找出重要内容、清理订阅,并开始通过展示日常操作消耗的 credit 成本以及节省 credit 的方式来实现自我改进。

6. 经济性、创始人与资本流向

  • 消费市场的空白地带包括 AI 进入“情感与人际关系领域”:你可以与 Claude、OpenAI 或 K3 对话,“并感受到情绪”;此前“40年的技术都在增强我们的智力……却没有任何东西真正触及人性”。初创公司可以探索实验室和大科技公司在文化上不适合构建的产品,例如一个会反驳你、或使用性暗示的陪伴者。消费者很乐意下载新软件;Anish 表示,与99美分时代不同,如今他们愿意每月支付200美元。
  • 融资姿态是:基本不投尚未产生收入的项目——“如今在任何阶段的路演中不展示真实上线产品,都是直接出局的理由,因为做出东西实在太容易了”;投资工作在于根据具有统计显著性的销售和产品表现,推断价格与隐含风险之间的关系;对于有才华且经验丰富的团队,则保留少量“万事俱无”前的期权式押注。过去“种子资本过多会毁掉公司”的判断也变得更细致:聚焦投入的1亿美元,可以带来不同于2000万美元的价值主张——这比金融科技公司“通过薄弱的承保能力间接补贴客户”更值得解决。
  • 创始人画像正在转变:“少一些 MBA,多一些研究人员”——商业成熟度较低,但技术成熟度显著更高,而后者“是一切好事的上游”。商业成熟度可以通过教育和观察获得,技术成熟度通常无法如此复制。Ben 在 offsite 上的一句话是:过去最大的风险是想法太大;现在最大的风险是想法太小。
  • 对于已经存在的 SMB,旧有渠道依然有效,但创始人越来越需要原创的网络效应——口碑——因为 Instagram、TikTok 和 X 让企业很难在既有渠道之上再建立一条新的分销渠道。最有意思的群体是新企业成立者:数量处于历史最高水平,除 COVID 期间的一个高峰外也处于最高水平——“不是那个55岁的水管工……而是一个25岁、正在为所在社区或高中开发 SaaS 的年轻人。”
Jen Kha

To help me break down all things around this incredible abundance, I'm going to bring up Anish Acharya.

Anish Acharya

Hi.

Jen Kha

Awesome. Awesome. Awesome. Hey, Anish. Anish and I were at the GP offsite earlier this week, and he shared with me that he's already running Grok Bot and has purchased a bunch of jeans for him. Anish, do you want to drop what you purchased?

Anish Acharya

True story. Yes. So, I'm going to reveal an important secret—protected IP—which is that I mostly wear Frame jeans. Frame is a great brand, and Grok Bots are an awesome product.

Actually, I'd say the defining characteristic of Grok Bots is resourcefulness. I went to bed a few nights ago and said, “Hey, buy me a pair of jeans that are inspired by these.” I took a photo of my current jeans, said, “Don't spend more than $500, and get it done,” and went to sleep. I woke up in the morning, and it had researched and found a pair with the same fit and a different wash, used my credit card, purchased them, and they were on the way.

I think that's going to be something that we see more and more of. We already have the capabilities, and now a lot of the unlock will come from resourcefulness, as well as product architecture delivered in a way that most consumers can understand.

1. Who Wins the AI Model Race in Three Years?

Jen Kha

Awesome. Awesome. Awesome. I told my team that I'm going to set my bot to finally take care of the pile of things I've been promising my husband I'm going to sell for the last 2 years. That is the project for this weekend.

Anish, we asked the question earlier: Which one of today's AI leaders will be the clear winner 3 years from now? What's your take?

Anish Acharya

I'm a many-winners guy, and I see I'm in good company with many of you. If you look at what's happened in the last 2 weeks, xAI went from not even being a real contender on the model side to being one of 3. We essentially went from a 2-horse race to a 3-horse race.

Even more broadly, over the course of the year, we went from Anthropic feeling like they were so dominant they could do no wrong to OpenAI, which has just had an excellent 3 months. The new models are exceptional. The new Codex harness and ChatGPT desktop app are very well done, and we're seeing the specialization of these labs in different directions.

They're both growing like crazy despite each other's continued successes. xAI is doing well, and open-weight models are doing well. So, I'm definitely in the many-winners camp.

Jen Kha

Yeah, it's interesting to see the sentiment also on X, which is not always a perfect weather vane for the future, but it's often an early indicator of at least where developer sentiment is. There's been a lot of pushback from Claude users recently in terms of token usage, and developers tend to be fair-weather fans on these things. They'll go where the latest, greatest, and very best model isn't.

2. What's Next in the Frontier of Intelligence

Over the last 6 to 8 weeks in particular, I think we're going to see some very interesting traction in terms of the flow of activity. Obviously, Anthropic is going public later this year, and there's a lot of keen interest in this.

With that, it brings us straight into the topic of discussion today: Where and what is next in the next frontier of intelligence?

Anish Acharya

Amazing. Thank you, Jen. Let me tee this up for everybody. Please hop in if you've got questions.

Let's first cover the macro and what's happening at a market level. Then we're going to hop into the application layer broadly and talk through why applications are the productization of the intelligence primitive. Finally, let's talk about the consumer. With the launch of Grok Bots and a few other products, it's actually been a very fun couple of weeks in consumer.

Hopefully, our dear friend Leopold doesn't mind me poking a little fun at him here with “Situational awareness, please. Next.”

Look, I think the case for this being a bubble is over-discussed, or at least fully discussed. I think the out-of-distribution topic that's less discussed is: What if we're insufficiently optimistic?

If you look at some of the underlying indicators, what they point to is essentially infinite demand and highly constrained supply. Things like B200s, which are non-cutting-edge GPUs, going up in price on a per-hour basis are very strange. Normally, we see these things be highly deflationary, and it points to very constrained supply and essentially infinite demand.

We're thinking and talking a lot about what's the informed case for optimism here, given some of these second-order indicators.

The SaaS bubble—or the SaaS whipsaw—was an interesting peek into market psychology. Back in February, when we saw a 30% to 40% drawdown in a bunch of SaaS names, we said that the market had oversold software. Lo and behold, here we are: Many of those names are back up 40%.

I'm not quite sure what we collectively accomplished, but I'll tell you what we said then, which is still true today: For the enterprise, software spend is 8% to 12%. It's just not a huge proportion of spend. So, the upside to vibe-coding your own payroll or CRM is not particularly high. The downside is essentially unlimited.

There are obviously all kinds of compliance implications of not getting things like payroll right. Most enterprise software today demands a level of precision that just isn't afforded by coding agents.

The one thing that has happened, though, is that the tide has receded. A lot of SaaS companies had a ton of SBC and things that distorted their economic performance. I think that's very much visible now, and they're going to have to accelerate or die.

So, it's less bleak for the SaaS market than perhaps we all collectively thought for a few months there, but there are still some existential questions to address.

There's been a huge discussion of moats. Are there any moats? There are no more moats. It's very funny because if you actually study moats, which I think are most famously codified in the book *7 Powers*, the vast majority of moats are not affected by abundant, low-cost intelligence.

When you think about network effects, scale effects, which show up in distribution, and brand effects—which we tend to discount in Silicon Valley—these things are as good as they've ever been. No amount of coding agents is going to make Nike not Nike. The power of Instagram was never the complexity of building the Instagram app. Of course, it was the network behind it.

You actually think the majority of moats are as good as they've ever been and, of course, are still critical to building compounding value.

There are a couple of moats that are exposed. For me, the integration moat is the most obvious one. SAP is so famously complex to integrate into and out of that it's an existential risk even to migrate from one version of SAP to the next. Coding agents make this dramatically better.

I think there's a bit of an existential question for SIs and GSIs as to what their value will be when they've historically been this point of integration. So, I do think this moat is a little bit at risk. But for the other traditional moats, they persist and they're as important as they've ever been.

Jen Kha

Yeah, I think this is a really important concept. As you start to think about what the job functions in the enterprise are that are alpha-creating, it's typically product, sales, engineering, and research. Conversely, what are the job functions in the enterprise that are supporting other functions? “Administrative” is maybe too bleak, but legal, HR, finance, and so on.

We really think that the rational architecture—and the one that is emerging—is that for jobs that have unlimited upside, like sales or product, you always want to use frontier tokens. The reason for that is you just don't know what the value of the new product feature or closing a customer account is. It's effectively unbounded, and therefore it's economically rational to pay almost any price for a model that's even 1 IQ point smarter—your Fable 5 or your Gro or your GPT56.

Conversely, when you talk about something like finance, the best way to close the books is accurately. You can't close them 10 times better than accurately. As a result, you have this bounded-upside problem where it makes sense to use open-weight models with reinforcement learning for the Pareto-efficient cost curve.

3. Why Open Source Is the Only Option for Some Startups

Maybe before we move on from this one—because this is a big debate, and again, when Kimmy dropped a few weeks ago, there was a lot of consternation about this topic, given the relative cost, which was the focus of the discussion—our founder, Jesse Zhang from Decagon, dropped this great post about the fact that, in some respects, and for a lot of companies like Decagon, open source is actually the only option.

It's not just cost. It's that they can actually localize it, train it, and fine-tune it. Maybe unpack a little bit of that configuration. Talk through the nuances there and why folks shouldn't be concerned, even though that is the case for startups, that there's a lot in the way of abundance around this topic.

Anish Acharya

Yeah. One of the big topics that we're seeing—or one of the big trends—is that there are comparative advantages of different models. The models often have areas of focus that are almost in tension with each other.

So you see a certain set of models that have a high degree of neuroticism. They’re sort of autistic models. GLM 5.2 and GLM 5.3 are great examples of this: they’re very literal, and they’ll only do exactly what you told them to do and nothing more. Then we’re seeing models like Kimi K3 that are much more open, presumptuous, and creative.

There are roles for both types of models in the organization, and often the shapes of those minds, if you will, are at odds with each other. That is one reason you actually want to have multiple models. Reinforcement learning is a really important point. If you actually have a problem that you can specialize the model around with your reasoning traces, you can start to create this compounding advantage in your domain for your customer base, where you’re able to shape the intelligence to be better than any general intelligence for your problem. I know Harvey has had some great results with this as well.

The trade-off of that kind of reinforcement learning is that you lose generality. If you have the best model fine-tuned for solving legal problems, it may not be great at solving theoretical math problems, and that’s okay for Harvey’s uses or in the case of Decagon’s customer support. This sort of open-weight specialization property is very unique and is one of the reasons our startups are selecting these models.

This is also a big topic. We’ve learned so much since January. We should really do this monthly, Jen.

Jen Kha

Yeah. I mean, honestly, there’s just so much changing.

Anish Acharya

So, in January and February, there was a lot of discussion, and it’s very idiosyncratic and interesting. Anthropic’s Claude released what is called a legal plugin. Plugins are just collections of skill files; you can think of them as a ZIP of skill files. Skill files are just prompts—long prompts. There was this huge panic, and all of a sudden Thomson Reuters and a bunch of other big legal names traded down dramatically.

But those were really just prompts. There was a lot of discussion about whether labs were going to vertically integrate up into the application layer. Instead, we’ve seen the very opposite: yes, they are vertically integrating, but they’re vertically integrating down into inference and compute. It’s actually logical now, in hindsight, because the workloads for inference are very homogeneous, so you can build enormous scale in one part of the value chain. Whereas when you think about the application layer, you’ve got so many idiosyncrasies and unique needs in terms of pricing, packaging, productization, and how the market wants to buy.

It’s actually a much more challenging and OpEx-heavy proposition to move into the application layer versus moving down into the inference layer. This is the point I alluded to earlier, which is this discussion of model commoditization. If you use the models every day, which I do—I hold myself to a standard of making something either small or big with every model that comes out—you start to appreciate the fact that these things are not commodities. They have comparative advantages at a domain level.

A great example is OpenAI’s new GPT models, which are just so good at knowledge work. The harness is also very well set up for knowledge work. If you’ve used the ChatGPT desktop app, you know what I mean. It’s very cool and interesting, and it’s the perfect product container—to use “harness,” I mean a product container like a browser—to do spreadsheets, slide presentations, written documents, and all of that type of work.

If you look at Claude Code, which many of you, I’m sure, have used, it’s oriented toward software engineering. It’s in a terminal UI. Everything from the small design decisions to the areas in which it specializes, like code planning and code testing, is oriented toward the software engineer. Both products make many trade-offs for their respective specializations.

So, one, you’ve got this domain-level specialization that’s already occurring. Then, two, as I mentioned earlier, you’ve got what I think of as the Big Five personality traits, if folks have studied that. You can’t be both highly open and highly neurotic. Sometimes, when you have an intelligence you’re applying to an accounting problem, you want neuroticism. When you’re applying it to a design problem, you want openness.

You actually have a need for both types of minds in the organization, which is why you would select something like a GLM 5.3 versus a Kimi K3. So, definitely not commodities in our view.

This is an important point. There are many product categories in which model aggregation delivers a greater-than-the-sum-of-its-parts outcome. A good metaphor for this is Expedia. It’s so much more useful to use Expedia than it is to go to United, then to Delta, then to Southwest. You just want a single place where you can benefit from seeing every airline’s inventory.

Similarly, in coding, we’re seeing this with Cursor a ton, where you want to use a frontier model for planning, for example, but then you can use a lesser model for execution. You really need to have one product harness, or sort of product architecture, that lets you use multiple models.

Creative tools are another great example, where you’ve got models that specialize in different modalities. You’ve got something like ElevenLabs, which of course is incredible at voice and music as well, and then you’ve got something like Black Forest Labs, which is doing such an excellent job in video and creative direction. The correct product is to bring all of these together into one shell.

Finally, there’s research and decisions. We see this all the time, where the models are often trained with non-overlapping data sets. You’re able to get more information by running the same query through many models adversarially and then having a separate model help you converge. This is a place where the application layer really shines, because labs are both incentivized and structurally only able to provide their own in-house models. You, as an application aggregator, can provide the best of breed.

Okay, let’s jump into the apps layer. The key point about the application layer is that intelligence is a primitive, just like buying cloud is a primitive. What does Salesforce do? It takes the AWS cloud primitive and turns it into CRM software that delivers an economic outcome for all of its customer segments. The same thing is true of the AI application layer.

It’s great to have the raw intelligence primitive, but you really need Harvey to turn that into an economic outcome for the legal industry. A similar example is credit unions, which are a really interesting market segment because they’re so idiosyncratic in how they want to buy products, how they want the product to be productized, and the shape of the ambition for their market.

Most credit unions don’t want to decrease their headcount by half. They want to double it, and they want to double it while having an economically performant business. It’s just a very specific way that they see the intelligence primitive playing out in their market segment, and the application layer’s opportunity is to be the one that delivers that.

This is a bit of an advanced concept, but I think it’s an important one. If you look at the way the evolution of AI use has gone, it’s gone from prompting models to putting models in loops. The term “agent” is overused, but an agent is just a model in a loop with tools and memory and a few other things.

A great example of this is coding. We’ve all seen this from software companies: a bug gets reported, it gets reproduced, a fix gets generated, and it gets verified. If it’s a low-risk fix, it gets integrated and shipped, and maybe the customer gets an email saying, “Your bug was fixed.” If it’s a high-risk change, perhaps a human reviews it. That way, every bug that actually gets reported to the enterprise now gets autonomously fixed through this coding loop.

As you start to take that idea and apply it to other parts of the business—things like price optimization and procurement—these are very natural business loops that can be fully automated by these models. Perhaps the most ambitious type of loop is the business loop: you make a change that’s very cross-cutting to the business, and the model comes back and says, “Hey, I think we need to open a branch in Tijuana.”

Now, the model can’t do that autonomously, but it can make a change at the surface level of the entire business, which is extraordinary. This is how enterprise automation is going to occur through AI. For me, coding has been, over and over again, an illustration of this.

4. One Dominant Personal Agent or Many Talking to Each Other?

Legal is another great example of an industry, not a market. This is something that Marc says, and he’s so right. If you look at intelligence as a primitive, let’s think now about coding intelligence as a primitive. All of these products are working in their respective areas of the stack.

Claude Code does such an excellent job of exposing the raw horsepower, so to say, to the developer, all the way up to Replit, which is a great abstraction layer for the average small-business owner who’s unfamiliar with code. These are variations of pricing, productization, and packaging for the coding and intelligence primitives, and all of them are working as a result. So, I think a big mental model shift for us is ensuring that we’re assessing these as industries, not necessarily simple markets.

Jen Kha

Okay, and consumers had a really cool couple of weeks. We've been saying for 3 years that this is going to be the consumer's quarter, but I think that this might be the consumer's quarter. Let's go into it.

The things that have actually held back consumers so far have been a couple of things. The first is that consumers don't love paying for software. We've learned this lesson over and over again. Unfortunately, unlike the sort of magic of software in the past, AI software has marginal costs of distribution and engagement, and the marginal cost can sometimes be very high.

I built an app I use to help me browse my X timeline, and it costs $250 to onboard a new user. So, if I'm a startup founder looking at a $250 CAC, even with a $0 onboarding cost, it's very hard to make a mass-market free product work. That is changing now because of open-weight models, which are dramatically cheaper and more performant.

The second is that we've never had an AI-native distribution channel. There's no App Store for AI. So, this actual product cycle for consumer looks more like Web 2.0, where you have to build the channel alongside the product, and less like mobile, where you have the central point of distribution for the entire ecosystem.

Then the final point, I think, is an important one: we're sort of in the DOS era of AI. For this technology and its capabilities to be fully embraced by consumers, we're going to need the Windows, so to say. I think there's just a ton of work to be done around product and design craft to ensure that consumers know how to consume all these magical new capabilities.

Two things are working. Coding agents are extraordinary, and I know they've been discussed. I think it's interesting to think about how they work for consumers. If you think of this concept of the digitally native entrepreneur, if you're not a programmer, the way that's historically shown up is that you're a YouTube creator.

There was a whole moral panic 10 years ago about how kids wanted to be YouTube creators, not astronauts. But I would interpret that instead as kids who grew up on the internet wanting to build businesses on the internet, and the only way to do it, again, was by being a creator. Now, with coding agents, you can build a software product that generates $100,000 of revenue a year or $1 million of revenue a year.

These are not venture-backable businesses, but it's a sort of mom-and-pop SaaS opportunity that's emerging, and I think it's very, very cool for the country.

Personal agents: we had this collective moment of excitement around OpenClaw in January, and it was an extraordinary composition of primitives, but it never really crossed over into consumer. It was sort of a developer-oriented thing, more of a Homebrew Computer Club kind of energy. We're starting to see, with the emergence of Grokbot and ChatGPT Work, personal agents being turned into software that consumers can use.

5. Redefining Consumer: When the Plumber Uses GrokBot

Anish, actually, do you mind just pausing on this before we go to the Town demo? You were a founder building in the last era of the consumer app experience, and when I even think about it, I'm like, gosh, how do you even define consumer today? The plumber who utilizes Grokbot to completely turn around their business end to end—is that consumer or is that enterprise? It's almost like a PLG movement, but it's coming from a consumer that then crosses over into enterprise.

Particularly, the last era of consumer applications was more towards entertainment as a way to monetize, so maybe unpack some of that, and particularly where you've been spending time as a part of that.

Anish Acharya

I mean, our simple rule is: if you cannot justify acquiring the customer through sales, which usually means a $15K ACV, you have to acquire them through marketing. We think of them as a consumer, which is most small-business owners. So, I think that plumber is definitely the consumer in our investing mind.

Entertainment is huge, and there are going to be a bunch of AI-native entertainment companies. I would argue Character was kind of an entertainment company. There's been a huge trend around short-form drama, mostly in Asia, and that's starting to come over here. Many of those are generative or generative-assisted.

I think entertainment is going to be massive. Most people want to spend time, not save time, and consumers are not that interested in productivity. So that's definitely going to happen, and it's probably worth a separate deep dive.

I think Town, for folks who have used it, is just such a magical experience. This is the No. 1 piece of advice I give to everybody—friends, family, folks in the industry: please just use the products, because it's so easy to build intuition when you see how they change day to day.

Town is an investment our partner Alex Rampell made. It's a really extraordinary productivity product, and you see how the compounding improvement of the product through memory advantages it over time. The first day you use a product, it doesn't know you that well. It's sort of like a new hire who's just getting up to speed.

By day 30, it's able to make excellent assumptions on your behalf because it has soaked in 30 days of context, memory, and skills. This is a pattern that we're seeing more and more: the compounding value being delivered to the end customer showing up as retention in the business and showing up as pricing power on a per-customer basis.

6. Town Demo: Personal Agents & Managing Chaos

Jen Kha

Yeah, this is a great one, because folks can utilize Town for their personal use case. It's a free trial; they give you, I think, something like 40 credits to start, or something around there, so you can see, once you plug in your personal email, how productive it actually is. On the professional front, I'm always at inbox zero. On the personal front—

Jen Kha

My inbox is like 20,000. David George is probably cringing on the inside here just because it's unacceptable. However, personal-life things are common. If you email me at my personal address, I will never respond to you.

However, I plug my email into Town, and I don't even check anymore. If there's something important, Town will surface it to me. It also does all the scrubbing of subscriptions and all the things that it can optimize, and it's starting to self-improve. It'll send you emails where it says, “Hey, this routine is costing this much; here's how you could actually save your credit spend.”

So, it's this unlock into what starts on the productivity side and, to your point, maybe people won't pay for that personally. But once it starts to get locked in and then expand in terms of the remit, you're like, “Okay, I'll pay whatever X bucks,” just because it helps to manage my life and I can put it on autopilot.

Anish Acharya

Yeah, it's such a great point, Jen. My mental model for this is just an experienced employee, a tenured employee versus a new hire. The new hire may be brilliant and may even cost less than a tenured employee, but we all know the value of a tenured employee. They're just able to make great assumptions on behalf of the organization and you.

This is a little philosophical, but I think this is where it all goes. Just as we talked about coding loops and business loops for the enterprise, we think there's a set of loops that are informally defined that really lay out a consumer's life. Think of family, friendships, money, and health. These are all areas where you have changing information, decisions, agency, execution, and then the loop continues.

We're starting to see some of these loops emerge around self-improvement. Health and finance are the 2 areas that OpenAI is focused on. We've seen a bunch of startups working on shopping, but we think that the way this ends up playing out is a dramatic quality-of-life improvement for the consumer, and that really follows the shape of past product cycles, where 80% of the surplus is delivered to the mass market.

Jen Kha

Maybe just going back to the last slide, there's a question here. When you think about these personal-agent examples, whether it be Town or Ethos, et cetera, they all point to one assistant having context, but it seems like there are many different options.

Do you think it'll end up being one dominant platform for this personal aspect of your life, such as time management? Or will it be like an operating system where you have many agents talking to each other and configuring on the back end?

Anish Acharya

The comparative vantage point kind of comes to mind. I think the characteristics you want from your CFA are different from the ones that you want from your party planner. The surface area is so broad that I think, yes, there's overlapping bits of context.

I think Grok Bots has done a nice job of illustrating this in product: you have many bots that are pointed in slightly different directions, and they all coordinate to deliver a globally optimal outcome.

7. Apps vs Model Companies: Who Captures the Value?

Jen Kha

There are a few questions. I'm going to go back to topics you've covered earlier. If the application layer captures economic outcomes, how do you think about the competition from the model companies, and what will allow value accretion to happen downstream? Are companies at the app layer able to compete with the frontier labs going after that particular market?

Anish Acharya

I mean, I think so. Again, I think we're underestimating the complexity of product, pricing, packaging, and how the end customer wants to buy.

The way that a teenager wants to consume the intelligence primitive is different from the way a marketing executive at a credit union wants to consume it. It's very heterogeneous. So to me, it just makes less sense for the labs to move up to the apps layer than to move down to inference.

That permission point is an interesting one. I think if we lived in a world of 2023, when it was one model to rule them all, it wouldn't even matter if you had permission, because the labs would just take 100% of your gross margin over time. But now, because you've got many options at all points on the Pareto frontier, the labs have a harder time actually doing things like that.

Jen Kha

Awesome. There was a question just on traction. Do you fund anything where there's no revenue at this point, given how quickly people have been making progress, or is it extremely difficult?

Anish Acharya

We try not to. I certainly have spent less time on that strategy. Look, I think the basket is mostly investments that are showing some signs of working. Certainly from a product-velocity perspective, that used to be something we measured pretty carefully. It's disqualifying not to be showing a live product in a pitch at any stage these days, because it's so trivial to build stuff.

So almost everything we're seeing is showing signs of some sort of breakout. My model is somewhat simplistic: once you have statistically significant sales and product, if we extrapolate from there, do we like the price that we have to pay to be a part of it and the risks that we're taking implicitly? I'd say that's the majority of the work that we do. We look for very talented, experienced folks, and we do take a small call option, which looks like a pre-everything round, but that's not the majority of what we do.

Jen Kha

When you think about the competitive landscape on the consumer side—this has been unloved for so long—are you seeing this reversion, given it's clear that apps are sort of this next layer of value creation? The model layer has been somewhat set, and I say that with a huge asterisk, because there might be new algorithmic breakthroughs, folks coming out from left field, as we have in the portfolio as well. But do you feel like the competitive dynamic is shifting more toward applications?

Anish Acharya

100%. It's sort of a renaissance for being a consumer builder because you've got this extraordinary primitive that you can work with. By the way, we now have a primitive that can operate in the emotional, interpersonal domain. You can have a conversation with Claude or OpenAI or K3 and feel feelings. We've had 40 years of technology that really boosted our intellect and productivity, but nothing that spoke to our humanity.

So it's a whole different technology surface, and it's very wide. I think there are a set of products that labs are just not culturally set up to go after, and big tech isn't set up to go after them either. Think about launching a companion product at Google that may disagree with you, that may have sexual innuendo in it. These are things that a thousand committees at Google are designed to prevent.

So startups have areas where they're uniquely capable. And then, finally, the consumer is excited to download new software, excited to pay for it. It's like Christmas 2009 with the iPhone. People want to try new apps, but unlike the 99-cent days, they're willing to pay $200 a month. So it's sort of a renaissance for consumer builders, and, yeah, I think things have changed.

Jen Kha

Anish, dropping Birkin bag–framed jeans. I had no idea you were such a fashionista. This is like your butt is helping you get up to C-suite here, my friend. Uh, for a guy secret is a good steward of capital. Okay, that's all that I am.

Anish Acharya

For a guy I only see in quarter-zips, I'm just saying.

8. Who's Actually Building Apps Today? Founder Archetypes

Jen Kha

Okay, maybe one question for you on the founders, because I don't know if you remember this conversation. This is probably 5 years ago or so, when most of the founders you saw had more diversity in their backgrounds, in part because the software and technology were way more sophisticated. So you had a lot of program managers spinning out of Google, for example, and starting a company. What are the types of founders you see building apps today? Do they tend to lean more technical, more researcher-derived? Are they product managers? What kind of archetype are you seeing in at least the early innings of apps coming out of the woodwork?

Anish Acharya

Yeah. Less MBAs, more researchers, and they both have their strengths and weaknesses. I think the business sophistication of the founders we're seeing today is lower, but the technical sophistication is dramatically higher, and the technical sophistication is upstream of all the good things that happen. Business sophistication can be taught and observed, but technical sophistication typically cannot.

So we're definitely seeing a much more technical, earlier-career founder, but the things they're doing are extraordinary because they don't have any preconceived notions about what's possible. And so much of what holds back senior founders that don't quite get to the other side of this product cycle is that they're not close enough to the technology, and they've got an idea that's rooted in the past of what the ceiling is.

I think the best thing about these young founders is they assume everything is possible. We are at an offsite where Ben was saying that the biggest risk with ideas in the past was that they were too big, and now the biggest risk is that the ideas are too small. But I think that's illustrative of the different founder archetypes.

9. Why Giving Founders Too Much Money Isn't Fatal Anymore

Jen Kha

Yep. And maybe on that similar thread, it used to be that if you gave a founder too much money, it would wreck the company, because the founder almost always has way too many ideas and is a visionary and doesn't have the talent to land all those ideas commensurately. We're seeing a whole new paradigm on that. Maybe unpack that idea a little bit more, because it was such a huge theme of the offsite.

Anish Acharya

Yeah, for sure. This was a historic wisdom. Why didn't we give every seed company $20 million, $50 million, or $100 million? It wasn't just the risk-reward, but rather, typically, the constraining factor was that they just didn't have enough talented people to work across $20 million of product surface at the same time. They really had to focus on 1 idea at a time, and the capital was a great way to enforce that focus.

What we're now seeing is that you could make different sorts of product and model trade-offs through more or less capital. And there is a case for a company that raises $100 million, uses it productively and in a focused way, and is able to deliver a different value proposition than the very same team would be able to do with $20 million.

So I think that, again, just as we talked about the fog of war around margins, this question of what is the optimal seed-round size and how much capital can you put to work effectively is a much more nuanced topic. It's sort of an embarrassment of riches, but I'd rather have this problem than the problem we had 5 years ago, which is: “Hey, my fintech company is indirectly subsidizing its customers through weak underwriting, and we don't know the path home.”

Jen Kha

Yeah. Yep. Yeah. The Chris Dixon model, which is: You always want the problem of supply, not of demand. Right now, we have to fix the supply part, right? But the demand is so abundantly there that undoubtedly the supply part will get fixed.

10. Go-to-Market for Startups Selling to SMEs

Maybe I'll close on this one last question for Anish. Double-clicking on the theme of sector adoption of AI, unlike large enterprise, the friction of adoption is much less because they require less change management. I agree with much of that, but not all. Small and medium-sized businesses sometimes have more habit change that you have to work through.

But the question is: how do you see the go-to-market playbook for startups targeting SMBs, and has that changed in the age of AI?

Anish Acharya

For existing SMBs, I think it's the same channels through which you historically reach them. I actually think one of the interesting things about marketing in the age of AI is that all of the existing networks have been so trained on the methodology of building new networks that they're very careful to ensure no one does it on their network. Instagram, TikTok, and X—it's very hard to build a new distribution channel off the backs of an existing one.

So what founders have to do is actually build a product that has the original network effect, which is word of mouth. We're definitely seeing more of a focus on word of mouth. The old channels for reaching them are still there.

I actually think the most interesting segment of the market, though, is new business formation, which is, by the way, at an all-time high. I think it's the highest it's been outside of a peak moment during COVID. These are people who would have never otherwise been business owners. It's not the sort of 55-year-old plumber; it's a 25-year-old who previously would have been a YouTube creator and now is building SaaS for their neighborhood, their city, their high school, or whatever else it is.

Jen Kha

Yep. Yep. Awesome. Well, thank you so much for listening. It's always great to have you on. Now, I know you're a fashionista, and we're going to be clipping that endlessly on the socials. But thank you for that. And if folks have any questions, you know where to find Anish. We'll follow up here on some of the questions we weren't able to get to as well.