Anthropic、Glean 与 OpenRouter:Menlo Ventures 的 Deedy Das 解读 AI 护城河如何构建
Glean 真正的护城河,是积累多年的企业级苦活,而不是“AI 搜索”这个口号。Deedy Das 说,2019 年在湾区派对上提到“企业搜索”,话题就会立刻冷场;但 Glean 用3年时间打磨权限、连接器、排序、新鲜度、评测和用户采用,最终让 ChatGPT 的出现加速了其商业化,而不是拯救一个从未奏效的产品。在70亿美元估值、营收已达数亿美元的情况下,他的总结很直接:“护城河就是我们把难活都做了。”
Anthropic 的收入曲线,甚至超过了投资人最乐观的预期。Menlo 在 Anthropic 尚无收入、估值约40亿美元时首次投资;Deedy 描述了这样一条爬坡路径:1年内从0增至1亿美元,之后又从1亿美元增至10亿美元,今年公开预测则是从10亿美元增至90亿美元。他认为 Anthropic 是“有史以来增长最快的软件公司”,但也明确承认,没有人预料到这一结果。
企业 API 份额显示,多模型格局可能长期存在,但 Deedy 不会仅凭份额为超过1700亿美元估值的融资下注。Menlo 的调研数据显示,2023年 OpenAI 约占50%、Anthropic 占12%;到2025年年中,两者分别为25%和32%。这里统计的是企业 LLM API 支出,而不是 token 数量。在当前规模下,更重要的是“收入、利润率和增长轨迹”,以及可信的新市场和新产品。
这场讨论同时衡量了模型层的防御性与应用层更薄的护城河。Deedy 的非对称测试是:Anthropic 进入某个应用品类,要比一家应用公司成长为 Anthropic 容易得多,尤其是在大多数 AI 应用上方仍缺乏足够“厚实的一层”时。Claude Code 通过使用量和数据飞轮强化了这一判断,但他不接受其被所有人普遍偏好的说法,并警告模型实验室最终可能会与那些为其创造 token 需求的企业正面竞争。
1亿美元的 Anthology Fund,是一支旨在避开传统企业创投激励的模型供应商生态基金。Menlo 在外部管理这支基金,因为内部基金往往会优先考虑“谁最常用我的东西”;Anthology 则可以投资战略重要企业、Claude 重度用户,或极其优秀的早期创始人,而不要求其依赖某个特定模型。基金已投资约40家公司,单笔支票从10万美元到2000万美元不等;Deedy 称,其被投公司进入后续融资轮的比例显著更高。
OpenRouter 是 Deedy 眼中由繁琐细节构成的 PLG 基础设施护城河典型。主持人称该公司大约抽取路由支出的5%;Deedy 则强调其开发者心智、供应商级性能数据、隐私路由,以及无需销售电话即可使用的产品。它面临的风险同样具体:token 价格下跌压缩收费池,业余用户流失,以及企业先用 OpenRouter 做评测、之后直接与模型供应商签约。
研究投资组合押注的是“未来可能变得必要”的场景,而不是相信每种架构都能胜出。Goodfire 的机械可解释性被称为“给 LLM 做脑外科”,目标是让影响重大的模型决策变得可检查;主持人认为,扩散语言模型可能以当前质量的80%–90%、十分之一的成本和延迟交付结果;而分布式训练、人才获取和更广阔的愿景,可能构成 Prime Intellect 的上行空间。Deedy 反复强调,强技术仍可能输给时机和市场结构。
编码代理同时制造了安全风险与人力资本风险:人们可能执行自己无法检查的代码,并逐渐失去推理能力。主持人讲述了一个据称是伪造的面试仓库:其中疑似把数据外泄链接藏进字节数组,Cursor 据报检测出了这一陷阱;同样的工具也可能变成一个不断发出“请修好”的“老虎机”。主持人提出的替代模式,是快速、有人参与的辅助:帮助工程师找到并阅读正确文件,同时让他们继续亲自编写和理解代码。
1. Glean 的护城河,是竞争对手不愿做的乏味苦活
Deedy 的回顾从2019年讲起:当时在湾区派对上提到“企业搜索”,话题会立即冷场。Glean 在这段不受欢迎的时期里搭建底层检索系统;2022年12月 ChatGPT 出现后,它加速的是销售推进,而不是拯救一个从未奏效的产品。
按照他的描述,这是一门具备传统企业软件吸引力的生意:自上而下销售、合同容易扩容、替换成本高昂,TAM 几乎覆盖所有知识工作者。他仍然愿意持有 Glean 的股票,部分原因在于其估值约70亿美元,“不是1000亿美元”,后续增长空间仍然很大。
Deedy 否定了“Glean 先做失败的搜索,再贴上 AI 标签”这种被压缩后的创投叙事。他给出的解释没那么光鲜:在品类变得拥挤之前,公司已经处理了集成、权限、客户特定例外和“最后一公里的杂事”:“这不是护城河。护城河就是我们把难活都做了。”
2. 数据把关者与前沿实验室,不会自动抹平 Glean
对于 SaaS 厂商限制 API 访问,Deedy 质疑其第一性原理上的商业逻辑。Glean 只向已经拥有相关权限的用户展示 Slack 结果;它既没有把一个授权席位转卖给1000人,也没有从 Slack 手中拿走收入。按他的说法,更多使用 Slack 数据,反而可能支持更多席位销售。
他的第二层防御是多元化:Glean 拥有数千个集成。Slack 对许多企业可能至关重要,但单一供应商关闭访问的伤害,远低于整个生态协调关闭——后者他承认“可能更麻烦”。
客户是第三个施压点:他们认为自己买下了软件,也拥有由此产生的数据。客户的反对理由很直接——“你不拥有这些数据。”因此,阻断一个把客户信息连接到另一款已购产品的 API,制造的是与客户的冲突,而不只是与 Glean 的冲突。
Anthropic 或 OpenAI 可以做出一款还算过得去的企业搜索工具,但 Deedy 怀疑它们是否会投入所需的深度。一笔10万美元、20万美元,甚至百万美元级的定制销售,对营收超过50亿美元的公司几乎无足轻重,却需要庞大的销售团队、大规模 FTE 团队和大量定制工作。“你加入一家大型 AI 实验室,是为了做模型,不是为了搭建 Google Drive 连接器。”
3. 企业搜索需要不同的排序信号,也需要强制分发
消费者搜索依靠海量行为数据持续优化,包括点击、悬停、停留时长和重复查询。但在1万人的企业里,即使每名员工每天搜索2到5次,也不足以驱动同一套机制,Glean 只能另造一套信号体系。
企业查询也更新鲜、更不符合头部词分布。除了“福利”或“薪资”等共享词,员工还会搜索与具体岗位绑定的高度特定材料,因此缺少消费者搜索中那种重复出现、便于传统排序的头部词。
评测变得异常不透明,因为工程师往往无法理解客户的查询、文档或正确排序。Deedy 回忆,团队研究客户的专门数据时会承认:“我们其实完全不知道自己在做什么。”这不是因为排序任意,而是因为真实答案藏在陌生的业务领域里。
用户采用和相关性同样难搞。生产力工具能留住用户,是因为员工喜欢它们,而不是因为买方能证明精确 ROI;但搜索没有 Slack 那样的网络效应。因此,Glean 反复追问自己要做什么才能“赢得拥有新标签页的资格”,并在自测效果更好后,用 Chrome 扩展替代原生的 Google Drive 搜索。
4. Anthropic 以非凡增长曲线,叠加异常宽松的文化
主持人记得,Claude 最早的界面是在 Slack 里给机器人打标签:Claude 1 于2023年3月推出,Claude 2 于2023年7月推出。这个笨拙的 Slack 起点,与后来那些传统产品管理可能根本不会提出的产品形成鲜明对比。
Menlo 在 Anthropic 尚无收入、估值约40亿美元时首次投资。Deedy 提到,Anthropic 1年内从0增至1亿美元,之后又从1亿美元增至10亿美元,今年公开预测则是从10亿美元增至90亿美元;结果“超出了我们最狂野的预期”。
Deedy 看到的是一支理想主义的研究团队,其结果分布异常宽广:同样的特质既可能让 Anthropic“彻底熄火”,也可能造就一家跨世代公司。主持人认为 Claude Code 是少见的聊天之后的产品创新,它通过终端交付,“是每个 PM 的噩梦”;Deedy 自己的解释则是,给优秀人才足够空间和大量 token,他们往往会做出好东西。
这家公司展现出比其他实验室更自由、更少规定的文化,据引用的估算,其员工1年留存率约为80%。Anthropic 可以不做图像生成,也可以没有 IMO 金牌模型;可以卖光“思考帽”,仍然因为坚持做自己而获得喜爱。与此同时,那句诚实的广告牌文案——“我的老板真的很想让你知道,我们是一家 AI 公司”——捕捉了职场中普遍存在的困惑。
5. 市占率衡量机会,经济性才支撑估值
Menlo 的企业调研显示,OpenAI 在 LLM API 支出中的占比从2023年的约50%降至2025年年中的25%,Anthropic 则从12%升至32%。主持人强调了两点:统计对象是 Anthropic 所瞄准的企业 API 支出,而不是 token 份额;数据来自对大量企业用户的调研估算。
更重要的结论是多元化,而不是 OpenAI 崩溃。企业现在有多个可信的前沿模型可选,一旦某个模型适配生产工作负载,它们往往会预留长期算力或专用实例。这种承诺让企业行为远比业余开发者更具黏性,后者可能每周都在不同模型之间切换,追逐当时看起来最好的选项。
当被问及当前投资人应如何评估 Anthropic 时,Deedy 拨开了虚荣指标:“这是收入,这是利润率,这是增长轨迹。”市占率主要帮助估算 TAM 的上限;要为一轮超过1700亿美元估值的融资承销,还必须相信公司正在进入或准备进入的市场真实存在。
他的心态仍然刻意保持偏执:眼下结果不错,第一反应是“很好,现在让它持续下去”,以及“下一步是什么?”价值在于未来的模型、产品、分发和进一步的份额增长,而不是把今天的百分比当成不可逆的“翻转”。
6. 编码让模型智能保持可变现,应用护城河却仍然很薄
Deedy 认为,更强的通用智能可能已经无法继续提高大多数消费者聊天用户的留存。需要前沿推理能力的用户或许不到1000万,而 ChatGPT 被引用的8亿用户中,很多人只是想修洗碗机或改写邮件——这些任务已经做得足够好。就这一限定意义而言,OpenAI 在消费者聊天领域“基本已经赢了”。
编码不同,因为它的质量前沿可能会无限向前。Anthropic 可以把更好的模型转化为更好的编码产品和更多收入;相比之下,再增加一档智能,未必能推动一个成熟的消费者助手。但 Deedy 仍然提醒,成本和质量—价格帕累托前沿依然重要。
他拒绝讨论 Claude Code 的利润率,但将 Anthropic 的大致策略概括为“快速扩张、保持便宜、让所有人都用上”。低价接入支持 Cursor、Devin、Cognition、Bolt、Lovable 等企业;同时也给 Anthropic 带来通过使用量和数据飞轮改进自身产品的机会。
主持人称 Claude Code 是使用 Claude 的最佳方式;Deedy 则反驳说,Cursor 和 Devin 仍拥有忠实用户。他更广义的护城河测试并未因这一分歧改变:当前应用层还不够厚,无法阻挡一家分发能力强的实验室进入;而应用公司却很难复刻模型实验室。长期看,类似 Amazon 的风险在于,生产环节的所有者最终会进入客户所在的品类。
7. Rahul Patil 的崛起,挑战了由学历驱动的上限
Deedy 认为,印度学术竞争在文化上类似美国体育,因为教育被广泛视为社会流动的通道。大约100万人参加 JEE 工程考试,前1万人进入 IIT,最终只有约200人进入计算机科学专业。这是一套极端的排名系统,其标签可能伴随一个人多年。
他担心的既是制度问题,也是心理问题。有些职场会根据员工过去取得的成就,而不是当前工作的质量来评判他们;人们也会内化早期被拒绝的经历:“我没进好大学,所以我很蠢,因此我不该那么努力。”
Rahul Patil 没有就读 Deedy 认为的印度顶尖大学,却成为 Anthropic CTO,这本身就是一个反例。主持人补充了重要限定:摆脱学历路径,同样需要进入强公司、抓住机会,还需要运气。Deedy 更狭义的主张是,任人唯贤的环境可以让持续努力推翻早期限制。
8. Anthology 将生态战略与投资判断分开
Menlo 与 Anthropic 在前一年年初前后设立了1亿美元的 Anthology Fund。他们有意将基金放在 Anthropic 之外,因为内部企业创投团队需要单独配置人员,而且可能会优先优化母公司产品的使用量,而不是投资回报。
目前组合约有40家公司,包括 OpenRouter、Goodfire、Prime Intellect 和 Wispr Flow。Deedy 表示,Anthology Fund 被投公司继续进入下一轮融资的比例显著更高。
基金的投资范围分为3类:具有战略重要性的企业、重度使用 Claude 且自身具备吸引力的公司,以及潜力很高的超早期创始人。Anthology 不要求使用某个特定模型,单笔投资可以从10万美元的跟投支票到2000万美元的领投。
较小的初始支票让 Menlo 可以先建立关系,之后再考虑领投后续轮次;基金举办的活动则让创始人与 Anthropic 的创始人和高管直接连接。Anthropic 已从一家无人知晓的实验室成长为大型平台,仅靠接近平台获得信息优势的价值下降,基金仍在重新思考如何继续为双方创造价值。
9. 研究投资,是从一个必然到来的未来倒推
Deedy 称,研究投资极其困难,但潜在回报也可能非同寻常。董事会层面反复出现的矛盾是:究竟应该把一项有前景的能力变现为数百万美元 ARR,还是继续资助可能带来更大产品的研究:“我是应该开始做点什么,还是继续让研究下注跑下去?”
他的方法是追随异常有能力的人,然后快进10年,追问届时什么事物高度可能存在。如果未来需求足够有说服力,而团队正沿着一条合理路径推进,他就能在今天的研究与未来业务之间画出一条“虚线”,但不会假装这条路径确定无疑。
Goodfire 的价值在于,影响重大的模型仍然是黑箱,现有评测描述的是输出,而不是内部原因。对于贷款、保险或法律决策,“模型就是这么说的”并不够。机械可解释性或许能检测模型内部的迎合、撒谎、窃取或说服行为;Deedy 的简化说法是“给 LLM 做脑外科”。规模不是瓶颈,拿到模型权重才是。
Prime Intellect 承担了最初让主持人否定分布式 AI 的那些风险,Deedy 也拒绝把它包装成必然成功。讨论认为,分布式训练、人才获取,以及算力之外尚未实现的更广阔愿景,都可能构成上行空间。在每3到4周就变化一次的市场里,Deedy 说,自己若要精确描述它最终会变成什么,反而是愚蠢的。
10. OpenRouter 通过运营细节不断积累开发者心智
OpenRouter 是 Deedy 眼中的“心头好交易”——他进入创投行业时最希望自己做出的公司。其创始人此前创建了 OpenSea,Deedy 称后者巅峰估值超过100亿美元;之后,他转而解决一个工程师起初以为很简单的问题:可靠、细致地接入一批不断变化的模型。
Deedy 认为,任何可行的网关都必须产品驱动:用户应当能够自行使用,不需要与销售沟通。OpenRouter 面向开发者的首页、使用数据,以及没有泛化企业导航的设计,都表明其创始人理解目标用户。Deedy 前往纽约,在遭到忽视后发送“情书”,承诺会促成未来的一轮融资。
主持人称,该公司的当前模式是抽取约5%的路由金额。最清晰的两个风险是模型价格下降导致收费池收缩,以及留存偏弱:业余用户会流失,企业则可能先用 OpenRouter 比较模型,之后直接与胜出的供应商签约。
针对 Vercel 的 AI Gateway,Deedy 认为,对于更广泛的平台,网关仍会处于次要位置;而 OpenRouter 已经拥有开发者心智,并补齐了被忽视的细节。用户可以只路由到不保留数据的供应商,也可以在上下文窗口、质量、延迟和吞吐量等维度比较同一个模型。排行榜提供了额外分发,但 Grok Code Fast 这类免费发布会抬高表面上的受欢迎程度。
11. Wispr 与扩散模型,检验执行力能否战胜商品化
Wispr Flow 处在看似已经商品化的语音听写领域,但 Deedy 认为它是速度最快、准确率最高、体验最令人愉悦的实现。按住功能键即可生成文本,像“I didn’t mean that”这样的自我纠正会被自动处理,其内部“零编辑率”据报超过80%。
主持人追问 Superwhisper、Granola、Notion 和 ChatGPT 都在加入邻近功能的问题。Deedy 没有给出绝对的护城河判断,而是指向用户喜爱、留存,以及可靠语音终于可能让“说话”成为舒适的主要界面——说话毕竟比打字更快。
另一个下注项目没有公开名称,只以“Stealth Co.”称呼。主持人将其描述为一种扩散模型路径,并估算当前扩散系统可以用十分之一的成本和延迟,达到当前质量的80%–90%。对于需要高速和可接受质量、而非前沿性能的高频应用,这一特性可能很有价值。
代码可能适合扩散模型,因为其依赖关系是双向的:程序员会在文件中上下移动,检查变量和结构,而不是严格从左到右推理。主持人则以 Transformer 的“硬件彩票”回应:如果沿一条研究路径多走4年,可能就再也追不回来。讨论最后认为,市场经常奖励拥有动能的技术,而不一定是内在上最优的想法。
12. 市场时机与资本,可以制造出它们预期的赢家
主持人用一名在隧道中奔向光亮的跑者来比喻市场动态,而隧道正朝光源方向收窄:即使是最快的跑者,也可能逃不出去。创始人可以拥有优秀想法和出色执行,却没有足够时间在更大的力量封死入口之前挤进市场。
MosaicML 展示了时机问题:它拥有强大的微调团队,却面临孱弱的开放模型、糟糕的客户数据和有限的专业能力;不过,被收购仍然可能带来很好的结果。新的 RL 环境和 RFT 或许会重新打开窗口,但原来的市场并不会因为技术优秀就自动准备好。
AI rollup 暴露了品类套利。一家买方可能以200万美元收购一家由人力运营、ARR 为100万美元的公司,承诺实现自动化,然后在交付之前就获得1亿美元的“AI 公司”估值。主持人的反驳很实质:客户和领域专业知识才是硬资产,而新股权可以为工程师提供资金,让最初的判断最终成真。
这就是反身性:资本可以验证一个叙事、招募员工、阻吓竞争对手,并制造出它原本假设会出现的赢家。算力领域也存在同样争论:讨论提到 OpenAI 一年支出70亿美元,其中20亿美元用于推理、50亿美元用于研发,并追问未来需求是否足以支撑正在建设的基础设施。
13. 算力充裕,仍然需要经济需求形成闭环
主持人看到的是一项巨大的实物投入——芯片、土地、电力、Anthropic 的 Amazon 基础设施,以及 Stargate 规模的计划——这说明复杂的市场参与者预期未来存在需求。与依赖人力的投机不同,数据中心投资可以作为有形基础设施进行建模。
Deedy 从需求端倒推:即便 ChatGPT 每周有8亿用户,他们目前也不需要那么多推理算力,因为大量请求只是基础问答。要实现“每个人一块 GPU”,就需要更多代理式工作,或显著更好的模型来创造新的使用场景。
尚未揭晓的下注是:更多算力带来更好的模型,更好的模型创造更多需求,更多需求再为更多算力提供资金。风险也很明确:如果新增研发支出无法带来有意义的智能提升或经济收益,那么仅凭今天的 Claude Code、Codex、ChatGPT、Sora 和 API 工作负载,庞大的产能未必合理。
因此,真正具有颠覆性的突破口可能是研究效率:就像 OpenAI 曾经冲击那些已经重金投入、却没有交付突破的大型 incumbents 一样,新的实验室也可能冲击 OpenAI。“你的利润就是我的机会”会变成“你的研发低效就是我的机会”,但没有人声称下一个实验室一定会成功。
14. 编码代理可能同时削弱软件安全与工程判断力
一位主持人讲述了一场据称真实的求职面试:候选人被要求克隆一个仓库、运行它并进行修改。Cursor 据报发现,其中一个字节数组会编译成一条链接,从而外泄私人信息。AI 识别出了这个陷阱,但开发者如果不检查就执行陌生的生成代码,会制造大得多的攻击面。
更深层的担忧是手艺。工程师过去通过长时间的挫败积累能力,再从解决难题中获得满足;代理则把这套循环替换成不断触发的“请修好,请修好,请修好”的“老虎机”。Deedy 认可这一担忧,并将其比作“给大脑抽烟”。
主持人不接受单纯禁欲式的答案:团队仍然必须关闭工单、合并 pull request。他们将 Claude Code 高度异步的行为,与那些在困难推理过程中始终和人类保持心智同步、提供不打扰式辅助的快速代理进行对比。
他们提出的绩效公式很简单:找到正确的文件,然后写出正确的文件。抬头显示式代理可以改善阅读和理解,把写作留给人类;Cursor 的可见 diff 和最终确认,则是另一种有人参与的模式。Deedy 最担心的,是一个还无法判断模型何时创建了4个多余文件的18岁学生,把这种错误当成正常工程实践学会。
I entered venture, and I was like, that is the company I would have built from 2019. I remember going to parties in the Bay Area and saying “enterprise search,” and that would shut down the conversation right there. Anthropic is the fastest-growing software company of all time. When we invested in the company, it had no revenue in India.
Academics holds the same sort of prominence that sports would hold in America. On average, people are quite poor, so education is seen as the means to social mobility. The way it works is similar to countries like China and others: you take a big exam and get ranked. A million people take the JEE engineering exam, the top 10,000 get in, and the top 200 get into computer science. That’s how hard it is.
Those top 10,000 get into IIT. Everyone’s heard of that. That’s where a lot of the great Silicon Valley people, from Sundar Pichai to many others, come from. You look at a guy like Rahul Patil, who’s become the CTO of Anthropic, and he’s not from a top university in India. He worked his way up to a position of such prominence, and that’s testament to the fact that even though you didn’t have the opportunities early, and even though you might not believe you could do it, if you work hard enough for a long time on things you care about, anything can happen.
Hey everyone, welcome to the Litter in Space podcast. This is Allesio, founder of Kernel Labs, and I’m joined by Swix, editor of Laid in Space.
Hello, hello. And today we’re finally joined by the epic return of Deedy Das. Welcome back.
Thank you for having me, guys. I’m so glad to see you. All of us have different jobs now.
All different jobs. Classic Bay Area. It’s been 2 years, right? Last time, it was April 2023; you joined us remotely, and you were still at Glean back then.
I was also looking at the Claude timeline. Claude 1 was March 2023, and Claude 2 was July 2023. It just feels like so long ago.
Man, I remember the time. I don’t know when your first experience using Claude was, but mine was—I remember early Glean, somebody from the company was like, “Hey, there’s this interesting new LLM that’s not OpenAI, and the only way you can talk to it is by tagging Claude in a Slack channel.” That’s a bizarre interaction model for a whole new product.
The best model.
And fast-forward to now, and I’m like, okay—
We’ve come quite a way.
Yeah. I think they only recently introduced Claude in Slack, right?
Like publicly?
Come back. The comeback.
Yeah, yeah, yeah.
It’s like how it started, and now Claude is in Slack—Claude and Slack. And so, since then, I wanted to start with Glean, obviously, because we’re going to cover a lot of startups in this episode. Glean was, like, $1 billion, I think, based on my research, and now it’s at $7 billion. So your options are good. What’s your take on how Glean’s going and the market in general?
1. Glean Builds A Search Moat
I would say that now, being on the venture side, I have a bit of a different take than I would have had at Glean. But broadly, one of the things that I love about Glean is that it’s such a boring, unsexy company that became sexy later.
From 2019, I remember going to parties in the Bay Area and saying “enterprise search,” and it would shut down the conversation right there. Nobody would ever ask a counter-question if you said enterprise search. They were like, “That sounds boring as hell. Leave me alone.”
Fast-forward to 2022, and enterprise search got more conversations. People were like, “Interesting. Tell me how you’re doing this search.” What was nice about that observation is that in those 3 years, we did a lot of work and didn’t take shortcuts on a lot of things that ended up generating a lot of value for us.
If you look at Glean from a high level, the business is top-down enterprise sales. It’s very hard to rip and replace, and we expand contracts very easily because the TAM is so large. Every knowledge worker could use a version of enterprise search. Then there’s the AI on top—I still call it search, but it’s information retrieval in the enterprise—and we solved a lot of critical problems in order to get there. I can go into that, too.
Then comes December 2022, the ChatGPT moment, and everything that’s happened since. When I look at Glean now, it’s a different world. We were very quick and correctly prioritized LLMs early on. It did a lot of good for our business.
But now there’s fire from a lot of angles. Everyone wants to be a part of the enterprise search story, and it makes sense. It’s a large, unconstrained TAM. LLMs are particularly useful for gathering information. Obviously, consumers are interesting, and enterprises are therefore interesting. How do you do this in the enterprise? Gather all the knowledge and then put an LLM on top.
That being said, I’m still very happy with Glean stock. Glean’s also valued at $7 billion, not $100 billion, so I think the company has a lot of growth. I think it’s done a lot of the hard work that nobody’s willing to do.
I also think VCs have a tendency, including myself now, to trivialize a problem into a 1-sentence narrative. With Glean, that narrative was often, “Oh, well, you guys built this enterprise search thing, which never worked, and then AI came along and it started becoming a thing.” I really think we did all the hard work to build search, and AI happened to accelerate our go-to-market motion at the right time. Now I see companies trying to tack on search. It’s not easy.
I know the last-mile stuff we did for some of our customers, and I just know that when I think about other companies, I’m like, would you really go all that distance? It’s not a moat. The moat is just that we did the hard work. I’m pretty happy. Things can go in any direction, but I’m pretty happy with the way Glean’s going right now.
And just to spell out the 2 main challenges, one is obviously Claude. I think it launched enterprise search today.
I was going to say, I have screenshots.
Did you see, like, “Hey, we’re introducing enterprise search”? I’m like, “Yeah, son of a gun.” And then on the other side, you have the data providers adding these rate limits, kind of like Salesforce has done with Slack. It feels like that part is more challenging than the competition from other companies. How do you think about that?
2. Enterprise Search Faces Real Friction
Two questions, I guess: competition and rate limits. On the rate-limiting side, this has happened for several SaaS tools.
I think one advantage that Glean has is—well, the first thing, let me address the premise of the argument. When I think about why SaaS tools would limit API access, inherently, it never made sense to me. I can see why you do it for business reasons. Maybe you want to launch a competing product, but Glean doesn’t eat into your revenue.
If you’re Slack and you’ve sold, call it, 100 seats at a company, and you have Glean at that company, Glean only shows Slack results to the 100 seats that you’ve sold. So we aren’t eating into your business.
From a first-principles business-logic perspective, I don’t see why you’d do it. If Glean is on Slack and more people are searching through Slack, it actually lets you sell more seats, not fewer, because we don’t reveal permissions to people who don’t have access.
If we were to do that, then I could see maybe a business case, like, “Oh, you’re taking the Slack data that I’ve only sold 1 license for and showing it to 1,000 people.” That’s problematic, but we’re only showing it to the licenses that you’ve sold.
The second thing is that we do have thousands of integrations. In a lot of enterprise customers, Slack is really important, and that’s a critical data source, but we also have many, many more. It’s just the law of large numbers. Maybe if everyone decides to shut it down, it could be more problematic, but if 1 person does, it’s less so.
The third thing I’ll say is that if you talk to the customers, they’re also super unhappy about this. They’re like, “Look, we bought your product. We own the data. You don’t own the data. If we want to buy another product to use our data in Slack, why can’t we do that? Why are you blocking the API?”
Those are the 3 prongs of the argument. I can’t know how this will all end up, but I don’t think it’s that sensible that it is like this, and I’m still optimistic that we’ll clear out some of those issues.
Yeah. Anything else you want to say? Obviously, we’re about to move to Anthropic, and Anthropic just launched enterprise search. What would you say, as a veteran of enterprise search, that Anthropic should take note of?
The question of the labs competing with Glean has always been a thing.
Sam Altman—like we were just discussing earlier—once came out and said, “If you’re an investor in OpenAI and 1 of these 5 companies, including Glean, we don’t want you as an investor.”
But yet, here’s what I see. Look at the revenue of Anthropic and OpenAI right now. These are billion-dollar-revenue-scale businesses. Glean is a several-hundred-million-revenue-scale business.
So the way I think about it—and this can even allude to how I think about startups, to compete and to win—is this: for Anthropic and OpenAI to build a deep enterprise search system, it doesn’t make them that much money. They have to put all this effort into making what, an incremental $100,000 sale, $200,000 sale, maybe even a 7-figure sale.
Is that moving the needle on your $5-plus billion in revenue, or $10-plus billion by the end of the year for OpenAI?
Not really. And the amount of effort it takes to get there is big sales teams, huge FTE teams, tons and tons of customization. My question is, in the long, long run, you could build a semireasonable enterprise search tool. If you really want to go deep, I don't think you will ever dedicate the people to do it.
The last thing I'll say is, think of it from an Anthropic engineer's perspective. You joined a big AI lab to work on models, not to build Google Drive connectors, right?
A meme like, you know, “I build the integrations.”
Build the integrations. I think I'm still very bullish, but, yeah, competition happens.
Yeah, yeah. Actually, I wasn't asking about competition. It was more about what the hard problems are that people don't appreciate.
Oh, okay. We can talk about that.
That was probably a safer category for you.
Basically, I'm in this boat as well. I've joined an enterprise AI company that has to worry about and build for these issues. I'll just give you one example. Until this point, we never had to deal with 2 Slacks. Enterprises have, like, when you acquire another company, different systems, and they all duplicate and overlap.
Yep. Oh man, I have some great stories about Devin. I'm sure there's some power-user version of this, but I still haven't figured out how to use Devin properly with 2 Slacks.
Wow. Cases like that remind me—that was a thing we had to address at Glean. I think every enterprise company has the same sort of hurdles.
We looked at each other and were like, “Oh, yeah, we're a real enterprise now. We have 2 of everything.”
That's funny. Okay, Glean: a bunch of interesting problems.
I'll talk about some of them. If you want to prod, feel free.
I think the number 1 most interesting thing to me when I joined the company was that consumer search was largely regarded as a solved problem. Not really, but largely. The way most consumer search systems work is by aggregating feedback data on how users use search—whether they click, hover, and how long they stay on a website. That's what powers ranking systems to get better over time. It's a very powerful, critical way that Google, Bing, and all of the above work.
In enterprise, if you take a 10,000-person company, even if every user issues 2 search queries a day—which is quite a lot, say even 5—that's just not enough volume to have any meaningful quantity of feedback for this to be relevant. On top of that, freshness is way more critical in the enterprise in certain ways. There are more freshness-seeking queries in enterprise than there are in consumer.
And number 2 is that the distribution of queries in consumer is very head-heavy. It's not that way in enterprise. In enterprise, maybe the query that everyone wants to search for is “benefits” or “payroll.” That's not really that useful. Every person is doing a different job, and they have different needs and different things they want to look up.
Given all of that, the techniques under the hood that work for consumer don't translate to enterprise. You have to invent a whole new set of signals that actually makes enterprise search work, and evaluation becomes very, very difficult too.
In consumer search, you have tons of data to pick and choose from to evaluate what's the right result to show for a query. In enterprise, we would look at some of our customers' data and look at each other and go, “We don't really understand what this query means. We don't really understand what these results are. We don't know what the right ranking is. We have actually no idea what we're doing.”
It's so out of domain for even us. Some of our customers are working on very, very specific problems. All of that is one huge, huge challenge: how do you make ranking work in enterprise in a great way?
The second interesting one is that selling productivity tools to enterprises is challenging because, no matter what ROI argument you make, people aren't actually buying tools for ROI. People buy productivity tools because their users like using them.
For example, when people buy Slack, I don't think any buyer is going, “Let's measure how much faster or how much more productive our team is getting by using Slack.” It's probably not even getting that much more productive. That's not what they're looking at. They're saying, “Everyone uses Slack. It's pretty useful. I'm going to keep Slack. I don't think we're going to churn that one.”
If you take that analogy to search and search systems, the issue is that search systems aren't inherently viral or growthy. Slack has a very clear virality moment: everyone is talking to everybody else, and so that's just how you have to speak.
In search, it's kind of a one-player game. You're not really sharing things, and you're not really talking to everybody else. The challenge for us was, how do you sell a productivity tool by getting everyone to love it on day 1 for a product like search? It's not easy.
If you look at how Google did it, they had Chrome. It was a great source of distribution. Get everyone to query, and then they'll hopefully learn to love it. We had to figure out what that meant in the enterprise as well, and how to get everyone to adopt, embrace, and love this new tool.
Yeah. So, 2 of the many good pointers. Just a question on that: was there anything—because you have a new search tool, it's like, “Go search,” and it's like, “What am I searching?” What was that blank-canvas onboarding for people? Anything good?
Several different things worked well for us. I can think of 2 at the moment, but I'm sure there were many, many more.
I'll say one of them was that, for a handful of companies—for many companies, actually—we would say, “We want to take over your new-tab page.” The critical part was, “Tell us what we need to do to earn the right to do that.” No one wants to give away their new-tab page.
We went the last mile. There were companies who were like, “Well, we have a new-tab page. We're pretty happy with it.” So we'd ask, “Do you have a search bar on it?” They'd be like, “Well, yes.” I'd say, “Okay, what is that using?” They'd be like, “Well, it's using our internal thing.” I'd say, “Do you like it?” Clearly not. That's why you're talking to us. So let's just rip and replace that.
Doing that extra mile was pretty important. So that's one: new tab.
The second one that we liked was a Chrome extension. When you were on your native product and issuing a search query, we ran a lot of evals. We thought we were better at every product's own search.
So if you were searching on Google Drive, we would do a Glean replacement of the search bar and the page pretty natively. It would teach people to use Glean and be like, “Okay, this is pretty useful. I think these results are great.” It automatically filters through Google Drive anyway, so functionality isn't lost. We would slowly get people into the ecosystem that way.
Yeah, superset adoption—something that OpenRouter also does.
Okay, so Anthropic: we have to obviously address the elephant in the room. You guys are huge, huge Anthropic investors. I think right after you maybe got promoted or became a partner, you guys led the Series D. What's the chronology there?
I think we did part of the Series C and then the D, and then every single round after that.
Yeah. Obviously, one of the greatest companies in AI. I honestly had no idea that we would be sitting here—Anthropic has 10×-ed in the time that you've been at Menlo—and I just... What's it like being an Anthropic investor? What are the considerations back then versus now?
3. Anthropic Defies Startup Expectations
Anthropic is the fastest-growing software company of all time? I think I can say that fairly. I haven't been disproven yet.
People say that, but everyone says, “We're first to $1 billion, first to $100 million.” I don't know. It's hard to tell.
I do believe the numbers are $0 to $100 million in 1 year, $100 million to $1 billion in 1 year, and this year it would be $1 billion to the public projection, which is $9 billion.
Even to this point, I know a lot of people—we've seen the graphs on Twitter—a lot of that is [?], some of that is GMV, all this other stuff. But in Anthropic's case, I think it's fairly legitimate revenue, and I do think it makes it the fastest-growing company, definitely at the $1 billion-plus scale. I can't think of too many examples.
So it clearly has outdone itself. I would say that when we invested in the company, it had no revenue. I mean, that's just a fact. When we wrote our first investment, it had no revenue. It was a $4 billion valuation, right?
It's been fascinating to see this company succeed. I couldn't have predicted it. All of us—this was beyond our wildest expectations. I think whether or not it continues to perform at this rate, I believe it will, but it is already somewhat of a generational company in many ways.
Kudos to the team for delivering these awesome results. One of the risks, kind of taking a tangent, with a company like Anthropic is that you essentially had a team of extremely idealistic researchers. Very often, the standard deviation of outcomes when you have teams like that, or similar to that, is quite large.
There was a world where maybe they would not have worked at all and would have absolutely fizzled to the ground. But I think the same qualities that gave them a high propensity to fail gave them a high propensity to succeed. And if you look at it, there are many other things they did right, but if you just look at a product like Claude Code, there aren't many product innovations in AI that I can think of that are as critical as something like that, because we had the whole chat era of RAG systems and ChatGPT.
That was a critical innovation, but since then there have been a lot of followers, a lot of Deep Research, which is, I would say, an addendum, and a couple of other things happening here and there. Agents—cool. But if you think about agents that actual end consumers use and gain value from, in my mind at least, Claude Code was the first time I saw that, in a terminal, in a weird interface. It was just weird.
It was like every PM's nightmare. No PM would have thought of that. And so it's—
Except for Cat Wu.
Yes, except for Cat Wu. And so, you know, it goes to show how Anthropic is able to function as a company and innovate like that, which is quite rare, especially at that scale.
To some extent, I think you just hire good talent and then let them loose with a lot of tokens, see what they come up with. They tend to build good stuff.
Well, it's interesting to talk about, because take OpenAI and DeepMind as a comparison point. I think we'd all agree they all have great talent, but they don't all innovate the same way. It's always been interesting, just as an academic exercise, to think about different leadership styles.
Maybe, from the outside looking in, you'd be surprised how little I actually know from an investor standpoint about how Anthropic operates, but it seems like a company that has such high employee-retention numbers because they are very free-spirited in how they let employees guide the direction of the product, versus other companies which are much more either top-down or prescriptive: “We need to go after this, and we need to go after that.” It's, “Let's see what happens. Try.”
Yeah, I think at my last conference, SignalFire had some stats. They track all the LinkedIn pages of everyone, and Anthropic has the best retention. It's a net gainer, whereas everyone else is a net donor of employees to something like that. I'm referring to the exact same article, where I think their one-year retention of employees is 80%, which in the AI world is quite wild.
Yeah. And I mean, Anthropic does not have image generation. They do not have an IMO gold-winning model. I feel like they don't—they just do their own thing. They do it great.
They have nice hats.
Yeah, they sell out Thinking Caps. So actually, I really wanted to discuss this, but I don't know how to. I think I need to get some marketing or PR agency person, because people actually forget that in 2024 they had out-of-home advertising campaigns, which sucked. Everyone was dog-piling on them, and then this year it's slightly changed.
It's still Anthropic, but slightly changed, and they decided to focus on thinking, and suddenly everyone loves them. They have the cafés and all that. It's a very interesting public-image rebrand, and I don't know if it's because the models are just better or it was actually PR. Which one comes first, chicken or egg—models or PR?
It's a good question.
Yeah.
4. Anthropic Finds Its Own Lane
It's a good question. I would say, though, ignoring the model side, I do think this one is aesthetically better.
Yeah. Purely, it looks nicer.
Yeah, and the vibes. I don't know—I have sat in those meetings, and someone's pitching you an idea and you're like, “I don't know. Looks good, okay.” Then it becomes one of the most hated campaigns of all time, and one year later someone else comes with a slightly different-looking idea.
The words are different in 4 ways—they chose slightly different words, but it's not that many words—and suddenly that one is the one that works.
Well, as somebody who writes online a lot, I can relate to how a couple of things being different can be the difference between something people care about and something they don't. Early at Glean, I had such run-ins with marketing, because the first campaign we actually did was just, “Really AI for work that works.”
Okay. Like—
Was that a hit?
No. I mean—
In enterprise, how does one even measure what is a hit and what is not? No one really cares enough, I feel, one way or the other. But we've all seen really cringe AI ads. If you've seen the Cisco ad in the airport, I hated that one for a while.
All kind of generic. So, anyway, I like the Anthropic one.
Okay, I'm going to sprinkle in some of your tweets. You had one ad about the billboard where the Reddit guy was like, “My boss really wants you to know that we're an AI company.” I thought that was the single most honest billboard I've seen in San Francisco.
Absolutely. I think it's a testament to all the comments of people going, “Yeah, I relate.” We've all heard it. Everyone, it feels like even on the technical side, is struggling to catch up and gain a sense of meaning again.
I've had developers go, “[expletive], man—is this it? What do I do anymore?” And even that's happening on the technical side, with people who semi-understand what's going on. On the non-technical side, people are like, “So there's this new thing. It's AI, and generally my boss literally just wants me to do something in it. I don't really understand, other than chat is quite helpful.”
Yeah, I have some charts. I don't know if you have any of these in mind, but I'm just going to bring up some of the Anthropic charts, which I think—I just want to put it on the record for people who are not paying attention, to understand.
In 2023, according to—these are Menlo numbers, right?—OpenAI's market share was 50%. And in mid-2025, you guys have OpenAI at 25% market share, and Anthropic was at 12%, now at 32%.
Enterprise API market share.
Correct. So I should clarify that that is enterprise LLM API spend—
The market that Anthropic happens to focus on. Yeah.
And critically, it's also spend numbers, not token numbers. So I think those clarifications are important, and also the methodology is based on surveying vast amounts of enterprise users on how they are doing their spend.
But that being said, yes, the point remains.
Market share for OpenAI has gone down. It's not a negative; obviously OpenAI has done super well. It's just that diversity has gone up. It used to be there was basically only 1 choice, and now there are 3 or 4 legitimate frontier labs, maybe more than that if you count all the open models as well.
But I think it's just super interesting and under-discussed still that you can actually build a sustainable advantage as a frontier lab.
You know, I'm sure you guys remember there was a lot of conversation at some point about the commoditization of models, and to an extent maybe it's happened. Models—a lot of the frontier models—are neck and neck on a lot of things.
But in practice, and this data was in that market map and market survey as well, once people like something and get used to it, they don't really churn off it once it fits their needs. And so we've seen a lot of that. There's a lot of churn in the hobbyist-developer-type category.
But in terms of enterprises, often what will happen is they'll buy up large chunks of long-term compute and dedicated instances, in which case you just don't churn, right? This is what you use. So I think that's part of the effect.
And to commend OpenAI, they were just focused on something else: they have launched the most incredible consumer product that we've seen since God knows when. So they were probably not focused on enterprise until now, again.
Yeah. How do you re-underwrite the company internally as you invest? Even since we talked about Claude Code, I think that was a pivotal moment in the trajectory of Anthropic. What are the things that matter to you when you're looking at a company like Anthropic? Does this market-share number matter?
How do you evaluate both the opportunity and the numbers that you really care about, versus, sure, higher market share, but that's not what we cared about?
I don't think the market-share number is the most important thing. It is more critical to understanding the TAM at that stage, to be very honest with you. At the stage that we invest in Anthropic now, the only things that would really move the needle on the decision are: here's the revenue, here's the margin, here's the trajectory, and here's the other markets we may be able to underwrite that they want to go into, that they may be early in or planning on going into.
I think it's really difficult to underwrite on market share, other than knowing what the potential cap of the TAM might look like. So the pie will also expand potentially, but other than that, I don't think it's more than just a nice vanity metric.
Yeah. In your mind, is it kind of like how people in crypto are always talking about flipping Ethereum and Bitcoin? Does it matter that Anthropic can go to 50%? Or is it that OpenAI was only at 50% at a moment in time when it was a new market? I'm curious how you think about that.
I don't want to color the way Anthropic—or the way all of us—probably think about this, but I just don't think it matters that much. In my view, I'm a very paranoid person with startups, companies, and technology. So, in my view, I'm like, "Great, now let's make it last." Or, "Great, but what's next?"
To me, it's nice to have. Look, if we're investing in a round right now that's north of $170 billion, sure, it matters. Some of the numbers matter, but the future of the company is where all the value really is: what we underwrite as the future. The future means that I'm more concerned about what's happening next. What are the new models? How do you gain market share? What has to be done? What are the new products that are going to be built?
I'm less concerned about where it's at right now in terms of market share. But that's just me. I don't want to speak for others.
Yeah, I think the new models are really good. Opus 4.1, Sonnet 4.5, and Haiku 4.5 were all released in the last few months. It's really interesting. I think OpenAI and Gemini are in a bit of a price war, with the Pareto frontier that I track in terms of LLMs versus pricing. Claude can still charge a premium and still have a lot of market share, obviously, and I think that's just because they have a better model. People naturally gravitate to it, especially for coding, but also for other things.
I just think articulating what makes a model good is very, very difficult. Obviously, there are benchmarks and evals, and everyone has this attitude: "Okay, today it's your turn to be best at SWE-bench, and tomorrow it's my turn." It's really stupid. We're just talking about 0.12 differences in SWE-bench.
But I wonder, if you're talking about, "Okay, I am investing $13 billion in Anthropic for Series F to underwrite Claude 5," what does that have to do with it? What kind of conversation does that look like? I have no idea. I'm not saying that, but I'm just—
I would say that, despite what you said about the premium, I still do worry. I think cost is a concern for a lot of people, and so the Pareto frontier does still matter. I'm glad Anthropic is where it's at right now, but who knows where that changes.
When it comes to Claude 5 and thinking about the future, one thing I think about that's really nice is that we can take for granted right now that furthering the intelligence of models in ChatGPT, a consumer product, does not lead to more users or more retention. It only really applies to a thin slice of users who care about very smart types of queries. I would say maybe under 10 million. That's just a random estimate, but most of the 800 million users on ChatGPT are asking, "How do I fix my dishwasher?" or, "How do I rephrase this email that I have sent to somebody?"
That's done. We know how to do that. So what's interesting is that now we're at a point in consumer where, maybe it's too early to say, but OpenAI has kind of won. How do you catch up to something where model quality is not going to be differentiated? You already have the users, you already have the retention, you already have a great product, and people are paying.
The interesting thing about Anthropic is that if you look at coding, that's probably never going to be the case. There is always an increasing frontier of how good you could be at a task like that, and we're nowhere close to that frontier. So it's more possible to underwrite the quality of future models versus OpenAI, where it wouldn't be as much of a revenue driver on their consumer business as it would be for Anthropic.
Yeah. Talking about coding, let's just talk about it, because I think this is also a fun discussion. One, there's the question of what the margins of Claude Code are. There are some numbers, and I don't want you to get yourself in trouble. But then there's also how you think about the Claude wrappers. We've talked to Bolt and Lovable, but then I'll put Cognition and Cursor in there as well. How do you think about this market? Basically, there's a whole ecosystem of startups, and they have all done really well building on top of Claude.
I think it's great. I mean—
Sustainable, was it?
5. Claude Builds An Ecosystem
I don't see why not. I don't want to allude to the margin question, which is: Can Anthropic continue to do this strategy? I'm not going to comment on the margins, but if you're trying to build out an enterprise-friendly business, there are 2 broad approaches. You have high customization and high price, which is usually less scalable, and then you have low customization and low price, which is very, very scalable. In a SaaS world, I guess it's a Slack–Palantir continuum.
This is kind of different, but generally Anthropic wants to play here: scale fast, keep it cheap, and get everybody on it. If we trust that most people, or a significant number of people, will stay on Claude if they continue to build products on top of it, then I think that's a win for the ecosystem and a win for Anthropic. I don't see why they would care.
I think the interesting thing—and again, I don't know what Anthropic's future plans are—is that Ben Thompson obviously talks about this classic strategy: Every time you own the means of production, you will end up getting into the markets that your users use you for.
The classic Amazon example is that first you're the marketplace where people sell. You find all the places where you can sell things that are commodities at high volume, and then you start creating batteries and Amazon-branded batteries. Then you push out a bunch of people who sell batteries. That's a risk, I think, for those companies that use Claude heavily and rely on Claude to think about.
But at this point in time, we're too early. I don't think Anthropic is anywhere near thinking about that, because they're still very much competing with other models on that layer.
Yeah, playing a different game.
Yeah.
Yeah. It's interesting. Would you rather be an investor? This is basically model layer versus app layer. So far, the model layer has won, and I think there was an app-layer summer, and then now it's very much back to models again.
I like the discussion. I was at a dinner where somebody was talking about this kind of question, and I was thinking about it more at that dinner. Maybe this is an ill-formed thought, so feel free to push back.
Yeah, we're riffing. But when I think about moats, it's classic VC-startup banter in my mind. I think the moat is whatever is the hardest to do in any part of the stack. When I think about people who tend to dismiss—there are other aspects to it too—but people tend to dismiss the idea that the app layers will capture all the value, if the app layer is easier to build, I think the model layer is harder and therefore will naturally capture all the value, net of competition from other model providers.
Put a different way, it is far easier for Anthropic to try to go into one of the app spaces than for an app to try to go into Anthropic's space, which makes me feel like one is more defensible than the other, all else equal. I think both can thrive, and that's ideally what everybody wants.
Yeah. I think very brutally, as an investor and as a human with my own limited time on Earth, if Anthropic can go from $3–4 billion to $183 billion in 2 years, then everything else is a waste of time. You know what I mean? You really do want to get this right. You can't just be like, "Everyone's great," and hedge your bets. Sometimes you have to go all in on the right thing, and you spend a lot of time and effort identifying the right thing. That's what I'm trying to do more of these days.
I think the means-of-production thing is interesting, because Claude Code only makes sense to be built if it's the best thing. If Claude Code is mid, they're better off promoting Devin and Cognition to sell more tokens.
So I'm curious: As the market gets more competitive, in one way it's, "We don't want you to use Devin, because Devin supports all the models, and we end up losing some of the revenue." But right now, Claude Code is obviously the best way to use the Claude models, so it drives usage. I'm curious whether, in the future, there will be more pressure on, "Hey, this product actually needs to be great to make sense for us to invest our resources into building it again."
Yeah. So, going from model lab to model lab plus product company, which is what OpenAI has done.
I would push back on that. First, I don't think everyone would agree that Claude Code is the best way to use Claude. I've heard multiple people, even in the last few months, say, "I'm a Cursor guy," or, "I'm a Devin guy." People have their preferences.
So I don't think it's set in stone. However, Claude Code is also a great way to use Claude, and there are nice flywheel effects because once you capture the way people are using Claude Code, you also get so much data to make Claude Code better over time. I think those are the 2 main reasons.
At this point, maybe this is oversimplifying, but I can't think of too many apps that have a very meaty layer on top of the model that's very impressive yet. There are somewhat meaty layers, and it's getting there. It's a time thing as well, right? Most of these companies haven't existed for more than 2 years.
I think it gets there, but I don't think we're at a point where we're like, “Holy shit, that app has so much stuff, interesting things, and technology built on top of the model that it becomes so difficult for the model company to compete.” I think if Anthropic or OpenAI decided tomorrow to take on another app, given their distribution and engineering, and the fact that these layers are still not as thick as you'd like them to be technically, they could. Whether they should or not is a different question, but they could, and that's something I do think about.
Thank you for engaging in all these very meaty discussions.
Yeah, you don't even work at Anthropic, so I know we put you on the spot.
Yeah, but this is what I want to get on the podcast because a lot of people don't get the chance to talk about this, but this is a normal San Francisco dinner. The last tidbit on Anthropic I'll point out, which is more fun, is that there was a new CTO joining Anthropic from Dropbox. You’re like the king of Indian posting. What's the significance of this? Last time you were on the podcast, you talked a lot about the Indian university system and all that, and I love to see this guy rise up.
6. Indian Meritocracy Shapes Careers
In India, academics largely hold the same sort of prominence as sports would hold in America. Everyone talks about it. It is part of Asian culture; it's top of everybody's mind, it is something a lot of people want to be good at, and it's an extremely competitive society with a very large population. On average, people are quite poor, so education is seen as the means to social mobility by a large number of people in India.
The way it works is similar to countries like China and some other countries: you take a big exam and get ranked. A million people take the core engineering exam, the top 10,000 get into IIT, and the top 200 get into computer science. That's how hard it is. That's pretty hard.
Everyone's heard of IIT. That's where a lot of the great Silicon Valley people, from Sundar Pichai to many others, come from. In India, something I'm generally very curious about is the motivation of humans and what determines the outcomes in their lives and careers.
One thing I've noticed a lot is that there are some societies that are inherently less meritocratic, where you're judged so much for what you've done in the past that you're not allowed to prosper later. I think many work environments in India and other places in Asia can be like that. Number 1, you're not judged on the merits of your work; you're judged on the merits of what you've done. Number 2, there's a very strong self-fulfilling-prophecy effect.
I've seen people underrate themselves because they think they couldn't be number 1 at something.
It's like your own mentality.
It's your own mental block, where you think, “I couldn't get into a good college, therefore I am stupid, and therefore I should not work that hard,” right? People in the Bay Area are also like this. The Bay Area is kind of like Asia. I know people who grew up thinking, “I couldn't get into a good college, therefore I am stupid, and therefore I should not work that hard.” It's inherent that they could be smart; they just believe they're not, and that also has a psychological effect on their long-term prospects.
You look at a guy like Rahul Patil, who's become the CTO of Anthropic, and he's not from a top university in India. Some people obviously debate that, but in general, I don't think it's a really well-known university in India. He's come to a society that is quite meritocratic and worked his way up to a position of such prominence.
I don't know him or everything else he's done, but it's a testament to the fact that even though you didn't have the opportunities early, and even though you might not believe you could do it, if you work hard enough in certain environments for a long time on things you care about, anything can happen. I think that's why it resonated with so many people and why I wanted to share it.
You choose to work at Stripe and Glean and do well. I think choosing the right company is also very important. If you're not going to do the credentials path, you have to be lucky and selective and work at good places. A lot of people make that mistake, and I definitely did. I had good credentials, and I worked at bad places, and that is very interesting.
You work at a pretty good place right now.
Yeah, but I took a long time to get there. This is funny: I have this automated podcast research, and when it sends me the email about you, it says, “Deedy has a strong presence in AI and immigration.” Those were the top 2 topics that it talked about.
Yeah, let's talk about the Anthology Fund. It's a $100 million fund in close partnership with Anthropic. Talk a bit about that. I think people are really curious about how close that actually is.
7. The Anthology Fund Expands Access
We set up the Anthology Fund when we invested in Anthropic around the beginning of last year. The idea was that Anthropic was a very different company back then. It was a much smaller company, and they said, “Look, there's an incentive for us to run our own fund. OpenAI runs its own fund, and there's a developer ecosystem that we want to create around this. It's really nice to have great startups that are using Anthropic, close to Anthropic, and building around Anthropic.”
We had a discussion about whether they wanted to have it inside Anthropic or outside Anthropic. Having it inside Anthropic would mean a corporate venture fund. You'd have to hire for that and have a whole role. Typically, if you look at corporate venture funds throughout history—obviously, besides OpenAI, as a notable exception—they tend not to be very good because all they prioritize is who uses their stuff the most. That's not a good way to invest in companies.
We thought this would be better, because the incentives in corporate venture funds are a little bit misaligned. We did that, and now we look back at this fund. Obviously, Anthropic is in a very different place. We've funded about 40 companies. The rate at which companies graduate from when we invested in them to their next round is significantly higher for Anthology Fund companies, and we write both small and lead checks.
Several notable companies from the Anthology program have been OpenRouter, Goodfire, Prime Intellect, and Wispr Flow. There are quite a handful of pretty interesting companies there.
The other really nice thing about it is that it allows us to move fast on companies where we may not feel immediately comfortable or ready to write a full check. We can participate in a round, get closer, and hopefully build a relationship and lead the next round in the future. It also lets them get really close to the Anthropic ecosystem.
We have all these events with the founders, executives, and things like that, and people really enjoy hearing it from the horse's mouth. Anthropic is in such a different place now; it's no longer an unknown entity. The program gets a lot of demand, but people kind of know what they need to know. We're still working on how to make this program more useful and more beneficial for founders and Anthropic alike.
Yeah. I want to highlight this for Latent Space as well: how does AI change venture? That's something Alessio was exploring too. I don't really know how to categorize the Anthology Fund, because it looks like a kind of what Conviction is doing, or maybe what YC is doing, but later stage. Some of these already have their Series C's, and some of these already have their Series A's. Abacus is in there—is that our Abacus? No, that's a different Abacus.
What's the model? What are the predecessors that you draw inspiration from for setting up this fund, or do you just see it as a corporate venture fund managed by Menlo and somewhat funded by Anthropic?
You can think of the companies that go into Anthology in 3 categories. One is companies that are strategically important to Anthropic, and those could typically be somewhat later-round, somewhat bigger companies. Two are companies that are using Claude heavily and are just great companies to be in. Three is very early-stage founders with very high potential who may potentially be using Claude models and Anthropic, and so on.
We don't require people to use a certain model or the other. We keep it pretty open, and we do everything from a $100,000 check to a $20 million check. I think it's really broad in terms of what we can do, and we wanted to intentionally keep it that way.
When it comes to where we draw the line, there are some old examples, but I don't think they're really relevant. There was a fund called the iFund that Kleiner Perkins did with Apple way back in the day, which was kind of similar.
How did that turn out?
I don't remember. I don't actually have enough data on that, but that's one example.
You know the answer.
No, I'm sure there are some great companies that came out of it. I just don't know the details about who was in it.
So, yeah, that's kind of how it's been for us, and I think it's been a really great program. I mean, we were excited about the companies that we could lead the rounds in as well.
Yeah. I wanted to get quick hits for people who maybe never heard of Goodfire. I know them because I've invited Mark to my conference, and I've been to a bunch of their events. Actually, I'll just give you that list. Goodfire and Prime Intellect are in your research category, right? There are others with diffusion-based language generation and novel architectures. It's all over the place. Research is the wildest west of this. How do you view research investing?
8. Research Bets Need Conviction
I can talk about any of those companies briefly as well. But the way I view research investing is that it is extremely hard to pull off, but when you pull it off, the results could be very remarkable. One of the hard parts is the tension between whether you keep investing in research, hoping for something that yields a better result that leads to a better product, or whether you try to monetize and scale what you already have.
That's tough. It's a really tough thing to do, and it's a really tough decision to make when you're working with those founders and you're on that board. It's somewhat anxiety-inducing when you're thinking about this, even from an investor standpoint. Do I just get to a couple million ARR? Do I start doing something, or do I keep the research bet going?
The way I think about research investing overall is, honestly, to follow where the talented people have the most competence, and then have an idea around how this could be useful in what I call a top-down way. It's not really top-down, but the way I frame it is: if I fast-forward 10 years into the future, what do I think is very likely to exist, and what are the ways I can get there?
If I do believe strongly that there's something like that, and I believe there's a team very strongly headed in that direction, I can sort of draw a dotted line and go, “Okay, maybe we can see something here.” So that's how I broadly think about it.
So, concrete example: Goodfire is the most interesting one. Mechanistic interpretability—I didn't even think that was a market that was worth investing in, but obviously Anthropic does. They seem like they have good vibes. What's the summary of your take on the company?
The way I think about the company is that right now, almost all frontier and many non-frontier AI models are complete black boxes. We don't understand why they produce the outputs they produce. All of the evals and studies on them are empirical studies, not intrinsic to the model. So it's like, “Hey, here's the outputs we saw, and therefore this is the benchmark score, or this is how we think it did.”
If we believe as a society that 5 and 10 years in the future, these models are going to be critically important for making pretty heavy decisions—anything from whether somebody should get a loan or insurance to a legal decision—then I don't think the black-box approach is long-term scalable.
It's just not how society can function, where you throw your hands up and say, “Well, this is what the model said,” and then I ask it, “Explain yourself,” and it says this other stuff. Great. That's kind of what we have today; that's the best thing that we have.
Mechanistic interpretability is really going into the weights of the model and trying to figure out why the model did what it did. One of the more concrete and relatable examples of this that you may be aware of is that GPT-4o had this phase of sycophancy that a lot of users really liked, but it's one of those things that's not as easily detectable in an eval unless you're specifically testing for it. Even then, it's quite hard. It's very personalized.
It's not like any keywords might arise, obviously, but it is something that's quite easy to tell with even current interpretability methods. You can tell when a model is being sycophantic. You can tell when a model is trying to lie. You can tell when a model is trying to steal or persuade you of something.
I think if we further that research direction 2 or 3 years into the future, we will be able to understand why models say what they say. “It's brain surgery for LLMs” is my catchphrase, but it doesn't apply to LLMs only—it applies to all models. That is a pretty important insight into deploying AI at scale.
Yeah. And you don't know the business model yet.
You don't need to know it, as long as we figure out what to do. There are some ideas we have, but we're not ready to talk about them publicly, and some of them are working as well. It's not right to discuss them publicly.
Does it feel worthwhile to do this on such small models? I think most of the work is done on the open-source releases. How much of a gap is there between what they're able to do and then translating that into doing it at scale?
They've shown that even for the biggest open-source models—even for DeepSeek models—they can do it. In general, scaling is not the bottleneck. Obviously, access to the weights would be a bottleneck.
But they're in the Anthology Fund, so they can work with—
Anthropic.
But they don't have Claude weight access, though.
For listeners who want to hear more about mechanistic interpretability, we did a podcast with the mechanistic interpretability team—Emmanuel from Anthropic—so that's your 101 there. We'll do something with Goodfire at some point.
Prime Intellect is another very hyped company. You don't have to say it, but I know it's very much in the water that they raised a very large round. So I ignored distributed AI for a long time. It's usually crypto people coming over saying, “Hey, we have these GPUs all over the place. We will somehow ignore the speed of light, and you can use our GPUs to train models.” That's why I ignored Prime Intellect. I was wrong. Tell me why I was wrong.
You may not be wrong. Look, I could be the kind of person who shills all of their companies and says, “This is the best thing ever, and if you don't think it's going to be a $10 billion company, you're wrong.” Every company has risks at this stage, and Prime Intellect has its fair share of risks. Whatever went through your mind went through my mind when I was looking at that company.
I do strongly believe in—I’m sure you've seen this quote too—the quote that pessimists are probably right often, but they rarely change things. It's an easy thing to say, but when you're investing, it's something to think about. There are a lot of things that could potentially be wrong with Prime Intellect, for sure, but the thing that I really liked that drew me to them is: if they were right about a couple of things, what could go fantastically right?
Distributed training is one of them. Access to talent, I think, is one of the things that I underwrote for them. The ability to hire fairly great people away from other labs is really hard, and I think they can do that. The third thing I think is that there's a broader vision to Prime Intellect that is not yet realized, where the first step of that was distributed compute. We'll see if they realize that.
Yeah.
Well, Will Brown's been on the podcast multiple times. They've launched kind of a verifiers SaaS platform or something, or a marketplace. I'm not really sure what exactly. I should probably try it out, but it's very interesting.
I mean, the other thing I'll just say out there is that everything in AI changes every 3 or 4 weeks. I'd be a fool to say that I could tell what this company is going to do.
Yeah. All I'm trying to do is capture for people who are not in the loop that these are the companies that people are talking about.
Okay, so let's at least hit on OpenRouter and maybe one more of your choice that is less known but you want people to know more about. OpenRouter we have to cover. Big deal. Obviously, I do think this is one where I was relatively early on. I saw the product, I saw what he was trying to do, and it clearly has done really well. I did not know he was taking investment, or I would have invested.
He wasn't.
Okay. Say more. Say more.
9. OpenRouter Finds The Sweet Spot
OpenRouter was sort of my—I don't want to make this about me; it's really about them—but in my mind, it was my darling deal.
You're proud of it.
Because I'm just like, man, I entered venture and I'm like, that is the company I want to have built.
I think we're skipping a bit. Let's explain who Alex is and what he did before.
Right. Let me give you the background on OpenRouter.
Alex is a phenomenal founder. He started a company called OpenSea, which was the NFT company. Obviously, at its peak, it was, I think, a $14 billion—more than $10 billion—company. It did not meet that valuation’s expectations, but there are many things out of your control in life.
Then Alex started this company called OpenRouter. What initially attracted me to it was two things. First, it was very clear from my time at Glean that this is a perfect problem where engineers all think it’s easy until it becomes incredibly annoying to keep maintaining. That’s the sweet spot, because no other person or company will gravitate toward it, yet it is a very thorny problem to maintain a portal that accesses a bunch of models. The nuances are quite tricky, annoying, and boring.
The second thing I liked is that I was pretty convinced that if there was a market for anything like this, it would have to be a PLG motion. I would go so far as to say that in any SaaS market, if there can be a PLG motion, the PLG motion will win. What I mean by that, if people aren’t familiar with venture words like PLG, is that all users have to be able to access and self-serve the product and try it without talking to anyone, rather than having to get on the phone through a classic SaaS website.
Those two things really drew me to the business. The third thing is just quality. There are these small details that make OpenRouter a beautiful website with a beautiful landing page. It’s not some SaaS trash of “Here’s what we do, product, solutions, about us.” I am so sick of that. You land on the page and it’s a developer page: “Here’s how many people are using it, and here are the models.” I love it. I’m like, “This guy knows what his users really want.”
All of those things were compelling. I went out to New York to talk to Alex, and he ignored me a bunch of times. I wrote him what I call love letters. I was like, “Hey, man. Love it, dude. It’s so cool. I don’t even want to invest. Just talk to me. I don’t really care. I just want to meet you. I have so many ideas and interesting things.” It was one of those companies where I genuinely felt that way.
When I did meet him, we started jamming on things. I don’t know the VC motions of how to sell, so I wasn’t really even trying to do that. But I told him, “Look, if you are ever going to raise, I will make it happen. I just love everything about this.” That’s how we ended up doing the round.
I think the company is interesting from a business-model perspective. I get this question a lot: How does this business model scale? I think right now the business is doing fairly well.
OpenRouter takes about 5% of everything.
There’s that business model, but then there is a reasonable threat vector: What if the spend on the network goes down over time as tokens go up? You do carry some risk that the prices of LLMs fall to a point where the business stops working, and I know many other companies take that risk as well.
That’s one risk of the business, based on pure consumer spend. The second risk would be keeping people on the platform. A lot of hobbyists use OpenRouter, and they tend to churn. A lot of enterprises will use OpenRouter to evaluate models and then pick one they want to settle with later. That’s a problem to fix. Those are the two risks, but overall, I think they’ve just been executing phenomenally.
Yeah. How do you think about the Vercel AI Gateway, for example? I think that’s been—I mean, I’m a fan of OpenRouter. I also use Vercel. When you already have Next.js, it’s like, “Well, I just use the AI SDK.” The AI SDK comes with AI Gateway, so it kind of makes sense to do it.
How do you think about this market, and how tied do you need to be to the actual application development versus just being Switzerland? OpenRouter doesn’t have a developer framework, for example. If we were in a partners meeting, that’s maybe what I would ask.
My simple answer is that I don’t think the AI gateways of other products are ever going to be their first priority. The other simple answer is that I think OpenRouter has this mindshare and momentum that just doesn’t go away overnight.
It would be similar to asking, “Hey, I’m OpenAI in 2020. What if somebody else does this?” Yeah, they could. Or in 2022, they could. But we are already so far ahead in some ways.
The last thing is that I think they have built a lot of smaller things that are non-obviously useful, which other people probably won’t sweat the details to build. It’s everything from a feature flag where you can choose to route only to certain LLMs that do not retain your data. They go to that level of granularity in thinking about what users actually want.
Another example is their level of detail about providers. Almost nobody has provider insights. There was a very interesting side study involving Kimi K2.
The providers.
The providers, okay. But I think that’s interesting. People don’t really acknowledge this, but the same open-source model can be served by different providers and have different context windows, different quality, different latency, and different throughput. Where would you go to see all that information? You see it on OpenRouter.
There are elements of scale, where enough people are using the different providers that you get that data. All of those things are somewhat defensible on OpenRouter, and hopefully more over time.
Yeah. I think their leaderboard charts are one of the best growth hacks because—
Very good graphics.
Especially people who are into open-source AI are always posting these things, saying, “Hey, open source is up. We’re back.”
One thing I used to joke about is that OpenRouter is the only non-Elon company that Elon has tweeted about the most, for obvious reasons.
Grok Code Fast is number 1 right now. I’m sure that’s because it’s free.
It was a good week where every day it was “OpenRouter, OpenRouter.” I was like, “Yeah.”
Yeah. And for those who don’t know, Grok Code Fast is a top model.
Yeah, because it’s free. There’s a lot of gaming of this stuff, where it’s like, “Oh, we’ll give it to you for free, but then we’ll say we’re very popular.” I’m like, “Yeah, you’re free because you’re popular, right?”
Yeah, the other way around. Okay, very cool. There are a bunch of others, and we’re not going to go through all 40. What comes to mind? What do you want to talk about? What do you think is a very interesting company in your portfolio that more people should know about?
I’ll talk about Wispr and Inception. Those are the 2 I want to talk about.
Inception isn’t even here.
We can talk about the company without saying the name.
Yeah. Okay, let’s try that.
Let’s talk about these 2 things. Wispr is a company that does, in many people’s eyes, something very commodity-like: voice dictation on your phone and laptop.
The things that really stood out to us about Wispr were that, in that quote-unquote commodity market, they are, in my mind, the fastest, best, and most delightful product. In many ways, they’ve set the frontier for the nuances of how to make this easy. Press your function key on your Mac and talk to it. It’s always on, and it has fantastic accuracy as you’re dictating.
If you ever stutter and go, “Oh, no. I didn’t mean that. I actually meant this,” it knows what you meant and corrects it. I find that they have this metric they use called zero-edit rate internally.
The number of times you don’t need to edit.
Correct. Their zero-edit rate, I think, is north of 80%, which is insane for a voice-dictation product.
There are many other risks to that business, too, but one thing I love is that users love it, users stay on it, and retention is great. It might make voice suddenly work, because if you think about computing, people type slower than they talk. It could be unlocking this new, faster way for people to feel comfortable talking to their computers. That really didn’t happen in voice dictation before.
It’s not just a Whisper model, which is a common question I get.
Yeah. For people who don’t know, it’s Wispr.
I mean, the question here is always: It’s the same thing, right? Voice is very commoditized. I actually happen to use Superwhisper, mostly influenced by Jeremy, actually. Granola is very popular, and Notion has this Notion Speech thing. What’s the plan?
This is why I’m not an investor: How do you survive? Basically, you’re trying to reason about why you should be the winner. Even ChatGPT desktop has shortcuts for stuff. I don’t know whether it does exactly the same thing, but it’s not that far away.
Anyway, you're excited about it. I do see a lot of tweets about Wispr, and it's one of those things where the FOMO is getting me, man. I like it. I'm like, “Should I switch?” I don't know. My thing's fine, but what if it feels better on the other side? I don't know.
Well, we'll see. We'll see how that pans out. There are some interesting plans to get it to be a cooler product, but we'll see.
Okay, we'll call this Stealth Co.
Stealth Co. One thing I find very interesting about Stealth Co. is that it's in the purview of research. We talk about different architectures all the time. One of the most compelling alternate architectures for AI is diffusion models.
One thing that I think is really interesting about it is that you talk a lot, Sean, about the Pareto frontier of latency, cost, and quality. Diffusion models today are, I would say, 80% to 90% of the quality at one-tenth the cost and latency. This has huge implications for, obviously, the stock market, which is kind of Nvidia, and many other things.
But there are clear examples of use cases where that might be very valuable, because there are many applications that work in volume that do not require high quality but definitely require better latency, and everyone could use some cheaper models. So I think there’s an interesting area of research there. Maybe it gets to frontier; maybe it doesn’t.
The one thing I want to draw attention to with diffusion, which I think is particularly interesting, is that left-to-right reasoning for code doesn't actually really make sense. In code, we might sometimes write code left to right, but after you write code, you go up and down and figure out, “Hey, is this variable set? Did I do this?” There are many bidirectional dependencies in code, so it has a natural tendency to lend itself to diffusion models.
You can imagine that as you're denoising, you fix partial issues in different parts of the code at once, versus this reasoning paradigm where you kind of have to figure everything out and then give your final answer.
Yeah. Yeah. I like that a lot, especially for syntax structures, like C-like languages, where you need to open and close a bracket and hold that state. I think the question is always the, quote-unquote, hardware lottery of Transformers. “Attention Is All You Need,” and diffusion is kind of a different branch off of that tree of research.
They're related, but we might be too far gone down the Transformers tech tree to come back and then go down diffusion, to the point where they might never be frontier because we've just had 4 more years of extra Transformers LLM research.
Yeah, it's true. I think about this all the time, thinking about, in the course of history, what are the significant moments where, if only something forked off a different way, maybe there would be a completely different paradigm or outcome.
Yeah. And usually the worst tech wins—Blu-ray, DVD, HD DVD, or something like that. I think there are a lot of variations of this. Even, I think, there was a discussion about AC versus DC currents back in Edison's days. There was this big fight between Tesla and Edison. I don't know if you—
I mean, I'm aware of the very, very basic details, but it's so interesting, right? You take something like this, and then the question becomes, “Okay, do we bet on it, or is the timing just off because something took off and we can't pull this rocket ship back to Earth, and so we've lost that fight?”
I don't know. I'm not a purist scientist anymore where I believe the best ideas and things win. I think in markets, it's very obvious that that's not true. A lot of things go into winning, and sometimes it's out of your control.
Yeah. Yeah. It's very true. Speaking of Anthropic and things that happened this year, MCP happened this year. When MCP came out, I was sleeping, and then when they came and did the workshop with me, I think I saw a lot more noise and was like, “Okay, there's something to this.”
Now it's basically the de facto interop layer for all the labs and all the models. There's no reason why this could have won versus anything else, apart from the fact that it was well specced out and backed by Anthropic. It's kind of a similar thing. But it was good enough.
Yeah, it happens. It happens so often. It makes it tricky not just in investing, but in general, to think about ideas. We see this with startups as well. It's very heartbreaking.
Every once in a while, you'll meet a founder where I'm like, “Your idea is fantastic. Your execution is great.”
I just don't see—
—it working because the market dynamics are not in your favor. Maybe I'm wrong about some of them.
When you say “market dynamics,” is it TAM or something else?
No, sometimes I just don't see that—
—you are a small group of people trying to wedge something into a market. We know how long that takes, and we know the other forces at play. I just don't see it. Imagine a single person running in a tunnel with a light at the end, with the tunnels closing in on you. You could be the fastest runner in the world, and you might not make it out of the tunnel. That's kind of the analogy.
And so you might be doing everything right. It's just that the window is not there, or at least I might not think that window is there. I do think a lot of companies fall into this bucket of ideas.
To me, in a way, I almost think of companies like MosaicML, which is like, “Hey, we have this amazing team. We can help you fine-tune models.” But nobody—you know, the market dynamic just—there's really nobody fine-tuning models.
Part of it is that the open models are not that good, and part of it is that people don't really have good data or the expertise. And again, if you go back now, there's RL environments and RFTs, like the next wave of that. Maybe they'll be able to get in the window, but it's just interesting how—
But then the other flip side of that is—and yet they get acquired for this amazing—
But yeah, because the market is just so big. I mean, even if you think about something like diffusion models for text, if you sell it for $1 billion, it's like 0.01% of Nvidia's market cap. So the amount of money being spent in the space is large enough to justify betting, yeah—
Like the same way Instagram was 1% of Facebook's market cap. It's similar, where it's like, man—
If Databricks is rich enough—
Exactly. It's like, you know—
They really want you to know that they're an AI company.
Exactly. And now they're worth $100 billion. I mean, without Mosaic—exactly, it's like, without Mosaic, maybe they're not on the same trajectory. I don't know. Maybe they are, because Ali Ghodsi is great and all.
Have you guys ever talked about the rollup companies, which is my favorite little—
The PE rollups.
Yeah. Yeah. Well—
I didn't know that was a topic of yours.
It's not really a topic of mine. I just find it quite interesting to see how, speaking of AI companies and markups, there are companies—obviously, I'm not going to name them—that go, “Hey, here's a small company that does $1 million in ARR completely with humans. I'll buy it for $2 million, and then I'll do some of it with AI. But now I'm an AI company, and $1 million of ARR in the AI company world is a $100 million valuation.”
So, cynically, it's pure multiple arbitrage on the category that you're in.
But yes, that's the cynical “ha-ha,” but then what if it actually works?
Because the hard part is getting the customers. The hard part is getting the domain expertise. You drop a bunch of software engineers in there and automate it, make it scalable, make it cheaper, and yeah, maybe it works.
No, you're right. You're right.
I think I just funded a company that bought a tax firm. So—
Yeah, a law firm, accounting firm—
A law firm. Law firm. Yeah.
If it works, it works. I just think what's interesting to me is that you can 50x the value of the company before you actually land anything with AI yet.
Yes, but then you use that funding and the equity to hire the people, and it's weird. So there's this concept I always talk about, which I'm surprised people don't really understand. It's reflexivity: the belief that something can be true can make it true, even though it's not true at the time that you believed it.
Yeah. That's venture capital.
Yeah.
Just give money and everybody's like, “Oh, they raised $300 million. It's a great company. I love that company.” It's like, “Yeah, I'm an investor in it, so I love it too.” And all the employees are like, “I love this company. My stock is worth a lot of money.”
There's also that effect that's very clearly in venture capital, where not just what you said—which I agree also happens—but imagine there are times where people funnel so much money into a company before it's really prime time that it dissuades anybody else from entering that market. Then they become the de facto winner of the market because they cancel the competition with funding.
And you can think—and I'm not going to name the categories—but you can think of innumerable categories in this market, in this paradigm, where that's already happened.
Yeah. And I feel like even in AI, maybe 2.5 years ago, when ChatGPT came out, this was cool, but a lot of enterprises were skeptical: Is this trend going to continue? But then, once you start seeing tens of billions of dollars being put into OpenAI and Anthropic, it's got to work.
Especially, you can deploy it in hardware.
Which, I think, at that point, you're building infrastructure. And infrastructure is very capital-intensive, and you can actually do the math: it's not humans anymore; it's machines and land.
Yeah. Exactly.
Power. Amazon is building all these training chips and all this infrastructure for Anthropic. Do you really think they're dumb? I think at some point, it's the same with Stargate. Do you think all these people are dumb, and you're saying the models are not that good?
Right?
Is it, by the way, an Anthropic-relevant thing?
Right.
But is that necessarily true? There could also be a world where that's just not true.
This is what makes it bitter: What if it doesn't apply to me this time? Right. Right. And I think, being in Sam Altman's place, that's absolutely the right chess move to play. But I do wonder what happens if all this investment in compute doesn't actually lead to economic gain, better models, or everything else.
But I feel like we've reached a point where the models are good enough that, even if the next generation isn't 10× better, we'll be able to use the compute.
That's the cope. They're writing it down for 30 years. So can you run GPT-5 Pro over the next 10 to 15 years, given the amount they're spending on compute?
And this is a general question—I'm not criticizing at all—but even if everyone were using Claude, Codex, or Claude Code all the time, inference demand isn't that big globally.
Right.
So you would have to believe—
What would you have to believe for that to be true? There are 800 million weekly active users. This is what Greg Brockman says: a GPU for every human. I'm somewhat shitposting, but they actually say this in their official communications, so I'm just repeating him.
I don't necessarily disagree. I'm just trying to work backward to what we need to believe to get there, because ChatGPT compute is not that much.
Correct?
Right. So they're not doing agentic stuff. Maybe they will in the future. Most people are doing basic Q&A-type queries.
By the way, I put it up on the chat, so if people watching on YouTube, they can see this. This year, OpenAI spent $7 billion on compute; only $2 billion of that was for all of their inference.
Right?
The remaining $5 billion was R&D.
So all of ChatGPT—all 800 million users—all of Sora, and all of the API volume: $2 billion. And they have 2.5 times that for R&D.
Right?
And so my point is, if inference is one thing, I don't know how that will scale to that volume, but then you'd have to believe that the rest of it goes into R&D and therefore produces models that are so much better that they therefore have more demand, et cetera. But in any case, if the incremental marginal benefit isn't that big, then that's the risk.
Yeah. So, to disrupt OpenAI, you need to have more efficient research, because right now it's pretty inefficient: you spend $5 to get $2.
So what OpenAI did to Google is what the next OpenAI has to do to OpenAI.
You know what I mean? Google was spending a lot of money, Facebook was spending a lot of money, and they didn't come up with anything. OpenAI did, and it was a small, tiny startup. They had GPTs, and I like Alec Radford, but someone else may or may not come up with that. It's that classic quote: “Your margin is my opportunity.” Google was milking those margins, and they didn't want to spend the compute for every search query, and so—
Now OpenAI is willing to—
So we've covered a lot of topics. Thanks for indulging me. For me, this is a survey episode: here's everything. We're also catching up with a former guest, which is always nice. Maybe we can end on this coding-interview thing, which you literally tweeted about today. What is the situation that engineers should be aware of? I think this maybe ties into LLM psychosis a little bit.
10. Coding Agents Need Human Judgment
So I tweeted about this guy who wrote a blog post. He was interviewing with a company. I didn't think it was a legitimate account; he thought it was a legitimate LinkedIn message where he was interviewing for the company. They sent him a coding interview: clone this repo, run this code, make this edit. Not untraditional—pretty run-of-the-mill. It happens.
In that interview, he claims he went to Cursor and asked whether the code had any vulnerabilities or anything he should be aware of. It revealed that it had a link—a byte array that compiled into a link that would go and take a bunch of private information from you. That was the TL;DR.
I tweeted about that, saying the interesting thing is that it was solved by vibe coding. But the world of vibe coders who don't really look at code could very easily be more susceptible to attacks like this in the future. It got me thinking about what attack vectors even look like if people aren't looking at code, what can go wrong, and what the implications are for model safety and how models behave in those environments.
So, that’s one, but I think the broader thing—and I’m curious what you guys think about this—is what I’ve been noticing more and more. I was having this conversation yesterday with some of my close friends where some of the joy of coding used to really be that you’re stuck on this annoyingly hard problem and you just bang your head against a wall and want to kill yourself, and then eventually you’re like, “I’ve figured it out,” and you solve it. That’s the muscle that you build when you improve and get better.
And now I find myself even doing this. It’s so hard to do if you just have a constant slot machine that might give you the right answer. Who knows if it will, who knows if it doesn’t, but you just pull it all day long: “Please fix, please fix, please fix.” What does that mean for the craft of engineering or software engineering in the future? I don’t know. This vibe-coding stuff—great for the rest of the world that was not an engineer, but I’m now seeing how it’s affecting trained software engineers, and it’s kind of like a drug for them. It stops them from living their own life.
Which is doing the engineering.
It turns your brain off.
Because it turns your brain off.
Yeah. I think people thought about this first with self-driving cars. This is why, when you drive your Tesla, you have to keep your eyes on the road: they don't want you to turn your brain off. We don't have that equivalent in developer environments yet. Maybe we should watch your eyes.
We removed one word in the code. Which one was it? Write it back.
So my answer—I happen to have shipped a model today, or two models. Part of that is what I've been calling the semi-async value death. A lot of it, I think, is my reflection on coding agents: we started with Copilot, which is tab autocomplete, and then went all the way to Claude Code, which is very async. It could take 30 minutes, it could take 30 hours—I don't know; it just runs.
I think something that Cognition is very interested in is fast agents, or something I've been writing about more. Fast agents are where, under a certain level, you actually want to be in a mind-meld with the human and AI, to have fast responses so that you can get helpful assistance. If it helps, you can use it; if it doesn't help, it gets out of the way.
That is actually where you do your hardest problems. When you're doing deep work and focused on a hard problem, you should be applying your human intelligence, augmented by AI, in an unobtrusive fashion. I think that's a pro-human message, but it's also a really interesting area of research for us.
But to play devil's advocate, that's almost like telling somebody, “I'm going to put the cigarettes right here. I know you love smoking, but please don't do it.”
It's not a cigarette.
It's right here.
It kind of is. There's an analogy to be made here: it's a cigarette for your brain, because you don't think anymore when you pull that button. And over time, I feel like the brain will get weaker if you don't use it for that task.
I like your message. Ideally, if I had a team of engineers, I would tell them the same thing. But I worry about the reality, which is that that's not what they do.
In many cases.
But I mean, you've got to ship the thing, right? I agree, but at some point you've got to close the ticket and merge a PR.
Mhm.
So how are you going to get that code done, right? It's like they're doing it, or they're going to get fired if they're just generating one way or the other.
Yeah, it's interesting. Okay, maybe I'll put it this way and see how you respond. We have the formula—the fundamental formula—for coding-agent performance. It basically is: find the right files and then write the right files. That's it. Read and write. Read the right files and write the right files. That's it, right?
So actually, what fast agents can do—or what I just did today—was basically the equivalent of a heads-up display. It gives you more information, but you still take all the actions. We help you read faster, read more efficiently, and read with more focus, but you still write.
And so I think that's still not a cigarette so much as we try to be helpful, and we're evaluated on the helpfulness of the reading and the comprehension, so that you can hold everything in your head.
That would be the pitch.
It's true. If I don't know what the product looks like, I would love to eventually play with it—with SWE-GPT and all of that stuff. But there's a world where I think the product decision also goes a long way toward how people use it. So if it is like that, then maybe. And I think when people use, for example, Cursor, a lot of people like the fact that they can see the code and then kind of have to hit the final accept.
Yeah, human in the loop.
Human in the loop. But I still worry. I worry the most about the younger kids, right? You think about the people growing up in college—
How would you ever get yourself to think if you just had this clearly more intelligent thing than you? At least, I don't want to rate myself too highly, but if I'm working in a domain that I understand, I can at least tell the model, "You're doing the wrong stuff. Definitely don't do that. Don't write that at all. That's a terrible file. Why are you creating 4 files for this?" But if you think about what it looks like to an 18-year-old freshman CS major, they're probably just like, "I guess that's how you do things." They can't hold it at that level. So their training is just a little bit different.
Cool.
Yeah.
Hi, D. Thanks for indulging us, and welcome back. Thanks for coming back.
Thank you, guys. Always fun chatting with you guys.