20VC:Anthropic 的编码业务值 2 万亿美元吗?| 美国企业是否应与中国开源模型合作?| 为什么未来 18 个月内 80–90% 的 Neo-Labs 会倒下?——Factory 联合创始人 Eno Reyes
- Eno Reyes 的核心定价逻辑是,AI 应按结果而不是 token 定价;按这个算法,“最聪明的模型实际上最便宜”。 一个用 1,000 个 token 就能完成代码审查的前沿模型,胜过一个烧掉 5000万个 token 才得到结果的廉价模型,这打破了“一个 token 就是一个 token”的框架,也指向模型的快速“物种分化”:一边是所有任务都能用的开放商品模型,另一边是企业后训练、完全留在内部的专用模型。
- 他认为,“前沿模型的 TAM 现在坦率地说被高估了”,而 Harry 将 Anthropic 的 2万亿美元估值理解为:在“最具竞争的应用市场之一、最挑剔的细分领域之一”——开发者工具——给 Claude Code 定价 2万亿美元。 2–4万亿美元估值背后,隐含的前提是实验室可以“把这些 token 的价格提高 2 倍,用户还会买单”;可能的突围路径只有监管俘获,或找到能匹配成本—质量前沿的应用和结果,而这很可能意味着接入更多模型,OpenAI 正在悄悄这么做。
- “未来 18 个月内,可能有 80%到90% 的新型实验室会倒下”——但这里的“倒下”往往意味着被收购、实现不错的退出,而不是走向毁灭。 他用 3 个问题检验业务能否长久:工作流是否持久,能否经受更强的前沿模型,能否经受全新的工作方式?法律符合全部 3 条;“通用电脑操作”以及 Excel/Jira 周边的知识工作则过不了关。
- “把开源模型叫作中国模型,是前沿实验室发起的一场心理战,基本是在诱导人们认为它们很可怕,并把它们他者化。” 他的预测是:“3 年后,99% 的工作流都会由开放模型完成。但其中 1% 的任务,可能占到未来智能经济价值的 30%、40%。”前沿应用——生物研究、国防和高级 AI 开发——“极其小众”,因此对 Global 2000 而言,成本才是决定性因素。
- 路由技术已经商品化;Stripe 以 800亿美元收购 OpenRouter,押注的是资本配置信息,而不是技术——“技术已经不再是护城河”。 真正的杠杆在 harness,也就是“某种意义上新的应用”:上下文窗口已经通过 agent 内部的压缩机制解决,闭环持续学习“尚未被开发出来,它不存在”,而学习会沉淀在 harness 层。
- 未来 5 年最关键的问题是:“谁是你智能的主权者?是你,还是另一家公司?” 两家最大的模型公司“已经明确表示,‘我们要进入每一个行业’”,这会推升本地部署需求(Factory Private),也让 Cursor 与 SpaceX 的合作变成潜在负担——当“他们会想推 Grok”时,很难继续保持模型独立。
- “娶微软、睡 NVIDIA、杀 Meta”;如果 SpaceX 估值能达到 2–3万亿美元,NVIDIA 3 年后达到 10万亿美元“很可能成立”。 微软凭借模型独立性和基础设施,成为“最有优势的超大规模云厂商之一”;债务周期对超大规模云厂商可控,但对 OpenAI/Anthropic 则“完全关乎生死”,它们“必须成为科技史上自由现金流最强的企业,才能活下去”。
- 在人才方面,“我们预计未来 100% 的招聘都会通过收购公司完成”;常春藤学历“几乎不能说明能力”,而表演式的 9-9-6 文化是危险信号,“几乎总是与某种其他短板相伴”。 Token 消耗应按项目而非人员分配——Factory 曾在单日内为一个基准测试(Program Bench)投入“接近七位数美元的额度”,并认为一些企业的这类预算轻松就会达到“八位数、九位数”。
1. 按结果而不是 token 定价——最聪明的模型可能最便宜
- Eno 开场提出的框架是:“一次代码审查要花多少钱,远比这次代码审查里的 token 要花多少钱更值得讨论。”一个复杂模型用约 1,000 个 token 就能拿到正确结果,胜过一个廉价模型消耗 5000万 token 反复摸索——“对于许多最具挑战的任务,我看到的世界是,最聪明的模型实际上最便宜。”
- Harry 追问:这难道不是与“一个 token 就是一个 token”相矛盾吗?Eno 的答案是模型正在发生物种分化:开放模型主导商品化任务;而对于少数高频专用任务,如果商品模型不够好、前沿模型又太贵,企业就会后训练一个中间层模型,并“完全留在内部”。不会有数百万个模型,“但数量肯定会相当多”。
2. 后训练将像软件开发一样民主化
- Harry 的反对意见是,如今的公司组织根本没有为后训练和落地实施做好准备。Eno 的反驳是,目前这套方法“主要存在于少数专业人士的头脑中”——这与 20 年前软件开发所处的阶段完全一样。很快,企业只需“打开一个平台,点几下按钮,描述你关心的任务,把平台指向那些工作流……然后一个模型就出来了”。
- 这也是对模型实验室的隐性回击:“有几家公司声称,递归自我改进和模型训练只能由它们掌握;但我认为现实会是,许多企业都能通过其他公司销售的软件服务获得这项技术。”
3. AI 的前沿,是在不存在验证机制的地方建立验证机制
- 可验证性是“当前 AI 系统能否成功的最重要属性”。在医疗、法律等模糊领域,评测设计者会请专家构建全新的判断方式;下一步则是:“AI 当前的前沿,是能在不存在验证机制的地方建立验证机制的 AI 系统”,从而进入人类认为 AI 太难完成的任务。
- 他的管理类比是:律所里的新手经理靠直觉判断招聘结果;但公司扩大后,必须把“什么是好”写下来——这又会改变行为。“如果你说好就是 A、B 和 C,那么你会得到大量的 A、B 和 C”,无论激励机制是否正确。
- Harry 接着说:“告诉我激励机制,我就能告诉你结果。”如果给 VC 设定的目标是每年完成 3 笔交易,得到的就是 3 笔交易,而不是一笔真正优秀的交易。
4. Harry 对 Mercor 的遗憾,以及数量级判断错误
- Harry 坦言,他在 Mercor 估值约 20–30亿美元时进行了投资,却跳过了约 200亿美元那一轮,因为当时想的是:“它还能大多少?也许能到 1000亿美元,但那只是 5 倍。”如今他认为自己“完全错了”,并认为随着未来对数据的需求增长,Mercor 有机会达到 2000–3000亿美元。
- Eno 同意市场误读了规模:那些认为 200亿、300亿、500亿美元估值荒谬的人,“低估了这场转型的巨大程度,低估了一个数量级”。但赢家不会长得像拥有护城河的企业——它们更像是“一群比其他人更清醒、更早看懂未来的人”。
5. 前沿模型的 TAM 被高估——Harry 将 Anthropic 的 2万亿美元视为押注 Claude Code
- “前沿模型的 TAM 现在坦率地说被高估了。”整个世界都假设会有 1–3 家公司主导智能时代;真正的问题是,在有如此多选择的情况下如何守住利润率,因为 2–4万亿美元估值隐含的前提是,你可以“把这些 token 的价格提高 2 倍,用户还会买单”。
- Harry 反驳称,Anthropic 刚刚实现首个盈利季度,利润率也在快速提升。Eno 的回答是,真正的利润驱动来自应用——模型利润率“肯定不如应用”。实验室的策略正在分叉:要么主导推理平台,要么向上游应用层延伸;“Anthropic 似乎在走应用路径,而 OpenAI 似乎在两边都试水。”
- 模型锁定会造成“糟糕的激励错配”:被锁定的供应商会给你自己的最佳模型,而不是全市场最佳模型。Harry 继续追问:这难道不是“给 Claude Code 定价 2万亿美元”,而 Claude Code “并不难切换”吗?Eno 承认这正是风险所在:这是“对最具竞争的应用市场之一、最挑剔的细分领域之一——开发者工具——下注 2万亿美元”。
- 两条突围路径分别是监管俘获,或找到能匹配成本—质量前沿的应用和结果。Eno 认为,这意味着接入更多模型。OpenAI 正在处理这个问题:“他们没有正式宣布,但显然在支持开放模型生态。”
6. 当代资本主义最糟糕的营销工作——以及被修正的预测
- 对于 Dario 的信息表达,Eno 认为 AI 营销“做了与你希望相反的事情:吓到每一个人,告诉他们这项技术极不可靠;一边大规模部署,一边用它威胁人们的生活和生计。”这些威胁确实存在,但“你一旦开始谈奇点和 AGI,把 AI 塑造成某种神一般的神话,就会吓到很多人”。
- Eno 认可 Sam Altman 修正自己预测的做法:“我完全低估了经济的动能……所以我预测了一个实际上并未发生的未来。”
- Eno 也修正了自己的判断:“建立一家企业,比制定计划更依赖即时反应。”最好的决策往往是在几秒内做出的,未来则在群聊中“实时被定义”。好消息是,“你距离更大的结果,基本上只差几个决策”。
7. 利润率、补贴,以及 Factory 拒绝用价格淹没市场
- “并不是所有 AI 企业的利润率都差。我知道,我们的利润率就不错。”Factory 不会在开发者工具上补贴消费者,因为这会消耗大量心智份额;“有两家实际上拥有无限资金的玩家,正在试图淹没市场”,而“正如我们已经学到的那样,补贴一撤,用户并不会留下来”。
- 长期看,1–3 年内最具成本效益的方案会是本地运行的开放模型,因此 Factory 现在就围绕本地、开放和本地部署进行优化——“最终,自助服务会找上门来……因为我们会拥有市场上最好的产品”。
- Factory 仍然每天稳定获得数万名自助服务用户。Eno 认为,少于 25万名用户可能就足以形成推动产品改进所需的反馈闭环;但数百万用户带来的草根社区、媒体关注和叙事能力,则更难复制。
- 给投资人的建议是,靠补贴换锁定用户是经典策略,但“如果没有提高利润率的路径,那就非常危险”——如果只涨价、不提升结果,“用户会流失,转向其他产品”。还有一条应牢记:“如果你的产品需要 100 名 FDE 才能部署,那你就是一个糟糕的产品。”
8. 路由已经商品化;harness 是新的应用层
- Harry 提出疑问:OpenRouter 能以 800亿美元卖给 Stripe,但 Ramp、Merged.dev、Requesty 等公司都在做路由。Eno 认为,技术本身没有差异化,Stripe 买的是一场关于资本配置的押注:能量变成智能,智能变成美元;Stripe 已经控制资金流,而 OpenRouter 展示了“这些模型将走向哪里”。“如果那是一家没有用户、只是技术更好的同一家公司,你不会支付 80亿美元……技术已经不再是护城河。”不过他也认为:“80亿美元确实很贵……80亿、100亿美元成了新的 10亿美元。”
- 网关路由能节省 10%–20% 的成本,但 agent 工作流需要的是有状态、任务内的智能资源分配。上下文窗口就是一个例子:它不是在模型端点解决的,而是在 agent 内部通过一种叫作 compaction 的机制解决的。“人们总希望问题能在别处解决,比如模型里或网关里,但我们越来越看到,真正解决这些问题的是 harness”——“harness 实际上就是某种新的应用”。
- 至于持续学习,LLM 的闭环版本“尚未被开发出来,它不存在”。相反,即使模型供应商也承认,学习发生在 harness 层,而企业会坚持把这一层掌握在自己手里。
9. 智能主权、本地部署,以及 Cursor 的 SpaceX 问题
- 未来 5 年的问题是:“谁是你智能的主权者?是你,还是另一家公司?”他的例子是,一家律所把每个案件都外包给 10 家供应商,5 年后会发现,“那些其他公司完全可以转过身来反过来拿捏你”。这也不是偏执——“如今提供模型的两家最大公司已经明确表示,‘我们要进入每一个行业’。”Palantir 和 Satya 都公开强调,企业应拥有自己的智能。
- 本地部署的核心是控制权,而不是技术本身。Factory Private 是其最受欢迎的产品之一,但许多买家仍选择 SaaS,因为知道必要时可以切换供应商,会带来更强的安全感。
- 对 Cursor 与 SpaceX 的合作,Eno 认为这对团队而言是非常好的结果,但一边与模型实验室绑定,一边还要保持模型独立,将会非常难以自圆其说——“他们会想推 Grok”。此外,企业也会担心信任问题;把软件开发生命周期交给一家模型锁定供应商,企业会“重新审视一次”。
10. 80%–90% 的新型实验室将在 18 个月内倒下——而能力进展存在行业滞后
- 这比此前 20VC 嘉宾的判断更悲观:“我认为未来 18 个月内,可能有 80%到90% 的新型实验室会倒下。”不过,“倒下”往往意味着通过收购实现“惊人的结果”;“这些企业可能不适合作为独立公司存在”。他的持久性测试是:业务是否绑定于持久工作流,工作流能否经受更强的前沿模型,能否经受“全新的思维方式”?法律满足全部 3 条;“通用电脑操作”以及 Excel/Jira 中间层知识工作则无法通过——“5 到 10 年后,我们可能不会再使用很多这类工具。”
- Harry 举了剪辑这期播客的例子,Eno 认为行业滞后确实存在,而且可能持续多年:媒体企业尚未提炼出其中的隐性知识——知道哪里是吸引人的开头,“归根结底依赖品味”,是一种人们“甚至无法描述”的直觉;而“一旦企业开始将这种差距资本化,进展就会极其迅速”。
11. “中国模型”是一场心理战——99% 的工作流将转向开放模型
- “把开源模型叫作中国模型,是前沿实验室发起的一场心理战,基本是在诱导人们认为它们很可怕,并把它们他者化。”判断任何模型都应问同样 3 个问题:它审查什么,能否解决你的问题,如果它消失了你能否切换。中国前沿模型“没有展示任何安全后门的案例”,美国模型也一样;差别只是创建者的偏见。他举的例子是:让模型写一份 10-K,描述一项涉及递归自我改进的战略,“它会把你拦下来。答案就是 Anthropic。”但他保留了这一限定:美国国家安全相关工作“绝对不应该使用中国模型”。
- Eno 预计,模型创建在相当长一段时间内仍会持续,甚至可能加速:模型构建越来越容易,而智能主权会催生拥有不同观点和视角的模型。模型路由器还会成为信息流;每当新模型发布,路由器也会获得免费的广告。
- 他的核心预测是:“3 年后,99% 的工作流都会由开放模型完成。但其中 1% 的任务,可能占到未来智能经济价值的 30%、40%。”Sam 和 Dario 反复强调的前沿应用——前沿生物、LLM 开发、国防——“极其小众”;问问 Global 2000 的企业今天在做什么,“99% 的时候都不需要真正处于智能前沿”。因此成本会占上风,而那些发布开放模型、再通过推理服务赚钱的实验室,“实际上可能会成为一门很好的生意”。
12. 微软胜过 Meta,债务关乎生死,以及雅虎时代
- “坦率地说,考虑到这种独立性,微软可能是最有优势的超大规模云厂商之一。”Satya “打了一场大师级的比赛”(Kevin Scott 找到了与 OpenAI 的交易),既捕获了上行空间,如今又能把 Azure 卖成 Anthropic、OpenAI 和开放模型共同的家。Zuck 推动开放模型“对人类是正确的”,但 Meta 必须用这些模型驱动自身业务,才能证明这笔投入合理。二选一时答案是:“肯定选微软”——因为基础设施在手,“无论上面运行的是什么模型,微软都会赢”。
- 债务周期在没有自由现金流时会很危险。超大规模云厂商可以承受 AI 现金流预期被下修;但对 OpenAI/Anthropic 而言,这“完全关乎生死……它们必须成为科技史上自由现金流最强的企业,才能活下去”。芯片——包括 Anthropic 据报道正在推进的“Jalapeño”芯片——是有道理的:纵向整合是摆脱巨额债务和负担的最强路径,但明确目标是取代现有供应商,“这与所有人都是竞合关系”。
- 对于市场泡沫,Eno 不认为会出现类似 2008 年的崩盘——但“我们可能正处于雅虎时代”,OpenAI 和 Anthropic 更接近 Netscape 类比。Factory 更看重 Jobs 留下的教训:“关键不是第一,而是做到最好。”
- Airtable 估值约 25亿美元,低于 110亿美元,说明“当代 SaaS 企业更像电影制片厂”:一部爆款远远不够,“如果你躺在这上面不动,那么 Bending Spoons 就会来把你吃掉”。未来会出现一轮大规模并购,收购对象是那些不错、但还远未达到 Stripe 级别的企业。
13. 人才:收购创始人,终结表演式文化,按项目分配 token
- “我们预计未来 100% 的招聘都会通过收购公司完成。”面对 Harry 对“使命一致”的质疑,Eno 解释说:使命一致“不是一个可以被建议或宣称的属性,它会在工作中极其明显地体现出来”——比如那些已经在构建软件开发 harness、拥有数万名日活用户的人。谈到 Cognition 的国际象棋冠军和数学天才时,Eno 说:“我上过常春藤。我亲身体会到,那几乎不能说明能力。常春藤学校里有很多笨蛋。”招聘应看重一个人能否在系统定义的规则之外行动。
- 创始人招聘中最大的错误,是“表演式工作文化”。用 9-9-6 表明自己很拼,“几乎总是与某种其他短板相伴”。Harry 澄清了自己被认为奉行 9-9-6 的含义:比如周日接客户电话,而不是严格工作这些小时。Eno 表示认同,但坚持认为,“任何时候,只要你创造了一种激励,让人展示自己在工作,而不是实际把工作做好,你基本上就是在激励错误的事情。”
- 对于 Jason Lampkin 提出的“给最优秀工程师 10万美元 token 额度”,Eno 认为,把额度分给个人“是一种非常奇怪的思考方式”——应按项目和结果分配。Factory 曾在单日内通过 1 个人完成的工作,为 Program Bench 投入“接近七位数美元的额度”;对一些企业而言,这类支出“轻松就会接近八位数、九位数”。
- 谈到定价人才,Harry 举了 Poolside 员工转投 NVIDIA、市场传闻交易价值 120亿美元的例子。Eno 认为:“不是某一个节点值 1亿美元,而是把所有节点放在一起后,它们形成的图谱可能价值数百亿美元。”对于一家处于 AI 之前时代的既有企业,正确更新这张图谱,“可能决定它是一家 2万亿美元公司,还是一家 4万亿美元公司。所以几乎任何代价都值得”。
14. 快问快答:威胁名单、Salesforce 收购,以及 200万人的祭司阶层
- 在 Meta、微软、NVIDIA 的“娶、睡、杀”选择中,答案是“娶微软、睡 NVIDIA、杀 Meta”——NVIDIA 是“当下的造王者”;Meta 在技术上是“正确的”(VR、开放模型),但只有“一头现金奶牛”。至于 NVIDIA 3 年后达到 10万亿美元:“如果我们允许 SpaceX 值 2–3万亿美元,那 NVIDIA 可能就值 10万亿美元……答案很可能是肯定的。”他还看好 Salesforce:“当人们说‘我讨厌这款软件’,但所有人都在买它时,那大概是一门很好的生意。”不过,敏捷开发本身可能受到冲击,从而对 Linear 和 Atlassian 都构成下一代系统记录的威胁。
- 威胁排名第一是 Claude Code(“每一次谈话都会被提到”),第二是 Codex——越来越多地表现为“Codex for Work”。而关闭 Claude Code,“对我们来说其实是最好的消息,因为这说明它基本上没有用户黏性”。第三是 Cognition,基本上是另一家唯一保持模型独立的企业级供应商;第四是 Cursor,市场仍将其视为 IDE,而不是企业级战略。4 家公司的共同点是,几乎都在推销最终由达到人类水平的 AI 劳动力替代人的愿景——一场“夺走印第安纳·琼斯角色”的替换;Factory 推销的则是“全新的开发方法论”,即人类与 AI 协同构建。
- 对于 Chamath 所说“硅谷过于以金钱为中心”,Eno 回答:“这听起来确实像 Chamath 会说的话,因为坦率地说,他就是这么赚到钱的。”旧金山聚集着一批在困难问题上“吞玻璃”的人;我们讲述什么故事很重要,因为“那项技术就是建立在我们讲述的故事之上的”。至于 Chamath 的软件创业项目,能否成功“取决于软件到底有多真实”。
- 企业销售的建议是:“别再把它当作说服工作……把它当作一次发现客户需求的机会。”他对未来 5 年的收尾预测是,人们会觉得不可思议:“我们竟然让一个大约 200万人的祭司阶层,决定全人类所有软件的命运。”到那时,伯利兹的一名船夫也会拥有完全定制的软件,“看起来比你在家里的 HR IT 软件还好”。
完整逐字稿
I see a world where the smartest model is actually the cheapest. People are thinking about outcomes in AI, and they're looking at 20, 30, 50, and they're saying, “That's ludicrous. That's crazy.” That is underestimating by an order of magnitude how massive a transformation this is going to be.
The TAM of frontier models is frankly overweighted right now. $8 billion and $10 billion is the new $1 billion. Two of the largest companies that provide models today have explicitly said, “We're going to go after every single one of these industries and businesses that we provide intelligence for.”
I think it could be 80 to 90% of neo labs die in the next 18 months. Calling open-source models Chinese models is a psyop by the frontier labs to basically trick people into thinking that they're scary and otherize them. In 3 years, 99% of workflows are going to be done on open models.
I am fed up with speaking to visionary, insightful leaders who aren't actually building today in the trenches at the cutting edge of infrastructure and AI. Today, we have CTO and co-founder of Factory, Eno Reyes. He is one of the most articulate and insightful thinkers about the value stack of AI that I've interviewed. Factory is one of the leading companies specializing in autonomous software development.
I co-invested in their round with Sequoia, and Eno is incredible in the show today. There are going to be a lot of notes taken in this discussion.
Eno, dude, it is so good to have you in the studio. I obviously had Matan in the studio, and I was chatting to Keith Rabois over the weekend about you, so thank you so much for joining me, dude.
No, thank you for having me. I'm super pumped to be here.
Now, I hate background stories. I'm sure you listen to podcasts. It's like, “How did you get into this?” and you're just like, “I'm bored of this.” But downstairs I asked you how you got into technology and fell in love with it, and your story was heart-wrenching and compelling. So I have to ask it. How did you first fall in love with computers, do you think?
Yeah. I think the beginning was my parents. Both my parents went to art school. They were always in love with technology and how that intersected with creativity and media, and I think that's where it started, especially with my dad.
He was born in San Francisco in the late ’60s. When he was 6 years old, he actually got hit by a bus. That changed the trajectory of his life in a pretty crazy way. He wasn't playing sports, and he wasn't able to do as much of that traditional 1960s kid stuff. Instead, he went to technology and computers, like the early days when you barely had screens and were instead tinkering.
He spent most of his life embracing technology as a way to extend his own reach beyond what I'd argue his physical body could do.
It's so interesting how life delivers you a hand, so to speak, and how consequential that hand is to who you are today.
100%.
It's very hard to transition from a father being hit by a bus to margins.
Right.
Only a venture capitalist could make that transition so swiftly.
Well, some margins in AI might make you feel like you've been hit by a bus.
Yes, absolutely. We chatted before, and you said the cheapest model isn't necessarily the cheapest system.
Mm-hmm.
I read this when I was doing the work over the weekend, and I was like, “Huh. Can we just unpack that? The cheapest model isn't necessarily the cheapest system.” What does that mean?
1. Outcome Based AI Pricing
Yeah. It really comes down to this idea that when you're thinking about price, you should not be thinking about the inputs to the price; you should be thinking about the outputs. So I think about the price of the outcome.
Let's take software as an example. How much a code review costs is far more interesting than how much the tokens inside of that code review cost. If you take a very sophisticated model and it's able to do that code review immediately without making any mistakes, getting the right outcome right away, searching the right phrases, and you use 1,000 tokens or whatever, it will be significantly cheaper than using a cheap model that spends time running and uses 50 million tokens.
Ultimately, that price difference for the full outcome makes the higher-quality model cheaper. That's not how all tasks go, but for many of the most demanding tasks, I see a world where the smartest model is actually the cheapest.
I totally hear you there, and it kind of goes against the “a token is a token” theory. But will we then have millions of specialized models, with every company having specialized models operating on its own data? Because, to your point, it'll be able to work much more efficiently.
2. Specialized Models Democratize Intelligence
Yeah. I think there's a real world where the speciation of models increases very rapidly. This is the world where Fireworks and the people who help make models possible, I think, win, because the alternative is that you only have a very few specialized providers that have models.
In our view, there is probably going to be a difference between the commodity task executors. This is just your everything model, and we think that will be dominated by open models.
Then you have businesses that will say, “We do a lot of commodity tasks, but there are a couple of very high-volume specialized tasks that only we do.” For those, your commodity model won't be good enough, and your frontier model will be too expensive. So they'll want something in between, where they can take a commodity model and make it good enough via post-training.
Ultimately, they'll run that, and they probably will be the only consumer of it. They won't even give it to the rest of the world; they'll just keep it entirely internal. I think that will lead to a lot of models. Not millions, but it'll definitely be quite a lot.
When you look at the post-training required, and when you look at the implementation required, I look at that and think that company structures and teams today are simply not equipped to do that. How will we solve for that? Is this just moving into an incredibly services- and implementation-heavy world where we have these insane AI teams coming into every company? How do we solve that?
Yeah. I think this is one of those things where, right now, the recipe for post-training and building models lives primarily in the heads of a specialized group of people. But that's also how software development was 20 years ago.
So in my mind, the same tools that are currently democratizing access to software are actually the tools that will be used to democratize access to intelligence in general. I see a world where you're a company, an enterprise, and you say today, “Well, we don't have the knowledge or the skill set to build our own specialized models.” Well, very shortly—and already, to a certain extent—you can open up a platform, go to its webpage, click a couple of buttons, describe the task you care about, point it toward those workflows that happen in your business today, and out comes a model. That model will be really good.
I think that right now there are a couple of companies that claim that recursive self-improvement and model training will be only their domain, and I think in reality, many businesses will have access to that technology via software services that other companies sell.
You said focusing on the outcomes and not the inputs. That is a great idea when it's a very verifiable output.
Mm-hmm.
Code review.
Yep.
When there is ambiguity to something, which could be a marketing conclusion—did it come through X channel or Y channel? My girlfriend's a lawyer, and different legal notes are ambiguous. Some people like it one way, some people like it another way. How do you think about the importance of verifiability in determining outcome quality?
3. Verification Defines AI Success
Verifiability is ultimately the single most important property of success with current AI systems, and I think that the way that we'll progressively address this is by building new ways to verify the work that we do. I'll be really concrete about that. What a bunch of the eval creators and model trainers have done in some of these domains, like healthcare and legal, where you don't really have a concrete set of verification strategies, is take experts, bring them in, and have them create effectively their own new forms of verification. Maybe they show 2 examples side by side and say, “Which one, based on your judgment, is better?”
Being able to then take intelligent models that have a lot of the grounded reasoning and knowledge that these domains have, combine them, and say, “Now you, as the model, go and build this similar form of verification”—I actually would argue the frontier right now of AI is AI systems that can build verification where there is none. Thus, they can progress into tasks that today humans consider to be too difficult for AI to resolve.
AI systems that progress into verification. What does that actually mean?
To make it as concrete as possible, imagine that you walk into a room at a law firm and they say to you, “Hey, you're going to start basically judging how to determine whether or not these new hires are good.” What does a very novice manager do? They go by their gut. They're looking, and if they are actually intuitive, they'll go by their gut and make good calls. They'll say, “That new grad is going to be big at this firm. I like them. I'm going to continue to promote them.”
But when you have a giant firm, or you systemize, you realize that that approach to management is very rare. In reality, you need to go to the firm and start writing stuff down: “Here are the things that we like about great new hires. Here are the things we don't like.” Any company that starts to learn what managing at scale looks like has to write down the things that it cares about and build a framework or a system to analyze the job to be done.
I think that great AI systems are basically relearning this management strategy of saying, “You can go by your gut.” Honestly, a great AI system can be right very often without this sort of structure or framework. But in reality, what you'll need to do is write down, “This is what good looks like, this is what bad looks like, and here's how we judge.”
What's interesting is that, yes, this makes the system better at determining what good looks like, but it also changes the incentives. When you write down, “This is what good looks like,” people read that and then start to act more like what good looks like. You have to be careful, because if you say, “Good looks like A, B, and C,” you're going to get a lot of A, B, and C. But if you built the wrong incentives, then that system will end up following that pattern regardless of whether it's actually good or not.
That is a brilliant statement: “Show me the incentive, and I'll show you the outcome.”
Mm-hmm.
It's the hardest thing about actually running venture firms as a business, because if I set you the goal of 3 deals per year, you'll give me 3 deals per year.
Yep.
I don't want 3 deals per year. I want a great deal. I want a factory. I don't care if it's 3 or 6 or 1.
100%.
I think I made a big mistake because I'm in Mercor. I think we did it at, like, $2 billion or $3 billion. My memory should be better, but I'm older than you. I didn't do the latest round at whatever, $20 billion, because I thought, “How much bigger can it be?”
Maybe $100 billion, but that's a 5X. It's just not that exciting. A 5X, and that's with no dilution. I'm now thinking that I'm completely fucking wrong and that there is a pathway to $200 billion or $300 billion in the data requirements that will be needed. How do you think about what I just said?
4. AI Outcomes Create Massive Value
No, I think that's totally true. People are thinking about outcomes in AI, and they're looking at $20 billion, $30 billion, $50 billion, and they're saying, “That's ludicrous. That's crazy.” That is underestimating by an order of magnitude how massive a transformation this is going to be.
However, I think that what people underestimate is that the types of businesses that are going to become massive do not look like businesses 20, 30, or 40 years ago, where they had a technology moat or some sort of key capability that no one else could replicate. Instead, it's basically collections of people that understand what the future looks like a little bit more clearly and more clear-eyed than other people.
These data companies, like you mentioned Mercor, yes, they sell data, but every person at that company understands how AI is going to look much more clearly than the average human, and that makes them worth significantly more than even what investors will say.
Okay. Going back to what we said about lots of specialized models and companies working on their own data, which is obviously proprietary, I'm confused. How does this not reduce the TAM for frontier models?
I think it might. I think that the TAM of frontier models is frankly overweighted right now. The world basically assumes that there are going to be 1 to 3 companies that have total domination over the intelligence era.
I think that is a silly proposition, because generally people don't like that sort of strong dominance by a couple of small companies or large companies. But the real question is, how do you defend your margins if you're a model lab when there are so many options? I think that shrinking margin profile is going to change the expected value of these businesses.
Baked into $2 trillion, $3 trillion, and $4 trillion valuations is an assumption that you can basically 2X the price of those tokens and people will buy them.
Is the margin profile looking shit? What I mean by that is Anthropic just produced their first quarter, I believe, of profitability. Margins are seeming to be ripping, and they're throwing off cash now. Is that not going counter to what you said?
5. Applications Capture Better Margins
No. I think also that one of the biggest drivers of that is actually the applications on top of the models. One of the things that I think is quite clear to all of these model businesses is that the model itself may be a fairly rough trade-off. Rather, the margin profile of the models is definitely worse than the applications.
In my mind, if you are a model provider, you're basically looking at 2 options. A, you want to dominate the platform era, in which case you want to be 1 of N companies that get really good at selling inference. Or B, you just want to move up to become an application-layer company that has really good models.
Anthropic seems to be following the application path, while OpenAI seems to be dipping its toes in both, but the platform commitment from them seems much stronger.
Which strategy do you think is right if you were to bet on 1?
I think that the application layer is going to be a much harder battle, because being model-locked is actually a huge disadvantage if you're trying to sell outcomes.
Why is that?
It's bad incentive alignment. If you are a model-locked provider, then you are inherently selling those tokens in order to make sure that your business gets the margin it needs. If you are going to a company and saying, “We can give you the best outcome,” you have to do that with only your models. So they can really give you their best model.
Meanwhile, someone who's not model-locked can give you the best model, right? That difference between their best model and the best model can be massive in the pricing. We see this right now in real time in coding. Anthropic can basically only deliver its model's outcomes.
And the best model is also highly subjective, depending on the consumer.
Oh, yeah.
It's totally different based on the task, the profile, and the risk-taking that people want to have. There are tons of different options.
Can you help me understand? Again, I'm very thick, but I don't like cynical questions. I like to be optimistic. I think it's fantastic we're seeing Anthropic potentially go out at $2 trillion, but you're essentially placing a $2 trillion price on Claude Code.
Mm-hmm.
Which is not that difficult to switch off of.
Yeah.
What am I missing? What should I know that I'm not getting? How should I think about that?
I think that is fundamentally the risk for an investor: you're making a $2 trillion bet on one of the most competitive application markets in one of the most finicky segments of the market, which is dev tools.
I do think that part of what needs to happen in order to make companies like Anthropic and OpenAI realize their value is that they either have to pursue regulatory capture—which they are—or figure out a way to build applications and outcomes that match the true Pareto frontier of cost and quality. I think that means opening up to more models.
It's at odds with the 2 strategies. You either capture it and keep the model, or open up to everybody. This is a very hard decision, and one that I think you can start to see OpenAI actually grappling with as they've let more models into their harness. They're not making it official, but they're clearly supporting an open model ecosystem in a more direct way.
Do you think Dario's marketing message has been mistaken?
I think that the marketing of AI in general was probably one of the worst marketing jobs done by contemporary capitalists. It basically did the opposite of what you want: scare every single person, tell them it's very unreliable, and basically threaten their well-being and livelihood with the technology while you roll it out at scale.
I think that the challenge is that the things Dario brings up are not only well-intentioned, but there are very real threats from unregulated and dangerous AI. I think there's a way, though, to communicate about this without maybe embellishing the economic ends.
The carrot-and-stick here can just be: one, the technology will be incredibly transformative; two, humans will have a huge role in that transformation, and you will have a huge role in that transformation; and three, if we don't do this well, then, like all other technologies, there are going to be risks. But the moment you start talking about the singularity and AGI and create this godlike mythology out of AI, you're going to scare a lot of people.
We're going to be the last remaining private company. Say, "Well, you flew here on United." I don't know how that's going to work.
Yeah, there's going to be a lot of change if people think that there's going to be 1 company. I think Sam Altman just did an interview where—and kudos to him, because it's hard to go back and say, "I was wrong"—he basically says, "I totally underestimated the momentum of the economy, the momentum of existing businesses, and so I predicted this future that actually has not come true."
I think that's a great reckoning, where you go and say, "I didn't think that this was going to happen, and it's clearly not going that way, so I'm revising my prediction."
What did you not think was going to happen where you had to revise your prediction?
I love that question. I think the biggest thing that surprised me, where I've had to go back and seriously correct my priors, is that building a business is much more reactive than planning.
Almost every one of our best decisions as a company has been in reaction to some information and a split-second decision, rather than a master plan that we forecasted 6 months ahead. Knowing that, you start to look at the rest of the world, hear from other business leaders how that's also how they make decisions, and realize that the world is just this constantly reactive feedback loop where people are talking to each other and no one actually has an answer to what the future is going to look like.
You are basically defining it in real time in group chats and in the actions that you take as a business. Learning that in real time has made me, first, question the existing world and structure. Basically, nothing's guaranteed and anything could change.
Second, it's quite empowering. It makes you realize you're basically a couple of decisions away from even greater outcomes and an even bigger business than you had prior.
I totally get that anything could change. One thing that I hope changes is margins and margin profiles. Help me understand. When we look across the wave of incredible businesses in AI, the margin profiles are still lower than they were previously.
Yeah.
They're at 30 to 35%, say, as a barometer, compared to 70 to 80% with SaaS.
Right.
Is that a momentary period in time where we're in a build-out phase and they expand over time, or is that just a net-new model where the revenues are going to be much larger?
Not all businesses in AI have bad margin profiles. I know we've got good margins.
Okay, great.
That actually comes in a way that I think is also customer-aligned, in that we are so focused on thinking about these outcomes themselves as the thing to be valued. We've gone away from a lot of common paths that we see other AI companies taking.
We don't subsidize consumers in the dev tool space. That hurts us in a lot of ways. We don't have the mindshare from self-service users.
So you have no PLG?
I would separate self-service from PLG. We have a lot of focus on PLG within companies that we've deployed to, to allow the adoption to increase. What we don't have is a consumer-facing plan that is, I would say, super-rational for a current consumer unless you're optimizing for quality.
In a land-grab environment like today, should you?
I think it's a really good question. In my mind, the biggest reason for this is that there are 2 players with effectively infinite money who are trying to flood the market, and their intent is: if we flood the market, we keep you.
But remember what we talked about just a second ago: you don't keep them after you pull back the subsidies, as we've learned. I think the consumer will continue to follow the most cost-effective solution.
What does that look like in 1, 2, or 3 years? I think it's open models. The most cost-effective solution for a model is going to be the cheap one that you can run locally on your computer. We want to make sure that we are the product that best fits that type of experience.
That's why we're so optimized for local and open models today for on-prem, but in the future it's for all consumers. We won't have to subsidize as much as we'll have to create an amazing experience for self-service.
I think of this as a long game where eventually self-service will come to us, but not because we subsidize, because we have the best product on the market.
If you're advising me as an investor on how to think about margin and how that plays into my decision to invest or not in a business, what would you say?
For us, this is a part of our strategy. It doesn't mean, though, that it's the only way to win. I do think that there are probably going to be businesses where they're able to lock you in because of a workflow or a system of record that they produce, and the margins—or the subsidies temporarily reducing margin—can be a route toward gaining a customer base.
This is a classic strategy, right? This isn't even new to AI. What I would say, though, is that if there's no path to increasing the margin profile, that is very risky.
A lot of investments are being made in businesses where the promise is simply that they will raise prices, but you won't see a commensurate increase in the value of the platform. It is so competitive in AI right now. If you are not also raising the outcomes and the value that you get out of the product while you raise that price, people will churn and move to another thing.
I do think it's quite tricky, and the margin profile actually matters a lot, but it's not the be-all and end-all.
Will you have that churn in enterprise sales? You work with some of the biggest companies in the world. I'm sure you sign year-long minimums.
Yep, yep.
You have pretty sticky client bases there, no?
I think so. People also see these year- or multi-year partnerships as just that: a partnership. Part of what makes it interesting to build right now is that a lot of what you're selling is not only the technology, but your knowledge about how to best use that technology.
I would carefully differentiate that from consulting or professional services. You don't actually have to go in and do all of the implementation. I honestly think if you have a product that requires 100 FDEs to get it deployed, you just have a bad product.
But instead, I think that if you have the advice and the knowledge of the direction that you think the world should go in, you're selling that with the product. People are willing to go and buy that. And I think that if they see it from you today, they sort of know that in a year you'll also still have that same knowledge and forward-thinkingness.
Now, obviously, lots of people can give that, but if you combine that with a product that then acts a little bit more as a platform or a system rather than a tool that people use, you can also get stickier by just being something that you build on top of.
We spoke about the different frontier model providers essentially having this really challenging dynamic of being locked into their own models—
Yep.
—when serving, say, Claude Code or any application that they choose to serve. One then thinks that the value becomes in the routing of models: the tasks to the model, what's optimized for each use case, cost, latency, function, whatever it is. And OpenRouter gets bought for $8 billion. All the value's in the routing. Great. And then everyone is doing routing. Ramp has a routing provider. One of my companies, Merged.dev, has one. We're in another startup, Requesty. And I'm like, "Well, the routing's completely commoditized."
Right.
Help me understand: What world do we live in?
6. Routing Is Not The Moat
Yeah. Well, I think that routing is a really interesting technology in that I don't think the technology itself is necessarily that differentiated. And so if somebody comes up and says, "Look, Stripe bought OpenRouter for $8 billion because of the technology," then I'd say either, A, if they have insider knowledge, that was a bad decision, or, B, if they're just assigning that to it, I think they're missing what I read when I read the letter to shareholders, which was that they see this as a bet on where the direction of capital allocation is going.
Think about it like this: What are tokens other than intelligence? And how do you get tokens? Well, you pay for them. How does the infrastructure layer get it? Energy. It's literally translating energy into intelligence, and you're just trading dollars along the way. All of this is basically converging toward one thing. Whether you call it allocating energy, allocating intelligence, or allocating money, businesses need to allocate whatever this is in order to grow and expand how they operate and grow and expand their bottom line.
If you're Stripe, you already control the flow of 1 of the 3. With OpenRouter, rather than the routing technology, you actually just gain the information of where these models are going, right? You start to understand: What are people doing with intelligence? How are they allocating it? What models are they using? So now you start to control the second of these 3 things. Maybe they'll make a play into energy infrastructure or data centers at some point, but just ownership over those 2 is a massive bet on how to think about where companies allocate resources.
And so for them, I think this is very reasonable. But what's interesting is you wouldn't pay $8 billion for the same company that had no users with better technology. That difference, I think, is really important, and it gets toward that broader idea that the technology's just no longer the moat.
Do you think it was a good buy?
At $8 billion, it would have to be really foundational to the team that becomes whatever this next bet on allocating capital is for Stripe—or, rather, helping the businesses that Stripe has as customers allocate capital. I would say that if they think that data gives them insight into how to run Stripe better as well, that could also potentially make it worth it. $8 billion is quite steep, though, so stranger things have happened. It feels like $8 billion and $10 billion are the new $1 billion.
I'm intrigued. You obviously have a routing product within Factory.
Mm-hmm.
What do you see? What insight do you get from that that maybe the world doesn't see?
Well, one of the most interesting things—and I've been talking about model routing for these products that you've mentioned—is gateway routing. That's sort of how it's referred to because, ultimately, the routing effectively happens outside of where the task is being completed.
A lot of companies will do this. They'll look at Ramp, Stripe, and OpenRouter, and they'll put in a model gateway. That gateway is just how all of the different tools and products of the company route to LLMs. What's interesting is that we've seen you can definitely get some nice cost savings doing this, like 10% or 20% from these types of products.
But you really need something fundamentally different when you have agentic workflows because, to actually take the most advantage out of models, you need to do something that's a little different from just routing. You need your agent or your system to dynamically understand the task it's working on and understand how to allocate intelligence in a much more stateful way.
And when I say stateful, all I mean is that you need to know not only what's going on today, but also what just happened and what's going to happen in the future. That can't happen outside of where the task is being completed. You have to be in there in the task.
Is that related to the importance of context-window expansion that everyone talks about?
I'd say that, in a way, taking the context window and expanding it was solved not at the model layer, or outside at the endpoint, but instead inside of the agent with something called compaction. It's exactly similar. People really want the problems to be solved somewhere else, like in the model or in the gateway, but more and more, we see it's the harness that solves these problems.
The thing that's underestimated about why the harness continues to solve this is that the harness is effectively the new application. It's just where all the logic happens. It's where the state is maintained. It's the easiest and best place to do work with AI.
And so, as that starts to accumulate more advantage or technology benefits, I think we're starting to see people ask questions like, "Should I be building my own harness? Should I get a really great harness? What is a harness?" That's something that I think at Factory we're trying to spend as much time as possible educating people about: What does well in the harness versus what can be done outside, and how to think about building versus buying a harness.
The context-window expansion is 1, and then continuous learning and the rise of the first truly continuous-learning model. Does a continuous-learning model help or hurt Factory's business?
Well, it's interesting because there are sort of 2 directions for this. I would say that there was an idea of continuous learning from over the last couple of years that said you would have a literal LLM-like model where all of the learning happens internal to this closed-loop system. That technology has not been developed. It doesn't exist. There are attempts and early looks at it, but in general, anybody who wanted to try and hold all of the learnings behind an API would be able to potentially accumulate an advantage that would make it harder for others to use that model in their product because, ultimately, they would be accumulating all of the learning.
But in reality, what has happened is quite the opposite. Basically, model providers have even acknowledged that all of the continual learning happens at the harness layer. That learning is basically something that businesses are going to find very critical to own, that they are the sovereign of. And I think that that is ultimately the question of the next 5 years of AI: Who is the sovereign of your intelligence? Is it you, or is it some other company?
Sorry, can you unpack that for me? What does that actually mean? Who owns the data outcomes that are generated from the tasks that you do?
7. Businesses Must Own Intelligence
Basically, who owns the learnings and the workflow that successfully achieved the outcomes for your business? A great example of this would be if you were a law firm and all you did was outsource every single one of your cases to some other company—in fact, maybe it was 10 different companies. Then, 5 years later, those other companies can just turn around and screw you over because they know exactly how to do your entire business.
What's interesting is pretty much every company in the world is thinking to themselves right now, "What's going to happen in a couple of years if I outsource all of my intelligence to someone else?" And what's interesting is that at least 2 of the largest companies that provide models today have explicitly said, "We are going to go after every single one of these industries and businesses that we provide intelligence for."
Do you see with large enterprises—the biggest companies in the world—are they scared of OpenAI and Anthropic coming after their business?
Not all of them necessarily think they're going to come after the business, but the largest companies in the world are very wary of the model labs coming in and promising them intelligence and sort of luring them into a trap.
That's something that I know many businesses are starting to become aware of. Palantir has been quite loud about this idea of owning your intelligence. Microsoft as well. Satya wrote a great piece about this. I think that all of that opining is very spot-on.
At the end of the day, if it's not your intelligence, then there is a real risk that either, A, they come after your business, or, B, if they disagree with what your business is doing, they have a little bit more leverage and control than I think the typical business owner would like.
And when we talk about sovereign intelligence and owning that outcome, is that why on-premise is so important?
I think that's a huge part of it. To most businesses, on-premise isn't even about the technology. It's just about the idea that, if I need to, I can take control and ownership over every dimension of this software, and we get to stay in control.
That element is interesting because, for us, we offer an on-premise offering. It's, in fact, one of our most popular offerings, Factory Private. Even though we offer this, a lot of the businesses that we talk to actually go with our SaaS model because they have the peace of mind of knowing that, if they need to switch, not only do we have it available, but they understand exactly how it would work.
I think that part of this story is just being able to share with people that we are incentive-aligned. If you need this, we have it, and you won't lose anything. On-premise and owning your intelligence are very similar stories.
How much does it help you or hurt you that Cursor was bought by SpaceX? It gives them some amazing scale benefits in terms of access to compute, but it does make them model-biased.
Yeah. That outcome for the folks at Cursor is obviously amazing.
Sure.
Needless to say. I think where it helps us is that it's going to be a very hard story to become model-independent—or rather, stay model-independent—when you're attached to a model lab. They're going to want to push Grok. The products are going to become increasingly oriented around Grok, and that, I think, will become a challenge.
There are also, to be honest, trust and enterprise-related concerns that they're going to have to deal with under their new brand. Ultimately, the team there is obviously incredibly competent, so I don't discount them as a player in this market. However, I do think that most enterprises are going to have a second look at the idea of ceding their software development life cycle to a provider that is, one, likely to be model-locked and, two, has an existing history or pattern of maybe struggling to operate in these larger and more secure environments.
Can I ask you, when we look at the cadence of model development today, it's just so fast? I actually use arena.ai as a discovery mechanism for new models.
Right.
I'm suddenly using these weird models that I've never heard of and would never have used before—
Right.
—and I'm loving the output. My question to you is, will we see the cadence of model creation sustain in the way that we are today? In other words, will the rate of new models keep coming, or is this a momentary period at the start of a new cycle?
8. Model Creation Keeps Accelerating
It will likely sustain for quite a long time. This actually gets to another interesting property of model routers. A lot of people treat model routers as effectively an information or news stream about which model is next. It's kind of a free advertisement every single time a model drops. You know, now Stripe can tweet, “New model on our router,” and you see Stripe's name with this news cycle.
I think it's very common for people to use new model drops and all this news as a way of keeping up with AI in general. New models will likely continue to drop as, one, it gets easier. It's just going to become fundamentally easy to build models. And two, with this idea of sovereign intelligence, just like humans, we're going to have models with tons of different opinions and tons of different perspectives.
That, I think, is going to play nicely into how people operate in today's world. A lot of the time, you don't go with a business because it's purely the best-performing. You go because you like the person who started it, or you want to buy from someone you saw on the news or on TV and agreed with. Models are going to be like that as well. They'll emit opinions and have takes that are different from the ones that are most popular, and people will gravitate toward those.
I see this getting much faster and even broader before it shrinks.
A lot of guests on the show before have made bold statements that 70% or 80% of the Neo labs we have today will die in a given time period—3 to 5 years, whatever you want to choose. Do you think that's true? And how would you advise me and other investors on the AI labs that will thrive versus die in this next wave?
I think it's plausible that it's even more. I think it could be 80 to 90% of Neo labs die in the next 18 months. “Die” is going to be a funny word to use because, for a lot of them, it'll probably mean incredible outcomes. I don't know if it's necessarily doom and gloom as much as these businesses may not make sense as independent businesses.
A lot of what I think matters for an AI lab is that you should ask questions like: one, is this business attached to a durable workflow? Two, is that workflow going to change if new frontier models get better? And three, if this workflow were to be introduced to a new business, would that new business figure out something even better?
Basically, is it durable to an entirely new way of thinking or a new way of working? Legal is a great example where I think, one, new models won't necessarily get better without access to the data; two, it's obviously a very proprietary workflow; and three, we're still going to have a legal system in 5, 10, or 20 years. So probably all the Neo labs focused on legal are going to have great outcomes.
By contrast, I would argue that there are some places, like a lot of knowledge work related to intermediate tasks—people operating in Excel and Jira—that just aren't going to be differentiated. The workflows are very common. I think we may not use a lot of tools like that in 5 to 10 years, so general computer use and all this other stuff may not be as valuable as an independent business.
The thing that strikes me is just the misalignment in capability progression. When you look at coding and customer service—
Yep.
—amazing, undeniable. Legal is good, but not at the same level as coding and customer service. And then other things, honestly, like marketing copy and visuals, are so far from being there. If I wanted to use AI to clip this show, it misses both of our faces because it goes through the middle.
Yep.
It has no understanding of how to align clips between an audio edit and a video edit. It's so far off. Will we see a real, multiyear time lag between different sectoral capabilities progressing in the same way that coding has?
Definitely. The biggest reason why, for example, clipping a podcast is still such a hard problem for models is likely because the handful of businesses that deal with media haven't devoted 100% of their time to taking the knowledge that lives inside people's heads and bringing it into AI.
The moment we start to see businesses capitalize on that delta, I think the progression will happen extremely quickly. It's purely a matter of time before most workflows become something where a business capitalizes on that first bit, which is a workflow that currently has proprietary data or proprietary knowledge.
You don't normally think of clipping a podcast as proprietary, but I think, for the most part, it's a real skill. A lot of people couldn't even describe how they know when to make the right clip, and that intuition—writing it down—is hard.
I totally think so. Knowing the hook, if you started 10 seconds earlier and the hook was 10 seconds in, your chance of virality goes down significantly. The skill of knowing what is a hit kind of goes to taste.
100%.
We always hear about the Chinese open-source ecosystem and the questions around security and everything in between. Do you think those are justified, or do you think we should leverage it and thank them for their capabilities?
9. Open Models Challenge Frontier Labs
Calling open-source models Chinese models is a psyop by the frontier labs to trick people into thinking that they're scary and to otherize them. In reality, in my mind, the open models that come from a bunch of different places are not any different from an existing frontier model from a lab. They just happen to have been created by people a couple thousand miles away.
There are some very real challenges with taking in models from really any business. What's interesting is that they're actually the same challenges as with models from OpenAI and Anthropic. You should ask questions about all of your models: one, what is potentially being censored by the creators of these models? Two, are these models going to be able to solve the problems that I care about? And three, if this model goes away in 6 months or 12 months, will I be able to switch to something else?
If the answer to all these things is yes, yes, yes, then I do think that there will be concerns about using those models. For me, the Chinese models specifically have demonstrated no examples where they have some sort of security risk or backdoor compared to American models. Instead, they are just biased in a way that American models are biased toward their own creators and preferences. These are things you have to be aware of, but they typically don't change the day-to-day.
So you don't think that American companies should be concerned about using open-source Chinese models?
As of today, the current Chinese frontier models should be analyzed by American companies for the tasks that they care about.
If you’re, for example, working on national security in the United States, you definitely should not be using Chinese models. But I do think that we have to be clear-eyed and say that for a code review, it’s very likely that a Chinese model and an American model will give you the same result.
I’ll give you a great example of this. If you are writing a 10-K that expresses your business’s current state and some of the examples in preparation for sharing financials with investors, let’s say that part of your strategy is about introducing recursive self-improvement to models, and you care about AI and your business is leaning heavily into it. If you use a model from this provider, it will block you. The answer is Anthropic.
And so if you are writing a memo about American defense or preparation, I highly recommend against using a Chinese model. But all of these are basically contextual, based on the preferences of the creator of the model. It’s just like any other technology. You have to be aware of who created it, and you have to be careful, because if the person who created it doesn’t want you doing the things that you’re going to do with that model, it is going to be harder. That’s something that I think people need to be aware of for all models, though.
Guillermo from Vercel tweeted last night, or yesterday, about the weighting or usage of open models significantly increasing at a much faster rate than tokens used on closed or frontier models. What percentage of workflows will be completed with open models in 3 years’ time?
In 3 years, 99% of workflows are going to be done on open models. But 1% of those tasks is probably going to be 30% or 40% of the economic value of the future of intelligence.
Wow, but you still think that 60% will flow to open models?
I think almost all usage of models in 3 years is going to be primarily open, but that difference between frontier and open is going to become actually larger than it is today. I think that this is actually pretty aligned with how—maybe that percentage is overweighted—but you’ve listened to how Sam and Dario talk. The use cases that they’re talking about their frontier models being used for are incredibly niche.
They’re talking about bio research at the frontier. They’re talking about super-advanced LLM and AI development. They’re talking about security and defense. These are use cases that really only fit a very specific profile of effectively the frontier of science and technology.
This is where I think labs like OpenAI and Anthropic can actually be incredibly differentiated, because they already have the muscle to work on those very frontier problems. But if you go into any business in the Global 2000 today and ask any random person, “What are you doing today?” it’s not something that needs the true frontier of intelligence 99% of the time. And so cost will dominate.
OpenAI and Anthropic could release an open model that they then run inference on, and I think that that could actually be a great business for them.
Given the commoditization of the model layer, like we’ve spoken about, people seemingly chastise Microsoft. Given that commoditization, do you think Microsoft has actually played a great hand in having a little bit of a bet through OpenAI, but not being tied down with an extensive model investment layer?
10. Microsoft Wins Model Independence
I think Microsoft might be one of the best-positioned hyperscalers, honestly, with respect to AI, because of this independence. Right now, Satya has played a masterful game of getting huge upside from the OpenAI investment and work. Kevin Scott as well, who I know was sourcing a lot of that deal. That team found a lot of the potential of what AI was going to be, but they currently realize that one provider is simply not sufficient to cover what intelligence is needed in the enterprise.
And so now they’ve moved towards more concretely expressing, one, that they support all AI developers, and Azure should be a place for inference on Anthropic, OpenAI, and, most importantly, open models. But two, even if you do want that frontier intelligence, you can come here.
I think that that duality, that model independence, is going to be massively valuable for a business that wants to accelerate work for all other businesses. I think it’s a very clever positioning, because they captured the upside with a bet, and now they’re capitalizing on the market as a whole.
Is Zuck wrong, then, to be putting as much money as he is into Spark and building out that program?
I think he’s right for humanity in that we need more open models like that, especially American-made open models. I think that a lot of people, even with respect to what I said earlier on Chinese models, are going to be biased. And so open models are great because it just increases adoption in America and abroad.
But two, I think that as a business, they’re going to have to build on top of that and take what they did with Meta, the monstrous consumer business that it is. They need to power all of their operations with these new models, and the investment will be well worth it.
So would you buy Meta or Microsoft today if you could only buy one?
If I could only buy one, Microsoft for sure.
Wow.
The biggest thing that Microsoft has going for it is that infrastructure. They own so many of these data centers. They’re spending so much on build-out. No matter what model runs on top of that, Microsoft is going to win.
Do you worry about the debt cycle? What I mean by that is, there is so much cash being put out into the data center build-out, and we’ve just never seen levels of debt like this before. You’re seeing that in bond pricing for Meta and other companies.
Do you worry about that sustainability, or do you just think, “Fuck it, we’re still so early”?
I think that it actually should start to concern investors. If you do not have a huge amount of free cash flow, then taking on massive amounts of debt is bad. It’s dangerous.
And so for the Microsofts and the Googles of the world, they have access to businesses that are durable, existing, and have huge barriers to entry. As a result, those businesses may take a hit from a future collapse in value, or even if it’s not a collapse, just a minor hit in the projected future cash flow from AI. Those businesses are still going to be around.
I think it’s a risk. But to be completely honest, if you’re OpenAI or Anthropic, the hundreds of billions in free cash flow that you need in order to pay back the debt that you’re taking on to accommodate these data center build-outs and get the next big training run—it’s totally existential for them.
And so they need to become the single greatest free-cash-flowing businesses in the history of technology in order for them to just live. It’s a—
It’s a high bar.
It’s a pretty massive bar, yeah.
Do you think they should be moving into the chip layer as they are? I mean, we’ve got the wonderfully named Jalapeño, and then we have Anthropic now reportedly working on its own chips. Do you think that is the right move?
I think so. I think verticalization is clearly the strongest way to free yourself from taking on a massive amount of debt and burden. In the future, if they become multitrillion-dollar companies, they will simply have to enter this market and own more of that infrastructure layer. So it seems quite obvious that they want to play there.
I think the biggest question is how this changes their relationship in the medium term with their current vendors, because ultimately their explicit goal is to replace their dependency on them. I think that’s going to be an interesting challenge to navigate.
I think it’s a little bit like competition in VC, which is, no one really has any loyalty anymore.
Yeah. No.
I’m so sorry to say that—
No, it’s true.
—but you know what I mean?
It’s true.
Yeah. And Jensen’s like, “I’m working on Nemotron, I’m buying Poolside.” Sam knows that fully.
It’s coopetition with everyone.
Yeah. Listen, this is our job. We have to survive.
Yep.
Do you worry about where we are in the market today? I have older, wiser friends being like, “Harry, this is peak froth.” And then I’m also like, peak froth, but also Cursor just sold for $60 billion after 4 years. That is cash that’s coming back to hospitals and foundations. That ain’t froth or IRR; that’s cash back.
Yeah. No, I think that what I am less concerned about is a 2008-style financial crisis or massive bubble or asset crash. I think that seems disconnected from the true reality of where this technology is and is going.
We’ve already started to see outcomes in science and in serious, real human-prosperity-style changes. It’s very early, but the technology pretty clearly, to almost all experts in the fields of science and technology, is on a trajectory towards accelerating human prosperity. That is hard to deny from an outcome perspective.
Now, what’s more interesting is which businesses are going to be the backbones of that transformation? We are probably in the Yahoo era, where we don’t actually have, or at least widely recognize, the Googles of the world—or the thing that comes after.
And so today, when I look at OpenAI and Anthropic, there are, I think, more analogies to Netscape and these technology companies that were first. But ultimately, it’s hard to navigate being first.
I think Steve Jobs always had a really great strategy at Apple—not being first, but being the best at almost everything they did. That’s something that we also like a lot.
You know, being first or best, but we heavily lean toward being best.
I’m continuously changing my mind on outcome sizes.
Mm-hmm.
We’re also an investor in a lot of the builders of the world. Suddenly, it’s a $13.5 billion business at $600–700 million of ARR. Fuck, dude, the trajectory of company growth is just unparalleled. Do you think investors need to change their mindset on outcome expectations and company growth expectations?
Yeah. I think that it is hard because there’s a current space where people are experimenting and buying a lot of technology in a fairly speculative way. And so it’s hard to index on AI for every other possible industry.
But you start to look at some of the other industries that are being influenced by this wave, even CPG, right? It feels like every other day you hear about a massive brand that got bought for billions of dollars that started 2 or 3 years ago. So I think that maybe it is just true that the world gets faster and grows bigger and better than ever before, and we’re starting to see the early days.
I think that a lot of people say this is what the singularity will feel like. Things will just move faster, things will grow bigger, there will be more, and we’ll start to normalize it and build models and say, “Oh yeah, that’s just the way it is today.” But I think it might just be us expanding as an economy.
Before we move to internals, which I do want to touch on, because you’ve got really interesting takes on hiring, we saw Airtable go for $2.5 billion, give or take. Listen, it’s a fantastic outcome, incredible, but it’s just a reduction from the $11 billion price before. Will we see a generation of SaaS companies sell and exit before we see this wave of cannibalization that could occur?
Yeah, I mean, I think that certainly we will. I heard this from a fellow founder who told me this, and I totally resonate. Basically, contemporary SaaS businesses are more like movie studios now, where you have to hit a blockbuster, and you have to keep hitting blockbusters in order to keep the attention of the world.
And if you are Airtable, you made that one movie that was a hit, and people loved it, so no doubt it’s a great business and a great outcome. But if you rest on that, then yeah, Bending Spoons will come and eat you. And I think that the new world is much more about continuously getting bigger and growing larger.
I won’t be surprised to see a huge wave of M&A of these businesses because they’re still good businesses fundamentally, or you can make them good businesses. They just aren’t going to be Stripe or these massive things that capture fundamental pieces of the economy.
I think a very good parallel is also gaming companies, where you have a banger of a game and a hardcore user base that will sustain for 5–7 years, and that life cycle is great, but you need another big, big hit.
100%.
I totally get you there. Listen, your hiring is candidly different. When we were chatting before, you said, “We expect 100% of our future hires to come through acquiring companies and bringing their founders and teams into Factory.” I read this and I was honestly like, “Wow, that’s hard.” Because often founders make bad employees. I would suck as an employee. How do you determine whether a team will thrive internally within Factory, or whether that’s an opinionated, pretty arrogant, egotistical founder who’s not a team player?
I think that’s a great question. To me, the most interesting thing about this is that the profile of an organization has changed so rapidly that acquiring an organization is no longer what it was 10 years ago.
What I mean by that is, if you are someone who just created an open-source project, then you are quitting your job and spending 5–8 months just building that one thing, and that shows so much more conviction than you can track in an interview. And so it’s much easier, actually, to spot these talented people, who are sometimes companies of one, who are ready to execute.
They want to be a part of a mission, or they’re already operating toward a mission, and basically what you’re offering them is more resources to do that. And so I think that a lot of what we’re looking for when we try to bring a team in is: Are you mission-aligned? Are you someone who’s going to operate independently and be able to take on a huge amount of responsibility?
What’s interesting is that all of these people are also aware of this lack of technology moat, and so they’re pretty willing and ready to either integrate everything they did in a couple of days or scrap what they’ve been working on in order to build something even bigger. And so for us, this profile is just such a match made in heaven for a team that already operates with a ton of former founders and a ton of people who were previously working at startups.
Can I be a dick?
Please.
“Mission-aligned.” Of course, no one wants someone who’s not mission-aligned.
Right, right.
And then also willing to take on responsibility. It’s not groundbreaking. What—do you know what I mean?
No, I totally get what you’re saying. In my mind, mission-aligned for us means that you’re literally working on the exact problem that we’re working on and doing it very well.
I think that when people come in and say, “I loved that we were working on payments,” in a way, payments is just like AI for software development. You’re kind of like, “All right, I’m sure that you absolutely could work on this,” and in fact, maybe we’ll actually hire that person, right? So I’m not suggesting that if you worked on payments, you can’t join Factory.
But the people that we’re looking at are already deep in the weeds of building harnesses for software development, where they’ve already built something that tens of thousands of people are using on a daily basis, and maybe they even say, “That’s better than the stuff we’re getting from Factory.”
Or they’ve thought so deeply about the problem of outcomes in AI and measuring that, and they’ve already sat down with business leaders to say, “I want to solve how to translate these AI inputs into AI outcomes.” And so mission-aligned to me isn’t a property that you can suggest or say. It’s actually extremely evident in the work of the founder.
That has made it very easy to stress-test whether you’re going to do well at a company, because you’re basically doing the same thing, except backed by us.
“Mission-aligned” goes against what Chamath has got a lot of heat over the weekend for saying.
Yep.
Which was that Silicon Valley has become too money-centric, and that’s now a problem. Do you think he’s right?
That is something that it sounds like Chamath would say because, frankly, that’s how he’s made his money. I think about what we’re doing and how we got started. When we created the concept of the software factory, something that he loves to use that term for as well, we said to ourselves, “We’re looking at this very hard, fuzzy, ambiguous problem. We’re eating so much glass. We’re going up to people saying, ‘This technology is going to exist. Here’s how it works.’”
And people would say, “Leave. Go away.” It has nothing to do with the pursuit of an outcome, because at that point, you’re literally trying to be right about a technology. And so it’s very producty, very engineering-heavy, and contrary to what the world is saying, or at least people are saying: maybe at some point in the future, they’re going to dismiss you.
In my mind, that sort of mindset—“I’m trying to build a thing. I see the way the future might look. I’m going to get super in the weeds. I’m going to have people tell me I’m wrong every single day for 3 years straight until they eventually agree with me”—I think that’s actually felt in San Francisco.
Everywhere you go in Silicon Valley, there are tons of people who just care about technology and about doing something that might change the world. And I think that what you get around that, though, are people who do want to profit on top of that, and their only way of participating, because their background might not be in building, is to try to build financial instruments around it.
That’s actually a healthy and important part of the ecosystem, but it’s not the only thing that happens in San Francisco.
I actually think it’s the paradox of what Chamath says, which is that I think the influx of money has led to 1% realizing, “I’ve got plenty of money. Whatever I do, I can always go back to a big company, to a great company, and get paid a lot.”
100%.
“So I’m going to choose to work on something that’s really interesting.” Do you see what I mean?
Yeah, and there’s just no shortage of people who are purely trying to realize a dream in San Francisco. It is amazing to see, and I think that it’s actually quite harmful to that culture that people really want to create a narrative that it’s purely financially seeking.
Because in today’s world, the words that you say and how you portray something become part of the story that the intelligent systems we’re building ingest. They’re world-model LLMs and the tools that we’re going to use to do work on a daily basis for the next decade. That technology is built on the stories that we tell.
So I’m always trying to share a little bit more of the optimistic side of how I perceive the world to be, because I think that actually helps make that world occur with a higher probability.
We’ve talked about team additions.
Mm-hmm.
Cognition places a lot of emphasis on the chess champion and the math prodigy. Do you think we are overweighting the importance of traditional certification, or do you think that is the right thing to focus on in a more verifiable, engineering-heavy hiring process?
Yeah, I think that it is conventional hiring wisdom to look at pedigree and achievements and say that's the right way to pick people who are going to be smart. I think, though, that in many ways the least agentic path that you could take is to only try to hit the goals that other people set in front of you. That tends to look like you go to the right school, do the right competitions, and follow the rules well enough that you then get recognized for how well you follow rules or operate within the system.
There are plenty of smart people who are going to do that because that's also how you almost guarantee a great outcome for your life. To be clear, you can still find many smart people who follow that path. But in our mind, the most important trait to hire for is how capable you are of operating outside the bounds of what today the system calls the rules. That is something that is very hard to measure for.
And so ultimately, I think you need to look outside of that traditional pedigree and start to look for people whom that system might have overlooked. I know that I'm a big proponent of this. I went to an Ivy League school, and I learned firsthand that that is barely a signal for competence. There are plenty of idiots who went to Ivy League schools, and I think the clearest signal for me that someone's done something great is that they have built something that they care about and that they want to tell the world about.
I think that you can see that all over, and that's why I say it's not just companies that we think will make up the people we, quote-unquote, “acquire”; it's also 1-person shops that just built something in their spare time that demonstrates they are going to go outside the boundaries of what traditional systems would reward.
How do you think about placing a value on those 1-person and small teams? We're seeing Poolside being bought and a lot of the employees moving over to NVIDIA at a rumored price of $12 billion. How the fuck do we put a price tag on heads?
No, it's a great question. This is ultimately probably one of the biggest challenges of capitalism in general: this desire to place a value on humans and talent, which is challenging. Some of the best outcomes in history have effectively come out of 1 person making a gut check or the right call.
People say this about Jeff Bezos. Is he worth $100 billion, $200 billion, or is it the company that built it? In my mind, there is something magical that happens in the connections between people. It's not that any 1 node is worth $100 million, but when you put all of these nodes together, the graph they make can be worth tens of billions of dollars.
I think that's actually one of the interesting parts about a talent strategy: you have to constantly be thinking about the graph you're building. Those nodes can come in, and there are traits and properties that you want them to have, but the best people are people who make that graph much stronger than it was before.
That, I think, can easily be worth $10 billion, $20 billion, $30 billion or more, especially to companies that are operating in a model where their talent was built pre-AI. Any business whose graph was decided before this insanely game-changing technology was created has to update its graph very rapidly. Bringing people in can be the difference between a $2 trillion company being a $4 trillion company. Almost anything's worth it.
Penultimate one before we do a quick fire. What do you see other founders make in terms of mistakes on hiring that makes you go, “Oh, no. Sarah or Simon, I wish you hadn't done that”?
I think the biggest thing is performative work culture: founders who look for people who say, “I'm going to grind 24/7, I'm going to be unstoppable, and I'm going to work myself to the bone.” I think a lot of founders see that, and they catch a signal that this person is going to be really productive. But what we have found is that this sort of performative work culture, this 9-9-6 attitude, is almost always correlated with making up for some other detractor or trait that basically means this person might not be a great hire.
I think that actually applies at the company level as well. All the companies that, for the most part, say, “We work people to the bone Saturdays and Sundays, and have to have people in all the time”—at different points in your company's life, you will work on weekends. You will work 15 hours in a row. That's going to happen. I think trying to make that your culture points out that your business doesn't make a ton of sense without it.
I get in a lot of trouble in the UK and in Europe for being the 9-9-6 guy.
Yeah.
I think people take 9-9-6 from me too literally. I definitely do not mean 9:00 a.m. to 9:00 p.m., 6 days a week. I mean a culture where, if I ping you on Sunday morning saying, “Hey, a big client has a problem. We need to jump on a call with the buyer there,” you jump on Sunday morning. It probably won't happen, but it's not, “I'm sorry, it's my weekend, and I will resume on Monday.”
Yep.
I am constantly talking to people on weekends, and we're doing stuff that indicates that the business operates outside of Monday through Friday.
Sure.
I think, though, that there's clearly a difference, and that's why I say “performative work culture.” Any time you create an incentive to show people that you're working rather than to actually do work, you're incentivizing the wrong thing.
Really focusing the business on outcomes is important. It's interesting that you just mentioned jumping on a call with a client in order to achieve something. It's pretty easy to see that you're not talking about the act itself; you're talking about the outcome you want to achieve.
Yeah.
That difference is pretty massive when you're hiring. I see a lot of people think that they're making the right call by only selecting for people who have that trait, and I think you miss out on a lot of great talent who know that that's bullshit.
I think the other thing is senior engineering talent, especially those with families, are instantly put off by the performative, often young hustle culture. I've learned that the leverage you get, especially when it comes to infrastructure engineering or architectural engineering, is very real.
Right now, being able to point a set of agents in the right direction and knowing from the beginning which direction to go in is worth not only more because it gets the job done faster, but it now translates into real dollars. If you spend 100 times more tokens trying to get an outcome because you just don't know as much, it doesn't matter that you worked harder. It basically means that you missed the ball the first time.
I had Brandon on the show from Macaw, and he said that they spend more on tokens than they do on engineering headcount. My dear friend Jason Lampkin from Sasta, who we do a weekly show with Rory, said, “We'll give $100,000 of tokens to our best engineers.”
Yep.
Where do you sit on that today, and how do you think that changes?
We actually don't even think about allocating tokens or credits toward people like that. I think that's a very weird way to think about it, and I think it gets at the fact that people are ultimately looking at inputs.
What we think about is how many tokens, or how much spend, we allocate toward projects and outcomes. I'll give you a great example. There's an evaluation that we've been hill-climbing against in order to see if we can build a system that beats it. It's incredibly difficult. It's called Program Bench.
One of the things that we've done is effectively allocated almost 7 figures of credits in 1 day on this benchmark. The work was currently being done by 1 person, so I guess you could say that we allocated 7 figures of credits to that person. But in reality, what we're trying to do is see whether our research pans out on this project. Of course, we're willing to spend that much in order to see if that outcome comes true.
Similarly, for a lot of our engineers, what we do is scope out these projects. Once you know the scale and scope of the project, all you have to do is share the bid for what you think it's going to cost, and then we send it off.
We have products called Agent Effectiveness that let you allocate and look at how many credits were spent on a given project in order to measure whether you're achieving the outcomes you'd like in your business, given the spend you're putting toward those areas.
I think that in this new world, the relationship of 1-to-1 mapping between agents and humans—or saying that an agent has a name—is a weird way to think about it. Really, you have an agent system, and you allocate capital toward projects.
I see that number for some businesses approaching 8 and 9 figures easily.
Final one before we do a quick fire.
Yeah.
What really good idea did you say no to that was very hard to say no to?
I think self-service is by far the hardest thing for us to continuously say no to. It is actually quite painful as a builder of products. I want more people to use our product, and there are certain types of adjustments that we could make. I think many of them would unfortunately come at odds with either making our business more successful or improving the experience for the enterprise.
That’s something that I don’t see as being permanent. I do think we’re going to hit a certain scale where we’re allowed to pursue many things at once, but it constantly nags at me that I can’t just hit a button and then have 10 million people using the product. When we look in the market at comparable solutions that have millions of users, basically the biggest difference is economics, and that’s it. We know from a product perspective that we have a lot there; we just have to make the hard decision not to subsidize.
If you own the outcomes of those consumers and you get that data back, could you not make an argument that the improvements that data would provide to the core product would outweigh the cost to serve those free self-service users?
That’s actually how we use self-service today. We have a product that, obviously, you can download off the internet and try. To be clear, we have a pretty steady stream of tens of thousands of people who use the product every day from that segment.
Those self-service users are giving us feedback and helping shape more of the experience for the individual user. At this point, I think a lot of that can be achieved with fewer than 250,000 people. Once you start to hit critical mass, you get that feedback loop and you have everything you need. You don’t need 5 or 10 million people to get those bug fixes in.
However, there is something really special about seeing the community build media, content, and storytelling around your product. That’s a part of the experience that we have to work really hard to create a comparable version of.
Yeah, a grassroots community—
Yes.
—that kind of grassroots brand is harder if you don’t have that identity.
Yeah, 100%.
We’re going to do a quick-fire round.
Let’s do it.
In the UK, we have a game, and I’m probably going to get in trouble for this, called “Shag, Marry, Kill.”
Right.
Brutal, but you get the theory, which is buy for the short term, buy for the long term, and sell hard.
Yep.
Meta, Microsoft, and NVIDIA.
Ooh, this is a fun one. I think I would have to say marry Microsoft, shag NVIDIA, and kill Meta, and I’ll tell you why.
Microsoft, to me, represents a software company that has become an everything company that really does sit in the lifeblood of almost every Fortune 500 company. There’s not a single business that doesn’t have something from Microsoft, and I think that means that company is going to be here for a very long time. I say this also as a former Microsoft employee for a year and a half.
NVIDIA is the current kingmaker of technology. They get to decide who is currently even sitting at the table, and I have no doubt that’s going to continue to grow massively. As for how durable that is, I think it’s just a matter of whether they accumulate power at the right rate. If they do it at the current rate, then at some point people might ask, “Are we going to let one company control the whole supply chain?” But today, I think the answer is probably, “What other choice do we have?” So they’re going to keep growing.
And Meta, I say this with love for the company, but I think that ultimately, from a technology perspective, they’re often quite accurate. I think Zuck’s push into VR was technologically correct. That felt like the right move. I think the push now into open models is technologically correct.
They have one cash cow, which is their ads business, and I think that means they’re going to have to work really hard to figure out if there’s anything other than that that can sustain the business. That makes it the weakest of the 3.
Will NVIDIA be a $10 trillion business in 3 years?
In 3 years?
Yeah.
I think there’s a real, serious chance that if we let SpaceX be worth $2 or $3 trillion, then NVIDIA probably is worth $10 trillion. So I think the answer is likely yes.
It depends on the buoyancy of multiples.
Yeah.
You said something about the entrenchment of Microsoft in businesses. Would you be a buy or a sell on Salesforce, given that?
Oh, I’m a buy on Salesforce.
Really?
I believe that the businesses that are going to be most durable are the ones that have a workflow and a system of record they’ve defined that everyone agrees is consensus. You look at the Salesforce of the world, and today, at least, you look at Atlassian.
I think the biggest thing is that when people say, “I hate that software,” and everyone buys it, that’s probably a pretty good business. They’re not buying the software; they’re buying what’s underneath it. To me, that’s actually much more durable than the technology itself.
I’m an investor in Linear. Linear are absolutely crushing, though.
Yeah.
Atlassian are crushing, and so are Linear. It just goes to show, again, that the market size is so much bigger than anyone comprehends. I think we think too much in zero-sum terms: I take from you, so you lose.
Yep.
But we’re both just crushing, actually.
I think that’s really true. One thing that’s important is that Linear and Atlassian are selling the same thing: the same workflow, agile, and a system of record that represents agile.
I will say, though, that one thing that’s pretty clear to me is that the way we build software is fundamentally changing. I think agile might be one of the things that gets hit by this new way of developing. That makes me think there’s a huge opening for what the next system of record and the next workflow look like.
I do think it’s going to have to be much more radical, and it’ll be very challenging for both Linear and Atlassian to transform into that new way of building.
Codex, Cursor, Cognition, Claw Code. Rank them 1 through 4 in terms of the threat level you feel from them.
Threat level is interesting. I think the current ranking would be—actually, I have to qualify this by saying I don’t feel an impending threat from the rise of these players. They’re all going in a different direction from what we’re building, and that’s really important. I’ll touch on that in a second.
In terms of relevance to the conversations when I’m talking to enterprise buyers, number 1 is Claude Code. It’s brought up in every single conversation, and we always have to share with people why we see ourselves as largely complementary to the Anthropic platform and suite.
The second is Codex, which has increasingly been referenced in deals and conversations where people are saying, “Look, we were on Claude Code, and now we’re switching to Codex.” That’s actually the greatest news for us because it shows how unsticky this is and it gives uncertainty. However, they’re coming up much more frequently now, and I think it’s because their work platform is better than Anthropic’s. It’s not Codex for coding, but rather Codex for Work that’s coming up more frequently, which is fascinating.
Then I would say Cognition, because they’re basically the only other model-independent vendor in the enterprise. One of the consistent feedback points we hear from them is that they’re building this cloud offering. It’s very futuristic, based on the idea of imitating a software engineer as a human. I think that makes people ask, “Is that different from or the same as your strategy?”
Then Cursor is present in a lot of these businesses. I don’t think anyone really perceives Cursor to be their primary enterprise software development strategy so much as an IDE, which is still a great business because they’re still going to get a lot of usage. That’s how I see them.
This is mostly informed by the idea that almost all 4 of these businesses are coming to enterprises and saying, “We’re going to build you an eventually human-level AI replacement for labor, and then we’re going to do an Indiana Jones swap and replace the people in your business.”
I think that’s so different from what we’re saying to them. You’re not going to replace humans with AI so much as you’re going to build a new system for developing software. Humans are going to build that new system alongside AI, and that new system is going to look very unfamiliar.
It’s not so much a 1-to-1 labor mapping as it is an entirely new development methodology, and that’s something they’re really only hearing from us right now.
And I think it sort of contextualizes the many tools in the space compared to Factory.
Will Chumath be successful with his... I can't remember, the 1809 or, but whatever it is?
Yeah. I think my answer would be that it depends on how real the software is. I haven't seen any examples of it working in an enterprise environment. If they're very focused on building software that works and delivers outcomes, I believe that he has just as good a chance as anyone and is very well connected. But ultimately, I do think that part of this is about building with the enterprise. So I think the biggest question is: Can they get enterprise traction in the markets that matter with people who take them seriously as a full-time software development opportunity?
Single biggest piece of advice on selling to large enterprises in today's world?
I think the biggest thing that I've learned about selling to enterprises is to stop treating it like persuasion, where you're trying to convince them that you're right, and instead treat it as a discovery opportunity to learn about what's currently the biggest problem they care about.
That approach difference is unique to new markets. So if we're selling something where the market's established, it's finite, zero-sum, and everyone knows it's a commodity, like databases where there are a million options, I do think persuasion is the strategy. You're trying to convince them that, all else being equal, they should buy from their friends.
I think that in this market, it's much more about trying to understand just how big of an opportunity it is and learning with the customer. People actually put a huge amount of value, especially in software, on people whom they perceive to be trying to problem-solve with them. If you're on the same team and you're both trying to problem-solve, then you're going to land on a real problem that that business has not yet solved, and that means almost 10 times out of 10 you're going to bring value if you can figure out a solution to their problem.
So I think that's something that people underrate. You're not trying to trick them or persuade them. You're just trying to help solve a problem for them.
Totally get you. Age-old enterprise sales doesn't change that much.
No, definitely not.
Final one. I like to ask the question: What seems ludicrous or strange today that will be incredibly commonplace in 5 years' time? I can give examples of finding your husband or wife online. Bizarre.
Yeah, that's a great one.
Putting your credit card details into your phone. Of course I'm not doing that. That's so dangerous. What today do we think is crazy that will just be obvious in 5 years' time?
I think the biggest thing that we're going to be surprised by is the fact that we let a sort of priestly class of maybe 2 million people decide the fate of all software for all of humanity. In 3 to 5 years, it'll be unthinkable that you couldn't just generate the thing that solved your problem with software on the fly, in the moment, for nearly any problem you have in front of you that can be solved by information manipulation.
Today, we see a little bit of that with Lovable and Bolts and these tools that let you build personal applications. But I think this will look quite interesting when you think about what problems not just an individual consumer has, but really what society at large has.
You'll be on vacation in Belize, and the boat operator that gets you from point A to point B will have software and an interface fully customized to them that looks better than your HR or IT software back at home. I think this total disbursement and distribution of amazing software to the entire world is going to make everything just feel way more futuristic, and that's going to happen very, very quickly, on the order of 3 to 5 years from now.
I love doing what I do because I genuinely get to pursue my own curiosity in a very natural way. So thank you so much for entertaining my curiosity, and you've been an amazing guest.
Thanks for having me. This was a fantastic conversation.