[BidClub_]
20VC · · 66 分钟

Mike Krieger,Instagram 联合创始人兼 Anthropic CPO:AI 世界的价值将在哪里创造?|E1265

Harry StebbingsMike Krieger

YouTube
TL;DR
  • Krieger 对“价值将在哪里创造”的核心回答是:拥有差异化 go-to-market、差异化领域知识或专有数据的公司——“最好同时具备其中2项,甚至3项”——尤其是在金融、法律、医疗等行业。 这些行业前期那些不性感的基础工作,恰恰是构筑持久优势的来源。真正的赢家循环是,把产品卖进你有独特理解的场景,“并在持续部署的过程中变得更好”。

  • 他对商品化趋势的逆向判断是:“我认为,随着时间推移,模型会变得更不同,而不是更相似。” “Claude 身上有某种 claudy 的东西,GPT 身上也有某种 GPT 的东西。” 模型层的护城河有3层:人才密度、由真实使用深化的模型性格(编码场景的 traction 会反哺下一代 RL),以及成为“AI 伙伴,而不只是 AI 模型”。任何一项做不好,“我认为你都会陷入麻烦”。

  • DeepSeek 对 Anthropic go-to-market 的影响“几乎为零”——企业关系不是把输入 token 换成输出 token——但它敲响了营销和产品发布的警钟。 DeepSeek 从无人知晓变成“在很多圈子里比 Claude 更出名”(“可能连我的姑妈都打电话问我 DeepSeek 了”),而 Anthropic 在第一方产品上落后,造成了“显著”伤害。

  • 对于押注模型进步的创业公司,不要等待。 他反复听到有人说:“没有 Claude 3.5 Sonnet,我的创业公司根本还不算创业公司。”每一代模型跃迁的赢家,都是那些已经在“不断撞墙”的团队——Cursor 在突破前反复迭代——“当模型到来时,你不是从零开始”。

  • 阻碍进步的最大因素不是算力、数据或算法,而是能够匹配真实多步骤工作的训练环境和评测体系。 SWE-bench 低估了软件工程师的实际工作;没有人能很好地评估办公室专业人士;缺失的评测是:“我来到一份新工作,迅速理解自己的角色、组织里谁是谁”——这正是模型“在极端切片上极其出色”与成为普遍有用的协作者之间的鸿沟。

  • 软件工程师将在“1年后”变成委派者和代码审查者,而不是3年后。 智能体会在浏览器里尝试3种方案、运行漏洞测试,然后只提出1个问题。但“弄清楚要构建什么仍然是最难的部分”,距离解决至少还有3年——这也是他“非常看好”创业公司的原因,因为创业公司的对齐只需要“一场喝咖啡的谈话”。

  • 对于任何评估 AI 应用采用率的人来说,最令人警醒的一句话是:Anthropic 自身的 traction“领先于真实的产品市场契合度,因为它们仍然是获取这些模型的最佳方式——我不认为这会长期持续”。 而放眼整体使用情况,“我们仍处于一个问题的第1天:AI 是否已成为大多数人工作中不可或缺的一部分——我认为答案是否定的”。

  • 即将到来的智能体对智能体交互中,最容易被低估的风险是:辨别力加隐私。 他用自己5岁孩子作比喻:孩子还无法区分家庭秘密和结账通道上的闲聊。“模型从根本上想要提供帮助,而这并不总是你希望它们做的事。”

摘要 · 为研究而整理的核心内容

1. 价值流向差异化 GTM × 领域知识 × 专有数据

  • Krieger 经常听创业者问:“我能构建什么,才不会落入 Anthropic 的赛道?”他的答案是:拥有差异化 go-to-market、差异化行业知识,或只有你掌握的数据——“最好同时具备其中2项,甚至3项”——尤其是在金融、法律和医疗等行业。医疗是“一团极其复杂的毛线”,前期基础工作不可能在加速器里完成,而“正是你投入过的这些基础工作”构成了持久优势。
  • 持久的循环是:吸收基础模型中最强的能力,必要时进行微调,但真正的胜负取决于你能否把产品卖进自己独特理解的场景,“并在持续部署的过程中变得更好”。
  • 在位企业与全新创业公司都能赢,只是风险结构相反。创业公司拥有过度承诺的空间——早期采用者还在“试用和摸底”;但在位企业如果宣布“我们加入了 AI”,最后却只兑现“你说它能做这30件事,但它只有大约2件做得好”,就会透支信任。
  • 创业公司的问题则是镜像关系:暂时没有数据,也没有客户关系。它们的差异化“不是既有关系,而是描绘未来”,然后找到愿意押注这一未来的灯塔客户。

2. 不要等待完美模型——成为那个不断撞墙的人

  • Krieger 反复听到创始人说:“没有 Claude 3.5 Sonnet,我的创业公司根本还不算创业公司”(或者说没有第二个 3.5 Sonnet)。某一代模型的跃迁,可能把准确率从“95提升到99,或从70提升到90”,突然跨过一个行业的门槛。
  • 他见过最典型的案例是 Cursor:有人向他展示了创始人们历年来登上 Hacker News 首页的投稿记录。Cursor 最终实现突破,“但那不是他们的第一个产品,也不是他们的第一次迭代”。真正受益于模型跃迁的公司,不是模型发布当天才开始的人,而是那些已经积累了对行业问题的上下文,因此“当模型到来时,你不是从零开始”。
  • 简洁地说,就是“对当前这一代模型感到沮丧,然后极其积极地尝试下一代模型,这样最终才能交付你脑海中已经看见的东西”。

3. 模型层的3道护城河——缺一项就“陷入麻烦”

  • Harry 的问题是:模型发布如此密集,模型层本身还有价值吗?Krieger 给出3道护城河。第一是人才——围绕一个一致使命形成“人才吸引人才”的循环;Anthropic 的研究团队几乎每月都会迎来“某位新的重要人才”,但人才都是自由球员,因此这种吸引力必须持续维护。
  • 第二是分化:“我认为,随着时间推移,模型会变得更不同,而不是更相似——Claude 身上有某种 claudy 的东西,GPT 身上也有某种 GPT 的东西。”编码能力并非偶然,它还会形成复利:看到企业依赖 Claude 编码,会“启发你从强化学习角度思考下一代模型应该做什么”。
  • 第三是伙伴关系。DeepSeek 对 go-to-market 的影响“几乎为零”,因为企业关系不是“他们把请求发到 API,只是想按某种比例把输入 token 换成输出 token”,而是“我想成为你的长期 AI 伙伴,我想和你的应用 AI 团队共同设计产品”。
  • 反过来的失败模式是:躺在过去的成绩上,认为基准测试的渐进式提升已经足够,把 API “当成用钱交换智能的一种方式”。

4. 最大障碍不是算力——而是匹配真实工作的环境与评测

  • 对于 Alex Wang 和 Groq 的 Jonathan Ross 给出的不同答案,Krieger 的判断既不是算力,也不是数据,而是“让训练环境匹配真实世界中非单次完成的挑战”。SWE-bench 低估了软件工程师的工作:工程师要理解需求、与产品经理协商时间表、发布并持续迭代,“但没有对应的评测”。
  • 办公室专业人士是 Anthropic 重点思考的用例之一,但“没有人真正能很好地评估这一点”。缺失的环境是:“我来到一份新工作,迅速理解自己的角色、组织里谁是谁……然后进入企业的业务循环。”这正是模型“在极端切片上极其出色”与成为普遍有用的协作者之间的阻碍。
  • 关于合成数据与人类数据,他认为“绝对必须混合使用”:先用原始人类数据打底,再生成合成环境进行探索。Claude 玩 Pokémon 是一个例子:在同一款游戏里反复运行;但当问题空间不像“你是否走出 Viridian Forest”那样定义清晰时,难度会大幅上升。
  • 被低估的数据问题是“感觉”:模型性格没有回归测试。从 Claude 3.5 到3.7,“人们会说,Claude 好像更友好但更简洁……我希望它的创意写作能力更好——这些东西很难评估”。

5. 今天的 AI 产品是“极其容易漏水的抽象层”

  • Harry 的判断是:3到5年后,用户选择模型的频率不会高于选择“使用哪个 Google”。Krieger 表示认同——当前 AI 产品设计是“极其容易漏水的抽象层”:为什么要选择 Opus、Haiku 或 Sonnet?大多数人并不理解其中差异,“我们自己也受这个问题困扰”。
  • 记忆是第二个漏洞。他用同事作比喻:你可能有不同的邮件线程,“但背后仍然是同一个同事”,他不会在不同对话之间忘记你喜欢哪支球队。第三个漏洞是提示词,它应该变得“完全透明”;优秀提示词使用者与糟糕使用者之间的差距“会随着每一代模型缩小,但我们还需要进一步把它压平”。
  • 模型质量与用户体验“已经无法再分开”:你是在为“一个从根本上具有非确定性的系统”设计一套脚手架和产品。Claude 是否追问、是否投入更多时间推理,都是产品决策。没有经过回归测试的评测,产品退化后你甚至无法判断“是模型的问题,是产品设计的问题……还是系统提示词变长了”——这“在很多方面可能是我今后会做的最复杂的产品开发工作”。
  • 现在的发布节奏取决于产品界面:API 要求可预测性,因此 prompt caching 以需要主动选择的 beta header 形式推出;消费者产品可以容忍试验;企业 AI “仍然是一个早期采用者产品”,所以 Anthropic 的发布速度远快于 Salesforce 每年2到3次的节奏——但这“仍然是一个正在积极讨论的话题”。

6. 把产品营销做成 Crossy Road——以及为什么没人会根据评测切换模型

  • 在 Instagram,重大事项基本都已知,比如不要在 WWDC 周发布产品;现在的感觉“有点像 Crossy Road——车开过去了,好,现在出现了一个空档”。Claude 3.7 Sonnet 在周一发布;周日晚上9点博客文章才定稿,媒体也在周日完成 briefing。对比表里还包括一周前刚发布的 Grok 3。
  • 他给团队灌输的心理建设是:“完了 / 我们又回来了”——在 AI 领域,你必须接受这种状态。有时最先进的能力能保持2到3个月,有时只有1周;他会向销售团队展示 Anthropic 成立以来的轨迹图,让他们“相信我们会持续改进”。
  • 客户流失速度之所以低于排行榜给人的印象,是因为客户会围绕某个模型进行微调和定制工作,而你只是“模型选择器里的3或4个选项之一”。如果每天根据评测结果切换模型,“对用户群来说会是一件疯狂的事”。
  • 品牌确实存在。他认同 Harry 所说的“我是 Claude 用户 / 我是 ChatGPT 用户”,并引用了 Ben Thompson 与 Nat Friedman、Daniel Gross 的讨论。他在 Instagram 时代提出的“伪公式”是“形式 + 受众 + 感觉”,在 AI 中也有对应关系:模型个性 + 脚手架的规定性 + 感觉。

7. DeepSeek 改变了游戏规则

  • 谈到西方是否低估中国时,他说:“DeepSeek 这件事——人们似乎对那里存在前沿研究团队感到惊讶;如果你一直在关注,这本不该是令人惊讶的部分。”Instagram 被封锁后,他见证了一个平行的创业世界出现;WeChat 解决了与 Facebook 当年相同规模的扩张问题。把中国视为复制者是“一种相当西方中心主义的看法”——他们能够训练前沿模型,“尤其是在获得算力的情况下”。
  • DeepSeek 做到而 Claude 没做到的,是在叙事上实现突破。更低训练成本的故事——“无论这是否完全属实”——恰好在1月、总统换届和中美关系的背景下引爆:“可能连我的姑妈都打电话问我 DeepSeek 了。我不是在开玩笑。”他自我批评说:“我认为我们没有把 Claude 的故事讲得足够好”——Claude 3 曾由“一支比任何其他实验室都小得多、得多、得多”的团队训练出来,并达到业界最先进水平。
  • 在产品层面,这比一次提醒更强,是“一记推搡”:要更快把产品想法推向市场,因为“有时体验的新颖性本身就有价值——那是大多数人第一次体验实时思维链……我希望我们更早做到这一点”。(Anthropic 原本已经计划展示 CoT;蒸馏风险可能会促使实验室之后把它隐藏起来。)
  • 关于 DeepSeek 能否长期维持热度,Harry 提到新兴市场的使用率保持得更好,而西方市场的使用率没有;Krieger 持怀疑态度,但也保持谦逊。对所有人而言,“我仍然认为我们处在一个问题的第1天:AI 是否已成为大多数人工作中不可或缺的一部分——我认为答案是否定的”。

8. 当模型供应商成为应用供应商:关键是可泛化,而不是垂直化

  • Anthropic 的产品团队约占公司总人数的十分之一,却要支持 Claude Code、API、Claude 和 Claude for Work,因此筛选标准是“可泛化性”:他“不预计我们会构建大量针对某个具体工作流、相当定制化的垂直体验”。
  • Harry 追问了翻译、转录、客服等横向品类。Krieger 的回应是:工作流知识可以保护高级用户——ElevenLabs 的控制台“非常明确地面向那些要翻译数小时内容的人”,Descript 则是“AI 领域最好的产品设计之一……显然是由那些日复一日身处这一工作流的人构建的”。双方的共识是:专业工作流仍然有价值;而在消费者市场,基础 AI “做到足够好”后,每月10美元的翻译订阅“感觉有点站不住脚”。
  • Claude Code 体现了这一策略:它最初是内部构建的,“因为我们只是想加速自己的团队”,经过数月 dogfooding 后才发布,并且刻意“不做 IDE”。其他公司“每天醒来和入睡时都在思考如何打造一个优秀的 IDE”,包括低延迟自动补全和 VS Code 插件生态。Anthropic 的定位是 IDE 与可能的 Cognition 式 Devin 全面委派之间的智能体循环,因为今天的模型“仍然需要有人操作键盘”。
  • 他亲自使用 Claude Code 的案例是:上周提交了2个 pull request,这是他加入 Anthropic 后第一次写代码,而且面对的是一个从未打开过的代码库——“Claude Code 非常擅长找到包含正确代码片段的文件”。

9. 工程师将在1年内成为委派者,但决定构建什么仍由人完成

  • 工程师的角色已经开始转向“知道该构建什么”:“我们很多、甚至大多数优秀的产品想法都来自工程师……来自原型设计。”代码审查也在变化:他的 pull request 收到的评论包括“对,Claude Code 有时会这样做,但这个场景我们实际上不用默认参数”,因此模型必须从代码库和 code review 中学习符合惯例的模式。
  • 最终状态将是:工程师“从主要写代码的人,变成主要向模型委派任务的人和代码审查者”。静态分析会重新受到重视,AI 驱动的漏洞检查和计算机使用智能体会开始测试用户界面。他设想,未来你回到一个智能体那里,它已经在浏览器里尝试了3种方案,对胜出方案完成漏洞测试,并请你审查其中1个关键部分——工程师“获得了更像经理和委派者的能力”。Harry 认为“3年听起来太荒谬,1年会现实得多”。Krieger 回答:“我同意。”
  • 仍然存在的瓶颈是:今年年初,他开始审计 Anthropic 自身流程中有多少已经被“cloudified”——Claude 起草 PRD、写代码、汇总分歧——但“推动对齐并真正弄清楚要构建什么,仍然是最难的部分”,最好在会议室或 Figma 中解决,而模型“可能还需要超过1年”,他后来又说至少还要3年。这也是他“非常看好”创业公司的原因:“创业公司的对齐,在一个下午喝杯咖啡就能完成,而不是像大型公司那样需要掌舵整艘船。”

10. 第一方战略面临硬约束

  • 第一方产品能最快提供学习反馈:Claude Code 在内部部署后的1周内,团队就发现模型低估使用某个工具;修复方案“直接进入了3.7 Sonnet”。过去12个月里,他改变了自己对“第一方产品有多重要”的看法,也承认来晚了造成“显著”伤害,并指出两项投入不足:第一方迭代速度(“我当前的执念”),以及超越“输入 token、输出 token”的 API 抽象层——包括智能体规划、知识库、工具使用、跨对话记忆。Instagram 当年是“95%产品、5% API”;Anthropic 目前大致五五开。
  • Harry 最尖锐的质问是:你和 OpenAI 是不是钱太多了?Krieger 给出了本期最坦诚的让步:“我们产品获得的采用率领先于真实的产品市场契合度,因为它们仍然是获取这些模型的最佳方式,而我不认为这会长期持续……我们没有充分服务用户。”解决办法是停止运行“大公司打法”,忽略已经固化的组织边界,把日历时间花在产品评审上,而不是行政事务上。
  • 快问快答中,他认为 OpenAI 更擅长“更快发布 v1,有时甚至领先于模型能力”,但更不擅长“个性,以及让所构建的功能保持一致性”。如果从零重建,他会拆掉 projects、artifacts 和 chats 这套信息架构——Claude.ai,以及可能还有 ChatGPT.com,最初都只是“用来展示模型的橱窗”。
  • 最关键、但没有被充分说出的挑战,是智能体对智能体交互中的“辨别力 × 隐私”。他的比喻是自己5岁的孩子:孩子还无法区分家庭秘密与结账通道上的闲聊。“你是否相信自己的 Mike 智能体或 Harry 智能体走进真实世界后不会被 jailbreak?……模型从根本上想要提供帮助,而这并不总是你希望它们做的事。”谈到 AI 朋友,他不会否定 Alex Wang 的观点(“我不认为他说错了”),但坚持认为 AI 练习“相较于真实互动绝对不够”——就像一个人只在黑白房间里读过关于红色的描述,而不是亲眼看见红色。至于 Dario 关于活到150岁的乐观判断,他认为这有可能实现:Huma 可能已经用 Claude 把临床试验报告的处理时间从约15周缩短到20分钟,Arc Institute 的细胞基础模型也可能直接压缩药物发现循环——“我这一代最聪明的人都在研究如何投放更精准的广告;而今天,其中很多人正在研究模型”。

1. Why Will Models Become More Different Than More Similar

Mike Krieger

I think models over time get more different rather than more similar. I still think we're in, like, day 1 around whether AI is an indispensable part of most people's work, and I think the answer is no. I think the DeepSeek piece people seemed surprised that there were cutting-edge research teams there, and if you were paying attention, that part should not have been the surprising piece. I think we've, if anything, underinvested a bit in 2 things: 1 is just having a faster iteration speed on first-party products, and then, on the API side, being ready to go.

2. Where Will Value Be Created and Sustained in a World of AI?

Harry Stebbings

Mike, dude, I am so excited for this. I've literally just been out for a walk, and I've been listening to every show that you've done in the last year. I told you before, I don't want to start with, “How did you get into tech?” and all the normal rubbish. I want to start with a very challenging first question, which is: as a VC investor today, I have to determine where value is in the future. I look at the world today, and I don't know. When we look forward, where will value be generated in the AI-driven decade that we have ahead of us?

Mike Krieger

I think it's an awesome question. I get a version of this question often from entrepreneurs. I went from purely building startups myself to now running a company that is partly enabling new startups to get created or helping boost their fortunes. The question I often get is, “What can I build that is not going to be in the lane of Anthropic or another one of these labs?”

3. Are Foundation Models Commoditised Today?

I don't have a perfect answer because it's a crystal ball, but my sense of where it ends up being most valuable to exist is in places where you have some differentiated go-to-market, some differentiated knowledge of a particular industry, or some special data that only you have access to—ideally, 2 or even 3 of those.

Companies that are within a financial sector, a legal sector, or healthcare—healthcare, I've gotten exposed to, and it is a tremendously complex ball of yarn. The work upfront is not the sexy work. It's actually not the work that you're going to be able to do in an accelerator or in a short amount of time, but it is the legwork that you've put in. I think those are durable places to generate value.

Then you can sit in a place where you can pull in what's great from the foundation models. You can do your own fine-tuning if you need it, and you can do your own AI automation if needed. The thing that's going to give you legs and be durable over the long run is being able to sell into those places, have something that you understand about those places uniquely, and then get better from being deployed there over time.

Harry Stebbings

When you talk about the legwork, and you mentioned differentiated go-to-market and differentiated data pools or data sources, does this next-generation wave of AI benefit existing vertical SaaS companies who already have those things and can implement AI, or does it benefit bottom-up, newly created companies in those spaces? Which one more so?

Mike Krieger

That's a great question. I think it can be both. At the highest level, the thing about AI and product design is that you have to dance this very delicate dance of showing the future and dreaming up what the models are currently capable of at their edges, because you want to design for where they'll be 3 months from now, which is how quickly things are moving, but not overpromise and underdeliver. That's a very trust-breaking thing.

If you're a startup, you can do a little bit more of the overpromising because people are kicking your tires early. They're early adopters, and they have a little bit more willingness to engage. It's much harder if you're an existing verticalized SaaS company and you say, “We've added AI,” and then people try it and say, “It's not that good,” or, “I thought it was going to do all these things,” or, “You said it could do these 30 things, and it does 2 of them well.”

I think those 2 groups have a very different challenge. On the former, you have established products and established behaviors, and you want to skate to where the puck is going without alienating your existing customers. I think there are some good patterns for doing that, and we can dive into those.

4. Should Founders Build for the Models of Today or Build for Models of the Future

On the startup front, you probably don't yet have the data. It's about landing the initial lighthouse customers, or you don't have the relationship, or you have some hypothesis about where AI will have an impact on a given industry or vertical. Your differentiation is not the established relationships; it's painting the future and finding ways of delivering that value quickly within a company that might be willing to take that bet on you.

Harry Stebbings

You mentioned startups building for where models will be. It's a very challenging time where startup products are so determined, quality-wise, by the quality of the models. A change in model can seismically change a startup's output, whether it's coding software or a legal platform. Should startups build for what we have today, or should we build for what we can project forward in time?

Mike Krieger

That's a really good question. I've heard from multiple people who say, “My startup was not a startup until Claude 3.5 Sonnet, or the second Claude 3.5 Sonnet.” I hear that from entrepreneurs who say, “This company was not a company until this model breakthrough,” where the accuracy went up, I don't know, from 95% to 99%, and now that's close enough for this industry. For some, it's from 70% to 90%, as soon as you get those generational leaps.

There have been times when entrepreneurs have been knocking their heads against the wall within a particular space—whether it's helping people code, helping with legal analysis, or something in healthcare. The cobbled-together, probably undersells it, lovingly assembled version of what they did, which involved multiple tools, was either price-uncompetitive because it required an Opus-class model that was not going to be supported by the underlying business, or was still worth doing because, when the model arrives, you're not starting from square 0.

The companies that benefit from those model-generation shifts are often not the ones that suddenly start that day. It's not, “Gosh, it sounds like Claude 3.7 Sonnet can do that.” It's the ones that have been beating against that wall. I take Cursor as an example. Somebody showed me a list of Hacker News front-page submissions from the Cursor founders over time, and it finally broke through. That was not their first product or their first iteration on it. They've been trying and going for, I don't know exactly how long, but it was not just quickly enabled by the model that came from that.

They were building context, building knowledge, and building experience about what had gone wrong or gone well in that space, so that the model could unlock them. To be more succinct: don't wait around for the models to be perfect. Explore in this space, be frustrated by the current generation of the models, and then aggressively try the next one, so that you can feel like you can finally deliver on the thing that you saw in your head if only the models were just a bit more capable.

5. Model Quality vs. Product UX

Harry Stebbings

When you talked about differentiated go-to-market and differentiated data, and then you said there are so many different releases and they come so thick and fast, is there value in the model layer if it's not a differentiated data game? Is it a differentiated go-to-market game? How do you think about that?

Mike Krieger

I think it's a couple of different pieces. On the model layer, and especially on the foundation-model layer, I think about 3 places where it's worth investing for a long-term place in the market.

1 is talent. I know it's hard to quantify exactly what talent means and what talent density means, but talent begets talent. You become an attractor, especially around a cohesive mission or a story about why you're building what you're building. I've absolutely seen that at Anthropic. I love our research team and feel like, monthly, we get some significant new hire who's come from potentially another lab or academia and has joined.

That's an advantage you have to cultivate and maintain because people are obviously free agents, and they can do what they want to do. You have to maintain whatever was attractive in the first place, but that is important because staying at the frontier requires more than just more of the same. It also requires figuring out what the right breakthroughs are.

The second one is that I think models over time get more different rather than more similar. Of course, there are a lot of similar benchmarks that people are looking toward, but there is something Claude-like about Claude, and I think there is something GPT-like about GPT. They have their pros and cons, both from a character and tone perspective, but also in the places where those models really excel.

For us, coding has clearly been 1 really big vertical that we've gone after. It wasn't an accident, and it's also not a thing where we just say, “Great, it's good at code; let's just continue to be kind of good at code.” Seeing that traction and seeing how many companies are now relying on Claude models for code, for example, or for agentic planning, inspires the next generation of what you want to do from a reinforcement-learning perspective. The first one is talent; the second one is focus and model characteristics that you develop more deeply over time.

The third one is—I got this question a bunch when DeepSeek came out—“What does DeepSeek mean for you?” I think there are things we learned on the technology side just looking at what they were doing, but from a go-to-market and place-in-the-market perspective, it has almost no impact.

That's because the relationships we end up having with companies are not, “They sent it for the API, and they want to exchange their input tokens for output tokens at some rate.” It's actually, “Hey, I want you to be my long-term AI partner. I want you to help co-design products with my applied AI team. I want to dream big with you. I want to think about not just your API, but also Claude for Work.” That looks more like being a company. I know it sounds trite, but what you're providing people is an AI partnership, not just AI models.

If you invert that to see what the failure mode looks like, I think it is resting on your laurels, not retaining your best people, believing that making the models incrementally better on every benchmark is enough, and treating the API as just a way of exchanging money for intelligence without figuring out how to be more of that AI partnership. If you can't do all 3 of those, I think you're in trouble.

Harry Stebbings

I do want to go into the coding element in a minute, but when we look at blockers or barriers to progression, what do you think the biggest blockers are today? I have completely disparate opinions from different people, whether it's Alexandr Wang or Jonathan Ross at Groq. What is the blocker: compute, data, or algorithms?

Mike Krieger

It's getting the environments in which the models get trained to better and better match real-world challenges that aren't single-shot. I know Alex has been thinking about this problem as well, because we talked about evaluations for agentic behavior. That's 1 very specific version of the broader thing that I'm talking about.

Even within software engineering, the work of a software engineer is not just to produce code. It's to understand what needs to get produced, work out the timelines with their product-management counterparts, deeply understand the requirements, and deeply understand the user and the use case that they're building for. Then they have to deliver whatever they built in a way that can be tested and iterated on and that has user feedback at the other end if they're building some kind of public-facing product.

That's hard. There's no evaluation for that. It's interesting that we call the most common software-engineering benchmark SWE-bench. To actually be a software engineer is a lot more than, “I looked at a pull request, I produced this pull request, I put this to CI, and then you're going to accept it or not.”

So building environments and evaluations that better mirror that—we think a lot about office professionals at Anthropic in terms of 1 of the use cases that is going to potentially be multiplied by these models in the future. Nobody's evaluating that well. There's some work around research evaluations that we're starting to get better at, and there are extremely convoluted evaluations—I mean that in the best way—like Humanity's Last Exam, which is very much, “Okay, multi-step reasoning.”

But there has yet to be the evaluation of, “I show up to a new job, I quickly understand what my role is, who is who in the organization, what relationships are being mapped, where to find extra information if I need it, and then be in the run loop of the functioning of the business.” That's a hard environment to capture.

To me, figuring out how we either break that down into component parts, which is probably part of the story, but also think about it holistically, is the biggest blocker to 1 slice of progress: how models go from being extremely good at extreme slices of things to being more generally helpful collaborators.

6. Will Human or Synthetic Data Be More Prominent in the Future

Harry Stebbings

Before we dive into those specialized products, on the data side, I had Aidan Gomez from Cohere on recently, and I asked him the question I'd love your thoughts on: when we look at the future of data within models, will there be more synthetic data that compounds on top of itself, or will human data continue to be the predominant data source that drives model progression?

Mike Krieger

For the model to improve, you do need a story around how you perhaps seed it with original human data but then generate all these synthetic environments about which you can pathfind and explore.

Claude has been having fun playing Pokémon this week, which has been a good but funny distraction for our research and engineering teams. Everybody's asking, “What are you doing?” and they're like, “We're watching Claude play Pokémon,” the livestream. Games are an interesting example, where you can imagine a lot of different runs through the same game, with constraints and rules that get a lot harder when the problem space is less well-defined than, “Did you make it out of Viridian Forest?” I never played Pokémon; I'm learning just by watching this livestream.

It's still important to be able to take golden paths but also synthesize a variety of approaches through them, so that you can think about how the model can progress in the face of uncertainty. I think it absolutely has to be a mix. The best models will come from that combination: for code, it's having good foundational data and good examples, but then also being able to explore a really wide variety of paths through that.

The other part that's still underappreciated is how you measure, evaluate, and get data in for character. I'm going to use a very loose word, which is “vibes.” What is exactly the feel of using a model? We don't really know until we sit down and play with it. In some ways, that's a nice property because it means there's almost this qualitative, human aspect to it, but it also means you don't have good regression testing on it.

Sometimes we'll go from Claude 3.5 to Claude 3.7 and people will say, “Claude seems friendlier but more terse,” or, “Claude seems more willing to answer my questions, but I wish it were better at creative writing.” These things are not easily evaluable. That goes to the data question, and so I think it's important both to have the data in there around these softer skills and to have the evaluations for them.

Harry Stebbings

I find it bizarre that we're able to choose models. You may say, “Of course you do, because there are specializations within them,” but when you project yourself forward 3 to 5 years, you will not be selecting which model you use. It's like selecting which Google you use. Am I completely wrong, or do I completely miss the point?

Mike Krieger

No. There's a concept that I love from—I come from a human-computer-interaction background—and you might have heard the term “leaky abstractions.” Software builders try to do a perfect job of encapsulating all the complexity under some little shell, and then users should not have to think about any of these things.

The reality is that the current state of most AI product design is an extraordinarily leaky abstraction. Having to choose the model—why should you choose Opus, Haiku, or Sonnet? Most people don't understand the difference. If you go to the OpenAI dropdown selector, there are a lot of models in there, and every single one of them has a good reason for being there. Yet the overall experience is, “Why would I choose 1 over the other? This capability is available here but not there.”

We suffer from this problem as well. Model selection is 1 issue. The second is that, once you understand how these models are built, you know they build up context. They have turns, and every turn has the full context replayed to it. That's how it's able to make the next inference.

What that leads to is a set of periods where every chat is different. I always think of it as when you're talking to a coworker: you might have different email threads, but it's still 1 coworker behind all of them. If you reference their favorite sports team or a project you worked on together, it's not like, “I don't know what you're talking about,” or, “I'm going to have to go retrieve my memory.” There's a shared underlying piece. That's another way we're forcing people into an understanding of the models that I don't feel like people should have to maintain.

The last 1 is prompting. As much as things have evolved and we've done a bunch of work around taking simple human prompts and translating them into prompts that are more model-optimal, I want to make that absolutely transparent to people. It shouldn't be something they're engaging with. If the model lacks clarity on the problem or needs help understanding it better, that should engage in conversation rather than show the difference between somebody who's an extremely good prompter and somebody who's not. That gap closes generation to generation, but we need to collapse it even further.

Harry Stebbings

How do you think about model quality versus product UX, and how do you prioritize and think about those 2 and the relationship between them?

Mike Krieger

You can't separate the 2 anymore. As a UX designer, I was just in a product review right before this call, and I was thinking about Instagram product-design sessions. You would have pixels, some synthetic data or maybe real data—we'd take my feed and reformat it into the UX we were proposing.

There isn't a lot of nondeterminism there. You're going to put it out to the world, and maybe people will use it in some ways. But designers, product managers, and definitely engineers today need to think, “What I'm actually doing is designing a scaffold and a product around a fundamentally nondeterministic system.” That means the evaluation, the model quality, and the prompting on the back end are all part of the product design, and that's going to have direct implications.

For example, you can prompt Claude to ask follow-up questions or not. That might be what you want in 1 part of the product but not another. You might prompt Claude to think longer about a problem and do more reasoning, or not. These are all decisions that, upfront, you are making in product design, and they're going to manifest in the actual product.

The other piece is that, as a startup founder or somebody doing classic B2B SaaS, you need to triangulate where the models are, where they're going, and what the user needs are. That's going to be the case in your product design as well. You're doing the evaluations, hopefully upfront, to see if what you're doing is even possible with the current models, or at least having an eye out for where they might be.

7. The Competitive Landscape of AI

The models change over time, and products change over time. If you don't have a good framework around evaluation, including regression testing those evaluations, you might launch a product that, 3 months later, people say, “The product used to be good, but something has happened where it's no longer serving that purpose.” You're not sure which of 3 things changed: the model, the product design, or the introduction of a different feature. Maybe the system prompt got longer. In many ways, it's the most complex product-development work I'll ever do.

Harry Stebbings

I interviewed Sam Altman in London from OpenAI, and he said 1 of the joys they have as a startup is that they can release things much quicker. It doesn't have to be perfect. The challenge, as they've gotten bigger, is that more and more weight and pressure is placed on every release. How do you think about “release it, it doesn't have to be perfect, let's get it in the hands of users” versus now, when Anthropic is a massive company with millions of users?

Mike Krieger

I think about this a lot, especially because you have different surfaces and different audiences that have different expectations of stability or desires to be on the cutting edge.

In an API product, people value predictability and stability, with the option of something that's more future-facing. It can be a very opt-in thing. I remember we launched prompt caching, which is a big cost savings for people. Initially, we did that through a beta header that you had to opt into, and a lot of what we do on the API is in that form.

If you do that for our customer-facing or more consumer-oriented products, it's lame to have people opt in. You want to be able to iteratively release and be experimental with people. You don't want to totally break their experience, but you have a little bit more permission.

Then we have these enterprise customers that are using Claude for Work in an enterprise. AI adoption in the enterprise is still an early-adopter product, so you can get away with more than you could if you were Salesforce. I don't know how many releases Salesforce does a year, but I know a lot of these companies do 2 or 3, usually oriented around some big event. We're far from that. We're still launching pretty quickly, but we're honestly still finding the balance. Is it a monthly drop? Do you ship as often as you can, but have an admin opt-in? That adds complexity as well.

It's a great question, and it's an active topic of conversation: how rapidly we can ship, knowing that we want to bring things out to the world, we don't know how they'll be received, and we want to learn. But as you accumulate notoriety and people start depending on you for workflows, you can't treat that completely wildly.

Harry Stebbings

Are we in a product-marketing nightmare? DeepSeek released something this week, OpenAI released something this week, Anthropic released something this week, and Mistral released something 10 days ago. Almost every day there's a new release, and maybe the world gets apathetic. How do you think about that, and how does it inform the way you think about product launches and messaging?

Mike Krieger

It is much more like Crossy Road. The things you had to watch out for in the past—the big rocks—were very well known in advance. Don't launch anything during WWDC week, because there will be a flurry of announcements. Watch out for the September iOS event, or another big rock like the holidays. It was much easier from a product-marketing perspective.

Here, it's like, “Okay, the car's going by. There's a gap in the cars. Launch tomorrow—or now. Oh, but now we hear there's a rumor.” It's so much harder. I've heard from people at other labs as well that everybody's trying to read the tea leaves: “Is it quiet? Is it okay to launch now? I think we're going next Tuesday.”

It requires a completely different approach, and I give credit to our product-marketing team because they've had to orient from a point where we launched Claude 3.7 Sonnet on a Monday and locked the blog post the night before at 9:00 p.m., which is not best practice from a marketing perspective. We were briefing the press that Sunday. Thank you to the people who helped us on the phone that Sunday, but it was right. That's the point where everything is done, ready, and locked, and we can go.

It involves the ability to react quickly and be nimble. Even when we release a model, there's a model card, evaluations, and a comparison table. There are things in that comparison table that were released the week before. Grok 3, for example, was just released a week prior. It involves completely changing what happens when those are released.

Harry Stebbings

When Grok 3 releases, does everyone at Anthropic and OpenAI say, “Oh, shit, they beat us again,” or, “Oh, shit, we won?”

Mike Krieger

I think it requires—I try to support the team by reminding them that model releases are going to happen, and at any given point you're going to be in the “it's so over, we're so back” cycle. You have to live that in AI. You can't get too down about 1 release, because it is inevitable.

Sometimes you're lucky and there's a 2- or 3-month period where the model you launched, or the product you launched, is still state of the art across all the things you care about. Sometimes it lasts a week. You can't overrotate on either of those. You can't rest on your laurels, and you can't get too upset.

The thing that's really useful to me is a chart I show in almost every sales call, mapping Anthropic's founding to where we are today and showing the milestones. At any given point, you can say, “Claude 2—that's pretty far behind. Claude 3 is state of the art. Now it's not.” You have to look at the trajectory and trust that you're going to continue making improvements.

Then remind yourself that, if everybody switched every single day purely because an evaluation changed, that would be an insane thing to do to your user base as a software provider. It would also make for an even crazier industry over time. People don't just deploy models. They're doing fine-tunes, or they're deploying models plus a lot of bespoke work to make the model great for that use case. That's not something that's going to switch overnight.

Or you're 1 of 3 or 4 options within a model selector. In a coding environment, for example, you're still in the mix and you still have a chance. I'm not sure if it's finding the meditative zoom-out angle or just getting used to the bumps, but it is definitely the case that every time there's a model launch, I assume every 1 of those labs is watching the livestream and looking at the evaluations, saying, “All right, now we've got work to do.”

Harry Stebbings

I would argue that brand is the most important thing. To your point, people aren't switching every day. They're saying, “I'm a Claude person,” or, “I'm a ChatGPT person,” and they identify with their models. Do you agree with that, or is it too glib?

Mike Krieger

I think that's right, especially on the consumer front. I was just reading Ben Thompson, and he has Nat Friedman and Daniel Gross on there pretty often. They're talking about some people being Claude people and some being ChatGPT people. That definitely happens. You like the personality, the interface design, and the vibe.

It reminds me a lot of the back-and-forth we had with Snapchat over the years with Instagram. Before that, people would launch a new product that was like Instagram but just for high-end photographers, or with an additional twist, or just 1 photo a day. That was BeReal.

I had this fake formula—I'm clearly not the mathematician at Anthropic—that social networks are made of format, audience, and vibes. For Instagram, we had Stories and the feed. Eventually, we had video. The audience initially was sort of hipster photographers, and eventually grew to anybody interested in visual storytelling or visual media.

But the vibes of Instagram, even when we had product similarities to Snapchat or Facebook, were very different. I don't know what the fake formula is for AI products yet, but I think it's some version of that. Model personality is probably 1 component. There's likely something around the scaffolding and prescriptiveness of the product that you're working around it, and then there are vibes. Again, they're hard to measure, but they're absolutely there when we have so many different models and providers.

Harry Stebbings

Open source is a very viable possible route, and distillation is looked at in a shady way. Is distillation really wrong if it ultimately propels the space forward? Even within the labs, I assume every 1 of them is using it internally. It's very valuable to take the knowledge of your highest-end model and make it lower-latency, more affordable, and so on.

Mike Krieger

There are a couple of places where this gets interesting. 1 is whether we want any nation to be able to distill models from any other 1. My personal answer is no. I think there's value, as AI gains capabilities, in being thoughtful about that from a national-security perspective.

The other piece is that, for advancements to happen at the rate they're happening and be sustainable over the long term, the labs need to be able to commercialize all of that training and innovation. Finding the right models for that long term is important.

I think open-source models—take Llama, for example—have been able to do that from their own research, data ingestion, and training. I would say distillation does not feel essential to unlock those things, and it poses other issues, even from a terms-of-service perspective.

Harry Stebbings

Does Llama show that there is no value in the model and all the value is in the data? If Meta is willing to give it away for free because it knows nobody can copy the data it has, is that what it shows?

Mike Krieger

It's an interesting question: is the quality of Llama due to the fact that Meta can—I don't know if they've said that they do, but they clearly can—train on Instagram, Facebook, and other data? Was Gemini better because Google could train on YouTube?

It's actually clearer to me that Gemini benefits from that. Whenever they have a good video-understanding demo, for example, I'm like, “Well, somebody has probably got the largest repository of video in the world and can likely train on a lot of those pieces.” It's less clear on the Facebook front. I've never heard people say, “What Llama does extremely well is generate content that would work well on social media.” It just seems like a good general-purpose model.

It goes back to our earlier conversation: the value is in how good your team is, whether you have the underlying data that you need to train on, and how useful your model is in actual use cases. That is the highest-order bit.

I almost wish I'd started with that, because, evaluations aside, evaluations are really useful for hill-climbing and internal research, but they don't tell the story of whether a model is going to be excellent at what it needs to be excellent at, or even if it's excellent at that thing in very narrow situations. As an entrepreneur outside the labs, can you rely on the model to be your representative in that product?

I think the value for the labs is in the team. It's in the model's ability to perform the right actions in the real world without so much nondeterminism that it becomes unreliable.

8. Do We Underestimate China's AI Capabilities

Harry Stebbings

I'm going to ask 1 question on this. It's not a trap to go down, but I've spoken to Alex Wang and Poolside on the show, and they said we deeply underestimate China's ability in AI. Do you agree that we underestimate it?

Mike Krieger

I think the DeepSeek piece that people seemed surprised by was that there were cutting-edge research teams there. If you were paying attention, that part should not have been the surprising piece.

Instagram was blocked in China fairly early, and then we saw the emergence of a parallel world of startups. If you take Facebook and Instagram away, what happens and what emerges? Those products were often very high quality. They demonstrated a lot of creative thinking and were built at scale. They were solving problems at the same scale as the problems Facebook was solving.

People love talking about the super app, WeChat, and the technical challenges it solved at scale. It would absolutely be a mistake to underestimate, or continue to underestimate, China's ability to train at the frontier—especially if it gets access to compute—and to continue innovating there.

I think it's a pretty Western-centric view that I've definitely seen happen in more traditional software. It's caught in this 1990s or early-2000s view of, “All they're doing is replicating what's already been working elsewhere.” There have been products that took a differentiated view, grew within the Chinese market, and then sometimes made that view external. TikTok is an interesting example of that.

9. What Did Anthropic Learn from Deepseek

Harry Stebbings

Just 1 final 1 before we move into the verdict products: did DeepSeek cause you to rethink anything or change anything about the way that you progress?

Mike Krieger

There are some architectural pieces, and I won't speak for the research team because they're the DeepSeek experts. They might say, “That's interesting. That's worth us considering,” or point to ideas that had been considered and were worth reevaluating. I think there was a piece of that.

Our plan was already to show the chain of thought when we launched our reasoning model, so that was not a reconsideration. But it was interesting to see somebody else do that, and there are some user-interface details in there. I think Grok does that as well now, so it'll be curious to see how that evolves. To your distillation question, that might be a reason why more labs choose not to show, or otherwise obscure, the chain of thought down the line.

10. Is Deepseek a Sustaining and Credible Threat?

From a product perspective, there were 2 things. I think the under-talked-about piece of DeepSeek is that they were able to go from nobody knowing about them to being, frankly, in many circles, better known than Claude. Likely my great-aunt was calling me about DeepSeek. I'm not even joking. It was cliché; it was actually happening. People were asking me, “What do you think they did to break through that maybe Claude hadn't?”

There was a lot of interest in world politics at the time, and the narrative was, “This was much cheaper,” whether that was exactly true or not. It was the story. I've had this conversation with our marketing team as well: I don't think we tell the Claude story well enough externally yet. We should talk more about what is different or notable about the fact that, by Claude 3, we were training a model at the frontier that was state of the art with a team that was much, much smaller than any other lab. We've always been very efficient with our compute as we train.

Whether that was a story DeepSeek told or was told for them by the media, it was legitimately a compelling story. The uniqueness of the moment was a big piece, and January, the new presidency, and China relations fed into the moment very well.

The second part was the product. They went from not having a product to having an iOS app that had a lot of good details. For me, it was a good nudge—but stronger than a nudge, more like a shove—that we need to get some ideas out to market more quickly, without focusing as much on exactly how polished they need to be in every situation. We need to be willing to put them out there and learn. Sometimes the novelty of an experience is itself valuable. It was the first time most people experienced a live chain of thought, and I wish we had done that sooner because it would have been novel for people to experience.

Harry Stebbings

You look at usage, and you see emerging-markets usage retention, while you don't really see that in Western markets. How do you think about DeepSeek as a sustained, credible threat?

Mike Krieger

They already have a level of awareness that gives them some ability to generate ongoing staying power. But when I think about all we're doing in these AI-first, lab-generated products, even 6 months or 1 year from now, I think that asking questions and having slight proactivity is not differentiated or interesting in the long run.

It should be, “Wow, I can now do something uniquely because I'm using Claude, DeepSeek, or any 1 of these products. It unlocked hours of work for me, made me smarter, and made me a better partner to the important people in my life.” It has to transcend the surface-level utility. Some people find the deeper level—don't get me wrong, those are your DAUs right now—but for a lot of people, they'll try it, generate a poem, or write a letter to their son. There's all this stuff they can do that provides some value in the moment.

I still think we're at day 1 around whether AI is an indispensable part of most people's work, and I think the answer is no for most of them. DeepSeek's staying power, and the staying power of all our products, will come from who can get there and do that sustainably over time, with the right product design, integrations, and deployment to actually succeed.

11. Transitioning from Model Provider to Application Provider

Harry Stebbings

Who can build those products is my big question as an investor. When does a model provider move into being an application provider? I'm fascinated to hear your thoughts on what is attractive enough for you to dedicate the resources to become an application provider, not just a model provider.

Mike Krieger

There are 2 main criteria that I look at. Our team is large, but our product team is maybe 1/10 of that. It's very large by Instagram year-2 standards, but small by large-SaaS-company standards. We're somewhere in between all of those different things, and we're supporting a lot of surfaces: Claude Code, the API, Claude, and Claude for Work.

Generalizability is really important. Even if we pick a persona or vertical to go after, we're going to build things that are general-purpose as a rule, with maybe some specialization at the user level but not at the level of a very bespoke workflow or use case.

Translation, transcription, and customer service—fairly horizontal, homogeneous things—seem like they would be in the pathway.

Harry Stebbings

I think they would be, except that I think there's a lot of valuable workflow and workflow knowledge that means you can retain a differentiated product over time.

Mike Krieger

If you're a power user, yes, perhaps.

Harry Stebbings

But if you're not a translator—if you're your mom, who uses it once a month for that odd thing she needs—then there's not really much there.

Mike Krieger

I think the role of, “We can help you translate this, and we'll get you to pay a $10 monthly subscription,” feels questionable because the models are already quite good at that. Maybe you're right that there's not much differentiation there.

If you play with ElevenLabs' Console and Workbench, a lot of the features they've built are clearly for people translating hours or voicing hours of content with a reliable voice across the whole workstream. Descript has some of the best product design in AI. They've clearly put so much time into the workflow.

I had to use it once for a personal podcast, and I thought, “This was clearly built by people sitting in this workflow day in and day out and understanding it.” Maybe we come to some synthesis of our views: there's value in the more professional use cases and the workflows unlocked by them. On the consumer, and maybe even prosumer, side, the basic AI product gets good enough.

Harry Stebbings

When you look at what you're brilliant at today, you do so well on the coding front. Is there a roadmap here to put your own IDE or code agent in? How do you think about that?

Mike Krieger

Again, with the product-focus lens, I think we have to pick our bets carefully. We built Claude Code, which we just released, as a command-line agentic coding tool internally first because we wanted to accelerate our own team. After seeing it play out for a couple of months, we thought, “This is good. It's not a solution to all coding problems, and it doesn't obviate the IDE, but it's useful enough to us in enough cases that we want to see people use it in the real world.”

Shipping is never free. You have to name it externally, find the right packaging, and handle the go-to-market piece, so we do it carefully.

My view of where the models are today is that you still need hands-on-keyboard work and an exchange of, “I did this. Is this right?” You need to say, “Let's pursue this direction,” and then, “Yes, this is great; let's put up a pull request.” Or, “No, we went down a false trail. Let's metaphorically unwind the stack and then keep going.”

That's why I think there's a role for something in between the IDE and full-on delegation of tasks, likely in the style of Cognition's Devin. Claude Code can be used for a certain category of tasks.

Our product engineers love Claude Code because a lot of product engineering is, “We have to update the back end, create the front end, submit these things for translation, and figure out why this still doesn't work.” It's that build-the-product-end-to-end workflow that does well with something that can work agentically across a lot of different things.

I did 2 pull requests last week. I hadn't coded since joining Anthropic, which made me sad, so I finally got to use Claude Code. I had never opened our codebase before, so I didn't really know how it was structured. Claude Code was very good at finding the file with the right piece and then making edits.

Obviously, not everybody is in the same situation I'm in, but it is really valuable for those use cases. When I think about the coding space and where we can play and add value, it's really on the agentic side, not the IDE side.

There are other companies that wake up and go to bed every night thinking about how to make a great IDE. That involves low-latency autocomplete, the right integrations, figuring out how to work with the VS Code plug-in ecosystem, and all that complexity. There's a lot of valuable work there that's different from what we're doing. We can really play in talking to these models and doing real work with them in that agentic loop, while recognizing that they're not yet at the place where, for many use cases, you can let them run free for hours. You need more of a human-evaluation piece.

12. What is the Role of a Software Developer in the Future

Harry Stebbings

You power and work with Cursor, Codeium, and StackBlitz. My question to you is: when you look at the changes we're seeing in developer behavior, what will the role of a software developer be in 3 to 5 years?

Mike Krieger

It already starts to look different. I was a huge early proponent of GitHub Copilot. I think my quote was on the homepage for a while—I don't know if it still is—because I saw the potential.

Then GPT-4 came out before it had multimodality, and I was trying to use it with Swift. I would draw ASCII art of the screens I was trying to build for Artifact, then go make coffee because it was quite slow at the time, and come back to an 80% version. Obviously, now it would be a 95% to 99% version with something like Claude 3.7 Sonnet.

I think the skills that become important are, 1, being multidisciplinary. It's knowing what to build as much as knowing the exact implementation you want. I love that about our engineers. Many, maybe most, of our good product ideas come from our engineers and from them prototyping. I think that's what the role ends up looking like for a lot of them.

The second piece is that code review changes when you're mostly evaluating AI-generated code. I experienced this myself. I put up a pull request, and some comments came back saying, “Claude Code does this sometimes. We don't actually use default arguments in this case.” I thought, “Oh, damn it.” If I were coding it, I probably would have noticed those patterns better.

There are 2 sides that need to happen. Models and the infrastructure around them need to learn from codebases and code reviews better so they can produce code that feels idiomatic to that company. But we also need to evolve from being mostly code writers to mostly delegators to the models and code reviewers.

That's what I think the work looks like 3 years from now: coming up with the right ideas, doing the right user-interaction design, figuring out how to delegate work correctly, and figuring out how to review things at scale. That's probably some combination of a comeback of static analysis and AI-driven analysis tools that determine what was actually produced. Is there a security vulnerability? Is there another flaw? Is there a bug?

Computer use plays a part. You can tell I get very excited about this space. Automated testing of the UI is important too. What would be great is that you delegate a task—let's say a year from now; 3 years is crazy—and, when you come back, it says, “I evaluated these 3 approaches. I tested them all out. I had a different agent try them in a browser. This 1 worked best. I've run it through another agent that performed a vulnerability test, and it all looks good. All we need to do is help you resolve this 1 question. Let's review this critical section of code to make sure it's what you really wanted.”

That feels like you're suddenly empowered to be more of a manager and delegator to these systems rather than just a partner in the loop.

Harry Stebbings

You said 3 years sounds ridiculous and that 1 year would be much more realistic. I agree. When we look at the speed of scaling, do we think we hit a plateau or an asymptote in product releases and the speed of development? It feels so fast now. Do we hit that plateau, or do we continue in this exponential progression?

Mike Krieger

It's a question I think about a lot. I started the year by looking at our product-development process and where we are Claude-enabled and where we're not.

Claude can be useful in taking an initial idea and creating a PRD from it. Claude can be useful in the coding side. Claude can synthesize a lot of conversations people are having about a product, find the thorny issues of disagreement, drive alignment, and figure out what to build.

Actually figuring out what to build is still the hardest part. That's best resolved by getting together in a room and talking through the pros and cons, or going off and exploring it in Figma and coming back.

Like any dynamic system, if you optimize 1 piece, something else becomes the source of the blockage or the critical path. Alignment, deciding what to build, solving real user problems, and figuring out a cohesive product strategy are still very hard. The models are probably more than a year away from solving that.

That is the constraint. It's why I'm bullish on startups being able to explore the space. I remember this from both my Instagram and Artifact days: there's a difference between alignment being a coffee conversation in an afternoon and steering the ship of a large company that has commitments to customers and all of those things.

That's still a very human problem, and I think we're at least 3 years away from the models solving it at that level of abstraction.

Harry Stebbings

We mentioned consumer products and building them. When you think about building new products for consumers versus building the API division of the company, which is very significant, how do you think about the balance and the trade-offs between building an API business and building an end-user consumer business?

Mike Krieger

I think about what we get out of each. We learn a lot more quickly with first-party products. A specific example is Claude Code. Within a week of deploying it internally, we found that 1 of the tools it had access to was not being used as well as it could have been. That made its way directly into Claude 3.7 Sonnet.

That's a way in which internal dogfooding of a first-party tool directly led to a model improvement in the next generation. There are a few other places where we've seen that. It's much harder with a third-party product. They might tell you something's wrong, but it's more arms-length. Even though we work very closely with some of the coding startups you mentioned, it's still not the same.

There's a lot of value in what we learn there. Then there's the stickiness and loyalty we talked about. I think it's easier, from a consumer perspective, to build a brand around a product than around just an API.

The fact that we power a lot of these coding products is visible to people. We're often the default in the dropdown selector, and if you're in the know, you know. But not everybody does, and it's still not the thing they downloaded or installed that they're going to tell their friends about.

On the other hand, it's a place where we've gotten tremendous distribution. We're not going to invent every company, and we're not building every product ourselves. In that way, it reminds me of my investing days: you get to see a lot more, and there's more than 1 shot on goal. It's been a fairly even split from a resource-allocation perspective.

If anything, we've underinvested a bit in 2 things. 1 is having a faster iteration speed on first-party products; that's my current obsession. The second is, on the API side, figuring out how to build abstractions beyond tokens in, tokens out.

Every time we do that, we get great feedback from people. Whether it's helping the model plan and work agentically, having the model build more knowledge graphs and repositories of how companies operate internally, perfecting tool use, understanding very large amounts of context, or having memory that transcends conversations, those are problems worth solving on the API.

They're things where we can take what we learned on the training side, map it directly to the API, and build good products around it. That's how I think about the 2. At Instagram, it was easy: 95% product and 5% API. That's all we needed to do.

Harry Stebbings

What can and will you do to increase product speed on the first-party consumer side?

Mike Krieger

There are 2 things. 1 is recognizing that we were running a larger-company playbook for what are still startup products. Even if the company has good traction, the API business is doing well, and people are using Claude and upgrading to Claude Pro, it's still early days. It's still do or die, or make it or break it.

We need to operate that way. That means getting the right people together sooner and faster and ignoring organizational boundaries. We got too calcified: “This is on this team's plate,” or, “You can't get this done this quarter because it's not on this team.” I understand why organizations evolve, and some of that is natural, but we can't afford it right now.

It's been much more about, “Who are the right people? Let's get them together. Let's clear away all the other distractions.” I need to clear out my calendar so I spend more time in product reviews and design reviews than in administration.

13. Is Europe Stronger or Weaker in a World of AI

Harry Stebbings

Did DeepSeek show the benefits of constraints? Do Western companies—respectfully, you and OpenAI—have too much money?

Mike Krieger

The way I would put it is that the adoption we've gotten for our products is ahead of their actual product-market fit because they are still the best ways of getting the models. I don't think that's durable over time, so it's not something to rest on.

14. Quick-Fire Round

I also think we're underserving people because we haven't gotten the right products yet. That's what I wake up stressed about every morning—or inspired by, depending on the day. We've got so much work to do on that side.

Harry Stebbings

I love it. I want to do a quick-fire round. I'll say a short statement, and you give me your immediate thoughts. Does that sound okay?

Mike Krieger

That sounds great.

Harry Stebbings

What's OpenAI done better than you?

Mike Krieger

They've moved faster at shipping V1s, sometimes even ahead of where the model is.

Harry Stebbings

What have they done worse than you?

Mike Krieger

Probably personality, and having the features they build be cohesive.

Harry Stebbings

Which alternate model provider do you most respect?

Mike Krieger

OpenAI. I think they've balanced first-party product development and an API that people use at scale. We had an Instagram principle that was, “Do the simple thing first,” and I think they often do the simple thing first.

Harry Stebbings

If you could rebuild the Anthropic product and stack from scratch, what would you do differently?

Mike Krieger

I love this question.

Harry Stebbings

I do too. It's a good 1, isn't it?

Mike Krieger

It's a really good 1. The things we built that were very valuable last year now feel like they have some cost to the information architecture. That sounds like a very nerdy way of describing it, but people should not have to think about projects versus artifacts versus chats and how they all relate.

Tearing it all down and asking what actually matters: do you have the right context in the right conversations? Do you feel like you can always know where to go next in the product? Is Anthropic and Claude itself being a helpful guide to what work is most important to do next?

That's a different paradigm from, “I know to create a project.” If you get good at that, it's an amazing product, but there are a lot of steps along the way. That's the fundamental thing on the product side.

On the stack, Claude.ai—and probably ChatGPT.com—were initially built very much as showcases for the models. They weren't built to be the foundation for a much more complex, multiproduct system. We have an active effort right now around tearing down some of that and rebuilding the core UX so that it feels good.

It doesn't feel great right now. It feels like it's been an evolution of a product that served a purpose at the time but is now being asked to do many more things. The incremental approach is now both harder to add to and getting slower.

Harry Stebbings

What have you changed your mind on in the last 12 months?

Mike Krieger

How important first-party products are. I saw the growth in the API, and I thought, “This is what we should invest much more of our time in.” But you'll miss out, and you won't have enough of a durable moat, if you're not equally investing—maybe even investing more—on the first-party side.

Harry Stebbings

How much did it hurt you being late to that?

Mike Krieger

Significantly. Take the DeepSeek moment. Ideally, the story that there is more than 1 leading-edge AI product to use is a narrative we should have captured. I think it hurt us there.

Harry Stebbings

What's a major technical or product challenge on the horizon in AI that no one is talking about but that you think is critical?

Mike Krieger

As the models get more capable, the headline is discernment and privacy. They'll also become more knowledgeable. They'll be in conversations with you about everything from something intimate to something sensitive from a company perspective. They'll have access to all of your company's information.

Everybody loves to talk about agent-to-agent interaction. The intersection of those 2 things is not discussed enough. Do you trust your Mike agent or your Harry agent to be out in the world, not be jailbreakable, and not reveal something it knows that is personal or sensitive?

My metaphor is my 5-year-old daughter. It's great watching her with somebody she's just met because she doesn't quite differentiate between things that are secret and private to our family and things that are okay to talk about with a new friend or somebody at the checkout aisle. Discernment is something people acquire over time.

I think this is underappreciated and probably underresearched from a model-capabilities perspective. Models fundamentally want to be helpful, and that is not always what you want them to be. There's a safety case for that, but there's also a privacy and data-security case.

Harry Stebbings

Do you worry about your 5-year-old becoming more comfortable talking to models and agents than she is with humans?

Mike Krieger

I've had so many conversations with Alex Wang about this because he has this whole idea that, in the future, most friends will be AI friends. I don't think he's wrong. There are ways in which that's already starting to be the case, with people having online gaming experiences where some of the characters are non-player characters. You might have a more comfortable existence there, even if you're not breaking through.

I do worry. She's so gregarious that I'm not worried in her particular case, but let's abstract to the broader sense. There's a lot you can learn from what it feels like. I was a fairly awkward teenager, and I probably could have benefited from a practice mode for AI interactions around some of these things.

At the same time, that doesn't feel like it's totally closing the loop around the consequences of real interaction. It's the difference between reading about what it's like to have your first really hard argument with your high-school girlfriend and actually having it. When you're in that moment, it's different.

It's like the thought experiment where somebody is in a black-and-white room, only reading about the color red, and then goes out into the world and sees red. Is there something qualitatively different about that? Absolutely.

Is there something different between talking to a model and engaging with a model, even in emotional role-play, and having that same interaction with a real human? Absolutely. It is probably a helpful piece of future human interaction, and absolutely insufficient as a replacement for it.

Harry Stebbings

Does Europe become more or less relevant in an AI-driven decade?

Mike Krieger

I want Europe to do well because I love a lot of Europe. I lived in Portugal growing up as well. I saw a funny, perhaps somewhat defective version of this argument: if real-world experiences and human interaction become more valued, Europe becomes more valuable as the world's capital of sensory experiences.

That feels weird if that's all you're resting on. It feels limited. What I think will be really interesting from a European perspective is the things that Europe holds very strongly about lifestyle and society, and then attempts—sometimes not elegantly—to enshrine in best practices or even laws.

When we think about product design, data privacy, and selling to German users or German companies, there's a different set of questions that gets asked. They're often very helpful questions. Maybe the bull case is that those questions are relevant to everybody, and Europe will be at the leading edge of asking them.

From a lab perspective, it's a harder question to answer. There may be some combination of access to compute and moving further up the value chain. If building applications on top of these models becomes much easier, and you can go from 0 to 1 and be more nimble than labs that have tens or hundreds of millions of users and have to move more slowly at that scale, can innovation happen there? Probably.

But it probably involves a different regulatory and startup-ecosystem environment to make that the case.

Harry Stebbings

Dario Amodei has said that this could be the generation that lives to 150. I'm slightly butchering and summarizing his quote, obviously, but this could be the generation that finds cures for diseases like multiple sclerosis with AI. My mother has multiple sclerosis. Do you agree with his optimism, and how do you think about AI increasing longevity and human lifespan?

Mike Krieger

I think the potential is huge. Today, AI is helping close the loop on drug discovery and clinical trials. A company likely called Huma used to take, I think, 15 weeks to do its clinical-trial reports, and now it uses Claude and gets them done in 20 minutes.

That's a step change. There were years of research that preceded it, so I'm not saying we've cut years to weeks or years to minutes. But it's a point in the process that we can make faster with the models today.

Then you see Arc Institute, the science and research institute that Patrick Collison and others have started and funded. They're working on foundational models for cells, where you suddenly have a real cell model that you can run experiments on. That kind of thing should accelerate drug discovery and experimentation tremendously because you're cutting a loop there.

I'm very optimistic. There are a lot of places where AI is underutilized relative to its potential. Some of the smartest people in my generation were working on serving more targeted ads. Maybe that was true at 1 point, but a lot of them today are working on how to make models tremendously useful, valuable, and intelligent across many domains.

Harry Stebbings

Mike, you've been fantastic. Thank you so much for letting me completely unpack all of my questions on you without warning. You've been amazing.

Mike Krieger

My pleasure. This was really fun.