[BidClub_]
20VC · · 66 min

Mike Krieger, Instagram CoFounder & Anthropic CPO: Where Will Value Be Created in an AI World?|E1265

Harry StebbingsMike Krieger

YouTube
TL;DR
  • Krieger's core answer to "where will value be created": companies with differentiated go-to-market, differentiated domain knowledge, or proprietary data — "ideally two or even three of those" — in sectors like finance, legal, and healthcare, where the unsexy upfront legwork is exactly what makes the position durable. The winning loop is to sell into places you uniquely understand "and then get better for being deployed there over time."
  • His contrarian call on commoditization: "I think models over time get more different rather than more similar" — "there is something claudy about Claude and there is something GPT about GPT." Model-layer moats are three: talent density, model character deepened by traction (coding traction feeds the next generation of RL), and being "an AI partner, not just AI models." Fail on any one and "I think you're in trouble."
  • DeepSeek had "almost no impact" on Anthropic's go-to-market — enterprise relationships aren't swapping input tokens for output tokens — but it was a marketing and shipping wake-up: it went from unknown to "in many circles better known than Claude" ("likely my great-aunt was calling me about DeepSeek"), and being late to first-party product hurt "significantly."
  • For startups riding model improvements: don't wait. "My startup was not a startup until Claude 3.5 Sonnet" is a pattern he hears from multiple people; the winners of each model-generation shift are the ones already beating against the wall — Cursor iterated repeatedly before breaking through — "when the model arrives you're not starting from square zero."
  • The biggest blocker to progress isn't compute, data, or algorithms — it's training environments and evals that match real multi-step work. SWE-bench undersells what a software engineer does; nobody evaluates office professionals well; the missing eval is "I show up to a new job, quickly understand my role, who is who" — the gap between models "extremely good at extreme slices" and generally helpful collaborators.
  • Software engineers become delegators and code reviewers "a year from now," not three — agents that try three approaches in a browser and run vulnerability tests before asking one question. But "figuring out what to build is still the hardest part," at least three years from being solved — which is why he's "really bullish" on startups, where alignment is "a coffee conversation."
  • The most sobering line for anyone underwriting AI app adoption: Anthropic's own traction "is ahead of their actual true product-market fit because they are still the best ways of getting the models — I don't think that's durable over time." And on usage broadly: "we are still in day one around is AI an indispensable part of most people's work — and I think the answer is no."
  • Underappreciated risk at the coming agent-to-agent intersection: discernment plus privacy — his 5-year-old metaphor, a child who can't yet distinguish family secrets from checkout-aisle chat. "Models fundamentally want to be helpful and that is not always what you want them to be."
Digest · the substance, structured for research

1. Value accrues to differentiated GTM × domain knowledge × proprietary data

  • Krieger fields "what can I build that won't be in the lane of an Anthropic" from entrepreneurs often. His answer: places with differentiated go-to-market, differentiated industry knowledge, or data only you have — "ideally two or even three of those" — in finance, legal, healthcare. Healthcare is "a tremendously complex ball of yarn" whose upfront legwork can't be done in an accelerator, "and it is the legwork that you've put in" that makes it durable.
  • The durable loop: pull in what's great from foundation models, fine-tune if needed, but win by selling into places you uniquely understand "and then get better for being deployed there over time."
  • Incumbents vs. net-new startups: both can win, with inverted risks. Startups get license to overpromise — early adopters are "kicking your tires" — while an incumbent that announces "we've added AI" and delivers "you said it could do these 30 things, it does like two of them well" breaks trust.
  • Startups' problem is the mirror image: no data, no relationships yet. Their differentiation "is not the established relationships, it's painting the future" and landing lighthouse customers willing to take that bet.

2. Don't wait for the perfect model — be the one beating against the wall

  • Krieger hears a recurring line from founders: "my startup was not a startup until Claude 3.5 Sonnet" (or the second 3.5 Sonnet) — a generational leap taking accuracy "from 95 to 99, or from 70 to 90," suddenly clearing an industry's bar.
  • His best specimen is Cursor: someone showed him the founders' Hacker News front-page submissions over time — it finally broke through, "but that was not their first product or their first iteration." The companies that benefit from model shifts aren't the ones that start that day; they're the ones with accumulated context about what goes wrong in the space, so "when the model arrives you're not starting from square zero."
  • The succinct version: "be frustrated by the current generation of the models and then be very aggressively trying the next one so that you can finally deliver on the thing that you saw in your head."

3. Three moats at the model layer — miss one and "you're in trouble"

  • Harry's challenge: with releases coming this thick, is there value in the model layer at all? Krieger's three: first, talent — "talent begets talent" around a cohesive mission; Anthropic's research team lands "some new significant hire" monthly, but people are free agents, so the attractor must be maintained.
  • Second, divergence: "I think models over time get more different rather than more similar — there is something claudy about Claude and there is something GPT about GPT." Coding wasn't an accident, and it compounds: seeing companies rely on Claude for code "inspires the next generation of what you want to do from a reinforcement learning perspective."
  • Third, partnership. DeepSeek's go-to-market impact was "almost no impact" because enterprise relationships aren't "they send it for the API, they want to just exchange their input tokens for output tokens at some rate" — it's "I want to be your long-term AI partner, I want to help co-design products with your applied AI team."
  • The failure mode, inverted: resting on laurels, believing incremental benchmark gains are enough, "treating the API as just a way of exchanging money for intelligence."

4. The biggest blocker isn't compute — it's environments and evals that match real work

  • Where Alex Wang and Groq's Jonathan Ross give Harry different answers, Krieger's is neither compute nor data: it's getting training environments to match real-world, non-single-shot challenges. SWE-bench undersells the job — a software engineer understands requirements, negotiates timelines with PMs, ships and iterates: "there's no eval for that."
  • Office professionals — a use case Anthropic thinks about heavily — "nobody's really evaluating that well." The missing environment: "I show up to a new job, I quickly understand what my role is, who is who in the organization... and then be in the run loop of the business." That's the blocker between models "extremely good at extreme slices" and generally helpful collaborators.
  • On synthetic vs. human data: it "absolutely has to be a mix" — seed with original human data, then generate synthetic environments to explore. Claude playing Pokémon is his example of many runs through one game; it gets much harder "when the problem space is less well-defined than did you make it out of Viridian Forest."
  • The underappreciated data problem is vibes: character has no regression testing. Going from Claude 3.5 to 3.7, "people will say oh, Claude seems friendlier but more terse... I wish it was better at creative writing — these things are not easily evaluable."

5. Today's AI products are "extraordinarily leaky abstractions"

  • Harry's bet: in three to five years you won't select models any more than you select "which Google you use." Krieger agrees — current AI product design is "an extraordinarily leaky abstraction": "why should you choose Opus, Haiku or Sonnet? Most people don't understand the difference... we suffer from this problem as well."
  • Memory is the second leak — his coworker analogy: you might have different email threads, "but it's still one coworker behind all of that" who doesn't forget your favorite sports team between conversations. Third is prompting, which should become "absolutely transparent"; the good-vs-bad prompter gap "closes generation to generation, but we need to collapse it even further."
  • Model quality and UX "can't be separated anymore": you're designing "a scaffold and a product around a fundamentally non-deterministic system," where whether Claude asks follow-up questions or reasons longer is a product decision. Without regression-tested evals, a product can degrade and you can't tell "is it the model, is it the product design... the system prompt got longer" — "in many ways the most complex product development work I'll ever do."
  • Ship cadence now varies by surface: the API demands predictability (prompt caching launched behind an opt-in beta header), consumer tolerates experimentation, and enterprise AI "is still an early adopter product," so Anthropic ships far faster than Salesforce's two-to-three releases a year — but "it's an active topic of conversation."

6. Product marketing as Crossy Road — and why nobody switches on evals

  • At Instagram the big rocks were known (don't launch WWDC week); now "it reminds me a little bit of Crossy Road — the car's going by, all right, there's a gap." Claude 3.7 Sonnet launched Monday; the blog post was locked Sunday 9 p.m. and press briefed that Sunday. The comparison table included Grok 3, released just a week prior.
  • The psychology he coaches: "it's so over / we're so back — that is, you have to live that in AI." Sometimes state-of-the-art lasts two or three months, sometimes a week; he shows sales the trajectory chart from Anthropic's founding — "trust that you are going to continue to make improvements."
  • Why churn runs cooler than leaderboards suggest: customers do fine-tunes and bespoke work on a model, and you're "one of three or four options within a model selector" — switching daily on evals "would be an insane thing to do to your user base."
  • Brand is real: he endorses Harry's "I'm a Claude person / I'm a ChatGPT person" framing, citing Ben Thompson's discussions with Nat Friedman and Daniel Gross. His Instagram-era "fake formula" — format + audience + vibes — has an AI analogue: model personality + scaffolding prescriptiveness + vibes.

7. DeepSeek changed the playbook

  • On whether the West underestimates China: "the DeepSeek piece — people seemed surprised that there were cutting-edge research teams there, and if you were paying attention, that part should not have been the surprising piece." He watched a parallel startup world emerge after Instagram was blocked; WeChat solved scale problems "of the same scale of challenges that Facebook was doing." Dismissing China as replication is "a pretty Western-centric view" — they can train at the frontier, "especially if they get access to compute."
  • What DeepSeek did that Claude hadn't: broke through on narrative. The cheaper-training story — "whether that was exactly true or not" — landed into January, a new presidency, and China relations: "likely my great-aunt was calling me about DeepSeek. I'm not even joking." His self-criticism: "I don't think we tell the Claude story well enough" — Claude 3 was state-of-the-art trained by "a team that was much, much, much smaller than any other lab."
  • On product it was "stronger than a nudge — a shove": ship ideas to market quicker, because "sometimes the novelty of experience is itself valuable — it was the first time most people experienced the live chain of thought... I wish we had done that sooner." (Anthropic had already planned to show CoT; the distillation risk may push labs to obscure it later.)
  • On DeepSeek's staying power, Harry notes that emerging-market usage retains while Western usage doesn't; Krieger is skeptical but humble, and for everyone, "I still think we are in day one around is AI an indispensable part of most people's work — and I think the answer is no."

8. When a model provider becomes an application provider: generalizability, not verticals

  • Anthropic's product team is roughly a tenth of the company yet supports Claude Code, the API, Claude, and Claude for Work — so the filter is generalizability: "I don't anticipate us building a lot of verticalized experiences that are fairly bespoke to a given workflow."
  • Harry probes horizontal categories — translation, transcription, customer service. Krieger's counter: workflow knowledge defends the power user — ElevenLabs' console is "very clearly for people translating hours of content," and Descript is "some of the best product design in AI... clearly built by people who are day in, day out sitting in this workflow." Their synthesis: professional workflows hold value; on consumer, basic AI "gets good enough" — a $10 monthly translation subscription "feels iffy."
  • Claude Code embodies the strategy: built internally first "because we just wanted to accelerate our own team," shipped after months of dogfooding, and deliberately not an IDE — other companies "wake up and go to bed every night thinking about how do we make a great IDE," low-latency autocomplete and the VS Code plugin ecosystem. Anthropic's lane is the agentic loop between the IDE and likely Cognition's Devin-style full delegation, because models today "still need hands on keyboard."
  • Proof point from his own hands: two pull requests last week, his first code since joining Anthropic, in a codebase he'd never opened — "Claude Code is very good at finding the file that has the right piece."

9. Engineers become delegators within a year — but deciding what to build stays human

  • The engineer's role is already shifting to knowing what to build: "many, maybe even most of our good product ideas come from our engineers... prototyping." Code review changes too — his own PR drew comments of "yeah, Claude Code does this sometimes, we don't actually use default arguments in this case," so models must learn idiomatic patterns from codebases and reviews.
  • The end state: "from mostly code writers to mostly delegators to the models and code reviewers" — with a static-analysis comeback, AI-driven vulnerability checks, and computer-use agents testing UIs. His scenario: you return to an agent that tried three approaches in a browser, vulnerability-tested the winner, and asks you to review one critical section — "empowered to be more of a manager and delegator." Harry: "three years sounds ridiculous, a year would be much more realistic." Krieger: "I agree."
  • The bottleneck that survives: he started the year auditing where Anthropic's own process is "cloudified" — Claude drafts PRDs, codes, synthesizes disagreements — but "driving alignment and actually figuring out what to build is still the hardest part," best resolved in a room or in Figma, and "probably more than a year away" for models — he later says at least three. That's why he's "really bullish" on startups: "alignment is a coffee conversation in an afternoon rather than steering the ship of a large company."

10. First-party strategy faces hard constraints

  • First-party products teach fastest: within a week of internal Claude Code deployment they found a tool the model under-used — the fix went "directly into 3.7 Sonnet." He's changed his mind on this in the past 12 months ("how much first-party stuff is important"), admits being late hurt "significantly," and names two underinvestments: first-party iteration speed ("my current obsession") and API abstractions "beyond tokens in, tokens out" — agentic planning, knowledge repositories, tool use, memory transcending conversations. Instagram was "95% product, 5% API"; Anthropic is roughly an even split.
  • Harry's sharpest jab — do you and OpenAI have too much money? — draws the episode's most honest concession: "the adoption that we've gotten of our products is ahead of their actual true product-market fit because they are still the best ways of getting the models, and I don't think that's durable over time... we're under-serving people." The fix: stop running "a larger company playbook," ignore calcified org boundaries, spend calendar on product review not administration.
  • Quickfire: OpenAI has been better at "shipping v1s faster, even ahead of where the model is sometimes," worse at "personality and having the features they build be cohesive." Rebuilding from scratch, he'd tear down the projects-vs-artifacts-vs-chats information architecture — Claude.ai—and probably ChatGPT.com—were "initially just built to be showcases of the models."
  • The critical unspoken challenge: discernment × privacy at the agent-to-agent intersection. His metaphor is his 5-year-old, who can't yet distinguish family secrets from checkout-aisle chat: "do you trust your Mike agent or your Harry agent to be out in the world and not be jailbreakable?... Models fundamentally want to be helpful and that is not always what you want them to be." On AI friends, he won't dismiss Alex Wang ("I don't think he's wrong") but insists AI practice is "absolutely insufficient" versus real interaction — the person who only read about red in a black-and-white room, versus seeing it. And on Dario's live-to-150 optimism: likely Huma cut clinical trial reports from ~15 weeks to 20 minutes with Claude, and Arc Institute's cell foundation models could cut the discovery loop itself — "the smartest minds of my generation were working on serving more targeted ads; a lot of them today are working on models."

1. Why Will Models Become More Different Than More Similar

Mike Krieger

I think models over time get more different rather than more similar. I still think we're in, like, day 1 around whether AI is an indispensable part of most people's work, and I think the answer is no. I think the DeepSeek piece people seemed surprised that there were cutting-edge research teams there, and if you were paying attention, that part should not have been the surprising piece. I think we've, if anything, underinvested a bit in 2 things: 1 is just having a faster iteration speed on first-party products, and then, on the API side, being ready to go.

2. Where Will Value Be Created and Sustained in a World of AI?

Harry Stebbings

Mike, dude, I am so excited for this. I've literally just been out for a walk, and I've been listening to every show that you've done in the last year. I told you before, I don't want to start with, “How did you get into tech?” and all the normal rubbish. I want to start with a very challenging first question, which is: as a VC investor today, I have to determine where value is in the future. I look at the world today, and I don't know. When we look forward, where will value be generated in the AI-driven decade that we have ahead of us?

Mike Krieger

I think it's an awesome question. I get a version of this question often from entrepreneurs. I went from purely building startups myself to now running a company that is partly enabling new startups to get created or helping boost their fortunes. The question I often get is, “What can I build that is not going to be in the lane of Anthropic or another one of these labs?”

3. Are Foundation Models Commoditised Today?

I don't have a perfect answer because it's a crystal ball, but my sense of where it ends up being most valuable to exist is in places where you have some differentiated go-to-market, some differentiated knowledge of a particular industry, or some special data that only you have access to—ideally, 2 or even 3 of those.

Companies that are within a financial sector, a legal sector, or healthcare—healthcare, I've gotten exposed to, and it is a tremendously complex ball of yarn. The work upfront is not the sexy work. It's actually not the work that you're going to be able to do in an accelerator or in a short amount of time, but it is the legwork that you've put in. I think those are durable places to generate value.

Then you can sit in a place where you can pull in what's great from the foundation models. You can do your own fine-tuning if you need it, and you can do your own AI automation if needed. The thing that's going to give you legs and be durable over the long run is being able to sell into those places, have something that you understand about those places uniquely, and then get better from being deployed there over time.

Harry Stebbings

When you talk about the legwork, and you mentioned differentiated go-to-market and differentiated data pools or data sources, does this next-generation wave of AI benefit existing vertical SaaS companies who already have those things and can implement AI, or does it benefit bottom-up, newly created companies in those spaces? Which one more so?

Mike Krieger

That's a great question. I think it can be both. At the highest level, the thing about AI and product design is that you have to dance this very delicate dance of showing the future and dreaming up what the models are currently capable of at their edges, because you want to design for where they'll be 3 months from now, which is how quickly things are moving, but not overpromise and underdeliver. That's a very trust-breaking thing.

If you're a startup, you can do a little bit more of the overpromising because people are kicking your tires early. They're early adopters, and they have a little bit more willingness to engage. It's much harder if you're an existing verticalized SaaS company and you say, “We've added AI,” and then people try it and say, “It's not that good,” or, “I thought it was going to do all these things,” or, “You said it could do these 30 things, and it does 2 of them well.”

I think those 2 groups have a very different challenge. On the former, you have established products and established behaviors, and you want to skate to where the puck is going without alienating your existing customers. I think there are some good patterns for doing that, and we can dive into those.

4. Should Founders Build for the Models of Today or Build for Models of the Future

On the startup front, you probably don't yet have the data. It's about landing the initial lighthouse customers, or you don't have the relationship, or you have some hypothesis about where AI will have an impact on a given industry or vertical. Your differentiation is not the established relationships; it's painting the future and finding ways of delivering that value quickly within a company that might be willing to take that bet on you.

Harry Stebbings

You mentioned startups building for where models will be. It's a very challenging time where startup products are so determined, quality-wise, by the quality of the models. A change in model can seismically change a startup's output, whether it's coding software or a legal platform. Should startups build for what we have today, or should we build for what we can project forward in time?

Mike Krieger

That's a really good question. I've heard from multiple people who say, “My startup was not a startup until Claude 3.5 Sonnet, or the second Claude 3.5 Sonnet.” I hear that from entrepreneurs who say, “This company was not a company until this model breakthrough,” where the accuracy went up, I don't know, from 95% to 99%, and now that's close enough for this industry. For some, it's from 70% to 90%, as soon as you get those generational leaps.

There have been times when entrepreneurs have been knocking their heads against the wall within a particular space—whether it's helping people code, helping with legal analysis, or something in healthcare. The cobbled-together, probably undersells it, lovingly assembled version of what they did, which involved multiple tools, was either price-uncompetitive because it required an Opus-class model that was not going to be supported by the underlying business, or was still worth doing because, when the model arrives, you're not starting from square 0.

The companies that benefit from those model-generation shifts are often not the ones that suddenly start that day. It's not, “Gosh, it sounds like Claude 3.7 Sonnet can do that.” It's the ones that have been beating against that wall. I take Cursor as an example. Somebody showed me a list of Hacker News front-page submissions from the Cursor founders over time, and it finally broke through. That was not their first product or their first iteration on it. They've been trying and going for, I don't know exactly how long, but it was not just quickly enabled by the model that came from that.

They were building context, building knowledge, and building experience about what had gone wrong or gone well in that space, so that the model could unlock them. To be more succinct: don't wait around for the models to be perfect. Explore in this space, be frustrated by the current generation of the models, and then aggressively try the next one, so that you can feel like you can finally deliver on the thing that you saw in your head if only the models were just a bit more capable.

5. Model Quality vs. Product UX

Harry Stebbings

When you talked about differentiated go-to-market and differentiated data, and then you said there are so many different releases and they come so thick and fast, is there value in the model layer if it's not a differentiated data game? Is it a differentiated go-to-market game? How do you think about that?

Mike Krieger

I think it's a couple of different pieces. On the model layer, and especially on the foundation-model layer, I think about 3 places where it's worth investing for a long-term place in the market.

1 is talent. I know it's hard to quantify exactly what talent means and what talent density means, but talent begets talent. You become an attractor, especially around a cohesive mission or a story about why you're building what you're building. I've absolutely seen that at Anthropic. I love our research team and feel like, monthly, we get some significant new hire who's come from potentially another lab or academia and has joined.

That's an advantage you have to cultivate and maintain because people are obviously free agents, and they can do what they want to do. You have to maintain whatever was attractive in the first place, but that is important because staying at the frontier requires more than just more of the same. It also requires figuring out what the right breakthroughs are.

The second one is that I think models over time get more different rather than more similar. Of course, there are a lot of similar benchmarks that people are looking toward, but there is something Claude-like about Claude, and I think there is something GPT-like about GPT. They have their pros and cons, both from a character and tone perspective, but also in the places where those models really excel.

For us, coding has clearly been 1 really big vertical that we've gone after. It wasn't an accident, and it's also not a thing where we just say, “Great, it's good at code; let's just continue to be kind of good at code.” Seeing that traction and seeing how many companies are now relying on Claude models for code, for example, or for agentic planning, inspires the next generation of what you want to do from a reinforcement-learning perspective. The first one is talent; the second one is focus and model characteristics that you develop more deeply over time.

The third one is—I got this question a bunch when DeepSeek came out—“What does DeepSeek mean for you?” I think there are things we learned on the technology side just looking at what they were doing, but from a go-to-market and place-in-the-market perspective, it has almost no impact.

That's because the relationships we end up having with companies are not, “They sent it for the API, and they want to exchange their input tokens for output tokens at some rate.” It's actually, “Hey, I want you to be my long-term AI partner. I want you to help co-design products with my applied AI team. I want to dream big with you. I want to think about not just your API, but also Claude for Work.” That looks more like being a company. I know it sounds trite, but what you're providing people is an AI partnership, not just AI models.

If you invert that to see what the failure mode looks like, I think it is resting on your laurels, not retaining your best people, believing that making the models incrementally better on every benchmark is enough, and treating the API as just a way of exchanging money for intelligence without figuring out how to be more of that AI partnership. If you can't do all 3 of those, I think you're in trouble.

Harry Stebbings

I do want to go into the coding element in a minute, but when we look at blockers or barriers to progression, what do you think the biggest blockers are today? I have completely disparate opinions from different people, whether it's Alexandr Wang or Jonathan Ross at Groq. What is the blocker: compute, data, or algorithms?

Mike Krieger

It's getting the environments in which the models get trained to better and better match real-world challenges that aren't single-shot. I know Alex has been thinking about this problem as well, because we talked about evaluations for agentic behavior. That's 1 very specific version of the broader thing that I'm talking about.

Even within software engineering, the work of a software engineer is not just to produce code. It's to understand what needs to get produced, work out the timelines with their product-management counterparts, deeply understand the requirements, and deeply understand the user and the use case that they're building for. Then they have to deliver whatever they built in a way that can be tested and iterated on and that has user feedback at the other end if they're building some kind of public-facing product.

That's hard. There's no evaluation for that. It's interesting that we call the most common software-engineering benchmark SWE-bench. To actually be a software engineer is a lot more than, “I looked at a pull request, I produced this pull request, I put this to CI, and then you're going to accept it or not.”

So building environments and evaluations that better mirror that—we think a lot about office professionals at Anthropic in terms of 1 of the use cases that is going to potentially be multiplied by these models in the future. Nobody's evaluating that well. There's some work around research evaluations that we're starting to get better at, and there are extremely convoluted evaluations—I mean that in the best way—like Humanity's Last Exam, which is very much, “Okay, multi-step reasoning.”

But there has yet to be the evaluation of, “I show up to a new job, I quickly understand what my role is, who is who in the organization, what relationships are being mapped, where to find extra information if I need it, and then be in the run loop of the functioning of the business.” That's a hard environment to capture.

To me, figuring out how we either break that down into component parts, which is probably part of the story, but also think about it holistically, is the biggest blocker to 1 slice of progress: how models go from being extremely good at extreme slices of things to being more generally helpful collaborators.

6. Will Human or Synthetic Data Be More Prominent in the Future

Harry Stebbings

Before we dive into those specialized products, on the data side, I had Aidan Gomez from Cohere on recently, and I asked him the question I'd love your thoughts on: when we look at the future of data within models, will there be more synthetic data that compounds on top of itself, or will human data continue to be the predominant data source that drives model progression?

Mike Krieger

For the model to improve, you do need a story around how you perhaps seed it with original human data but then generate all these synthetic environments about which you can pathfind and explore.

Claude has been having fun playing Pokémon this week, which has been a good but funny distraction for our research and engineering teams. Everybody's asking, “What are you doing?” and they're like, “We're watching Claude play Pokémon,” the livestream. Games are an interesting example, where you can imagine a lot of different runs through the same game, with constraints and rules that get a lot harder when the problem space is less well-defined than, “Did you make it out of Viridian Forest?” I never played Pokémon; I'm learning just by watching this livestream.

It's still important to be able to take golden paths but also synthesize a variety of approaches through them, so that you can think about how the model can progress in the face of uncertainty. I think it absolutely has to be a mix. The best models will come from that combination: for code, it's having good foundational data and good examples, but then also being able to explore a really wide variety of paths through that.

The other part that's still underappreciated is how you measure, evaluate, and get data in for character. I'm going to use a very loose word, which is “vibes.” What is exactly the feel of using a model? We don't really know until we sit down and play with it. In some ways, that's a nice property because it means there's almost this qualitative, human aspect to it, but it also means you don't have good regression testing on it.

Sometimes we'll go from Claude 3.5 to Claude 3.7 and people will say, “Claude seems friendlier but more terse,” or, “Claude seems more willing to answer my questions, but I wish it were better at creative writing.” These things are not easily evaluable. That goes to the data question, and so I think it's important both to have the data in there around these softer skills and to have the evaluations for them.

Harry Stebbings

I find it bizarre that we're able to choose models. You may say, “Of course you do, because there are specializations within them,” but when you project yourself forward 3 to 5 years, you will not be selecting which model you use. It's like selecting which Google you use. Am I completely wrong, or do I completely miss the point?

Mike Krieger

No. There's a concept that I love from—I come from a human-computer-interaction background—and you might have heard the term “leaky abstractions.” Software builders try to do a perfect job of encapsulating all the complexity under some little shell, and then users should not have to think about any of these things.

The reality is that the current state of most AI product design is an extraordinarily leaky abstraction. Having to choose the model—why should you choose Opus, Haiku, or Sonnet? Most people don't understand the difference. If you go to the OpenAI dropdown selector, there are a lot of models in there, and every single one of them has a good reason for being there. Yet the overall experience is, “Why would I choose 1 over the other? This capability is available here but not there.”

We suffer from this problem as well. Model selection is 1 issue. The second is that, once you understand how these models are built, you know they build up context. They have turns, and every turn has the full context replayed to it. That's how it's able to make the next inference.

What that leads to is a set of periods where every chat is different. I always think of it as when you're talking to a coworker: you might have different email threads, but it's still 1 coworker behind all of them. If you reference their favorite sports team or a project you worked on together, it's not like, “I don't know what you're talking about,” or, “I'm going to have to go retrieve my memory.” There's a shared underlying piece. That's another way we're forcing people into an understanding of the models that I don't feel like people should have to maintain.

The last 1 is prompting. As much as things have evolved and we've done a bunch of work around taking simple human prompts and translating them into prompts that are more model-optimal, I want to make that absolutely transparent to people. It shouldn't be something they're engaging with. If the model lacks clarity on the problem or needs help understanding it better, that should engage in conversation rather than show the difference between somebody who's an extremely good prompter and somebody who's not. That gap closes generation to generation, but we need to collapse it even further.

Harry Stebbings

How do you think about model quality versus product UX, and how do you prioritize and think about those 2 and the relationship between them?

Mike Krieger

You can't separate the 2 anymore. As a UX designer, I was just in a product review right before this call, and I was thinking about Instagram product-design sessions. You would have pixels, some synthetic data or maybe real data—we'd take my feed and reformat it into the UX we were proposing.

There isn't a lot of nondeterminism there. You're going to put it out to the world, and maybe people will use it in some ways. But designers, product managers, and definitely engineers today need to think, “What I'm actually doing is designing a scaffold and a product around a fundamentally nondeterministic system.” That means the evaluation, the model quality, and the prompting on the back end are all part of the product design, and that's going to have direct implications.

For example, you can prompt Claude to ask follow-up questions or not. That might be what you want in 1 part of the product but not another. You might prompt Claude to think longer about a problem and do more reasoning, or not. These are all decisions that, upfront, you are making in product design, and they're going to manifest in the actual product.

The other piece is that, as a startup founder or somebody doing classic B2B SaaS, you need to triangulate where the models are, where they're going, and what the user needs are. That's going to be the case in your product design as well. You're doing the evaluations, hopefully upfront, to see if what you're doing is even possible with the current models, or at least having an eye out for where they might be.

7. The Competitive Landscape of AI

The models change over time, and products change over time. If you don't have a good framework around evaluation, including regression testing those evaluations, you might launch a product that, 3 months later, people say, “The product used to be good, but something has happened where it's no longer serving that purpose.” You're not sure which of 3 things changed: the model, the product design, or the introduction of a different feature. Maybe the system prompt got longer. In many ways, it's the most complex product-development work I'll ever do.

Harry Stebbings

I interviewed Sam Altman in London from OpenAI, and he said 1 of the joys they have as a startup is that they can release things much quicker. It doesn't have to be perfect. The challenge, as they've gotten bigger, is that more and more weight and pressure is placed on every release. How do you think about “release it, it doesn't have to be perfect, let's get it in the hands of users” versus now, when Anthropic is a massive company with millions of users?

Mike Krieger

I think about this a lot, especially because you have different surfaces and different audiences that have different expectations of stability or desires to be on the cutting edge.

In an API product, people value predictability and stability, with the option of something that's more future-facing. It can be a very opt-in thing. I remember we launched prompt caching, which is a big cost savings for people. Initially, we did that through a beta header that you had to opt into, and a lot of what we do on the API is in that form.

If you do that for our customer-facing or more consumer-oriented products, it's lame to have people opt in. You want to be able to iteratively release and be experimental with people. You don't want to totally break their experience, but you have a little bit more permission.

Then we have these enterprise customers that are using Claude for Work in an enterprise. AI adoption in the enterprise is still an early-adopter product, so you can get away with more than you could if you were Salesforce. I don't know how many releases Salesforce does a year, but I know a lot of these companies do 2 or 3, usually oriented around some big event. We're far from that. We're still launching pretty quickly, but we're honestly still finding the balance. Is it a monthly drop? Do you ship as often as you can, but have an admin opt-in? That adds complexity as well.

It's a great question, and it's an active topic of conversation: how rapidly we can ship, knowing that we want to bring things out to the world, we don't know how they'll be received, and we want to learn. But as you accumulate notoriety and people start depending on you for workflows, you can't treat that completely wildly.

Harry Stebbings

Are we in a product-marketing nightmare? DeepSeek released something this week, OpenAI released something this week, Anthropic released something this week, and Mistral released something 10 days ago. Almost every day there's a new release, and maybe the world gets apathetic. How do you think about that, and how does it inform the way you think about product launches and messaging?

Mike Krieger

It is much more like Crossy Road. The things you had to watch out for in the past—the big rocks—were very well known in advance. Don't launch anything during WWDC week, because there will be a flurry of announcements. Watch out for the September iOS event, or another big rock like the holidays. It was much easier from a product-marketing perspective.

Here, it's like, “Okay, the car's going by. There's a gap in the cars. Launch tomorrow—or now. Oh, but now we hear there's a rumor.” It's so much harder. I've heard from people at other labs as well that everybody's trying to read the tea leaves: “Is it quiet? Is it okay to launch now? I think we're going next Tuesday.”

It requires a completely different approach, and I give credit to our product-marketing team because they've had to orient from a point where we launched Claude 3.7 Sonnet on a Monday and locked the blog post the night before at 9:00 p.m., which is not best practice from a marketing perspective. We were briefing the press that Sunday. Thank you to the people who helped us on the phone that Sunday, but it was right. That's the point where everything is done, ready, and locked, and we can go.

It involves the ability to react quickly and be nimble. Even when we release a model, there's a model card, evaluations, and a comparison table. There are things in that comparison table that were released the week before. Grok 3, for example, was just released a week prior. It involves completely changing what happens when those are released.

Harry Stebbings

When Grok 3 releases, does everyone at Anthropic and OpenAI say, “Oh, shit, they beat us again,” or, “Oh, shit, we won?”

Mike Krieger

I think it requires—I try to support the team by reminding them that model releases are going to happen, and at any given point you're going to be in the “it's so over, we're so back” cycle. You have to live that in AI. You can't get too down about 1 release, because it is inevitable.

Sometimes you're lucky and there's a 2- or 3-month period where the model you launched, or the product you launched, is still state of the art across all the things you care about. Sometimes it lasts a week. You can't overrotate on either of those. You can't rest on your laurels, and you can't get too upset.

The thing that's really useful to me is a chart I show in almost every sales call, mapping Anthropic's founding to where we are today and showing the milestones. At any given point, you can say, “Claude 2—that's pretty far behind. Claude 3 is state of the art. Now it's not.” You have to look at the trajectory and trust that you're going to continue making improvements.

Then remind yourself that, if everybody switched every single day purely because an evaluation changed, that would be an insane thing to do to your user base as a software provider. It would also make for an even crazier industry over time. People don't just deploy models. They're doing fine-tunes, or they're deploying models plus a lot of bespoke work to make the model great for that use case. That's not something that's going to switch overnight.

Or you're 1 of 3 or 4 options within a model selector. In a coding environment, for example, you're still in the mix and you still have a chance. I'm not sure if it's finding the meditative zoom-out angle or just getting used to the bumps, but it is definitely the case that every time there's a model launch, I assume every 1 of those labs is watching the livestream and looking at the evaluations, saying, “All right, now we've got work to do.”

Harry Stebbings

I would argue that brand is the most important thing. To your point, people aren't switching every day. They're saying, “I'm a Claude person,” or, “I'm a ChatGPT person,” and they identify with their models. Do you agree with that, or is it too glib?

Mike Krieger

I think that's right, especially on the consumer front. I was just reading Ben Thompson, and he has Nat Friedman and Daniel Gross on there pretty often. They're talking about some people being Claude people and some being ChatGPT people. That definitely happens. You like the personality, the interface design, and the vibe.

It reminds me a lot of the back-and-forth we had with Snapchat over the years with Instagram. Before that, people would launch a new product that was like Instagram but just for high-end photographers, or with an additional twist, or just 1 photo a day. That was BeReal.

I had this fake formula—I'm clearly not the mathematician at Anthropic—that social networks are made of format, audience, and vibes. For Instagram, we had Stories and the feed. Eventually, we had video. The audience initially was sort of hipster photographers, and eventually grew to anybody interested in visual storytelling or visual media.

But the vibes of Instagram, even when we had product similarities to Snapchat or Facebook, were very different. I don't know what the fake formula is for AI products yet, but I think it's some version of that. Model personality is probably 1 component. There's likely something around the scaffolding and prescriptiveness of the product that you're working around it, and then there are vibes. Again, they're hard to measure, but they're absolutely there when we have so many different models and providers.

Harry Stebbings

Open source is a very viable possible route, and distillation is looked at in a shady way. Is distillation really wrong if it ultimately propels the space forward? Even within the labs, I assume every 1 of them is using it internally. It's very valuable to take the knowledge of your highest-end model and make it lower-latency, more affordable, and so on.

Mike Krieger

There are a couple of places where this gets interesting. 1 is whether we want any nation to be able to distill models from any other 1. My personal answer is no. I think there's value, as AI gains capabilities, in being thoughtful about that from a national-security perspective.

The other piece is that, for advancements to happen at the rate they're happening and be sustainable over the long term, the labs need to be able to commercialize all of that training and innovation. Finding the right models for that long term is important.

I think open-source models—take Llama, for example—have been able to do that from their own research, data ingestion, and training. I would say distillation does not feel essential to unlock those things, and it poses other issues, even from a terms-of-service perspective.

Harry Stebbings

Does Llama show that there is no value in the model and all the value is in the data? If Meta is willing to give it away for free because it knows nobody can copy the data it has, is that what it shows?

Mike Krieger

It's an interesting question: is the quality of Llama due to the fact that Meta can—I don't know if they've said that they do, but they clearly can—train on Instagram, Facebook, and other data? Was Gemini better because Google could train on YouTube?

It's actually clearer to me that Gemini benefits from that. Whenever they have a good video-understanding demo, for example, I'm like, “Well, somebody has probably got the largest repository of video in the world and can likely train on a lot of those pieces.” It's less clear on the Facebook front. I've never heard people say, “What Llama does extremely well is generate content that would work well on social media.” It just seems like a good general-purpose model.

It goes back to our earlier conversation: the value is in how good your team is, whether you have the underlying data that you need to train on, and how useful your model is in actual use cases. That is the highest-order bit.

I almost wish I'd started with that, because, evaluations aside, evaluations are really useful for hill-climbing and internal research, but they don't tell the story of whether a model is going to be excellent at what it needs to be excellent at, or even if it's excellent at that thing in very narrow situations. As an entrepreneur outside the labs, can you rely on the model to be your representative in that product?

I think the value for the labs is in the team. It's in the model's ability to perform the right actions in the real world without so much nondeterminism that it becomes unreliable.

8. Do We Underestimate China's AI Capabilities

Harry Stebbings

I'm going to ask 1 question on this. It's not a trap to go down, but I've spoken to Alex Wang and Poolside on the show, and they said we deeply underestimate China's ability in AI. Do you agree that we underestimate it?

Mike Krieger

I think the DeepSeek piece that people seemed surprised by was that there were cutting-edge research teams there. If you were paying attention, that part should not have been the surprising piece.

Instagram was blocked in China fairly early, and then we saw the emergence of a parallel world of startups. If you take Facebook and Instagram away, what happens and what emerges? Those products were often very high quality. They demonstrated a lot of creative thinking and were built at scale. They were solving problems at the same scale as the problems Facebook was solving.

People love talking about the super app, WeChat, and the technical challenges it solved at scale. It would absolutely be a mistake to underestimate, or continue to underestimate, China's ability to train at the frontier—especially if it gets access to compute—and to continue innovating there.

I think it's a pretty Western-centric view that I've definitely seen happen in more traditional software. It's caught in this 1990s or early-2000s view of, “All they're doing is replicating what's already been working elsewhere.” There have been products that took a differentiated view, grew within the Chinese market, and then sometimes made that view external. TikTok is an interesting example of that.

9. What Did Anthropic Learn from Deepseek

Harry Stebbings

Just 1 final 1 before we move into the verdict products: did DeepSeek cause you to rethink anything or change anything about the way that you progress?

Mike Krieger

There are some architectural pieces, and I won't speak for the research team because they're the DeepSeek experts. They might say, “That's interesting. That's worth us considering,” or point to ideas that had been considered and were worth reevaluating. I think there was a piece of that.

Our plan was already to show the chain of thought when we launched our reasoning model, so that was not a reconsideration. But it was interesting to see somebody else do that, and there are some user-interface details in there. I think Grok does that as well now, so it'll be curious to see how that evolves. To your distillation question, that might be a reason why more labs choose not to show, or otherwise obscure, the chain of thought down the line.

10. Is Deepseek a Sustaining and Credible Threat?

From a product perspective, there were 2 things. I think the under-talked-about piece of DeepSeek is that they were able to go from nobody knowing about them to being, frankly, in many circles, better known than Claude. Likely my great-aunt was calling me about DeepSeek. I'm not even joking. It was cliché; it was actually happening. People were asking me, “What do you think they did to break through that maybe Claude hadn't?”

There was a lot of interest in world politics at the time, and the narrative was, “This was much cheaper,” whether that was exactly true or not. It was the story. I've had this conversation with our marketing team as well: I don't think we tell the Claude story well enough externally yet. We should talk more about what is different or notable about the fact that, by Claude 3, we were training a model at the frontier that was state of the art with a team that was much, much smaller than any other lab. We've always been very efficient with our compute as we train.

Whether that was a story DeepSeek told or was told for them by the media, it was legitimately a compelling story. The uniqueness of the moment was a big piece, and January, the new presidency, and China relations fed into the moment very well.

The second part was the product. They went from not having a product to having an iOS app that had a lot of good details. For me, it was a good nudge—but stronger than a nudge, more like a shove—that we need to get some ideas out to market more quickly, without focusing as much on exactly how polished they need to be in every situation. We need to be willing to put them out there and learn. Sometimes the novelty of an experience is itself valuable. It was the first time most people experienced a live chain of thought, and I wish we had done that sooner because it would have been novel for people to experience.

Harry Stebbings

You look at usage, and you see emerging-markets usage retention, while you don't really see that in Western markets. How do you think about DeepSeek as a sustained, credible threat?

Mike Krieger

They already have a level of awareness that gives them some ability to generate ongoing staying power. But when I think about all we're doing in these AI-first, lab-generated products, even 6 months or 1 year from now, I think that asking questions and having slight proactivity is not differentiated or interesting in the long run.

It should be, “Wow, I can now do something uniquely because I'm using Claude, DeepSeek, or any 1 of these products. It unlocked hours of work for me, made me smarter, and made me a better partner to the important people in my life.” It has to transcend the surface-level utility. Some people find the deeper level—don't get me wrong, those are your DAUs right now—but for a lot of people, they'll try it, generate a poem, or write a letter to their son. There's all this stuff they can do that provides some value in the moment.

I still think we're at day 1 around whether AI is an indispensable part of most people's work, and I think the answer is no for most of them. DeepSeek's staying power, and the staying power of all our products, will come from who can get there and do that sustainably over time, with the right product design, integrations, and deployment to actually succeed.

11. Transitioning from Model Provider to Application Provider

Harry Stebbings

Who can build those products is my big question as an investor. When does a model provider move into being an application provider? I'm fascinated to hear your thoughts on what is attractive enough for you to dedicate the resources to become an application provider, not just a model provider.

Mike Krieger

There are 2 main criteria that I look at. Our team is large, but our product team is maybe 1/10 of that. It's very large by Instagram year-2 standards, but small by large-SaaS-company standards. We're somewhere in between all of those different things, and we're supporting a lot of surfaces: Claude Code, the API, Claude, and Claude for Work.

Generalizability is really important. Even if we pick a persona or vertical to go after, we're going to build things that are general-purpose as a rule, with maybe some specialization at the user level but not at the level of a very bespoke workflow or use case.

Translation, transcription, and customer service—fairly horizontal, homogeneous things—seem like they would be in the pathway.

Harry Stebbings

I think they would be, except that I think there's a lot of valuable workflow and workflow knowledge that means you can retain a differentiated product over time.

Mike Krieger

If you're a power user, yes, perhaps.

Harry Stebbings

But if you're not a translator—if you're your mom, who uses it once a month for that odd thing she needs—then there's not really much there.

Mike Krieger

I think the role of, “We can help you translate this, and we'll get you to pay a $10 monthly subscription,” feels questionable because the models are already quite good at that. Maybe you're right that there's not much differentiation there.

If you play with ElevenLabs' Console and Workbench, a lot of the features they've built are clearly for people translating hours or voicing hours of content with a reliable voice across the whole workstream. Descript has some of the best product design in AI. They've clearly put so much time into the workflow.

I had to use it once for a personal podcast, and I thought, “This was clearly built by people sitting in this workflow day in and day out and understanding it.” Maybe we come to some synthesis of our views: there's value in the more professional use cases and the workflows unlocked by them. On the consumer, and maybe even prosumer, side, the basic AI product gets good enough.

Harry Stebbings

When you look at what you're brilliant at today, you do so well on the coding front. Is there a roadmap here to put your own IDE or code agent in? How do you think about that?

Mike Krieger

Again, with the product-focus lens, I think we have to pick our bets carefully. We built Claude Code, which we just released, as a command-line agentic coding tool internally first because we wanted to accelerate our own team. After seeing it play out for a couple of months, we thought, “This is good. It's not a solution to all coding problems, and it doesn't obviate the IDE, but it's useful enough to us in enough cases that we want to see people use it in the real world.”

Shipping is never free. You have to name it externally, find the right packaging, and handle the go-to-market piece, so we do it carefully.

My view of where the models are today is that you still need hands-on-keyboard work and an exchange of, “I did this. Is this right?” You need to say, “Let's pursue this direction,” and then, “Yes, this is great; let's put up a pull request.” Or, “No, we went down a false trail. Let's metaphorically unwind the stack and then keep going.”

That's why I think there's a role for something in between the IDE and full-on delegation of tasks, likely in the style of Cognition's Devin. Claude Code can be used for a certain category of tasks.

Our product engineers love Claude Code because a lot of product engineering is, “We have to update the back end, create the front end, submit these things for translation, and figure out why this still doesn't work.” It's that build-the-product-end-to-end workflow that does well with something that can work agentically across a lot of different things.

I did 2 pull requests last week. I hadn't coded since joining Anthropic, which made me sad, so I finally got to use Claude Code. I had never opened our codebase before, so I didn't really know how it was structured. Claude Code was very good at finding the file with the right piece and then making edits.

Obviously, not everybody is in the same situation I'm in, but it is really valuable for those use cases. When I think about the coding space and where we can play and add value, it's really on the agentic side, not the IDE side.

There are other companies that wake up and go to bed every night thinking about how to make a great IDE. That involves low-latency autocomplete, the right integrations, figuring out how to work with the VS Code plug-in ecosystem, and all that complexity. There's a lot of valuable work there that's different from what we're doing. We can really play in talking to these models and doing real work with them in that agentic loop, while recognizing that they're not yet at the place where, for many use cases, you can let them run free for hours. You need more of a human-evaluation piece.

12. What is the Role of a Software Developer in the Future

Harry Stebbings

You power and work with Cursor, Codeium, and StackBlitz. My question to you is: when you look at the changes we're seeing in developer behavior, what will the role of a software developer be in 3 to 5 years?

Mike Krieger

It already starts to look different. I was a huge early proponent of GitHub Copilot. I think my quote was on the homepage for a while—I don't know if it still is—because I saw the potential.

Then GPT-4 came out before it had multimodality, and I was trying to use it with Swift. I would draw ASCII art of the screens I was trying to build for Artifact, then go make coffee because it was quite slow at the time, and come back to an 80% version. Obviously, now it would be a 95% to 99% version with something like Claude 3.7 Sonnet.

I think the skills that become important are, 1, being multidisciplinary. It's knowing what to build as much as knowing the exact implementation you want. I love that about our engineers. Many, maybe most, of our good product ideas come from our engineers and from them prototyping. I think that's what the role ends up looking like for a lot of them.

The second piece is that code review changes when you're mostly evaluating AI-generated code. I experienced this myself. I put up a pull request, and some comments came back saying, “Claude Code does this sometimes. We don't actually use default arguments in this case.” I thought, “Oh, damn it.” If I were coding it, I probably would have noticed those patterns better.

There are 2 sides that need to happen. Models and the infrastructure around them need to learn from codebases and code reviews better so they can produce code that feels idiomatic to that company. But we also need to evolve from being mostly code writers to mostly delegators to the models and code reviewers.

That's what I think the work looks like 3 years from now: coming up with the right ideas, doing the right user-interaction design, figuring out how to delegate work correctly, and figuring out how to review things at scale. That's probably some combination of a comeback of static analysis and AI-driven analysis tools that determine what was actually produced. Is there a security vulnerability? Is there another flaw? Is there a bug?

Computer use plays a part. You can tell I get very excited about this space. Automated testing of the UI is important too. What would be great is that you delegate a task—let's say a year from now; 3 years is crazy—and, when you come back, it says, “I evaluated these 3 approaches. I tested them all out. I had a different agent try them in a browser. This 1 worked best. I've run it through another agent that performed a vulnerability test, and it all looks good. All we need to do is help you resolve this 1 question. Let's review this critical section of code to make sure it's what you really wanted.”

That feels like you're suddenly empowered to be more of a manager and delegator to these systems rather than just a partner in the loop.

Harry Stebbings

You said 3 years sounds ridiculous and that 1 year would be much more realistic. I agree. When we look at the speed of scaling, do we think we hit a plateau or an asymptote in product releases and the speed of development? It feels so fast now. Do we hit that plateau, or do we continue in this exponential progression?

Mike Krieger

It's a question I think about a lot. I started the year by looking at our product-development process and where we are Claude-enabled and where we're not.

Claude can be useful in taking an initial idea and creating a PRD from it. Claude can be useful in the coding side. Claude can synthesize a lot of conversations people are having about a product, find the thorny issues of disagreement, drive alignment, and figure out what to build.

Actually figuring out what to build is still the hardest part. That's best resolved by getting together in a room and talking through the pros and cons, or going off and exploring it in Figma and coming back.

Like any dynamic system, if you optimize 1 piece, something else becomes the source of the blockage or the critical path. Alignment, deciding what to build, solving real user problems, and figuring out a cohesive product strategy are still very hard. The models are probably more than a year away from solving that.

That is the constraint. It's why I'm bullish on startups being able to explore the space. I remember this from both my Instagram and Artifact days: there's a difference between alignment being a coffee conversation in an afternoon and steering the ship of a large company that has commitments to customers and all of those things.

That's still a very human problem, and I think we're at least 3 years away from the models solving it at that level of abstraction.

Harry Stebbings

We mentioned consumer products and building them. When you think about building new products for consumers versus building the API division of the company, which is very significant, how do you think about the balance and the trade-offs between building an API business and building an end-user consumer business?

Mike Krieger

I think about what we get out of each. We learn a lot more quickly with first-party products. A specific example is Claude Code. Within a week of deploying it internally, we found that 1 of the tools it had access to was not being used as well as it could have been. That made its way directly into Claude 3.7 Sonnet.

That's a way in which internal dogfooding of a first-party tool directly led to a model improvement in the next generation. There are a few other places where we've seen that. It's much harder with a third-party product. They might tell you something's wrong, but it's more arms-length. Even though we work very closely with some of the coding startups you mentioned, it's still not the same.

There's a lot of value in what we learn there. Then there's the stickiness and loyalty we talked about. I think it's easier, from a consumer perspective, to build a brand around a product than around just an API.

The fact that we power a lot of these coding products is visible to people. We're often the default in the dropdown selector, and if you're in the know, you know. But not everybody does, and it's still not the thing they downloaded or installed that they're going to tell their friends about.

On the other hand, it's a place where we've gotten tremendous distribution. We're not going to invent every company, and we're not building every product ourselves. In that way, it reminds me of my investing days: you get to see a lot more, and there's more than 1 shot on goal. It's been a fairly even split from a resource-allocation perspective.

If anything, we've underinvested a bit in 2 things. 1 is having a faster iteration speed on first-party products; that's my current obsession. The second is, on the API side, figuring out how to build abstractions beyond tokens in, tokens out.

Every time we do that, we get great feedback from people. Whether it's helping the model plan and work agentically, having the model build more knowledge graphs and repositories of how companies operate internally, perfecting tool use, understanding very large amounts of context, or having memory that transcends conversations, those are problems worth solving on the API.

They're things where we can take what we learned on the training side, map it directly to the API, and build good products around it. That's how I think about the 2. At Instagram, it was easy: 95% product and 5% API. That's all we needed to do.

Harry Stebbings

What can and will you do to increase product speed on the first-party consumer side?

Mike Krieger

There are 2 things. 1 is recognizing that we were running a larger-company playbook for what are still startup products. Even if the company has good traction, the API business is doing well, and people are using Claude and upgrading to Claude Pro, it's still early days. It's still do or die, or make it or break it.

We need to operate that way. That means getting the right people together sooner and faster and ignoring organizational boundaries. We got too calcified: “This is on this team's plate,” or, “You can't get this done this quarter because it's not on this team.” I understand why organizations evolve, and some of that is natural, but we can't afford it right now.

It's been much more about, “Who are the right people? Let's get them together. Let's clear away all the other distractions.” I need to clear out my calendar so I spend more time in product reviews and design reviews than in administration.

13. Is Europe Stronger or Weaker in a World of AI

Harry Stebbings

Did DeepSeek show the benefits of constraints? Do Western companies—respectfully, you and OpenAI—have too much money?

Mike Krieger

The way I would put it is that the adoption we've gotten for our products is ahead of their actual product-market fit because they are still the best ways of getting the models. I don't think that's durable over time, so it's not something to rest on.

14. Quick-Fire Round

I also think we're underserving people because we haven't gotten the right products yet. That's what I wake up stressed about every morning—or inspired by, depending on the day. We've got so much work to do on that side.

Harry Stebbings

I love it. I want to do a quick-fire round. I'll say a short statement, and you give me your immediate thoughts. Does that sound okay?

Mike Krieger

That sounds great.

Harry Stebbings

What's OpenAI done better than you?

Mike Krieger

They've moved faster at shipping V1s, sometimes even ahead of where the model is.

Harry Stebbings

What have they done worse than you?

Mike Krieger

Probably personality, and having the features they build be cohesive.

Harry Stebbings

Which alternate model provider do you most respect?

Mike Krieger

OpenAI. I think they've balanced first-party product development and an API that people use at scale. We had an Instagram principle that was, “Do the simple thing first,” and I think they often do the simple thing first.

Harry Stebbings

If you could rebuild the Anthropic product and stack from scratch, what would you do differently?

Mike Krieger

I love this question.

Harry Stebbings

I do too. It's a good 1, isn't it?

Mike Krieger

It's a really good 1. The things we built that were very valuable last year now feel like they have some cost to the information architecture. That sounds like a very nerdy way of describing it, but people should not have to think about projects versus artifacts versus chats and how they all relate.

Tearing it all down and asking what actually matters: do you have the right context in the right conversations? Do you feel like you can always know where to go next in the product? Is Anthropic and Claude itself being a helpful guide to what work is most important to do next?

That's a different paradigm from, “I know to create a project.” If you get good at that, it's an amazing product, but there are a lot of steps along the way. That's the fundamental thing on the product side.

On the stack, Claude.ai—and probably ChatGPT.com—were initially built very much as showcases for the models. They weren't built to be the foundation for a much more complex, multiproduct system. We have an active effort right now around tearing down some of that and rebuilding the core UX so that it feels good.

It doesn't feel great right now. It feels like it's been an evolution of a product that served a purpose at the time but is now being asked to do many more things. The incremental approach is now both harder to add to and getting slower.

Harry Stebbings

What have you changed your mind on in the last 12 months?

Mike Krieger

How important first-party products are. I saw the growth in the API, and I thought, “This is what we should invest much more of our time in.” But you'll miss out, and you won't have enough of a durable moat, if you're not equally investing—maybe even investing more—on the first-party side.

Harry Stebbings

How much did it hurt you being late to that?

Mike Krieger

Significantly. Take the DeepSeek moment. Ideally, the story that there is more than 1 leading-edge AI product to use is a narrative we should have captured. I think it hurt us there.

Harry Stebbings

What's a major technical or product challenge on the horizon in AI that no one is talking about but that you think is critical?

Mike Krieger

As the models get more capable, the headline is discernment and privacy. They'll also become more knowledgeable. They'll be in conversations with you about everything from something intimate to something sensitive from a company perspective. They'll have access to all of your company's information.

Everybody loves to talk about agent-to-agent interaction. The intersection of those 2 things is not discussed enough. Do you trust your Mike agent or your Harry agent to be out in the world, not be jailbreakable, and not reveal something it knows that is personal or sensitive?

My metaphor is my 5-year-old daughter. It's great watching her with somebody she's just met because she doesn't quite differentiate between things that are secret and private to our family and things that are okay to talk about with a new friend or somebody at the checkout aisle. Discernment is something people acquire over time.

I think this is underappreciated and probably underresearched from a model-capabilities perspective. Models fundamentally want to be helpful, and that is not always what you want them to be. There's a safety case for that, but there's also a privacy and data-security case.

Harry Stebbings

Do you worry about your 5-year-old becoming more comfortable talking to models and agents than she is with humans?

Mike Krieger

I've had so many conversations with Alex Wang about this because he has this whole idea that, in the future, most friends will be AI friends. I don't think he's wrong. There are ways in which that's already starting to be the case, with people having online gaming experiences where some of the characters are non-player characters. You might have a more comfortable existence there, even if you're not breaking through.

I do worry. She's so gregarious that I'm not worried in her particular case, but let's abstract to the broader sense. There's a lot you can learn from what it feels like. I was a fairly awkward teenager, and I probably could have benefited from a practice mode for AI interactions around some of these things.

At the same time, that doesn't feel like it's totally closing the loop around the consequences of real interaction. It's the difference between reading about what it's like to have your first really hard argument with your high-school girlfriend and actually having it. When you're in that moment, it's different.

It's like the thought experiment where somebody is in a black-and-white room, only reading about the color red, and then goes out into the world and sees red. Is there something qualitatively different about that? Absolutely.

Is there something different between talking to a model and engaging with a model, even in emotional role-play, and having that same interaction with a real human? Absolutely. It is probably a helpful piece of future human interaction, and absolutely insufficient as a replacement for it.

Harry Stebbings

Does Europe become more or less relevant in an AI-driven decade?

Mike Krieger

I want Europe to do well because I love a lot of Europe. I lived in Portugal growing up as well. I saw a funny, perhaps somewhat defective version of this argument: if real-world experiences and human interaction become more valued, Europe becomes more valuable as the world's capital of sensory experiences.

That feels weird if that's all you're resting on. It feels limited. What I think will be really interesting from a European perspective is the things that Europe holds very strongly about lifestyle and society, and then attempts—sometimes not elegantly—to enshrine in best practices or even laws.

When we think about product design, data privacy, and selling to German users or German companies, there's a different set of questions that gets asked. They're often very helpful questions. Maybe the bull case is that those questions are relevant to everybody, and Europe will be at the leading edge of asking them.

From a lab perspective, it's a harder question to answer. There may be some combination of access to compute and moving further up the value chain. If building applications on top of these models becomes much easier, and you can go from 0 to 1 and be more nimble than labs that have tens or hundreds of millions of users and have to move more slowly at that scale, can innovation happen there? Probably.

But it probably involves a different regulatory and startup-ecosystem environment to make that the case.

Harry Stebbings

Dario Amodei has said that this could be the generation that lives to 150. I'm slightly butchering and summarizing his quote, obviously, but this could be the generation that finds cures for diseases like multiple sclerosis with AI. My mother has multiple sclerosis. Do you agree with his optimism, and how do you think about AI increasing longevity and human lifespan?

Mike Krieger

I think the potential is huge. Today, AI is helping close the loop on drug discovery and clinical trials. A company likely called Huma used to take, I think, 15 weeks to do its clinical-trial reports, and now it uses Claude and gets them done in 20 minutes.

That's a step change. There were years of research that preceded it, so I'm not saying we've cut years to weeks or years to minutes. But it's a point in the process that we can make faster with the models today.

Then you see Arc Institute, the science and research institute that Patrick Collison and others have started and funded. They're working on foundational models for cells, where you suddenly have a real cell model that you can run experiments on. That kind of thing should accelerate drug discovery and experimentation tremendously because you're cutting a loop there.

I'm very optimistic. There are a lot of places where AI is underutilized relative to its potential. Some of the smartest people in my generation were working on serving more targeted ads. Maybe that was true at 1 point, but a lot of them today are working on how to make models tremendously useful, valuable, and intelligent across many domains.

Harry Stebbings

Mike, you've been fantastic. Thank you so much for letting me completely unpack all of my questions on you without warning. You've been amazing.

Mike Krieger

My pleasure. This was really fun.

Mike Krieger, Instagram CoFounder & Anthropic CPO: Where Will Value Be Created in an AI World?|E1265 | BidClub