[BidClub_]
20VC · · 45 分钟

Kevin Scott,Microsoft CTO:对 DeepSeek 的评估,以及我们如何低估了中国人

Harry StebbingsKevin Scott

YouTube
TL;DR
  • Scott 对价值归属的核心判断是:模型不是产品。 模型“极具价值”,但只有通过产品把模型连接到用户所需之处,价值才能兑现;“归根结底,我认为产品最重要”。基础设施和算力会在过程中实现变现,但“我们建设基础设施,不是为了基础设施本身”。Harry 的困惑很有代表性——互联网、移动互联网等每次范式转移初期都是如此,“真正具备持久价值的想法少之又少”。
  • 关于规模化撞墙的说法,他认为“这简直荒谬”。 “我非常清楚我们现在在做什么、下一步要做什么,也看不到规模化规律的尽头。”他的保留意见在经济层面,而非技术层面:最终会出现边际回报递减,直到“成本高到我们会认为再多花1美元、让系统聪明一档并不值得”;但“现在还没进入瞄准镜”。
  • 关于 DeepSeek R1,他说:“我们有过比 DeepSeek R1 更有意思的模型,但选择了根本不发布。” R1 是“扎实的技术工作”,只是“价格性能提升曲线上的一个点,可能其他所有人都没看见”;而推理价格性能的提升,更多来自软件,而非硬件每代约2x的增益。真正值得警惕的是外界的惊讶:“我们确实应该高度尊重中国创业者、科学家和工程师的能力……这本不该令人惊讶。”
  • 代理的关键缺口是“明显缺失”的记忆。 这让代理“非常事务化”;他预计记忆能力未来1年会大幅提升,也希望出现更多异步调度——“把代理派出去做事,而你自己暂时不用关注”。路径是:5秒任务→5分钟任务→像交代工作给同事一样完成委派。而且是多个代理,不是一个代理,因为产品经理可能必须成为领域专家,帮助搭建代理的反馈闭环。
  • 5年后,95%的净新增代码将由 AI 生成。 “很少会有代码是人逐行写出来的”,但“创作权中更重要、更有意思的部分,仍将完全属于人类”。这又一次对应从汇编到高级语言的过渡——“以前有些老家伙会说,不会写汇编就不算真正的程序员……现在没人再谈这个了”。他希望小团队更容易做成大事:“10名真正优秀、动力十足的工程师,配上强大的工具,就能做很多事。”
  • 快速问答透露出他的几项立场:他最尊重的竞争对手可能是 Anthropic,前沿模型已经可能比普通全科医生更擅长健康诊断。 他在节目第1年最大的体会是:“如今最强的前沿模型能做什么,与它们实际被用于做什么之间的差距,比甚至2年前还要大。”
摘要 · 为研究而整理的核心内容

1. 模型不是产品——产品连接用户之处才是价值池

  • Harry 开场坦言,自己很久以来第一次不知道价值究竟在哪里,这被 Scott 视为典型现象:“每次重大技术范式转移初期,都会发生同样的事。”早期互联网和早期移动互联网阶段,“所有人都在设想什么会有价值,但真正具备持久价值的想法少之又少”。
  • Scott 过去两年的核心判断是:“模型不是产品。”当被追问这是否意味着模型没有价值——Andrew(很可能来自 Cerebras)曾告诉 Harry,算力才是赢家——他斩钉截铁地说:“不不不,模型极具价值,但只有通过产品把模型连接到用户所需之处,价值才能兑现……归根结底,我认为产品最重要。”好模型、基础设施和高效算力都会变现——“人们在构建这些产品时,就会需要消耗你的平台”——但“我们建设基础设施,不是为了基础设施本身”。
  • 面对困惑,答案不是坐等。 “如果你有创业精神,这是有史以来最好的时代。”推出产品、收集数据,并且“对自己狠一点……不能爱自己的想法爱到看不见现实”。每轮早期周期都有一个典型陷阱:“技术人员沉迷于技术细节,忘了真正重要的唯一一件事,是做出好产品。”
  • 初创公司与既有巨头之间,历史周期显示价值创造地点“往往相当均衡”。Microsoft 的做法,是把新能力应用到现有客户身上;但“任何像 Microsoft 这样的实体……都不可能拥有足够的想象力和视角,知道每一件有趣的事是什么”,而工具从未像现在这样“便宜且易于获得”。

2. 规模化撞墙:“简直荒谬”——极限在经济层面,而且尚未进入视野

  • 对于“我们正在触及规模化极限”的说法,他回应:“这简直荒谬。我非常清楚我们现在在做什么、下一步要做什么,也看不到规模化规律的尽头。”他的保留意见有明确边界:直觉上,他认为最终“必然”会有极限——不像有些人认为,智能可以突破人类20瓦大脑的限制,持续扩展到“奇怪的领域”——但他预期的渐近线是成本决策:“成本高到我们会认为再多花1美元、让系统聪明一档并不值得。”目前,“还没进入瞄准镜”。
  • 数据方面,合成数据的占比正在上升;与低质量数据相比,高质量数据加上专家级人工反馈,对后训练的价值高得多——可以把它们“放大成一组正确的 token”,其价值远高于“在互联网上漂浮的、未经区分的 token”。
  • 他最希望得到答案的开放问题,是数据价值至今没有一套科学:“很难定量知道,新增数据对模型质量的增量价值究竟是多少……人们的大多数判断都没有依据”;而测量显示,有些人认为数据价值所在,与数据实际贡献之间存在“相当大的脱节”。其底层错误是把模型当成“世界上最糟糕、最昂贵的数据库”——模型的重点是对信息进行推理,而不是把信息回忆出来,因此训练推理需要的是“不同的 token”。

3. DeepSeek 只是微软早已看见的价格性能曲线上的一个点

  • Scott 对 R1 的判断,建立在过去几年“优化模型性能方面令人难以置信的进展”之上:模型规模不断扩大,API 调用却越来越便宜。硬件“如果幸运,每一代大概只能带来2x的价格性能收益”,软件则能带来“大得多的提升”。R1 是“价格性能提升曲线上的一个点,可能其他所有人都没看见”,但对那些深度参与系统研发的人来说并不意外——“而且它不会是最后一个点”。
  • 他说这话时语气平淡:“我惊讶的是,大家觉得它有这么有意思。我们有过比 DeepSeek R1 更有意思的模型,但选择了根本不发布。”他认可这项工作——“扎实的技术工作”,而且将其开源发布“非常酷”;真正令他意外的是公众的反应,而不是模型本身。
  • 他从中得到的启示是:“开发者希望拥有大量选择……我们必须比过去给人们更多‘如何做’的选择。”至于中国,他没有任何保留:“我们确实应该高度尊重中国创业者、科学家和工程师的能力……所有人看起来都那么惊讶,这本不该令人惊讶。”
  • 他也承认自己改变了对开放与封闭的看法:“读研究生时,我是彻头彻尾的开源狂热分子……后来变得务实得多。”他对未来的参照系是搜索:开源搜索引擎存在,搜索即服务存在,Bing 和 Google 也存在,但“搜索领域所有的经济价值,都流向了那个搭建起巨大基础设施的人”,并且这个基础设施拥有自己的反馈闭环。他预计这里也会是“两者都会有很多”。

4. 200年的交互范式正在断裂——代理不是一个,而是很多个

  • 他最喜欢的框架是:“我们使用计算设备的范式,实际上已经持续了大约200年,可能从 Ada Lovelace 写出第一段程序算起。”你要么是程序员,要么依赖一个提前猜到你需求的程序员。AI 改变了这一点:它“能够理解你想让计算设备替你做什么,并想办法让这件事发生”。时间表也很明确:“我不认为这是明年会发生的事,但可能也不会等10年。”
  • 软件不再需要那么多预判性工作——团队猜测用户的细分需求、写代码、把代码挂到界面上,再不断根据反馈打磨。“这类工作不再需要那么多。”工程师仍然是能力基础设施的建设者;而把这些能力呈现出来的界面,“可能会是代理”。
  • 他的反直觉判断是:“我不相信一个代理包打天下的理论。我认为你会拥有大量代理。”原因在于,产品经理可能必须成为医学、药物研发,甚至“早期轮次风险投资”等领域的专家,搭建反馈闭环,让产品经理与用户共同教会代理如何完成这些工作。

5. 记忆与异步调度,让代理变成同事

  • Harry 质疑,大家可能高估了近期采用速度:全球最大公司真的会在1到3年内运行代理吗?Scott 的回答是:“使用总是跟随效用。”软件开发代理已经让开发者从怀疑者变成了“你休想从我冰冷的尸体上拿走它”。
  • 对于缺乏锁定效应的担忧,他以搜索为例:“搜索没有锁定效应……你可以把下一个查询发给另一家搜索引擎——但你不会这么做。”留存来自每天持续“把代理做得更好”;只要做得好,“用户就会继续选择你”。
  • 缺失的关键能力是:“它们明显缺少记忆,这让它们非常事务化。”记忆“未来1年左右会好得多”,从而带来更强的抽象能力和组合性:一个问题解决1次、记录下来,以后不必再从第一性原理重新推导。他也希望未来12个月出现“更多异步任务”:把代理派出去,让它在无人关注时自行工作。路径是:5秒任务→5分钟任务→“越来越重的工作……像交代给同事一样”。他的具体标准是:早上5点的代理消化一夜邮件,起草紧急回复,而他自己喝着咖啡。
  • 对代理怀疑者,他说:“如果你不打算采取行动,我不知道怀疑一件事能赢得什么奖品。”他引用了同事一本书的标题《悲观主义没有奖品》(There's No Prize for Pessimism)。Harry 补充了一句,大概来自 Collison 兄弟中的一位:“悲观者是对的,乐观者赚钱。”Scott 回应:“他们没有错。”

6. 95%的代码将由 AI 生成——抽象层级上升,创作仍属于人类

  • 他的判断是:5年后,“95%的代码将由 AI 生成……很少会有代码是人逐行写出来的。但这并不意味着 AI 在承担软件工程师的工作”,因为创作权“仍将完全属于人类”。他的类比是:他12岁写出第一段程序,如今52岁,“已经41年了”;在从汇编到高级语言的过渡期,“以前有些老家伙会说,不会写汇编就不算真正的程序员……现在没人再谈这个了”。GUI 构建器生成“大量垃圾样板代码”,也是同一逻辑。
  • 顶尖程序员仍然会深入底层:“极其优秀的程序员……一路理解到底”,当 AI 生成的代码出错时,他们会“钻进更低的抽象层级”。在 Bolt、Lovable 这类工具的世界里,人人都是程序员吗?“我认为是。”但“世界上最难的计算问题”仍然需要计算机科学家,只是他们如今拥有极其强大的工具。
  • 团队结构也会变化:“我希望小团队做大事能变得更容易……10名真正优秀、动力十足的工程师,配上真正强大的工具,就能做很多事。”他在 Microsoft 的内部目标是:“我不希望工程师的野心与他们尝试实现野心的能力之间存在任何距离。”
  • 他带领工程团队20年后最大的敌人是技术债——“就像金融债务一样,它会产生利息”。AI 有机会把这个“完全零和的问题”变成非零和问题;Microsoft Research 约1年前成立了一个实验室,唯一使命就是“规模化消除技术债”。第1年的经验具有普遍性:“如今最强的前沿模型能做什么,与它们实际被用于做什么之间的差距,比甚至2年前还要大。”

7. 快速问答:尊重 Anthropic、直方图法则,以及“我们走得够快吗?”

  • 他最尊重的竞争对手,在 Google、Anthropic 和可能的 Meta 之间选择:“可能是 Anthropic……我认为 Dario 做得不错。”谈到物理约束,他说自 GPT-4 以来,“我们一直以每小时1000英里的速度建设基础设施……已经在以实际上可能的最快速度前进”;真正的瓶颈是“混凝土浇筑的速度,以及电网扩容的速度”。
  • 他得到过的最佳建议,是把能力想象成一张从白痴到天才的直方图:通过极大努力,你也许能把自己向右移动“1个、最多2个桶位”,所以“你花在努力变得平庸上的所有时间,都没有用来做你真正擅长的事”。他承认自己的弱项是官僚主义:“如果我愿意,我大概可以成为一个非常平庸的官僚。”这条领导原则很可能来自 Satya Nadella:“创造能量,产生清晰度。”
  • 他把一个科幻式预测当作事实:“前沿模型可能已经比普通全科医生更擅长健康诊断。”他提到自己家人在弗吉尼亚州中部乡村地区获得的医疗照护并不充分;世界需要“意识到它们确实很强,这样我们才能部署这些东西”。
  • 他很少被问到的一个问题是:“我们走得够快吗?”他的答案是否定的。更快意味着在教育上“进行极大投入”,让每个孩子都觉得这些工具就是为他们打造的;同时把激励机制对准 AI 的部署,投向“我们目前认为存在稀缺的地方”——医疗、气候、教育。“让我们把投资投到这些地方。”
Kevin Scott

This is the best time to be alive if you have an entrepreneurial spirit. I can very clearly see what we’re doing now and what we’re doing next, and I don’t see the limit to the scaling laws. I don’t believe in this one-agent-for-everything sort of theory.

I think you’ll have a lot of agents. The reason I think you’re going to have a lot of agents is because your product managers are probably going to have to be domain experts. The agents will definitely be less transactional and less session-oriented going forward.

Harry Stebbings

Kevin, I am so excited for this. I was just telling you, I was listening to you and Dwarkesh on my run. I don’t think I’ve ever run quite as fast, which clearly means the conversation was brilliant.

1. Where is Enduring Value in a World of AI

I’ve never done a 10K so fast. I wanted to start with a super-easy question, which is: my job as a venture investor is to try and determine where value lies in different given moments. I look at the world today and, for the first time in quite a long time, Kevin, I don’t know. Where does value lie sustainably in this next generation of AI, do you think?

Kevin Scott

I think the thing that you just described, which is that all of a sudden things have gotten a little less clear than they have been, is exactly the thing that happens at the beginning of every big technological paradigm shift and every new cycle that’s driven by it.

It was super confusing in the early days of the internet, and I think it was super confusing in the early days of mobile, where everybody had these ideas about what was going to be valuable, and very few of those ideas were actually the durable ones that proved all the way through.

Harry Stebbings

In those moments of transition, where there is this confusion, what have you learned is the right action to take? Is it to be active, to iterate and learn, knowing that you’ll make so many mistakes that you regret, or should you sit on your hands and watch others make those mistakes?

Kevin Scott

You definitely shouldn’t do the latter. This is the best time to be alive if you have an entrepreneurial spirit.

The thing I think you have to do in these moments is not forget the things that you’ve learned from past moments about what works. It’s not, “Okay, well, here’s the specific thing to do.” It’s about how you go about doing the exploration that you just described.

Product matters. I’ve been saying this for the past couple of years: models aren’t products. Everybody was just so fascinated by the infrastructure itself, and this is also a characteristic of the beginning of these cycles: you have technical people who get swept up in the technical bits, and they kind of forget that the only thing that really matters is making good product.

That’s where we’re at right now. You have to make good product, and you have to have ideas and conviction. Then you have to go get stuff done really fast so that you can see whether you’re full of crap or not about the conviction that you have.

You have very few patterns at the beginning of a cycle to snap to. You’re not looking at someone else’s success and saying, “Okay, well, I’m going to do that, but just a little bit better.” You’re trying to figure out something completely new, and the only way to figure that out is to launch stuff, gather data and iterate. You have to be super, super brutal with yourself about what you’re seeing. You can’t love your idea so much that you overlook what you’re seeing in the data and the feedback that you’re getting.

Harry Stebbings

You said that models aren’t products. I just had Andrew—likely from Cerebras— on the show, and I asked him this question: if we think about compute, or hardware, then we think about models, and then we think about apps, where does the value lie? Naturally, he said compute.

When you think about that three-pronged tier of value, and you say models aren’t products, if they’re not products, does that mean they’re not valuable?

Kevin Scott

No, no. They’re super valuable, but they’re only valuable to the extent that you can connect them to things that users need via product. In the limit, I think product is the most important thing.

If you build good models, good infrastructure around models, efficient compute and all of these other things, you’re going to get lots of ability to monetize all of those things. As people build those products, they will need to consume your platform and your infrastructure, and all of that’s good.

But most of the value has to be in the products. We don’t build infrastructure just for the sake of infrastructure. We build infrastructure so people can make products.

Harry Stebbings

This is a leading question. If you think about those products, who benefits most? Is it startups that are able to integrate new technologies very easily from the bottom up, starting from nothing, or is it Microsoft integrating AI into incredible distribution already, or Google doing the same? Who benefits most in that respect?

Kevin Scott

Again, if you look to past cycles, you’ve got a pretty good mix of where value gets created across startups and new ventures and existing enterprises.

I think everybody’s kind of doing the same thing: you’re trying to discover the new. If you’re a big company like Microsoft, with a long tradition and a bunch of successful things already in the market, the thing you’re trying to do is figure out which of the things that you already know super well, and which of the customers that you’re already serving super well, you can serve with this new set of capabilities that you can provide.

Hopefully, you know that I run, among other things, Microsoft Research. I also have a charter to say, “Hey, can we go try to shine flashlights in places that no one else has shown them before?” We try to discover some super-disruptive, brand-new things.

That’s the job of the startup ecosystem as well. I’m also an angel investor, I advise startups, and I’ve worked at startups, so I think it’s really important that you’ve got lots of people hunting for those interesting new things.

I have super-high conviction on that in the AI platform transition that we’re going through right now, because it’s impossible for any entity, like Microsoft or any other big company, to have enough imagination and enough perspective to know what every interesting thing is. Having this vibrant ecosystem, with lots and lots of people exploring where value exists, is incredibly exciting and necessary.

I also think that there has never been a moment where the tools, infrastructure and platform have been as cheap, accessible, available and easy to use as they are right now. It’s super easy to pick stuff up and get cracking.

2. Why Scaling Laws are BS

Harry Stebbings

I was listening to your show, as I mentioned earlier, and you pushed back on the idea that we’re reaching scaling laws and this kind of asymptote of efficiency or effectiveness, when many people suggest that we are hitting a scaling asymptote soon. First, why do you think we’re not?

Kevin Scott

That’s a ridiculous statement. I can very clearly see what we’re doing now and what we’re doing next, and I don’t see the limit to the scaling laws.

If you’re just thinking about the raw capability of the models, and how well you can condition them to reason over increasingly complicated things, I’m sure that at some point there will be a limit. I intuitively feel like there must be.

There are some people who don’t believe that there’s a limit. The limit that human beings have on intelligence is that you’ve got so many neurons packed into your skull and about a 20-watt power envelope, and that’s the limit. Some people believe that, with AI, there is no such limit, and things will continue to scale into weird territory.

I don’t necessarily believe that. I believe we will get to some point where we’ll hit a scaling asymptote, and there will be diminishing marginal returns. It will be so expensive that we will decide it’s not worth spending that next dollar to make this thing one unit smarter, because we haven’t figured out how that translates into something useful for the people who are using the tool.

I think that point will come. I just don’t see it yet. It’s not in the viewfinder right now.

Harry Stebbings

When we think about the 3 core elements that make up efficiency in this respect—data, compute and algorithms—when we drill into data, what are your biggest observations on data efficiency? How do you think about the importance of quality of data versus quantity, and synthetic versus human data, today?

Kevin Scott

The mix of synthetic data is going up. High-quality data is becoming much more useful, especially in the post-training parts of the model-production pipeline, than low-quality data.

I think we’re clearly at the point now where, if you have the right infrastructure, super-high-quality data and super-high-quality expert human feedback, you can amplify that into the right set of tokens for training bigger and bigger models. That stuff has way more value than just the undifferentiated tokens floating around on the web.

Harry Stebbings

What questions about data and its usage do we not know the answers to that would be most helpful to know?

Kevin Scott

There’s a super-interesting thing right now that we don’t have around assessment. It’s very hard to know, quantitatively, what the incremental value of a token of data is to the quality of a model generated using that data in its training.

If you’re asking, “Okay, I think my data is super valuable, and if this gets used in a model, it’s going to make the model better,” most of those assertions are unfounded in any kind of science. The measurements we do have show that there’s a pretty big disconnect between what some people think valuable data is and how valuable it actually is to producing capabilities and models that are legitimately useful.

Most of what’s legitimately useful is models. People think of models as repositories of factual information, and they’re treating them like the world’s worst and most expensive databases. That’s not super useful. We’ve got search indexes and databases, and those things are plenty good enough for retrieving information.

What you want models to be able to do is reason over information. If you give them access to information, how well are they able to reason over a set of information to go do something that’s useful for you? You need different tokens for training a model to make it good at reasoning than you do for making it a recaller of facts.

3. In 10 Years Time: What % of Data Usage will be Synthetic

Harry Stebbings

It’s so funny you said that about reasoning, because it made me think about inference. I get really annoyed by the word “inference.” I wish we’d just delete it and call it “usage.” There’s training and there’s usage.

You’ve been very clear about the transition in emphasis and importance from training, which has had the focus over the last few years, to inference. What are we not talking about or seeing in inference that we need to spend more time thinking about?

Kevin Scott

The thing that most people miss, although the DeepSeek R1 launch a few weeks back clued everybody into it, is that we just have an incredible track record over the past handful of years—and it’s many years now—of repeated, year-over-year, mind-boggling progress in optimizing the performance of models.

The performance of inference is getting better and better and better. Over time, the models have gotten bigger and the API calls have gotten cheaper. A little bit of that is because you get maybe a 2x benefit in price performance from hardware every generation, if you’re lucky, but you get a much bigger improvement in price performance from all of the things that you’re doing in software.

There’s a ton of work happening there. The DeepSeek R1 stuff, which was good work, is a point on a line of price-performance improvement that maybe was invisible to everyone else, but not invisible to the people who are neck-deep in optimizing these systems. It’s not the last point. It just marches on.

Harry Stebbings

What was the internal sentiment when it came out? To what extent were you surprised by the public reaction?

Kevin Scott

We’ve had models more interesting than DeepSeek R1 that we didn’t even choose to launch. I was surprised at how interesting people thought it was.

They did good work, so don’t get me wrong. It was solid technical work, and it was super cool that they chose to release the thing and make it open source. It was really interesting seeing how the public reacted to what they did.

Harry Stebbings

Is there anything that you learned from how the public reacted to the release that you’ll take with you in your own releases?

Kevin Scott

Even when you’ve made it as easy and cheap as humanly possible for folks to go do something, they still have super-strong preferences about the how. We’re paying very close attention to that.

We have to give people more “how” than we have been, because I think developers want lots and lots of choice.

Harry Stebbings

What did you believe that you no longer believe, or what have you changed your mind on, in the last 12 to 24 months?

Kevin Scott

When I was a graduate student, I was a complete open-source zealot. As I’ve gotten older, I’ve become a lot more pragmatic. It’s probably more important for me to make a set of pragmatic decisions about how I’m going to build these things rather than singly optimizing for my curiosity.

Harry Stebbings

When you look forward to the next 3 to 5 years, or 3 to 10 years, how do you think about the pervasiveness of open versus closed? Which will be more dominant than the other?

Kevin Scott

I think there’s going to be lots of both. Part of it is that we should forget about AI, which is the controversial thing at the moment. It’s an area where industry structure hasn’t settled yet, and we don’t know exactly what it’s going to be.

Just pick a previous example, like search. There are a whole bunch of open-source search-engine projects out there. People who want to have a search feature in their application, or who want to build a search engine themselves, have lots of options. They can grab something open source as a starting point and stand up a product.

They can also load their data into something like Azure Cognitive Search, which is a search-as-a-service platform. Google has one, Amazon has one, and they’re readily available. Then you still have search engines like Bing and Google that are out there.

All of those things exist. All of the economics in search go to somebody who stood up gigantic infrastructure and is running a whole search business with its own feedback loop.

I think we’re probably going to have similar sorts of things happening here. For the infrastructure layer, you’re going to have lots of open-source infrastructure products, and people are going to use them in lots of different ways.

4. How Will AI Agents Evolve Over the Next Five Years

You’re also going to have a lot of people who don’t want to stand up their own infrastructure from scratch, or take an open-source project and build it out where it’s lacking for the things they need it to do. It’s good to live in a universe where you have both of those things.

Harry Stebbings

You mentioned earlier the centrality of product. Taking that into account with the current conversation, how should we think about whether chat is the right UI for the next paradigm of this product realm? OpenAI and ChatGPT have made it the default. To what extent do you think it is the right default, and how will we see that change?

Kevin Scott

I think it’s a reasonable step in the right direction. The thing I’ve been saying for a few years now is that one of the most interesting things happening with AI is that we’ve had one paradigm for using computing devices for effectively 200 years, since likely Ada Lovelace wrote the first program.

If you want a computing device to do something for you, you have to be a programmer yourself, which is a pretty high barrier to entry for a lot of people. Or you have to rely on the fact that a programmer has anticipated some need that you might have and packaged up a piece of software into an application that you’re able to run.

Those are the 2 ways you can get a computing device to do something for you—until now. The thing that changes with AI is that it can understand something you want your computing device to do for you, and it can figure out a way to make that thing happen. You don’t have to be a programmer.

It’s a profound change, because it basically means—and I don’t think this is next year, but it’s probably not going to be 10 years—that this whole notion of teams of people whose job is to anticipate a bunch of very granular user needs in some narrow space, write a bunch of code, figure out how to hang that code onto some user experience, and hope that they’ve done a good enough job anticipating the needs and designing the user interface in the right way will change.

They’ve gotten all of the code right, and they just grind away on figuring out what that feedback loop is. That’s going to change. You just aren’t going to need as much of that anymore.

What you’re going to need instead doesn’t mean that it completely goes away. You will still need all of the capabilities that these applications provide, but you’re probably going to want some kind of agent actuating those capabilities on your behalf, rather than having to do this weird impedance matching that we have right now between how a user has a set of expectations and how a product team has imagined what those expectations are.

Harry Stebbings

Is there a role in engineering or product teams today that you think, in 20 years’ time, people will look at and say, “What? You had secretaries who typed out voice-recorded notes from a doctor?”

Kevin Scott

You’re still going to have to have people who build capability infrastructure—make this thing happen in the real world, provide access to this weirdly situated repository of information, and build the capabilities that people will need.

The user interface that surfaces those capabilities will probably be agents. Product managers—I don’t believe in this one-agent-for-everything sort of theory. I think you’ll have a lot of agents.

The reason I think you’re going to have a lot of agents is because your product managers are probably going to have to be domain experts: people who deeply understand something like medicine, drug discovery, early-round venture investing, or just pick your thing.

They will have to deeply understand the idiosyncrasies of that domain, and they will have to help set up the feedback loops that help agents assisting people with those tasks do their jobs better and better.

It will be a combination of the product manager and the users of the agents teaching the agents how to be better and better at the things you’re trying to get them to assist you with.

Harry Stebbings

I often think that we overestimate adoption in the near term, or in a year or the short term, and underestimate it in the long term. When I look at the hype around agents, I share the excitement, but I question the immediate adoption—or the expectation that some of the world’s largest companies will be using agents in the next year or even 3 years.

To what extent do you think I’m right, or to what extent do you think this wave is different, given the distribution of a company like Microsoft?

Kevin Scott

I think usage always follows utility. You make useful things, and they get used a lot. Clearly, with software-development agents, we’re getting a lot of adoption right now.

We’ve gone very quickly from developers being skeptical about these tools to, “You’ll get this from my cold, dying fingers. This is one of the most essential tools in my toolkit, and I will never give it up.” The agents are becoming more and more powerful, and I even see—

Harry Stebbings

Can I ask you to what extent there’s lock-in there? When I look at them, and when I speak to people about them, you’re right: there’s user love, but everyone says, “There’s no lock-in. I’d happily switch to the next person tomorrow.” To what extent does that mean it’s valuable?

Kevin Scott

There’s no lock-in in search. You can send your next query to a different search engine than the one you’re using right now, and yet you don’t.

The reason that’s true is that it’s our job, building these agents, to grind and grind and grind, and every day try to make the agent better and better and better, and do more and more and more of value for our users. If you do that, and you do it well, they will continue to choose you.

Harry Stebbings

When you think about a 5-year time horizon, what will the interaction model look like between humans and agents?

Kevin Scott

The thing that’s missing right now with our agents is that they are conspicuously missing memory, which makes them awfully transactional. Even in the places where agents have memory, it’s a pretty limited form of memory.

5. What is the Bottleneck Today: Data, Compute or Algorithms

I think one of the things that’s going to happen, because I know lots of people are working on it right now, is that memory is going to get a lot better over the next year or so. As you’re using an agent, it will remember more and more about your past interactions with it, and it will be able to conform itself more and more to your preferences.

It will be able to do things that we do very naturally. You solve a problem once and record the solution to that problem, and then you don’t solve it from first principles over and over and over again.

Memory even gives you the ability, with these agents, to have some kind of abstraction and compositionality, where you can build up more and more powerful ways of doing things inside the agent over time because it’s remembering the past things that it’s done and learned.

I think the agents definitely will be less transactional and less session-oriented going forward. I hope we get more asynchronous things happening over the next 12 months. Right now, it’s very interactive: you go to your agent, send a prompt in, it goes and does something immediately, and gives you the response back. It’s, “Yep, I’ve done it.”

I think there’s going to be more of you dispatching your agent to go do something, and it goes and works while you’re not paying attention to it. The thing you want with agents, by the way, is that we should never lose the plot on where we’re going.

The first generation of agents are good at 5-second tasks, and the generation after that will be good at 5-minute tasks. What we’re going toward are things that you can delegate increasingly complicated tasks and increasingly substantial work to over time, the same way that you would to a coworker.

That’s how I think about the future. That’s what everybody’s going to want, and that’s where the capabilities are headed. How do you think about building product around where the future is almost certainly going to be, and what do you need to augment these systems with to allow them to do more of this thing that is ultimately what we want?

You don’t want a thing that’s just a good email summarizer. You want something that you can tell, “I get up every morning at 5:00 a.m.”

At 5:00 a.m. every morning, digest all of the email that came in overnight, draft responses to anything urgent, and show them to me while I drink my coffee. That ought to be an entirely possible, doable thing.

Harry Stebbings

Is there anything that many people think about agents that you often hear that you think is wrong?

Kevin Scott

I often think skeptics who say, “Oh, this is hard or impossible,” are probably wrong. But I’m not unique in thinking they’re wrong; there are plenty of optimists out there who think the technology is going to get more capable. It’s fine to have skeptics. I don’t know what skin in the game they have.

I have a colleague who wrote a book called There’s No Prize for Pessimism, and there really isn’t. I don’t know what prize you win for being skeptical about something if you’re not going to go do something about it.

6. The Future of Software Development

Harry Stebbings

Well, I think it’s John or Patrick Collison—I’m never quite sure which one—but likely one of the Collisons said, “Pessimists are right and optimists make money.” They’re not wrong.

We mentioned software development being one of the most widely adopted usage mechanisms that we’re seeing today. When we look forward 5 years, what percentage of net-new code do you think will be AI-created versus human-created?

Kevin Scott

95% is going to be AI-generated. I think very little of it is going to be line-by-line human-written code.

That doesn’t mean that the AI is doing the software engineering job. I think the more important and interesting part of authorship is still going to be entirely human. What does authorship mean in a world where you’re not the input master, you’re a prompt master?

It’s just raising the level of abstraction. Are you a programmer? No. We’ve accepted that for the past 35 or 40 years.

When I was a kid, I wrote my first program when I was 12. I’m 52 years old, so I’ve been doing this for 41 years already. By the time I was 12, which was in the ’80s, you were mostly writing your code in a high-level language. But the thing that runs on the machine is not a high-level language. It’s not even assembly language; it’s some machine encoding of assembly-language instructions that run on the hardware.

Nobody bemoans the fact that they’re not writing in machine code. There was a period when this was true, during the transition from assembly-language programming to high-level-language programming, when some old farts would say, “You’re not a real programmer if you don’t know how to write in assembly language. That’s the only real coding, and that’s the right way to do things.” Nobody talks about that anymore.

I think this is going to be a similar distinction. In the same way that GUI builders and things that have been around for 20 years have changed how we work—for example, when you’re designing an iPhone application in Xcode, you don’t write all of the code. You drag a whole bunch of user-experience elements around on the screen, and the system emits a huge amount of boilerplate code for you—this is, in my mind, just the same trend.

We’re raising the level of abstraction. We’re changing the interface that programmers use to communicate to the machine: “Here’s a problem that needs to be solved.”

One thing that is true is that the extraordinarily good programmers right now, even when they’re using tools at a very high level of abstraction, understand all the way down. If something’s broken, you can go into the machine code, look at the boilerplate that your development environment is generating, muck around, and figure out what’s going on.

The same will almost certainly be true when you’ve got mostly AI-generated code. The very best programmers are going to be able to say, “The thing emitted this, but something’s off. Let me spelunk down into the lower levels of abstraction.”

Harry Stebbings

Is everyone a programmer in a world where you have Bolt or Lovable, which allow you to create simple websites with pure prompts?

Kevin Scott

I think so. But it also doesn’t mean everybody is solving the same sorts of programming challenges.

Again, this is raising everyone’s level, so it makes everybody a programmer. You no longer have to get someone to make a website for you. But if you’re trying to solve the world’s hardest computational problems, I still think you’re going to need computer scientists. They’re going to use these tools insanely well to solve problems that were just harder than they could solve before.

Harry Stebbings

Will the structure of engineering teams be fundamentally different in the future?

7. The Thing That Most Excites Me in AI is Tech Debt

Kevin Scott

I think so, but maybe not in the way that people think. I’m guessing, and I’m hoping, that it will get easier for small teams to go do big things. The reason that’s important is that I think small teams are just faster than big teams. You can do a lot with 10 really great, super-motivated engineers and powerful tools.

Harry Stebbings

What would you most like to do but, because of scale, decision-making, or whatever it is, you’re not able to do? Scale is usually tough for 2 reasons in a technology company, but it does mean that sometimes you’re slower than you would like to be. Sometimes slow is necessary, but sometimes slow is a side effect of being big. Where have you been slow where you would like to be fast?

Kevin Scott

I want to be fast all the time. I want more product happening.

There are things that can’t go faster than they go because the laws of physics are attached to them. We’ve been running 1,000 miles an hour building infrastructure over the past 2.5 years since GPT-4, and we are literally going as fast as is possible to go. You just wish you could change the rate at which concrete can be poured, power grids can be augmented, and all of that sort of stuff. I wish it could go a little bit faster.

What I would love to be able to do in an ideal world, at Microsoft and everywhere else, is eliminate any space between an engineer’s ambition for what they want to do or a good idea they want to try and their ability to go try it. A lot of our internal use of AI right now is focused on figuring out how to enable that for all of our people at Microsoft.

There’s another thing, too. If you’ve ever managed any size engineering team, one of the nastiest problems you have, which has traditionally been very zero-sum, is the accumulation of tech debt.

At some point, you’re almost always confronted with a painful tradeoff: “I’ve got to get this thing out, which means I can’t quite get the technical bits of it in exactly the state that I want them to be. I’m going to launch now, and I’ll fix this thing later.” The minute you’ve done that, you’ve incurred technical debt.

Technical debt is just like financial debt: It carries interest, and you have to pay the interest payments. If you don’t pay down the tech debt, plus the interest, you’ll be in trouble at some point because it accumulates to a large extent. Then things start failing in your infrastructure.

One of the things I’m absolutely most excited about with AI is that I think we can turn this very zero-sum problem of tech-debt accumulation into something non-zero-sum, where you don’t have to make those trades in the same way that you have in the past.

We started a big research initiative at Microsoft Research about a year ago. The whole mission of the lab is to eliminate tech debt at scale using these new AI tools. It’s super exciting stuff. I’ve been leading engineering teams for 20 years now, and tech debt is just my mortal enemy.

Harry Stebbings

What have you learned from doing that program over the last year?

Kevin Scott

That the AI tools are more capable than people think they are. This is the thing in general: I think, honestly, right now there’s a bigger gap than there was even 2 years ago between what the most capable frontier models can do and what they’re being used for.

8. Quick-Fire Round

Harry Stebbings

Kevin, I could talk to you all day. I’d love to move into a quick-fire round, if that’s okay. Let’s start with a tough one: Which competitor do you most respect—Google, Anthropic, or likely Meta—and why?

Kevin Scott

If I had to pick one, maybe Anthropic.

Harry Stebbings

Out of interest, why?

Kevin Scott

I think Dario’s doing a good job.

Harry Stebbings

What was the best advice you’ve ever received?

Kevin Scott

I had a mentor one time who told me that you can imagine an individual’s or a team’s competencies on a histogram, where the bucket all the way on the left is “idiot” and the bucket all the way on the right is “genius,” with the middle buckets being mediocre or average.

Their assertion was that you could take everything that you do and everything that you’re trying to do and assign it to one of those buckets. The mistake people make is that, with great effort, you can take something and move it up 1, maybe 2, buckets to the right on the histogram.

The mistake people always make is focusing on trying to improve at the things they’re worst at. If you believe this theory, the best you’re ever going to do if you’re an idiot at something is get mediocre at it. All of the time you spend trying to get to mediocre is time you’re not spending doing the things that you’re a genius or very good at.

I think that’s very good advice because the thing about everything that’s worth doing is that you probably have to do it with a team. It’s super easy to construct a team where you complement people.

Harry Stebbings

What are you bad at that you’ve consciously decided not to get mediocre at?

Kevin Scott

I’m bad at so many things. I’m super impatient with bureaucratic things. I hate budgets and facilities and all of the mechanical parts of being an engineering leader. Bureaucratic things just bug me, and I could probably be a very mediocre bureaucrat if I wanted to be. I’m just terrible at it.

Harry Stebbings

I love that. I’m the same. Delegation is the secret to life.

9. Leadership Lessons from Satya Nadella

Likely Satya Nadella is one of the most incredible leaders of our generation. What have been your biggest lessons from working so closely with him and seeing him operate?

Kevin Scott

I think his core leadership principle is that you have to simultaneously create energy for people and produce clarity.

You really do have to make sure—and he’s very good at this—that the energy of conversations is positive and that people walk out of reviews and conversations, and anything that we’re doing, carrying energy with them that’s going to help them go do the hard thing ahead of us.

10. DeepSeek Evolution: Do We Underestimate China

At the same time, you can’t just produce a bunch of rah-rah and not clarify for folks what the most important things are.

Harry Stebbings

We mentioned DeepSeek earlier. Do we underestimate China’s ability in AI?

Kevin Scott

I don’t think we should. We should really, really, really respect the capabilities of Chinese entrepreneurs, scientists, and engineers. They are very good. If you are underestimating them, you shouldn’t.

Maybe some people did. That’s another interesting thing about the DeepSeek reaction: how surprised everyone seemed to be. “Oh my God, this is coming from China.” That shouldn’t have been surprising.

Harry Stebbings

What’s the crazy AI prediction that most people would call science fiction that you believe to be true?

Kevin Scott

It’s already the case that I think the frontier models are probably better health diagnosticians than your average GP. It’s a good thing to realize and act on as quickly as possible because we have a whole world of people who have inadequate access to high-quality health care, including my own family in rural Central Virginia, where it’s just not good.

There are a bunch of things like this where the models are already really good. We basically need the whole world to wake up to the fact that they’re good so that we can deploy this stuff. We should deploy it because the thing we really care about is the good of the public, not trying to sustain some status quo.

Harry Stebbings

Kevin, a lot of people ask you a lot of questions—team members, journalists, you name it. What question are you not often asked that you think is an important question that you should be asked?

Kevin Scott

I don’t know. Are we going fast enough?

Harry Stebbings

Do you think we’re going fast enough?

Kevin Scott

No.

Harry Stebbings

Is it possible to go much faster?

Kevin Scott

Yeah, I think so.

Harry Stebbings

How could we go faster?

Kevin Scott

I think we could go faster in a bunch of different ways. The thing that I would want in my ideal world is for us to invest super heavily in education. I would love to see every child feel as if these new tools that we’re building right now are for them, accessible to them, and expressly built for them to accomplish the things that they think are most important.

I want billions of human beings taking all of this creative energy that we all have and doing the most amazing thing with the best tools that they possibly have. I don’t want anybody feeling constrained by anything.

I would also love to make sure that, across the public and private sectors, we’re creating every incentive we possibly can to deploy these tools to produce good. Whether it’s health care, climate change, education, or whatever your thing is where we don’t think we have enough of what everybody thinks there ought to be, if I had a piece of technology that could create abundance in this thing where we currently think there’s scarcity, let us go invest in that.

Harry Stebbings

Kevin, I’ve so enjoyed talking to you. I really appreciate your tolerance with the wide range of questions and future pontifications. You’ve been fantastic, so thank you so much.

Kevin Scott

Thank you for keeping me company on my runs. This has been awesome.

Harry Stebbings

You’re very welcome.

Kevin Scott

Thank you for having me.

Kevin Scott,Microsoft CTO:对 DeepSeek 的评估,以及我们如何低估了中国人 — 文字稿与摘要 | BidClub