[BidClub_]
The a16z Show · · 61 分钟

AI 使用的经济学与 SaaS 的下一步 | a16z 上的 Benedict Evans

Erik TorenbergBenedict Evans

YouTube
TL;DR
  • 智能体编程是 AI 首个毋庸置疑实现产品市场匹配的场景,但几乎所有更广泛的市场结构问题仍没有答案。 Evans 说,它已经从“有点有用”变成了“真正改变一切”,客户几乎是在把产品从供应商手里拽走。但由于 6 个月前这件事还几乎行不通,预测 3 年后的工程团队结构、初级工程师招聘或软件职业走向,都是“疯了才会这么做”。

  • 基础模型在结构上似乎注定走向商品化的经济模式,除非其提供商能够证明持久的差异化,或向价值链上游迁移。 Evans 看不到明确的网络效应,除了砸钱之外也看不到多少差异化;客户不太可能关心一个 SaaS 产品底层由哪个模型驱动——正如他们很少会问它托管在哪家云服务商上。他有意保留余地地提出挑战:按现有逻辑,这些东西“看起来确定会成为商品”,所以“请解释为什么它们不会”。

  • 今天的 token 经济学属于过渡性的稀缺状态,并不是永久定价权的证据。 用户可以花 20 美元获得“价值 1 万美元的 tokens”,也可能意外收到 1 万美元账单,这让人想起 2009–10 年前后的移动数据。随着可能到来的 1 万亿美元–2 万亿美元资本开支,以及模型效率每年提升“100x、200x”,供给、使用、价格和 ROI 必须寻找新的均衡。

  • 移动网络的先例表明,巨量使用和基础设施投资未必能转化为有吸引力的回报。 移动流量增长约 1,500–2,000 倍;网络运营商合计年收入约 1 万亿美元、每年资本开支约 2,000 亿美元,但相关股票 20 年来基本横盘,而“所有酷东西都是别人造的”。投资者真正要问的是,模型会变成低毛利基础设施,还是获得类似操作系统的杠杆——Evans 指出,模型目前并不具备这种能力。

  • AI 很可能创造“多得多的软件”,但这并不能说明哪些既有 SaaS 公司能够存活。 更低的开发成本、此前不可能实现的功能,以及概率模型与确定性系统的新组合,都将扩大供给并加剧竞争。Evans 预计,一定比例的 SaaS 公司会被淘汰,但认为投资者无法有足够把握地识别它们,不足以支持把整个板块估值一律下调 50%。

  • 最大的机会将来自制造此前不可能存在的产品,而不只是用 AI 重做旧软件。 Evans 的例子从在图片中找到一件外套,进阶到推荐替代品,最后从用户的 Instagram 中挑出一件能改变其风格“但不要变化太大”的外套;企业系统则可能综合通话、邮件、遥测数据和分析,推荐有助于降低客户流失率的价格调整。“重要的不是把旧事做得更多,而是做一件用旧东西做不到的新事。”

  • 财务约束会在技术可行性之前遏制 AI 资本开支,而由此产生的大部分价值可能会在竞争中转化为消费者剩余。 Microsoft、Meta 和 Google 的资本开支占收入比例都将超过 50%,四大公司的指引合计约为 7,000 亿美元;Evans 说,世界根本承受不起每年 10 万亿美元的 AI 基础设施投入。即便 AI 把一份耗时 1 周的 DCF 变成 10 秒钟的任务,公司也可能在不额外收费的情况下做 50 次分析——直到今天的“魔法”变成计算机仿佛一直都会做的事情。

摘要 · 为研究而整理的核心内容

1. 编程智能体率先跨过产品市场匹配门槛

  • Evans 自上次演讲以来最大的更新是,产品战略已经不再围绕“更快地做出更大的模型”,而是发生了分化;编程智能体已经成为唯一真正实现产品市场匹配的用例:“客户正从你手里把它拽走”(The customers are pulling it out of your hands)。

  • OpenAI 看起来像是在“同时做所有事情”,几乎像是让 ChatGPT 列出 15 种在基础设施之上创造价值的路径,然后一条不落地全部推进。资本更少的 Anthropic 则集中攻克编程并让它跑通了——究竟是有意为之的战略,还是摸索出来的结果,仍不清楚。

  • 编程在方向上是可预见的,因为开发者是最早一批用户:人们拿到 PC 后做的第一件事就是编写计算机程序,而如今拿 LLM 做的第一件事——“从某种意义上说,LLM 就是计算机”——也是创造更多算力。突破发生的时间和智能体能力突然跃迁的方式,并不能确定性地预测。

  • 在编程之外,采用情况仍从两端分化:一端是硅谷用户在 Mac Studio 集群上持续运行 OpenClaw,另一端是只会说“嗯,有点用,我上周拿它做了件事”的人。更普通的企业需求正按一条条工作流落地,例如一家低毛利大宗商品公司用 LLM 预测小型生产商何时会付清发票款。

  • ChatGPT 的各项数据、模型规模、资本开支和使用量仍在增长,但 Evans 说,模型是否会出现赢家、模型能否沿价值链上移并攫取价值、它们究竟能完成多少工作,以及消费者会不会从每周使用转为每日使用,仍是未解问题。

2. 仅有 6 个月的可用技术无法揭示 3 年后的劳动力市场

  • Erik 问编程智能体对初级工程师、高级工程师、团队组织和职业发展意味着什么。Evans 直截了当地回答:“我觉得我们什么都还没学到”——6 个月前,这项技术还不是这样运作的,所有人仍在仓促摸索如何解读它。

  • 把此前分配给初级工程师的工作自动化,会让旧问题立刻变得现实:公司为什么要招聘初级员工?他们被招来究竟是做那些工作,还是做别的事情?Evans 说,这些问题现在已经是真问题,而不再只是理论问题。

  • Evans 拒绝根据聚会轶闻或当前定价扭曲得出自信预测。实践需要数年才能稳定下来;任何声称知道 3 年后软件工程职业会是什么样的人,都是“疯了才会觉得现在就能知道”。

3. 移动数据揭示,爆炸式需求可能摧毁基础设施的定价权

  • 比较采用率时,必须把累积形成的基础设施纳入考虑。Marc Andreessen 做 Netscape 时,全球只有数千万台 PC,因此不可能支撑每周 9 亿用户;后来的每一代平台都会继承此前建成的设备、网络和消费者习惯。

  • 早期平台迁移也往往从不稳定的产品,以及一小群愿意投入精力把它们调通的用户开始。把这种产品变成用户只需按一下按钮就能使用的产品,需要时间。

  • AI 当前的定价困境类似 2009–10 年前后的移动数据:有些用户收到了 5,000–10,000 美元的账单,而 AT&T 的固定费率 iPhone 套餐鼓励用户用 3G 看视频,直到网络容量不堪重负。如今,20 美元可以买到“价值 1 万美元的 tokens”,几天试验就可能产生 1 万美元账单。

  • 移动运营商最终通过流量上限、套餐、公平使用政策和限速,使成本、价格与用户感知价值对齐。此后流量增长约 1,500–2,000 倍;如今该行业年收入约 1 万亿美元,每年资本开支约 2,000 亿美元,但其股票 20 年来基本横盘。

  • 令人不安的类比是,运营商建成了改变全球的基础设施,但价值却向上游迁移。Evans 对 AI 的核心问题是,模型会成为按接近边际成本出售的商品化基础设施,还是成为“能够决定什么会被建造”的操作系统层;与 Windows 或 iOS 不同,模型目前没有可比的网络效应。

4. 模型商品化论很强,但并非定论

  • Evans 看不到让一个模型长期、从根本上优于所有其他模型的明显路径。产品可以突出不同特性,用户也可能偏好某个模型,但表面上真正的竞争杠杆主要是“你愿意花多少钱”,而不是 Instagram、YouTube 或搜索那样的网络效应。

  • 聊天机器人也是一个“奇怪且受限的 V1 界面”。大多数专业任务都需要正确的数据、配置、控制、工具和界面;会做一项工作,不等于懂得如何设计产品,正如优秀的财务顾问未必就是应该设计 TurboTax 的人。

  • Skills 和模板可以弥合部分缺口,但用户最终会超越它们,就像部门会超越 Excel 一样。模型实验室不可能构建每一个应用,正如 Microsoft 或 Apple 不可能构建每一个 Windows 或 iPhone 应用;企业买家通常也不会要求第三方 SaaS 统一使用 Claude 或 OpenAI。

  • Evans 设想,也许会有 3–6 家前沿模型提供商,加上边缘模型和开源模型,年投入规模在 2,000 亿美元–2 万亿美元之间,但具体数字并不确定。Google 的广告业务也使其面对与 OpenAI 不同的定价激励。不过,最终只剩 2 家实验室、产品价值被模型层吸收,或出现新的价值链上游杠杆,都可能推翻他的商品化判断:历史类比“不具备预测价值”。

5. AI 的关键问题正从硅谷转移出去

  • 一旦平台结构变得清晰,Evans 就会转向别处:“当你理解了某件事,知道它如何运作以及将会发生什么时,就是你该转向其他事情的时刻。”此时“一切都还未定”,就像 Android 与 iPhone 之间“开放胜过封闭”的结论,曾经也显得显而易见。

  • AI 对法律、银行、咨询、广告和专业服务金字塔的影响,越来越是行业问题。Evans 将其比作 Netflix:技术让这家公司成为可能,但真正重要的决定——节目、人才、奖项、电影和体育——最终成了洛杉矶的问题,而不是旧金山的问题。

  • 与早期平台迁移不同,生成式 AI 没有可知的物理边界。Evans 说,录音结束后,人们可能看一眼手机,就发现某个新模型因为意外突破而只按“原价的 2%”收费,尽管他认为这种情况不太可能。

  • 其他未决问题包括,旧模型、开源模型或端侧模型是否会变得足够好,以及多少算力能够迁移到设备上。由于编程目前是唯一明确实现产品市场匹配的领域,另一个问题是,下一个突破会出现在法律、银行,还是其他领域。

  • Evans 提到,run rate 已从约 90 亿美元升至 470 亿美元,随后强调,所说的增长全部来自软件:“这全都是软件,对吧?”

6. 自动化的真正回报,在于释放原本不存在的需求

  • Evans 为每个行业提出几项检验:成本下降后,是用更低成本获得同样产出、以同样成本获得更多产出,还是让产出和支出都继续增加——这就是杰文斯悖论所问的问题。它是否消除了进入壁垒、释放了某种商业模式,或让此前无法想象的活动变得可行?

  • 拥有一台印刷机曾是报纸的护城河;去掉这项成本会改变竞争。更进一步,蒸汽机让火车成为可能,而再多的马也无法复制它;Spotify 则让人们每月花 15 美元就能获得所有录制音乐——这在此前从物理上不可能实现。

  • 即使横向预测正确,也可能掩盖行业特有的结果。互联网摧毁了实体分发的价值,却重创报纸,对电影制片厂的改变小得多。“要看情况”在这里不是推诿,而是核心机制。

  • 在新平台上重做旧产品,是与之相伴的谬误:Google Docs 在 Web 上复刻了 Office,可能拿下约 20% 的份额,但那不是最深层的机会。卓越的创业公司“填补宇宙中的一个空洞”(fills a hole in the universe)——它们往往解决的是行业自身都不知道存在的问题,这正是通用模型无法简单“把整件事做完”的原因。

7. 商业场景展示从相关性到高阶意图的跃迁

  • 广告市场规模约 1 万亿美元,零售市场规模 25 万亿美元,因此,更好地理解产品具有重大的经济意义。Google、Meta 和 Amazon 传统上知道 SKU、出版商元数据和共同购买关系,却不知道人们为什么购买某件东西;所以才会把一次购买马桶圈理解成建立一套收藏的起点。

  • Evans 对消费者需求的递进描述从一张外套照片开始:识别它并说出哪里可以买到;然后推荐 10 件不同价格、列出优缺点的相似外套;最后查看 Instagram,推荐一件能改变用户风格“但不要变化太大”的冬季外套。3 年前,最后一步还是科幻;如今看起来已经做得出来。

  • 企业版则是把录制的 Zoom 通话、Salesforce 中的邮件、产品遥测和分析数据结合起来,回答:“我们应该如何调整价格,才能改善客户流失?”这比给愤怒来电者做情绪评分高出一个抽象层级;而 Google 和 Facebook 转化率、广告指标的上升,已经反映出 AI 正进入推荐和预测系统。

8. AI 将让软件倍增,同时重排 SaaS 格局

  • 软件会变得更便宜、更快构建,此前不可能实现的功能将变得可用,竞争也会加剧。利润率和定价结构仍未确定:按结果收费听起来很有吸引力,但把企业软件的每一次按钮点击都对应到损益表,往往并不现实。

  • Evans 将今天的企业软件版图分为几类:SAP、Workday 这类大型横向系统;CRM、人力资本管理和薪资软件;垂直应用;以及由邮件、电子表格和共享文件系统临时拼成的中间地带。一家大型美国公司可能拥有 300–400 个 SaaS 应用,另外还有约 1,000 个在本地运行的自建或采购应用。

  • 规模决定选择:PwC 大规模招聘应届毕业生,值得配备专用软件;一家只招聘 5 名应届毕业生的公司则可以使用邮件和共享 Google Sheet。LLM 又增加了更多选项——直接完成任务、扩展 Salesforce、增强垂直应用,或生成一个定制工具,而它可能成为明天那张来历不明的 10 MB 电子表格。

  • 模型可以位于底层,成为受控的 Salesforce 功能;也可以位于顶层,综合 Salesforce、Workday、邮件和分析数据。答案可能是两者兼有:概率型软件负责解释,确定性数据库保留可靠状态。结论是“更多软件——多得多的软件”(more software—like, way more software),尽管会有多少既有公司被淘汰、占比多少仍不可知。

9. 隐性工作流与财务约束限制这场转型

  • 软件公司和战略咨询顾问都会观察企业如何运作,并提出更好的工作流;前者把它编码进软件,后者则把它落实为流程、组织架构、培训和激励机制。大量实际运作并未被记录,包括员工奖金制度奖励他们无视既定战略的情况。

  • 这正是 Bain、BCG 和 McKinsey 仍有价值的原因:它们可以跨越组织边界,发现工作实际如何发生,并提供一个管理层可以采用、也可以甩锅的外部答案。这类隐性知识很难封装进一个只会说“做一份 PowerPoint”的 Claude skill。

  • 资本面临更硬的边界。Microsoft、Meta 和 Google 的资本开支都将超过收入的 50%,而电信行业为 15%–20%;四大公司的指引合计 7,000 亿美元,相比之下,电信行业约 3,000 亿美元,石油和天然气行业为 7,000 亿美元–1 万亿美元。“我们不可能每年花 10 万亿美元”,因为这样的资本根本不存在,而一款前沿模型可能只在“3–6 个月、6–9 个月,或者你想怎么说都行”的窗口内保持相关性。

  • ROI 可能以消费者剩余,而不是更高利润的形式出现:如果一份 DCF 从 1 周缩短到 10 秒,分析师可能在不额外收费的情况下做 50 份;咨询顾问可能在相同收费和成本基础上完成 5 倍的分析。Evans 最后引用 IBM 在 1950 年代初的承诺:电子计算器“能给你增加 150 名工程师”;如今的魔法最终会变得不可见——“计算机一直都能做到这些”(Computers have always done that.)。

Erik Torenberg

Benedict, welcome back to the a16z podcast.

Benedict Evans

Thank you.

Erik Torenberg

Last time you were here, we were discussing the first iteration of your presentation, “AI eats the world.” You wrote it almost a year and a half ago at this point. You always begin your presentation with your big questions, but before getting into the questions going forward, I want you to reflect on what we’ve learned since you originally made the presentation. What’s played out? Let’s reflect on what’s changed in the last year.

Benedict Evans

I think we have much more of a sense of diverging product strategy. We have much more of a sense of competitive tension that goes beyond just, “Make a bigger model faster with more compute.” We’ve had several iterations of OpenAI’s strategy, in particular, from sort of everything all at once to, “Oops, no, maybe we should double down on coding.”

Clearly, agentic coding started working, and so all the focus in tech has narrowed massively onto that as something that has absolute product-market fit, in the sense that customers are pulling it out of your hands. And, of course, that comes with the supply crunch around capacity and the price imbalance—the imbalance of supply, demand, capacity, capex, and pricing that we see at the moment.

So that’s the big shift: We went from, “This is kind of working and kind of exciting, but we’re not quite sure what we’re going to do with it,” to, “Right, it works for coding. Will it work for anything else?” Almost certainly, but that’s what’s working right now, and so we’ve got this much narrower focus.

Otherwise, the ChatGPT numbers keep coming up, the models keep getting bigger, the capex keeps growing, the usage keeps growing, and people are using this more. But most of the fundamental questions you might have had 2 or 3 years ago still don’t really have answers. We don’t know if there’ll be a winner in the models. We don’t know if they can capture value up the stack. We don’t know how much the models can do. We don’t see a way that consumers will use this daily rather than weekly with the technology we have right now. So all of those questions are still open.

Erik Torenberg

And just on coding, could we have figured—could we have foreseen—that that would have been the use case that really would have taken off? What’s a reflection on that?

Benedict Evans

Well, deterministically, you could have said, “Look, who’s messing about with this stuff? Software developers. What are software developers going to try and make work? Software development.” So, at a simplistic, naive level, yes, the stuff that should work first is software development.

I often compare this moment to the internet in 1997–98, but it’s also like PCs in the early ’80s or the late ’70s. It’s incredibly exciting, but it’s not quite clear what it’s for, and it doesn’t quite work yet. Clearly, the first thing that people did with PCs was make computers. And the first thing that people are doing with LLMs—in a sense, LLMs are computers—is make more compute. So that’s not terribly surprising.

I think the shift at the beginning of this year was that agentic coding went from being kind of useful to really changing everything. I’m not sure you could have—clearly, there were people who were going to say, “Well, this is going to be able to do absolutely anything,” and so they’ll say, “Yes, look, I told you.” But I don’t think anyone could have deterministically predicted exactly when that was going to happen, or that coding would be the first thing where it worked.

Erik Torenberg

And what have we learned about—say more about what this means for engineers, junior engineers, senior engineers, the jobs discussion, how teams are organized, et cetera. What have we learned so far?

Benedict Evans

I don’t think we’ve learned anything. This didn’t work 6 months ago.

Erik Torenberg

Yeah.

Benedict Evans

Everyone is scrambling around trying to work out what it means. You can get very into the noise and the detail: “What did somebody say at a party yesterday? Oh, my God, that’s how it’s all going to work.”

It’s going to take a couple of years for this all to settle down, if nothing else because of the pricing. We’ve got this enormous crunch between the demand and the supply, and hence the pricing. So we don’t know what a team is going to look like.

I think people are asking new questions around the obvious one: Do you hire junior people, and if so, what are they doing? Why were you hiring junior people in the past? Were you actually hiring them to do the thing that they did, or were you hiring them to do something else? If you automate away a class of stuff that used to get done by people, what will happen?

That becomes much more real now in software development because you actually are automating a bunch of stuff that used to be done by people. So those questions are real now rather than theoretical. But I don’t think anybody can possibly say they know what the market structure is going to look like, or what the career of a software engineer is going to be in 3 years’ time. I think you’d be insane to think that you could know that yet.

Erik Torenberg

Yeah. Talk about OpenAI. Talk about what’s most surprised you, or how have you made sense of their strategy development and the questions that they have going forward?

Benedict Evans

Well, it’s always been such a tranquil, drama-free environment. Obviously, they’ve had the issue with Fidji Simo having to take medical leave, which shuffled things up a bit.

Clearly, in the last quarter of last year, their question was, “Right, well, the models are the models, but what else? And how do we get people to do other stuff with this?” “Ask ChatGPT for 15 ideas for what we could do to build value on top of infrastructure, and then we’ll do all of them.” That’s almost literally what it looked like.

And then Anthropic, having raised less capital, said, “No, we’re going to focus on coding.” And they got coding working. Whether that was a deliberate strategy or they stumbled into it is for other people to say, but clearly that worked.

But the question still remains: The stuff that’s working right now is software development and some things in some other fields. Then there are a lot of people who are excited about using this around the edges and using it for some things. There’s clearly a very wide spread between people in the Valley who bought a cluster of Mac Studios and are running OpenClaw all day versus those other 40% of people who say, “Yeah, it’s kind of useful. I used it last week for something.”

I’m like, how do you bridge that? I don’t think that question is solved. Software is a place where people have really, really jumped over that bridge.

There are a lot of other places where people are scratching their heads and using it up to a point. And then there are a lot of places where corporations are using it to automate some specific back-office process. You’re not asking the user to work out what they do with the new tool; instead, you’re saying, “Okay, here’s a problem that we can solve.”

I go and talk to companies outside America and outside tech, and talk to consultants and investors. They’re looking at those one-at-a-time point solutions. I was speaking a couple of days ago to a commodities company, and they want to use LLMs to get better predictions on their cash flow because they deal with all sorts of small producers and don’t necessarily know when their invoices are going to get paid. It’s a very low-margin business, so that’s a big deal, and they want to use LLMs to get better cash-flow forecasting.

That’s a very different thing from going to ChatGPT or Claude and saying, “Hey, give me a summary of my meetings this week.”

Erik Torenberg

Yeah. How did this compare with mobile or other sorts of platforms in terms of early user adoption, in terms of weekly or daily users?

Benedict Evans

I think there are a bunch of different ways to answer this. One of them is that we’re always standing on the shoulders of giants, and growth is always compounding.

So mobile didn't need to wait for the internet or cellular networks, like mobile data. Mobile internet did need to wait for cellular data, but it didn't need to wait for the internet to happen, and the internet didn't need to wait for PCs, and PCs didn't need to wait for consumer electronics and semiconductors, and so on. So you've always got this accelerating adoption. When your boss—my old boss, Marc Andreessen—was working on Netscape, there were double-digit millions of PCs on the entire planet. So, no, you couldn't have 900 million weekly active users because there weren't 900 million PCs.

So there's always that acceleration. That's one point. I think the second point is that, at the early stage of any of these shifts, it's not really clear how it's going to work, and nothing works. I'm just about old enough to remember this. I'm not sure how old you are, but anyone in their 30s doesn't really remember a time when it was completely normal that you'd be working, everything on the screen would just freeze, and you'd have to crawl under your desk and unplug the computer, and then pray that some of what you'd done in the last hour might still be there.

That just doesn't happen anymore. Go back to the '80s: you bought a sound card. Well, that's $300. You want to have sound on your computer? Okay, that's $300, and it's like a weekend to make that work. I remember trying to get this stuff to work.

It's the same thing with the internet. You've got to get a floppy disk with TCP/IP on it, and it's slow, and none of the stuff that you need to do existed. It's the same with mobile. We're kind of at that stage, and of course it's not clear which of these things are going to work. Is a browser going to work? Is it going to be this? Is it going to be that? How's this all going to fit together?

There's a gap between what's incredibly exciting and the small number of people who are willing to put the work in to get something to work, and just turning that into a thing where you can press a button in real hands.

I think the third point here is a much more tangible observation: the pricing crunch that we've already mentioned looks to me a lot like what happened with mobile data in 2009–10, where suddenly people got bills for, like, $5,000 or $10,000 of data. On the one side, and on the other hand, if you had flat-rate data—which is kind of what happened in the U.S. with the iPhone—AT&T, or Cingular, launched the iPhone with flat-rate data, and then everybody bought iPhones and started using 3G. People started watching YouTube, and the whole network went down because they just didn't have the capacity to do that.

It's funny: there are still people in tech who don't understand that cellular networks have marginal costs. They have to add more capacity, and that costs more money. The networks had to scramble to get the cost curve aligned with the pricing system, aligned with the underlying cost and aligned with perceived value, which they kind of did with data caps, bundles, fair use, throttling, and so on.

The other side of that comparison—and this is exactly where you see it now—is that, on the one hand, you're paying $20 a month and getting $10,000 worth of tokens, and on the other hand, you mess about for a couple of days and get a $10,000 bill, and you're like, “What the hell is this?” You literally see these stories now, which is exactly what happened in 2009, 2010. It's also what happened in 2001, 2002, and 2003 with GPRS.

But I think the other interesting part of that analogy, or that comparison, is that since then, mobile data traffic has risen by something like 1,500 to 2,000 times. The mobile networks collectively have revenue of about $1 trillion, and they spend about $200 billion a year on capex. Their stocks have been flat for 20 years, and all the cool stuff got built by somebody else.

They kind of all thought that they were going to build all the cool stuff. I worked for a phone company that had a banking license because they thought they would do mobile banking, which now seems absolutely insane. But that's kind of the point: they built this amazing piece of global, incredibly sophisticated, very expensive infrastructure, with enormous growth in use all the time, and it changed all of our lives. We all pay for it, and they didn't make any money from it because all the value moved up the stack.

This is, of course, the absolutely central question for LLMs: can the model do the whole thing, or do you have to have 300 apps built on top of it? Can you just go to the model and say, “Do my taxes for me,” or do you need to have a tax thing that might use some AI in 10 different ways inside it? If not, then what does it mean to be a foundation-model provider?

Is this just commodity infrastructure that gets sold at marginal cost? That somehow seems to be a very difficult concept for people to grasp right now, because you can sell all the tokens you can make, so you can price it at ROI. But over the next couple of years, we've got $1 trillion to $2 trillion of capex coming down the pipe, and the models get 100x or 200x more efficient every year. Then there are new models, and will the models use more tokens or fewer tokens? But we'll get to a different equilibrium. Why would that equilibrium be one where the model companies have pricing power when the models are all kind of the same, doing kind of the same thing with the same chips? Why would they have pricing power?

I think that's the long answer to your question. You go back and look over time: chip companies didn't capture the value, ISPs didn't capture the value, and mobile network operators didn't capture the value. Windows and iOS did, but they were doing something else. They had all these levers to go up the stack, and of course they had network effects, which models don't have.

So that's sort of the question: do they end up like the infrastructure layers, or do they end up like the operating-system layers and capture value, and actually get to decide what gets built? Or do they end up—I mean, the irony of this is Netscape, where Marc Andreessen famously said that he was going to turn Windows into a set of badly debugged device drivers, and Microsoft kind of crowbarred its way into the market. But it turned out that web browsers weren't the point, because all the value was somewhere else.

I think that's the more—or the kind of—a swirling mass of questions about how this settles out, which comes back right through to all my answers to your question. Some of this stuff you know, but you don't know how it's going to work.

Erik Torenberg

Yeah. It's unclear whether it looks more like the internet or software, where a lot of the value—or just better margins—happen at the application layer, or more like the cloud, where the value sort of existed at the hardware layer. Right now, so far, it seems like NVIDIA—and going up, it seems like they have better margins and are capturing a lot of the value. But it's unclear if that will remain the same, or if there will be applications, if we'll look more like the internet. How would you even begin to predict the answer to this?

Benedict Evans

Well, there are 2 answers to this. There are all these sorts of quotes about how history works, and my favorite one is, “History teaches us nothing except that something will happen.” You can always explain afterward why it worked out like that, but it generally wasn't obvious at the time.

In particular, I remember that about 15 years ago, a lot of really, really clever people in tech looked at the iPhone and Android and said, “This is open versus closed again, and Android is going to crush the iPhone,” which of course isn't what happened. I can go and explain why, but all of these comparisons are useful, and none of them are predictive. It's always obvious in hindsight.

I've done a couple of podcasts recently, and I've published this presentation. There's a class of comment on this stuff that says, “Benedict, you're not doing your job. You're supposed to tell us what's going to happen. You're supposed to make predictions, and all you seem to do is say, ‘Well, we don't know.’” There are 2 problems with that. One is that there are a bunch of places where I actually do say, “I don't think this is going to work. I think it's going to work like that.” I don't think foundation models are a product. I don't think a chatbot is a product. I think the value will be further up.

But the other side of this is that, when you're at this stage in the cycle, there are many paths, and you don't know which of those paths it's going to take. To try to say, “Well, I think it's going to be that one,” you might be right, but you do have to be conscious of how uncertain this is and how many different paths it could take. That's the nature of this part of the cycle: all bets are open.

We get to the point where the S-curve kind of curves up and it narrows in. There was a moment when Windows Phone might have worked. In hindsight, no, it probably wasn't going to work.

But there was a moment when it wasn't clear how mobile was going to work. And there's a moment when it was clear: right, this is what's happening. Now we move on to the next question.

One of the characteristics of tech is that the moment you understand something—how it works and what's going to happen—is the moment you should move on to something else. You should always be looking for the places where we don't know what the answers are, because I haven't updated my Apple spreadsheet in 5 years because we know what happened. I don't care what next year's iPhone looks like. I don't pay attention to Apple's market share in China; it happened. Next question.

Erik Torenberg

You mentioned the prediction that you don't think foundation models are the product; you think it'll move up. Explain the reasoning there a bit and what that could look like.

Benedict Evans

I think there are 3 or 4 building blocks you can put on the table. One of them is that it's not clear how you could build a model that was fundamentally better than everybody else's model in some sort of sustainable, differentiated way. There doesn't seem to be a network effect. There don't seem to be levers you can pull, or a strategy like the one Instagram has, or YouTube, or Google Search. We don't see an equivalent of that for LLMs.

Now, you have different emphases. Maybe this one's better than that one; maybe you like this one more than that one. But there doesn't seem to be a fundamental differentiation, a fundamental competitive difference between the models, except your willingness to spend money.

The second problem is that the chatbot itself is a kind of weird, limited V1 UI. There are some things, some people, and some kinds of tasks where it works really well, but for most of the others, you need a bunch of other stuff. You need tooling, it needs to be set up right, it needs to have the right data, it needs to be configured and controlled, and it needs to have the right user interface.

People need to have sat down and thought about how this should work, because generally, people who are good at using the tool and doing the job that needs the tool are not the same people who are good at deciding what the tool should be. People who are really, really good at designing print publications are not the people who should create and design the software. That's a different set of skills. People who are really, really good at giving financial advice are not the right people to design TurboTax. Those are different people with different skills.

You're kind of groping around in the middle of this. You now have Claude for this, Claude for that, and skills and so on. To me, one question is, who builds the skill? Another question is, that seems to be a bit like what you get if you do File > New in Excel. These are templates, and they'll take you so far, but at a certain point, people outgrow the templates.

There's a slide in my presentation that is a quote from somebody who said to me on Twitter years ago. They said they were a consultant, and half of the jobs were telling people who used Excel to use a database, and the other half were telling people who used a database to use Excel. So there's this kind of fuzzy, swirly place of: do you need dedicated software? Do you need horizontal software? Do you need vertical software?

You can't just do everything in Excel. We've all seen the department that runs along on a 10-megabyte file. I run my business in Numbers, on a spreadsheet, but there's a certain point where you outgrow that.

Following that on, can the model labs build all of that? Of course not, no more than Microsoft or Apple could build every Windows app or every iPhone app. So then, do the model labs have leverage? Are they Windows? Are they iOS?

Again, is there a network effect? If you're a law firm right now and you buy a piece of software—you see all the pieces of enterprise software that a16z is invested in—how often does the law firm, or the manufacturing company, or the bank say, “Oh, does this use Claude or does it use OpenAI? We standardize on Claude.”

Well, no, that's not how it works, any more than it worked like that for cloud. You didn't say, “Our company standardized on AWS.” You don't even know what company, what cloud, that SaaS product runs on. That's the whole point: it's abstracted away. It's not your problem.

The foundation models seem to look more like that. They look more like the hyperscalers in that sense. They might have competitive advantages, but further up the stack, you don't have leverage, you don't have a network effect, and you don't have control.

That prompts me, incidentally, to say that maybe the right comparison here is with semiconductors, where with each generation it just gets more expensive, and so you have fewer players. All of that taken together—the models are kind of commodities, the chatbot isn't the right UI or the right product, and the companies aren't going to be able to build all of that stuff themselves—means that they're low-level infrastructure.

So then, do they have pricing power? You're going to have, pick a number, 3 to 6 companies making a frontier model, spending—no one knows, no one honest knows—something between $200 billion and $2 trillion a year on building these models. Plus, there'll be a bunch of edge models and a bunch of open source.

Where is this going to settle down? It might be half a dozen companies that are all competing to sell this stuff. Where is the price discipline going to come from, particularly when some of them have whole other business models as well? Google sells ads, so it has a different attitude to pricing than OpenAI.

I think the challenge here is that there's a difference between where we are right now and where this should end up, which is kind of a first-year economics student conversation. Right now, we're in this period of extreme disequilibrium of supply and demand, price, capex, and capacity.

Just because demand for tokens is infinite, that doesn't mean you can't get to a different price equilibrium, because of course that's what happened with mobile data. Demand for bits is infinite. It's grown 1,500x in the last 15 years, but you still get your supply-and-demand price equilibrium, and you still get a murderous price war between telcos in most parts of the world.

Fundamentally, you're selling a commodity to people who will swap back and forth, and developers will also swap back and forth. I'm happy to say that this might be completely wrong. It may be that we get to a world in which there are only 2 companies that can make an LLM and they have pricing power, or we get to a world in which most of what we do gets subsumed into the model, or they have leverage further up the stack.

It's my point about iOS versus Android. Just because you can say, “It worked like that the last 3 times,” that doesn't prove what's going to happen this time. But it does mean that you should ask the questions, and you should certainly pay attention to them.

I'll just say, as a primary observation, that this situation right now is transitory. We're in this extreme scarcity, and then we have a pricing system, we have a free market, and we have a surge of capex—like $1 trillion of capex. Those multiples are going to move around. And then what?

Erik Torenberg

Going back to your point earlier—it's a good segue to your point that, hey, we know Apple's Apple—what are some of the next questions that you're most focused on, or that we should be paying most attention to?

Benedict Evans

One way to answer that is that some of the questions we've already talked about are: how far do the models go? Can the models differentiate? I think another question is, at what point do we see more and more classes of use cases where the models are good enough and we don't need the most expensive, fastest, biggest, heaviest model in the cloud?

You can use an older model, you can use an open-source model, or you can have a model running on-device. Obviously, this is what Apple is going to be talking about in a couple of weeks. How much can you push onto the device, where the compute is free—or free to you, anyway—and doesn't have marginal cost for the developer?

Another classic question is that it's almost as if the question has moved out of technology. If you're looking at a law firm, a consultancy, or basically anyone in professional services, where you traditionally have this pyramid structure and you can automate a great chunk of what the people at the bottom of the pyramid were doing, what happens?

The only thing you can say there is that if you've never worked at a law firm, or never worked at Bain, BCG, or McKinsey, you're probably not going to have a good idea of how this works. You probably don't really know what it is that all those associates are doing, and you also don't really know what it is that the client is paying for, or how those things get reconfigured.

What does AI mean for finance, both for that internal hiring structure and the kind of products you can create and the margin structure? What does it mean for consultants? What does it mean for the Big 4, for the Big 3, for Accenture, for big law firms, and for advertising? You can probably answer some of those questions.

But if you’re not in that industry, you don’t really know the answers. This reminds me a lot of something I wrote when I was at a16z, which I called “Content Isn’t King.” I also wrote something that said “Netflix Isn’t a Tech Company.”

The point I was getting at is that if you looked at Netflix, this whole thing is enabled by stuff the tech industry built. But all the questions for Netflix are LA questions: What shows? How many shows? What kind of shows? What should you pay the talent?

Should you aim for awards? Should you do movies? Should you buy sports? What kind of sports? These are all Los Angeles questions. These are not San Francisco questions. No one in San Francisco even knows what the right questions are.

They’re media industry questions. This was kind of my point: all the questions that matter to Netflix have become media industry questions. This is obviously the great tension point about Tesla. Is it a car company? Is it a technology company?

What I’m getting at is that “What does this stuff mean for law?” is a question for lawyers as much as it is for people who understand a lot about law firms—how they actually work, what they’re actually doing, and what the clients are actually buying from them. The same thing goes for, “What does generative video mean for Hollywood?”

Ben Affleck probably knows a lot more about this than I do. He built a company and sold it for around $100 million, so obviously he does. That’s kind of a second question: the questions move outside of AI, and they become half-AI questions and half-something-else questions.

The third level—and I probably should have said this earlier—is that the way all of this is fundamentally different from previous platform shifts is that with 3G, or the iPhone, or the web, you didn’t know what was going to happen next, but you knew the physical limits.

In 1995, you knew that telcos weren’t going to give everybody in the world broadband the next week, and you knew that everyone in the world wasn’t going to go out and buy a PC because a PC cost around $3,000. So you knew the basic physical limits of what could and couldn’t happen.

With generative AI, we don’t know those things. We might look at our phones when we get off this recording and see a push notification saying that OpenAI’s new model is out and it’s 2% of the price because they worked something out. I don’t think that’s very likely at this point, but we don’t know those kinds of things.

How much bigger will the models get? How much better, faster, or cheaper? How much pickup will there be? In what ways will the characteristics of the models change? We don’t know.

That is different from previous platform shifts, where you did know the fundamental constraints. That will spin off more questions. In a sense, this is something I pointed to earlier: the place that has product-market fit right now is coding. Nothing else has equivalent product-market fit right now.

I think I’m pretty safe in saying that when swap has gone from whatever it was—$9 billion in run rate at the end of last year—to a $47 billion run rate, that’s all software, isn’t it? So what happens when someone else in some other field gets something working?

Erik Torenberg

Yeah.

Benedict Evans

Which field? Law, banking? I don’t know where.

Erik Torenberg

If you had to guess, what are the use cases outside of coding that could potentially yield daily activity?

Benedict Evans

So, there’s a presentation that I published a couple of weeks ago. There are 3 sections. One of them is about capital, CapEx, infrastructure, foundation models, and differentiation, which is the stuff we talked about.

The second is: How would you build software with this? What does this do for the software industry? What would software look like, and what happens to the margins and the companies and everything else?

The third section I called “Change,” which gets to this point. I opened it with what appears to upset a certain category of person, where I used the Yogi Berra quote that “predictions are hard, especially about the future.”

Benedict Evans

I think there’s a sort of backtest point here: Imagine asking these kinds of questions about the internet in 1997. What would you have gotten? What would you not have gotten?

One way you can look at this is to say, “This is automation.” It makes a class of things that people used to do, which couldn’t be automated, something you can now automate. Then you can ask, “What does that mean?”

I proposed 3 or 4 sorts of buttons to press. The first one is: Is this just price elasticity, which is really what Jevons’s paradox is? If you make it cheaper to do stuff, do you do the same amount of stuff for less money? Do you do more for the same money? Or do you do more for more money?

Because it becomes so much cheaper, was there something that you couldn’t do before that now becomes cheap? Was there something that was expensive and served as a barrier to entry, like owning a printing press as a newspaper?

Is there something that was a cost-based barrier to entry that now goes away? Is there something that gets unlocked in your business model or in your competitive space because this thing became cheap?

The final question would be: What stuff was just completely impossible—totally cost-prohibitive—so that nobody even thought about it, and now that’s within reach?

The example I used to give here was that steam engines make trains possible. It wouldn’t matter how many horses you bought; you couldn’t have a train or an express train.

A more contemporary example would be to point to something like YouTube, or indeed to Spotify. Spotify says that when you look at the last 25 years of the music business, the first half is what happens if you don’t have to buy a $15 CD to get that track.

But the second half is: What if $15 a month gets you all the music there is? That was something that was just completely impossible.

The problem with making predictions like this is that, on the one hand, you’re going to say stuff that’s clever and obvious, but you don’t actually know what it’s going to mean industry by industry.

If we’d been back in the late 1990s and said, “We know the internet will destroy the value of physical distribution,” it turned out that meant completely different things for newspapers and movie studios. Newspapers got completely screwed by this, and movie studios kind of haven’t really changed very much.

So again, it depends. The other part of this is that there are some places where you can ask more useful questions. The one that intrigues me is: How does this change advertising, e-commerce, brands, marketing, and everything that we buy?

Advertising is $1 trillion, and retail is $25 trillion, so it’s a reasonable-sized TAM. The thing that I always used to think about was that Google, Meta, and Amazon don’t really know what the product is.

They know it’s a SKU. They know what the publisher typed into the metadata field, and they know that people who bought this also bought that. But they don’t know why, and they don’t really know what those things are.

That’s why you get these jokes about, “Hey, Amazon, I bought a toilet seat cover. I’m not collecting toilet seats.” It doesn’t really know what a toilet seat is, and it doesn’t know that people don’t buy 2. Actually, it should know that. That should be frequency analysis, but it doesn’t.

With an LLM, in principle, you would know what those things are, why people buy them, and what other things people buy. Obviously, “know” is a difficult, tricky term to use. What do you mean when you say “know”?

But at a minimum, it’s a very different level of statistical correlation from what an AI system would be able to do. That’s why you see the ad numbers and the conversion rates shooting up every quarter at Google and Facebook, because they’re rolling all of this into their ad systems, recommendation engines, and prediction algorithms.

You get shown more stuff that you would like, and the ads that you’re seeing are more likely to be things you’d like to buy. So they have this enormous, sudden acceleration in their ad revenue.

All of which is to say, you look at how these systems work, and right now they say, “People who bought that could buy this.” You should now be able to say, “Here’s a picture of a coat. What is it? Where can I buy that?”

Ten years ago, that certainly wouldn’t have worked. Five years ago, it probably wouldn’t have worked. Now, that should work.

Then you can say, “Suggest 10 other coats like that at different prices, tell me where I can buy them, and suggest the pros and cons of each one.” You’ll kind of get that, too.

Then you can push one step further and say, “Look at my Instagram and suggest a winter coat I should buy that will change my look, but not too much.”

Again, 3 years ago, that would have been total science fiction. Now you think, “Yeah, you could probably build something like that.”

That would kind of work. And those kinds of shifts in what the computer knows, what it can automate, and what suggestions it can make—going back right to the beginning—whenever you get a new technology, you start by doing the old thing but more: more spreadsheets, more PowerPoints, more email, better email. But the important stuff is not doing the old thing but more. It's doing something new that you couldn't have done with the old thing. It's a pretty banal observation, but we kind of lose sight of it.

And so what are the new things that you can only do with this, as opposed to automating the old stuff? I mean, the enterprise version of this would be: you've got all Zoom calls with clients recorded, all the flows of emails in and out of Salesforce, and you can see all of the telemetry, metrics, and analytics of how people use our product. So how should we change our prices to improve our churn? Again, that's something that an LLM might be able to do, which is very different from saying, “Do sentiment analysis on calls into the call center and tell me which customers are angry.” You get multiple shifts in the layer of abstraction around what analysis you can do. And, of course, that then creates new companies and destroys old companies, and creates new businesses and everything else.

But again, we're in 1997 and I'm trying to predict Uber and Airbnb. If I could actually do that, there's a general point here: if we could actually predict what was going to happen, we'd live in a parallel universe. VCs would have—it wouldn't be a 1-in-10 hit rate; it would be a 10-out-of-10 hit rate.

Erik Torenberg

It seems like one of the questions we're now asking is: what was unreasonably expensive to do before that is now possible? Maybe, I don't know, something crazy like rebuilding YouTube from scratch or rewriting Linux from scratch?

Benedict Evans

Yeah, it's funny. The other paired fallacy, of course, is that the new thing comes along and says, “Well, we're going to build the old thing with the new thing.” Of course we're going to build Office with open source; we're going to rebuild it on the web. And, you know, it turns out—guess what? Look at Google Docs. It's got like 20% of the market because that's actually not the point. What's interesting is to do something else, to do something new. It's to shift that level of abstraction and to kind of spot problems that have never existed.

I mean, the experience you get sitting in pitches all day at a venture firm is that there's some stuff where you think, “Well, that sounds kind of useful,” and stuff where you think, “I'm not quite sure why that would work.” But there are some things that kind of fill a hole in the universe, and as soon as somebody explains it to you, you think, “Wow, why did nobody do that before? Why did no one see that that thing existed?” And that's part of the fun of looking at startups. People will suddenly work out that that problem existed, and no one—including the people who have that problem—realized that problem existed. And then they'll go out and make a thing to solve it.

This is also, incidentally, going back to an earlier point: this is why I don't see the—I think that this is the problem with the idea that the model will do the whole thing. If you go back and think about all the pitches you've seen since you joined a16z, how many of them were things where people in the industry knew that was a problem? Quite often, the answer is actually no. No one in the industry thought that was a problem. It actually took like 2 years to explain to them and persuade them that that problem existed at all, and that this new thing would fix that for them. And that's kind of the problem with the idea that a middle manager in finance is going to use this tool to solve this big global industry problem. No, because no one knew that industry problem was there, let alone could work out the right way to build a tool to solve it.

Erik Torenberg

Does this imply a less consolidated SaaS environment than before AI? Maybe less bundling or single behemoths like the Microsoft enterprise? Gosh, way to bring me back down to earth. Is the SaaS industry going to be less consolidated, Benedict? That's all great, but tell us about the stocks. What are the kind of building blocks that we can put down here?

Benedict Evans

So, obviously, it's going to be way cheaper and quicker to build software. Obviously, there's going to be a bunch of stuff you could do with software that you just couldn't do before at all. And so there will be more competition there, and of course this comes with a new margin structure. But, as per our conversation earlier, we don't really know what that margin structure is going to look like.

Are you going to go to outcome-based pricing? It's really hard to tie each button press in a piece of enterprise software to P&L. Sometimes you can in Salesforce or something, but there's an awful lot of software where it would be really hard to say, “Well, the work I did today did this to P&L; therefore, this is what we should pay for it. This is what we should pay for that piece of software.” I don't think that makes sense in the long run. Anyway, what does the pricing structure look like over time versus now? There will be more competition. It will be easier and quicker to build stuff.

The way that I thought about it, I suppose, is that there are maybe 2 framings to think about this that are useful. One of them is to say that if you think about the enterprise software fleet today, you've got 3 buckets. You've got your big-iron horizontal systems—SAP and Workday, your CRM and your human capital management software, your payroll management software, and so on. Then you've got vertical software. A typical big U.S. company has like 300 to 400 SaaS apps, and then another 1,000 apps that they bought or built themselves internally, running on-prem.

And in the middle, you've got this fuzzy, improvised space of Excel and email and the shared file system and so on. Stuff kind of moves back and forth between those. In principle, every SaaS app is doing something that you could have done in SAP or Excel. You could have managed your graduate recruiting in Workday, but at a certain point—like in our conversation the other day—if you're PwC and you hire however many thousand graduates every year to train to be accountants, you probably have a piece of dedicated software that you built for yourself, or maybe you hired Accenture to build. You probably hate it, but anyway, you've got this piece of dedicated hiring software, or you bought something.

If you're a company that hires 5 graduates a year, you're doing that in email and a shared Google Sheet, because why would you buy software for that? Then there's a space in the middle. Do you do it in Workday? Do you do it in Excel? Do you do it in a dedicated app? And now you add ChatGPT to that. Do you do that in an LLM? Is there an LLM tool that means you can do that in Salesforce where you couldn't do it before, or you can do it in your vertical app where you couldn't do it before? Do you use the LLM to build yourself a tool for that, just as you might have a company department that runs on a 10-meg Excel spreadsheet that someone built 15 years ago? No one knows it works. No one knows how it works, but they're still using it. So it arrives within this broad, fragmented, complicated landscape, and it's another set of options for how you would do that task.

So this is one framing to think about it. I think the other framing to think about this is: does the LLM go at the top of the stack or the bottom of the stack? On the one hand, the bottom of the stack is a feature inside Salesforce. You're in Salesforce: look at the history with this customer, look at the context of every other sales call we've done, look at our business objectives, and suggest an email—or suggest what I should do here and what I should say on the call to the customer. So it's a feature; it's a button that's controlled and has tooling and guardrails and everything else that are driven by that particular use case.

The other way to look at it is the example I gave earlier: go look at Salesforce and Workday and all of our email and Google Analytics, and synthesize something that you couldn't have done before. The tension in both cases is: where do you put the probabilistic software that can make mistakes, and where do you put the deterministic system software that can't answer these kinds of questions? Where do you put the database, and where do you put the LLM? Which is at the top and which is at the bottom? The answer is probably both, depending on what you're doing and where it goes.

All of which is a long way of saying: what does this do to software? The answer is more software—way more software. I mean, all software companies exist to solve problems created by other software companies. That was the joke in security: all security software exists to solve problems created by other security software. Clearly, that's what we went through with SaaS. SaaS gave us an order of magnitude—2 orders of magnitude—more software. We should probably expect that with this.

What that gets to, with the SaaS apocalypse, is that all the investors are looking at all these companies and saying, “We don’t really know which of these companies are going to get screwed by all of this.” Some of them must be. Obviously, some percentage of all the SaaS companies that are out there are going to get wiped out by this, but you don’t know which ones, so you probably shouldn’t derate the whole thing by 50%.

But clearly you’re going to go, “I’m not sure I’m going to be long software at the moment until I have some idea of what the hell’s going on.”

Erik Torenberg

You said in your talk with Ben Thompson that software is someone sitting down and designing a workflow and saying, “This is the right way of doing this from now on.” But you also said that a process grows out of the way a business runs. Does that just take time, or do you think we need more experimentation and iteration from these vertical AI startups to get this into the right shape of software for the future?

Benedict Evans

In a sense, maybe an interesting turn on this is that this is both what strategy consultants and software companies do. They look at what’s going on inside a company and say, “This is a crap way of doing it. This would be a better way of doing it. It would achieve your objectives better.” A software company encodes that in software, and a strategy consultancy encodes that in workflows, charts, processes, training, and objectives. It might tell them to buy some software to do that thing, or now, increasingly, maybe build them that software as well.

Another thing to talk about here is how much of what’s done inside an organization is implicit and not documented, not in the training data, and not something that anybody in that company could actually sit down and draw you a flowchart of and explain to you. That’s a big chunk of the value of Bain, BCG, and McKinsey: They have a license to come into a company and talk to everyone, including the people you’re not allowed to talk to because they’re in a different organization and might get fired. They can work out how this actually works, as opposed to how it’s supposed to work, and why people aren’t doing the strategy. Because, actually, guess what? Their bonus targets depend on them not doing the strategy.

They can work all of that out and be a team ready to come in from the outside and give you the answer. Then you can blame them or have that kind of pre-baked solution. These are problems in organizational management and how people function, and in how people can explain what they do. They’re very hard to write down and very hard to bake into a Claude skill and say, “There you are. Make a PowerPoint.”

There’s a broader “How does this always work?” challenge here: How do you get people to use these technologies? How do people adopt new tools? How do you work out how to help people adopt new tools and work out what new things you would do with them? That’s also what happened with cloud, the web, mobile, the internet, PCs, spreadsheets, and so on.

Erik Torenberg

To that end, do you think there’s some kind of coevolution between AI-native software and new types of interfaces? For example, new customer-service AI platforms that might not have had as much human-facing UI, or systems-of-record software being built without a front end at all because its primary user will be AI agents querying it. I think these are interesting ideas. They’re things I struggle to have a strong opinion on because they’re not deep into the weeds of how enterprise infrastructure gets bought.

Benedict Evans

I wonder how new some of these questions are. I remember Chris Dixon saying 10 or 15 years ago that APIs are the new BD, and software companies wouldn’t need to come—software companies could just open up their APIs. Well, what’s old is new. You don’t need an API anymore. You just have an MCP server, and people will plug into that. The agent will just plug into it.

I don’t know. I think the challenge with a lot of this stuff is that all the decisions are really exception handling. The question is always: What can you not automate? What requires someone to make a decision, exercise some judgment, and have an opinion about it because maybe that hasn’t been written down, or that didn’t happen before, or it doesn’t look quite the way it happened before?

There are various ways of thinking about separating out what gets automated and what doesn’t. The way I used in the deck was to talk about what’s a task versus what’s a job. The tasks used to accomplish the job might change without the job itself changing very much, or without the thing that the job is selling to the client changing very much. If you think about what accountants did 50 years ago and what accountants do today, they spend almost none of their time doing the same things. To the client, though, it’s kind of the same thing. It just gets done in a completely different way, with a whole bunch of different tasks.

One of the more profound, or perhaps abstract, ways to think about this is: Where is it that you want the average? Where is it that what you want is the way everybody would do this? That’s the way everyone would do it. That’s what anyone would say. That’s what anyone would make. That’s what any associate would make. That’s what anybody would give me. That’s the answer anyone would give.

Versus where is that not what you want? Where is it that you want the answer to a new question, or a different answer, or a different idea? LLMs are going to be very good at anything where you can describe how people do it, and where what you want is the way anybody would do that. They’re not going to be as good at things where you can’t really explain why you did it like that, and where you’re doing it differently from the way people would normally do it.

Erik Torenberg

Various people, including the CEO of Google, have said that the risk of underinvesting is riskier than overinvesting. Is there any level of capex where that stops being true, and are we getting there now?

Benedict Evans

There’s a financial-gravity problem in that Microsoft, Meta, and Google are all in line to spend over 50% of revenue on capex this year. We think of telecoms as being capital-intensive. Telecoms spend 15% to 20% of revenue on capex.

The guidance from the big 4 companies this year is $700 billion. Telecom is $300 billion, mobile is $200 billion, and total telecom is $300 billion. Oil and gas, depending on which definition you use and which parts of it you’re counting, is anything from $700 billion to $1 trillion. I think, from memory, it depends exactly on who you ask.

$700 billion a year is not an impossibly large amount of money in terms of what big global infrastructure costs. It’s just a lot of money. Clearly, those companies could not spend $1.5 trillion next year. If they did, they’d have to borrow it, and they certainly couldn’t sustain that level of spending for any length of time. There’s a certain point at which that growth has to slow down because there isn’t any more money.

You can talk about ROI and your ability to produce returns from that investment. Clearly, the capital markets are willing to fund that up to a point. But pick a number at random: We can’t spend $10 trillion a year on AI infrastructure because there isn’t $10 trillion a year there to spend on it. There are finite, almost laws-of-physics caps on the amount of money that’s available. I’d hesitate to say something more tangible than that at the moment.

I almost go back to what I said at the beginning: We’ve got a bunch of multiples. There’s far more demand than supply. On the other hand, efficiency is increasing massively. We don’t know what the next model will be. We don’t know where edge or open source come in yet, or when they come in yet. Meanwhile, you’re always chasing the next model.

This is the line that runs across all of it: The model is only relevant for 3 to 6 months, 6 to 9 months, or whatever you want to say. The model costs how many billions of dollars? How much infrastructure do you need to do that? I don’t think the math has really shaken out yet.

Obviously, there are a bunch of very clever semiconductor analysts who spend lots of time trying to put numbers on this. It’s kind of like trying to put numbers on internet bandwidth in the late 1990s. You know what the rows in the spreadsheet are, but you don’t really know what the values are. All you can really say is, “Well, look, it can’t be huge.” There are clearly physical limits on this.

Another way to answer the question is that if you’re Google, Meta, or Microsoft—and, to some extent, Amazon, and to some extent, Apple—this is an existential problem. You have a FOMO problem. On the one hand, your returns on the investment at the moment are hugely positive. On the other, you can’t let other people get away with this without you participating, because then your company is gone.

You don’t want to end up like Microsoft in the 2000s, IBM in the 1990s, or indeed Meta in the 2010s, where they were continually getting shafted by Apple. If this is the future of compute, then you need to be participating in it.

Obviously, at the same time, the CFO is sitting there saying, “Well, yeah, that’s great, but how much participation are we talking about here?” It’s clear that, at a certain point, that curve is going to have to taper off, because there’s nowhere else it can go.

Erik Torenberg

Do you think there’s going to be a reckoning around token maxing? Is it possible that companies have been overshooting AI usage and, when they do proper ROI studies, they’ll pull back?

Benedict Evans

Well, obviously, you’ve had people using the most expensive model to dick around on the internet, which is kind of what happened with mobile in 2010. You got a $10,000 bill and would have said, “Wait, wait, I thought this was a flat-rate bundle. What happened?” So you’ve obviously got a bunch of silly and meaningful stories.

I think what’s slightly more interesting as a question is that, clearly, there’s going to be a point at which—as I’ve said several times—we’re at a moment of massive disequilibrium. The pricing has got to get back into alignment with the cost, and the usage has got to get into alignment with the pricing and the ROI.

The challenge is that it’s a bit tricky at this early stage. It’s quite hard to know what the ROI is. It’s rather like giving everybody the internet in the late ’90s and saying, “Okay, go off and be more productive.” If you ask CFOs where they’ve seen the benefits, most of the benefits so far have been stuff that’s pretty hard to measure.

There’s a survey from Deloitte, and there’s also a survey from the Fed that’s in my presentation. The benefits are things like better analytics, better customer support, and more productivity. You can make more slides more quickly, and you can do analysis more quickly. It’s kind of tough to put a financial value on that. It has a financial value, but it’s not the same as saying, “We made this new thing with AI, and it had this revenue, or it saved us this much money.”

Those things just take longer. It’s harder to build a new revenue line than to give this to everybody and have them use it to make spreadsheets more quickly. So there’s a little bit of, “Well, how long does this take?”

I think the other answer to the problem here, of course, is consumer surplus, which is kind of what happened with Excel. If a DCF takes you a week, then you probably only do 1 or 2 DCFs. If a DCF takes you 10 seconds, then you do 50 DCFs, but you probably can’t charge any more money for that.

Some of what happens is that these things become competitive necessities, and everybody has to buy and use them. But the cost saving or the productivity gain that you get from them just gets competed away, so you don’t get to charge more for it.

I mean, if you’re at McKinsey, Bain, or BCG, and a piece of analysis used to take a week and now takes a day, you probably do 5 times more analysis and charge your customer the same amount. Your cost base hasn’t changed either. That’s exactly the way to think about what happened with investment banks and financial analysis. You just went and did way more analysis with probably fewer people and charged customers the same amount of money.

Erik Torenberg

Part of your big thesis is this idea that models are going to end up as commodities, and yet the layer that’s raising the most money—in the fastest time in history—is these foundation-model companies. Given that, what advice might you have for them, either collectively? We can pick on someone individually in order to adapt.

Benedict Evans

It’s not that I know they’re going to become commodities. My position is more, “Well, here is a chain of argument that says that, deterministically, it looks like these things will be commodities. Explain to me why they won’t.” That’s as far as I would commit to that.

I think the raising of all this money kind of goes back to my point about mobile, which again has no predictive value but is a worthwhile observation. The mobile industry is very big, spends a lot of money on infrastructure, isn’t very profitable, and all the cool stuff is done by somebody else.

Then you ask, “What’s the return on capital?” The answer is, well, it depends which market—whether you’re in America, Europe, India, or China. But meanwhile, that was a worthwhile thing to do, and it produced a return for somebody. It just ended up not controlling the whole thing, and other people ended up getting more value from that than they did.

I don’t have the number in my head. What was Google’s net income last year—$50 billion or something? What was the net income for the total telecom industry? I should really subscribe to Bloomberg; then I could just answer these questions instantly.

But it’s a pretty safe bet that Google, Meta, Amazon, Microsoft, and Apple produce more profits than the entire telecom industry. This is a puzzle: You’re driving the frontier forward, but you’re caught in this trap that you have to keep competing because otherwise they’ll do it and you’ll fall behind.

You’ve also got this thing that we haven’t talked about at all: Aren’t we just building AGI? We’re going to build God in a box, which some people do believe, although it’s kind of hard to analyze. So you’re going to carry on building this stuff, but the practical question is, how do you get things that people want to use that aren’t software development? I mean, that’s a good business.

Is that the only business? There are, you know, many hundreds of billions of dollars to be made making the software industry more productive. Great—but then what? How do you expand this into the rest of the economy, into everybody else?

That’s why you get these conversations about private equity partnering with consultancies. As we’ve been discussing, it’s actually quite hard to work out what to do with this stuff if you’re running a real company. So you go to Bain, BCG, McKinsey, Infosys, Cognizant, IBM, Accenture, or private equity shops.

There’s this sense that, on the one hand, you’re building these bigger and bigger models and you feel like you’ve got to keep doing it. But on the other hand, what are people doing with it?

Erik Torenberg

Why do most people look at ChatGPT and not really think of anything to do with it today?

Last question: Is there anything from the presentation that you want to make sure listeners leave with?

Benedict Evans

The thing that I used last year and used again is an IBM ad I found from the early ’50s, which has a picture of a sea of engineers all holding up slide rules. It’s an IBM ad, and it says, “An IBM electronic calculator gives you 150 extra engineers.” How many pictures have you seen at a16z where that was the pitch?

We kind of remember that we go through these waves of fundamental technology changes every 10 or 15 or 20 years. They’re all amazing, change everything, and are completely unlike anything that’s happened before.

AI is amazing and transformative and completely unlike anything that’s happened before. Mobile was quite a big deal, too, and so was the internet, and so were PCs, and so was computing. Those were all also very big deals where it was hard to tell what was going to happen.

We should presume as a base case that we’re going to go through that again. That will produce a bunch of things that ruin people’s lives and put a bunch of people out of work. There’ll be a bunch of stuff that we’re not very happy about, and there’ll be a bunch of stuff that we all think is great.

Then, in 20 years’ time, we’ll kind of forget that there was a world when computers couldn’t do that. I mean, here we are. We’ve been on this call for an hour, and our computers didn’t crash. We’re streaming HD video to each other, and it’s like, of course that worked.

In fact, I’m also doing it with my iPhone. My iPhone is streaming to my Mac over Wi-Fi, streaming video here, and it just works. It’s like magic, and we don’t notice it anymore. I think that’s really my one-line description of how all of this is going to end up: It’s going to be magic, and in 20 years’ time we’ll just say, “Well, of course that’s how it is. Computers have always done that.”

Erik Torenberg

Yeah, that’s a great place to wrap. The presentation is called “AI Eats the World,” and it’s on Benedict Evans’ website. Benedict, this has been a great conversation. Thanks so much for coming to the podcast.

Benedict Evans

Thanks. Great to chat.

AI 使用的经济学与 SaaS 的下一步 | a16z 上的 Benedict Evans — 文字稿与摘要 | BidClub