[BidClub_]
The a16z Show · · 62 分钟

AI 吞噬世界:Benedict Evans 谈下一轮平台变迁

Benedict EvansErik Torenberg

YouTube
TL;DR
  • 生成式 AI 可能带来平台级变革,但 Evans 目前看不到它超越互联网或智能手机的证据。他的“中间派”立场是:AI“和互联网或智能手机一样重要,但也就和互联网或智能手机一样重要”。上行空间依然异常难以判断,因为我们既没有关于智能的可用理论,也不了解这些模型为何表现得如此出色。

  • AI 需求规模巨大,却高度不均衡,使工作流采用成为核心商业问题。ChatGPT 拥有“8亿或9亿”周活跃用户,但只有约5%付费;Evans 估计,发达国家用户中约10–15%每天使用,另有20–30%每周使用。投资者真正要问的是:为什么多出5倍的人已经理解这个产品,却“这周或下周都想不出能拿它做什么”。

  • 泡沫大概率会出现,但无论时点还是最终基础设施需求,都无法有把握地建模。Evans 认为,“如果现在还不是泡沫,那也会是”,而实时区分1997、1998和1999年式的市场状态根本不可能。即使使用量爆发,算力需求也可能每年下降20倍、30倍或40倍,重演1990年代末那种无法落地的带宽预测。

  • 模型之上的产品机会,很可能集中在能够编码工作流、验证机制和机构知识的专业化产品。“人们买的是解决方案,不是技术”:律所要的是法律取证软件,而不是翻译和情感分析 API 调用。讨论指向的是围绕通用模型打造的专用应用,而不只是出售裸模型访问权限。

  • OpenAI 的8亿至9亿周活跃用户代表分发能力,但还不是持久护城河。Evans 看到了品牌和默认入口地位,却看不到明确的网络效应、功能锁定、专有基础设施或成本优势:“每个月 Satya 都会给你寄来一张账单。”OpenAI 必须同时冲刺两件事:建立可防御的产品生态,以及联合 NVIDIA、Broadcom、AMD、Oracle 和新一轮资本池打造基础设施。

  • 对在位者的影响并不对称:Google 可以吸收 AI,Meta 和 Amazon 面临更深层的产品问题,而 Apple 可能继续保持隔离,除非计算本身发生变化。Google 能够为前沿模型提供资金,并把 AI 变成搜索和广告的功能;Amazon 则可能终于改善超越商品零售的发现体验。如果应用消失,Apple 会受到威胁;但即使进入 LLM 优先的世界,用户可能仍然需要“一块漂亮的大彩色屏幕”、摄像头和电池——换句话说,仍然需要某种和 iPhone 非常相似的东西。

  • 最深层的颠覆,会暴露那些利润依赖路由、捆绑或摩擦,而非依赖名义产品本身的企业。Evans 的判断路径是:先采用功能,再出现新能力,最终可能“把整个行业从里到外翻过来”;报纸后来发现,自己部分上是轻工业、本地分发和卡车运输公司,而 LLM 可能抹掉建立在繁琐行政流程上的防御。不过,要说 AI 比互联网更大,他还需要看到狭窄护栏之外“真正就是一个人”的能力:“我们现在拥有的还不是那个。”

摘要 · 为研究而整理的核心内容

1. AI 可能成为平台级变革,但历史无法揭示赢家

  • Evans 用两个问题搭建整场演讲:当平台发生变化时,技术内部会发生什么;以及哪些外部行业会被改造,而不只是得到辅助。互联网彻底改变了报纸行业,对水泥行业却“只是有点用”,并未从根本上改变其商业模式。

  • 他刻意保持克制的判断是,生成式 AI 可能“和互联网或智能手机一样重要,但也就和互联网或智能手机一样重要”。这些变革已经创造并摧毁行业,重排科技公司的主导格局,并催生出新的千亿美元和万亿美元级企业。

  • 今天的“AI”标签本身可能只是过渡性的。就像数据库、Web 和智能手机一样,成熟能力最终会隐入普通产品:Otis 曾把红外光束宣传为“电子礼貌”,如今却“只是一部电梯”。Evans 说,在日常语境中,“AI 似乎意味着新东西”,而 AGI 意味着“新的、可怕的东西”。

  • 历史分类法有参考价值,却没有预测力。移动互联网把计算从 Web 推向应用,让智能手机进入50亿至60亿人的手中,而消费级 PC 用户还不到10亿,并由此催生 TikTok 和现代网络约会。但在1990年代中期,“你可以知道它,却不知道它意味着什么”:Amazon 当时只是书店,Netscape 刚刚上线,Google 和 Facebook 的创始人还没有创办自己的公司。

2. AI 不可知的上限打破传统科技预测

  • 以往的平台变革虽有不确定性,却存在清晰的物理边界:电信运营商不可能在1995年铺设覆盖所有人的千兆光纤,iPhone 也不会突然多出一年的续航、展开成投影仪或飞起来。AI 没有对应的路线图,因为“我们并没有真正理解它为什么表现得这么好”,也不了解人类智能本身。

  • Evans 从 OpenAI 一场直播中看到了一个意味深长的矛盾:直播先承诺次年推出达到人类水平、博士水平的 AI 研究员,随后又推广能够让数百或数千名软件开发者“像使用 Windows 一样”使用的 API。AI 要么会变成消灭软件公司的“盒中之神”,要么会成为一个新的软件底座,在其上构建出更多产品;而行业经常同时主张这两种说法。

  • AGI 争论让他想起那个神学家的笑话:要么弥赛亚已经到来,但世界几乎没有明显变化;要么它的到来永远还差5年。Sam Altman 说博士级研究员已经出现,Demis Hassabis 则不接受这种描述;Andrej Karpathy 把时间范围放在大约10年。由于没有关于上限的模型,预测最终只剩下“我的感觉是”,Evans 没有给出可证伪的日期。

3. 即使每个投资者都理性行事,过度投资仍可能发生

  • Evans 的确定性判断是:“非常新、非常大、非常令人兴奋、会改变世界的事物,往往会催生泡沫。”Marc Andreessen 的区分——1997年不是泡沫,1998年不是,1999年是——恰好说明时点问题:“如果现在还不是泡沫,那也会是”,但没人知道眼下更像哪一年。

  • 算力预测类似于估算1990年代末的全球带宽需求。分析师可以在表格中填入用户数、页面大小、视频码率和观看时长,再推算路由器销量,但只要使用合理的输入,结果就可能相差100倍。AI 需求同样受到能力、效率和使用强度等不稳定变量的共同影响。

  • 超大规模云厂商仍然可以理性投资,因为 AI 已经在提升搜索、广告、云和消费产品的价值。它们公开的计算逻辑是:不投资的下行风险高于过度建设——这一论点“总是在失效前表现良好”,尤其是在杠杆、交叉杠杆和循环收入叠加、市场却开始下跌之后。

  • 效率可能已经以每年20倍、30倍或40倍的速度提升,下一次模型演进甚至可能用当前算力的百分之一实现同等结果;但与此同时,使用量也在上升。Zuckerberg 提出 Meta 可以转售多余产能,却忽略了其中的相关性:如果 Meta 不再需要这些产能,“其他所有人也会有大量闲置产能”。

4. 采用分化为明显的重度用户和等待产品的人群

  • Evans 认为,软件开发、营销、弹性知识工作和狭窄的企业级解决方案会率先落地。营销人员过去制作30个素材,现在可以制作300个;Accenture、Bain、McKinsey 和 Infosys 等咨询公司则可以把模型应用到大型企业内部的具体流程中。

  • 大众市场的数据呈现出另一幅图景:ChatGPT 拥有“8亿或9亿周活跃用户”,约5%付费;发达国家用户中可能有10–15%每天使用,另有20–30%每周使用。对于一个每天使用数小时的人来说,真正值得追问的是:为什么大约多出5倍、且已经了解产品的人,却找不到本周有用的任务。

  • Excel 提供了 Evans 的类比。对会计师来说,修改一份10年期 DCF 的折现率,可以把数天的重新计算压缩到几分钟;对律师来说,电子表格有用,却不是日常工作的核心。Torenberg 补充说,很多人可能需要把这些能力嵌入工作流、UX、工具和产品中,由产品告诉他们该做什么。

  • 验证决定概率性输出究竟能否节省劳动力。生成200张营销图片、从中挑出10张是高效的;如果从 PDF 中抄录200个数字,却必须由人逐一核对,那就不是。Evans 以 OpenAI Deep Research 生成的移动市场分析为例:数字既有转录错误,也有来源选择错误,所以“我还不如自己做”。

5. 胜出的产品会告诉用户该问什么

  • 如果因为生成式 AI 无法执行所有现有任务就否定它,就像在1970年代末因为 Apple II 跑不了银行系统而否定 Apple II。问题不只是新系统能否完成在位系统的核心工作负载,更在于它能否带来过去无法完成的事情。

  • 因此,机会不只是自动化旧任务,也包括发现过去没人尝试过的行动。创业者可以识别出一种新能力,理解一个行业,再围绕它做出一个按钮;用户不必从空白聊天框开始推导工作流。Torenberg 称之为“拆解 ChatGPT”,而 Evans 的例子说明了产品层为何重要。

  • 他用 Everlaw 说明了技术栈逻辑。机器学习可以提供翻译和情感分析,但律所仍然购买云端法律取证软件,而不是自己拼装 AWS API 调用。“人们买的是解决方案,不是技术”,因为领域专属的流程、界面和分发能力仍然位于模型之上。

  • 图形界面不只是暴露数百项功能:每个页面上的7个相关按钮,都凝结了多年机构知识,告诉用户下一步该做什么。原始提示词则把这项工作重新丢回给用户,就像收到“无限多个实习生”,却要自己告诉他们什么是风险投资,以及任务需要的来源到底是季报、Bloomberg 还是 PitchBook。“它是在问你所有事情。”

6. 模型趋同让 OpenAI 拥有规模,却只有脆弱的护城河

  • Torenberg 对 a16z 的反思是,公司最后悔的是“没有把规模做得更大”:语音、图像生成和其他专业方向诞生了比预期更多的独立赢家。即使在同一品类,巨大的市场也能容纳多家公司;而品类本身也会被捆绑、拆分和重组。1995年,Evans 手里有4到5个浏览器,因为当时连 Web 的用途都还没有确定。

  • 如今,通用基准显示前沿模型之间的差距已经相当小,但消费者使用却极不均衡。重度用户能分辨 Claude 的语气,或 GPT-5.1 和“GPT-4.9,或者它到底叫什么鬼”的差异;每周使用者通常分辨不出来。按 Evans 的说法,Claude 几乎没有消费者使用量,而 ChatGPT 的领先幅度超过 Meta 和 Google,尽管几者基准表现大体相当。

  • 对轻度用户而言,底层模型因此可能成为商品。OpenAI 的8亿至9亿周活跃用户依靠的是“默认入口的力量和品牌”,却还没有形成成熟的网络效应、无法复制的记忆、广泛生态、自有基础设施或成本优势。它必须紧急向上进入浏览器、应用、社交视频和平台,同时向下锁定算力供应:“我们昨天就要把所有这些都建起来。”

7. 每家超大规模云厂商面对的战略方程式都不同

  • 对 Google 而言,前沿能力可能只是继续做 Google 的必要成本。Gemini 与 GPT-5.1 可能每月轮流占据基准测试领先位置;维持这一地位的成本,按 Evans 有意保持宽泛的估计,可能是每年1000亿美元或2500亿美元。Google 付得起这笔钱,可以改进搜索和广告,发明定义 AI 的界面,也可以像 Android 复制智能手机模式那样复制它。

  • Meta 面临的是内容、社交体验和推荐等更大的问题,因此掌控模型具有战略必要性。Amazon 可以销售商品化基础设施,同时用 LLM 改善发现体验:它非常擅长交付用户点名要的 SKU,却“非常不擅长告诉你想要哪个 SKU”。AI 或许能够推断意图、创造需求,而不只是关联过去的购买记录。

  • 出版商、品牌和营销人员可能还不知道自己该问什么。如果 LLM 直接回答一道菜谱请求,原本依赖 Google 路由流量的菜谱网站会怎样?如果购物者把手机对准客厅,询问应该买什么,发现路径以及谁能捕获商业意图,都可能与搜索或 Amazon 当前的目录体系根本不同。

  • Apple 的问题取决于 AI 是一种服务,还是计算本质的变化。Craig Federighi 的挑战——Apple 也不拥有 YouTube 或 Uber——比听起来更有力;Microsoft 失去了开发环境,却仍然让 Windows PC 的销量增长了一个数量级,因为访问 Web 仍然需要 PC。即使应用消失、融入 LLM,用户可能仍会为最好的屏幕、摄像头和电池付费,从而让 iPhone 的硬件地位保持得更久。

8. AI 会揭示每个企业真正卖的是什么

  • Evans 的采用阶梯分为3个阶段:把 AI 做成功能,用它创造新东西,然后看着新进入者可能“把整个行业从里到外翻过来”。对于湾区或华盛顿特区的 Walmart 门店经理,这可能从“帮我找那个指标”,进展到“给我做一个仪表盘”,再到黑色星期五当天问:“我应该担心什么?”

  • Amazon 的对应跃迁,是从用户买灯泡后推荐封箱胶带,变成推断买家正在搬家,并展示家庭保险报价——后者可能是购买相关性无法捕捉的意图。内容行业也面临同样的区分:用户想要的是 Bolognese 菜谱,还是 Stanley Tucci 谈意大利烹饪;一份演示文稿,还是 Bain 合伙人提供一周的建议?

  • 平台变革会暴露隐藏的工作和护城河。报纸强调自己卖的是新闻,却发现轻工业、本地分发和卡车运输在经济上同样关键。Evans 以美国健康保险为例,提出一个刻意保留余地的思想实验:如果盈利部分依赖于把流程做得“无聊、困难、耗时”,那么消除令人厌烦工作的 LLM,就会攻击一种从未被承认的防御机制。

  • 3G 的“杀手级应用”最终证明是随时随地拥有互联网,而不是分析师试图事先点名的那些狭窄应用。AI 可能也会带来同样的事后清晰度,但 Evans 认为,要称其规模超过互联网,门槛更高:必须具备突破狭窄护栏、 “真正就是一个人”的能力。今天的系统有时能出色完成类似人的任务,但“我们现在拥有的还不是那个。它会成长到那个程度吗?我们不知道。”

Benedict Evans

ChatGPT has 800 or 900 million weekly active users. If you’re the kind of person who’s using this for hours every day, ask yourself why 5 times more people look at it, get it, know what it is, have an account, know how to use it, and can’t think of anything to do with it this week or next week. The term AI is a little bit like the term technology. When something’s been around for a while, it’s not AI anymore. Is machine learning still AI? I don’t know. In actual general usage, AI seems to mean new stuff, and AGI seems to mean new, scary stuff.

Erik Torenberg

AGI seems to be a bit like this: either it’s already here and it’s just software, or it’s 5 years away and will always be 5 years away. We don’t know the physical limits of this technology, and so we don’t know how much better it can get. You’ve got Sam Altman saying we’ve got PhD-level researchers right now, and Demis Hassabis says, “No, we don’t. Shut up.” Very new, very big, very exciting, world-changing things tend to lead to bubbles. So, yeah, if we’re not in a bubble now, we will be.

Benedict, welcome back to the a16z podcast.

Benedict Evans

Good to be back.

Erik Torenberg

We’re here to discuss your latest presentation, “AI Eats the World.” For those who haven’t read it yet, maybe you can share the high-level thesis and contextualize it in light of recent AI presentations. I’m curious how your thinking has evolved.

Benedict Evans

Yeah, it’s funny. One of the slides in the deck references a conversation I had with a big-company CMO who said, “We’ve all had lots of AI presentations now. We’ve had the Google one and the Microsoft one. We’ve had the Bain one and the BCG one. We’ve had the one from Accenture and the one from our ad agency. So now what?”

It’s sort of 90-odd slides, and there are a bunch of different things I’m trying to get at. One of them is to say: if this is a platform shift, or more than a platform shift, how do platform shifts tend to work? What are the things that we tend to see in them, and how many of those patterns can we see being repeated now?

Some of the patterns are things like bubbles, but others are that lots of stuff changes inside the tech industry. There are winners and losers; people who were dominant end up becoming irrelevant, and then there are new billion- and trillion-dollar companies created. But there’s also the question of what this means outside the tech industry, because if we think back over the last waves of platform shifts, there were some industries where this changed everything and created and destroyed industries, and others where this was just a useful tool.

If you’re in the newspaper business, that had a very different impact. The last 30 years look very different from if you were in the cement business, where the internet was just kind of useful but didn’t really change the nature of your industry very much.

What I tried to do is give people a sense of what’s going on in tech: how much money we’re spending, what we’re trying to do, what the unanswered questions are, and what might or might not happen within the tech industry. But then, outside technology, how does this tend to play out? What seems to be happening at the moment? How is this manifesting in tools and deployment, new use cases, and new behaviors?

As we step back from all of this, how many times have we gone through all of this before? It’s funny: I went on a podcast this summer, and my opening line was something like, “Well, I’m a centrist. I think this is as big a deal as the internet or smartphones, but only as big a deal as the internet or smartphones.” There were about 200 YouTube commenters underneath saying, “This is more, and he doesn’t understand how big this is.” And I think, well, it was kind of a big deal.

Erik Torenberg

It was kind of a big deal.

Benedict Evans

I finish the day by looking at elevators, because I live in an apartment building in Manhattan and we have an attended elevator. That means there’s a person; there are no buttons, there’s an accelerator and a brake, and the doorman gets in and drives you to your floor. It’s a streetcar.

In the 1950s, Otis deployed automatic elevators. You get in and press a button, and they marketed it by saying, “Ah, it’s got electronic politeness,” which meant the infrared beam. Today, when you get into an elevator, you don’t say, “Ah, I’m using an electronic elevator. It’s automatic. It’s just a lift.”

That’s what happened with databases, the web, and smartphones. Databases certainly aren’t AI. I’ve done a couple of polls on this on LinkedIn and Threads, asking, “Is machine learning still AI?” I don’t know. There’s obviously an academic definition where people say, “This guy’s an idiot.” Of course, I’m going to explain the definition of AI, but in actual general usage, AI seems to mean new stuff.

Erik Torenberg

Yeah, and AGI seems like new, scary stuff.

Benedict Evans

Yeah, it’s funny. There’s an old theologian’s joke that the problem for Jews is that you wait and wait and wait for the Messiah, and he never comes. The problem for Christians is that he came and nothing happened. The world didn’t change; there is still sin. For all practical purposes, nothing happened.

AGI seems to be a bit like this. Either it’s already here, and it’s just more software, so you’ve got Sam Altman saying we’ve got PhD-level researchers right now, and Demis Hassabis says, “No, we don’t. Shut up.” Or it’s 5 years away and will always be 5 years away.

Erik Torenberg

Yeah, yeah, it’s interesting. Let’s compare this with previous platform shifts, because some people look at something like the internet and say, “Hey, there were net-new trillion-dollar companies—Facebook and Google—that were created from it, and all sorts of new emerging winners.”

Whereas they look at mobile and say, “There were big companies like Uber, Snap, Instagram, and WhatsApp, but these were billion-dollar outcomes or tens-of-billions-of-dollar outcomes. Really, the big winners were, in fact, Facebook and Google.”

In some sense, mobile perhaps was sustaining. Feel free to quibble with the definition of sustaining versus disruptive, but sustaining in the sense that maybe more of the value went to incumbents, or companies that existed prior to the shift.

I’m curious how you think about AI in light of that. Is it enabling? Are more of the gains going to come from net-new companies like OpenAI and Anthropic and others that follow? Or are more of the gains going to be captured by Microsoft, Google, Facebook, and Meta—companies that existed prior to the shift?

Benedict Evans

There are several answers to this. One of them is that you kind of have to be careful about framings and structures and things, because you end up arguing about the framing and the definition rather than arguing about what’s going to happen. They’re all useful, but they’ve all got holes in them.

What mobile did was shift us in several fundamental ways. It shifted us from the web to apps, for example, and it gave everybody in the world a phone. It gave everybody in the world a pocket computer. Even today, there are fewer than 1 billion consumer PCs on Earth, and there are somewhere between 5 billion and 6 billion smartphones.

It made possible things that would not have been possible without it, whether that’s TikTok or, arguably, things like online dating. You can map those against dollar value, but you can also map them against structural change in consumer behavior and access to information. You could certainly argue that Meta would be a much smaller company if it weren’t for mobile, for example. You can argue the puts and calls on this stuff a lot.

Not all platform shifts are the same, and you can do the standard typology: there were mainframes, then PCs, then the web, and then smartphones. But you kind of want to put SaaS in there somewhere, and you kind of want to put open source in there. Maybe you want to put databases in there, too.

These are useful framings, but they’re not predictive. They don’t tell you what’s going to happen; they just give you one way of understanding some of the patterns we’ve seen.

The big debate around generative AI is whether this is just another platform shift or something more than that. The problem is that we don’t know, and we don’t have any way of knowing other than waiting to see. This may be as big as PCs, the web, SaaS, or open source—or it may be as big as computing itself.

Then you’ve got the very overexcited people living in group houses in Berkeley who think this is as big as fire or something. Well, great.

But does this create new companies? I mean, you go back to mobile. There was a time when people thought blogs were going to be different from the web, which seems weird now. Google needed a separate blog search.

This seriously was a thing. There was a time when it was really not clear, and I think you can kind of generalize his point. You go back to the internet in the mid-’90s: we kind of knew this was going to be a big thing, but we didn’t really know it was going to be the web. Before that, we didn’t know it was going to be the internet. We knew there were going to be networks, but we didn’t know it was going to be the internet.

Then it wasn’t clear that it was going to be the web, and it wasn’t really clear how the web was going to work. When Netscape launched, Mark Zuckerberg was in junior high or elementary school or something, Larry and Sergey were students, and Amazon was a bookstore. So you can know it but not know it.

You could make the same point about smartphones. We knew everyone was going to have an internet-connected thing in their pocket, but it wasn’t clear that it was basically going to be a PC from the PC company from the ’80s and a search engine company. It wasn’t clear that it wasn’t going to be Nokia or Microsoft. I think you have to be super careful about making deterministic predictions about this. What you can do is say, “When this stuff happens, everything changes,” and that’s happened 5 or 10 times before.

Erik Torenberg

I’m curious how you got conviction in this idea, or what got you to the prediction that AI is going to be as big as the internet—which, of course, is pretty big. I’m not yet, Benedict, at the conviction that it’s going to be any bigger. I’m curious what sort of inspires that statement, and what might change your mind either way: that it might not be as big as the internet, because the internet was obviously very big, but also that perhaps it might be bigger.

Benedict Evans

I don’t want to—I remember I made a diagram of S-curves going up slightly, and someone said, “Well, what’s the axis on this diagram?” I don’t want to get into whether this is 5% bigger than the internet or 20% bigger. I think the question is more like: Is it another of these industry cycles, or is it a much more fundamental change in what technology can be? Is it more like computing or electricity, as a sort of structural change, rather than, “Here’s a whole bunch more stuff we can do with computers”?

I think that’s the question, and there’s a funny disconnect in looking at debates about this within tech. I watched one of the OpenAI livestreams a couple of weeks ago, and they spent the first 20 minutes talking about how they were going to have human-level, PhD-level AI researchers next year. Then the second half of the stream was, “Here’s our API stack that’s going to enable hundreds and thousands of new software developers, just like Windows,” and they literally quoted Bill Gates. You think, “Those can’t both be true.” Either I’ve got a thing that is a PhD-level AI researcher—which, by implication, is like a PhD-level CPA—

Erik Torenberg

Yeah.

Benedict Evans

—or I’ve got a new piece of software that does my taxes for me. Which is it? Either this thing is going to be human-level, and that’s a very, very challenging, problematic, complicated statement, or this is going to let us make more software that can do more things than software could do before.

I think there’s a real schizophrenia in conversations around this, because it’s, “Scaling laws, and it’s going to scale all the way,” while meanwhile I’m hearing, “Look how good it is at writing code.” Again, is it writing code, or do we not need software anymore? Because, in principle, if the models keep scaling, nobody’s going to write code anymore. You’ll just say to the model, “Hey, can you do this thing for me?”

Erik Torenberg

Is it a little bit of a hedge, or is it a sequencing thing?

Benedict Evans

Some of it’s a sequencing thing, but in principle, if you think this stuff is going to keep scaling, why are you investing in a software company?

Erik Torenberg

Yeah.

Benedict Evans

We’ll just have this god in a box that can do everything. I think this is the funny challenge, and I think this is the fundamental way that this is different from previous platform shifts. With the internet, or with mobile, or even with mainframes, you didn’t know what was going to happen in the next couple of years. You didn’t know what Amazon would become, you didn’t know how Netscape was going to work out, and you didn’t know what next year’s iPhone was going to be.

Ten years ago, when we cared about that, you kind of knew the physical limits. You knew, in 1995, that telcos were not going to give everybody gigabit fiber the next year, and you knew that the iPhone wasn’t going to have a year’s battery life, unroll, have a projector, and fly or whatever. But we don’t know the physical limits of this technology because we don’t really have a good theoretical understanding of why it works so well. Nor, indeed, do we have a good theoretical understanding of what human intelligence is, and so we don’t know how much better it can get.

You could do a chart and say, “Well, this is the road map for modems, and this is the road map for DSL, and this is how fast DSL will be.” Then you could make some guesses about how quickly telcos will deploy DSL, and say, “Clearly, we’re not going to be able to replace broadcast TV with streaming in 1998.” But we don’t have an equivalent way of modeling this stuff to know what its fundamental capability is going to look like in 3 years, which gets you to these slightly vibes-based forecasts where no one really knows.

Geoffrey Hinton says, “Well, I feel like,” and Demis Hassabis says, “Well, I feel like,” but no one knows.

Erik Torenberg

And then Andrej Karpathy goes on our podcast and says, “I feel like it’s a decade out.”

Benedict Evans

Yeah, I know. I saw this meme of—what’s his name?—saying, “The answer will reveal itself.” Somebody like me would say it’s photoshopped, but of course it wouldn’t have been photoshopped. Somebody had turned him into a Buddhist monk wearing an orange outfit: “The future will reveal itself.”

But this is the problem. We don’t know, and we don’t have a way of modeling this.

Erik Torenberg

Yeah. And so let’s connect this to the upfront investment that some of these companies are making. We don’t know—is there a risk of overinvestment leading to some potential bubble-like mechanics? How do you think about that question?

Benedict Evans

Well, deterministically, very new, very big, very exciting, world-changing things tend to lead to bubbles.

Erik Torenberg

Yeah.

Benedict Evans

I don’t think anybody would dispute that you can see some bubbly behavior now. You can argue about what kind of bubble, but again, that doesn’t have very much predictive power. One of the features of bubbles is that when everything’s going up, everything goes up all at once, everyone looks like a genius, and everyone leverages and cross-leverages and does circular revenue. That’s great until it isn’t, and then you get a kind of ratchet effect as it goes back down again.

If we’re not in a bubble now, we will be. I remember Marc Andreessen saying, “1997 was not a bubble. 1998 was not a bubble. 1999 was a bubble.” Are we in 1997 now, or 1998, or 1999? If we could predict that, we’d live in a parallel universe.

I think there are maybe 2 more specific, more tangible answers to this. The first is that we don’t really know what the compute requirements of this stuff are going to be. Forecasting that feels a lot like trying to forecast bandwidth use in the late ’90s. Imagine if you were trying to do the algebra on that: This many users, how much bandwidth does a web page use? How will that change? How will that change as bandwidth gets faster? What happens with video? What kind of video? What bit rate of video? How long do people watch a video? How much video?

You could build the spreadsheet, and it would tell you what global bandwidth consumption would be in 10 years. Then you could try to use that to back-calculate how many routers this is going to sell. You could get a number, but it wouldn’t be the number. There’d be a hundredfold range of possible outcomes from that. You could make the same point about the algebra of consumption now.

Right now, we have a bunch of rational actors saying, “This stuff is transformative and a huge threat. We can’t keep up with demand for it now, and as far as we know, the demand is going to keep going up.” We’ve had a variety of quotes from all of the hyperscalers basically saying that the downside of not investing is bigger than the downside of overinvesting. That kind of thing always works well until it doesn’t.

Erik Torenberg

Yeah.

Benedict Evans

I saw a slightly strange quote from Mark Zuckerberg saying, “Well, if it turns out that we’ve overinvested, we can just resell the capacity.” I thought, “Let me just stop you there, Mark, because if it turns out that you can’t use your capacity, everybody else is going to have loads of spare capacity as well.”

Erik Torenberg

Yeah.

Benedict Evans

All these people who are desperate for more capacity—if it turns out we can get the same results for a hundredth of the compute—

Erik Torenberg

That will be true for everyone else too, not just you.

Benedict Evans

Yeah. So, in an investment cycle like this, you tend to get overinvestment, but after that, there are very limited predictions you can make about what's going to happen. I think the more useful way to look at this is: you've got these transformative capabilities that are already increasing the value of your existing products if you're Google, Meta, or Amazon, and you're going to be able to use them to build a bunch more stuff. Why would you want to let somebody else do that rather than you doing it, as long as you're able to keep funding and selling what you're building?

Erik Torenberg

Yeah.

Benedict Evans

And it may well turn out that we have an evolution of models in the next year that means you can get the same result for 1/100th of the compute that you're using today. Bearing in mind that it's already going down—depending, pick your numbers—20, 30, 40 times a year.

Erik Torenberg

Yeah.

Benedict Evans

But then the usage is going up. So you're in this very—as I said, it's like trying to predict bandwidth consumption in the late '90s and early 2000s. You can throw all the parameters out, but it doesn't get you to something useful. You just need to step back and say, “Yeah, but is this internet thing any good?”

Erik Torenberg

Well, yeah, because I'm curious if you see the bottlenecks as being more on the supply side or the demand side—more technical constraints—or is it just, is AI any good? Are there enough use cases to justify the type of spend? What are you seeing, and what are you predicting?

Benedict Evans

So, maybe 2 answers to this question. The first of them is, I think we've had this sort of bifurcation of what all the questions are. There are now very, very detailed conversations about chips, and then very, very detailed conversations about data centers and funding for data centers, and then about what a new enterprise SaaS company built on AI—what margins will it have and how much money does it need to raise. So there are venture capital conversations, and there are many different conversations within which I don't know anything about chips. I can spell “ultraviolet,” but I don't know what an ultraviolet process is. It's more violets, I don't know.

And so you've got this—it's like the Milton Friedman line: no one knows how to build a pencil. I think the second answer might be that there are 2 kinds of generative AI deployment. One of them is where it's very easy and obvious right now to see what you would do with this, which is basically software development, marketing, point solutions for many very boring, very specific enterprise use cases, and also people like us, who have very open, free-form, flexible jobs with many different things and who are always looking for ways to optimize that.

Erik Torenberg

Yeah.

Benedict Evans

And so you get people in Silicon Valley who are like, “I spend all my data time in dbt. I don't use Google anymore. I've replaced my CRM with this.” And then, obviously, people who write code—if you're writing code, this works really well. If you're in marketing, there are all these stories of big companies where they're making 300 assets where they would have made 30. And Accenture, Bain, McKinsey, Infosys, and so on are sitting and solving very specific problems inside big companies.

Then there's a whole bunch of other people who look at it and are like, “It's okay.” You go and look at the usage data, and you see that ChatGPT has 800 or 900 million weekly active users, and 5% of people are paying. Then you go and look at all the survey data, and it's very fragmented and inconsistent, but it all sort of points to something like 10% or 15% of people in the developed world using this every day. Another 20% or 30% of people are using it every week. If you're the kind of person who is using this for hours every day, ask yourself why 5 times more people look at it, get it, know what it is, have an account, know how to use it, and can't think of anything to do with it this week or next week.

Erik Torenberg

Why is that?

Benedict Evans

Yeah. Is it because it's early? It's not a young-people thing, either, incidentally. Is that just because it's early? Is it because of the error rates? Is it because you have to map it against what you do every day?

One of the analogies I always used to use—which isn't in the current presentation, but I've used it in previous presentations—is: imagine you're an accountant and you see spreadsheet software for the first time. This thing can do a month of work in 10 minutes, almost literally.

Erik Torenberg

Yeah. You want to change—you want to recalculate that 10-year DCF with a different discount rate. I've done it before you finished asking me to. And that would have been like a day or 2 days or 3 days of work to recalculate all those numbers. Great. Now imagine you're a lawyer and you see it. You think, “Well, that's great. My accountant should see it. Maybe I'll use it next week when I'm making a table of my billable hours, but that's not what I do all day.” Excel doesn't do things that a lawyer can do every day.

Benedict Evans

Yeah. So you've got a whole swirling matrix of how you map this against existing problems. But the other side of it is: how do you map this against new things that you couldn't have done before? And this comes back to my point about platform shifts, because I see people looking at ChatGPT or looking at generative AI and saying, “Well, this is useless because it makes mistakes.” I think that's like looking at an Apple II in the late '70s and saying, “Could you use these to run banks?” To which your answer is no, but that's kind of the wrong question.

Erik Torenberg

Right? Really? Isn't that what they're doing? They're unbundling ChatGPT, just as the enterprise software company of 10 years ago was unbundling Oracle or Google or Excel. Do you have the view that what Excel did for accountants, AI is now doing for coders and developers, but it hasn't quite figured out that daily critical workflow for other job positions? So it's unclear for people who aren't developers why they should be using this for many hours a day.

Benedict Evans

I think there's a lot of people who don't have tasks that work very well with this.

Erik Torenberg

Yeah. And then there's a lot of people who need it to be wrapped in a product and a workflow and tooling and UX, and someone to come and say, “Hey, have you realized you could do it with this?”

Benedict Evans

I had this conversation in the summer with Balaji Srinivasan, who's another former a16z person, and he was making this point about validation: Can you—because these things still get stuff wrong, and people in the Valley often hand-wave this away—there are questions that have specific answers where it needs to be the right answer, or one of a limited set of right answers. Can you validate that mechanistically? If not, is it efficient to validate it with people?

With the marketing use case, it's a lot more efficient to get a machine to make you 200 pictures and then have a person look at them and pick 10 that are good than to have people make 10 good images, or even 100. Even if you're going to make 500 images and pick 100 that are good, that's a lot more efficient than having a person make 100 images.

But on the other hand, if you're doing something like data entry—and I wrote something about this about OpenAI's launch of Deep Research—its whole marketing case is that it goes off and collects data about the mobile market. I used to be a mobile analyst. The numbers are all wrong. Its use case is, “Look how useful this is,” but the numbers are wrong.

In some cases, they're wrong because they've literally transcribed the number incorrectly from the source. In other cases, they're wrong because they've used a source that they shouldn't have used. But if I'd asked an intern to do it for me, an intern would probably have picked that. And to my point about verification, if I'm going to ask a machine to copy 200 numbers out of 200 PDFs and then I'm going to have to check all 200 of those numbers, I might as well just do it myself.

Erik Torenberg

Yeah. So you've got a whole swirling matrix of how you map this against existing problems. But the other side of it is: how do you map this against new things that you couldn't have done before? And this comes back to my point about platform shifts, because I see people looking at ChatGPT or looking at generative AI and saying, “Well, this is useless because it makes mistakes.” I think that's like looking at an Apple II in the late '70s and saying, “Could you use these to run banks?” To which your answer is no, but that's kind of the wrong question.

Benedict Evans

Right. And a lot of the question is, okay, it may not be very good at doing—there's a class of old tasks that generative AI is good at.

Erik Torenberg

There's also a lot more old tasks that generative AI is maybe not very good at. But then there's a whole bunch of other things that you would never have done before that generative AI is really, really good at. And then how do you find those or think of those? And how much of that is the user thinking of it, faced with a general-purpose chatbot? How much of that is the entrepreneur saying, “Hey, I've just realized that there's this thing that I can do that you couldn't do before, and here you are. I've given you a product with a button that will do it for you”?

Benedict Evans

Right.

Erik Torenberg

And it's why there are software companies, right?

Benedict Evans

Right.

Erik Torenberg

And on mobile, some of the new use cases were getting in strangers' cars—we mentioned Lyft and Uber—or dating people you met via an app, or lending your spare bedroom out, et cetera. Those were net-new companies that were built around those behaviors.

And I think for AI, there are still questions of what those net-new behaviors are. We're starting to see some in terms of people engaging and talking with chatbots instead of humans, or in addition to humans. And then there's a question of whether these are done by the model providers that currently exist, or by net-new companies, both in enterprise and consumer.

Benedict Evans

Well, this is always a question: How far up the stack does the new thing go? I was talking about this with another former a16z person, who pointed out that in the mid-'90s, people kind of argued that the operating system does all of it, and the Windows apps are basically just thin Win32 wrappers.

And Office is basically just a thin Win32 wrapper. All the important stuff is being done by the OS, whether it's document management, printing, storage, and display—all stuff that used to be done by apps. On DOS, the apps had to do printing; the apps had to manage the display. We moved to Windows, and 90% of the stuff that the app used to do is now being done by Windows, so Office is just like a thin Win32 wrapper, and all the hard stuff has been done by the OS. And it turns out, well, that was again—frameworks are useful, but that's maybe not a useful way of thinking about what's going on.

And the same thing now: How much does this need a single, dedicated understanding of how that market works, or what that market is, and what you would do with that? I remember when we were at a16z, there was an investment in a company called Everlaw, which is legal discovery in the cloud.

Erik Torenberg

Yeah.

Benedict Evans

And so machine learning happens, and now they can do translation. Are they worried that lawyers are going to say, “Well, we don't need you guys anymore. We're just going to go and get a translation app and a sentiment analysis app from AWS”? No, that's not how law firms work. Law firms want to buy a thing that sells legal discovery software and management. They don't want to go and write their own or do API calls. Very, very big law firms might, but a typical law firm isn't going to do that. People buy solutions; they don't buy technologies.

And the same thing here: How far up the stack do these models go? How much can you turn things into a widget? How much can you turn things into an LLM request? And how much does it turn out that you need that dedicated UI?

The funny thing is you can see this around Google, because Google had this whole idea that everything would just be a Google query and Google would work out what the query was. And guess what? Google Flights is not a Google query.

One of the interesting things about this is thinking about what a GUI is doing. The obvious thing a GUI is doing is that it enables Office to have 500 applications, 500 features, and you can find them all. At least you don’t have to memorize keyboard commands. You can now have effectively infinite features, and you can just keep adding menus and dialog boxes. Eventually you run out of screen space for dialog boxes, but you can have hundreds of features without people needing to memorize keyboard commands.

But the other side is you’re in that dialog box or screen in that workflow in Workday or Salesforce or whatever the enterprise software is, or the airline website or Airbnb. There aren’t 600 buttons on the screen. There are 7 buttons because people at that company have thought: what should the user be asked here? What questions should we give them? What choices should there be at this point in the flow? That reflects institutional knowledge, learning, testing, and careful thought about how this should work.

Then you give somebody a raw prompt and say, “Tell the thing how to do the thing,” and you’ve got to think from first principles: how does all this work? I always used to talk about machine learning as giving you infinite interns. Imagine you’ve got a task and an intern, and the intern doesn’t know what venture capital is. How helpful are they going to be? They don’t know that companies publish quarterly reports, that we’ve got a Bloomberg account that lets us look up multiples, that you should probably use PitchBook for this data rather than Google. This is my point about Deep Research: you should use this source and not that source. Do you want to work that out from scratch, or do you want people who know a lot about this stuff to have spent 5 years working out what the choices should be on the screen for you to click on? It’s the old user-interface saying: the computer should never ask you a question that it should know by itself. You go to a blank raw chatbot screen, and it’s asking you literally everything. It’s not just asking you one question; it’s asking you absolutely everything about what you want and how you’re going to work out how to do it.

Erik Torenberg

And so, you're mentioning ChatGPT, right, about how ChatGPT isn't so much a product as a chatbot disguised as a product. I am curious: When we look back at this platform shift, do you think that there will be another iPhone-esque or Excel-esque product that kind of defines the future—the platform shift—in a way that ChatGPT won't? Or is it that the world has to catch up to how to use ChatGPT, or something like ChatGPT?

Benedict Evans

So both of these can be true, because it took time to realize how you would use Google Maps and what you could do with Google and how you could use Instagram, and all of these products have evolved a huge amount over time. So some of it is that you grow toward realizing what you could do with this. You realize that's just a Google query now. You realize that you could just do it like that, and you realize, “I spent hours doing this, and I just realized, oh, I could actually just make a pivot table.”

The other side of it is that you're still then expecting people to work it out themselves from first principles. And it's kind of useful to have 100, 1,000, 10,000 really clever people sitting and trying to work out what those things are and then showing it to you as a product. I think another side to this is that there were always these precursors. There were lots of other things before Instagram.

Erik Torenberg

Yeah.

Benedict Evans

You know, YouTube didn't start as YouTube. It started as video dating, I think. There were lots of attempts to do online dating that all kind of worked until Tinder kind of pulled the whole thing inside out. And so there were always lots of things—what's the phrase? Local maxima. In fact, this is where we were, particularly with the iPhone, before, because I was working in mobile for the previous decade.

It didn't feel like we were waiting for a thing. It felt like it was kind of working: Every year the networks got faster, the phones got better, and it got a little bit better every year. We had apps, we had app stores, we had 3G, we had cameras, and stuff seemed to be a bit better every year. And then the iPhone arrives, and it just blows the chart: You've got this line doing this, and then there's a line that does that. Although remember, also, the iPhone took like 2 years before it worked, because the price was wrong, the feature set was wrong, and the distribution model didn't quite work.

And so, yeah, you can think everything's going well, and then something comes along and you realize, “No, oh, no, no, no.” That's the same for Google: Search was a thing before Google; it just wasn't very good. There was lots of social stuff before Facebook, and that was the thing that catalyzed it. So I just think, deterministically, this whole thing is so early that it feels like, of course, there are going to be dozens, hundreds of new things. Otherwise, a16z should just kind of shut down and give the money back to the LPs, because the foundation models will just do the whole thing.

And I don't think you're going to do that. At least I hope not.

Erik Torenberg

No, no, no. If we have any regrets from the last few years, it's not going bigger. I think we didn't fully appreciate how much specialization there would be across whether it's voice or image generation, or take any sort of subsector, that there would be net-new companies created that would be better than the model providers, that there would be even multiple model providers in every category.

One thing we've always—in the Web 2 era, we always bet on the category winner, and the category winner would take most of the market. But these markets are so big, and there's so much expertise and specialization, that there can be winners in every category. It's not just that the model providers take everything; even in every category, including the model providers, there can be multiple winners, with increasing specialization, and the markets are just big enough to contain multiple winners.

Benedict Evans

I think that's right. And I think the categories themselves aren't clear, right?

Erik Torenberg

And many things—you think this is a category, and it turns out, no, it was actually that whole other thing. The categories kind of get unbundled and bundled and recombined in different ways. I remember I was a student in 1995, and I think I had 4 or 5 different web browsers and web servers on my PC.

Tim Berners-Lee's original web browser had a web editor in it because he thought this was kind of like a network drive and a sharing system, and he didn't realize it wasn't really a publishing system. So, you would have your web pages on your PC, you'd leave your PC turned on, and that would be how your colleagues would look at your Word documents or your web pages.

So, again, we just don't know how. I keep coming back to this point: I feel like most of the questions we're asking at the moment are probably the wrong questions. Picking up on a strand within what you just said, though, one of the interesting things I'm thinking about a lot is looking at OpenAI.

Because I'm fascinated by disconnections, we've got this interesting disconnect now. If you look at the benchmark scores, you've got these general-purpose benchmarks where the models are basically all the same. If you're spending hours a day in them, then you've got this opinion about, "Oh, I like Claude's tone of voice more than I like ChatGPT, and I like GPT-5.1 more than GPT-4.9, or whatever the hell it's called." If you're using this once a week, you really don't notice this stuff. The benchmark scores are all roughly the same, but the usage isn't.

Basically, Claude has no consumer usage, even though on the benchmark score it's the same. Then it's ChatGPT, and halfway down the chart it's Meta and Google. The funny thing is, you read all the AI newsletters, and Meta's lost, they're out of the game, they're dead. Mark Zuckerberg is spending $1 billion per researcher to get back in the game. But from the consumer side, well, it's distribution.

The interesting thing here is that what I'm kind of circling around is: if the model for a casual consumer user certainly is a commodity, and there are no network effects or winner-takes-all effects yet—those may emerge, but we don't have them yet—and things like memory aren't network effects; they're stickiness, but they can be copied, how is it that you compete?

Do you just compete on being the recognized brand and adding more features and services and capabilities, and people just don't switch away? Which is kind of what happened with Chrome, for example. There's not a network effect for Chrome, but it's not actually much better—maybe it's a bit better than Safari—but you use Chrome because you use Chrome.

Or is it that you get left behind on distribution or network effects that emerge somewhere else, and meanwhile you don't have your own infrastructure? I suppose what I'm getting at is: you've got these 800 or 900 million weekly active users, but that feels very fragile because all you've really got is the power of the default and the brand.

You don't have a network effect. You don't really have feature lock-in. You don't have a broader ecosystem. You also don't have your own infrastructure, so you don't control your cost base. You don't have a cost advantage. You get a bill every month from Satya.

So you've kind of got to scramble as fast as you can in both of those directions: on the one side, build product and build stuff on top of the model, which is our earlier conversation. Is it just the model? Yeah.

Benedict Evans

Now, you've got to build stuff on top of the model in every direction. It's a browser. It's a social video app. It's an app platform. It's this; it's that. It's like the meme of the guy with the map with all the strings on it. It's all of these things: we're going to build all of them yesterday.

And then in parallel, it's infrastructure. We've got to deal with NVIDIA, Broadcom, AMD, Oracle, and petrodollars. Because you're kind of scrambling to get from this amazing technical breakthrough and these 800 or 900 million weekly active users to something that has really sticky, defensible, sustainable business value and product value.

Erik Torenberg

Yeah. And so, as you're evaluating the competitive landscape among the hyperscalers, what are the questions that you're asking that you think are going to be most important in determining who's going to gain durable competitive advantages, or how this competition is going to play out?

Well, this kind of comes back to your point about sustaining advantage. We talked about Google: if we think about the shift to mobile—particularly the shift to mobile for Meta—this turned out to be transformative. It made the products way more useful.

Benedict Evans

Yeah.

Erik Torenberg

For Google, it turned out mobile search is just search.

Benedict Evans

And Maps changed, probably, and YouTube changed a bit, but basically, for Google, search is search. Web search just means more people doing more search, more of the time. Yeah.

Erik Torenberg

And the default view now would seem to be: Gemini is as good as anybody else. Next week, the new model—I haven't looked at the benchmarks for GPT-5.1, which is out today. Is it better than Gemini? Probably. Will it still be better next month? No.

So that's a given: you've got a frontier model. Fine. What does that cost? It costs you—pick a number—$250 billion a year, $100 billion a year. What's this? This is our earlier conversation about capex.

Okay, so Google can pay that because they've got the money. They've got the cash flow from everything else. And so you do that, and your existing products get you to optimize search. You optimize your ad business. You build new experiences. Maybe you invent the new iPhone of AI. Maybe there is no iPhone of AI. Maybe someone else does it, and you do an Android and just copy it.

So, fine, it's the new mobile. We'll just carry on. Search is search. AI is AI. We'll do the new thing. We'll make it a feature. We'll just carry on doing it.

For Meta, it feels like there are bigger questions about what this means for search, or what it means for content and social and experience and recommendation, which makes it all the more imperative that they have their own models, just as it is for Google.

For Amazon, okay, on the one side, it's commodity infrastructure, and we'll sell it as commodity infrastructure. And on the other side—maybe we can step back—if you're not a hyperscaler, if you're a web publisher, a marketer, a brand, an advertiser, or a media company, you could make a list of questions, but you don't even know what the questions are right now.

Benedict Evans

What is this? What happens if I ask a chatbot a thing instead of asking Google? Even if it's Google—from Google's point of view—well, I'll ask Google's chatbot. It's fine. But as a marketer, what does that mean?

What happens if I ask for a recipe and the LLM just gives me the answer? What does that mean if my business is having recipes?

Erik Torenberg

Yeah.

Benedict Evans

Do you have a kind of split? And this is also an Amazon question: how does a purchasing decision happen? How does this decision to buy a thing that I didn't know existed before happen? What happens if I wave my phone at my living room and say, "What should I buy?" Where does that take me, in ways that it wouldn't have taken me in the past?

Erik Torenberg

Yeah.

Benedict Evans

Do LLMs mean that Amazon can finally do really good recommendation, discovery, and suggestion at scale, in ways that it couldn't really do in the past because of this kind of pure commodity retailing model that it has?

Apple's sort of off on one side. Interestingly, they produced this incredibly compelling vision of what Siri should be 2 years ago. It just turned out that they couldn't make it. Interestingly, nobody else could have made it either.

You go back and watch the Siri demo that they gave and you think, okay, so we've got multimodal, instantaneous, on-device, tool-using, agentic, multiplatform e-commerce in real time, with no prompt-injection problems and zero error rates. Well, that sounds good. Has anyone got that working? No.

OpenAI and Google don't have that working. I don't think Google or OpenAI could deliver the Siri demo that Apple gave 2 years ago. They could probably do the demo, but they couldn't consistently and reliably make it work. That demo, that product, isn't in Android today.

Apple, to me, has the most intellectually interesting question. I saw Craig Federighi make this point: “We don't have our own chatbot. Fine. We also don't have YouTube or Uber. Explain why that is different.” That's a harder question to answer than it sounds like.

Of course, the answer is: If this fundamentally changed the nature of computing, then it's a problem. If it's just a service that you use, like Google, then that's not a problem. That's kind of the point about where Siri goes.

The interesting counterexample here would be to think about what happened to Microsoft in the 2000s. The entire development environment gets away from them, and no one builds Windows apps after 2001 or something. But you need to use the internet, and to use the internet, you need a PC. What PC are you going to buy? Apple wasn't really a player at that time and was just getting back into the game. Linux obviously wasn't an option for any normal person, so you bought a Windows PC.

Basically, Microsoft loses the platform war and sells an order of magnitude more Windows PCs as a result of this thing that Microsoft lost. It takes until mobile for them to lose the device as well as the development environment. So here's the question: If all the new stuff is built on AI and I'm accessing an app that I download from the App Store, to what extent is this a problem for Apple? You would need a much more fundamental shift in what was happening for that to be a problem for Apple.

Even if you take not the full rapture arriving and we all go to sleep in pods like the guys in WALL-E—maybe we'll be those people; maybe we'll be like that—in which case, fine. There's a sort of a mid-case, which is that the whole nature of software changes, there are no apps anymore, and you just go and ask the LLM a thing. Fine. What is the device on which you ask the LLM a thing?

It's probably going to have a nice, big color screen and a 1-day battery life. It probably needs a microphone and a good camera. It kind of sounds like an iPhone. Am I going to buy the one that's a tenth of the price and just use the LLM on it? No, because I'll still want the good camera, the good screen, and the good battery life.

There are a bunch of interesting strategic questions when you start poking away. What does this mean for Amazon? Those are completely different questions from what it means for Google, Apple, Facebook, or Salesforce. What does it mean for Uber? Right back to what we were saying at the beginning of this conversation: What does this mean for Uber? Their operations get X% more efficient, and now the fraud detection works. Maybe their autonomous cars—that's a different conversation, but presume no autonomous cars. Otherwise, as Uber, what does this change? Not a huge amount.

Erik Torenberg

You've been doing these presentations for a while now. You bumped them up to 2 times because there's so much changing. One of the things you do in each presentation is ask really great questions and chronicle what the important questions are to be asking.

I'm curious, as you reflect—maybe post-ChatGPT in 2022, or GPT-3, rather—the questions you were asking then and compare them to now: To what extent do we have some direction on some of those questions? To what extent are they the same questions, or new and different questions? If I woke up in a coma after reading your original presentation—let's say the one after the GPT-3 launch came out—and then saw this one now, what were the most surprising things, or the things that we learned that updated those questions?

Benedict Evans

I think we have a lot of new questions this year. You could make a list of what might be half a dozen questions in spring of 2023: open source, China, NVIDIA, does scaling continue, what happens to images, and how long does OpenAI's lead remain?

Those questions didn't really change in 2023 and 2024, and most of those questions are still there. The NVIDIA question hasn't really changed. The answer on China, the answer on how many models there will be—the answer is, okay, anybody who can spend a couple hundred million dollars can have a frontier model. That was pretty obvious in early 2023. It took a while for everyone to understand that.

And big models and small models: Will we have small models running on devices? No, because the capabilities keep moving too fast for small models to shrink down onto the device. Those questions kind of didn't change for 2 or 2½ years.

I think we now have a bunch of more product-strategy questions, as you see real consumer adoption and OpenAI and Google building things in different directions, Amazon going in different directions, and Apple trying—and obviously failing—and then trying again to do things. There's a sense that there is something more going on in the industry than just, “Let's build another model and spend more money.”

Erik Torenberg

Yeah.

Benedict Evans

There are more questions and more decisions. Now there are also more questions outside of tech, certainly on the retail-media side, about how you start thinking about what you would do with this.

The classic framing in my deck is that step 1 is you make it a feature, absorb it, and do the obvious stuff. Step 2 is you do new stuff. Step 3 is maybe someone will come and pull the whole industry inside out and completely redefine the question.

You could imagine step 1 as you're a manager at a Walmart in the Bay Area or D.C., or whatever it is: “Find me that metric.” Step 2: “Build me a dashboard.” Step 3: “It's Black Friday, and I'm managing a Walmart outside D.C. What should I be worried about?”

That might be the wrong example, but step 1 for Amazon is that you bought light bulbs, so here's some packing tape. What Amazon should actually be doing is saying, “This person is moving home. We'll show them a home-insurance ad,” which is something Amazon's correlation systems wouldn't get because they wouldn't have that in their purchasing data.

We're still starting to think about step 1, but what would step 2 and step 3 be? What would new revenue be for this, other than simple, dumb automation? What new things would we build with this? Where might this actually redefine or change what the market looks like? That's obviously a big question for anyone in the content business.

Erik Torenberg

Yeah.

Benedict Evans

What does it mean if I can just go and ask an LLM this question? What kinds of content were predicated on Google routing that question to you? What kind of content isn't really about that question?

Do I want a Bolognese recipe, or do I want to hear Stanley Tucci talking about cooking in Italy? Do I just want the SKU, or do I want to work out which product I should buy? Amazon is great at getting you the SKU, but terrible at telling you what SKU you want.

Do I just want the slide deck, or do I want to spend a week talking to a bunch of partners from Bain about how I could think about doing this? Do I just want money, or do I want to work with a16z's operating groups?

Erik Torenberg

What is it that I'm doing here?

Benedict Evans

I think the LLM thing is starting to crystallize that question in lots of different ways. What am I actually trying to do here? Do I just want a thing that a computer can now answer for me, or do I want something else that isn't? The LLMs can do a bunch of stuff that computers couldn't do before, right?

Erik Torenberg

Right?

Benedict Evans

Is that thing that the computer couldn't do before my business?

Erik Torenberg

Yeah.

Benedict Evans
Erik Torenberg

We're about to figure out, in a much more granular way, what the true job to be done is for many, many of these.

Benedict Evans

Yeah. Going back to the internet, there was the observation about newspapers: Newspapers looked at the internet and talked about expertise, curation, journalism, and everything else. They didn't really say, “Well, we're a light-manufacturing company and a local-distribution and trucking company.”

Erik Torenberg

Yeah.

Benedict Evans

That was the bit that was the problem. Until the internet arrived, that wasn't a conversation you thought about. Then the internet suddenly makes that clear and suddenly creates an unbundling that didn't exist before.

And so there will be those kinds of things where you didn't realize you were that before, until someone comes along with an LLM and says, “I can use this to do this thing that you didn't really realize was the basis of your defensibility or the basis of your profitability.” I mean, it's like the joke about US health insurance: the basis of US health insurance profitability is making it really, really boring, difficult, and time-consuming. That's where the profits come from. Maybe it isn't—I don't know that industry—but, for the sake of argument, say that's your defensibility. Well, an LLM removes boring, time-consuming, mind-numbing tasks.

So, what industries are protected by having that? And they didn't realize that. And, you know, it's like you could have asked these questions about the internet in the mid-'90s or about mobile a decade later. Generally, half of the questions you asked would have been the wrong questions in hindsight. I mean, I remember, as a baby analyst in 2000, everyone kept saying, “What's the killer use case for 3G? What's a good use case for 3G?” And it turned out that having the internet in your pocket everywhere was the use case for 3G.

Erik Torenberg

But that wasn't the question people were asking. And I'm sure that will be the thing now: so much will happen and get built where you go and realize, “Oh, that's how you would do this. You can turn it into that.”

Benedict Evans

Yeah.

Erik Torenberg

No, 100%. My last question to get you out of here: if we're talking 2 or 3 years from now, and you're doing a presentation and you say, “Oh, this is actually bigger than the internet,” or maybe, “This is like computing,” what would need to be true? What would need to happen? What would evolve our thinking?

Benedict Evans

I mean, I kind of come back to my point about the Jews and Christians: the Messiah came, nothing happened. We forget—I mean, there are maybe 2 very brief ways to think about this. One of them is that I think we forget how enormous the iPhone was and how enormous the internet was. You can still find people in tech who claim that smartphones aren't a big deal. And this was the basis of people complaining about me: “This idiot thinks generative AI is as big as those silly phone things. Come on.”

I think another answer would be that I don't want to get into the argument about what the reasoning capability is, the benchmarks, and all that. You see lots of 5-hour-long podcasts of people talking about this stuff, but the stuff we have now is not a replacement for an actual person outside of some very narrow and very tightly constrained guardrails, which is why I agree with Demis's point that it's absurd to say that we have Ph.D.-level capabilities now. What we would have to be seeing is something that would really shift our perception of the capability of this stuff.

Erik Torenberg

Yeah.

Benedict Evans

So that it's actually a person, as opposed to something that can kind of do these people-like things really well sometimes, but not other times. And it's a very tough conceptual kind of thing to think about because I'm deliberate. I'm conscious that I'm not giving you a falsifiable answer. But I'm not sure what a falsifiable answer would be to that. When would you know whether this was AGI?

You know, it's the Larry Tesler line: AI is whatever doesn't work yet. As soon as people say it works, people say, “Well, that's just not AI. That's just software.” It becomes a slightly drunk philosophy grad student kind of conversation as much as it is a technology conversation. Like, have you ever considered, Erik, that—

Erik Torenberg

Maybe we're not either.

Benedict Evans

That's a thought. All I can say to give a tangible answer to this question is that what we have right now isn't that. Will it grow to that? We don't know. You may believe it will. I can't tell you that you're wrong. We'll just have to find out.

Erik Torenberg

I think that's a good place to wrap. The presentation is “AI Eats the World.” We'll link to it. It's fantastic. Benedict, thanks so much for coming on the podcast to discuss it.

Benedict Evans

Sure. Thanks a lot, Erik.

AI 吞噬世界:Benedict Evans 谈下一轮平台变迁 — 文字稿与摘要 | BidClub