Benedict Evans
ChatGPT has 800 or 900 million weekly active users. If you’re the kind of person who’s using this for hours every day, ask yourself why 5 times more people look at it, get it, know what it is, have an account, know how to use it, and can’t think of anything to do with it this week or next week. The term AI is a little bit like the term technology. When something’s been around for a while, it’s not AI anymore. Is machine learning still AI? I don’t know. In actual general usage, AI seems to mean new stuff, and AGI seems to mean new, scary stuff.
Erik Torenberg
AGI seems to be a bit like this: either it’s already here and it’s just software, or it’s 5 years away and will always be 5 years away. We don’t know the physical limits of this technology, and so we don’t know how much better it can get. You’ve got Sam Altman saying we’ve got PhD-level researchers right now, and Demis Hassabis says, “No, we don’t. Shut up.” Very new, very big, very exciting, world-changing things tend to lead to bubbles. So, yeah, if we’re not in a bubble now, we will be.
Benedict, welcome back to the a16z podcast.
Benedict Evans
Good to be back.
Erik Torenberg
We’re here to discuss your latest presentation, “AI Eats the World.” For those who haven’t read it yet, maybe you can share the high-level thesis and contextualize it in light of recent AI presentations. I’m curious how your thinking has evolved.
Benedict Evans
Yeah, it’s funny. One of the slides in the deck references a conversation I had with a big-company CMO who said, “We’ve all had lots of AI presentations now. We’ve had the Google one and the Microsoft one. We’ve had the Bain one and the BCG one. We’ve had the one from Accenture and the one from our ad agency. So now what?”
It’s sort of 90-odd slides, and there are a bunch of different things I’m trying to get at. One of them is to say: if this is a platform shift, or more than a platform shift, how do platform shifts tend to work? What are the things that we tend to see in them, and how many of those patterns can we see being repeated now?
Some of the patterns are things like bubbles, but others are that lots of stuff changes inside the tech industry. There are winners and losers; people who were dominant end up becoming irrelevant, and then there are new billion- and trillion-dollar companies created. But there’s also the question of what this means outside the tech industry, because if we think back over the last waves of platform shifts, there were some industries where this changed everything and created and destroyed industries, and others where this was just a useful tool.
If you’re in the newspaper business, that had a very different impact. The last 30 years look very different from if you were in the cement business, where the internet was just kind of useful but didn’t really change the nature of your industry very much.
What I tried to do is give people a sense of what’s going on in tech: how much money we’re spending, what we’re trying to do, what the unanswered questions are, and what might or might not happen within the tech industry. But then, outside technology, how does this tend to play out? What seems to be happening at the moment? How is this manifesting in tools and deployment, new use cases, and new behaviors?
As we step back from all of this, how many times have we gone through all of this before? It’s funny: I went on a podcast this summer, and my opening line was something like, “Well, I’m a centrist. I think this is as big a deal as the internet or smartphones, but only as big a deal as the internet or smartphones.” There were about 200 YouTube commenters underneath saying, “This is more, and he doesn’t understand how big this is.” And I think, well, it was kind of a big deal.
Erik Torenberg
It was kind of a big deal.
Benedict Evans
I finish the day by looking at elevators, because I live in an apartment building in Manhattan and we have an attended elevator. That means there’s a person; there are no buttons, there’s an accelerator and a brake, and the doorman gets in and drives you to your floor. It’s a streetcar.
In the 1950s, Otis deployed automatic elevators. You get in and press a button, and they marketed it by saying, “Ah, it’s got electronic politeness,” which meant the infrared beam. Today, when you get into an elevator, you don’t say, “Ah, I’m using an electronic elevator. It’s automatic. It’s just a lift.”
That’s what happened with databases, the web, and smartphones. Databases certainly aren’t AI. I’ve done a couple of polls on this on LinkedIn and Threads, asking, “Is machine learning still AI?” I don’t know. There’s obviously an academic definition where people say, “This guy’s an idiot.” Of course, I’m going to explain the definition of AI, but in actual general usage, AI seems to mean new stuff.
Erik Torenberg
Yeah, and AGI seems like new, scary stuff.
Benedict Evans
Yeah, it’s funny. There’s an old theologian’s joke that the problem for Jews is that you wait and wait and wait for the Messiah, and he never comes. The problem for Christians is that he came and nothing happened. The world didn’t change; there is still sin. For all practical purposes, nothing happened.
AGI seems to be a bit like this. Either it’s already here, and it’s just more software, so you’ve got Sam Altman saying we’ve got PhD-level researchers right now, and Demis Hassabis says, “No, we don’t. Shut up.” Or it’s 5 years away and will always be 5 years away.
Erik Torenberg
Yeah, yeah, it’s interesting. Let’s compare this with previous platform shifts, because some people look at something like the internet and say, “Hey, there were net-new trillion-dollar companies—Facebook and Google—that were created from it, and all sorts of new emerging winners.”
Whereas they look at mobile and say, “There were big companies like Uber, Snap, Instagram, and WhatsApp, but these were billion-dollar outcomes or tens-of-billions-of-dollar outcomes. Really, the big winners were, in fact, Facebook and Google.”
In some sense, mobile perhaps was sustaining. Feel free to quibble with the definition of sustaining versus disruptive, but sustaining in the sense that maybe more of the value went to incumbents, or companies that existed prior to the shift.
I’m curious how you think about AI in light of that. Is it enabling? Are more of the gains going to come from net-new companies like OpenAI and Anthropic and others that follow? Or are more of the gains going to be captured by Microsoft, Google, Facebook, and Meta—companies that existed prior to the shift?
Benedict Evans
There are several answers to this. One of them is that you kind of have to be careful about framings and structures and things, because you end up arguing about the framing and the definition rather than arguing about what’s going to happen. They’re all useful, but they’ve all got holes in them.
What mobile did was shift us in several fundamental ways. It shifted us from the web to apps, for example, and it gave everybody in the world a phone. It gave everybody in the world a pocket computer. Even today, there are fewer than 1 billion consumer PCs on Earth, and there are somewhere between 5 billion and 6 billion smartphones.
It made possible things that would not have been possible without it, whether that’s TikTok or, arguably, things like online dating. You can map those against dollar value, but you can also map them against structural change in consumer behavior and access to information. You could certainly argue that Meta would be a much smaller company if it weren’t for mobile, for example. You can argue the puts and calls on this stuff a lot.
Not all platform shifts are the same, and you can do the standard typology: there were mainframes, then PCs, then the web, and then smartphones. But you kind of want to put SaaS in there somewhere, and you kind of want to put open source in there. Maybe you want to put databases in there, too.
These are useful framings, but they’re not predictive. They don’t tell you what’s going to happen; they just give you one way of understanding some of the patterns we’ve seen.
The big debate around generative AI is whether this is just another platform shift or something more than that. The problem is that we don’t know, and we don’t have any way of knowing other than waiting to see. This may be as big as PCs, the web, SaaS, or open source—or it may be as big as computing itself.
Then you’ve got the very overexcited people living in group houses in Berkeley who think this is as big as fire or something. Well, great.
But does this create new companies? I mean, you go back to mobile. There was a time when people thought blogs were going to be different from the web, which seems weird now. Google needed a separate blog search.
This seriously was a thing. There was a time when it was really not clear, and I think you can kind of generalize his point. You go back to the internet in the mid-’90s: we kind of knew this was going to be a big thing, but we didn’t really know it was going to be the web. Before that, we didn’t know it was going to be the internet. We knew there were going to be networks, but we didn’t know it was going to be the internet.
Then it wasn’t clear that it was going to be the web, and it wasn’t really clear how the web was going to work. When Netscape launched, Mark Zuckerberg was in junior high or elementary school or something, Larry and Sergey were students, and Amazon was a bookstore. So you can know it but not know it.
You could make the same point about smartphones. We knew everyone was going to have an internet-connected thing in their pocket, but it wasn’t clear that it was basically going to be a PC from the PC company from the ’80s and a search engine company. It wasn’t clear that it wasn’t going to be Nokia or Microsoft. I think you have to be super careful about making deterministic predictions about this. What you can do is say, “When this stuff happens, everything changes,” and that’s happened 5 or 10 times before.
Erik Torenberg
I’m curious how you got conviction in this idea, or what got you to the prediction that AI is going to be as big as the internet—which, of course, is pretty big. I’m not yet, Benedict, at the conviction that it’s going to be any bigger. I’m curious what sort of inspires that statement, and what might change your mind either way: that it might not be as big as the internet, because the internet was obviously very big, but also that perhaps it might be bigger.
Benedict Evans
I don’t want to—I remember I made a diagram of S-curves going up slightly, and someone said, “Well, what’s the axis on this diagram?” I don’t want to get into whether this is 5% bigger than the internet or 20% bigger. I think the question is more like: Is it another of these industry cycles, or is it a much more fundamental change in what technology can be? Is it more like computing or electricity, as a sort of structural change, rather than, “Here’s a whole bunch more stuff we can do with computers”?
I think that’s the question, and there’s a funny disconnect in looking at debates about this within tech. I watched one of the OpenAI livestreams a couple of weeks ago, and they spent the first 20 minutes talking about how they were going to have human-level, PhD-level AI researchers next year. Then the second half of the stream was, “Here’s our API stack that’s going to enable hundreds and thousands of new software developers, just like Windows,” and they literally quoted Bill Gates. You think, “Those can’t both be true.” Either I’ve got a thing that is a PhD-level AI researcher—which, by implication, is like a PhD-level CPA—
Erik Torenberg
Yeah.
Benedict Evans
—or I’ve got a new piece of software that does my taxes for me. Which is it? Either this thing is going to be human-level, and that’s a very, very challenging, problematic, complicated statement, or this is going to let us make more software that can do more things than software could do before.
I think there’s a real schizophrenia in conversations around this, because it’s, “Scaling laws, and it’s going to scale all the way,” while meanwhile I’m hearing, “Look how good it is at writing code.” Again, is it writing code, or do we not need software anymore? Because, in principle, if the models keep scaling, nobody’s going to write code anymore. You’ll just say to the model, “Hey, can you do this thing for me?”
Erik Torenberg
Is it a little bit of a hedge, or is it a sequencing thing?
Benedict Evans
Some of it’s a sequencing thing, but in principle, if you think this stuff is going to keep scaling, why are you investing in a software company?
Erik Torenberg
Yeah.
Benedict Evans
We’ll just have this god in a box that can do everything. I think this is the funny challenge, and I think this is the fundamental way that this is different from previous platform shifts. With the internet, or with mobile, or even with mainframes, you didn’t know what was going to happen in the next couple of years. You didn’t know what Amazon would become, you didn’t know how Netscape was going to work out, and you didn’t know what next year’s iPhone was going to be.
Ten years ago, when we cared about that, you kind of knew the physical limits. You knew, in 1995, that telcos were not going to give everybody gigabit fiber the next year, and you knew that the iPhone wasn’t going to have a year’s battery life, unroll, have a projector, and fly or whatever. But we don’t know the physical limits of this technology because we don’t really have a good theoretical understanding of why it works so well. Nor, indeed, do we have a good theoretical understanding of what human intelligence is, and so we don’t know how much better it can get.
You could do a chart and say, “Well, this is the road map for modems, and this is the road map for DSL, and this is how fast DSL will be.” Then you could make some guesses about how quickly telcos will deploy DSL, and say, “Clearly, we’re not going to be able to replace broadcast TV with streaming in 1998.” But we don’t have an equivalent way of modeling this stuff to know what its fundamental capability is going to look like in 3 years, which gets you to these slightly vibes-based forecasts where no one really knows.
Geoffrey Hinton says, “Well, I feel like,” and Demis Hassabis says, “Well, I feel like,” but no one knows.
Erik Torenberg
And then Andrej Karpathy goes on our podcast and says, “I feel like it’s a decade out.”
Benedict Evans
Yeah, I know. I saw this meme of—what’s his name?—saying, “The answer will reveal itself.” Somebody like me would say it’s photoshopped, but of course it wouldn’t have been photoshopped. Somebody had turned him into a Buddhist monk wearing an orange outfit: “The future will reveal itself.”
But this is the problem. We don’t know, and we don’t have a way of modeling this.
Erik Torenberg
Yeah. And so let’s connect this to the upfront investment that some of these companies are making. We don’t know—is there a risk of overinvestment leading to some potential bubble-like mechanics? How do you think about that question?
Benedict Evans
Well, deterministically, very new, very big, very exciting, world-changing things tend to lead to bubbles.
Erik Torenberg
Yeah.
Benedict Evans
I don’t think anybody would dispute that you can see some bubbly behavior now. You can argue about what kind of bubble, but again, that doesn’t have very much predictive power. One of the features of bubbles is that when everything’s going up, everything goes up all at once, everyone looks like a genius, and everyone leverages and cross-leverages and does circular revenue. That’s great until it isn’t, and then you get a kind of ratchet effect as it goes back down again.
If we’re not in a bubble now, we will be. I remember Marc Andreessen saying, “1997 was not a bubble. 1998 was not a bubble. 1999 was a bubble.” Are we in 1997 now, or 1998, or 1999? If we could predict that, we’d live in a parallel universe.
I think there are maybe 2 more specific, more tangible answers to this. The first is that we don’t really know what the compute requirements of this stuff are going to be. Forecasting that feels a lot like trying to forecast bandwidth use in the late ’90s. Imagine if you were trying to do the algebra on that: This many users, how much bandwidth does a web page use? How will that change? How will that change as bandwidth gets faster? What happens with video? What kind of video? What bit rate of video? How long do people watch a video? How much video?
You could build the spreadsheet, and it would tell you what global bandwidth consumption would be in 10 years. Then you could try to use that to back-calculate how many routers this is going to sell. You could get a number, but it wouldn’t be the number. There’d be a hundredfold range of possible outcomes from that. You could make the same point about the algebra of consumption now.
Right now, we have a bunch of rational actors saying, “This stuff is transformative and a huge threat. We can’t keep up with demand for it now, and as far as we know, the demand is going to keep going up.” We’ve had a variety of quotes from all of the hyperscalers basically saying that the downside of not investing is bigger than the downside of overinvesting. That kind of thing always works well until it doesn’t.
Erik Torenberg
Yeah.
Benedict Evans
I saw a slightly strange quote from Mark Zuckerberg saying, “Well, if it turns out that we’ve overinvested, we can just resell the capacity.” I thought, “Let me just stop you there, Mark, because if it turns out that you can’t use your capacity, everybody else is going to have loads of spare capacity as well.”
Erik Torenberg
Yeah.
Benedict Evans
All these people who are desperate for more capacity—if it turns out we can get the same results for a hundredth of the compute—
Erik Torenberg
That will be true for everyone else too, not just you.
Benedict Evans
Yeah. So, in an investment cycle like this, you tend to get overinvestment, but after that, there are very limited predictions you can make about what's going to happen. I think the more useful way to look at this is: you've got these transformative capabilities that are already increasing the value of your existing products if you're Google, Meta, or Amazon, and you're going to be able to use them to build a bunch more stuff. Why would you want to let somebody else do that rather than you doing it, as long as you're able to keep funding and selling what you're building?
Erik Torenberg
Yeah.
Benedict Evans
And it may well turn out that we have an evolution of models in the next year that means you can get the same result for 1/100th of the compute that you're using today. Bearing in mind that it's already going down—depending, pick your numbers—20, 30, 40 times a year.
Erik Torenberg
Yeah.
Benedict Evans
But then the usage is going up. So you're in this very—as I said, it's like trying to predict bandwidth consumption in the late '90s and early 2000s. You can throw all the parameters out, but it doesn't get you to something useful. You just need to step back and say, “Yeah, but is this internet thing any good?”
Erik Torenberg
Well, yeah, because I'm curious if you see the bottlenecks as being more on the supply side or the demand side—more technical constraints—or is it just, is AI any good? Are there enough use cases to justify the type of spend? What are you seeing, and what are you predicting?
Benedict Evans
So, maybe 2 answers to this question. The first of them is, I think we've had this sort of bifurcation of what all the questions are. There are now very, very detailed conversations about chips, and then very, very detailed conversations about data centers and funding for data centers, and then about what a new enterprise SaaS company built on AI—what margins will it have and how much money does it need to raise. So there are venture capital conversations, and there are many different conversations within which I don't know anything about chips. I can spell “ultraviolet,” but I don't know what an ultraviolet process is. It's more violets, I don't know.
And so you've got this—it's like the Milton Friedman line: no one knows how to build a pencil. I think the second answer might be that there are 2 kinds of generative AI deployment. One of them is where it's very easy and obvious right now to see what you would do with this, which is basically software development, marketing, point solutions for many very boring, very specific enterprise use cases, and also people like us, who have very open, free-form, flexible jobs with many different things and who are always looking for ways to optimize that.
Erik Torenberg
Yeah.
Benedict Evans
And so you get people in Silicon Valley who are like, “I spend all my data time in dbt. I don't use Google anymore. I've replaced my CRM with this.” And then, obviously, people who write code—if you're writing code, this works really well. If you're in marketing, there are all these stories of big companies where they're making 300 assets where they would have made 30. And Accenture, Bain, McKinsey, Infosys, and so on are sitting and solving very specific problems inside big companies.
Then there's a whole bunch of other people who look at it and are like, “It's okay.” You go and look at the usage data, and you see that ChatGPT has 800 or 900 million weekly active users, and 5% of people are paying. Then you go and look at all the survey data, and it's very fragmented and inconsistent, but it all sort of points to something like 10% or 15% of people in the developed world using this every day. Another 20% or 30% of people are using it every week. If you're the kind of person who is using this for hours every day, ask yourself why 5 times more people look at it, get it, know what it is, have an account, know how to use it, and can't think of anything to do with it this week or next week.
Erik Torenberg
Why is that?
Benedict Evans
Yeah. Is it because it's early? It's not a young-people thing, either, incidentally. Is that just because it's early? Is it because of the error rates? Is it because you have to map it against what you do every day?
One of the analogies I always used to use—which isn't in the current presentation, but I've used it in previous presentations—is: imagine you're an accountant and you see spreadsheet software for the first time. This thing can do a month of work in 10 minutes, almost literally.
Erik Torenberg
Yeah. You want to change—you want to recalculate that 10-year DCF with a different discount rate. I've done it before you finished asking me to. And that would have been like a day or 2 days or 3 days of work to recalculate all those numbers. Great. Now imagine you're a lawyer and you see it. You think, “Well, that's great. My accountant should see it. Maybe I'll use it next week when I'm making a table of my billable hours, but that's not what I do all day.” Excel doesn't do things that a lawyer can do every day.
Benedict Evans
Yeah. So you've got a whole swirling matrix of how you map this against existing problems. But the other side of it is: how do you map this against new things that you couldn't have done before? And this comes back to my point about platform shifts, because I see people looking at ChatGPT or looking at generative AI and saying, “Well, this is useless because it makes mistakes.” I think that's like looking at an Apple II in the late '70s and saying, “Could you use these to run banks?” To which your answer is no, but that's kind of the wrong question.
Erik Torenberg
Right? Really? Isn't that what they're doing? They're unbundling ChatGPT, just as the enterprise software company of 10 years ago was unbundling Oracle or Google or Excel. Do you have the view that what Excel did for accountants, AI is now doing for coders and developers, but it hasn't quite figured out that daily critical workflow for other job positions? So it's unclear for people who aren't developers why they should be using this for many hours a day.
Benedict Evans
I think there's a lot of people who don't have tasks that work very well with this.
Erik Torenberg
Yeah. And then there's a lot of people who need it to be wrapped in a product and a workflow and tooling and UX, and someone to come and say, “Hey, have you realized you could do it with this?”
Benedict Evans
I had this conversation in the summer with Balaji Srinivasan, who's another former a16z person, and he was making this point about validation: Can you—because these things still get stuff wrong, and people in the Valley often hand-wave this away—there are questions that have specific answers where it needs to be the right answer, or one of a limited set of right answers. Can you validate that mechanistically? If not, is it efficient to validate it with people?
With the marketing use case, it's a lot more efficient to get a machine to make you 200 pictures and then have a person look at them and pick 10 that are good than to have people make 10 good images, or even 100. Even if you're going to make 500 images and pick 100 that are good, that's a lot more efficient than having a person make 100 images.
But on the other hand, if you're doing something like data entry—and I wrote something about this about OpenAI's launch of Deep Research—its whole marketing case is that it goes off and collects data about the mobile market. I used to be a mobile analyst. The numbers are all wrong. Its use case is, “Look how useful this is,” but the numbers are wrong.
In some cases, they're wrong because they've literally transcribed the number incorrectly from the source. In other cases, they're wrong because they've used a source that they shouldn't have used. But if I'd asked an intern to do it for me, an intern would probably have picked that. And to my point about verification, if I'm going to ask a machine to copy 200 numbers out of 200 PDFs and then I'm going to have to check all 200 of those numbers, I might as well just do it myself.
Erik Torenberg
Yeah. So you've got a whole swirling matrix of how you map this against existing problems. But the other side of it is: how do you map this against new things that you couldn't have done before? And this comes back to my point about platform shifts, because I see people looking at ChatGPT or looking at generative AI and saying, “Well, this is useless because it makes mistakes.” I think that's like looking at an Apple II in the late '70s and saying, “Could you use these to run banks?” To which your answer is no, but that's kind of the wrong question.
Benedict Evans
Right. And a lot of the question is, okay, it may not be very good at doing—there's a class of old tasks that generative AI is good at.
Erik Torenberg
There's also a lot more old tasks that generative AI is maybe not very good at. But then there's a whole bunch of other things that you would never have done before that generative AI is really, really good at. And then how do you find those or think of those? And how much of that is the user thinking of it, faced with a general-purpose chatbot? How much of that is the entrepreneur saying, “Hey, I've just realized that there's this thing that I can do that you couldn't do before, and here you are. I've given you a product with a button that will do it for you”?
Benedict Evans
Right.
Erik Torenberg
And it's why there are software companies, right?
Benedict Evans
Right.
Erik Torenberg
And on mobile, some of the new use cases were getting in strangers' cars—we mentioned Lyft and Uber—or dating people you met via an app, or lending your spare bedroom out, et cetera. Those were net-new companies that were built around those behaviors.
And I think for AI, there are still questions of what those net-new behaviors are. We're starting to see some in terms of people engaging and talking with chatbots instead of humans, or in addition to humans. And then there's a question of whether these are done by the model providers that currently exist, or by net-new companies, both in enterprise and consumer.
Benedict Evans
Well, this is always a question: How far up the stack does the new thing go? I was talking about this with another former a16z person, who pointed out that in the mid-'90s, people kind of argued that the operating system does all of it, and the Windows apps are basically just thin Win32 wrappers.
And Office is basically just a thin Win32 wrapper. All the important stuff is being done by the OS, whether it's document management, printing, storage, and display—all stuff that used to be done by apps. On DOS, the apps had to do printing; the apps had to manage the display. We moved to Windows, and 90% of the stuff that the app used to do is now being done by Windows, so Office is just like a thin Win32 wrapper, and all the hard stuff has been done by the OS. And it turns out, well, that was again—frameworks are useful, but that's maybe not a useful way of thinking about what's going on.
And the same thing now: How much does this need a single, dedicated understanding of how that market works, or what that market is, and what you would do with that? I remember when we were at a16z, there was an investment in a company called Everlaw, which is legal discovery in the cloud.
Erik Torenberg
Yeah.
Benedict Evans
And so machine learning happens, and now they can do translation. Are they worried that lawyers are going to say, “Well, we don't need you guys anymore. We're just going to go and get a translation app and a sentiment analysis app from AWS”? No, that's not how law firms work. Law firms want to buy a thing that sells legal discovery software and management. They don't want to go and write their own or do API calls. Very, very big law firms might, but a typical law firm isn't going to do that. People buy solutions; they don't buy technologies.
And the same thing here: How far up the stack do these models go? How much can you turn things into a widget? How much can you turn things into an LLM request? And how much does it turn out that you need that dedicated UI?
The funny thing is you can see this around Google, because Google had this whole idea that everything would just be a Google query and Google would work out what the query was. And guess what? Google Flights is not a Google query.
One of the interesting things about this is thinking about what a GUI is doing. The obvious thing a GUI is doing is that it enables Office to have 500 applications, 500 features, and you can find them all. At least you don’t have to memorize keyboard commands. You can now have effectively infinite features, and you can just keep adding menus and dialog boxes. Eventually you run out of screen space for dialog boxes, but you can have hundreds of features without people needing to memorize keyboard commands.
But the other side is you’re in that dialog box or screen in that workflow in Workday or Salesforce or whatever the enterprise software is, or the airline website or Airbnb. There aren’t 600 buttons on the screen. There are 7 buttons because people at that company have thought: what should the user be asked here? What questions should we give them? What choices should there be at this point in the flow? That reflects institutional knowledge, learning, testing, and careful thought about how this should work.
Then you give somebody a raw prompt and say, “Tell the thing how to do the thing,” and you’ve got to think from first principles: how does all this work? I always used to talk about machine learning as giving you infinite interns. Imagine you’ve got a task and an intern, and the intern doesn’t know what venture capital is. How helpful are they going to be? They don’t know that companies publish quarterly reports, that we’ve got a Bloomberg account that lets us look up multiples, that you should probably use PitchBook for this data rather than Google. This is my point about Deep Research: you should use this source and not that source. Do you want to work that out from scratch, or do you want people who know a lot about this stuff to have spent 5 years working out what the choices should be on the screen for you to click on? It’s the old user-interface saying: the computer should never ask you a question that it should know by itself. You go to a blank raw chatbot screen, and it’s asking you literally everything. It’s not just asking you one question; it’s asking you absolutely everything about what you want and how you’re going to work out how to do it.
Erik Torenberg
And so, you're mentioning ChatGPT, right, about how ChatGPT isn't so much a product as a chatbot disguised as a product. I am curious: When we look back at this platform shift, do you think that there will be another iPhone-esque or Excel-esque product that kind of defines the future—the platform shift—in a way that ChatGPT won't? Or is it that the world has to catch up to how to use ChatGPT, or something like ChatGPT?
Benedict Evans
So both of these can be true, because it took time to realize how you would use Google Maps and what you could do with Google and how you could use Instagram, and all of these products have evolved a huge amount over time. So some of it is that you grow toward realizing what you could do with this. You realize that's just a Google query now. You realize that you could just do it like that, and you realize, “I spent hours doing this, and I just realized, oh, I could actually just make a pivot table.”
The other side of it is that you're still then expecting people to work it out themselves from first principles. And it's kind of useful to have 100, 1,000, 10,000 really clever people sitting and trying to work out what those things are and then showing it to you as a product. I think another side to this is that there were always these precursors. There were lots of other things before Instagram.
Erik Torenberg
Yeah.
Benedict Evans
You know, YouTube didn't start as YouTube. It started as video dating, I think. There were lots of attempts to do online dating that all kind of worked until Tinder kind of pulled the whole thing inside out. And so there were always lots of things—what's the phrase? Local maxima. In fact, this is where we were, particularly with the iPhone, before, because I was working in mobile for the previous decade.
It didn't feel like we were waiting for a thing. It felt like it was kind of working: Every year the networks got faster, the phones got better, and it got a little bit better every year. We had apps, we had app stores, we had 3G, we had cameras, and stuff seemed to be a bit better every year. And then the iPhone arrives, and it just blows the chart: You've got this line doing this, and then there's a line that does that. Although remember, also, the iPhone took like 2 years before it worked, because the price was wrong, the feature set was wrong, and the distribution model didn't quite work.
And so, yeah, you can think everything's going well, and then something comes along and you realize, “No, oh, no, no, no.” That's the same for Google: Search was a thing before Google; it just wasn't very good. There was lots of social stuff before Facebook, and that was the thing that catalyzed it. So I just think, deterministically, this whole thing is so early that it feels like, of course, there are going to be dozens, hundreds of new things. Otherwise, a16z should just kind of shut down and give the money back to the LPs, because the foundation models will just do the whole thing.
And I don't think you're going to do that. At least I hope not.
Erik Torenberg
No, no, no. If we have any regrets from the last few years, it's not going bigger. I think we didn't fully appreciate how much specialization there would be across whether it's voice or image generation, or take any sort of subsector, that there would be net-new companies created that would be better than the model providers, that there would be even multiple model providers in every category.
One thing we've always—in the Web 2 era, we always bet on the category winner, and the category winner would take most of the market. But these markets are so big, and there's so much expertise and specialization, that there can be winners in every category. It's not just that the model providers take everything; even in every category, including the model providers, there can be multiple winners, with increasing specialization, and the markets are just big enough to contain multiple winners.
Benedict Evans
I think that's right. And I think the categories themselves aren't clear, right?
Erik Torenberg
And many things—you think this is a category, and it turns out, no, it was actually that whole other thing. The categories kind of get unbundled and bundled and recombined in different ways. I remember I was a student in 1995, and I think I had 4 or 5 different web browsers and web servers on my PC.
Tim Berners-Lee's original web browser had a web editor in it because he thought this was kind of like a network drive and a sharing system, and he didn't realize it wasn't really a publishing system. So, you would have your web pages on your PC, you'd leave your PC turned on, and that would be how your colleagues would look at your Word documents or your web pages.
So, again, we just don't know how. I keep coming back to this point: I feel like most of the questions we're asking at the moment are probably the wrong questions. Picking up on a strand within what you just said, though, one of the interesting things I'm thinking about a lot is looking at OpenAI.
Because I'm fascinated by disconnections, we've got this interesting disconnect now. If you look at the benchmark scores, you've got these general-purpose benchmarks where the models are basically all the same. If you're spending hours a day in them, then you've got this opinion about, "Oh, I like Claude's tone of voice more than I like ChatGPT, and I like GPT-5.1 more than GPT-4.9, or whatever the hell it's called." If you're using this once a week, you really don't notice this stuff. The benchmark scores are all roughly the same, but the usage isn't.
Basically, Claude has no consumer usage, even though on the benchmark score it's the same. Then it's ChatGPT, and halfway down the chart it's Meta and Google. The funny thing is, you read all the AI newsletters, and Meta's lost, they're out of the game, they're dead. Mark Zuckerberg is spending $1 billion per researcher to get back in the game. But from the consumer side, well, it's distribution.
The interesting thing here is that what I'm kind of circling around is: if the model for a casual consumer user certainly is a commodity, and there are no network effects or winner-takes-all effects yet—those may emerge, but we don't have them yet—and things like memory aren't network effects; they're stickiness, but they can be copied, how is it that you compete?
Do you just compete on being the recognized brand and adding more features and services and capabilities, and people just don't switch away? Which is kind of what happened with Chrome, for example. There's not a network effect for Chrome, but it's not actually much better—maybe it's a bit better than Safari—but you use Chrome because you use Chrome.
Or is it that you get left behind on distribution or network effects that emerge somewhere else, and meanwhile you don't have your own infrastructure? I suppose what I'm getting at is: you've got these 800 or 900 million weekly active users, but that feels very fragile because all you've really got is the power of the default and the brand.
You don't have a network effect. You don't really have feature lock-in. You don't have a broader ecosystem. You also don't have your own infrastructure, so you don't control your cost base. You don't have a cost advantage. You get a bill every month from Satya.
So you've kind of got to scramble as fast as you can in both of those directions: on the one side, build product and build stuff on top of the model, which is our earlier conversation. Is it just the model? Yeah.
Benedict Evans
Now, you've got to build stuff on top of the model in every direction. It's a browser. It's a social video app. It's an app platform. It's this; it's that. It's like the meme of the guy with the map with all the strings on it. It's all of these things: we're going to build all of them yesterday.
And then in parallel, it's infrastructure. We've got to deal with NVIDIA, Broadcom, AMD, Oracle, and petrodollars. Because you're kind of scrambling to get from this amazing technical breakthrough and these 800 or 900 million weekly active users to something that has really sticky, defensible, sustainable business value and product value.
Erik Torenberg
Yeah. And so, as you're evaluating the competitive landscape among the hyperscalers, what are the questions that you're asking that you think are going to be most important in determining who's going to gain durable competitive advantages, or how this competition is going to play out?
Well, this kind of comes back to your point about sustaining advantage. We talked about Google: if we think about the shift to mobile—particularly the shift to mobile for Meta—this turned out to be transformative. It made the products way more useful.
Benedict Evans
Yeah.
Erik Torenberg
For Google, it turned out mobile search is just search.
Benedict Evans
And Maps changed, probably, and YouTube changed a bit, but basically, for Google, search is search. Web search just means more people doing more search, more of the time. Yeah.
Erik Torenberg
And the default view now would seem to be: Gemini is as good as anybody else. Next week, the new model—I haven't looked at the benchmarks for GPT-5.1, which is out today. Is it better than Gemini? Probably. Will it still be better next month? No.
So that's a given: you've got a frontier model. Fine. What does that cost? It costs you—pick a number—$250 billion a year, $100 billion a year. What's this? This is our earlier conversation about capex.
Okay, so Google can pay that because they've got the money. They've got the cash flow from everything else. And so you do that, and your existing products get you to optimize search. You optimize your ad business. You build new experiences. Maybe you invent the new iPhone of AI. Maybe there is no iPhone of AI. Maybe someone else does it, and you do an Android and just copy it.
So, fine, it's the new mobile. We'll just carry on. Search is search. AI is AI. We'll do the new thing. We'll make it a feature. We'll just carry on doing it.
For Meta, it feels like there are bigger questions about what this means for search, or what it means for content and social and experience and recommendation, which makes it all the more imperative that they have their own models, just as it is for Google.
For Amazon, okay, on the one side, it's commodity infrastructure, and we'll sell it as commodity infrastructure. And on the other side—maybe we can step back—if you're not a hyperscaler, if you're a web publisher, a marketer, a brand, an advertiser, or a media company, you could make a list of questions, but you don't even know what the questions are right now.
Benedict Evans
What is this? What happens if I ask a chatbot a thing instead of asking Google? Even if it's Google—from Google's point of view—well, I'll ask Google's chatbot. It's fine. But as a marketer, what does that mean?
What happens if I ask for a recipe and the LLM just gives me the answer? What does that mean if my business is having recipes?
Erik Torenberg
Yeah.
Benedict Evans
Do you have a kind of split? And this is also an Amazon question: how does a purchasing decision happen? How does this decision to buy a thing that I didn't know existed before happen? What happens if I wave my phone at my living room and say, "What should I buy?" Where does that take me, in ways that it wouldn't have taken me in the past?
Erik Torenberg
Yeah.
Benedict Evans
Do LLMs mean that Amazon can finally do really good recommendation, discovery, and suggestion at scale, in ways that it couldn't really do in the past because of this kind of pure commodity retailing model that it has?
Apple's sort of off on one side. Interestingly, they produced this incredibly compelling vision of what Siri should be 2 years ago. It just turned out that they couldn't make it. Interestingly, nobody else could have made it either.
You go back and watch the Siri demo that they gave and you think, okay, so we've got multimodal, instantaneous, on-device, tool-using, agentic, multiplatform e-commerce in real time, with no prompt-injection problems and zero error rates. Well, that sounds good. Has anyone got that working? No.
OpenAI and Google don't have that working. I don't think Google or OpenAI could deliver the Siri demo that Apple gave 2 years ago. They could probably do the demo, but they couldn't consistently and reliably make it work. That demo, that product, isn't in Android today.
Apple, to me, has the most intellectually interesting question. I saw Craig Federighi make this point: “We don't have our own chatbot. Fine. We also don't have YouTube or Uber. Explain why that is different.” That's a harder question to answer than it sounds like.
Of course, the answer is: If this fundamentally changed the nature of computing, then it's a problem. If it's just a service that you use, like Google, then that's not a problem. That's kind of the point about where Siri goes.
The interesting counterexample here would be to think about what happened to Microsoft in the 2000s. The entire development environment gets away from them, and no one builds Windows apps after 2001 or something. But you need to use the internet, and to use the internet, you need a PC. What PC are you going to buy? Apple wasn't really a player at that time and was just getting back into the game. Linux obviously wasn't an option for any normal person, so you bought a Windows PC.
Basically, Microsoft loses the platform war and sells an order of magnitude more Windows PCs as a result of this thing that Microsoft lost. It takes until mobile for them to lose the device as well as the development environment. So here's the question: If all the new stuff is built on AI and I'm accessing an app that I download from the App Store, to what extent is this a problem for Apple? You would need a much more fundamental shift in what was happening for that to be a problem for Apple.
Even if you take not the full rapture arriving and we all go to sleep in pods like the guys in WALL-E—maybe we'll be those people; maybe we'll be like that—in which case, fine. There's a sort of a mid-case, which is that the whole nature of software changes, there are no apps anymore, and you just go and ask the LLM a thing. Fine. What is the device on which you ask the LLM a thing?
It's probably going to have a nice, big color screen and a 1-day battery life. It probably needs a microphone and a good camera. It kind of sounds like an iPhone. Am I going to buy the one that's a tenth of the price and just use the LLM on it? No, because I'll still want the good camera, the good screen, and the good battery life.
There are a bunch of interesting strategic questions when you start poking away. What does this mean for Amazon? Those are completely different questions from what it means for Google, Apple, Facebook, or Salesforce. What does it mean for Uber? Right back to what we were saying at the beginning of this conversation: What does this mean for Uber? Their operations get X% more efficient, and now the fraud detection works. Maybe their autonomous cars—that's a different conversation, but presume no autonomous cars. Otherwise, as Uber, what does this change? Not a huge amount.
Erik Torenberg
You've been doing these presentations for a while now. You bumped them up to 2 times because there's so much changing. One of the things you do in each presentation is ask really great questions and chronicle what the important questions are to be asking.
I'm curious, as you reflect—maybe post-ChatGPT in 2022, or GPT-3, rather—the questions you were asking then and compare them to now: To what extent do we have some direction on some of those questions? To what extent are they the same questions, or new and different questions? If I woke up in a coma after reading your original presentation—let's say the one after the GPT-3 launch came out—and then saw this one now, what were the most surprising things, or the things that we learned that updated those questions?
Benedict Evans
I think we have a lot of new questions this year. You could make a list of what might be half a dozen questions in spring of 2023: open source, China, NVIDIA, does scaling continue, what happens to images, and how long does OpenAI's lead remain?
Those questions didn't really change in 2023 and 2024, and most of those questions are still there. The NVIDIA question hasn't really changed. The answer on China, the answer on how many models there will be—the answer is, okay, anybody who can spend a couple hundred million dollars can have a frontier model. That was pretty obvious in early 2023. It took a while for everyone to understand that.
And big models and small models: Will we have small models running on devices? No, because the capabilities keep moving too fast for small models to shrink down onto the device. Those questions kind of didn't change for 2 or 2½ years.
I think we now have a bunch of more product-strategy questions, as you see real consumer adoption and OpenAI and Google building things in different directions, Amazon going in different directions, and Apple trying—and obviously failing—and then trying again to do things. There's a sense that there is something more going on in the industry than just, “Let's build another model and spend more money.”
Erik Torenberg
Yeah.
Benedict Evans
There are more questions and more decisions. Now there are also more questions outside of tech, certainly on the retail-media side, about how you start thinking about what you would do with this.
The classic framing in my deck is that step 1 is you make it a feature, absorb it, and do the obvious stuff. Step 2 is you do new stuff. Step 3 is maybe someone will come and pull the whole industry inside out and completely redefine the question.
You could imagine step 1 as you're a manager at a Walmart in the Bay Area or D.C., or whatever it is: “Find me that metric.” Step 2: “Build me a dashboard.” Step 3: “It's Black Friday, and I'm managing a Walmart outside D.C. What should I be worried about?”
That might be the wrong example, but step 1 for Amazon is that you bought light bulbs, so here's some packing tape. What Amazon should actually be doing is saying, “This person is moving home. We'll show them a home-insurance ad,” which is something Amazon's correlation systems wouldn't get because they wouldn't have that in their purchasing data.
We're still starting to think about step 1, but what would step 2 and step 3 be? What would new revenue be for this, other than simple, dumb automation? What new things would we build with this? Where might this actually redefine or change what the market looks like? That's obviously a big question for anyone in the content business.
Erik Torenberg
Yeah.
Benedict Evans
What does it mean if I can just go and ask an LLM this question? What kinds of content were predicated on Google routing that question to you? What kind of content isn't really about that question?
Do I want a Bolognese recipe, or do I want to hear Stanley Tucci talking about cooking in Italy? Do I just want the SKU, or do I want to work out which product I should buy? Amazon is great at getting you the SKU, but terrible at telling you what SKU you want.
Do I just want the slide deck, or do I want to spend a week talking to a bunch of partners from Bain about how I could think about doing this? Do I just want money, or do I want to work with a16z's operating groups?
Erik Torenberg
What is it that I'm doing here?
Benedict Evans
I think the LLM thing is starting to crystallize that question in lots of different ways. What am I actually trying to do here? Do I just want a thing that a computer can now answer for me, or do I want something else that isn't? The LLMs can do a bunch of stuff that computers couldn't do before, right?
Erik Torenberg
Right?
Benedict Evans
Is that thing that the computer couldn't do before my business?
Erik Torenberg
Yeah.
Benedict Evans
Erik Torenberg
We're about to figure out, in a much more granular way, what the true job to be done is for many, many of these.
Benedict Evans
Yeah. Going back to the internet, there was the observation about newspapers: Newspapers looked at the internet and talked about expertise, curation, journalism, and everything else. They didn't really say, “Well, we're a light-manufacturing company and a local-distribution and trucking company.”
Erik Torenberg
Yeah.
Benedict Evans
That was the bit that was the problem. Until the internet arrived, that wasn't a conversation you thought about. Then the internet suddenly makes that clear and suddenly creates an unbundling that didn't exist before.
And so there will be those kinds of things where you didn't realize you were that before, until someone comes along with an LLM and says, “I can use this to do this thing that you didn't really realize was the basis of your defensibility or the basis of your profitability.” I mean, it's like the joke about US health insurance: the basis of US health insurance profitability is making it really, really boring, difficult, and time-consuming. That's where the profits come from. Maybe it isn't—I don't know that industry—but, for the sake of argument, say that's your defensibility. Well, an LLM removes boring, time-consuming, mind-numbing tasks.
So, what industries are protected by having that? And they didn't realize that. And, you know, it's like you could have asked these questions about the internet in the mid-'90s or about mobile a decade later. Generally, half of the questions you asked would have been the wrong questions in hindsight. I mean, I remember, as a baby analyst in 2000, everyone kept saying, “What's the killer use case for 3G? What's a good use case for 3G?” And it turned out that having the internet in your pocket everywhere was the use case for 3G.
Erik Torenberg
But that wasn't the question people were asking. And I'm sure that will be the thing now: so much will happen and get built where you go and realize, “Oh, that's how you would do this. You can turn it into that.”
Benedict Evans
Yeah.
Erik Torenberg
No, 100%. My last question to get you out of here: if we're talking 2 or 3 years from now, and you're doing a presentation and you say, “Oh, this is actually bigger than the internet,” or maybe, “This is like computing,” what would need to be true? What would need to happen? What would evolve our thinking?
Benedict Evans
I mean, I kind of come back to my point about the Jews and Christians: the Messiah came, nothing happened. We forget—I mean, there are maybe 2 very brief ways to think about this. One of them is that I think we forget how enormous the iPhone was and how enormous the internet was. You can still find people in tech who claim that smartphones aren't a big deal. And this was the basis of people complaining about me: “This idiot thinks generative AI is as big as those silly phone things. Come on.”
I think another answer would be that I don't want to get into the argument about what the reasoning capability is, the benchmarks, and all that. You see lots of 5-hour-long podcasts of people talking about this stuff, but the stuff we have now is not a replacement for an actual person outside of some very narrow and very tightly constrained guardrails, which is why I agree with Demis's point that it's absurd to say that we have Ph.D.-level capabilities now. What we would have to be seeing is something that would really shift our perception of the capability of this stuff.
Erik Torenberg
Yeah.
Benedict Evans
So that it's actually a person, as opposed to something that can kind of do these people-like things really well sometimes, but not other times. And it's a very tough conceptual kind of thing to think about because I'm deliberate. I'm conscious that I'm not giving you a falsifiable answer. But I'm not sure what a falsifiable answer would be to that. When would you know whether this was AGI?
You know, it's the Larry Tesler line: AI is whatever doesn't work yet. As soon as people say it works, people say, “Well, that's just not AI. That's just software.” It becomes a slightly drunk philosophy grad student kind of conversation as much as it is a technology conversation. Like, have you ever considered, Erik, that—
Erik Torenberg
Maybe we're not either.
Benedict Evans
That's a thought. All I can say to give a tangible answer to this question is that what we have right now isn't that. Will it grow to that? We don't know. You may believe it will. I can't tell you that you're wrong. We'll just have to find out.
Erik Torenberg
I think that's a good place to wrap. The presentation is “AI Eats the World.” We'll link to it. It's fantastic. Benedict, thanks so much for coming on the podcast to discuss it.
Benedict Evans
Sure. Thanks a lot, Erik.