Erik Torenberg
Benedict, welcome back to the a16z podcast.
Benedict Evans
Thank you.
Erik Torenberg
Last time you were here, we were discussing the first iteration of your presentation, “AI eats the world.” You wrote it almost a year and a half ago at this point. You always begin your presentation with your big questions, but before getting into the questions going forward, I want you to reflect on what we’ve learned since you originally made the presentation. What’s played out? Let’s reflect on what’s changed in the last year.
Benedict Evans
I think we have much more of a sense of diverging product strategy. We have much more of a sense of competitive tension that goes beyond just, “Make a bigger model faster with more compute.” We’ve had several iterations of OpenAI’s strategy, in particular, from sort of everything all at once to, “Oops, no, maybe we should double down on coding.”
Clearly, agentic coding started working, and so all the focus in tech has narrowed massively onto that as something that has absolute product-market fit, in the sense that customers are pulling it out of your hands. And, of course, that comes with the supply crunch around capacity and the price imbalance—the imbalance of supply, demand, capacity, capex, and pricing that we see at the moment.
So that’s the big shift: We went from, “This is kind of working and kind of exciting, but we’re not quite sure what we’re going to do with it,” to, “Right, it works for coding. Will it work for anything else?” Almost certainly, but that’s what’s working right now, and so we’ve got this much narrower focus.
Otherwise, the ChatGPT numbers keep coming up, the models keep getting bigger, the capex keeps growing, the usage keeps growing, and people are using this more. But most of the fundamental questions you might have had 2 or 3 years ago still don’t really have answers. We don’t know if there’ll be a winner in the models. We don’t know if they can capture value up the stack. We don’t know how much the models can do. We don’t see a way that consumers will use this daily rather than weekly with the technology we have right now. So all of those questions are still open.
Erik Torenberg
And just on coding, could we have figured—could we have foreseen—that that would have been the use case that really would have taken off? What’s a reflection on that?
Benedict Evans
Well, deterministically, you could have said, “Look, who’s messing about with this stuff? Software developers. What are software developers going to try and make work? Software development.” So, at a simplistic, naive level, yes, the stuff that should work first is software development.
I often compare this moment to the internet in 1997–98, but it’s also like PCs in the early ’80s or the late ’70s. It’s incredibly exciting, but it’s not quite clear what it’s for, and it doesn’t quite work yet. Clearly, the first thing that people did with PCs was make computers. And the first thing that people are doing with LLMs—in a sense, LLMs are computers—is make more compute. So that’s not terribly surprising.
I think the shift at the beginning of this year was that agentic coding went from being kind of useful to really changing everything. I’m not sure you could have—clearly, there were people who were going to say, “Well, this is going to be able to do absolutely anything,” and so they’ll say, “Yes, look, I told you.” But I don’t think anyone could have deterministically predicted exactly when that was going to happen, or that coding would be the first thing where it worked.
Erik Torenberg
And what have we learned about—say more about what this means for engineers, junior engineers, senior engineers, the jobs discussion, how teams are organized, et cetera. What have we learned so far?
Benedict Evans
I don’t think we’ve learned anything. This didn’t work 6 months ago.
Erik Torenberg
Yeah.
Benedict Evans
Everyone is scrambling around trying to work out what it means. You can get very into the noise and the detail: “What did somebody say at a party yesterday? Oh, my God, that’s how it’s all going to work.”
It’s going to take a couple of years for this all to settle down, if nothing else because of the pricing. We’ve got this enormous crunch between the demand and the supply, and hence the pricing. So we don’t know what a team is going to look like.
I think people are asking new questions around the obvious one: Do you hire junior people, and if so, what are they doing? Why were you hiring junior people in the past? Were you actually hiring them to do the thing that they did, or were you hiring them to do something else? If you automate away a class of stuff that used to get done by people, what will happen?
That becomes much more real now in software development because you actually are automating a bunch of stuff that used to be done by people. So those questions are real now rather than theoretical. But I don’t think anybody can possibly say they know what the market structure is going to look like, or what the career of a software engineer is going to be in 3 years’ time. I think you’d be insane to think that you could know that yet.
Erik Torenberg
Yeah. Talk about OpenAI. Talk about what’s most surprised you, or how have you made sense of their strategy development and the questions that they have going forward?
Benedict Evans
Well, it’s always been such a tranquil, drama-free environment. Obviously, they’ve had the issue with Fidji Simo having to take medical leave, which shuffled things up a bit.
Clearly, in the last quarter of last year, their question was, “Right, well, the models are the models, but what else? And how do we get people to do other stuff with this?” “Ask ChatGPT for 15 ideas for what we could do to build value on top of infrastructure, and then we’ll do all of them.” That’s almost literally what it looked like.
And then Anthropic, having raised less capital, said, “No, we’re going to focus on coding.” And they got coding working. Whether that was a deliberate strategy or they stumbled into it is for other people to say, but clearly that worked.
But the question still remains: The stuff that’s working right now is software development and some things in some other fields. Then there are a lot of people who are excited about using this around the edges and using it for some things. There’s clearly a very wide spread between people in the Valley who bought a cluster of Mac Studios and are running OpenClaw all day versus those other 40% of people who say, “Yeah, it’s kind of useful. I used it last week for something.”
I’m like, how do you bridge that? I don’t think that question is solved. Software is a place where people have really, really jumped over that bridge.
There are a lot of other places where people are scratching their heads and using it up to a point. And then there are a lot of places where corporations are using it to automate some specific back-office process. You’re not asking the user to work out what they do with the new tool; instead, you’re saying, “Okay, here’s a problem that we can solve.”
I go and talk to companies outside America and outside tech, and talk to consultants and investors. They’re looking at those one-at-a-time point solutions. I was speaking a couple of days ago to a commodities company, and they want to use LLMs to get better predictions on their cash flow because they deal with all sorts of small producers and don’t necessarily know when their invoices are going to get paid. It’s a very low-margin business, so that’s a big deal, and they want to use LLMs to get better cash-flow forecasting.
That’s a very different thing from going to ChatGPT or Claude and saying, “Hey, give me a summary of my meetings this week.”
Erik Torenberg
Yeah. How did this compare with mobile or other sorts of platforms in terms of early user adoption, in terms of weekly or daily users?
Benedict Evans
I think there are a bunch of different ways to answer this. One of them is that we’re always standing on the shoulders of giants, and growth is always compounding.
So mobile didn't need to wait for the internet or cellular networks, like mobile data. Mobile internet did need to wait for cellular data, but it didn't need to wait for the internet to happen, and the internet didn't need to wait for PCs, and PCs didn't need to wait for consumer electronics and semiconductors, and so on. So you've always got this accelerating adoption. When your boss—my old boss, Marc Andreessen—was working on Netscape, there were double-digit millions of PCs on the entire planet. So, no, you couldn't have 900 million weekly active users because there weren't 900 million PCs.
So there's always that acceleration. That's one point. I think the second point is that, at the early stage of any of these shifts, it's not really clear how it's going to work, and nothing works. I'm just about old enough to remember this. I'm not sure how old you are, but anyone in their 30s doesn't really remember a time when it was completely normal that you'd be working, everything on the screen would just freeze, and you'd have to crawl under your desk and unplug the computer, and then pray that some of what you'd done in the last hour might still be there.
That just doesn't happen anymore. Go back to the '80s: you bought a sound card. Well, that's $300. You want to have sound on your computer? Okay, that's $300, and it's like a weekend to make that work. I remember trying to get this stuff to work.
It's the same thing with the internet. You've got to get a floppy disk with TCP/IP on it, and it's slow, and none of the stuff that you need to do existed. It's the same with mobile. We're kind of at that stage, and of course it's not clear which of these things are going to work. Is a browser going to work? Is it going to be this? Is it going to be that? How's this all going to fit together?
There's a gap between what's incredibly exciting and the small number of people who are willing to put the work in to get something to work, and just turning that into a thing where you can press a button in real hands.
I think the third point here is a much more tangible observation: the pricing crunch that we've already mentioned looks to me a lot like what happened with mobile data in 2009–10, where suddenly people got bills for, like, $5,000 or $10,000 of data. On the one side, and on the other hand, if you had flat-rate data—which is kind of what happened in the U.S. with the iPhone—AT&T, or Cingular, launched the iPhone with flat-rate data, and then everybody bought iPhones and started using 3G. People started watching YouTube, and the whole network went down because they just didn't have the capacity to do that.
It's funny: there are still people in tech who don't understand that cellular networks have marginal costs. They have to add more capacity, and that costs more money. The networks had to scramble to get the cost curve aligned with the pricing system, aligned with the underlying cost and aligned with perceived value, which they kind of did with data caps, bundles, fair use, throttling, and so on.
The other side of that comparison—and this is exactly where you see it now—is that, on the one hand, you're paying $20 a month and getting $10,000 worth of tokens, and on the other hand, you mess about for a couple of days and get a $10,000 bill, and you're like, “What the hell is this?” You literally see these stories now, which is exactly what happened in 2009, 2010. It's also what happened in 2001, 2002, and 2003 with GPRS.
But I think the other interesting part of that analogy, or that comparison, is that since then, mobile data traffic has risen by something like 1,500 to 2,000 times. The mobile networks collectively have revenue of about $1 trillion, and they spend about $200 billion a year on capex. Their stocks have been flat for 20 years, and all the cool stuff got built by somebody else.
They kind of all thought that they were going to build all the cool stuff. I worked for a phone company that had a banking license because they thought they would do mobile banking, which now seems absolutely insane. But that's kind of the point: they built this amazing piece of global, incredibly sophisticated, very expensive infrastructure, with enormous growth in use all the time, and it changed all of our lives. We all pay for it, and they didn't make any money from it because all the value moved up the stack.
This is, of course, the absolutely central question for LLMs: can the model do the whole thing, or do you have to have 300 apps built on top of it? Can you just go to the model and say, “Do my taxes for me,” or do you need to have a tax thing that might use some AI in 10 different ways inside it? If not, then what does it mean to be a foundation-model provider?
Is this just commodity infrastructure that gets sold at marginal cost? That somehow seems to be a very difficult concept for people to grasp right now, because you can sell all the tokens you can make, so you can price it at ROI. But over the next couple of years, we've got $1 trillion to $2 trillion of capex coming down the pipe, and the models get 100x or 200x more efficient every year. Then there are new models, and will the models use more tokens or fewer tokens? But we'll get to a different equilibrium. Why would that equilibrium be one where the model companies have pricing power when the models are all kind of the same, doing kind of the same thing with the same chips? Why would they have pricing power?
I think that's the long answer to your question. You go back and look over time: chip companies didn't capture the value, ISPs didn't capture the value, and mobile network operators didn't capture the value. Windows and iOS did, but they were doing something else. They had all these levers to go up the stack, and of course they had network effects, which models don't have.
So that's sort of the question: do they end up like the infrastructure layers, or do they end up like the operating-system layers and capture value, and actually get to decide what gets built? Or do they end up—I mean, the irony of this is Netscape, where Marc Andreessen famously said that he was going to turn Windows into a set of badly debugged device drivers, and Microsoft kind of crowbarred its way into the market. But it turned out that web browsers weren't the point, because all the value was somewhere else.
I think that's the more—or the kind of—a swirling mass of questions about how this settles out, which comes back right through to all my answers to your question. Some of this stuff you know, but you don't know how it's going to work.
Erik Torenberg
Yeah. It's unclear whether it looks more like the internet or software, where a lot of the value—or just better margins—happen at the application layer, or more like the cloud, where the value sort of existed at the hardware layer. Right now, so far, it seems like NVIDIA—and going up, it seems like they have better margins and are capturing a lot of the value. But it's unclear if that will remain the same, or if there will be applications, if we'll look more like the internet. How would you even begin to predict the answer to this?
Benedict Evans
Well, there are 2 answers to this. There are all these sorts of quotes about how history works, and my favorite one is, “History teaches us nothing except that something will happen.” You can always explain afterward why it worked out like that, but it generally wasn't obvious at the time.
In particular, I remember that about 15 years ago, a lot of really, really clever people in tech looked at the iPhone and Android and said, “This is open versus closed again, and Android is going to crush the iPhone,” which of course isn't what happened. I can go and explain why, but all of these comparisons are useful, and none of them are predictive. It's always obvious in hindsight.
I've done a couple of podcasts recently, and I've published this presentation. There's a class of comment on this stuff that says, “Benedict, you're not doing your job. You're supposed to tell us what's going to happen. You're supposed to make predictions, and all you seem to do is say, ‘Well, we don't know.’” There are 2 problems with that. One is that there are a bunch of places where I actually do say, “I don't think this is going to work. I think it's going to work like that.” I don't think foundation models are a product. I don't think a chatbot is a product. I think the value will be further up.
But the other side of this is that, when you're at this stage in the cycle, there are many paths, and you don't know which of those paths it's going to take. To try to say, “Well, I think it's going to be that one,” you might be right, but you do have to be conscious of how uncertain this is and how many different paths it could take. That's the nature of this part of the cycle: all bets are open.
We get to the point where the S-curve kind of curves up and it narrows in. There was a moment when Windows Phone might have worked. In hindsight, no, it probably wasn't going to work.
But there was a moment when it wasn't clear how mobile was going to work. And there's a moment when it was clear: right, this is what's happening. Now we move on to the next question.
One of the characteristics of tech is that the moment you understand something—how it works and what's going to happen—is the moment you should move on to something else. You should always be looking for the places where we don't know what the answers are, because I haven't updated my Apple spreadsheet in 5 years because we know what happened. I don't care what next year's iPhone looks like. I don't pay attention to Apple's market share in China; it happened. Next question.
Erik Torenberg
You mentioned the prediction that you don't think foundation models are the product; you think it'll move up. Explain the reasoning there a bit and what that could look like.
Benedict Evans
I think there are 3 or 4 building blocks you can put on the table. One of them is that it's not clear how you could build a model that was fundamentally better than everybody else's model in some sort of sustainable, differentiated way. There doesn't seem to be a network effect. There don't seem to be levers you can pull, or a strategy like the one Instagram has, or YouTube, or Google Search. We don't see an equivalent of that for LLMs.
Now, you have different emphases. Maybe this one's better than that one; maybe you like this one more than that one. But there doesn't seem to be a fundamental differentiation, a fundamental competitive difference between the models, except your willingness to spend money.
The second problem is that the chatbot itself is a kind of weird, limited V1 UI. There are some things, some people, and some kinds of tasks where it works really well, but for most of the others, you need a bunch of other stuff. You need tooling, it needs to be set up right, it needs to have the right data, it needs to be configured and controlled, and it needs to have the right user interface.
People need to have sat down and thought about how this should work, because generally, people who are good at using the tool and doing the job that needs the tool are not the same people who are good at deciding what the tool should be. People who are really, really good at designing print publications are not the people who should create and design the software. That's a different set of skills. People who are really, really good at giving financial advice are not the right people to design TurboTax. Those are different people with different skills.
You're kind of groping around in the middle of this. You now have Claude for this, Claude for that, and skills and so on. To me, one question is, who builds the skill? Another question is, that seems to be a bit like what you get if you do File > New in Excel. These are templates, and they'll take you so far, but at a certain point, people outgrow the templates.
There's a slide in my presentation that is a quote from somebody who said to me on Twitter years ago. They said they were a consultant, and half of the jobs were telling people who used Excel to use a database, and the other half were telling people who used a database to use Excel. So there's this kind of fuzzy, swirly place of: do you need dedicated software? Do you need horizontal software? Do you need vertical software?
You can't just do everything in Excel. We've all seen the department that runs along on a 10-megabyte file. I run my business in Numbers, on a spreadsheet, but there's a certain point where you outgrow that.
Following that on, can the model labs build all of that? Of course not, no more than Microsoft or Apple could build every Windows app or every iPhone app. So then, do the model labs have leverage? Are they Windows? Are they iOS?
Again, is there a network effect? If you're a law firm right now and you buy a piece of software—you see all the pieces of enterprise software that a16z is invested in—how often does the law firm, or the manufacturing company, or the bank say, “Oh, does this use Claude or does it use OpenAI? We standardize on Claude.”
Well, no, that's not how it works, any more than it worked like that for cloud. You didn't say, “Our company standardized on AWS.” You don't even know what company, what cloud, that SaaS product runs on. That's the whole point: it's abstracted away. It's not your problem.
The foundation models seem to look more like that. They look more like the hyperscalers in that sense. They might have competitive advantages, but further up the stack, you don't have leverage, you don't have a network effect, and you don't have control.
That prompts me, incidentally, to say that maybe the right comparison here is with semiconductors, where with each generation it just gets more expensive, and so you have fewer players. All of that taken together—the models are kind of commodities, the chatbot isn't the right UI or the right product, and the companies aren't going to be able to build all of that stuff themselves—means that they're low-level infrastructure.
So then, do they have pricing power? You're going to have, pick a number, 3 to 6 companies making a frontier model, spending—no one knows, no one honest knows—something between $200 billion and $2 trillion a year on building these models. Plus, there'll be a bunch of edge models and a bunch of open source.
Where is this going to settle down? It might be half a dozen companies that are all competing to sell this stuff. Where is the price discipline going to come from, particularly when some of them have whole other business models as well? Google sells ads, so it has a different attitude to pricing than OpenAI.
I think the challenge here is that there's a difference between where we are right now and where this should end up, which is kind of a first-year economics student conversation. Right now, we're in this period of extreme disequilibrium of supply and demand, price, capex, and capacity.
Just because demand for tokens is infinite, that doesn't mean you can't get to a different price equilibrium, because of course that's what happened with mobile data. Demand for bits is infinite. It's grown 1,500x in the last 15 years, but you still get your supply-and-demand price equilibrium, and you still get a murderous price war between telcos in most parts of the world.
Fundamentally, you're selling a commodity to people who will swap back and forth, and developers will also swap back and forth. I'm happy to say that this might be completely wrong. It may be that we get to a world in which there are only 2 companies that can make an LLM and they have pricing power, or we get to a world in which most of what we do gets subsumed into the model, or they have leverage further up the stack.
It's my point about iOS versus Android. Just because you can say, “It worked like that the last 3 times,” that doesn't prove what's going to happen this time. But it does mean that you should ask the questions, and you should certainly pay attention to them.
I'll just say, as a primary observation, that this situation right now is transitory. We're in this extreme scarcity, and then we have a pricing system, we have a free market, and we have a surge of capex—like $1 trillion of capex. Those multiples are going to move around. And then what?
Erik Torenberg
Going back to your point earlier—it's a good segue to your point that, hey, we know Apple's Apple—what are some of the next questions that you're most focused on, or that we should be paying most attention to?
Benedict Evans
One way to answer that is that some of the questions we've already talked about are: how far do the models go? Can the models differentiate? I think another question is, at what point do we see more and more classes of use cases where the models are good enough and we don't need the most expensive, fastest, biggest, heaviest model in the cloud?
You can use an older model, you can use an open-source model, or you can have a model running on-device. Obviously, this is what Apple is going to be talking about in a couple of weeks. How much can you push onto the device, where the compute is free—or free to you, anyway—and doesn't have marginal cost for the developer?
Another classic question is that it's almost as if the question has moved out of technology. If you're looking at a law firm, a consultancy, or basically anyone in professional services, where you traditionally have this pyramid structure and you can automate a great chunk of what the people at the bottom of the pyramid were doing, what happens?
The only thing you can say there is that if you've never worked at a law firm, or never worked at Bain, BCG, or McKinsey, you're probably not going to have a good idea of how this works. You probably don't really know what it is that all those associates are doing, and you also don't really know what it is that the client is paying for, or how those things get reconfigured.
What does AI mean for finance, both for that internal hiring structure and the kind of products you can create and the margin structure? What does it mean for consultants? What does it mean for the Big 4, for the Big 3, for Accenture, for big law firms, and for advertising? You can probably answer some of those questions.
But if you’re not in that industry, you don’t really know the answers. This reminds me a lot of something I wrote when I was at a16z, which I called “Content Isn’t King.” I also wrote something that said “Netflix Isn’t a Tech Company.”
The point I was getting at is that if you looked at Netflix, this whole thing is enabled by stuff the tech industry built. But all the questions for Netflix are LA questions: What shows? How many shows? What kind of shows? What should you pay the talent?
Should you aim for awards? Should you do movies? Should you buy sports? What kind of sports? These are all Los Angeles questions. These are not San Francisco questions. No one in San Francisco even knows what the right questions are.
They’re media industry questions. This was kind of my point: all the questions that matter to Netflix have become media industry questions. This is obviously the great tension point about Tesla. Is it a car company? Is it a technology company?
What I’m getting at is that “What does this stuff mean for law?” is a question for lawyers as much as it is for people who understand a lot about law firms—how they actually work, what they’re actually doing, and what the clients are actually buying from them. The same thing goes for, “What does generative video mean for Hollywood?”
Ben Affleck probably knows a lot more about this than I do. He built a company and sold it for around $100 million, so obviously he does. That’s kind of a second question: the questions move outside of AI, and they become half-AI questions and half-something-else questions.
The third level—and I probably should have said this earlier—is that the way all of this is fundamentally different from previous platform shifts is that with 3G, or the iPhone, or the web, you didn’t know what was going to happen next, but you knew the physical limits.
In 1995, you knew that telcos weren’t going to give everybody in the world broadband the next week, and you knew that everyone in the world wasn’t going to go out and buy a PC because a PC cost around $3,000. So you knew the basic physical limits of what could and couldn’t happen.
With generative AI, we don’t know those things. We might look at our phones when we get off this recording and see a push notification saying that OpenAI’s new model is out and it’s 2% of the price because they worked something out. I don’t think that’s very likely at this point, but we don’t know those kinds of things.
How much bigger will the models get? How much better, faster, or cheaper? How much pickup will there be? In what ways will the characteristics of the models change? We don’t know.
That is different from previous platform shifts, where you did know the fundamental constraints. That will spin off more questions. In a sense, this is something I pointed to earlier: the place that has product-market fit right now is coding. Nothing else has equivalent product-market fit right now.
I think I’m pretty safe in saying that when swap has gone from whatever it was—$9 billion in run rate at the end of last year—to a $47 billion run rate, that’s all software, isn’t it? So what happens when someone else in some other field gets something working?
Erik Torenberg
Yeah.
Benedict Evans
Which field? Law, banking? I don’t know where.
Erik Torenberg
If you had to guess, what are the use cases outside of coding that could potentially yield daily activity?
Benedict Evans
So, there’s a presentation that I published a couple of weeks ago. There are 3 sections. One of them is about capital, CapEx, infrastructure, foundation models, and differentiation, which is the stuff we talked about.
The second is: How would you build software with this? What does this do for the software industry? What would software look like, and what happens to the margins and the companies and everything else?
The third section I called “Change,” which gets to this point. I opened it with what appears to upset a certain category of person, where I used the Yogi Berra quote that “predictions are hard, especially about the future.”
Benedict Evans
I think there’s a sort of backtest point here: Imagine asking these kinds of questions about the internet in 1997. What would you have gotten? What would you not have gotten?
One way you can look at this is to say, “This is automation.” It makes a class of things that people used to do, which couldn’t be automated, something you can now automate. Then you can ask, “What does that mean?”
I proposed 3 or 4 sorts of buttons to press. The first one is: Is this just price elasticity, which is really what Jevons’s paradox is? If you make it cheaper to do stuff, do you do the same amount of stuff for less money? Do you do more for the same money? Or do you do more for more money?
Because it becomes so much cheaper, was there something that you couldn’t do before that now becomes cheap? Was there something that was expensive and served as a barrier to entry, like owning a printing press as a newspaper?
Is there something that was a cost-based barrier to entry that now goes away? Is there something that gets unlocked in your business model or in your competitive space because this thing became cheap?
The final question would be: What stuff was just completely impossible—totally cost-prohibitive—so that nobody even thought about it, and now that’s within reach?
The example I used to give here was that steam engines make trains possible. It wouldn’t matter how many horses you bought; you couldn’t have a train or an express train.
A more contemporary example would be to point to something like YouTube, or indeed to Spotify. Spotify says that when you look at the last 25 years of the music business, the first half is what happens if you don’t have to buy a $15 CD to get that track.
But the second half is: What if $15 a month gets you all the music there is? That was something that was just completely impossible.
The problem with making predictions like this is that, on the one hand, you’re going to say stuff that’s clever and obvious, but you don’t actually know what it’s going to mean industry by industry.
If we’d been back in the late 1990s and said, “We know the internet will destroy the value of physical distribution,” it turned out that meant completely different things for newspapers and movie studios. Newspapers got completely screwed by this, and movie studios kind of haven’t really changed very much.
So again, it depends. The other part of this is that there are some places where you can ask more useful questions. The one that intrigues me is: How does this change advertising, e-commerce, brands, marketing, and everything that we buy?
Advertising is $1 trillion, and retail is $25 trillion, so it’s a reasonable-sized TAM. The thing that I always used to think about was that Google, Meta, and Amazon don’t really know what the product is.
They know it’s a SKU. They know what the publisher typed into the metadata field, and they know that people who bought this also bought that. But they don’t know why, and they don’t really know what those things are.
That’s why you get these jokes about, “Hey, Amazon, I bought a toilet seat cover. I’m not collecting toilet seats.” It doesn’t really know what a toilet seat is, and it doesn’t know that people don’t buy 2. Actually, it should know that. That should be frequency analysis, but it doesn’t.
With an LLM, in principle, you would know what those things are, why people buy them, and what other things people buy. Obviously, “know” is a difficult, tricky term to use. What do you mean when you say “know”?
But at a minimum, it’s a very different level of statistical correlation from what an AI system would be able to do. That’s why you see the ad numbers and the conversion rates shooting up every quarter at Google and Facebook, because they’re rolling all of this into their ad systems, recommendation engines, and prediction algorithms.
You get shown more stuff that you would like, and the ads that you’re seeing are more likely to be things you’d like to buy. So they have this enormous, sudden acceleration in their ad revenue.
All of which is to say, you look at how these systems work, and right now they say, “People who bought that could buy this.” You should now be able to say, “Here’s a picture of a coat. What is it? Where can I buy that?”
Ten years ago, that certainly wouldn’t have worked. Five years ago, it probably wouldn’t have worked. Now, that should work.
Then you can say, “Suggest 10 other coats like that at different prices, tell me where I can buy them, and suggest the pros and cons of each one.” You’ll kind of get that, too.
Then you can push one step further and say, “Look at my Instagram and suggest a winter coat I should buy that will change my look, but not too much.”
Again, 3 years ago, that would have been total science fiction. Now you think, “Yeah, you could probably build something like that.”
That would kind of work. And those kinds of shifts in what the computer knows, what it can automate, and what suggestions it can make—going back right to the beginning—whenever you get a new technology, you start by doing the old thing but more: more spreadsheets, more PowerPoints, more email, better email. But the important stuff is not doing the old thing but more. It's doing something new that you couldn't have done with the old thing. It's a pretty banal observation, but we kind of lose sight of it.
And so what are the new things that you can only do with this, as opposed to automating the old stuff? I mean, the enterprise version of this would be: you've got all Zoom calls with clients recorded, all the flows of emails in and out of Salesforce, and you can see all of the telemetry, metrics, and analytics of how people use our product. So how should we change our prices to improve our churn? Again, that's something that an LLM might be able to do, which is very different from saying, “Do sentiment analysis on calls into the call center and tell me which customers are angry.” You get multiple shifts in the layer of abstraction around what analysis you can do. And, of course, that then creates new companies and destroys old companies, and creates new businesses and everything else.
But again, we're in 1997 and I'm trying to predict Uber and Airbnb. If I could actually do that, there's a general point here: if we could actually predict what was going to happen, we'd live in a parallel universe. VCs would have—it wouldn't be a 1-in-10 hit rate; it would be a 10-out-of-10 hit rate.
Erik Torenberg
It seems like one of the questions we're now asking is: what was unreasonably expensive to do before that is now possible? Maybe, I don't know, something crazy like rebuilding YouTube from scratch or rewriting Linux from scratch?
Benedict Evans
Yeah, it's funny. The other paired fallacy, of course, is that the new thing comes along and says, “Well, we're going to build the old thing with the new thing.” Of course we're going to build Office with open source; we're going to rebuild it on the web. And, you know, it turns out—guess what? Look at Google Docs. It's got like 20% of the market because that's actually not the point. What's interesting is to do something else, to do something new. It's to shift that level of abstraction and to kind of spot problems that have never existed.
I mean, the experience you get sitting in pitches all day at a venture firm is that there's some stuff where you think, “Well, that sounds kind of useful,” and stuff where you think, “I'm not quite sure why that would work.” But there are some things that kind of fill a hole in the universe, and as soon as somebody explains it to you, you think, “Wow, why did nobody do that before? Why did no one see that that thing existed?” And that's part of the fun of looking at startups. People will suddenly work out that that problem existed, and no one—including the people who have that problem—realized that problem existed. And then they'll go out and make a thing to solve it.
This is also, incidentally, going back to an earlier point: this is why I don't see the—I think that this is the problem with the idea that the model will do the whole thing. If you go back and think about all the pitches you've seen since you joined a16z, how many of them were things where people in the industry knew that was a problem? Quite often, the answer is actually no. No one in the industry thought that was a problem. It actually took like 2 years to explain to them and persuade them that that problem existed at all, and that this new thing would fix that for them. And that's kind of the problem with the idea that a middle manager in finance is going to use this tool to solve this big global industry problem. No, because no one knew that industry problem was there, let alone could work out the right way to build a tool to solve it.
Erik Torenberg
Does this imply a less consolidated SaaS environment than before AI? Maybe less bundling or single behemoths like the Microsoft enterprise? Gosh, way to bring me back down to earth. Is the SaaS industry going to be less consolidated, Benedict? That's all great, but tell us about the stocks. What are the kind of building blocks that we can put down here?
Benedict Evans
So, obviously, it's going to be way cheaper and quicker to build software. Obviously, there's going to be a bunch of stuff you could do with software that you just couldn't do before at all. And so there will be more competition there, and of course this comes with a new margin structure. But, as per our conversation earlier, we don't really know what that margin structure is going to look like.
Are you going to go to outcome-based pricing? It's really hard to tie each button press in a piece of enterprise software to P&L. Sometimes you can in Salesforce or something, but there's an awful lot of software where it would be really hard to say, “Well, the work I did today did this to P&L; therefore, this is what we should pay for it. This is what we should pay for that piece of software.” I don't think that makes sense in the long run. Anyway, what does the pricing structure look like over time versus now? There will be more competition. It will be easier and quicker to build stuff.
The way that I thought about it, I suppose, is that there are maybe 2 framings to think about this that are useful. One of them is to say that if you think about the enterprise software fleet today, you've got 3 buckets. You've got your big-iron horizontal systems—SAP and Workday, your CRM and your human capital management software, your payroll management software, and so on. Then you've got vertical software. A typical big U.S. company has like 300 to 400 SaaS apps, and then another 1,000 apps that they bought or built themselves internally, running on-prem.
And in the middle, you've got this fuzzy, improvised space of Excel and email and the shared file system and so on. Stuff kind of moves back and forth between those. In principle, every SaaS app is doing something that you could have done in SAP or Excel. You could have managed your graduate recruiting in Workday, but at a certain point—like in our conversation the other day—if you're PwC and you hire however many thousand graduates every year to train to be accountants, you probably have a piece of dedicated software that you built for yourself, or maybe you hired Accenture to build. You probably hate it, but anyway, you've got this piece of dedicated hiring software, or you bought something.
If you're a company that hires 5 graduates a year, you're doing that in email and a shared Google Sheet, because why would you buy software for that? Then there's a space in the middle. Do you do it in Workday? Do you do it in Excel? Do you do it in a dedicated app? And now you add ChatGPT to that. Do you do that in an LLM? Is there an LLM tool that means you can do that in Salesforce where you couldn't do it before, or you can do it in your vertical app where you couldn't do it before? Do you use the LLM to build yourself a tool for that, just as you might have a company department that runs on a 10-meg Excel spreadsheet that someone built 15 years ago? No one knows it works. No one knows how it works, but they're still using it. So it arrives within this broad, fragmented, complicated landscape, and it's another set of options for how you would do that task.
So this is one framing to think about it. I think the other framing to think about this is: does the LLM go at the top of the stack or the bottom of the stack? On the one hand, the bottom of the stack is a feature inside Salesforce. You're in Salesforce: look at the history with this customer, look at the context of every other sales call we've done, look at our business objectives, and suggest an email—or suggest what I should do here and what I should say on the call to the customer. So it's a feature; it's a button that's controlled and has tooling and guardrails and everything else that are driven by that particular use case.
The other way to look at it is the example I gave earlier: go look at Salesforce and Workday and all of our email and Google Analytics, and synthesize something that you couldn't have done before. The tension in both cases is: where do you put the probabilistic software that can make mistakes, and where do you put the deterministic system software that can't answer these kinds of questions? Where do you put the database, and where do you put the LLM? Which is at the top and which is at the bottom? The answer is probably both, depending on what you're doing and where it goes.
All of which is a long way of saying: what does this do to software? The answer is more software—way more software. I mean, all software companies exist to solve problems created by other software companies. That was the joke in security: all security software exists to solve problems created by other security software. Clearly, that's what we went through with SaaS. SaaS gave us an order of magnitude—2 orders of magnitude—more software. We should probably expect that with this.
What that gets to, with the SaaS apocalypse, is that all the investors are looking at all these companies and saying, “We don’t really know which of these companies are going to get screwed by all of this.” Some of them must be. Obviously, some percentage of all the SaaS companies that are out there are going to get wiped out by this, but you don’t know which ones, so you probably shouldn’t derate the whole thing by 50%.
But clearly you’re going to go, “I’m not sure I’m going to be long software at the moment until I have some idea of what the hell’s going on.”
Erik Torenberg
You said in your talk with Ben Thompson that software is someone sitting down and designing a workflow and saying, “This is the right way of doing this from now on.” But you also said that a process grows out of the way a business runs. Does that just take time, or do you think we need more experimentation and iteration from these vertical AI startups to get this into the right shape of software for the future?
Benedict Evans
In a sense, maybe an interesting turn on this is that this is both what strategy consultants and software companies do. They look at what’s going on inside a company and say, “This is a crap way of doing it. This would be a better way of doing it. It would achieve your objectives better.” A software company encodes that in software, and a strategy consultancy encodes that in workflows, charts, processes, training, and objectives. It might tell them to buy some software to do that thing, or now, increasingly, maybe build them that software as well.
Another thing to talk about here is how much of what’s done inside an organization is implicit and not documented, not in the training data, and not something that anybody in that company could actually sit down and draw you a flowchart of and explain to you. That’s a big chunk of the value of Bain, BCG, and McKinsey: They have a license to come into a company and talk to everyone, including the people you’re not allowed to talk to because they’re in a different organization and might get fired. They can work out how this actually works, as opposed to how it’s supposed to work, and why people aren’t doing the strategy. Because, actually, guess what? Their bonus targets depend on them not doing the strategy.
They can work all of that out and be a team ready to come in from the outside and give you the answer. Then you can blame them or have that kind of pre-baked solution. These are problems in organizational management and how people function, and in how people can explain what they do. They’re very hard to write down and very hard to bake into a Claude skill and say, “There you are. Make a PowerPoint.”
There’s a broader “How does this always work?” challenge here: How do you get people to use these technologies? How do people adopt new tools? How do you work out how to help people adopt new tools and work out what new things you would do with them? That’s also what happened with cloud, the web, mobile, the internet, PCs, spreadsheets, and so on.
Erik Torenberg
To that end, do you think there’s some kind of coevolution between AI-native software and new types of interfaces? For example, new customer-service AI platforms that might not have had as much human-facing UI, or systems-of-record software being built without a front end at all because its primary user will be AI agents querying it. I think these are interesting ideas. They’re things I struggle to have a strong opinion on because they’re not deep into the weeds of how enterprise infrastructure gets bought.
Benedict Evans
I wonder how new some of these questions are. I remember Chris Dixon saying 10 or 15 years ago that APIs are the new BD, and software companies wouldn’t need to come—software companies could just open up their APIs. Well, what’s old is new. You don’t need an API anymore. You just have an MCP server, and people will plug into that. The agent will just plug into it.
I don’t know. I think the challenge with a lot of this stuff is that all the decisions are really exception handling. The question is always: What can you not automate? What requires someone to make a decision, exercise some judgment, and have an opinion about it because maybe that hasn’t been written down, or that didn’t happen before, or it doesn’t look quite the way it happened before?
There are various ways of thinking about separating out what gets automated and what doesn’t. The way I used in the deck was to talk about what’s a task versus what’s a job. The tasks used to accomplish the job might change without the job itself changing very much, or without the thing that the job is selling to the client changing very much. If you think about what accountants did 50 years ago and what accountants do today, they spend almost none of their time doing the same things. To the client, though, it’s kind of the same thing. It just gets done in a completely different way, with a whole bunch of different tasks.
One of the more profound, or perhaps abstract, ways to think about this is: Where is it that you want the average? Where is it that what you want is the way everybody would do this? That’s the way everyone would do it. That’s what anyone would say. That’s what anyone would make. That’s what any associate would make. That’s what anybody would give me. That’s the answer anyone would give.
Versus where is that not what you want? Where is it that you want the answer to a new question, or a different answer, or a different idea? LLMs are going to be very good at anything where you can describe how people do it, and where what you want is the way anybody would do that. They’re not going to be as good at things where you can’t really explain why you did it like that, and where you’re doing it differently from the way people would normally do it.
Erik Torenberg
Various people, including the CEO of Google, have said that the risk of underinvesting is riskier than overinvesting. Is there any level of capex where that stops being true, and are we getting there now?
Benedict Evans
There’s a financial-gravity problem in that Microsoft, Meta, and Google are all in line to spend over 50% of revenue on capex this year. We think of telecoms as being capital-intensive. Telecoms spend 15% to 20% of revenue on capex.
The guidance from the big 4 companies this year is $700 billion. Telecom is $300 billion, mobile is $200 billion, and total telecom is $300 billion. Oil and gas, depending on which definition you use and which parts of it you’re counting, is anything from $700 billion to $1 trillion. I think, from memory, it depends exactly on who you ask.
$700 billion a year is not an impossibly large amount of money in terms of what big global infrastructure costs. It’s just a lot of money. Clearly, those companies could not spend $1.5 trillion next year. If they did, they’d have to borrow it, and they certainly couldn’t sustain that level of spending for any length of time. There’s a certain point at which that growth has to slow down because there isn’t any more money.
You can talk about ROI and your ability to produce returns from that investment. Clearly, the capital markets are willing to fund that up to a point. But pick a number at random: We can’t spend $10 trillion a year on AI infrastructure because there isn’t $10 trillion a year there to spend on it. There are finite, almost laws-of-physics caps on the amount of money that’s available. I’d hesitate to say something more tangible than that at the moment.
I almost go back to what I said at the beginning: We’ve got a bunch of multiples. There’s far more demand than supply. On the other hand, efficiency is increasing massively. We don’t know what the next model will be. We don’t know where edge or open source come in yet, or when they come in yet. Meanwhile, you’re always chasing the next model.
This is the line that runs across all of it: The model is only relevant for 3 to 6 months, 6 to 9 months, or whatever you want to say. The model costs how many billions of dollars? How much infrastructure do you need to do that? I don’t think the math has really shaken out yet.
Obviously, there are a bunch of very clever semiconductor analysts who spend lots of time trying to put numbers on this. It’s kind of like trying to put numbers on internet bandwidth in the late 1990s. You know what the rows in the spreadsheet are, but you don’t really know what the values are. All you can really say is, “Well, look, it can’t be huge.” There are clearly physical limits on this.
Another way to answer the question is that if you’re Google, Meta, or Microsoft—and, to some extent, Amazon, and to some extent, Apple—this is an existential problem. You have a FOMO problem. On the one hand, your returns on the investment at the moment are hugely positive. On the other, you can’t let other people get away with this without you participating, because then your company is gone.
You don’t want to end up like Microsoft in the 2000s, IBM in the 1990s, or indeed Meta in the 2010s, where they were continually getting shafted by Apple. If this is the future of compute, then you need to be participating in it.
Obviously, at the same time, the CFO is sitting there saying, “Well, yeah, that’s great, but how much participation are we talking about here?” It’s clear that, at a certain point, that curve is going to have to taper off, because there’s nowhere else it can go.
Erik Torenberg
Do you think there’s going to be a reckoning around token maxing? Is it possible that companies have been overshooting AI usage and, when they do proper ROI studies, they’ll pull back?
Benedict Evans
Well, obviously, you’ve had people using the most expensive model to dick around on the internet, which is kind of what happened with mobile in 2010. You got a $10,000 bill and would have said, “Wait, wait, I thought this was a flat-rate bundle. What happened?” So you’ve obviously got a bunch of silly and meaningful stories.
I think what’s slightly more interesting as a question is that, clearly, there’s going to be a point at which—as I’ve said several times—we’re at a moment of massive disequilibrium. The pricing has got to get back into alignment with the cost, and the usage has got to get into alignment with the pricing and the ROI.
The challenge is that it’s a bit tricky at this early stage. It’s quite hard to know what the ROI is. It’s rather like giving everybody the internet in the late ’90s and saying, “Okay, go off and be more productive.” If you ask CFOs where they’ve seen the benefits, most of the benefits so far have been stuff that’s pretty hard to measure.
There’s a survey from Deloitte, and there’s also a survey from the Fed that’s in my presentation. The benefits are things like better analytics, better customer support, and more productivity. You can make more slides more quickly, and you can do analysis more quickly. It’s kind of tough to put a financial value on that. It has a financial value, but it’s not the same as saying, “We made this new thing with AI, and it had this revenue, or it saved us this much money.”
Those things just take longer. It’s harder to build a new revenue line than to give this to everybody and have them use it to make spreadsheets more quickly. So there’s a little bit of, “Well, how long does this take?”
I think the other answer to the problem here, of course, is consumer surplus, which is kind of what happened with Excel. If a DCF takes you a week, then you probably only do 1 or 2 DCFs. If a DCF takes you 10 seconds, then you do 50 DCFs, but you probably can’t charge any more money for that.
Some of what happens is that these things become competitive necessities, and everybody has to buy and use them. But the cost saving or the productivity gain that you get from them just gets competed away, so you don’t get to charge more for it.
I mean, if you’re at McKinsey, Bain, or BCG, and a piece of analysis used to take a week and now takes a day, you probably do 5 times more analysis and charge your customer the same amount. Your cost base hasn’t changed either. That’s exactly the way to think about what happened with investment banks and financial analysis. You just went and did way more analysis with probably fewer people and charged customers the same amount of money.
Erik Torenberg
Part of your big thesis is this idea that models are going to end up as commodities, and yet the layer that’s raising the most money—in the fastest time in history—is these foundation-model companies. Given that, what advice might you have for them, either collectively? We can pick on someone individually in order to adapt.
Benedict Evans
It’s not that I know they’re going to become commodities. My position is more, “Well, here is a chain of argument that says that, deterministically, it looks like these things will be commodities. Explain to me why they won’t.” That’s as far as I would commit to that.
I think the raising of all this money kind of goes back to my point about mobile, which again has no predictive value but is a worthwhile observation. The mobile industry is very big, spends a lot of money on infrastructure, isn’t very profitable, and all the cool stuff is done by somebody else.
Then you ask, “What’s the return on capital?” The answer is, well, it depends which market—whether you’re in America, Europe, India, or China. But meanwhile, that was a worthwhile thing to do, and it produced a return for somebody. It just ended up not controlling the whole thing, and other people ended up getting more value from that than they did.
I don’t have the number in my head. What was Google’s net income last year—$50 billion or something? What was the net income for the total telecom industry? I should really subscribe to Bloomberg; then I could just answer these questions instantly.
But it’s a pretty safe bet that Google, Meta, Amazon, Microsoft, and Apple produce more profits than the entire telecom industry. This is a puzzle: You’re driving the frontier forward, but you’re caught in this trap that you have to keep competing because otherwise they’ll do it and you’ll fall behind.
You’ve also got this thing that we haven’t talked about at all: Aren’t we just building AGI? We’re going to build God in a box, which some people do believe, although it’s kind of hard to analyze. So you’re going to carry on building this stuff, but the practical question is, how do you get things that people want to use that aren’t software development? I mean, that’s a good business.
Is that the only business? There are, you know, many hundreds of billions of dollars to be made making the software industry more productive. Great—but then what? How do you expand this into the rest of the economy, into everybody else?
That’s why you get these conversations about private equity partnering with consultancies. As we’ve been discussing, it’s actually quite hard to work out what to do with this stuff if you’re running a real company. So you go to Bain, BCG, McKinsey, Infosys, Cognizant, IBM, Accenture, or private equity shops.
There’s this sense that, on the one hand, you’re building these bigger and bigger models and you feel like you’ve got to keep doing it. But on the other hand, what are people doing with it?
Erik Torenberg
Why do most people look at ChatGPT and not really think of anything to do with it today?
Last question: Is there anything from the presentation that you want to make sure listeners leave with?
Benedict Evans
The thing that I used last year and used again is an IBM ad I found from the early ’50s, which has a picture of a sea of engineers all holding up slide rules. It’s an IBM ad, and it says, “An IBM electronic calculator gives you 150 extra engineers.” How many pictures have you seen at a16z where that was the pitch?
We kind of remember that we go through these waves of fundamental technology changes every 10 or 15 or 20 years. They’re all amazing, change everything, and are completely unlike anything that’s happened before.
AI is amazing and transformative and completely unlike anything that’s happened before. Mobile was quite a big deal, too, and so was the internet, and so were PCs, and so was computing. Those were all also very big deals where it was hard to tell what was going to happen.
We should presume as a base case that we’re going to go through that again. That will produce a bunch of things that ruin people’s lives and put a bunch of people out of work. There’ll be a bunch of stuff that we’re not very happy about, and there’ll be a bunch of stuff that we all think is great.
Then, in 20 years’ time, we’ll kind of forget that there was a world when computers couldn’t do that. I mean, here we are. We’ve been on this call for an hour, and our computers didn’t crash. We’re streaming HD video to each other, and it’s like, of course that worked.
In fact, I’m also doing it with my iPhone. My iPhone is streaming to my Mac over Wi-Fi, streaming video here, and it just works. It’s like magic, and we don’t notice it anymore. I think that’s really my one-line description of how all of this is going to end up: It’s going to be magic, and in 20 years’ time we’ll just say, “Well, of course that’s how it is. Computers have always done that.”
Erik Torenberg
Yeah, that’s a great place to wrap. The presentation is called “AI Eats the World,” and it’s on Benedict Evans’ website. Benedict, this has been a great conversation. Thanks so much for coming to the podcast.
Benedict Evans
Thanks. Great to chat.