Erik Torenberg
People are spending a lot on these models. They're presumably doing this because they're getting value from them. You can argue, “I don't think that value is real. I think people are just playing around,” whatever. But they're paying for it, and that's a pretty solid sign.
We're almost giving you here the useful answer of, “I don't think it's a bubble because it hasn't burst yet. When it has burst, then you'll know it's a bubble.” People often make the case, “AI hasn't been profitable yet, and they're spending more to make it profitable.” In reality, they'll have paid off the cost of all the development they've done in the past very soon. It's just that they're doing more development for the future.
Will they regret that spending? How much are they spending? You can look at NVIDIA and how much they're selling each year, and you can see whether it keeps growing and whether things are looking good enough to continue.
Speaker 3
Math seems unusually easy for AI. I'm going to be honest. People often make claims about it being this intuitive, deep thing—that it would mean AI has achieved some huge level of intelligence for it to solve. I think in practice this is just like making a piece of art. It turns out to be farther down the capabilities tree than people might have guessed.
We sort of had this with chess decades ago, right? Computers solved chess very well, and everyone was thinking of this as the pinnacle of reasoning. As a result, everyone concluded, “Of course computers can do chess.”
The interesting scenario to think about—a 20% or 30% chance of something like this happening in the next decade—is a 5% increase in unemployment over a very short period of time, like 6 months, due to AI. The public's reaction to this will determine a lot. There will be very, very strong feelings about AI once this happens. I think there will be a bunch of very strong consensus on what to do, including on things that we don't normally think of as things people are considering.
I know when this happened with COVID, there was a several-trillion-dollar stimulus package in a matter of weeks to days. It was breakneck speed. I don't know what that will look like for AI, but I think it's like everything else in AI: it's exponential, which means it will pass the point of people sort of caring about it to people really caring about it quite fast. I just expect wherever we end up, there will be this certain thing that we would have considered unimaginable a year ago.
Erik Torenberg
Guys, there's a lot of conversation about the macro—are we in a bubble? How should we even think about this question? We're going to get into forecasting later on, but why don't you just take your first stab at how you approach such a big, general question?
Speaker 2
Yeah, for me at least, the way that I thought about this a little bit is that I look at the big indicator being how much people are spending on stuff like compute. I guess maybe some sense of whether they will regret that spending is relevant.
The how-much-are-they-spending thing—you can look at NVIDIA and how much they're selling each year, and you can see whether it keeps growing and whether things are looking good enough to continue. The will-they-regret-it side, we'll actually have to wait and see. It does seem as if most compute gets spent on inference that companies don't so far regret using to offer their products. On that side, I'm thinking, not too bubbly yet. But I have low confidence; there's other stuff to think about.
Speaker 3
Yeah. Right now, the amount of money companies are actually earning in profit, not including the cost to develop the models initially, seems to be very positive. If they stop developing bigger and bigger models and just stick with the ones they've had, they would have earned a profit pretty quickly at the current margins. In this sense, it doesn't seem bubbly.
On the other hand, at any given time, they are investing in building even larger and larger models. If that goes well, they'll earn more money, and if that doesn't go well, then no matter how profitable they are right now, it'll be a small amount of money compared to how much they would have spent. So I think right now there are no financial signs that there's actually a bubble.
A lot of people worrying about bubbles just aren't necessarily used to the level of spending and the level of success that has happened in scaling. But if there is a bubble, it could happen very suddenly and be pretty bad.
Erik Torenberg
Yeah, I think we're almost giving you here the useful answer: I don't think it's a bubble because it hasn't burst yet. When it has burst, then you'll know it's a bubble.
Speaker 3
Yeah. Yeah. I do think you could imagine a world in which there's all this spending and the current level of success does not continue. People often make the case, “AI hasn't been profitable yet, and they're spending more to make it profitable.” In reality, they'll have paid off the cost of all the development they've done in the past very soon. It's just that they're doing more development for the future.
So I think there's this underlying financial success so far that I wouldn't expect to see if there were, at the very least, an obvious bubble.
Erik Torenberg
People are spending a lot on these models. They're presumably doing this because they're getting value from them. You can argue, “I don't think that value is real. I think people are just playing around,” whatever. But they're paying for it, and that's a pretty solid sign.
I guess one quick question related to this: You talked in the report AI 2030 about not having seen signs of these models basically plateauing—the capabilities keep increasing—and you have the benchmarks, the amount of data going in, and the amount of compute. Do you think phases or parts of the models are plateauing? For instance, pre-training: are we seeing some sort of plateauing in that, or do you think people are still exploring innovations at that stage? I'm curious what you think about that.
Speaker 2
Yeah, I think this gets a bit harder to look at. We get to an area where there isn't as much public data to say a lot, right? It seems as if pre-training is comparatively less of a focus than it was before, partly because you have this exciting new—or newish—direction of post-training, where they've done so much with reasoning, whatever.
But I don't necessarily take that as evidence that pre-training couldn't scale further. It seems as if there's meaningfully more data out there. Plausibly, a lot of this stuff is quite synergistic: you develop a better model, use post-training to make it better, and get a load of data from the model actually being used successfully or not. A lot of that can probably go into pre-training next time.
Erik Torenberg
You're not projecting a software-only singularity where AI is able to automate AI research through an automated feedback loop. Why not?
Speaker 2
Yeah, I mean, I might answer this, and then I'll say more. For me, that report isn't one person's forecast or prediction. It very specifically looks at the current trends, whether there are reasons they clearly couldn't continue or might not, and, if they do continue, where they lead.
I think whether you see this self-improvement thing is very hard to determine from a trend-extrapolation basis. Currently, AI does help AI R&D at least a little, in terms of coding or selecting and creating datasets, whatever. But it's quite hard to actually measure, and it's not really helping in some big way, as this kind of self-improving thing would suggest.
There are reasons you might think it could be very hard. People have discussed before how, possibly, if things do just depend a lot on scaling up compute, then automating a lot of the R&D isn't that helpful. I find that somewhat compelling, but I think it's also just pretty uncertain. It's hard to speculate about something that's quite out of regime like that.
One thing that needs to happen in order for a software-only singularity to occur is that you need to be in this world where scaling up the amount of researcher R&D time basically allows you to improve AI enough that it makes up for the lack of being able to scale experimental compute or pre-training.
I think something you would expect to see if this were the case is maybe not that much experimental compute being used in practice, and instead all of the money going toward researchers. Now, there's a very good case that a very large amount of money is going toward researchers.
But as far as we can tell, experimental compute, which you seem to need to do research, is also receiving a similar amount of money. In fact, it's receiving many times more money than the final training runs of the models that are actually being released.
I think this, in my mind, is a strong update toward, “Oh, you actually need to do very large-scale experiments to do research,” and that we don't really have good evidence that researchers—and just researchers—would be able to speed things up without doing more experiments.
However, there are pretty good arguments on either side of this. I tend to lean toward no—you actually need to do more experiments, and that means you can’t get this software-only singularity. But I don’t think the people who claim otherwise are crazy. They have very reasonable differences, and we’re both speculating on something where the data is currently pretty sparse.
Erik Torenberg
Actually, related to that, what do you think? If you have some of the exploration that researchers are trying—obviously, people are exploring a lot with RL, trying to go beyond verifiable domains—what do you think about the argument, for instance, that gradient descent is really good at learning in the current data set that you’re giving it, and if you keep training this over and over, it’s going to start forgetting things that it was trained on before, like catastrophic forgetting?
Is there an argument that kids don’t learn that way? Maybe there’s some imitation learning that kids do, maybe there’s some sort of exploration that they do, and I wonder what you think about it. If kids really would just learn via imitation learning, I think parents would have a great time raising kids, but it seems like the reason they have such a hard time raising kids is because they explore all these different things.
What do you think about it in terms of the algorithms and the things we need to keep improving these models over and over, beyond the data and the compute?
Speaker 2
I am cautious about comparing how AI learns to how humans learn. Not because I don’t think they’re comparable, but because I think we know a lot more about how AIs learn right now than we know about how humans learn. People like making assumptions about how human learning works and saying, “Oh, AI doesn’t do it that way.” I don’t know; maybe that’s true.
Maybe human kids learn via RL. I don’t have strong opinions on whether you need to change to a method that’s more like what we think kids do right now. I suspect people will find some method that works to use the compute available because they’ve been able to do this in the past.
Speaker 3
I’m also reluctant. It’s one of those things where, when we point to particular issues—like the example of catastrophic forgetting—it’s sort of, well, okay, but as we’ve scaled up, we have managed to do quite well at having models that remember more and more things.
This isn’t to say that the problem is solved, that we’re done, or that no more innovations are necessary, but I’m not exactly going to write it off.
Speaker 4
I definitely don’t think we’ve seen any slowdown yet in capabilities from any of these concerns people have. I think people always have these sorts of concerns. I’m reluctant to believe any given one of them until this actually shows up in numbers I can see on a graph, which I just don’t think has happened yet.
Erik Torenberg
Dario at Anthropic said in March 2025 that within 6 months AI would write 90% of code, and of course that hasn’t happened yet. He also said we could have AI systems equivalent to a country of geniuses in a data center as soon as 2026 or 2027. How do you evaluate why Dario at Anthropic is so bullish, or what is the crux of the difference between what they believe and perhaps what you believe?
Speaker 2
My model, at least—which I don’t know if it’s right—is that they think a bit more like the people who believe in automating R&D, and that gives you very quick takeoff. They see it as: We’re working on these AIs that are great for research-engineering-type coding, and at some point they’re going to be useful. That’s going to rapidly accelerate us to develop the next ones, and then it’s going to be quick progress.
I think it’s hard to tell the extent to which their views of this software-only takeoff are wrong, insofar as they will take a little bit longer to get to the minimum level of competence for AI to get you there. That definitely seems to be the case, but I don’t know. It’s hard to tell the extent to which we’ve actually had significant updates on this.
I know Dario often qualifies what he says by saying “as soon as” or something like this. So this is maybe more so the faster timelines he gives, although I’m not sure. There has also been sort of Talmud-style commentary where people are carefully looking at his exact wording and then at the wording of other people’s discussion of how many lines of code are generated by some teams at Anthropic by Claude Code, and whether this does or doesn’t satisfy what you said. So it gets a bit tricky.
I remember there was the uplift paper that was claiming that models would slow you down, but I think it mattered a lot what models they were using at the time, because they were pretty outdated by the time the report came out. In my personal experience, you definitely become way faster, and it just does so much more for you. Having the whole context on your codebase is such a huge advantage that I think it would be really hard for a human to do.
Far more than 90% of the code I write is written by AI these days. But I know I’m not the average coder at all, and I don’t think it’s a wild prediction at this point that 90% of code is going to be written by AI. For all I know, somewhere at OpenAI there’s someone with AlphaCode doing evolutionary algorithms, running tons and tons of trials, trying to million-shot some hard problem. It’s really unclear how many lines of code are actually being written by AI right now.
By a lot of people’s intuitive sense—“Oh, is 90% of the job of a programmer being done by AIs?”—definitely not. There’s this more complicated sense of how much is being written by AI. Probably not 90%, but it’s hard to tell.
Speaker 3
I think that is a very meaningful distinction. If you were to measure how many lines of code are being written, quote-unquote, by tab completion, then it’s probably quite high, but you don’t necessarily expect that it’s taking on that much of the programmer’s really hard work.
That uplift paper you mentioned—I find it really interesting and really good, and it’s also surprisingly recent in a way. You mentioned the models are outdated, but this was early 2025, so these were models that people actually did think were helping them.
In the paper, they even got people to say ahead of time how much they thought this would speed them up, and they said, “Yeah, I think it will speed me up.” They then asked them afterward how much they thought this sped them up, and they were like, “Yeah, yeah, it sped me up.” I feel it does reveal that it might be hard for us to judge whether we were sped up or not.
Speaker 4
One thing that might be happening here is that a lot of the code that’s getting written by AI is code that wouldn’t have been written otherwise. So it’s not really speeding up things that would normally happen. There are a lot of simple graphs or simulations I run that might not have gotten written otherwise, so it’s hard to tell exactly what’s going on here in terms of the impacts.
I think at the end of the day, the most reliable indicator is going to be how much money these people are making from programmers and from subscriptions in general. And it’s a lot of money. I think there are definitely indications that people are finding a use for them, and probably a decent amount of that use is for coding, but not exactly by the metric of doing 90% of an existing coder’s job.
There’s this phrase that’s been used a lot: AI isn’t end-to-end; it’s middle-to-middle. It’s meant to imply that we’re going to need a lot more human involvement than some people typically think. What is your mental model of what AI is going to do for labor markets, either on the lower end or on the higher end, in the next decade, let’s say?
Speaker 2
In the next decade? At the higher end, I definitely expect new jobs to be created. Everyone could still be influencers. But at the higher end, there aren’t very many individual things that you can point to where it’s very obvious that AI can’t automate that job at this point.
You could argue that there are some unknowns, and I think that’s pretty reasonable, but sometimes AI gets up against its limits, we figure out what they are, and then it learns to surpass them. At the higher end, it definitely seems plausible that it could automate basically all existing jobs, with the exception of ones that require manual labor that people actually care about being done by a human. It just does not seem at all implausible to me that that can happen, or that it could happen very fast, with the caveat that there would probably be some regulatory pushback if that happens.
On the lower end, I don’t know. It could just be a bubble and not have any impact. The interesting scenario to think about—which I don’t know, maybe there’s a 20% or 30% chance something like this will happen in the next decade—is a 5% increase in unemployment over a very short period of time, like 6 months, due to AI being released. That’s something I think will have a very substantial impact on the world, both in terms of how people think about AI and how much attention it gets. It seems plausible to me, but it’s far from guaranteed.
Erik Torenberg
Yeah, I strongly agree with being highly uncertain. It seems very plausible to me that you end up more or less with this generation actually being exactly where we run out of progress. It would be kind of crazy, but it could happen. Then it’s like, “Oh, okay, everything is very much just generating more jobs for technical people to try to integrate it into doing useful but janky things for all of the existing work people do.”
The scenario where it becomes a crazy runaway thing that lets you really automate large swaths of remote work is the one where my timelines are probably a bit longer than the others, but it seems hard to rule out that something really big happens in a decade. A decade is quite a long time.
Speaker 2
I think I would be surprised if there were not 5% of jobs that exist now that AI has automated away over the course of the next decade. Honestly, I’d be surprised if it’s not 10% of the jobs that exist now. I think how fast that happens, and the extent to which those people find other jobs, is something for which I don’t think I’ve seen compelling evidence either way.
It probably depends on how fast various things go and exactly what jobs are automated. I think that 10% of current jobs over the next decade seems like a pretty reasonable lower bound—not quite my lower bound, but a pretty reasonable number. But this might not show up in overall employment numbers. Yeah.
Speaker 3
This is interesting. To the extent there is a mainstream economics view of this stuff, it would probably be that automation happens at the level of tasks rather than occupations, and occupations can, as a result, go down quite a bit. But a lot of the time, you’re automating similar tasks across lots of jobs. I think this is compatible with what you’re saying; it’s just that some jobs get really hit by it.
I don’t know. I find it quite hard to think about. I’m not sure what even the historical base rate for jobs ceasing to exist is. I know there are problems with the historical employment data series. There is actually, I believe, quite a high base rate of just the tasks in a job changing, jobs themselves changing, and jobs kind of going away and coming in. So even this 5% thing, I don’t know what to think. That would be a big effect—or, yeah, that’s actually roughly the size of the effect you’ve already seen from something like software. I don’t know.
David Owen
Yeah, probably 5% of jobs that existed before software no longer exist. It seems pretty reasonable, but I’m not confident in this. It’s definitely something where I don’t know. I expect, especially if revenue trends continue, to know a lot more about this in a year or two—probably within the next year—because it will just be the case that, okay, we’ll have AI earning enough to be a substantial part of the economy.
If it’s not showing up in unemployment, then we’ve learned something about what it’s doing: We’ve learned that it’s able to do this without showing up in unemployment numbers. Or maybe it will show up in unemployment numbers, and we’ll see exactly what has happened. There has been some early work looking at indicators of this. There are a lot of things that complicate looking into this, because interest rates also have effects on the sorts of things you might care about, as does normal churn.
It’s also possible that tech companies may lay off a bunch of programmers so that they have the capital to build data centers. Are those programmers being laid off because of AI? I don’t know.
Erik Torenberg
If you had a kid who was a freshman in college and they were asking, “Hey, what should I major in if I want to have a great career?” what might you tell them? And if they asked you about computer science or math or…
Marco Mascorro
Prompt engineer.
Yafah Edelman
I’d probably say not prompt engineer. In general, I think people get better at using AI; it’s very easy to use. Yeah, I think it’s a good question. If they’re majoring in programming, or computer science, the thing they should be looking for is not being a person who’s going to—
The skills that are going to be useful are not going to be knowing a programming language. They’re going to be more general-purpose skills: the ability to work with other people, communication skills, this sort of thing. I don’t really know entirely if this points to a particular major. Most majors are probably not majors that are actually relevant for your job.
David Owen
Yeah, I guess I’d sort of be like, well, there’s not too much that you can do to plan around the super-crazy futures. So I guess go for something that you’re passionate about that’s useful in the worlds, but don’t go crazy in that way. I actually think that computer science and math, if you’re passionate about them, are very good because you’ll learn interesting things that are valuable in many worlds.
But I don’t know. I gave advice to a younger relative recently, and they chose to study drama instead. [Laughter.]
Marco Mascorro
I do think that one of the things is that if you have a better time in college—that’s 4 years of your life—you had a better time during, and at the end of the day, it’s a crapshoot which of those things is actually going to give you a better time in the future. Planning for the present is a lot easier.
Erik Torenberg
Yeah. I mean, it’s definitely becoming really hard to know, right? And I remember the prompt engineer was obviously a joke because everyone believed 2 years ago that that was some sort of viable thing, and obviously models are phenomenally better at just being great prompters. So that’s one thing that’s been happening: It’s really hard to predict what’s happening as these models keep getting better.
One question I have related to this is that, obviously, code is such a big market and has had such a big impact—one that I’m very excited about—but it’s still much earlier, I think, is computer use, right? It’s basically automating all the digital tasks that you’re doing on your computer, and there are very few benchmarks around this, whether it’s WebArena or OSWorld.
You talk a little bit in your report about benchmarks. I’m curious: What do you think is missing in that space? Why haven’t we seen that moment yet—for example, the moment when Claude 3.5 Sonnet came out, or Claude Code, or Codex, when we saw significant improvement in coding? In general, we haven’t had that moment for computer use. What do you think is missing there?
Yafah Edelman
Interesting. I mean, there have been improvements in computer use for sure. I do have—maybe I’m going out on a limb here slightly—but I do think that there is a sense in which models are a little bit artificially hobbled by their vision capabilities. It does seem as if a common pattern you see when you try to get models to do stuff with a GUI is that they kind of get a bit confused about manipulating it, in a way where it’s like, okay, this is interacting with your general propensity to get confused in long, difficult coding problems, but it’s kind of exacerbated because you’re not able to just easily look back at the thing and see, “I was wrong.” Instead, you go down some awful dead end of, “I’m just going to click this again and again and again.”
So I think that’s part of it. I think there’s something here, too, probably about long-context coherence. Those tokens to represent the GUI are pretty big, and then you’re filling up your context window as you go with, “Oh, yeah, well, I had all of this stuff that’s happened before,” and you seem to just run into a kind of spiral of increasingly less sensible outputs. I feel like these are 2 of the big things, but I don’t know if that answers your question.
David Owen
I found computer use—I know this was the first year that I found computer use actually useful. We use ChatGPT agent in our data center research, because a lot of what we have to do is find permits, which are all going to be on janky county-by-county databases of air permits for the county that Abilene, Texas, is in.
I don’t know what databases exist for every county in the U.S. ChatGPT normally can’t search them, because they’re these actual user interfaces; you can’t just search them with URLs, because they definitely don’t work that well. It’s able to navigate them such that I can just ask it to find me permits on a data center in a particular city, and it will come back with air pollution permits, tax abatement documents, and all of this stuff that lets me learn a huge amount.
This is just because of the improvements we’ve seen in computer use over the past year or so. I’m excited to see it get better from there, but I’ve definitely found it starting to get to the point where it’s actually useful.
Erik Torenberg
What’s your mental model, more broadly, for what is going to happen to productivity, or just sort of economic statistics in general? Some people say GDP growth would be 5%. I think that’s a Tyler Cowen view. I think some people would say, “No, no, we should get up to 10% growth, or maybe even higher, if we truly have AGI, in terms of how we understand it.” What’s your model of what happens to productivity?
I think my baseline guess would be: If revenue keeps growing the way it has, in theory, for it to be worth spending that much on those chips to do that inference, you should be getting something kind of similar to that value out of those chips by then.
David Owen
So then you could draw from that: “Oh, okay, extrapolating to 2030, I think in the report it was on the order of a percentage-point GDP increase in a few years, right?” That’s not presuming AGI. That’s presuming NVIDIA’s revenues keep growing as they have previously, and that you assume they make roughly as much compute from it as before, and so on.
If you actually get something—I mean, AGI is used to mean different things—but if you actually get something that can do any task that humans can do remotely, then presumably you see a lot of growth. It feels difficult to guess exactly what kind of lag you’re going to see. There are reasons to think, “Well, maybe people will be slow to adopt this. How do they learn to trust it?” There are other reasons to think, “Well, they’re already using these technologies.” A lot of it might actually be quicker than most growth. Indeed, adoption has been quicker for LLMs than for many previous technologies.
So, yeah, I think it gets hard to model at that point. At some point on our site, we had some rough numbers: What if you doubled the virtual labor force? What if you multiplied it by 10? Then you see these crazy GDP boosts. I don’t know whether that’s the most reasonable way to think about it. A lot of it comes down to whether you imagine that you really get something that can do everything, versus getting something that can do a meaningful fraction of remote tasks but maybe can’t do an entire bucket of them, so it bottlenecks you more.
My best guess on current trends is this fairly well-defined few-percent-of-GDP-in-2030 thing, which is already pretty crazy by economic standards. But once you go much further, it’s like, God, my predictions are just going to be even crazier. I’m reluctant to make them.
Yafah Edelman
I’m going to be slightly less reluctant and make some claims. That’s what we’re here for.
Assuming that in the next 10 years we get AI capable of doing any remote job as well as any human, I think 30% GDP growth seems like a lower bound on something that’s reasonable. This is a big assumption; there’s a lot going on in it. But assuming that happens, I think you either get 30% GDP growth or negative 100% GDP growth because everyone’s dead.
At the end of the day, it seems like you’re going to have AI that can scale. But if you have AI that can scale, you could probably have AI that scales even further. The economic models I’ve seen of what happens if you get this sort of full replacement, where you can automate a job, either show an extremely fast, wild takeoff, or you have some people attempting to model it who then say—and you look down through the paragraphs, and it’s like—“Assuming AI is as capable as GPT-3.”
I think the smaller numbers are either nearer-term predictions or predictions that aren’t looking at the full, more upper end of what capabilities you might see in the next 10 years.
David Owen
Yeah. It does seem hard to imagine a world where you have this supply of virtual labor that literally can do anything humans can do, and then it doesn’t lead to crazy things. I definitely agree with that. I guess there may be worlds in which things don’t go crazy after that, perhaps in some sort of heavy-regulation situation.
Erik Torenberg
There are. They do, yeah.
David Owen
I think there exist worlds in which things don’t go crazy after that. It does seem like those worlds are not in an indefinitely stable state. It’s not impossible, but it does seem like the default is that you either go crazy up or you go crazy down. It’s probably going to be one of those two if you get to a world where AI can genuinely do any job as well as any human.
It seems wild to me to claim that, given that, your default case should be not super ridiculous changes. That’s a lot of things your AI can do right there, and it seems like it should have fundamentally changed the economy in one direction or another.
My intuition is that a lot of the disagreement probably does come down to cached beliefs people already have. But I also think that when people talk about AGI—AI that can do a remote job, whatever—even though we feel like we’re talking about the same thing, sometimes we’re not. I’ve certainly had conversations where someone says, “Yeah, AI that can do any remote job,” and then they discuss stuff that it can’t do. And the stuff that it can’t do is like, “Well, no, that’s also a remote job. That’s the kind of thing people currently do.” So I think there is some of this.
Erik Torenberg
What do you think—I mean, you talk about benchmarks in your report—but I wonder, in 2027 or 2028, what are going to be the right benchmarks to measure progress? More than economic growth, what are the capabilities and the intelligence of the model?
We had AlexNet in 2012, and obviously that was solved long ago. But that was probably not a measure of AGI by any means. Do you think the same would happen with the current benchmarks we have? SWE-bench, MMLU—let’s say we max out on those benchmarks. What comes after that? How do we measure it? Is it sort of GDP growth with these models? Is it breakthroughs in science? What do you think is the right measure going forward?
David Owen
Yeah, I think most of what we have is likely to be solved. Indeed, the examples you gave are already pretty close. MMLU is basically solved. SWE-bench is possibly close; it depends a bit on how ambiguous some of the questions are and on some other details, but it’s really getting there.
Some directions are obvious. You do similar things but harder and a bit better, trying to make them more realistic, and people are doing this. There are harder software benchmarks that people have made more of an effort to curate and that cover larger tasks, for example.
I think there’s also perhaps some question of the budgets involved. Obviously, if you just burn money, it doesn’t intrinsically make the benchmark better, but you are probably going to see a situation where you have to devote more resources on average to them. If you’re trying to prove a higher level of capabilities to a higher standard of proof, it will probably involve more effort in developing the benchmarks.
I also think you’re going to see examples of relatively small numbers of things that are just very impressive, and these will be a valuable signal. When you see LLMs being able to do things like, “It just refactored this entire codebase, and it was really useful,” that’s going to be useful. Even if it’s not yet formalized into a benchmark, if you’ve seen it for yourself, it’s going to be useful evidence. People will probably make benchmarks that cover things like this to try to systematize them.
Erik Torenberg
I want to go back to our question on timelines and ask you about a few different milestones to get your perspective. First, what is a rough timeline for a major unsolved math problem being solved by AI?
David Owen
Oh, I actually wondered about that, because you had a few of these that you said for us to look at. When you say that it solves this, is it entirely unassisted? Or is it a news report, or someone tweets, “Hey, I dumped this into GPT and it solved it”? And what counts as major?
Erik Torenberg
Something that we would all agree is a substantive version of it, not just an anecdotal person describing it.
David Owen
But does it have to solve it on its own?
Erik Torenberg
Yeah, let’s go with that. Sure. Yes, honestly.
David Owen
Oh, yeah. There are already cases of LLMs being useful. People are debating it a little bit, but mathematicians who seem trustworthy are saying, “Wow, I used this, and it was really helpful during my proof.”
I would not be surprised if AI solved a major unsolved math problem, like the Riemann hypothesis or something similar, in the next 5 years. I’m not saying that’s necessarily my median case, but I definitely wouldn’t be that surprised. Right now, math doesn’t look that hard for AI. Some things turn out to be hard and some things don’t, and math is one of the domains where reinforcement learning seems to work pretty well. In most other domains, it’s not at the point where it’s useful to a full professor to the same extent. I think it is for math, or is getting very close to it for math.
It’s also very unclear to what extent certain capabilities that it has unusually well might turn out to be very, very useful. Maybe there are 4 papers out there that it knows about, with obscure results in them, that, when combined, solve some big conjecture. That might be much more feasible to figure out with AI than for a human to figure out.
Erik Torenberg
There's a lot of uncertainty here, but it just does not currently seem like something that AI is actually going to struggle with. People often make claims about it being this intuitive, deep thing, and that it would mean AI had achieved some huge level of intelligence for it to solve. I think in practice, this is just making a piece of art. It turns out AI could do that before it could do a lot of other things—before it could remember things for more than a couple of days, or whatever. It turns out to be farther down the capabilities tree than people might have guessed.
Speaker 2
Yeah, I think I'm also bullish. I do think it's one of those things where it's tricky, and you really probably do need to define it quite well to get a good forecast on it—or hope to get a good forecast on it. We've had this experience with benchmarking mathematics: we got mathematicians to come up with problems that I think aren't as difficult as the kind of problems you're talking about, but nevertheless, they were like, “Yeah, if AI could solve this, it would be a big deal for AI progress. It would mean something to me.” Then AI solved them, and usually their response has been kind of like, “Oh, yeah, that updates me a bit. Although, man, when I look at it, I just realize, yeah, you can kind of brute-force this. You can kind of cheese this. You can get through.” And it's a bit like, “Oh, okay.”
What if there's a problem that, for humans, we consider, “Oh, this would be quite big,” and then AI solved it? “Ah, well, it solved it. Whatever.” We sort of had this with chess decades ago, right? Computers solved chess very well, and everyone was thinking of this as the pinnacle of reasoning. Then they did, and everyone, as a result, kind of concluded, “Oh, well, of course computers can do chess.”
So, I don't know. I suspect that math is quite nice for AI to do. I'm reluctant to go out and assert, “Oh, yeah, definitely AI is going to solve some of the Millennium Prize Problems in the next few years,” but it would not at all surprise me if it solves quite impressive-seeming things in the next few years.
Speaker 3
What about a breakthrough in biology or medicine? We've already seen some of that with—what's it called?—AlphaFold.
Erik Torenberg
Math seems unusually easy for AI, I'm going to be honest. To the extent that I'm asking whether it's going to do the same exact level of, “Oh, it, on its own, did this huge thing,” that seems like a much bigger stretch to me. It definitely seems plausible, but there are a lot of other concerns: it needs to be able to actually do experiments, get data, and interact with the real world for a lot of these in a way that does not need to happen at all for math. In particular, for certain things, they just seem farther off.
What's more plausible to me is that we see it become ubiquitous—that some tools using AI in some aspect of biology or chemistry, or something useful like that, enhance certain aspects of it. It's also possible that AI will make incredible strides without humans, but it's harder.
Speaker 2
Yeah, I think again it's a bit tricky where you draw the line. I mean, I think you're not counting tools like AlphaFold, because if you were, then probably you'd argue for that, right? The inventors co-won the Nobel Prize.
I guess there's kind of different directions in biology. You could have AI being able to predict quite specific things like that, or you could have something that's more general-purpose, this so-called AI co-scientist, or whatever they want to call it, where it's more about, “Oh, it was able to look through the literature and have good ideas.” There are different extents of human involvement.
There already seem to be some results where impressive stuff is happening. I've not vetted them enough to really have a sense of whether this would already count as having satisfied the sort of level of impressiveness you're looking for. I sort of assume that finding things that end up being meaningful will happen pretty soon, if it hasn't already happened. But then maybe there's a question of, “Okay, but is it doing as well as human researchers at actually prioritizing the best few ones to work on?” I think most of these co-scientist results have probably had pretty involved humans prioritizing, though again, I've not looked enough to say.
Speaker 3
Lastly, how about real superintelligence, for your definition of superintelligence?
Speaker 2
I think I am on the record as saying that my modal timeline—which might be on the early side compared to my median—is 2045. When I did the podcast with Himeme, we discussed our forecasting breaking down and everything going bananas—that's the terminology I have used—and that looks like superintelligence.
I think that if we get AI that can do every single job that a human can do as well as any human can do that job in the near future, then this means that scaling just works to get things much, much better. It probably means that you're not that many steps—you're just a bit more scaling away—from getting AI that could do anything humans do vastly better than humans.
It gets hard to predict, and I think as well it gets to be one of these things where the predictions get a bit unmoored from the stuff that you can properly model. My guesses—my judgmental forecasts, to use the fancy term—for AI being able to do any remote-work tasks probably have a median of about 20–25 years. I struggle to imagine a world where that happens and people are deploying it and doing research, yet they're not making further progress toward being able to do stuff much better. So I guess I'd have to say it would be not too much longer after that, for some definition of superintelligence. But, yeah, it's all very uncertain, and it seems to break down a bit.
Erik Torenberg
You talk a lot about progress in data centers, benchmarks, and biology, and there was one interesting part that I noticed: the field of robotics is making a lot of progress with, let's say, world models and the physical space. I'm curious about what your take is here. It seems like a lot of the problems in robotics can be solved purely with imitation learning; you might not need a lot of breakthroughs in math or whatever. You can basically learn it from a lot of data, and I think the last couple of years have been remarkable just in robotics and world models overall. I'm curious about your take on this, and whether you did any kind of research in the space.
David Owen
We've looked into what sort of amount of compute is actually being used to do these training runs. What we found is that the training runs being used for robotics are 100 times smaller than the training runs being used for frontier models. So there's a lot of scaling you can do there. I don't think that, until plausibly very recently, there have been serious attempts to gather data for robotics at a massive scale. You could hire a bunch of people to move around in motion-capture suits if you need to, and there have been a lot of attempts to do that, although I think this might be changing.
I think of robotics as mostly a hardware problem—a hardware and economics problem. If it costs $100,000 to build a robot, then it's not necessarily better than a human who could work for $20,000 a year, or a very cheap human in certain countries, where you might be able to afford labor for something like minimum wage. It's just not obvious to me that there is a software problem here.
The hardware does seem unclear. It's very unclear to me how much of a hardware problem is left. In particular, there are certain tasks that robots might be able to do, but are they actually the tasks that you care about a robot being able to do? If you want your robot to be able to nimbly walk around while lifting heavy things, moving fast, and reacting, then that's hard. That's a hardware problem that I don't think we've seen solutions for yet.
Marco Mascorro
Yeah, I think my impression roughly matches this. People fairly often talk about this distinction between remote work and physical work. I think that's because there's a perception of robotics progress lagging behind a bit, and there even is some intuition that maybe this physical manipulation stuff is actually just harder. But I wouldn't conclude that with much certainty. As you said, it feels like you'd also want to see what happens if it gets scaled up in a similar way, to even get a sense of, “Okay, was it actually harder, or was it just deprioritized?”
Erik Torenberg
Is there anything we didn't get to that you feel is important that we leave our audience with?
David Owen
We did discuss the data centers release you just did. I'm not sure if there's a good way to leave the audience with that.
Erik Torenberg
Yeah, let's get into it. Okay, so you guys just did a release—a project. Why don't you talk a little bit about what you were trying to achieve there and what you hope people take from it?
Marco Mascorro
Yeah. So, we took 13 of the largest data centers we could find.
David Owen
These include several from each of the major labs in the US. We found permits and took satellite images, including new satellite images, of all these data centers. We figured out how to determine how much compute is in them based on the cooling infrastructure they're building, as well as when they're coming online and their future timelines. So we understand this real-world data, and it's all available online on our website for free. This is to give insight into this giant infrastructure buildout that's happening and the pace of it.
There are some things about it that surprise me a lot. For instance, we learned that the most likely candidate to have the first gigawatt-scale data center is Anthropic, which would not have been my pick. Anthropic and Amazon's new Project Rainier development seems on track to come online in January, followed shortly thereafter by Colossus 2.
We also learned a lot about what the largest concrete plans are, rather than just marketing plans. Some people will throw around numbers, but the one we found that's actually seriously underway, has permits, and is setting up the electrical infrastructure for is one by Microsoft, which is going to be used by OpenAI, at least in part, in Mount Pleasant. They're calling it Microsoft Fairwater. That one's going to use a sizable amount of power—not quite as much as New York City, but I think more than half the . What's stopping us from significantly increasing the cluster? Is it cost? Is it supply lead times? Are there any other engineering breakthroughs required? Power.
Marco Mascorro
I think people are approximately wrong that there's something stopping us, and we are scaling up as fast as there's money to scale up, approximately. I suppose they could want all of the clusters literally today, but they're scaling up really quite fast. You're seeing these data centers, which are using—I think the one I mentioned for Anthropic and Amazon is using nearly as much power as the state capital of Indiana, which is where it's located.
The timelines on some of these, like Colossus 2, are 2 years or less, which is just an insane thing: to build this thing that's using as much power as a city. I think that, plausibly, you don't want to buy chips now. You want to wait for there to be better chips. I think people think there's a lot of noise about things being difficult and scaling up, and I think this is because people are having to spend a little bit more than they would ordinarily have to spend.
You can't use the ordinary sort of power pipeline, which is designed to deliver affordable infrastructure at a slow pace. You have to buy things that you wouldn't ordinarily have to buy and spend more than you would ordinarily have to spend, but not buy enough to slow it down. All of these things pale in comparison to the cost of your GPUs.
So my actual takeaway from a lot of this has been, oh, we're not having too much trouble scaling up. These plans are going really quite fast, and it's not obvious that people would actually have the finances and desire to do them faster.
Yafah Edelman
When people are talking about energy as a major potential bottleneck, or having to increase our capabilities significantly, you're not worried that that's going to be a durable, sustainable bottleneck?
Marco Mascorro
I think people like complaining because they can't just use the traditional pipeline of plugging into the grid for cheap, affordable power 4 years down the line. At the end of the day, there are expensive technologies that exist right now. You could pay for solar power plus batteries. These have fairly short lead times. It might cost twice as much as normal power, but that's still way less than your GPUs, so you're going to do it if you have to.
You see people doing these sorts of emergency things that cost them a bit more, starting up their data centers. A common thing we see is people starting their data centers before they're connected to the grid. I think Abilene was an example. xAI's Colossus 1 is a prominent example of just finding ways around this that are expensive, and you complain about it because it would be nice if you could do it the cheaper way, and no one's used to having to do it this expensive way.
At the end of the day, though, there just seem to be enough solutions—especially if you are as willing to pay as people are in AI—that I don't really expect it to be a significant bottleneck.
Yafah Edelman
Maybe let's close with this. If these systems get as powerful as we're discussing, I'm curious how the political system is going to respond. I'm curious if you're sympathetic to Leopold Aschenbrenner's view that there's some potential nationalization that occurs. But, in general, how do you expect governments to respond? It's kind of remarkable how little it is in the political discourse, given how powerful it is already. I'm curious how you think about that.
Marco Mascorro
Calling back to what I mentioned earlier—the concept of a potential 5% increase in unemployment in 6 months—I think the public's reaction to this will determine a lot. There will be very strong feelings about AI once this happens. I think there will be a very strong consensus on what to do, including things that we don't normally think of as things people are considering.
I know when this happened with COVID, there was a several-trillion-dollar stimulus package passed in a matter of days to weeks. It was breakneck speed. I don't know what that will look like for AI, but I think it's like everything else in AI: exponential, which means it will move from people caring about it to people really caring about it quite fast if things keep going.
I just don't know where we're going to end up. I just expect that wherever we end up, it will look like, oh, everyone suddenly agrees that we should do this certain thing, which we would have considered unimaginable a year ago. I don't know what that will look like. It might look like nationalization. It might look like pausing. It might look like—I don't know—going faster, guaranteeing better unemployment benefits. Who knows? I just think there's going to be some sort of strong response, and it's going to be very fast.
Yafah Edelman
Yeah. I mean, you make the point that governments are maybe less interested than you'd expect now, but the current impacts, I think, aren't really that large. I feel like the attention is getting larger, but it's not that AI, as of right now, is that powerful. And yet governments are already talking about it a lot, right? You have people meeting with heads of state from various hardware manufacturers and AI companies, and countries talking about their AI strategy—stuff like this.
So I feel clearly that national governments are going to be quite involved. It's just a question of how, and, yeah, I also am a bit unclear on that. I think that right now we've seen this thing in revenue and finances where it's been doubling or tripling every year, and my default assumption is that the attention that AI gets from policymakers and governments is going to follow a similar trend, where it will double and triple every year.
This means that in the future, if trends continue, there could be a huge amount of attention. It means that right now there's a lot more attention than last year, but you don't suddenly skip from very little attention to all of the attention. Although you do move quite—we are moving, I think, quite fast.
Erik Torenberg
I think we made enough predictions that we'll have to have you back next year and, at the end of the year, check in and see where we're at, and then make predictions for next year. Yeah, David, thank you so much for coming to the podcast.
David Owen
Thank you. Thank you, too.
Marco Mascorro
Thanks so much for having us.