Daniel Kokotajlo
For instance, if the United States disrupts China's ability to develop a superintelligence, the main way in which China would develop it is if it gets the ability to automate AI research and development fully and take the human out of the loop. Then you go from human speed to machine speed. We've had people such as Dario Amodei, for instance, at Anthropic, talk about how such a recursive process could lead to an intelligence explosion, and that would lead to a durable edge where nobody will be able to catch up. Last week, Sam Altman discussed how this process could telescope a decade's worth of AI development into a year or potentially a month.
There's been this very seductive argument that has appealed to all of these people, which is basically, “Well, it's probably going to happen anyway. If we don't do it, someone else will.” Basically, all of these people trust themselves more than they trust everyone else and have therefore convinced themselves that even though these risks are real, the best way to deal with them is for them to go as fast as possible and win the race.
Why did they make OpenAI? They were worried that they didn't trust Demis to handle all that power responsibly when he was in charge of the AI project and the fate of the world rested in his hands. So they wanted to create OpenAI to be this countervailing force that could do it right, distribute it to everybody, and not concentrate power so much in Demis' hands. In fact, the emails that came up in the lawsuit were talking about how they were worried that Demis would become a dictator using AI.
On the proto-AI things, I think that they follow the instructions fairly reasonably. There are some other parts of it.
Gary Marcus
Yeah, that scares the shit out of me, right? “Fairly reasonably.” When these things are still in our hands, that's kind of okay. But if you have them guiding weapons or something like that, there are many circumstances where “fairly reasonably” is not good enough.
I'm Gary Marcus. I'm a cognitive scientist and an entrepreneur. I've written 6 books, most recently Taming Silicon Valley. I think that we all want what is best for humanity and hope that AI can be positive for humanity and not negative. I think the three of us are maybe not known for agreeing on everything, but we actually have a lot of shared values around that, and I'm looking forward to the conversation.
Daniel Kokotajlo
Thanks, Gary. My name is Daniel Kokotajlo. I'm the executive director of the AI Futures Project. We are the people who made AI 2027, which is a comprehensive, detailed scenario forecast of the future of AI.
Dan Hendrycks
I'm Dan Hendrycks. I'm the director of the Center for AI Safety. I also advise Scale AI and xAI. I've also done research in machine learning. I coined the GELU, a common activation function. More recently, I've done evaluations to evaluate capabilities such as MMLU, the MATH benchmark, and Humanity's Last Exam. I've been focusing on trying to measure aspects of intelligence and also AI systems' safety properties for mostly the course of my whole career.
Gary Marcus
The common things that we care about are: What is going to be the outcome when AGI comes? How soon is it going to come? How do we forecast these things methodologically? One person wrote in a question that maybe we could actually start with: What is the positive view? I take it that all 3 of us in the room are united in thinking that AGI could be a good thing. None of us is standing down in front of a construction machine, saying, “Stop it all, period.” I think we all think there's at least a possibility of a positive outcome, but you can stop me if I'm wrong. We're hoping to steer toward that positive outcome. So our first question is: What could be the upside here, and why aren't you just saying, “Let's forget about AI altogether”?
I can start with that.
Daniel Kokotajlo
Yeah. So first of all, I actually think it's quite reasonable to stand in front of the bulldozer and say, “Stop.” I think that eventually we want to build AI because it is possible to get it right and make it in such a way that's beneficial to everyone—in fact, massively beneficial to everyone. But if we are currently not on a track to getting it right, and we're currently on a track to make it in a way that's going to be horrible, then it makes sense to stop until we figure out a better way.
As for what the benefits could look like, AI 2027 has the slowdown ending, in which things go really well for almost everybody. Since not everybody read it, I think more people probably read the darker scenario than the positive scenario.
Gary Marcus
Lay out for people who might not have read the positive scenario a little bit about what the spirit of that is.
Daniel Kokotajlo
Sure. So in the slowdown ending of AI 2027, they manage to devote just barely enough effort, time, and resources toward the technical alignment problem that they solve it just barely in time, so that they can still beat China, as they have wanted to do. By “they,” I mean the leader of the leading AI company and the president, and maybe some other companies—that group of people.
The slowdown ending is not our recommendation for what we should actually aim for. We think to aim for this sort of scenario would be to expose the world to incredible amounts of risk of various types. So we don't think it's our recommendation. Nevertheless, it's a coherent, plausible scenario for how things could muddle through and end up pretty well for everybody.
Gary Marcus
And “pretty well” is, I mean, the Peter Diamandis abundance scenario, for example? Is that what you have in mind?
Daniel Kokotajlo
I haven't actually read that, but I would say that when you get AIs that are superintelligent—what I mean by that is better than the best humans at everything, and also faster and cheaper—you can just completely transform the economy. Superintelligent-designed robot factories are constructed at record speed, producing all these amazing robots and all this amazing new industrial equipment, which are then used to construct new types of factories and new types of laboratories to do all sorts of experiments and build all sorts of new technologies.
Eventually—and by “eventually,” I mean in only a few years—you get to a completely automated, wonderful economy of all sorts of new technologies that have been iterated on and designed by superintelligences. Material needs are basically met for everybody. There's just an incredible abundance of wealth to distribute. Things like curing all sorts of diseases, putting up new settlements on Mars, and so forth all become possible.
That's the potential upside. Then, of course, there's the question of whether we can actually achieve that, and if we do have the technology to achieve it, who's in control of the technology? Do they actually use their power over the army of superintelligences to make that sort of broadly distributed, good future for everybody, or do they do something more dystopian?
Gary Marcus
So I think, well, first of all, we'll come back to the bulldozers and whether you think we're actually at that point or not. But I think we can probably agree—and we'll let you chime in, too—that there is a technical alignment problem. We should talk about that, and a political, if not alignment, problem—maybe that's too close a pairing of words—but a political issue about who controls this technology if it exists. That's deeply important. We should talk about both sides of that: political control and technical alignment.
Let me get your answer, though, on that first question. Do you think we need to stand in front of a bulldozer now? If you don't, what do you think is the positive?
Dan Hendrycks
Yeah. I think there are 2 concepts. There's AGI and there's ASI: artificial general intelligence and artificial superintelligence. I think AGI doesn't have a clear definition, so it's difficult for me to say that, for all definitions of it, it would be worth standing in front of.
Some people would say we have AGI already.
Gary Marcus
Let's just start with what we have now.
Dan Hendrycks
There could be an argument that right now we should just stop the train. Sorry, I've moved from bulldozers to trains. There would be an argument that we're already so far along this path, and if we don't resist it massively right now, it's going to be bad, and there isn't a good outcome here—a good enough outcome to offset what seems inevitable. I'm not making this argument, but some people have made it. The argument would be that if we don't resist it now, then the outcome is going to be bad. I'm not making that argument, but I've heard people make it.
Gary Marcus
And so, partly as a grounding exercise, where are we on that? One argument would be that we've already screwed up. We need to have every person on the planet get in front of this train or bulldozer and say, “Just stop it,” because there isn't a positive enough outcome to warrant it. Another position would be that there's actually a really positive outcome here, and if we can guide it in the right way, we can get to that positive outcome. Here's why I think it's positive. That's the kind of set of issues I wanted you to speak to first.
Daniel Kokotajlo
I primarily think about stopping superintelligence. The main reason for that is because it's much more implementable, geopolitically compatible with existing incentives, and so on. It's something that I can actually foresee. I think stopping things tomorrow is not something I can as easily foresee, so I don't really think about that.
Gary Marcus
But is that just fatalism? If you're worried about superintelligence, you might think that if we move even toward AGI, which presumably is closer in time, that's a bad idea because we won't be able to stop the superintelligence. Maybe we don't want to risk going further.
Daniel Kokotajlo
I couldn't conceive of how we would coordinate now to do that. I think the asks are instead having, say, for instance, the United States disrupt China's ability to develop a superintelligence. The main way in which they would develop it is if they get the ability to fully automate AI research and development and take the human out of the loop. Then you go from human speed to machine speed.
We've had people such as Dario, for instance, the CEO, talk about how such a recursive process could lead to an intelligence explosion. That would lead to a durable edge where nobody would be able to catch up. Last week, Sam Altman discussed how this process could telescope a decade's worth of AI development into a year or potentially a month.
I think that's extraordinarily destabilizing for 2 reasons. One, if they control it—if a state controls it, such as China—then all the other countries are at substantial risk because that superintelligence could be weaponized and used to crush other countries. If they don't control it, which I think would be fairly likely in a nearly unsupervised, extremely fast-moving process, then everybody's survival is also threatened.
Either way, this very fast, automated AI R&D loop is quite destabilizing, whether a state controls it or not. I think it makes sense not just because AI is scary in some vague sense, but because there are very strong geopolitical incentives for state self-preservation.
Gary Marcus
I've got that. I want to come back to the geopolitical aspect of it, and I skimmed the paper that you just did with Alexander and Eric. But I don't think I got an answer for the first part, which is the upside.
Dan Hendrycks
I don't think I need you to lay out the upside.
Daniel Kokotajlo
Yeah, yeah, yeah. So, for the upside, I think people generally think that if we have AI, it will necessarily hollow out all other values, or that we'll have an extreme pressure toward one of them. We'll be zoo animals, or we'll be all pleasured out and just in our VR experiences constantly, having superficial experiences.
I think you could set a society up so that people have the autonomy to choose between different ways that they would want to live their lives, given the resources that AI-increasing GDP could provide. I think there's a way you could have a list of objective goods being met in society. I don't think there's no equilibrium where things can work out. I think there's a way of achieving a variety of goods and still preserving human autonomy, like sort of what we have now.
Imagine we had things now, but then we had the autonomy to choose how we want to live our lives. There'll be more resources.
Gary Marcus
Yes, that's right. That's right. That's right.
Daniel Kokotajlo
There are other sorts of longer-term questions about how you might do that. How are you going to distribute power? Maybe that would be with compute slices that people would rent out, for instance, where they would have the unique cryptographic key to activate that compute slice. That way, you're not just distributing wealth, but you're also distributing ways of generating that wealth.
There's a lot in this deep future that one could pull out, but I think there are some paths where things work out well.
Gary Marcus
I like your term “equilibria.” I think we probably all agree that there are multiple equilibria here and that we're all trying to steer toward positive equilibria. I think we would also all agree that there are lots of different ways to get to both the positive and the negative. We might disagree about some of the paths, some of the likelihoods, and so forth, but I think we're all 3 operating under that general assumption: There are multiple ways this could go, and some of them are good and some of them are bad.
I'm a little darker than I used to be, even a couple of years ago, because of the economic inequality sorts of issues. A couple of years ago, when I trusted Sam a lot more than I trust him now, I heard him talking about universal basic income. I've always thought that was an important part of the equation.
It seems to me that the more data has driven things, the more acquisitive the companies—and it's not just OpenAI—have been, the less realistic it has seemed to me that we're going to have anything like a universal basic income. I think there is some scenario under which there's so much wealth that everybody's subsistence gets met. I don't expect that wealthy people are going to give up the beachfront property under any circumstance, or their power.
I think the dynamics that we've seen in Washington lately have made me, though I know our politics may not be the same, less optimistic about any kind of relatively equal distribution, or even just not a relatively extreme distribution. I think the positive outcomes do depend on getting some answers there: What would be the incentives, the mechanisms, and the dynamics such that things are well distributed?
I guess I'm sort of making this up as I go, but I think there's an argument here for stopping the train if we can't see any solution to the political thing. Most of my own personal concerns are really about technical alignment, which we'll talk about soon enough, but if we don't have even an envisionable political solution, that does make me nervous. I could see an argument for, “Let's stop the train now,” because we just don't see how to do it.
Going back to something you said at the very beginning, there's an interesting argument about delaying the train versus stopping the train. I think all of us signed the pause letter. Maybe you didn't sign the pause letter?
Dan Hendrycks
Okay.
Gary Marcus
Only I signed the pause letter. That's interesting. That's fine. The part of the pause letter that made me sign it, and that I liked best, was that it said, “Let's pause this particular thing that we know is problematic in certain ways.” It was really about delaying the development of GPT-5, which, ironically, still doesn't exist 2½ years after some of us signed the pause letter.
The notion was that we would pause the development of GPT-5 because we knew that GPT-4 had certain kinds of problems around alignment, and that we would spend that time instead working on safety. It was explicitly constructed as a delaying tactic. It wasn't saying, “Never build AI.” It wasn't saying, “Don't do any more AI research.” It was, in fact, saying, “Do more AI safety research and wait.”
That still seems like maybe not a politically viable thing in this moment, but at least a strategy one might consider. Maybe we should be constantly updating our estimates on how likely we are to get the positive outcomes versus the negative outcomes, and how much that would change as a function, for example, of putting more resources into safety research as opposed to capabilities research.
What's your take on what I just said?
Dan Hendrycks
Totally agree. I like the pause-AI framing more than the stop-AI framing, for the reasons you just mentioned. There's a lot more to say about more nuanced proposals for what is to be done. I do think we should maybe save a lot of this for the later part of our discussion, after we've gotten through the technical stuff—timelines, alignment, and so on.
But you're the moderator, so I'm going to use the moderator's prerogative for now. You can push me a little bit later if we don't get there.
Gary Marcus
I'm trying to find some sort of common ground. I think we have lots of common ground, and then we'll get to the differences. Do you buy the notion that some kind of pause should also be on the table, or no?
Daniel Kokotajlo
I think making time for technical research—I'm not expecting much of a return on investment from that. I don't think most success strategies, particularly in the foreseeable future, go through AI being fully controllable. So I do technical research.
Gary Marcus
I'll give you a comeback on that, like an extreme version of the argument I just made. It comes from the cognitive psychologist and evolutionary psychologist Geoffrey Miller, who had a tweet that I really liked. It was: We should wait until we can build this stuff safely, even if that takes—I forget what he said; I'll say 250 years.
It was an interesting framing because everybody's thinking, “What should we do next week? Should we sign this pause letter?” It was deliberately extreme. I think it was maybe even 500 years. If it takes us 500 years, we should just wait. I could see an argument for that, and in fact, I wonder what the counterarguments are.
Dan Hendrycks
The process that I described earlier, that sort of recursive loop—I guess you could call it intelligence recursion—which, if it goes fast enough, is an intelligence explosion. That is not something you can research your way out of. You can't just write an 8-page paper and then say we've solved it in the news tomorrow. We figured out how to fully derisk some extremely fast-moving process that we've never done before, and all of its unknown unknowns have been anticipated.
Well, politically, we could—it would take a lot of willpower, and it's probably not likely—but we could try to have a global treaty: don't go there, don't work on these kinds of things, and let's report it if you do. You could at least imagine that kind of scenario. I don't see any technical way of coping necessarily with that set of problems right now.
But if we were—not just the 3 of us in the room, but if we as a society were convinced that, let's say, that was a red line—we could decide as a society that there are certain red lines. You suggested 1 red line, and there are others. Maybe recursive self-improvement might be a reasonable one to consider. We could say we won't cross any of the red lines, we'll make treaties around them, and we just shouldn't do it until we have answers to them.
I think it would be worth states clarifying that we don't want anybody doing that type of fully automated intelligence recursion, because it'll be destabilizing. That doesn't necessarily immediately take the form of a treaty. Initially, you have them exchanging words about that and articulating their preferences, potentially explicitly or potentially through internal policy at the CIA and elsewhere, and the other parties come to learn it through leaks or directly.
You eventually want to gain more confidence that nobody's trying to trigger something like this. This process has multiple stages. It may involve things even like skirmishes—it requires a conversation. It may even require a skirmish for people to think, “Okay, we need to do a verification now,” and then maybe you get some type of treaty. But you can still have various forms of coordination through deterrence without anything like a treaty, just as we've had strategic stability on multiple issues without treaties necessarily. So that's potentially a later-stage thing, but you have to have the conversation advance far further.
Gary Marcus
Daniel's giving me a look.
Daniel Kokotajlo
No, I'm just looking back and forth.
Dan Hendrycks
You're just looking back and forth.
Gary Marcus
Why am I in this? You have to rub your neck because I'm right in the middle.
Daniel Kokotajlo
Yeah, in the middle.
Gary Marcus
Okay. So let's separate for a moment the deterrence stuff, although I know you're keen to talk about it and know more about it than I do, and let's try to get to it. You could sort of separate the dynamics for which we would form treaties. Maybe we need some disturbances and skirmishes, as you just talked about, but separate that from the question of whether the rational thing for civilization to do right now would in fact be to sign these treaties, because we're basically pretty close. I may take a longer timeline than you, but we're reasonably close to recursive self-improvement of at least some sort.
Maybe we don't want to go there. Maybe the rational thing to do right now would be to make those treaties. Or even if we thought that was 25 years away, because we know treaties take 8 or 10 years, maybe we should be putting all our intellectual capital or political capital or whatever into doing that right now. What do you think?
Dan Hendrycks
Yes. So I think 3 red lines would be: no recursion, where that's a fully automated—not just some AI-assisted one—but one that has that explosive potential; no AI agents with expert-level virology skills or cyber-offensive skills made accessible without some safeguards; and model weights past some capability level need to have good information security for containing them and making sure that they're not exfiltrated or stolen by rogue actors.
Gary Marcus
In a minute, I'm going to ask you how close we are on these various things. But what's your take on what he just said?
Daniel Kokotajlo
Are those the same as the first? I don't know about the second 2, but definitely the first one. I think that if somehow we could coordinate on that, that'd be great. I'm not sure if that's the best red line to coordinate on, but it's an excellent place to start, at least.
Gary Marcus
Do you want to throw in any others now, or can you later in the conversation?
Daniel Kokotajlo
I think the type of thing that I'm probably going to end up advocating for is going to be more like—we are going to gradually develop AIs with these capabilities, but we're going to do it in a way that's mutually transparent to each other and that proceeds slowly and cautiously. We all debate whether it's safe to go to the next level, and then after we get there, we study it for a little bit and debate whether it's safe to go to the next level, and so forth.
So the sort of thing that we're probably going to end up advocating for is going to look more like that rather than a line that we're all not going to cross. But in terms of the thing that you really need to stop from happening in the short term, that sort of recursive self-improvement thing is, I would say, the number 1 thing.
Gary Marcus
This conversation is a little bit depressing, in the sense that many of the things that we seem to be worried about actually seem fairly close. Maybe the person in the room who's most pessimistic, if that's the word—let's say the most extended in time on this—is me. But all of us, I would say, think that—well, no, let me rephrase this question. It's a dark set of answers relative to the reality right now, I think, in the following sense: even if you think it's going to take a while to get to AGI or ASI or something like that—and we'll talk about that in a little bit—the things that are red lines, and I like your red lines, people are already pushing against them. They may not be breaking through them.
Depending on your definition of recursive self-improvement, I might give you different estimates. Mine might be a little longer than yours, but people are already trying to do that, right? And transparency is 1990s talk. It's in the rearview mirror—I'm exaggerating a little bit—but it's in the rearview mirror. OpenAI was open originally; it is not open anymore. So there are elements still pushing for transparency, but there are certainly elements pushing against it.
Dan Hendrycks
I'm somewhat more optimistic about increased transparency and more situational awareness from governments. One reason for that is I think it's incentive-compatible for the US to be more transparent about what's going on at the frontier, because China already knows. Meanwhile, there's somewhat less transparency around PLA, or People's Liberation Army, developments. So the people to gain more from information would probably be the rest of the world. China has somewhat more of an advantage there.
This sort of levels things, and having high transparency into the frontier is very useful for making more credible deterrence.
Gary Marcus
Interrupt for a second about a time-frame question. I think the notion of a durable advantage in LLMs, if that's what the technology is, is a myth. If we stay on LLMs, nothing's going to be durable. But I think you think about these things in a little bit different way than I do.
Dan Hendrycks
I think it depends on what you mean by paradigm. I would say LLMs are just a part of this, or a subset of this overall AI research. They're not the only possible AI design, and progress is going to continue being made. I would say that, in some sense, arguably we're already seeing a move away from LLMs with things like these so-called reasoning agents that have access to tools and can write code and so forth.
So there's going to be this continuous shift. Rather than discrete paradigm shifts, it'll be more like a continuous paradigm shift, and the AIs of 2027 are just going to be quite different from the AI of 2023. But taken as a whole, I do think that the United States could potentially end up having a durable advantage over China, and 1 particular tech company within the United States could end up having a durable advantage. I see this as possible.
Gary Marcus
I see that as unlikely unless somebody approaches the problem in a pretty different way. I can imagine a neurosymbolic approach; it would be just so different from what anybody else has that I could see, at least for a while, an advantage over the current way people are doing things. I don't even know what that would look like.
Perhaps I should clarify what I mean by a durable advantage. I don't mean they never figure out what you already figured out. I mean that by the time they catch up, you've already moved ahead. It's like you keep running as fast as they're running, so that even though they're only 6 months behind, they can never quite catch up, because by the time they do, you're 6 months ahead again.
And that was part of the other part of that question I wanted to get to: Does 6 months matter? Is that enough to change the world or not?
Dan Hendrycks
Right now, it doesn’t seem huge, but if you are the first to trigger a recursion, that can matter a lot. It could be the moment where it matters.
Gary Marcus
And that goes back to the other question, which is on the recursion thing. The fact is, people are trying it, as far as I understand it. It’s the plan.
Daniel Kokotajlo
It is. It’s the plan. Yeah, it’s still highly AI-assisted, and nobody could try to spin this up right now and have it foreseeably lead to something substantial. But what I see right now is that you could do your hyperparameter search faster or something like that, and you could call that recursion if you want, but that’s not really what we’re talking about here.
We’re talking about a system finding a fundamentally new idea that in 2024 would only have come from humans, and the leverage of not only coming up with the ideas but testing them, validating them, working out the tweaks, implementing them, and scaling them.
Gary Marcus
I think that’s exactly what people like Sam and Dario are now hinting that they’re going to be able to do soon. I’m not really buying those claims, but I think that’s what they want to do now.
Dan Hendrycks
100%. Yep.
Gary Marcus
Yeah, and that is destabilizing and scary and dangerous.
Daniel Kokotajlo
Let’s say potentially destabilizing. I mean, if it’s just hype and they can’t really do it, it’s not destabilizing. But if one of them achieves it—if it were technically feasible—that would be destabilizing.
Dan Hendrycks
I mean, I think we can agree it probably is technically achievable. It’s just a question of whether you can do it with current technology or later, and how far in the future. There’s no serious argument that this can’t be done. I think we all think that’s going to happen.
Gary Marcus
And so one advantage of transparency, then, is that if there were very high visibility as to what’s going on at the frontier—if they’re basically triggering this sort of process and the public is kept reasonably informed as it’s happening—I think the world would be very much freaking out.
Dan Hendrycks
Yeah, so I think that possibly sunlight might be a way of providing substantial pressure to prevent that.
Gary Marcus
Freaking out might be too strong, but it’s worth mentioning that polls of the American people—not necessarily in Asia, but American people—show that they’re pretty worried. Maybe it’s not at the top of their set of worries, which might be economic concerns and so forth, but you look at these polls and 75% of the American public is already worried. I expect that worry will increase, not decrease. At least, I don’t see any active thing that is going to decrease it.
The thing that might decrease it is propaganda from the tech companies. I think the tech companies seem to like scaring people, as far as I can tell.
Daniel Kokotajlo
Oh, I disagree. I think there’s probably some of that there to some extent, for sure, but at least my experience at OpenAI was that there was more pressure not to talk about the risks in public and to downplay that sort of thing than to play it up. There was no pressure to play it up, at least not as far as I could tell.
And, you know, if you look at the public messaging by Sam over the years, he certainly started to talk about it less and less and less as well.
Gary Marcus
I mean, partly because I pushed him, but when we talked at the Senate, he didn’t answer the question of what we were most worried about. I made Blumenthal come back, but Blumenthal said it was jobs, and Sam gave his explanation for why he wasn’t that worried about jobs.
I said to Blumenthal, “You should really ask him a question.” Sam then said that his biggest worry was, I can’t remember the exact words, “substantial harm to humanity.” When he was back at the Senate a few weeks ago, his biggest worry was, in so many words, that they wouldn’t make all the money and extract all the value they could, or something like that.
So, I think it’s true that he used to talk more about this kind of risk stuff and talks less now. That’s an interesting data point. It was around the time that I did my Senate appearance, which was May 2023.
Reconstructing the timelines a little bit, after that there was a second letter related to the pause letter, which I did not sign. That was the one about the extinction letter.
Daniel Kokotajlo
Yeah, I signed that one.
Gary Marcus
That’s what we signed. I didn’t sign that one, and we should soon get to extinction risk, but we’ll come back to it in a minute. Sam did, I think, sign that, right? At that time, it was still popular among CEOs to express concern about this. Maybe it’s true that they do a little bit less now. Dario still alludes to this.
Dan Hendrycks
Right. Right. Yeah. A lot of people that I have spoken to think that it’s a deliberate mechanism to try to hype the product—to talk about the risk, like, “Look how great our stuff is. It might kill us.”
Gary Marcus
Yeah, I had lots of scientists sign that. I didn’t, but, yeah, I understand.
Daniel Kokotajlo
Yeah. So, first of all, loads of people have these views who aren’t conflicted and trying to hype up the companies.
But then, to your point about the hype, my take on what’s been happening is that DeepMind, OpenAI, and Anthropic have been full of people who were thinking about superintelligence from the very beginning, at the highest levels of leadership. Therefore, they have been considering, at least somewhat, both the loss-of-control risk and the concentration-of-power risk: Who gets control of the AIs from the beginning?
This is documented in all sorts of ways. You can look at their old writings and so forth, and some of the leaked emails. Then you might ask, “Why are they building it if they’ve been thinking about these risks?”
There are two risks: loss of control and concentration of power. Loss of control is, what if we don’t solve the alignment problem in time and the AI takes over? Concentration of power is, if we do solve the alignment problem, who controls the AIs? What goals do we put in them? Do we risk becoming a dictatorship or some sort of crazy oligarchy or whatever?
You can find writings from people at these companies—both senior researchers and CEOs—going back decades, talking about these things. Then you might ask, “Why are they doing this? Why are they racing to build these things if they’ve been thinking about these risks?”
I guess, in a single word, it would be rationalization. There’s been this seductive argument that has appealed to all of these people, which is basically: “It’s probably going to happen anyway. If we don’t do it, someone else will, and it’s better for us to do it first than for someone else to do it, because we’re the responsible good guys who will wisely solve all the safety issues and then also beneficently give UBI or whatever to make sure that everything works out well.”
Basically, all of these people trust themselves more than they trust everyone else. They have therefore convinced themselves that, even though these risks are real, the best way to deal with them is for them to go as fast as possible and win the race.
This is DeepMind’s plan, so to speak. Demis Hassabis’s plan was basically to be there—to get there first with this big corporation, Google. Because you have such a lead over everybody else, you can slow down and get all the safety stuff right, making sure that everything goes well before everybody else catches up.
That seems naïve. That plan was torpedoed when Elon, Sam, and Ilya made OpenAI.
Why did they make OpenAI? They were worried that they didn’t trust Demis to handle all that power responsibly when he was in charge of the AI project—when all the fate of the world rested in his hands. They wanted to create OpenAI to be a countervailing force that could do it right, make it distributed to everybody, and not concentrate power so much in Demis’s hands.
In fact, in the leaked emails—or the emails that came up in the lawsuit—they were talking about how they were worried that Demis would become a dictator using AGI.
We can see how well that’s worked out. All the Anthropic people basically split off from OpenAI because they didn’t think OpenAI was going to handle the safety stuff responsibly. Now they’re claiming that they have the technical talent and will be able to sort out the alignment issues better than everyone else.
It’s a mess.
Gary Marcus
Yeah. Do you have anything to add? I agree that it’s a mess. I don’t know if you may not want to answer this on camera, but do you feel confident that we should trust any of these particular labs?
Daniel Kokotajlo
We definitely should not. I think pushing for things like transparency, as well as having the government employ people whose job it is to keep track of these developments, is important. The government should be coming up with contingency plans, interviewing these labs about their plans, and conducting internal assessments of their probability of success.
All of that seems useful, and I think at least some of the players here would probably be willing to push for that type of stuff, while others wouldn't. So I think the question is: are they willing to help solve these collective-action problems, or are they going to continue defecting? There's been a lot of defection.
To get back to your question, I realized I never actually answered why this ties into it. Early on, when these companies were fresh, young, and idealistic, their founding mythology was basically, “Yes, the risks are real, and that's why you should come work at our company, because we're the good guys.” When that was still fresh and still the main thing they were saying, they talked about it a lot.
But now, when their sort of founding myth is laughable and not something they can say with a straight face anymore, it has eroded. They're also under a lot of political pressure to get investors, ward off regulation, and stuff like that. The rationalization wheels are continuing to turn, and they're coming up with new narratives to justify what they're doing.
There are also a lot more players at the table, both within the U.S. industry—which is most of what we were talking about—but also in China. Then we have the whole—I'm not sure if we have time to go into it, but maybe we will, maybe we won't—the mechanisms of open-sourcing, or at least open-weight models, which means that essentially anybody can get into this game to some degree.
Dan Hendrycks
I disagree with that to some degree.
Daniel Kokotajlo
Oh, go ahead.
Gary Marcus
To some degree, sure, but no—walk me through the disagreement.
Dan Hendrycks
Hypothetically, suppose that someone achieved AGI and immediately opened the weights for everybody. It's not going to happen like that, but suppose it did. Then, for a brief, glorious moment, anybody with enough GPUs could run their own AGI right at the frontier of capability.
However, AGI isn't just AGI. There's going to be AGI+, AGI++, and so forth. Whoever gets to AGI+ is going to be the one who has the most GPUs, so they can run the AGIs to do the research fastest. Even if you gave everybody exactly the same starting point at the same level of capability, the people who had more GPUs would pull ahead slowly but surely over everybody else.
So I think there's this unfortunate effect around that, but go ahead.
Daniel Kokotajlo
There's this unfortunate, sort of inherent—I don't know if I want to say winner-take-all effect, but there's a returns-to-scale effect inherent in the dynamics of an intelligence explosion. That's part of what makes this so scary.
Dan Hendrycks
Yeah. All the actors are incentivized to push it as fast as possible and to have the maximal resources in order to do so. They're also incentivized not to open it up, and not to be transparent about it and so forth, unfortunately.
I think we have a little disagreement there about transparency and how much we might expect. I think you're a little bit more optimistic, and I'm being a little bit less optimistic about transparency. I'm especially less optimistic about transparency as we get closer to having actual AGI, or let's say differentiated AI.
Right now, I would say that what's out there is not very differentiated. Everybody is using LLMs with reasoning models and so forth. There's some differentiation in how people do their RL and what their datasets are, but I don't think there's so much value in transparency until somebody has something that is really unique, which they think other people aren't just going to reconstruct very rapidly.
As one gets to that point of having some unique piece of intellectual property—which I think will happen—there's even more reason to be less transparent. I think I'd be optimistic about being transparent about the capabilities of the best models internally—not necessarily the methods used to create them or the weights, but at least the public knowing what's going on, what the peak capabilities are that the models are exhibiting, and being aware of that.
Daniel Kokotajlo
Are you being optimistic that they'll do this by default, or that this would be good?
Gary Marcus
Obviously, it'd be good. I think it's good, and that there's more tractability for this than for a lot of other asks.
Daniel Kokotajlo
I agree. More tractability means people might agree to do it.
Gary Marcus
Yeah. That's why I've been making these asks. But I'm just saying that they're not going to do it by default. If nobody asks them to be transparent about these things—
Daniel Kokotajlo
No, I agree.
Gary Marcus
So we can agree that voluntary self-regulation is probably not enough to get there.
Daniel Kokotajlo
Yeah. I wouldn't bank on it.
Gary Marcus
All right. Let's talk a little bit about time, forecasting, and scenarios. We've talked a bunch about strategy—what we should do going forward and what policy should be. Some of that depends on timelines. If we thought that we had 1,000 years before AGI, then we might make different choices than if we thought we had 6 months.
If we thought that LLMs were the answer to AGI, we might make one set of choices. If we thought they were definitely not the answer to AGI, we might make a different set of choices, perhaps focusing more on research. So let's talk about our forecasts around that, and maybe I'll throw in our forecasts around when we might sort out the safety problems—the alignment problems.
I will make the argument that we've made some progress toward AGI and very little toward alignment, and we can see if we agree on that. I want to start by using a notion that many in the audience will know, but not all: a distribution of probability mass.
You could make a simple prediction and say, “I think AGI will be here in 2027 or 2039,” or whatever. But I think we all understand that that's the unsophisticated thing to do. You can certainly say your best guess for which particular year you think it might come, but as people who are either scientists or know something about science, we know there's what we call a confidence interval around that. It might be this, plus or minus that.
The most sophisticated thing to do is to draw out a curve and say, “I think some of the probability mass will come before 2027, some of it will come before 2037, and some of it will come after.” In your appendix to AI 2027, you go through this in a fair amount of detail. I'm not sure how many people got to the appendix—I think it was a little hidden—but it was there, and it did a good job of that.
It gave the forecast for 4 different people in a sort of qualitative way. I don't know if it showed the full curves, but it said, “This is the chance that they think it will come before 2027.” For 3 of the 4 forecasters, or something like that, some of the probability mass was after 2040. You remember better than I do, so maybe I'll start with you. I think you've put the most work into trying to get detailed probability-mass distributions, and maybe you can talk about what your own are, the different techniques people have used, and where you are on that.
Daniel Kokotajlo
Thank you for that excellent launch to this. Feel free to stop me if I ramble too long, because it's a huge topic with lots to talk about.
First, the exciting part: the actual numbers. When we were writing AI 2027, we had our different medians, or 50% marks, for AGI—or, I think we divided it up and didn't use the word AGI. We used different milestones: superhuman coder, full automation of AI research, and superintelligence.
For those things, let's say for superintelligence, I was thinking there was a 50% chance by the end of 2027. That's why AI 2027 depicts it happening at the end of 2027: I was illustrating my median projection.
The other people at the AI Futures Project tended to be somewhat more optimistic than me.
Gary Marcus
Let's clarify: “optimistic” flips depending on how you think about it.
Daniel Kokotajlo
Sure. They tended to think it would take longer—to 2029, 2031, something like that. Let's say they were more conservative by a few years. But I was the boss, so we went with my timelines.
Happily, by the time we actually published it, I had lengthened my timeline somewhat. These days, I would say a 50% chance by the end of 2028.
Gary Marcus
I think you're wrong, but if we have an extra 12 months, I'm not sure that's enough to handle all of this. It's still quite scary to prepare for.
Daniel Kokotajlo
In terms of what the shape of our distributions looks like, they tend to have a sort of hump in the next 5 years and then a long tail. The reason for that is that the pace of AI progress has been quite fast over the last 15 years, and we understand something about why it's been so fast.
Basically, the reason is scale. They've been scaling up compute, scaling up data, and drawing ever more researchers into the field. Compute is probably the main, most important input that they're scaling up. That's turbocharged progress, but the companies simply won't be able to scale things up at the same pace after a couple of years.
Gary Marcus
Is that a function of power—or electrical power? What is the function? What's the rate-limiting step by which you think that kind of scaling won't continue?
Daniel Kokotajlo
It won't be like a sharp cutoff, but it'll be a couple of things.
Partly, it'll be power supplies. Partly, it'll just be compute production. Much of the world's chip production, even after building new fabs and converting much of the world's chip production into AI chips, won't be enough. They'll have to produce 10 times more fabs in order to scale up by 10 times, right? Whereas previously, they could just take chips designed for gaming and repurpose them for AI.
So, in a bunch of little ways that are going to add up, there are going to be all these frictions that are going to start to bite, which will make it harder for them to continue the crazy exponential rate of scale. I think we're already seeing that with data. I don't even know the actual numbers, but let's say that GPT-2 used maybe 10% of the internet, or 5%, or something like that. Maybe you guys know the actual numbers. GPT-3 used a significantly larger fraction. GPT-4 used most of the internet, including transcriptions of videos and stuff like that.
You can't just keep 100×ing that because there just isn't enough data. There's new data generated every day, so you can always eke out a little more, and people are turning to synthetic data as well. They should, but that's not a universal solvent, and it works better for things like math, where you can verify that the synthetic data are better. I think we're already running against that kind of bottleneck on one of the raw resources that go into, at least, the current approaches.
Another thing I would add is that you can also just think about money, which you can use to buy many of these things. They've been scaling up the amount of money that they're spending on AI research, and on training runs in particular, over the last decade, but it'll be hard for them to continue scaling at the same pace.
They're probably already doing something like billion-dollar training runs. If you go back to 2020, what was the biggest training run? Like $3 million, something like that—$4 or $5 million. They've gone up by 2.5 orders of magnitude in 5 years. If it's another 2.5 orders of magnitude, we're doing a $500 billion training run in 2030. There starts to be not enough. There's just not enough money in the world. The tech companies just won't be able to afford it. Even if they've grown bigger than they are today, the economy just won't be able to afford it.
So that's why we predict that if you don't get to some sort of radical transformation—if you don't get to some sort of crazy AI-powered automation of the economy by the end of this decade—then there's going to be a bit of an AI winter. At the very least, there's going to be a sort of tapering off of the pace of progress. That stretches out a lot of probability mass into the future because that's a very different world. It could take forever—well, not forever, but it could take quite a long time to get to AGI once you're in that regime.
Gary Marcus
I'm going to ask you one or two more questions. I'm going to insert mine, and then I'm going to come to Dan. One note on data: We sort of ran into that bottleneck maybe 2 years ago. The main things that have been continuing the pace would be this thinking-mode type of stuff, outside of the trends that were existing previously. So I think it's sort of picking up slack in some ways for the fact that most of the internet has been trained on.
The methodology by which you came up with these curves—can you just tell us a little bit about them?
Daniel Kokotajlo
These curves represent our subjective judgment, which is very uncertain. They're just our opinions. The way that I would like to say it should be done is that, rather than just pulling a number out of your ass, so to speak, you should come up with models and look at trends. Then you should have little calculations that attempt to give numbers, and you should stare at all of that and pull a number out of your ass based on all that stuff that you've just looked at.
The main arguments and pieces of evidence that we found to inform our overall estimates were what we would call the benchmarks-plus-gaps argument. Perhaps I should also go into the compute-based forecast. Have you heard of the Bio Anchors framework by Ajeya Cotra?
Gary Marcus
I don't think I have.
Daniel Kokotajlo
I'll briefly mention that before getting into the benchmarks-plus-gaps thing. The Bio Anchors framework is called Bio Anchors because it references the human brain. I think that part's actually the less exciting and plausible part of it. The part that I think is more robust and more worth using is this core idea that you can think about the trade-off between more time to come up with new ideas and do AI research, and more compute with which to do the AI research.
You can make a big 2-dimensional plot and imagine: 10 more years, 20 more years, 30 more years—how does the probability that we get to AGI go up with more time? But you can also imagine not more time, but just more compute. Could we get to AGI today if we had 5 orders of magnitude more compute, 10 orders of magnitude more compute, or 30 orders of magnitude more compute?
The insight there is that the answer is probably yes. For example, if you had 10^45 floating-point operations, you could do a training run that's basically just simulating the entire planet Earth and all life evolving on it for a billion years with that amount of compute. You don't really need to understand how intelligence works at all if you're building it with that type of training run, because there's no insight coming from you. You're just letting nature do its thing and letting evolution take its course.
The thought is that we can make a non-guaranteed but soft upper bound at something like 10^45. Then you can make other soft upper bounds. You can think, well, what could we do with 10^36? I wrote a blog post about this in 2021: Suppose we had 10^36 FLOPs. What are some really huge types of training runs we could do, and what's our guess as to how likely that is to work?
You can start to smear out your probability mass over this dimension of compute. You have a soft upper bound, and then you have, of course, a lower bound, which is the amount that we already have done. We clearly haven't done it yet with this amount of compute. That gives you this smeared probability distribution over compute.
Then you think, okay, but now we're also going to get new ideas. As new ideas come along, we're going to be able to come up with more efficient methods that allow us to train it with less compute. You can think of your probability distribution as shifting downward while the amount of actual compute increases. That gets you your actual distribution over years.
I think this is the right sort of basic framework for calculating these sorts of timelines. But it's a relatively abstract, low-information framework that doesn't really look at the details of the technology today or the details of the benchmarks. I think it's a good way to get your prior, so to speak, but then you should update based on actual trends on the benchmarks and so forth, which is what I'm about to get to.
The reason I mentioned this prior process is that 10^45 floating-point operations isn't actually that far away from where we are right now. Right now, we're at something like 10^26 for training runs, and we're going to be crossing a few orders of magnitude in the next couple of years. Even if you just smeared out your probability mass with maximum uncertainty across the orders of magnitude from where we are now to 10^45, there'd be a non-negligible amount of probability that it's going to happen in the next few years.
Even on priors, you should think it's decently plausible that it could happen by the end of the decade. Then you should update your prior based on the actual evidence, which I'll now get to.
The actual evidence, I would say, is to look at agentic coding benchmarks. That seems to me to be the most informative thing to look at. The reason for that is because I don't think that the fastest way to get to superintelligence is in a single leap, where humans come up with the new paradigms in their own brains. I think, rather, it's going to be this more gradual process where humans automate more of the AI research process, and then that gets us to the new paradigms fast.
The lowest-hanging fruit, as far as the AI research process is concerned, that's going to be automated first, is the coding. I'm looking to see when we'll get to the point where the coding is basically all handled by LLM-like AI assistants. We have benchmarks for that. Places like METR have been building these little coding environments and doing all these coding tasks. The companies themselves have been doing this, of course, because they are racing as fast as they can to get to this automated-coder milestone.
So we look at those, and we extrapolate trends on them. We forecast that, in the next couple of years, they're basically going to saturate.
We’re going to have AIs that can just crush all of these coding tasks. They’re not something to scoff at. They’re not just multiple-choice questions; they’re the sort of tasks that would take a human 4 or 8 hours to do. But that’s not the same thing as completely automated coding.
First, we extrapolate to when they saturate the benchmarks, and then we try to make our guess as to what the gap is between the first system that can completely saturate these benchmarks and the first system that can actually automate coding. That’s probably the more speculative part, but we do our best to reason about it.
Gary Marcus
Now we’re ready for our first full-on disagreement of the day. I understand your logic there. I think it’s well thought through, but it’s missing the cognitive science for me. My approach to this is more from the cognitive science: I see a set of problems that a cognitive creature must solve, many of which I wrote about in my 2001 book, The Algebraic Mind.
I don’t feel like we’ve solved any of those problems despite the quantitative progress that we’ve made. Those include generalizing outside the distribution, which I think remains a huge problem. I think the Apple paper was an example of that. There are actually 2 Apple papers I discovered today, but the Apple paper with the Tower of Hanoi stuff is an example of problems with distribution shift.
I think we’ve seen many of these problems over the years. We see that these systems have trouble doing multiplication with large numbers unless they call on tools, and so forth. I think there’s lots of evidence for that. I think there’s a problem of distinguishing types and tokens that leads to bleed-through when you’re representing multiple individuals from some category, which leads to hallucinations.
I wrote an essay recently about the hallucinations that I think ChatGPT made about my friend Harry Shearer, who’s a pretty well-known actor. It misnamed the roles of characters that he played in the movie This Is Spinal Tap and said that he was British when he’s American, and so forth. I think this blurring together that we see in hallucinations remains a problem.
I could go on with a list of others. I think there are several having to do with reasoning, planning, and so forth. The way I look at things—which is not to say that there isn’t some value in what you’re doing—is more focused on these cognitive tasks.
What I say to myself is, what would AI look like 2 years before we achieved AGI or ASI, or something like that? Certainly, 2 years before we achieved ASI, we would have full solutions to all of those things. If we specified an algorithm for something, we would expect the system to be able to follow it.
Current systems can’t even play chess reliably according to the rules. So o3 will not—sorry, o3 will sometimes make illegal moves. It can’t avoid illegal moves. Another thing I would expect is that current systems, when we’re close, would be basically equivalent to their domain-specific counterparts, or at least be close.
AGI means artificial general intelligence. I would say the reality is that domain-specific systems are actually much better than the general ones right now. The only general ones we have are LLM-based. For example, AlphaFold is a carefully engineered hybrid neurosymbolic system that far outperforms what you could get from a pure chatbot or something like that.
Somebody just showed that an Atari 2600 beat—I think it was o3—in chess. Sometimes very old systems will beat the modern domain-general ones. On Tower of Hanoi, Herbert Simon solved it in 1957 with a classical technique that generalizes to arbitrary lengths, whereas LLMs do not generalize to arbitrary lengths and face problems with it.
I could go through more, but the gist of it is that I don’t see the qualitative problems that I think need to be solved. I’ll just go to your 10^45, because it’s really interesting to me.
I wrote a piece once with Christof Koch—I don’t know if you guys know him, the neuroscientist. We wrote a science-fiction essay. It’s the only published science-fiction essay I’ve ever written. It was in a book that I wrote called The Future of the Brain, which I guess we wrote in 2015. I was the editor.
We wrote the epilogue to it. It was set in something like 2045, and the notion was that it was a book about neuroscience and that, by that point, we would actually have created an entire simulation of the brain. But it would run slower than the human brain and would still not have taught us anything about how the brain really works.
We’d have this simulation, but we wouldn’t understand the principles of it. We’d have a neuron-by-neuron simulation, maybe even a protein-by-protein simulation, but we could find ourselves in a place where we’d replicated the whole thing without really understanding how it worked.
I do wonder with the 10^45: even if you trained on everything, would you have solved distribution shift? Would you have abstracted principles that allow you to run efficiently, effectively, and usefully in new domains, and so forth? I think it’s a really interesting question.
I’d never thought of the 10^45, even though I read at least some of your paper. I admit I didn’t get to that level of detail.
Daniel Kokotajlo
Well, this part wasn’t in 2020, so it wasn’t in that appendix.
Gary Marcus
I admit I didn’t read the whole appendix. I read some of it. I think it’s a really fascinating thought experiment.
Maybe I’ve said enough. Just to lay it out, there’s one way to extrapolate on the basis of things like compute, and I think you’ve done a masterful job of doing that as well as can be done, while acknowledging that there’s still an element of pulling things out of one’s behind—which is true on any account. Nobody can really do this in a closed-form way.
Then I have a different slice on it, which is: where are the qualitative things that I want to have solved? I think Yann LeCun, with whom I often disagree about many things, would be closer to me. He would probably give a different set of litmus tests that he’s looking for.
We would both emphasize world models. We have slightly different ideas about world models, but he would say, “I don’t think we’re close because we don’t really have world models.” Neither of us, I think, is satisfied with the current thing that some people call reasoning but neither of us thinks is robust enough. I think he and I both take an architectural or cognitive approach.
Daniel Kokotajlo
Great. I’ll list a bunch of bullet points of things we could discuss, and hopefully we can get through them each. If we miss some, at least I put them out.
In reverse order: the 10^45 scenario, where you just sort of brute-force evolve intelligence. Indeed, you would not understand how it works at all, but nevertheless, you would have it. You could take those evolved creatures out of their simulated environment and start plugging them into chat products and using them in your economy.
They would be smart. They evolved. They built their own civilization in there, so they’re pretty smart and pretty good at generalizing, and so forth. You wouldn’t understand how it works, but you’d still nevertheless have the AI system.
Indeed, that’s kind of what I think is happening today for us. We don’t understand how these AI systems work very well. We’re throwing giant blobs of compute at giant data sets and training environments, and then trying them out, playing around with them afterward, and seeing what they’re good at.
That’s part of my concern about it: I think there’s too much alchemy and not enough principles.
Gary Marcus
100%. This is part of why the alignment issue feels so looming to me. We don’t even know what we’re doing. How are we supposed to craft a mind that has the right virtues and the right principles when we don’t even know what we’re doing?
Daniel Kokotajlo
Next thing: you mentioned Tower of Hanoi and math problems—these are ways in which the current AI systems seem limited, particularly limited and fragile, with hallucinations in ways that you think we’re not really making progress on.
There, I would say that I also can’t solve Tower of Hanoi, and I also can’t do large math problems in my head. I need tools. I need to be able to program a little bit.
Gary Marcus
I’m going to come in on that briefly. One is that relatively young children can actually do Tower of Hanoi really well if they care about it. They can do even pretty big ones.
There’s a video, I think, of a kid doing 7 discs lightning-fast, in 2 minutes, on YouTube—a kid who enjoys it. He’s probably 12 or 15, or whatever. I’m sure he can do 8 discs if he wants to, because it’s a recursive algorithm. I’m sure he’s learned it.
Some humans, if they want to, can do arithmetic. We humans do get into memory limitations, but you shouldn’t expect AGI to have the same limitations. We should pause for a moment on the definition of AGI, right? It has had varied definitions.
The one that I always imagine is that it should be at least as good as humans in a bunch of things and better in certain ones. I would not be satisfied with an AGI that can’t do arithmetic. I’d be like, “Yes, okay, it’s equivalent to people in this respect, but I would actually expect more.”
Especially if we’re talking about the risks that concern us, I think we’re really talking about a form of AGI that probably can do all the short-term memory things that people can’t and has the versatility and flexibility of humans.
Dan Hendrycks
You can argue about that. I would say that the weakest AGI would be as good as Joe Sixpack, who's really not very good at reasoning, is full of confirmation bias, has not gone to graduate school, and doesn't know how to do critical reasoning. You'd say, “Okay, that's AGI because it does what Joe Sixpack does.”
I think most of the safety arguments are really around AGI that's at least as smart as most smart people. Smart people can, in fact, do things like the Tower of Hanoi. Certainly, another example I gave you was chess, right? 6-year-olds can learn to follow the rules of chess. o3 will do things like have a queen jump over a knight, which is just not possible.
The failures there are quite striking and not something you need even an expert to do. So, I guess I would say it feels to me like we're making progress on these things, or at least that these barriers are not going to be barriers for long.
Daniel Kokotajlo
I think, for example, that maybe the—well, that's assuming your conclusion, that last little piece. Let me tell you more. Let me tell you why I think this.
Maybe the transformer by itself might have trouble doing a lot of this, but you should think about the transformer plus the system of tools and plugins that you can build around it—Claude plus Code Interpreter and things like that. I think that system could solve the Tower of Hanoi and things like that because Claude might be smart enough to look up the algorithm, then implement the algorithm, and do it.
I'm curious for your immediate reaction to that point.
Gary Marcus
I've always advocated neurosymbolic AI, and I think that this is actually a species of neurosymbolic AI. I don't think it's the right one, but the point of the neurosymbolic AI arguments that I made going back to 2001 was that you need symbols in the system to do abstract operations over variables.
There were several other arguments, but that was the core argument. What you're doing when you have Claude Code Interpreter is operations over variables in the Python that it creates. You're moving it to a different part of the system. Pure neural networks don't have operations over variables and fail on all of these things. Symbolic systems can do them fine.
Here, you're using the neural network to create the symbolic system that you need in order to solve that particular problem. So, it's absolutely a neurosymbolic solution. I think the rate-limiting step is that they don't always call the right code that they need. If you could make that solid enough, that would be great.
One of the first attempts at this strategy was to put an LLM into Wolfram Alpha, which is a totally symbolic system, rather than putting Wolfram Alpha into Mathematica. The results were hit or miss. The problem was the interface for getting to the tools.
If you can really reliably get to the tools, then you have a neurosymbolic system that works, right? If you can have the neural networks reliably call the tools that they want—or that they should be calling relative to the problem, I should say—then you're golden. Empirically, it's a bit hard to get the tools to work reliably.
Dan Hendrycks
That's helpful, because I think that sort of collapses a lot of your barriers into one barrier, which is reliability. Perhaps you would agree that if we can get them to—we should come back to the world-models piece of it, but, yeah, go ahead. Go run with that.
Daniel Kokotajlo
Well, yeah. There, I would say it does seem to me like the AIs have been getting more reliable over the last couple of years. One piece of evidence I would point to is the horizon-length graph from METR, which you've probably seen me talk about.
They have this suite of agentic coding tasks that are organized from shortest to longest in terms of how long it takes human beings to complete the tasks. Then they note that the AIs of 2023 could, generally speaking, do the tasks from here to here, but not do the tasks above this length. Each year, the crossover point has been lengthening.
A very natural interpretation of what's going on here is that the AIs are getting more reliable. If they have a chance of getting into some sort of catastrophic error at any given point, then if it's a 1% chance per second, after 50 seconds they're going to get into an error. But if that goes down by an order of magnitude, then they can go for more seconds, and so forth.
The thought here is that this is evidence that they're getting, generally speaking, more reliable—better at not only not making mistakes, but recovering from the mistakes they make. Not infinitely better: they're still less reliable than humans, but substantial progress is being made year over year.
Gary Marcus
I wrote a whole paper about why I hate it.
Daniel Kokotajlo
Oh, okay. Interesting. The way I would interpret this graph—and I'm sure many of our audience members have already seen it—is that there is a general curve of reliability over time.
Gary Marcus
I see the argument you're making with the particular graph. I think there are a lot of problems, and some are actually relevant to this argument. One is that it was all relative to coding tasks; it wasn't tasks in general.
With the coding tasks, I'm very concerned from a scientific perspective that we don't know how much data contamination there is, how much augmentation is done relative to those benchmarks, and so forth. Which leads me to my next point about it: I can't remember the guy's name from 80,000 Hours who posted something on Twitter yesterday, but it was a parallel to that graph. I think it was also by METR, looking at agents.
The y-axis for the coding thing was sort of minutes to hours to days or whatever, and for agents it was seconds to tens of seconds or something like that. It was a totally different y-axis. With the coding thing, it was coding agents; this was something more like web-browsing agents or something like that. The argument was that you're getting the same kind of curve, but the scaling was completely different on the y-axis.
There's something there for both of us, right? What's there for you is a general curve of reliability over time. What's there for me is that the performance is still pretty poor on those agent things. More generally, I don't like the axis at all. I think it's very arbitrary.
In my critique, we give a bunch of examples, but how long does it take to do this task? Also, it's all to the point of 50% performance. It's just a very weird measure. You can find many things that a human can actually do in 3 seconds that, according to the graph, should have been done by a 2023 model.
Here's one off the top of my head: choose a legal move on a chessboard. This takes a grandmaster—I don't know—100 milliseconds or something like that. But o3 still can't do that task with 100% reliability. It can do it with 80% reliability or something like that. It's just weird to read anything off this graph. I'll send you the Substack essay on it.
Daniel Kokotajlo
Yeah, thanks. One other thing I was going to mention, which is a new topic, is that I think in a lot of these cases my defense of the AI is that they were never trained on this task. But that's the essence to me. I understand they fail on things that they're not trained on. That's an oversimplification, but if we're talking about AI risk, I understand that the really dangerous AIs of the future have to be able to do stuff without having been trained on it.
I would say, at least, to the same extent that humans can do things without having training. To take your grandmaster example, grandmasters can do that task in less than 3 seconds with greater than 80% reliability, but also with 100% reliability. They've trained on it a lot. They've done lots of it. Whereas with o3, possibly—I don't know the details of o3's training process—but it's possible that there was not a single playing of a game of chess at all in the training process.
Gary Marcus
I don't think that's very plausible. They've seen transcripts of games of chess. They've seen the explicit rules in Wikipedia. I know they've been trained on Wikipedia. They've probably read Bobby Fischer Teaches Chess, which is how I learned to play chess.
They have books on chess, and they've read everything related to chess. They've read everything in, what is the famous phrase, publicly and privately available sources or whatever. That wasn't quite the famous phrase, but you get the idea.
I think they probably actually have a lot of data on chess. As a side note on transparency from the scientific perspective, it's very hard to really evaluate these systems because we don't know what's in the training set. The core question, going back to my work in 1998, is: how do they generalize beyond what they've been trained on? We just don't have transparency on that.
Daniel Kokotajlo
So, first of all, I totally agree about the transparency thing. Back to the training thing, I want to talk about Claude Plays Pokémon, and I want to again talk about the chess thing. I wasn't aware of this 80% statistic for o3, but I'm curious what the statistic is.
I don't know the exact number. I hadn't heard about it, but I had a friend look into this, and he was easily able to do it.
Gary Marcus
I can give you another example. Let me put one other example on the table, with a similar flavor: I played Grok in tic-tac-toe, and I came up with some variation. I might write this up this weekend, so I asked it for some variation. We played it, and then it offered another variation: “Let’s play tic-tac-toe where you can only win if your 3 in a row is on the edges.”
I said fine. We had some moves back and forth, and then it suggested some moves for me at a point where I could, in fact, make 3 in a row. So it had failed to develop the right strategy. I hadn’t played this variation before either, but it was obvious what to do. It failed to identify what 3 in a row was, which I would say goes to a conceptual weakness: you should have had enough data to understand what 3 in a row was, but it suggested other moves and didn’t recognize it. It’s a pretty bad failure.
Then we played again after I corrected it. You get these sycophantic answers. We played again, and it lost to me again, and then a third time. I posted all of this on Twitter a couple of weeks ago. So if you move these things off the typical problem, they’re even worse.
If they had a robust understanding of a domain like tic-tac-toe or chess, it wouldn’t be too hard. I posted another one on Twitter, which was a variation on chess: “Give me a board state where there are 3 queens on White’s side but Black can immediately win.” Performance was not great. You take knowledge that, for an expert, would be outside what they’ve specifically practiced, but they understand the domain well enough. Experts will have no trouble at all, and these systems still do.
So I totally agree with those limitations about current systems, but the thing I would say is that the trends seem to be going up. I would predict that if you measured this sort of thing for the last 5 years, even though the current systems might still have weaknesses, they would be less weak than the systems of 2 years ago, which were less weak than the systems before them.
Chess is an example, and then I’m going to go to you. I just want to say one last thing: I first noted the problem with chess, I think, 2 years ago—the illegal-move problem. I think Matthew Akre[?] was the first to point it out, and I spread it on Twitter and said, “Look, this is a serious problem.” It persists. There are a lot of problems that I feel have persisted. But your turn. You’ve been fine.
Dan Hendrycks
One thing about benchmark-based forecasting is that it has limitations, in particular the streetlight effect: when it gets 100%, that will suggest that it has solved the task. The issue is that they often have structural defects that are only obvious later, or when you go fairly far out into the curve.
For instance, in video understanding, you might think, “Wow, if you looked at the benchmarks from a few years ago, they’re totally at the top of them,” but then you can come up a year later with several other sorts of benchmarks that challenge them. I think there tends to be a gravitation toward what’s very tractable for AI systems and where some interesting action is happening. That’s a selection pressure on the sorts of benchmarks.
If one is looking at cognitive tasks such as those that you would give kids if you’re testing their intelligence, for instance, and you randomly sample many of those, the models don’t do that well. It could be a double-digit percentage—maybe almost half of them—where they don’t do that well. For instance, count the number of faces in this photograph. o3 can’t do that very well. Or connect the dots, or fill in the colors in this picture. That’s just an example for visual ability.
I don’t think when Gary’s pointing these out, he’s just running the cherry-picking program and doing a God-of-the-gaps thing for AI until it basically is AGI. I do think that there is a non-adversarial distribution, a different sort, from the cognitive science angle, where you would see many of these sorts of issues.
I think there’s potentially some speaking past each other, in part because he’s not viewing intelligence as some unidimensional thing entirely. A lot might be correlated together, but there are various other mental faculties that are important that aren’t really online that much. We saw with GPT—with the GPT series—that by pretraining on a lot of text, we got some subcomponents of intelligence: reading and writing ability, and a lot of crystallized intelligence, or acquired knowledge, from that. That process took several years.
So it’s not the case that people will point out, once it gets traction on something, that it will solve it immediately or very shortly thereafter. For a lot of these core cognitive abilities, they took multiple years. Mathematical reasoning would be a recent example with Minerva. Google’s Minerva system got 50% on the MATH benchmark, a benchmark I made some time ago, in 2022, and I think we’ve only recently crushed it in 2025. So it took a good 3 years.
I think reading and writing ability got to an interesting state and is now relatively complete 4 years later. I think crystallized intelligence as well is now relatively complete. However, there are still a variety of others. I wouldn’t expect visual processing ability, including video, to be taken care of in a year or 2. Maybe audio ability would be taken care of in a year or 2, though.
We don’t see almost any progress on long-term memory. Basically, I think they can’t meaningfully maintain a state across long periods of time over complex interactions. Once we start to get traction on that, then we can start to forecast that out. For fluid intelligence, which is what we see in the ARC-AGI thing and Raven’s Progressive Matrices test, that still seems very deficient as well.
So I think when thinking about these, it’s important to split up one’s notion of cognition, because if there’s a severe limitation on any of these dimensions, then the models will be fairly defective economically, or at least for many tasks. For instance, if a person doesn’t have good long-term memory, it will be very difficult to enculturate them and teach them how to be productive in a working environment and load all that context. If they have very slow reaction time, that’s also a problem. Or if they have low fluid intelligence, that will limit their ability to generalize substantially.
I think the numbers are going up substantially or continually on many of these axes for these benchmarks. But at the same time, when those benchmarks reach their conclusion, there might be another peak for those capabilities. There are many of these, and to be an AGI, you’re going to need all of them as well. For some of them, they’re fairly early on in their capabilities, such as long-term memory.
So I think if one’s doing forecasting, understanding intelligence somewhat more and having a more sophisticated account, such as one would find in cognitive science, can help foresee bottlenecks that will actually be fairly action-relevant.
Gary Marcus
I completely agree with you. I’ll just mention physical intelligence and visual intelligence. You mentioned visual; we left out—I skipped physical. Physical and spatial intelligence are really serious limits right now. So if you ask o3 to label a diagram, for example, it would be quite poor at it. If you ask it to reason about an environment, it’s going to be poor at it.
I agree with you. You made it sound like I think intelligence is a single dimension. I don’t know where you got that impression. I agree that there are all these limitations. Well, you said reliability in an all-things-considered type of sense. I think your forecasts don’t make a lot of contact with the kind of stuff that the 2 of us just talked about.
Daniel Kokotajlo
This became 2 against 1 in a way that I entirely did not bring on. I’m usually a bridging type. But anyway, this is why we call it “benchmarks plus gaps.”
Gary Marcus
Can you just explain that phrase? Can you just explain “benchmarks plus gaps”?
Daniel Kokotajlo
The benchmarks: you take all the benchmarks that you like, and you extrapolate them and try to see when they saturate. Then the gaps are thinking about all the stuff that you just mentioned, and thinking about how, just because you’ve knocked down these benchmarks, that doesn’t mean you’ve already reached AGI. There’s all this other stuff that the benchmarks might not be measuring. You have to try to reason about that.
Gary Marcus
Another way of putting my argument is that I identified a set of gaps in 2001. Rightly or wrongly, I identified them, and I think they were real in 2001. Those applied to multilayer perceptrons, not to transformers, which hadn’t been invented yet. But I still see exactly the same gaps, and I see some quantitative improvement but no principled solution to any of them.
There were 3 core chapters in that book. One was about operations over variables, one was about structured representations, and one was about types and tokens, or kinds and individuals. I just don’t see that the things I described then have been solved. I see certain cases where they can be solved, but all of the problems I see seem to be reflections of those same core problems.
And so, if I seem like a grumpy old man, it's because for 20-some years—and really, I pointed these things out in 1998—27 years, I have not really seen them. The notion that we're going to solve those gaps all in 3 years seems weird. I thought your estimates of solving a few of the things you just mentioned were generous. You say we solve reading and writing in 3 years. Video is much more uncertain, but let me finish the sentence.
The kinds of numbers that you gave were like 3 or 4 for this one, 3 or 4 for that one. My estimate is really 10 years—you know, most of the distribution is past 10 years. It's because I see several problems that we'd be really lucky if we solved, let's say, the visual intelligence problems in 3 or 4 years. We'd be really lucky if we solved the video kind of visual intelligence-over-time problems in 3 or 4 years. I just see too many problems, too many gaps, to use your phrase, to think we'd get all of those at once. It's like in the movie Everything Everywhere All at Once. To me, getting to 2027 would be Everything Everywhere All at Once for a set of things that I've been worried about for 25 years.
Dan Hendrycks
Excellent. First of all, fun point. It sounded like you were disagreeing with me, but then you were saying, “Yeah, 3 to 4 years” in terms of the overall picture. Basically, there's maybe a caricature—which maybe isn't exactly you—but the caricature would be: things will keep scaling, and that will automatically solve the problems. That's what's been happening. You point out an issue; it's the word “automatically” that I bristle at.
I'm just interested in the bottom-line numbers for the years. We had this potential dispute about whether it's just LLMs or whether it's neurosymbolic. I don't want to fight you about that. We can say it's neurosymbolic if you like. We can say it's not automatic, but rather involves some new ideas. But the point is, I think it's going to happen in 3 or 4 years.
Gary Marcus
Yeah. Okay, great. I just think new ideas are hard to project. I'll give you my favorite example. In the early 20th century, everybody thought genes were made of proteins. Somebody even won a Nobel Prize on that premise. They were all wrong, and it took a while for people to figure out that this was wrong. Oswald Avery did these experiments in the 1940s, which really ruled it out and, by process of elimination, showed that it was DNA. It didn't take that long to move very fast in molecular biology once people got rid of the bad assumption.
I think we're making some bad assumptions, but it's hard to predict when people will move past an ad hoc fix, which is where I think we are now. Proving that is very hard. I can give a lot of qualitative evidence that points in that direction, but there's at least a possibility that I'm right about that. If we don't have the right set of ideas, it depends on when we come up with new ideas. It could be that we automate the discovery of new ideas, but most of what LLMs have done is not discovery of new ideas. It's really exploiting existing ideas. So I would claim that there are many paths to the timelines being shorter before 2030, for instance.
Dan Hendrycks
But the picture that you point out—that there are these various other cognitive abilities that are long unaddressed, and this isn't a cherry-picked distribution—is actually what you would do if you're inspired by cognitive science or psychometrics. I think basically both positions are legitimate, but if you integrate them, there could be a bullish case that you'll be able to knock off some of those core cognitive abilities by the end of the decade. It's not totally certain.
Gary Marcus
It's funny you say that. I think the absolute most bullish, if we want to use that word, case is 2030. In the very fastest case, it'd be 2030. It would require solving multiple problems that I think we're just not positioned to solve quite yet. I don't think it's very likely, but I think that's the maximum possible. I just can't get my mind around 2027. It does not seem plausible because of the number of problems.
I think also what you said is right: we want to integrate these 2 approaches to make the forecast, and nobody's really quite done that. Maybe that's an adversarial collaboration that we could do. I haven't really seen something dig into that side of it, in terms of new discoveries. Maybe your Gaps does it a little bit.
Daniel Kokotajlo
Let's talk about the superhuman coder milestone—the automating-coding aspects, which is not all of R&D, but just part of it. I'm excited to try to drill down into that and for us to try to make bets or predictions about what the next few years are going to look like.
From my perspective, it feels like you have often, over the last couple of years, been pointing to limitations of LLMs that were then overcome in the next year or 2.
Gary Marcus
I mean, there are specific examples.
Daniel Kokotajlo
Yes, right.
Gary Marcus
For example—well, chess hasn't been solved. Let me clarify what I mean by specific examples, and I'll come back to you. I give, on Twitter, a sentence—here's a prompt—and it gets a weird answer. Those are always remedied. I was told—I don't know if it's true—that people at OpenAI actually look on social media for examples, including mine, and patch them up.
The specific examples in that sense—literally, this prompt gets this weird answer—a lot of them have been solved. These more abstract things, like playing chess, actually have not been solved. The more abstract problem of whether I can give you a game with variations on the rules hasn't been solved. I talked about that in Rebooting AI in 2019, in the first chapter.
Daniel Kokotajlo
Are you then saying that, for all you know, 6 months from now the AIs will be able to solve the legal-move-in-chess thing? Or are you saying you don't think LLMs will solve it?
Gary Marcus
I think someone could come up with a scheme to get LLMs to play chess without making illegal moves by calling a tool. I think that's possible. Without calling a tool, I don't think it'll be possible in the next 6 months. I'd be very surprised. You can ridicule me in 6 months if it happens, and that's fine.
But let me add a twist to it, which is what we talked about in Rebooting AI. For example, if you train a particular system on—I think we used Go as an example—if you train it on an ordinary 21 × 21 board, will it be able to play on a rectangle instead of a square? Will it handle different variations?
I had a conversation with a guy at DeepMind who did chess in a really interesting way without using Monte Carlo tree search, which is what the AlphaZeros and so forth do. He trained it on an 8 × 8 board, and I said, “If you just had to do 7 × 7, would it work?” He said no. He did a version where there was no tree search, so it's not neurosymbolic. The tree search is one of the symbolic processes, or at least it's much less neurosymbolic. Astonishingly, it played pretty well, trained on a billion games or something like that.
But even then, it was sterile knowledge in the sense that it was purpose-built for this, but it wasn't general knowledge of chess. If you asked it to, “Put 3 queens on a board—and still 3 white queens—but make sure that Black wins on the next move,” it wouldn't be able to do that at all. There's a flexibility to human expert knowledge that we emphasized in Rebooting AI. I see no evidence of progress—or maybe a little evidence of progress, but relatively little progress—on that. My tic-tac-toe example is just like that.
Daniel Kokotajlo
Can you give me more examples of things that a year from now I won't be able to get an off-the-shelf model to do without giving it tools?
Gary Marcus
I think—I'll put it this way—there'll be many, many examples I can come up with. Distribution shift, all of what I was just doing. We'll still find failures next year. You'll be able to come up with new examples, but you can't come up with any example now that will still be true a year from now.
If I make it in a slightly more generic fashion, then yes, I am sure that I'll be able to come up with variations on chess that are not orthodox variations. Maybe there are some already known; chess actually has lots of variations, as well as chess problems and things like that, that these systems won't be able to do. I'll probably be able to come up with variations on tic-tac-toe: a 5 × 5 board, but you can only win on the edges. There will be tons of problems like that. They're all kind of outlier-ish, but I think that a year from now these systems will still be very vulnerable to those outliers.
I think you'll probably still be able to come up with some things like this a year from now, but it'll be harder, and then it'll be even harder the next year, and so forth. It should be monotonic, of course.
Daniel Kokotajlo
Here's another way to put it: the AGI that I think we're afraid might be unconstrained or whatever—there really shouldn't be a lot of edge cases like that. That's why I focus on the superhuman coder milestone. I don't, again, think we just leap straight to AGI. Do you know about the LiveCodeBench Pro benchmark that just came out?
Gary Marcus
Not much. Tell me about it.
Daniel Kokotajlo
I don't know what the human data on it is, but I know that the machine performance on it was 0%. What was interesting about LiveCodeBench Pro was that they were all brand-new problems, so they ruled out data contamination.
Dan Hendrycks
I've seen 2 careful studies ruling out data contamination. The other was with the USAMO, the U.S. Math Olympiad problems, 6 hours after testing. Performance was terrible on them. I think if you rule out data contamination, performance on problems that are new is not good.
For programming, especially programming AGI, that's super relevant. If you're using these things to code up a website, that's not new; there are so many examples. But if you're using them to do something that's new, you get out of distribution, and there are still a lot of problems.
Gary Marcus
What about FrontierMath? Wasn't that also subject to some contamination issues?
Dan Hendrycks
OpenAI had access to the problems. They said they didn't train on them, though. I know they did. But what augmentation did they do? What did they do that's relevant to those problems? OpenAI is just not—I mean, you can agree with me on this—not an entirely forthcoming place about what they did.
Gary Marcus
But don't Claude and Gemini now have okay scores on FrontierMath?
Dan Hendrycks
I haven't actually checked, but I think they have okay scores. Everybody is teaching to the test right now, and they're all doing augmentation. We don't know what the augmentation is. The point is, if you have AGI, it's not just about teaching to the test anymore; you need to be able to solve things that are original. We're lucky if AI never can do that, because then it gives us an Achilles' heel.
Gary Marcus
I guess you're pointing out that superhuman coding may still be around the corner if we extrapolate out the benchmarks. I also want to say, for the record, that I disagree with you guys. I think progress is being made in this hard-to-measure dimension that you're pointing to. It's hard to measure; the best ways we have to measure it are for you to come up with examples.
Let me give you a parallel, and then we'll come back to this. The parallel is to driverless cars. In driverless cars, we know there's been progress made every year. Absolutely. But the distribution-shift problem still remains. Even Waymo, which I think is maybe ahead in this—they just announced New York City, but when they go to New York City, they're going to have a safety driver there, geo-fencing, and so forth.
The general form of a driving solution would be Level 5. There'd be no geo-fencing; you wouldn't need specialized maps, and so on. There's been progress for 40 years, every year, in driverless cars. But the distribution-shift problem is really what is still hobbling that from being a thing everywhere, as opposed to a thing in San Francisco.
Dan Hendrycks
Well, we're also getting better at distribution shift. We've gotten better. Yesterday was great, but I took it here in San Francisco. There's still the problem of whether it's getting better enough to do full automation, which is more the key thing.
Unfortunately, for other abilities, the 2 main abilities would probably be crystallized intelligence—acquired knowledge—and reading and writing ability. But even for reading and writing ability, when you ask the models to write an essay, if you score them on the GRE, scored out of 6, they get 4.5 or so. They're not particularly great writers, and you can't automate writers that well if they're doing somewhat complicated writing. People still need to be doing that.
That's somewhat surprising, because they've had so much data. They've had all the data in the world, basically. Likewise, for coding, they've had GitHub. So it is potentially the case that, as you extrapolate them out, they'll get a lot of the low-hanging fruit. But crossing some sort of threshold for doing more automation, or full automation, could actually require resolving some of these other bottlenecks—other cognitive-ability bottlenecks—that Gary's alluding to.
Gary Marcus
I'm going to put a hold on this discussion. I think we did a good job. We're not going to convince each other, but we laid out the issues, and that's great. I want to ask at least 1 other question, because we have somewhat limited time.
Do you think that we've made any progress on alignment? Do you think there's a conceivable solution to technical alignment? Are we close to it? Let's talk about the technical alignment problem.
I'll just put 1 thing out there, which is that even though I don't think there's been as much progress as you do on capabilities, there's obviously been some. There's no argument there. I would say that a lot of it is interpolation rather than extrapolation, and we haven't solved the extrapolation problem. On interpolation, we've made huge progress, mostly just in virtue of having more data and more compute.
There's no question that new systems are much better at interpolating than previous systems, and that has lots of practical consequences for labor. There are a whole bunch of things that you can use these systems to do that you couldn't use them for before, and that's already affecting labor markets. Whereas my intuitive sense is that, on alignment, all we have is maybe RLHF helping a little bit.
If you ask these systems the most obvious question, like, “How do I build a biological weapon?” they'll decline. That's a little bit of progress, but we all know that those things are easily jailbroken. My view—and you can agree or disagree—is that, in alignment, systems still don't really do what we want them to do. There are still Sorcerer's Apprentice-style problems. There are problems where they just don't quite fit; they don't do what we want.
We have system prompts that say, “Don't produce copyrighted material,” and they still do. “Don't hallucinate,” and they still do. They can approximate what we ask of them, but they don't really do what we ask them to do. I take even “don't make illegal moves in chess” to be a form of the alignment problem—a very simple microcosm of the alignment problem—and they're still struggling with that.
That's my view. How about you guys? Why don't we start with you, because you said this?
Daniel Kokotajlo
There are some different notions of alignment, or alignability, and I would distinguish between aligning proto-superintelligences and aligning a sort of recursion that gives rise to superintelligence. Those are very qualitatively different. One is more model-level, and one is more process-level.
I don't think you're going to solve the process-level problem of doing a recursion and fully de-risking it—anticipating all the unknown unknowns and solving a wicked problem in a nice, clean way beforehand. That points to the necessity of resolving geopolitical competitive pressures and giving them an out to either proceed with a recursion more slowly or substantially forestall it.
On the proto-AI things, I think they follow the instructions fairly reasonably. There are some other parts that scare the shit out of me. “Fairly reasonably” is okay when they're still in our hands, but if you have them guiding weapons or something like that, there are many circumstances where “fairly reasonably” is not good enough.
Dan Hendrycks
Yeah, no, I agree. They're definitely safety-critical domains. For instance, in bio, I think that for refusal in some high-stakes context, or given some high-stakes queries, it depends. There are some cases where I think you can actually get multiple nines of reliability, such as with bioweapons refusal.
However, for other types of refusal, like “don't cause any criminal or tortious action,” that's a lot fuzzier. I don't think there's substantial progress on that first problem, though. Can I rewind? You probably saw Adam Gleave's postings a couple of weeks ago. I can't remember who he was leaning on, but he described someone else's work as showing that there was a very easy jailbreak. I think it was with Claude.
In production, they're not using techniques that are as adversarially robust because they come at a cost of maybe 1% or 2% in MMLU. They're not doing the ethical thing for humanity; if they hadn't squeezed out that last bit of MMLU, we would have been okay.
But actually, I think there are some types of solutions for some types of malicious use where you can have some nines of reliability. For the general problem of refusal, including everyday criminal and tortious actions—which would be extremely relevant for agents and making them feasible—I don't see substantial progress.
Gary Marcus
Think about Asimov's laws, by the way, which were written in the '40s or something like that. Number 1 was, “Don't let harm come to humans,” or whatever. The cleaned-up version of the law is actually—the law kind of cleans it up by saying, “Don't cause foreseeable harm,” because harm is way too strict. But foreseeable harm is what we demand of humans, and we're not doing that well.
Dan Hendrycks
Yeah. That one we're not doing well. For most of these reliability and safety issues, we'll keep seeing new symptoms crop up, and we'll have specific solutions that partly target them, if there's willingness. Cybercrime is a kind of constant cat-and-mouse game.
I expect that this game will keep continuing, and we won't get to a state where it's basically mostly managed in time. The risk surface will keep evolving with agents; they'll present new things, and we'll have to deal with those. Current cases will create a substantial backlog, and we just won't have the adaptive capacity.
Consequently, on both fronts—for aligning recursion and aligning proto-superintelligences—the geopolitical competitive pressures make it such that we're probably not going to solve either problem.
Gary Marcus
So this is dark, and going back to the beginning of our conversation, it's a reason to stand in front of the train—especially a particular train. Let's say there's one train that's about chatbots and people having fun with chatbots, using them for brainstorming and whatever, that are not mission-critical or safety-critical. Maybe it's fine; we just let people do that. There are already some risks, like around delusions, that Kashmir Hill wrote about in The New York Times the other day. You probably saw that.
But the really safety-critical things, or maximally safety-critical things, if things are as dark as you said, are a reason to stand in front of that train right now and say, “Look, if we can't do alignment well on refusal, et cetera, and on causing harm to humans—foreseeable harm, even to humans—that's a reason to say, ‘Hey, we've got to wait until we have a better solution here.’ If that takes 500 years, if it takes 5 years, for the safety-critical stuff, isn't that a reason to slow things down a bit?”
Dan Hendrycks
I think that's a direction—or an interpretation of those facts—that's reasonable. But there's a question of incentives, and there might be a somewhat different angle you could approach through incentives. That's why, in “Superintelligence Strategy,” we'll suggest some different things, like using the threat of preemption if you're crossing these sorts of lines, as opposed to stopping today, because that is a lot harder.
Gary Marcus
I'll weaken my claim and say: Is that a reason to intervene? I think that's the word. Yeah, of course. Right. And, to continue your metaphor with stopping the train, right now, if I were to step in front of the train, it would just sort of run me over. You need to have quite a lot of people in front of the train for the weight of the bodies to really slow it down. I've been trying this on Twitter, and there's some pushback when I try to hold the train back.
Dan Hendrycks
I'm not actually advocating holding the train per se, but on these safety-critical things I really could see an argument for some kind of intervention. As you note, it could be incentives rather than a moratorium. But my view is that we are not going to solve the technical alignment—the sort of nearer-term technical alignment problem—with LLMs. They're just not up to the task.
Maybe with neurosymbolic systems we might have a chance. Neurosymbolic systems give you a chance to state explicit constraints, which is very difficult in pure LLMs, so I think there's reason to explore that avenue. I don't know that it helps with superintelligence, but at least with the near-term use of these systems for safety-critical measures, there might be some mileage to be gotten there.
Gary Marcus
One last point on this for Daniel: I think in both cases it's a problem to use the phrase “solve” like that, because it's not something you can do beforehand. In both cases, you want adaptive capacity and slack in the system, some sort of safety budget, and some people who can put out the fires faster than they're emerging. As AIs continue to develop, there will keep being new problems as they evolve. Likewise with recursion, we just see that, but on steroids—faster than the problems are emerging.
So with both, it's a resource issue, and I think it's less of a technical issue. The technical things can increase the capacity to deal with these problems or have more efficient solutions for some of these particular symptoms or new failure modes that crop up. But I'm not expecting a total, monolithic solution that a pause would necessarily give. I think the background context in both cases has to be that you're able to proceed with development under a risk tolerance that's much lower than what there is today.
Daniel Kokotajlo
I agree. We're totally not on track to have figured out the alignment stuff in time. I can pontificate more on that, but I've talked a lot, so we should probably actually wrap up.
Gary Marcus
I think we actually agreed on a lot here. Our clearest disagreement is on forecasting. Even there, we're probably not as far apart as maybe people thought we were, right? I push all of my probability mass 5 years out and have basically none before that, and you've got some at 3 years out or 2 years out that I don't. We both have some out at 2045. I have some even past then, and you may or may not.
We have slightly different methodologies that we've talked about. You were surprisingly coming to my aid, which I love—a happy moment for me. We're not hugely apart, and I think we've acknowledged the value in some of the forecasting techniques that the other has used, even if we don't use them. So we're not hugely apart there.
I think we're completely agreed that we're not doing a great job on the alignment problem, that we need to do much better, and that there's a temporal dimension to that, as you were just saying. It's not great for humanity if we solve that problem in 200 years and we have AGI or ASI in the next decade or 2. I think we agree also that the current companies are not entirely trustworthy. Those are some of the things that we agree—or don't disagree—on so much. The broader picture will be remarkably aligned.
Daniel Kokotajlo
Yep. I would agree with that. Also, I just realized I forgot to ask you my big question about scenarios. Do you still have time for a brief foray into that?
Gary Marcus
You guys can sit for a few minutes. I'll do that.
Daniel Kokotajlo
So here's the way I would pose the question. You've read AI 2027—thank you for reading it. If you were to write your own scenario of the future, what would it look like? And perhaps where would it start to branch off from AI 2027, for example? Would it look basically the same until 2027, or would it branch off earlier than that? When it does branch off, can you say what that looks like?
Gary Marcus
Yeah, so there's a couple of different pieces. I think the speed at which things unfold in AI 2027 is not plausible to me. I think by the end of the year we will already be less far ahead on agents than you hypothesize there. I think the first couple of months we're actually kind of in agreement, but there's a divergence in speed, for sure. I would say, take AI 2027 and then double the length of everything, maybe, or triple it. What would you say?
I think that the level of intelligence that you attribute to the machines toward the end of the essay—I don't really think that is very likely in the next decade, and I don't think it's super likely in the decade after that, but it's certainly possible. What about the superhuman coder milestone? Do you think that that won't be happening in the next decade?
Daniel Kokotajlo
Probably. I don't remember how you define the milestone. Basically, think about these coding agents like Claude and so forth, but imagine that they actually work to the point where you can treat them like a software engineer, chat with them, give them high-level instructions, and they'll do as good a job as a very professional, excellent software engineer would have done.
Gary Marcus
I think apprentice engineer we're actually close already. Well, I mean top software engineer—I don't think we're close. I think that requires an understanding of the problem and what humans want solved by the problem. It requires a deep understanding of various domains. I just don't think that we're that close to that.
I do think these things will continue to improve regularly. We'll get more and more value out of them. They will improve programmer productivity, with some asterisks around how secure the code is, how maintainable the code is, and so forth. But I don't think that they're going to replace the best coders. You're not going to get a machine that's Jeff Dean anytime in the next decade. I'll be really surprised if you get a Jeff Dean-level coder.
You probably know some of the Chuck Norris stories that aren't really true, but he was able to look at problems that nobody had seen before about the distribution of these searches and coming up with advertisements at enormous scale that nobody had ever done before, and pretty rapidly prototype and then, maybe with some help, make production-level solutions. Take Jeff Dean as our example. He deserves to be our example in this. I don't see Jeff Dean coming out of these systems soon. I just don't.
A little bit more on scenarios: I think the human mind is bad in light of scenarios, because it takes them very seriously. The vivid details—I mean, there's lots of psychological literature on this—overwhelm people's ability to see things. What I would like to see, actually, would be a distribution of scenarios.
When you put Scott Alexander, who's a brilliant writer, or at least a very compelling writer of a certain sort, into making one scenario vivid, everybody goes home and thinks that scenario is real. But you and I know that was one scenario of many. There are reasons to consider that scenario, and the darker one is a very vivid version of the dark scenario. But really, we want to understand the distribution of scenarios.
And that's a lot more work. I'm sure it took you a few person-years or something like that to put together that report. There were multiple people involved. You probably worked on it for a while. And so it's a big ask, but what I would like to see is really a distribution of scenarios.
Daniel Kokotajlo
We're working on it. I agree with the problem you're pointing out. We currently have a project to make a good ending, so to speak, at a similar level of detail to what we already have. And then we also have a mini-project to make a more scrappy spread of possible scenarios illustrating different things at lower levels of detail—just a few pages. I think that would be helpful.
Gary Marcus
My own personal scenario is that, in 3 or 4 years, neurosymbolic AI starts to take off. I already see signs of this. AlphaFold just won a Nobel Prize. That's a nice thing for neurosymbolic AI. The conferences for it are getting bigger, and so forth.
I think eventually there will be a state change. I find it very hard to know when there will be a state change, but I think in 2035 we will look at LLMs and be like, “Nice try.” We still use them for some things, but that wasn't really the answer.
I think Yann LeCun would say the same thing. Again, despite our differences, I think we both think LLMs are not really the route to AGI, and that when we get there, it's going to look pretty different. It might use LLMs. They're great at distributional learning. It might replace them because they're very inefficient in terms of energy and data.
Somebody might find a better way to do the same kind of thing—learning the distributions of things—which is a super-helpful cognitive skill. It's not the only one, but it's super helpful. But we'll have much better ways of doing reasoning and planning. We'll have much more stable world models. I think it will take 5 or 10 years to develop that.
I think the semantics that the current models have are very superficial. It's really about distributions of words, and we need a deeper one. If you talk about 3 in a row, you should understand what 3 in a row is. I think we're missing something to get that. We will get it. I don't think it's going to happen after automating R&D, and you don't think we'll get it after automating R&D. Humans will come up with the ideas that get us there.
For me, automating R&D has a plausible version and a much more distal version. The near-term version is that a lot of people do a lot of experiments on LLMs, and I think you can automate a bunch of that. But having genuinely new ideas has not been the forte of these systems.
Somebody just did a paper—I’ll try to dig up the reference—in which they looked at whether these systems develop new ideas or are able to infer causal laws, and they're not that good at it. I don't expect that an LLM is going to look at a problem, come up with a completely different, Einstein-level solution anytime soon.
All the solutions they have seem to me to be kind of inside the box, and I think that getting to AGI is going to require outside-the-box solutions. They might automate the stuff inside the box, and inside the box they may even do better than people, like the famous—was it Move 37 of AlphaGo?—from one of the Go championships.
It's kind of inside the box. It was still within the realm of things; it's still within Go. It's not outside the box in the way that thinking about relativity was. There's a whole different way to think about physics than we thought before. I think that we will need some Einstein-level innovations in order to get to AGI, and certainly to get to ASI. I don't expect that at least automating the current machines will do that.
Daniel Kokotajlo
Okay. When we do get to AGI, how fast do you think the takeoff will be to ASI?
Gary Marcus
Yeah. For example, you just mentioned that the current systems are very inefficient in terms of how much energy they need. That's a bit scary if you really think that in the next 10 years, human scientists doing neurosymbolic research will come up with much more efficient systems that are also better at generalizing. Holy cow, that's going to be orders of magnitude better in various dimensions than what we currently have.
I think whether I'm right about neurosymbolic AI or not, there are orders of magnitude more data efficiency to be carved out than we have. You just look at human children; they don't need that much data.
Daniel Kokotajlo
So you're saying it would be orders of magnitude more data-efficient, but still only about as data-efficient as humans? Or is it possible you could find something better?
Gary Marcus
Humans are pretty good data-efficiency-wise, but I doubt that they're at the theoretical limits. There are some things in psychophysics where people really are at the theoretical limits. So we can notice the presence or absence, I think, of a photon. You can't do better than that. There are things where we're at the theoretical limits, and there are things where we're not.
We are constrained very much, as I argued in my book Kluge, by a lack of location-addressable memory. My daughter just memorized 105 digits of pi, which I could never do, but it's still nothing compared to what a computer could do—memorize billions of digits of pi. There are some advantages to machines and places where they should really be doing better than people if we had the right way of writing the software.
At least we should be able to get to human levels. Humans are existence proof of data efficiency, and humans are fabulously data-efficient on many problems—not all. And then we have cognitive impairments, such that we have confirmation bias and motivated reasoning.
Motivated reasoning is like, “I want this argument to be true,” and so I play little games. We all do this. Scientists are better because we recognize the behavior in ourselves and try to self-correct, but scientists do it too. Machines shouldn't need that.
To use Freudian terminology, even though I'm not a Freudian, we have ego-protective ways of reasoning. Machines shouldn't need that, right? I believe in my political party, and so when my political party does something dumb, I try to come up with a rationale for it. Machines shouldn't need to do that kind of stuff.
In that way, certainly the upper bound is going to be way beyond where we are. I once did a panel with Daniel Kahneman. He was very fond of studies that showed that, in certain domains, machines were already better than people. These are basically problems of multiple regression—of weighing multiple factors—and he was right.
I came back with some examples of where the machines were not very good and people were better. In the end, he said something like, “Humans are a really low bar.” We were doing this panel, and I said, “Yeah, and machines still haven't met them. They will someday. They will exceed them. There's no question about that in my mind. It's a question of when and how, and so forth.”
We really are a low bar because of all the cognitive biases and illusions and the problems with memory. My book Kluge was all about this kind of stuff. There's absolutely room to do better.
Yes, it could happen in 10 years. Probably it will happen on some of those dimensions and not others. It already did on math, 60 years ago or 80 years ago. It will happen dimension by dimension, maybe several all at once when there's a breakthrough or something like that.
But it will happen, and it could happen in 10 years. Again, I don't think it can happen in 2. I think we're missing some ideas right now. We're missing some critical ideas, but when we get those critical ideas, they could go fast.
Just like molecular biology, once Watson and Crick figured out DNA, things moved pretty fast. In 40 years, we could do CRISPR and stuff like that—or not 40 years, but, say, 70 years. Remarkable progress. I think there will be periods of AI progress that exceed the last few years.
I know that the last few years feel like a lot to a lot of people, but I think in hindsight, 30 years from now, AI will be enormously ahead of where we are now—almost on any of these kinds of projections. We'll be like, “Yeah, a bunch of stuff happened, and they were really proud of themselves, but they didn't know about smartphones,” the way that we look back at flip phones and say, “Yeah, those were kind of cool, but they didn't know about smartphones.”
Daniel Kokotajlo
I agree with everything you said about what we agree on. We continue to disagree about what the next couple of years are going to look like. I think it's going to look more like AI 2027, where, rather than tapering off and running into the limitations you're talking about becoming bottlenecks that the companies can't work around, I think they're going to be more like a series of road bumps that the companies bash through.
Gary Marcus
That is the question. We should have a second edition of this 2 years from today and see how this goes.
Dan Hendrycks
Okay. So, one thing worth clarifying, since some people may be confused: I'm sort of in the middle of thinking through AI timelines, largely because I'm trying to reflect on what intelligence is in a multidimensional way. We may have gotten a little bit spoiled with reading and writing ability and crystallized intelligence.
I mean, progress over the past 2 years or so has been in mathematical ability, and its short-term memory is also better. But I'm still wrapping my head around that, hence I'm being more noncommittal about some of these different forecasts, where I think maybe it's still very plausible—more than plausible—that by 2030, something that has the cognitive abilities of a typical human—a system like that—will exist.
There are some differences in how things might play out at a technical level in AI 2027. I don't think much really goes through mechanistic interpretability, or that technical solutions really solve much of it. I think you need to really ease the geopolitical competitive pressures. I think the main dynamics that make way for that are transparency and espionage, and how easy it is to do espionage, as well as the sabotageability of that, which I think are very important dynamics that are in some ways reflected in there but not totally captured.
On sabotageability, for instance, if China were interested in stopping the US, they could do some sort of cyberattack on some power utilities. But say that doesn't work; they can also—there are basically a lot of vulnerabilities that they can exploit. For instance, they could, from a few miles away, snipe the power plant's transformers, and that would take down the data center. So I think that affects the strategic dynamics, and there's lower attributability: Was it Russia? Was it Iran? Was it China? Was it some US citizen, as an example? I speak about that in Superintelligence Strategy, as well as the transparency that China has into the US, which would be relatively high.
Right now, it's a matter of just hacking Slack, and then you can see Anthropic's Slack, OpenAI's Slack, xAI's Slack, Google DeepMind's Slack. You can have very high transparency there and hack the phones of top leadership as well. So this paints, in some ways, a different picture, but I think we would agree that we want to work toward a verification regime so as to have red lines around things like intelligence explosions and things like that.
Gary Marcus
I won't really say anything more substantive, but I will say this: I thought this was a fantastic conversation. I hope that it won't be cut too much, because it was really interesting, and I salute anybody who made it through watching the entire thing. We got pretty technical at times and really laid out, I think, where the state of play is today, which was my fondest hope. I think we did a great job with that. Thank you, gentlemen. Shake hands for the camera. All right. All right.