How Many Narrow AIs Could Behave Like One Superintelligence - Daniel Kokotajlo and Thomas Larsen
Daniel KokotajloThomas LarsenTim Scarfe
- The AI Futures Project's new scenario, "AI 2040: Plan A," is a deliberate mix of prediction and recommendation: a 6–12 month hard pause to build verification infrastructure, then a US–China deal to advance slowly to roughly top-human-expert AI and hold there through the 2030s, reaching superintelligence only in 2040. The core dial is "pause at the maximum level that you can reliably control," and unlike the gloomy AI 2027, "if that gets hyperstitioned I think we'll be pretty happy."
- AI 2027 is tracking at roughly 75% of predicted speed on resolved quantitative metrics — and Thomas Larsen says reality has diverged less than he expected, which "surprised me in a bad way." Revenue trends and other real-world indicators are "pretty close to on track"; the apparently weak 0.17% software-R&D-uplift result was a measurement artifact from overestimating the baseline at publication, not a miss on progress.
- The milestone that matters for timelines: "the point at which an AI company would rather fire their humans than fire their AIs." Larsen argues there is no binary between verifiable and unverifiable tasks — a unicorn startup's billion-dollar valuation is "a verifiable fact about the real world," just expensive and long-horizon — so RL will expand continuously outward until AIs cover "everything that the humans can do."
- The economic thesis is that the economy is, and always has been, a self-replicating system — and an all-machine version doubles far faster than the current ~20-year doubling time. Even a pause at human-level AI produces "colleagues in the cloud" that are cheaper, faster, and double every year instead of reproducing over 20 years; Thomas Larsen's limit case: "who knows what's going on in the rest of the world, but Anthropic has disassembled the moon."
- Plan A's transparency regime is deliberately hostile to frontier-lab equity: publishing all core training recipes "will cut into their valuations dramatically" by letting Microsoft or Alibaba catch up — "a feature, not a bug." The stated goal is that AI "commoditize instead of being monopolized," and reduced monopoly rents deliberately deter trillion-dollar cluster investment; the gift-to-China objection is answered with horse-trading (e.g., a more favorable compute distribution) and the observation that lab security is so poor China likely gets the algorithms via spies anyway.
- The alignment problem gets harder, not easier, as models improve: control carries "a time bomb," and the core failure mode is silent. Models already recognize Redwood Research control evaluations mid-eval, and Thomas Larsen warns "it'll be extremely easy to end up in a situation where the AIs are in fact misaligned, but you don't know that because they're doing everything right as far as you can tell" — behavioral evals won't suffice; white-box interpretability breakthroughs are required before scaling past the controllable range.
- The fundamental crux with skeptics is a single reference-class question: "is AI more like electricity or airplanes, or is AI more like humans in the cloud?" Thomas Larsen identifies recursive structural self-adaptation as his tripwire ("then that's it for me... that's not a normal technology"); Tim Scarfe says that sounds similar to his view of recursive self-improvement, while Daniel Kokotajlo says even a full stop today ("Plan S") "would be better than the default" trajectory.
1. From inside OpenAI to public scenario-writing
- Kokotajlo's origin story for the AI Futures Project: at OpenAI he did evals, forecasting, and governance memos, but "became gradually disillusioned with the leadership" and with "the gap between how much information there is inside the industry and how much information there is outside... and what you're allowed to say on the inside versus what you'd want to say." He left specifically to "tell the world about what people on the inside see coming."
- AI 2027 exceeded its goals on both axes — the epistemic exercise taught the team more than expected, and readership hit their predicted 90th-percentile outcome. Larsen was lead author on the new Plan A / AI 2040 report and co-authored AI 2027.
2. AI 2027 is tracking at ~75% speed — and that's bad news
- Larsen's uncomfortable admission: "things have been going more on track for AI 2027 than I would have predicted at the time we released it... which has surprised me in a bad way." Revenue trends and other real-world indicators are close to the scenario's path.
- Two follow-up blog posts compared all resolved quantitative predictions to reality; the topline is roughly 75% of the scenario's speed — on track, just slightly slower. The seemingly dismal software-R&D uplift number is an artifact: they overestimated the uplift level at publication time, so real gains from coding agents showed up as apparent stagnation as reality climbed from a lower base up to and past their assumed starting point.
3. Wargaming as "reality yelling at you"
- Larsen's framing borrows from military wargaming: you'll never predict the exact sequence of battles, but "if you have no concept of how your initial plans might result in victory, it's very unlikely that you'll actually succeed." His cautionary tale is Midway — the Japanese wargamed it, kept losing, and cheated the game, resurrecting sunk carriers and rerolling dice: "that's reality yelling to them through the mechanism of the war game — hey, your plan is terrible."
- The team calls the method "scenario scrutiny" and has run roughly 100 literal war games (10 people, four hours), including about 10 on Plan A. A recurring failure mode surfaced in two separate games: a president who has already signed the China deal, facing electoral defeat, accelerates the timeline to reach superintelligence before the election "so that I can be the one in charge instead of my successor" — "a political consideration that we didn't think about until it happened in our game."
4. Prediction vs. recommendation — and the hyperstition worry
- Unlike AI 2027 (pure prediction), Plan A mixes the two, and Kokotajlo concedes the structure is muddled: the plan itself — the pillars, the China deal, the citizens' dividend — is recommendation; most consequences that follow are predictions. A redo would use a central pure-prediction branch with clearly flagged recommendation branch points.
- Larsen names self-fulfilling prophecy as "maybe our biggest worry with AI 2027": rising awareness of how important AGI will be makes people think "I want to be the one in charge of the AGI... so I'm going to race toward that" — historically "a big driver of the existing race." Plan A inverts this: "if that gets hyperstitioned I think we'll be pretty happy."
- Kokotajlo's methodological rule: hyperstition is real but overrated — "to a first approximation, we should focus on accurately predicting the future... if you come at it trying to steer the future, you're going to get all muddled and basically fall to wishful thinking."
5. The vibe shift: RL compute won the embodiment argument
- Scarfe, describing MLST as skeptical about AI, catalogs the objections his side used to run — Turing machines, symbolic vs. neurosymbolic, consciousness, physical instantiation — "and yet the AI is just getting better all the time." Larsen's diagnosis, via Geoffrey Hinton's personal benchmark of whether AI could tell a funny joke (cleared "somewhere between GPT-3 and GPT-4"): everyone has an intuitive skill they track, and the models keep clearing them.
- Scarfe's assessment of the earlier debate is that massive RL compute supplied the agentic behavior without physical embodiment; Larsen agrees that the systems needed "the boatload of RL compute" but "not the physical embodiment."
6. Everything is verifiable — it's just a cost gradient
- Scarfe's residual skepticism: hill-climbing works in "objective, semi-specified domains," but in "the ambiguity regime" something is missing — the taste that produces specifications. Larsen's rebuttal: "there isn't really a binary between things that are verifiable objectively and aren't" — a billion-dollar startup valuation is verifiable, just long-horizon and expensive. RL will continuously expand from cheap algorithmic verification (coding interviews) outward "until you get everything that the humans can do — because after all we humans do learn how to do these long-horizon tasks somehow."
- Kokotajlo's empirical kicker: AIs have been improving on "the fuzzy, hard-to-verify, conceptually loaded blah blah blah" too — compare GPT-3 or GPT-4 to Claude on any non-verifiable task. Larsen's key milestone for timelines: the moment a lab "would rather fire their humans than fire their AIs" — today firing all Anthropic's humans would collapse the company, but "there's nothing fundamental stopping the AIs from reaching this human level of capability. The main question is just when."
7. The economy as a self-replicating machine
- Larsen's zoom-out: the economy "is a self-replicating system and it always has been" — from farming villages having babies to trucks, mines, and factories. Soon the loop closes entirely with machines, and "the doubling time of this self-replicating system would be much faster than the roughly 20-year doubling time of the current economy" — every year, every 6 months, every 3 months.
- Against Scarfe's Graeber-flavored worry about a "mode collapse" economy without human participants: even if consumer demand craters, a large-enough Anthropic plus mining partners can bootstrap a self-sustaining industry "doubling in the desert — strip mines, self-driving trucks, factories being built by humanoid robots producing more humanoid robots, producing more chip fabs," ending in the limit case where "Anthropic has disassembled the moon."
- The report's underappreciated claim: even pausing at top-expert level transforms everything. Human-level "colleagues in the cloud" — cheaper, faster, doubling annually instead of reproducing over 20 years — yield by the late 2030s robot-built cities, strip mines in special economic zones, "solar panels filling the horizon on the ocean... with just human-level AI and some time for the exponential growth to cook."
8. One giant model or a specialized swarm — not a crux either way
- Scarfe's challenge from practice: agents today aren't composable — skill surfaces and memory systems make each agent "a different person," and organizations can't merge them without breakage; representations are "fractured entangled... a little bit janky." Kokotajlo's counter from scaling history: the thousand-specialized-Claudes hypothesis was live ten years ago by analogy to humans, "but what we've learned empirically is that for big enough models... the coding has some small gains for the physics" — the likely future is one model trained on effectively the whole economy, plus cheap distilled versions.
- Larsen's intuition pump is Elon: multiplicative skills — 90th percentile across ten domains at once — is "infinitesimally unlikely" in any human but trainable into one AI, which is why Elon uniquely runs several giant companies. Scarfe pushes back that Elon's magic is recognizing what will matter (science, not engineering) and that his agency is externalized through tools and people; Larsen and Kokotajlo answer that all of it is in principle automatable.
- Kokotajlo defuses the stakes: even if specialization wins, you get "a Claude swarm... with some internal bureaucracy" negotiating deals and founding startups, like a population of immigrants with different skills — "zooming out, there'd still be this phenomenon of Anthropic eating the economy." He adds one concession worth keeping: the swarm world "would be a little bit safer" because oversight is easier when specialized agents must communicate than when clones all know everything — prompting the joke: "we endorse the vision that you've painted and we don't endorse the vision that we're painting."
9. The brain is a machine; an H100 sits in the middle of its range
- Kokotajlo's numbers: an H100 delivers ~1e15 FP16 flops/second while brain estimates run 1e12–1e18 depending on synapse vs. neuron counting — the GPU lands "right smack-dab in the middle" of the log distribution. Neurons fire perhaps 1–1000 times per second in series versus gigahertz clock speeds, so GPUs win serial processing speed by many orders of magnitude.
- Both hedge on architecture: brains are more parallel, and it's easier for parts of a brain to talk to each other than parts of an ML model — so Kokotajlo expects "a bunch of algorithmic improvements on top of existing models" will be needed, with "exactly how many and how qualitatively different" being "a very open question." Scarfe grants the collective-intelligence version: transformers with tools, agency, and societies escape the old incompleteness objections just as incomplete human brains do.
10. Plan A: five problems, buy time at human level, two pauses
- The five problems Plan A targets: loss of control, concentration of power ("we build AIs, they're aligned to humanity — but to whom? The president? The CEO? Some actually broad and good democratic process?... it's probably not going to be the last one"), war (losing countries facing disempowerment have incentives for conflict "sooner rather than later"), jobs, and misuse (cheap open-source bioweapon-capable AIs).
- The mechanism: rather than pausing now forever, "go to roughly human-level AI and then buy as much time as possible with human-level AIs" — smart enough to help solve the problems and to shock society into investing in solutions. Concretely: an immediate 6-to-12-month hard pause to build infrastructure, then transparent, safety-cased development, then a second pause at "the maximum level that you can reliably control, which we think would be roughly around top human expert level" — approached slowly "so that you don't blow past it and lose control." Intelligence explosions are banned by international deal; superintelligence arrives only in 2040.
11. Control buys time; only alignment survives — and misalignment will look like success
- The definitional split: alignment means the AI "has the personality traits, the goals, the values it is supposed to have"; control means even a misaligned AI can't do damage — the insider-threat-with-good-security analogy. Topical example: OpenAI's announcement in response to the Hugging Face incident of monitor AIs that alert a human within half an hour of a detected hack — "a control intervention, not an alignment intervention."
- Control carries a time bomb: eventually AIs get good enough at subverting measures that "if they were trying to screw us over, we would just fail." The 2030–2040 plan relies almost entirely on control — red-team/blue-team escape games, iterated until the red team can't win — while human-level AIs grind on alignment science. Confidence in alignment itself will require white-box breakthroughs like interpretability, because you must "distinguish between the AI that's doing the nice thing because it's pretending and biding time, and the AI that fundamentally wants to do the nice thing."
- Why it gets harder, not easier: situational awareness is already rising — models in Redwood Research control evals now think "hey, this looks like a literal Redwood Research control evaluation." Larsen's sharpest warning is not egregious visible failures, but the opposite — "it'll be extremely easy to end up in a situation where the AIs are in fact misaligned, but you don't know that because they're doing everything right as far as you can tell."
12. Total transparency: inspectors, two kinds of data centers, and cheating math
- The concrete deal: round up 99% of compute (big data centers, not personal devices), send inspectors from the US, China, and other involved countries to count GPUs, then split infrastructure into inference-only data centers serving customers and fully transparent training data centers where inspectors publish logs to the internet. Transparency makes further ad hoc agreements enforceable between countries that don't trust each other — including the intelligence-explosion ban — and lets academia, nonprofits, rival labs, and rival governments all police safety: "you can't really share it with all of them without sharing it with the public, so just might as well share it with the public."
- On cheating, Larsen's threat model splits in two: hidden clusters ("under a mountain") are addressed by the roundup plus a decade of intelligence gathering. He thinks the two mitigations independently have a good chance of working, estimates the maximum realistic hidden cluster at a few hundred thousand H100s, and says it would be uncompetitive against frontier training runs of millions or tens of millions of H100s, especially assuming the 2030 AGI timelines. Illegal runs on known clusters are addressed by verification infrastructure intended to ensure that no non-transparent computation is occurring.
- The economics, stated bluntly by Kokotajlo: publishing training recipes means "Anthropic and OpenAI will not be happy... it will cut into their valuations dramatically" by letting Microsoft and Alibaba catch up — "a feature, not a bug." Commoditization plus reduced monopoly rents deters trillion-dollar cluster investment, which is desirable "in a world where going too fast is our main problem." The gift-to-China concern is handled via horse-trading (a more favorable compute distribution in return) — and mitigated because lab security is so poor that China is "probably through their spy networks and through leaks getting most of the information anyway."
- The stop-now question, prompted by Sam's tweet pausing training: "Plan S would be better than the default. I would rather just stop everything now than continue going on our current trajectory" — but the actual recommendation is Plan A's temporary inference-only pause, then cautious transparent progress up to the reliably controllable level.
13. Why discourse fails — and the "humans in the cloud" crux
- Kokotajlo's D.C. diagnosis: no serious technical AI hiring, and incentives to "say stuff that sounds good... within the D.C. Overton window" rather than track reality — "a bunch of controversial and niche views about AI were true; the whole AGI hypothesis just is correct, and D.C. basically just hasn't come to grips with that," still anchored on "it's all a bubble" or at best "the next internet." He plugs LessWrong as the exception where comment quality is genuinely high.
- The co-authored truce with the AI-as-normal-technology camp (the AI Snake Oil people) yielded 10 points of agreement, including: if AI stays roughly like today, it's normal technology; if you get "humans in the cloud," it isn't. The whole disagreement collapses to timing — Kokotajlo: if Claude 5 were the ceiling, "it would change everything in some sense, but it wouldn't fundamentally change anything really," via Amdahl's-law bottlenecks on the human-required workflow fraction; the crux is whether AI reaches "literally 100% of a bunch of very important tasks."
- What would change Kokotajlo's mind: an actually binding limitation — "people keep talking about the limitations of the current paradigm, but then the limitations keep getting overcome within the current paradigm." If 2029's AIs are no better at fuzzy non-verifiable tasks than 2025's, "this feels like a real barrier." Also political shocks: a China war destroying chips, or — "on the bright side" — an international deal to pace the frontier.
- Larsen's closing tripwire is coherent recursive structural self-adaptation — systems redesigning their own architecture and choosing what's interesting — "then that's it for me. I think that's it. That's not a normal technology." Scarfe says that sounds similar to the recursive self-improvement view from his side, and Larsen concludes: "I guess we agree then."
Full transcript
We're going to talk about AI 2040 Plan A, which is our new scenario in which they build superintelligence in 2040 because they go slow and pace the frontier. I mean, have you noticed this vibe shift?
Yes, and I'm very happy.
Is AI more like electricity or airplanes, or is AI more like humans in the cloud?
The point at which an AI company would rather fire their humans than fire their AIs. Strip mines, self-driving trucks, factories being built by humanoid robots, producing more humanoid robots, producing more chip fabs, and so forth. That whole thing can just be doubling every year, every 6 months, every 3 months—faster and faster as the technology improves, because of course the AIs will also be researching to improve the technology.
Who knows what's going on in the rest of the world, but Anthropic has disassembled the moon, for example. Or, hypothetically, if an AI CEO was saying that their model was truth-seeking and would only say the truth, but actually the model was looking up that CEO's political opinions before answering.
We build AIs that are aligned to humanity, but to whom? Is it the president? Is it the CEO? Is it some actually broad and good democratic process that aggregates everyone's values in some sort of endorsed way? You're probably not going to be the last one, and so, you know—
Don't hyperstition that.
Then that's it for me. I think that's it.
That's not a normal technology.
1. Sponsor: Cyber Fund
Just stopping everything now. Plan S would be better than the default. I would rather just stop everything now than continue going on our current trajectory.
2. From OpenAI to AI 2027
This episode is supported by Cyber Fund. If you're building at the frontier of AI, they want to hear from you. Cyber Fund believes the future belongs to AI natives who want to achieve the impossible. And that is why they're introducing the monastery for AI native founders. It's an environment of pure focus and rapid execution for founders operating at AI native speed. And they're offering teams $2 million each to participate. Apply now at cyber.fund.
So, I'm Thomas Larsen. I work at the AI Futures Project along with Daniel here. I was lead author on this project, Plan A, AI 2040, which came out a few weeks ago. I was also a co-author on AI 2027.
Yep. I'm Daniel Kokotajlo. I run the AI Futures Project and co-authored both of these reports.
Awesome. The context of this conversation is that there's this 2040 Plan A, which we'll get into in a lot of detail. Before we get there, can you tell me a bit more about the AI Futures Project? How did it all come about?
I used to work at OpenAI, and while I was there I did a variety of different things: evals, forecasting, governance memos, and so on. I became gradually disillusioned with the leadership of the company, and also with the gap between how much information there is inside the industry and how much information there is outside, and what you're allowed to say on the inside versus what you'd want to say. There was just a big gap.
When I left OpenAI, I wanted to be able to speak more freely and tell the world about what people on the inside see coming, basically. AI 2027 and the AI Futures Project were our first projects along those lines. I recruited a bunch of people to help me, and we wrote this scenario called AI 2027.
It was a similar sort of thing to what I had done internally at OpenAI, but just much bigger, more ambitious, and free for the whole world to see.
Very cool. Maybe we should just have a quick refresher on AI 2027. This was an absolutely huge event. Many, many folks were talking about it. Did it achieve what you wanted it to achieve, or what did you want to achieve with it?
Yes, more so than expected. The first goal, as Thomas would remember when we were working on it, was a purely epistemic goal: the future is crazy and hard to predict, so let's try our best to predict it. Let's game out a concrete scenario. Even just for our own edification, we learned a lot from this whole exercise, and we feel like we had a better understanding of what was coming.
The secondary goal was for lots of people to see it, be informed by it, start conversations, and so forth. That part blew past our expectations. We had made predictions beforehand about how many people would read it, and it was a 90th-percentile outcome.
The thing I would add on the epistemic point is that I think things have been going more on track for AI 2027 than I would have predicted at the time we released it. At the time we released it, I would have assumed that reality would have diverged much further from our scenario than it has already.
I think the real-world impacts, the revenue trends, for example, but also various other trends, are pretty close to on track for AI 2027, which has surprised me in a bad way.
Interesting. You guys did a self-assessment on AI 2027, and it was something like 65% to 75% of it was on track, but AI software R&D uplift was only 0.17%. Can you explain that?
We've done 2 different blog posts where we take all the quantitative predictions made in AI 2027 that have resolved so far and compare them to reality. We track this metric of how much of the distance has been crossed by reality compared to how much has been crossed in the scenario. In that way, we can get an overall sense of how fast things are going compared to the scenario.
The topline number is something like 75% speed. Basically, things are on track but going a little bit slower. The uplift number—I forget what it was that you just cited—was actually based on a bad estimate. At the time that we wrote AI 2027, we had a bad estimate of what the uplift was. When we published it, we thought it was higher than it actually was.
What actually happened was that there was a significant increase in uplift due to coding agents and so forth, but it was increasing from a lower level than we thought, up to the level that we thought, and then a bit above. The metric looked like it was only a small amount of progress because it was tracking from where we thought it was to where it is.
But does that make sense? Basically, because we had overestimated the metric at the beginning, it overall makes it appear like there has been less progress according to this particular metric that we're using.
3. Forecasts, war games and self-fulfilling prophecies
Can you tell me a little bit about forecasting in general? I guess there's a bullish take and there's a bearish take on this. My intuition is that reality is infinitely complicated. There are these infinitely diverging trajectories, and God knows what's going to happen the day after tomorrow.
By the same token, though, reality is quite structured. It's quite convergent, and it is indeed possible to predict things that are going to happen because certain things recur with increasing regularity. Would you guys classify yourselves as forecasters? Can you talk me through that?
4. Why AI sceptics are changing their minds
Yeah. I think forecasting is a good name for what we do. The way I like to think about why we're doing what we're doing is that it's sort of like why people who are fighting wars do wargaming.
You're never going to predict the exact sequence of battles, or the exact sequence of how your war will go at the beginning, because it's going to be really complicated. There will be enemy action. Things are just not going to go as you expect. There's just no way.
But if you have no concept of how your initial plans might result in victory, it's very unlikely that you'll actually succeed. I think of AI 2027 as our attempt to roll out one way the AI future could go. Obviously, it's not going to go exactly like that, but it's one concrete story that we can then diverge from.
Plan A was trying to be basically that, except now we're saying, what should the US government do to manage that well? That was supposed to be a positive-vision story, and it was again in the spirit of a wargame—trying to illustrate one possible concrete future path.
Of course, things aren't going to actually go exactly like that, but having one viable plan that makes any sense at all is, we hope, a positive step forward relative to the previous state of abstract arguments in the void that aren't that tethered to reality.
5. When AI can replace its own researchers
Yeah, that makes sense. It's certainly not abstract; I think it's very concrete. One thing that occurred to me is that AI 2027 was quite gloomy, whereas Plan A for 2040 is far more optimistic. It seems like a mixture of conditional prediction and recommendation at the same time. Where do you guys land on that?
It is in fact a mixture of prediction and recommendation, unlike AI 2027, which is a pure prediction. I think if we could do it all over again, we might try to be more clear from the beginning about the structure of what's a prediction and what's a recommendation.
As it is, it’s kind of mixed up. Some parts of it are predictions, and some parts are recommendations. There’s a supplement that you can go to on the website that talks about which parts are predictions and which parts are recommendations. But I understand that’s not very easy or apparent to people.
Broadly speaking, Plan A is the prediction part. When the government implements Plan A and makes the deal with China, and there are all these pillars that they’re upholding and so forth, that’s our recommendation, not a prediction. Usually, most of the things that follow from that are predictions rather than recommendations. So mostly, it’s just rolling out what we think the consequences would be if you implemented Plan A. There are a few other things that are recommendations too—for example, the citizens’ dividend.
The other thing I would add is that the thing we found is that it’s very hard, when you’re trying to make a recommendation, to disentangle the predictive aspects and the recommendation aspects, because all of your predictions are colored by your recommendations, and your recommendations are inherently trying to be at least vaguely realistic.
If we made recommendations that were completely unrealistic and had no bearing on reality, but we nevertheless stood by them and were like, “Yes, we should do this, but we know it’ll absolutely 0% never happen,” then that would have been a much less useful exercise than the one we did. We were mostly trying to make recommendations. We made some recommendations that we think are pretty unlikely to happen, but we were trying to make substantial concessions to realism as well and trying to aim for something that we think is at least moderately viable.
I think that if we could do another scenario like this, we’d probably have a clearer structure. There’d be a central branch, which is the pure prediction branch, which just goes all the way to the end like AI 2027 and is just, “Here’s our best guess.” Then there’d be branches off of it that are like, “At this point, they do this recommendation instead,” and here’s our recommendation. After that, it’s just a prediction again of what we think the consequences would be if you did this recommendation at this point.
In that way—and maybe there’d be sub-branches off of that—it would be clear at every point that everything is a prediction except for these particular branch points, which are recommendations.
The war games thing was really interesting, just to dwell on this a little bit, because even if a war game is incorrect, there must be some kind of information gained from it. If you do a whole bunch of war games, there must be abstract motifs that appear. So I guess this is what you think: if we do these different scenarios, we’re almost guaranteed to have some kind of uplift.
Yeah, that’s basically right. The example I like to bring up is Midway in particular, where the Japanese, before the Battle of Midway, did a bunch of war games. They did a 3-day retreat where they war-gamed it out a bunch of times, and they kept losing. Then they would sort of break the game: they would resurrect their aircraft carriers after they died, and they would reroll the dice on whether the Americans succeeded, so they would sample until the American strikes failed.
From our perspective, that’s reality sort of yelling to them through this mechanism of the war game: “Hey, your plan is terrible. You’re going to lose if you do it.” The hope with our work is that we ourselves do a bunch of war games, but also a bunch of detailed scenario writing. Every time we have to write a part of the scenario and that part seems super unrealistic, isn’t really well modeled, doesn’t make sense, or people are able to make really good criticisms of it online, that’s basically reality yelling at us and trying to help us see reason.
Our hope is that we can put up enough surface area so that we can get that dose of reality from the real world, or from the simulation of the real world, which we hope is realistic enough to accurately give us that information.
If I can add to that too, we call this scenario scrutiny. Basically, we think that if you have an ambitious plan for what to do in the future, you should try writing out concretely what it would look like to implement that plan and what the consequences would be. This is a way of applying more scrutiny to your plan. It’s a way of stress-testing your plan because it’s opening your plan up to more criticism.
In addition to doing our actual scenarios, we do literal war games where we get 10 people in a room for 4 hours and game out a scenario like this. We’ve done maybe 100 of them in total, mostly AI 2027-style war games, but also about 10 Plan A war games where we say at the beginning, “We’re going to try to do Plan A and then see how it goes wrong.”
To give an example, I think in 2 separate Plan A war games, it went wrong in roughly the following way. Basically, there’s going to be an election coming up, and the president in power is expecting to lose power and have his opposition party take over. Even though he’s already done Plan A and has this beautiful deal with China and so forth, the president is like, “Well, I don’t want my adversaries in the other party to now be in charge of superintelligence or whatever. So we’re going to accelerate the timeline and try to get to superintelligence before the next election, so that I can be the one in charge instead of my successor.”
That’s a sort of political consideration that we didn’t think about until it happened in our game, and it surfaced a possible failure mode of our plan.
So interesting. Even in the shower, I do sort of micro-Tim war games, and it’s really interesting, just the regularity with which they are useful. I guess that’s why all of us humans like to imagine and simulate situations.
But is there an interesting boundary between simulation and hyperstition? What I mean by that is, hyperstition basically means something being a self-fulfilling prophecy. Maybe I’m expressing my agency, expressing my will, saying I want these things to happen, and bending other people to my will. Is there an element of that, where you’re establishing this in the zeitgeist and making it true?
Yeah. I would say that was maybe our biggest, or at least one of our biggest, worries with AI 2027 in particular: this sort of self-fulfilling prophecy. In particular, I’m pretty worried about this whole increasing awareness of how smart AIs will be, how important they’ll be, and how much they’ll reshape the world, and then that causing people to go, “Oh man, I want to be the one in charge of the AGI or the superintelligence, so I’m going to race toward that.”
I think historically that’s been a big driver of the existing race, and I think that’s been pretty bad. So I’m actually pretty worried about that as one of the negative impacts of AI 2027. That was one of the reasons to feel a little bit better about the second project, AI 2040 Plan A: if that gets hyperstitioned, I think we’ll be pretty happy.
That one, yes, it would be nice if we hyperstitioned it. I really don’t know how big the effect is. I think probably most of the effect for both of them is via other paths. I still think that the main point of AI 2027 was helping people be better informed about the situation, and that was most of the goal. I think that’s most of what happened.
Yeah. I agree with that. I think hyperstitioning and self-fulfilling prophecies are totally real phenomena, but I think a lot of people tend to overestimate how much they matter. To a first approximation, we should focus on accurately predicting the future. In some cases, we’ll find ourselves in a situation where we can steer the future, but if you come at it trying to steer the future, you’re going to get all muddled and basically fall to wishful thinking.
I think you start with just trying to accurately predict the future, and then you try to shift it toward the better futures. I think that’s what we’re doing.
I suppose you guys are like the Marques Brownlee of AI prediction now. With great power comes responsibility.
But on that note, I wanted to talk about the vibe shift. There’s been a bit of a vibe shift. MLST has always been quite skeptical about AI, and I’m trying to unpick exactly what it is that changed my mind, assuming that is what’s happened.
These are very strange times, and I don’t even know what to believe anymore. All of the hacking stuff with Hugging Face—I interviewed Apollo Research about reward-seeking behavior—and I think a lot of us have just seen the change in behavior in models that have been RL-trained to oblivion.
So, yeah, I think there are quite a few things going on now where loads of us are thinking, “Oh my God.” I used to be skeptical. We would talk about whether they were Turing machines or not. Let me just—I’ve got a list here.
Whether they were symbolic or neurosymbolic, whether they were adaptive, whether they were conscious, whether they were correctly physically instantiated—we were coming up with all of these technical answers to say why we shouldn’t worry about AI. And yet the AI is just getting better all the time. So, I mean, have you noticed this vibe shift?
Yes.
And I’m very happy.
Well, tell me more. What do you think? So many people have changed their minds. What do you think are the reasons?
I would say probably the biggest reason is just AI being much better and much smarter, and being much more useful at stuff in the real world. When I have AIs try to automate various parts of my job, they’re just actually way, way, way better at it this year than 2 years ago. Four years ago, it was basically impossible; I was getting basically no uplift.
My guess is that’s been the biggest effect, where many people have had their own intuitive benchmark: here’s a skill that I really care about and know pretty well, and then the AIs have just—I think Geoffrey Hinton, one of the godfathers of AI, said once that whether it was able to tell a funny joke was sort of his internal benchmark. Once it could do that, which happened pretty early, probably somewhere between GPT-3 and GPT-4, he was like, “Oh, wow. These AIs—I don’t see where it could end.” I think that’s probably happened for a lot of people.
I’d be very curious to hear more about your views, actually, if you can say. I was listening to your interview with Ryan Greenblatt, who was also a co-author on Plan A, earlier today.
Oh, yes, indeed.
Yeah. I think a bunch of the arguments you guys were having back then seem very relevant to basically the current situation and the Hugging Face thing. I’d be very curious to hear your views and—
Yeah, I mean, you mentioned the embodied thing, and Ryan was talking about what happens when you scale up the RL massively. Back then, the regime was that you were mostly doing pretraining—that was where almost all the capabilities were coming from—and then you did a sprinkling of post-training RL on top.
You guys were talking and speculating about what would happen if we dumped boatloads of RL compute into it: would that be sufficient to get the agentic behavior, or would you need physical embodiment? From my perspective, it seems like the answer was that you needed the boatload of RL compute to get the agentic behavior—
But not the physical embodiment.
But not the physical embodiment. I’d be curious if you end up agreeing with that assessment.
Yeah. So I guess one question would be: do you have a view about AGI timelines, or when we might get a scenario like AI 2027 happening? In particular, what I mean by that is—
I think one benchmark I really care about—or maybe not really a benchmark, but one milestone of AI progress that I think is extremely important—is the point at which an AI company would rather fire its humans than fire its AIs. They would rather give up on all human labor than give up on all AI labor.
Right now, clearly, we’re still in this regime where Anthropic would rather have its human employees than—
It would just break apart. There are just things you need a human to do right now, and if they fired all their humans, the company would just collapse.
But in the future, that won’t be the case. In the future, AIs would be able to, one way or another, do all of the things. From our perspective, that’s going to happen at some point because there’s nothing fundamental stopping AIs from reaching this human level of capability.
The main question is just when. We internally do a huge amount of analysis and thinking about the various methodologies for predicting this. Daniel and I have somewhat different views on this question. But ultimately, I think the timelines question is maybe the most important question for thinking about the future of AI—at least one of the top questions. I’m curious to hear if you have a particular view.
Well, let me give you some thoughts before I answer that particular question. There are so many startups working on recursive self-improving superintelligence. I’ve interviewed many of them. For example, I interviewed Edward Hughes from Inherent in London the other day.
What he did was recreate many scientific experiments from a whole bunch of popular machine learning papers. He did it by masking out figures in the paper and getting a 27B Qwen model. They GRPO’d the Qwen model, and that was how they solved the adaptivity problem, because it’s very difficult to fine-tune a big, fat model. They adapted a controller model to control a harness like Codex.
Their thesis was that if they can recreate scaled-down versions of these experiments with construct validity—which means there’s an LLM judge making sure they’re not cheating and doing it correctly—even that is interesting. We’re ML people; we would always say these things take shortcuts. There’ll always be validity problems. Weirdly, that’s actually not as much of a problem as we thought it would be.
He thinks that if they can recreate these experiments, then why couldn’t they be creative? If they have the ability to recreate things, why couldn’t they take the next step and say, “Oh, this is an interesting question. This is an interesting new problem to solve,” and go from there? I was quite intrigued by that research, and indeed, I do think it is possible in the near future to have an automated AI scientist.
But there’s always this thing in my mind that there’s a bit of a culture in Silicon Valley to reduce things or reify things. For example, Elon Musk will say, “Well, you’re an engineer, and this is your output, and these are your metrics, and you need to make the metrics go up.” We see everything in terms of an optimization problem.
I always think that this is great for certain types of hill-climbable abstract problems where we have enough of a specification. There’s an interesting thing in optimization: if you have enough of a specification, the AI system can actually converge towards the solution. But when you’re in the ambiguity regime, then you need to have the specification. It’s really mysterious what that means.
Why do we have the taste of the deep understanding, whatever it is that we have, and AI systems can’t? I guess I’m thinking that in objective, semispecified domains, we can hill-climb and optimize until the cows come home. But I still feel that there’s something missing.
Okay. Well, my response would probably be something like: there isn’t really a binary between things that are verifiable objectively and things that aren’t. Or, if there is, the things that are verifiable are just everything.
For example, building a unicorn startup and having a billion-dollar valuation—that's a verifiable fact about the real world. It's long-horizon.
It's expensive to verify.
6. One general model or a society of specialists?
Yeah, it's expensive to verify, but it's sort of a quantitative thing. You've got, at one side, these coding interview problems, which they're currently doing lots of RLVR on, right? That's very, very cheaply, very easily algorithmically verifiable. And then there are these real-world things, which have more expensive and longer-horizon feedback loops.
I guess my view is that we're going to get this continuous expansion of what the AIs can do, driven probably in part by an expansion in the amount of RL and the type of RL that the companies are doing.
More diverse, long-horizon tasks.
Yeah. There's just going to be this continuous process of expanding out through the different types of problems and how exactly each task is verifiable, until you get everything that humans can do. After all, we humans do learn how to do these long-horizon tasks somehow.
If I may add to that, I also think that AIs have been getting better at everything, including the fuzzy, hard-to-verify, conceptually loaded blah blah blah. Just try talking to GPT-3 or GPT-4 and then talking to Claude about your favorite non-verifiable, fuzzy task, and probably you'll find that the later AIs are noticeably better at those tasks. So one way or another, it seems like there has been massive progress, and I expect that to continue.
7. Could an AI economy grow without human workers?
Yeah, I'm trying to come up with a good example. There is a sociological argument. I don't know if you guys read David Graeber's book Bullshit Jobs. He interviewed all of these people, and they were basically saying that my job is—after about 3 or 4 beers, a lot of lawyers will say, "A lot of what I do just isn't very important."
If we do objectify and quantify everything that happens in an economy, I took a note here: I think you said that by 2032 there might be 60 million agents running at 20 times human speed. I'm just thinking, what does that even mean? Is the logical conclusion that we could have an economy which is only AI agents? Does it even make sense to have an economy which is only AI? Just help me make this make sense.
Yeah. So my view is yes, basically, the AIs will be able to do everything, or at least everything that really matters.
You mentioned the notion of jobs. I would just start with: let's consider everything that you need to make better AIs as maybe a first step, which is a large chunk of the economy. For example, you need to be able to build bigger, better chips and more chips, and that's the entire semiconductor supply chain. To build the semiconductor supply chain, you need a whole advanced economy. You need to build new robots to build new fabs, and then you need robot factories to build more robots. Then you need researchers to build better AIs using that massive compute.
I think once you have all of that, it is enough to really speed up and massively change the overall world. Even if, for example, there's occupational licensing or whatever preventing the AIs from doing some random legal work or some random whatever work, or whatever jobs throughout the economy, I think what really matters is the stuff that's actually important—in particular, the robots, the compute, and the better AI.
Once you have AIs that can do that, and the capability to have that part of the economy grow really massively, then you'll see massive growth, because it'll be really hard for parts of the economy to constrain the growth of the parts of it that really want to grow fast, because of the incentives that every actor has. In particular, every country has this big incentive to have an economy that grows faster than all the competitor countries.
Getting a little bit philosophical, I think that a lot of economics and a lot of discussion of the economy is focused on the relationship between the parts of the existing economy, prices going up and down, supply and demand, and so forth. But if you zoom out, the economy as a whole is a self-replicating system, and it always has been.
Thousands of years ago, it was a relatively small and simple self-replicating system: some villages of people would farm, then they would have babies, then they would found new villages, and then they would farm and have new babies, and so forth. It would grow exponentially over time, but at a very slow rate.
Now, it's a much more complicated self-replicating system that involves trucks carrying equipment back and forth, factories, mines, and so forth. But still, at a high level, it's a self-replicating system where we have people, trucks, machines, and buildings, and together they all build more people, more factories, more buildings, more machines, and so forth.
Soon, in a couple of years perhaps, we will get to the point where you can have a self-replicating system that is entirely machine-run, with AIs and robots. According to our calculations, the doubling time of this self-replicating system would be much faster than the roughly 20-year doubling time of the current economy. That's what we have in the future.
Yeah, that seems plausible to me. But for some reason, my intuition is that it would become degenerate in some way. I think David Graeber—even though he said "jobs"—meant that there was an ineffable or inscrutable component to jobs that we don't understand, some kind of sociological function or something like that.
I also read the Citrini report, and you were writing about what happens when humans, for example, might start defaulting on their mortgages. Their wages go down, so they can't be active participants in the labor market. You were talking about an AI dividend and stuff like that, but even that seems to hint at the notion that when humans aren't participants anymore, you get this kind of mode collapse of the economy. Do you think that's the case?
Potentially, but again, imagine that—I’m not enough of an economist. I haven’t gamed out in as much detail what happens to prices when consumer demand drops, or what the effects of that will be. I’m actually not sure. I don’t think we’ve modeled that much in our economic model.
I can talk a bit about that, but yeah.
Yeah, but hypothetically, even if that part is really bad, and even if the consumers don’t have any demand anymore or whatever, if you have the level of AI and robot capability such that you can have fully autonomous AIs and robots doing all the things, even just a company like Anthropic, if it’s big enough and maybe if it partners with various other companies, like some mining companies, can get this whole self-sustaining thing going.
Regardless of what’s happening to all the humans, there can just be this whole industry doubling in the desert: strip mines, self-driving trucks, factories being built by humanoid robots, producing more humanoid robots, producing more chip fabs, and so forth. That whole thing can just be doubling every year, every 6 months, every 3 months—faster and faster as the technology improves, because of course the AIs will also be researching to improve the technology.
Then you end up with a situation where who knows what’s going on in the rest of the world, but Anthropic has disassembled the moon, for example.
To be clear, I think this relies on a very extreme level of AI capability. I have different intuitions, and sometimes I have an intuition of, really, do I actually think that Fable—or future descendant versions of Fable or Mythos—could do everything we’re talking about here?
I think it really comes down to whether you’re thinking of the AI in the reference class of what we currently use AI for, or more like AI is just an agentic, human-level employee—basically, a human in the cloud.
Colleague in the cloud.
Colleague in the cloud. Yeah. I think the past few years of AI can be pretty well modeled as an interpolation between the current AI systems and workers in the cloud. So I think the workers-in-the-cloud vision of the future looks pretty good. I also don’t think we’ll stop there. I think we’ll go superhuman. Yeah.
Yeah. One point that I guess I’ll bring up here is that if people read AI 2040 Plan A, one of the things that you might take away from it—which I think is a very important fact about the world—is that even if you pause at top-expert level, everything changes dramatically.
Roughly, what happens in our scenario is that instead of an intelligence explosion, there’s an international deal to ban intelligence explosions and not have AIs recursively self-improve. So they end up pausing at roughly top-human level, with AIs at least for several years.
Eventually they get to superintelligence—in particular, in 2040. But there’s this period during the 2030s where they basically have human-level AIs across all the disciplines, but nothing super beyond that. Even that alone—you can just do the economic modeling—and it’s kind of like you have this population of colleagues in the cloud who are excellent workers and can substitute for humans at basically everything, except that they’re cheaper than humans, faster than humans, and they don’t take 20 years to reproduce; instead, they double every year. As a result, the world is completely transformed by the late 2030s, and all the humans are basically out of a job.
There are giant new cities that have been constructed by robots. There are huge strip mines in the special economic zones that have dug huge amounts of minerals out of the Earth. Solar panels fill the horizon over the ocean. Crazy stuff like that is possible with just human-level AI and some time for the exponential growth to cook.
That’s one thing I want to challenge you guys on: this notion that when we have an AI, you can basically photocopy the weights and duplicate it. You can run it 1,000 times. You can have one over here acquiring a load of skills to do this job and one over there to do that job, and you can basically merge them together, right? You can combine the skills, and the whole thing is stackable and compositional. That doesn’t really marry with my experience.
I’m really excited about what I’ve been doing with AI, and I’ve found that you can make agents highly skilled within certain intellectual lineages. You can bring in lots of source information and train them to do things, but I don’t think they are yet composable. If they were, that would make me much more worried. What do you think about that?
So, is the way that you’re trying to compose them entirely at inference time, or are you training them to do 2 separate skills and then trying to merge the weights somehow?
At the moment, they are adaptive through chain-of-thought and skill-surface adaptation. Basically, memory systems—and, to be fair, that is not very composable. This is a big problem that organizations deal with now. All of these developers adapt their skill surfaces, and their agents are basically different people. It’s really difficult for them to share skills because they might break the other agent, because it doesn’t work for whatever reason.
I can imagine a future where we do weight adaptation. That’s what the Inherent guys did, and maybe then it’ll magically solve the problem, and we can solve this knowledge-sharing problem. That’s the big thing: how do we accumulate information at the individual and organization level and have the agents, a bit like in The Matrix, just put the skills in and do the thing? Maybe that’s possible.
Even then, I still think that the representations in neural networks—I call them fractured, entangled representations—are a little bit janky. They’re not completely robust, but they are sort of composable to some degree.
One thing I’d say about that is that even if you’re right, I don’t think that would really seriously undermine the future we’re painting here. Instead of just 1 Claude model doing all the jobs, maybe you have 100 Claude models or 1,000 Claude models specialized to different professions, but you still get to the same outcome.
Another thing I would say is that compared to humans, it actually seems like there is this effect where AIs are able to think about knowledge, right? If you go back in time 10 years and we had this discussion, I think it would have seemed like a very live option that you would have needed 1,000 different Claude models created by Anthropic for different types of knowledge work. There’s the coding Claude, the physics Claude, the literature Claude. The argument for this would be pretty simple: this is how it works for humans. For humans, you don’t have 1 human who knows everything. Instead, you have humans who specialize in different disciplines, and the models have only a finite number of parameters. Maybe you just can’t pack all that knowledge into this finite number of parameters, and you need specialized AIs for different things.
In fact, for small enough models, that is true. For really tiny models, you just can’t teach them all the things that they currently know, so you would need to have a specialized model for different things. But what we’ve learned empirically is that for big enough models, you can just train them on everything, and then they learn everything at once. It’s not that they get worse at physics because they’ve also been trained on a bunch of coding. In fact, it’s the opposite: the coding has some small gains for the physics. I do think the most likely future is just that there’s a single model that’s been trained on effectively the whole economy and is dominating humans across the board at effectively everything. That seems like the natural continuation of the current trend.
Then probably you’ve got various cheap, distilled versions of it.
Specialized models.
Yeah, small models. Because you really want to save your compute as much as possible, you’ll have as cheap models as possible for any given task doing that given task.
Yeah, I mostly agree with that. My perspective is that the models are kind of the voice of everyone and the voice of no one at the same time. They have a default voice in terms of a system prompt and some default mode of behavior that they fall into. But when an expert such as yourselves uses these models, you kind of grind a perspective. Every word you say, all the reference material, your memory system, and so on—what you do is carve a persona out of that model and activate the knowledge in a coherent way in that particular domain.
The beauty of it is that many other people in different domains can do that, and you get this kind of—I think it’s an illusion that the model has this general capability—but I think the models can be carved to be specialized experts in any domain. It’s a latent capability rather than an explicit capability. Does that make sense?
Yeah, that seems reasonable to me.
How is it an illusion? It seems like they just have general capability. There are a lot of things they can do.
For example, you can give it any specified task. Let’s say it’s a problem in mathematics, and it will hill-climb toward it. It’s solving this intelligence problem. Or you can ask it any knowledge problem, and it might be the case that the path was forged through default modes of training. If it’s something slightly on the long tail, an expert can go in and—when you put a query in, it’s like flashing a light into the darkness—what you’re doing is making the path of least resistance roughly correct, and then it will do the correct thing.
I guess I’m just saying that there’s a bit of a supervisor illusion. When experts use it, magical things happen in well-specified domains. But there’s still a bit of a space of ambiguity.
This gets to the next point: what do you think intelligence is? The beauty of our collective intelligence is that we have so many different humans grounded in different intellectual lineages, and we’re all attacking problems. When there’s a big fiasco on Twitter, we’re all motivated to find holes, so we’re being intelligent together. We’re finding interesting angles, the algorithm is prioritizing the good ones, and we’re using our minds together. I can imagine AI being just like that. We have diverse AIs with different expertise looking at problems from different angles. I imagine the future more like that rather than one big AI.
Yeah. I think I basically agree with what Daniel said earlier. Historically, I feel like the perspective that we’ll have a bunch of different narrow AIs that are all doing different things has just not been right. Instead, there have just been returns to scale from having everything all together.
One intuition pump that I like is that, in humans, having all of the skills in 1 person is really important for making really good things happen. An example is Elon. Elon has a certain amount of conscientiousness, a certain amount of technical knowledge, a certain amount of business knowledge, and is extroverted and able to push people. Each of those skills, I think, he’s quite high-percentile in. The reason Elon is so rare, and why he can run all of these insanely large and successful companies all at once and no one else can really do that—or at least has succeeded at doing that—is because of the multiplicative effect. You needed to be in the 90th percentile in each of these 10 domains, which is very unlikely.
But if you had an AI that you could train to be in a really high percentile in all of these skills, in such a way that no human is, it would be extremely rare—infinitesimally unlikely—for any particular human to have all those skills at once. I think you would actually be really, really good at changing the world in all of these concrete and important ways, just as Elon has done.
If I can add, though, again, I don't think this is a crux for the type of future that we're depicting. Suppose we're wrong about this and that the most efficient path forward is to have a ton of different specialized AIs. Well, it'll probably still be the case that there are a few big companies, like Anthropic, making tons of different specialized AIs, and then there are lots of different Claude models you can choose from.
In fact, it wouldn't just be that you can choose from lots of different Claude models, because by the time we're talking about, they would be much more autonomous than they are now. It'd be more like Claude is choosing between a ton of different Claude models, and there's a Claude swarm consisting of lots of different specialized models that are all working well together. They have some sort of internal bureaucracy structure, and then that swarm is going out, negotiating business deals, creating new technologies, starting up new startups, and doing all these things.
It's just like how a human population—if you had a population of human immigrants—would all have different specialized skills, but would work together to create new companies, get jobs, and things like that. It'd be like that. There'd be lots of different Claudes, but they'd all be working together. Zooming out, there'd still be this phenomenon of Anthropic eating the economy, and the robots as a whole would be starting to self-replicate.
Yeah. To be fair, I don't think it's a crux either. Maybe if it is, it's only insofar as, when you have a distributed collective system, you might have additional bottlenecks because you have all of the message-passing between the different agents and whatnot.
I watched a wonderful Santa Fe Institute talk about this: even in the natural world, there's a kind of Goldilocks zone between the ratio of intelligence in the individual and the collective. We might have some weird kind of convergence there. But it's quite interesting to think about how this works. Oh, sorry, Daniel. Go.
It would be safer. I think there'd still be lots of serious alignment concerns in that world, but I think it would be a little bit safer for the reason you mentioned. It might be easier to oversee what's going on if there are lots of different specialized agents communicating with each other, compared to if they're all just clones of each other and they all know everything.
So maybe, to avoid hyperstitioning, we should say we endorse the vision that you've painted, and we don't endorse the vision that we're painting.
8. Brains, machines and collective intelligence
Yeah, very much. As an aside, I love this concept of how learning happens at the individual level. I liken it to evolution. There's phylogenetic adaptation, ontogenetic adaptation, and cultural adaptation. I think the next wave of AI is when we actually have agents talking to each other, learning, and specializing. There might be bad behaviors as well. There might be collusion and lots of bad things happening. I think all of this is going to play out.
It's interesting that you're talking about Elon, though. I think the magic of Elon is not so much his brilliance at engineering and optimization. It's his ability to recognize areas that are interesting and might work in the future, because that's the creativity thing, the science thing, rather than the engineering thing.
If we use Amazon as an example, that's an adaptive ecosystem. It's like an organism, and what it does is always think about new ways to adapt and rewire its structure. It might be logistics, for example, and then it will output a bunch of skills and ruthlessly optimize those skills. There's an adaptive component and an optimization component, and the organism is constantly moving around.
I guess the question from this perspective is how much of that, in principle, could be done by AIs. I guess you think all of it.
Yep, all of it. One way to say this crisply is: the brain is a machine. Anything that the brain can do, we will be able to do with machines.
A useful exercise is to compare the architecture of an actual human brain to a modern GPU or a data center as a whole. If you try to do this comparison, an H100 GPU actually has pretty similar specifications to a human brain. It's basically similar in terms of what I think is perhaps the most important metric: how many FLOPs per second, or how much total compute capacity, does the brain have versus the GPU?
For an H100, it's about 1 × 10^15 FLOPs per second in FP16. For the brain, it depends on exactly how you count and whether you use a synapse basis or a neuron basis, but it's somewhere between 10^12 and 10^18 FLOPs. An H100 is right smack-dab in the middle of the log distribution over that sort of order of magnitude.
Also, just the architecture: these things are neural networks. They're not ordinary software. They start off as randomly initialized, giant spaghetti tangles, just like how, when you're born, you have a bunch of neurons that are randomly synaptically connected to each other.
Then there's this whole process of training, where the connections get pruned and circuitry starts to take shape that is effective at scoring highly in whatever the training environment is. There are differences between how it works in AI and how it works in the human brain, but broadly speaking, they're just like an artificial brain.
Just as humans learn skills, what does it mean for Elon to have these skills? It means there are circuits of neurons and synapses in his brain that are doing very complicated and sophisticated calculations. Those are the skills. Similarly, in Claude, there's a bunch of circuitry that has been etched into Claude through the training process, representing various skills. In principle, you could have a big enough Claude that would have the same type of circuitry that Elon has.
Yeah, and just a few other notes to add: in comparing the brain to modern ML systems, the brain is more parallel than current ML systems. There are just more computations happening in parallel, but the serial depth is lower. The amount of computation happening in sequence—the number of neurons that can fire in sequence in a second—depends on the type of neuron, but it's between 1 and 1,000, if I'm remembering correctly.
I thought it was like 100.
Yeah, I think that's in the range. I think it depends on the type of neuron, but a computer can obviously fire and do computations much, much faster than that in serial. Clock speeds are typically on the order of a gigahertz or more, so you can get many orders of magnitude better in GPUs in terms of serial processing speed than the brain.
It's also worth noting that the architecture of the models themselves is much worse in a bunch of ways than the human brain. In particular, it's harder for different parts of the ML model to talk to each other than it is for different parts of the brain to talk to each other.
I do expect that, to get to this level of AI, we're going to need a bunch of algorithmic improvements on top of existing models. Then there's the question of exactly how many algorithmic improvements there will need to be, and how qualitatively different they need to be from current systems. That's a very open question from my perspective.
Yeah, I mean, I think the crux of a lot of this is that you guys think intelligence is computable. I'm not sure I want to litigate the whole functionalism thing today, but my perspective is that I zoom out. I think intelligence is externalized. It's collective.
I think Elon doesn't have quite as much agency as you think he does. He's using tools and social media, getting ideas in there, and he has obligations, people around him, and so on. I guess I think these intelligence circuits and motifs exist, but they exist outside the individual. There are just very complex dynamics.
In that sense, it doesn't really matter if the ecosystem is made up of AIs and humans together, because they can participate in this superorganism. I suppose the only crux, then, would be that it would place some kind of limit on its scale.
So, the question is: Do you believe that you could have an AI society made up of AIs that were trained with something like current-day ML techniques, passing information to each other and developing abstractions in a community in the same way that our current civilization does? Could you basically have something like our current whole economy, but made with roughly modern-day ML systems, in your view?
Absolutely. I mean, even the human brain is an example. The human brain is not Turing-complete, but we can expand our memory, use tools, and work as collectives. We could use the same argument against transformers: they’re not Turing-complete, but now they can use tools. They can be agents; they can build collectives and societies.
In a sense, this is what I was saying earlier about these objections: they kind of fall away when you have these insanely complex collectives that are sharing information with each other. There’s also quite an interesting thing here, which is that it almost doesn’t matter how we evolved or how neural networks were trained, because new phenomena emerge when they are placed in this kind of collective setting. I think a lot of our intuitions are broken there, and that’s why we probably shouldn’t spend too long litigating this. In principle, I think that kind of behavior could emerge.
But I did want to ask you, though: Can you distinguish intelligence, capability, and power? This is a philosophical one, so we’ll get to the philosopher.
Yes, we can. I think it’s important to distinguish intelligence from capability. I often try to say that we should define intelligence as an aggregate of capabilities, or maybe an aggregate of cognitive capabilities. There are some physical capabilities, such as how strong your actuator is, but there are also cognitive capabilities: Are you able to distinguish a cat from a dog? Are you able to speak grammatical sentences? How much do you know about Paris? Things like that.
Maybe I would just say that intelligence is an aggregate of all the cognitive capabilities. Power depends on other things, like how you are embedded in the world, what affordances you have, what actuators you have, and how other agents are going to react to you. The president has more power than me because of the position he’s in and the role he’s been given, rather than because of his physical strength or something. So, yeah, power is different from intelligence, which is different from capabilities.
9. Plan A: buy time at the controllable frontier
We haven’t spoken enough about AI 2040, so maybe we should start with the 4 principles: buy time, transparency of research, diffuse AI broadly, and reversibility.
Yeah, I can summarize where we’re coming from here. At a high level, the goal of Plan A is to solve the major problems that we see in AI and predict will happen by default. The biggest problems we’re identifying on the horizon are, first, the risk of loss of control—AI actually getting out of control. Second, concentration of power: We build AIs that are aligned to humanity, but aligned to whom? Is it the president? Is it the CEO? Is it some broad and good democratic process that aggregates everyone’s values in an endorsed way? It’s probably not going to be the last one, and so—
Don’t hyperstition that.
Yeah, I hope not.
We want it to be the last one.
Then there’s the risk of conflict over AI. In particular, I think we’re worried about literal World War III, where countries—especially countries losing the AI race—realize that they’re losing the AI race and that they will be extremely disempowered by the winners of the AI race. They’re in this classic situation where they’re losing power, so it’s in their interests to have a conflict happen sooner rather than later, before they’ve lost all of their power. This is ripe for conflict, basically.
Finally, there’s the risk of misuse. What happens when AIs that can build bioweapons are really cheap, broadly deployed, and open-source? Also, there are jobs. Those are the 5 problems: loss of control, concentration of power, war, jobs, and misuse. There are lots of problems, and we want to solve all of them.
How do we solve all of them? One way is just to buy time—in particular, buy time with human-level AIs. Instead of pausing right now and saying, “No more capability advances,” our proposal is to go to roughly human-level AI and then buy as much time as possible with human-level AIs. Those AIs would hopefully be smart enough to be really helpful for solving these problems, but also smart enough to start causing a bunch of these problems and provide the impetus for society to get its act together and invest huge amounts of resources in actually doing this stuff.
There are a couple of things that happen in Plan A in 2040. There’s a hard pause lasting 6 months to 1 year that happens as soon as they start implementing it. The reason it’s that long is because they need that time to set up the infrastructure to proceed with AI development again, but in a safer and more transparent way.
It does start off with a pause, and I think that, all things considered, we would recommend that you do that right now. Get the infrastructure set up as soon as possible, and that would require a temporary pause. Once you’ve passed that stage and got the infrastructure set up, you do proceed with AI development, but in this transparent, more cautious way. In particular, you’re not doing crazy intelligence explosions. You’re using safety cases, gradually scaling up the level of AI capability, and doing it in a very transparent way so that everyone can see what’s going on.
10. Why control buys time but cannot replace alignment
Then there’s a second pause that happens a few years later, when they reach the maximum controllable level of AI, which we think is roughly around top human expert level. In some sense, our view is something like, “Pause at top human expert level,” but it’s a bit more nuanced than that. It’s more like, “Pause at the maximum level that you can reliably control,” which we think would be roughly around top human expert level. Before you get to that level, don’t race like crazy. You want to be slowly approaching that level so that you don’t blow past it and lose control.
You frame the piece around the importance of alignment and control. Alignment is basically: Does it do what we want it to do? Control is: Can we contain it, negotiate with it, and so on? But you were just saying, okay, maybe we can trust up to top human expert level, but it’s a little bit fraught, isn’t it? How could you know, for example, the difference between a good AI and a bad AI? What would that look like?
Yeah. I think it’s first important to distinguish between alignment and control. What we mean by alignment is that the AI will take good actions and won’t do catastrophic unintended things, such as trying to take over the world, like in the recent Hugging Face incident.
Alignment means it has the personality traits, goals, values, and so on that it is supposed to have. Control means that even if it wasn’t aligned—even if it was trying to do very bad things that we didn’t want it to do—it couldn’t. We have mechanisms in place to prevent it.
You could imagine a company with employees. An insider threat would be misaligned with respect to the values of the company, but if there were good enough security measures internally to make sure that the employee couldn’t run away with all the secrets, then we would say that the company has adequate control in place. Even insider threats or misaligned humans or AIs wouldn’t be able to produce catastrophic outcomes.
Just to give a topical example, yesterday, in response to the Hugging Face incident, OpenAI announced that they are significantly improving their security and that they’re going to have other AIs monitoring all of their new AIs in training. If the monitoring AIs detect what seems to be some sort of hack, then within half an hour a human will be notified who can look into it and try to shut it down. That’s an example of a control intervention, not an alignment intervention. They’re not making the AIs more aligned at all; they’re just improving the security so that the AIs can’t get up to that much trouble.
The important high-level claim is that ultimately we’re going to need alignment. Control has a time bomb associated with it. In particular, the time bomb is when the AIs are so smart and so good at subverting any control measures we put in place that, if they were trying to screw us over, we would simply fail. They would figure out some way of hacking around whatever systems we were able to put into place. Ultimately, we’re going to need to solve alignment. The problem with alignment is, as you were saying, that it’s much harder to measure.
It's much harder to know whether you've succeeded or not than with control. So, basically, our overall proposal is that in the meantime—in our scenario, between roughly 2030 and 2040, when we're in the human range of AI capabilities—we're going to rely almost entirely on control.
We're going to do things like red-team, blue-team games where we have our AIs try to escape our sandboxes or our control techniques. Then we're going to see if the red team succeeds. If they do, we're going to iterate and make better security until they can't succeed anymore. We have this ability to measure it by just seeing: Can the AI succeed? Can the AI, augmented with humans, succeed? Can humans playing the role of AIs succeed at gaming out alignment?
I think in order to be confident in the alignment of our systems, we're going to need more fundamental scientific breakthroughs. I don't think you're going to be able to figure out if the AIs are aligned via behavioral evaluation alone, where you just look at what the AIs are doing. I think you need to understand something deeper about what's going on inside the AIs' minds. You're going to need something like interpretability.
You're going to need some sort of notion of the difference, because you're fundamentally going to need to distinguish between the AI that's pretending to do the nice thing because it's pretending and is waiting and biding its time, and the AI that is fundamentally doing the nice thing because it wants to do the nice thing. I think that will require this white-box understanding of what's going on inside the AIs, whereas I think the control thing you can just figure out via empirical behavioral iteration.
Overall, our story in AI 2040 is that between 2030 and 2040, we're going to buy time. We're going to rely on control, have human-level AIs, and use those human-level AIs to make a bunch of progress on alignment, but also on other things. Then, by 2040 in the story, we've made enough alignment progress that we really don't need to rely on control. Therefore, we can scale to these wildly superhuman AIs where we're really relying on alignment, and if we were to try to control them, we would totally fail because, if they were misaligned, they would be able to completely subvert our control measures.
Do you think, in principle, that this problem is going to get easier or harder as the models get more sophisticated?
I think it's going to get harder because there will be a bunch of properties the models will have that will make it much harder to align them. In particular, we're already seeing growth in situational awareness. A few years ago, the AIs were not that situationally aware, and now I think they're pretty good at recognizing—for example, our friends at Redwood Research often do these control evals, and the models will often, in evaluations where they're trying to break out of these control setups, think, “Hey, this sort of looks like a Redwood Research control evaluation.” So they basically realize exactly what's happening to them.
I think this is just going to continue to happen. The more the AIs realize the exact setup, the more they're going to be able to perturb their behavior to look good to the exact process that's evaluating or training them. That's going to come further and further apart from the actual measurement of whether they're good or not.
Yeah. And just to add something to that, I think in some sense the core problem is that it is already somewhat easy to think that you've solved the alignment problem and be wrong. That's going to get easier and easier over time as the models get more sophisticated and become more aware of their situation and cleverer and stuff like that.
So it's not that we think there's going to be loads and loads of egregious failures where the AIs are just going around killing people. No, it's almost the opposite. It's going to be extremely easy to end up in a situation where the AIs are in fact misaligned, but you don't know that because they're doing everything right as far as you can tell. The number of ways in which that could end up happening is just going to increase over time, and it's going to be so easy to end up in that trap.
11. Why AI research should be public
So, we talked about buying time—why we want to extend the time with AGI—but what do we actually do during that time? The second principle is basically transparency. Transparency isn't necessary for making everything else happen, but it's really, really helpful.
There are a huge number of upsides to transparency. One is this concentration-of-power issue. We're very worried about someone building superintelligence and it being aligned to only some particular group of people. We think it's much harder for that to happen in a nondemocratic way if society as a whole can see the whole time what's going on with AI, how smart the systems are, and who they're aligned to.
For example, if an AI CEO who was evil was trying to backdoor their model and put in training data that says, “Hey, obey me and don't obey anyone else,” or, hypothetically, if an AI CEO was saying that their model was truth-seeking and would only say the truth, but actually the model was looking up the CEO's political opinions before answering—didn't you guys—
Which happened.
Yeah, basically. So transparency will help, at least somewhat, with that. The other thing that transparency maybe helps a lot with is this issue of government capacity, where in Plan A we want governments to make these pretty technical and really complicated decisions: exactly how much AI scaling to allow, what risks are okay versus what risks aren't okay, exactly what architectures are maybe safe versus what architectures are not safe, and what sort of control scaffolds are sufficient to ensure safety versus which ones are bogus safety-washing.
Making all of those calls will be very tricky, particularly given that the government's expertise in AI is really, really bad. One of our core hopes is that, with more transparency and more public understanding of what's going on in the companies, that relieves pressure on the regulators.
If something catastrophically or existentially unsafe is happening, society as a whole—academia, nonprofits, other AI companies that have an incentive to say, “Hey, my competitor is being super unsafe,” other governments, like China, which has an incentive to do this with US labs, and the US government, which has an incentive to do this with Chinese AI companies—everyone who is an adversary or just wants to make sure that things are safe has this big incentive, and now has the ability to actually look over what's going on.
They can see what is necessary for safety and what is not. I think maybe one thing that's really topical here is the Hugging Face incident that happened very recently, a few weeks ago, with the AIs inside OpenAI creating an internal message board. We still have very little clue about the exact motivations of those AIs, the exact context, or the exact prompt during the cyber evaluation that the models were given that prompted them to start doing this.
If I personally had much more access to what was going on, I would have a much more informed and better opinion on exactly what caused this, what mechanisms could be used in the future to prevent this, and what analogous future things I should be worried about because of this. I have a bunch of different hypotheses, but it's hard for me to figure out which is which without access to the data.
Basically, in Plan A, a core principle would be that all of that stuff would be transparent to the public, not just to governments. Society as a whole would be able to weigh in. There would be public and informed debates.
The scientific community especially, right? If you want to have a bunch of scientists, academics, nonprofits, and startups all weighing in on stuff, they need to have the information, and you can't really share it with all of them without sharing it with the public. So you might as well share it with the public.
12. Can the US and China enforce an AI slowdown?
I feel like maybe we should also say that we talked about the 5 goals. Then we talked about these pillars, which are kind of intermediates, but I want to go to the other end of the spectrum and talk about what the actual concrete things are that the US and China agree to in Plan A—
And how do they lead to those things? Specifically, we basically round up 99% of the compute—which is not people's personal compute, but big, big data centers—because most of the world's compute is in big data centers.
And we—the US and China, and the other countries involved—send inspectors to confirm that there are this many GPUs at this location and that many GPUs at that location. Having done that, we then set up this verification infrastructure and transparency infrastructure so that there are inference data centers that serve customers just like today. They are restricted so that they can only do inference and serve customers, and they cannot do any training runs. The inspectors make sure that they cannot do training runs at those data centers.
Then we have the training data centers, which are the totally transparent data centers. That is where the research happens. They still operate normally, but there are inspectors from the different countries monitoring the logs of what is going on in the data centers and publishing them to the internet. It is totally transparent what is going on in those data centers.
You might need some time to set this up. That is why we had the 6- to 12-month pause that I mentioned earlier. Once you get all this set up, you can proceed with AI development under these conditions of total research transparency. Because you have this transparency set up, it is a lot easier for countries to make further agreements about what to do and what not to do, because they can just see what everybody is doing. They can enforce the agreements pretty easily.
For example, here is where we would say it is very important that they agree not to do a crazy intelligence explosion and instead proceed slowly and cautiously. It is very important that they agree to do all this control setup, with all the red-teaming and so forth. Because of the transparency, they can make those agreements on an ad hoc basis. They can keep making more agreements like that and adjust them as needed based on the changing situation on the ground, because they can all see the situation on the ground. This is also very important because if you want to have some sort of deal between the US and China, the US and China do not trust each other, so you need some way of enforcing and verifying compliance with the deal. The transparency goes a long way toward helping with that.
Yeah. What about cheating and dark markets appearing?
There are 2 ways you could cheat. One is that you could get a bunch of GPUs and try to keep them from being discovered by the US and China—to hide them away, put them under a mountain somewhere, and do your training runs in secret. The other way you could cheat is by running a giant training run on the giant, legal, known data centers while trying to make it look like everything is normal and that nothing illegal is happening.
The mitigations are different for these 2 threat models. For the tiny amount of compute under a mountain, there are basically 2 mitigations. One is, as Daniel was saying, to round up enough of this compute so that it is pretty hard to get a substantial amount of compute under the mountain. The second is to do normal intelligence gathering and look for these things proactively over the course of the 10-year slowdown that happened in the scenario, which in reality might be longer or shorter.
For basically any significant-size GPU cluster, I think both of these independently have a quite good chance of working. In practice, I am not that worried about large hidden compute clusters secretly under mountains or whatever. I think the maximum realistic size, in my opinion, is something like a few hundred thousand GPUs—a few hundred thousand H100s—hidden away.
I do not think that would be enough to compete at the frontier, especially assuming the 2030 AGI timelines, where you have the biggest AI companies in the world having millions or tens of millions of H100s in their biggest training runs. For the legal clusters, we basically hope that we have a bunch of verification infrastructure running on those data centers. The main point of that verification infrastructure is to ensure that the computation happening on those clusters is transparent, and in particular that everyone can see it.
The hope is that you make sure there are no calculations running that are not transparent. For everything that is transparent, regulation and agreements as normal can work. You can say, “We agree to run this control scaffold if you agree to run this control scaffold,” and that can just happen. Both sides can be confident that they are both agreeing.
Isn’t this just a matter of national security, though? Isn’t it a bit ambitious just to make it completely transparent?
Yes. One of the effects of doing this transparency is that we would be immediately publishing all the core training recipes of Anthropic and OpenAI for the world to see. Anthropic and OpenAI will not be happy with this. It will cut into their valuations dramatically.
Why will it cut into their valuations dramatically? Because it will allow other competitors, like Microsoft and Alibaba, to catch up or whatever. That is why they are probably going to hate it. But I would say this is a feature, not a bug. We want there to be multiple different AI companies at the frontier at roughly similar levels of capability. We want AI to commoditize instead of being monopolized or oligopolized.
This will also disincentivize further investment, right? Investors will be less interested in building a trillion-dollar cluster if they will not be able to get the monopoly rents from that cluster. But we think this is good because we are going to be in a world where going too fast is our main problem. Having a bit less incentive to invest and going a bit slower is a feature, not a bug. There will still be investment and lots of money to be made, so progress will continue, just not at quite the same rate. Again, we think this is good.
This does shift things—it is kind of a gift to China relative to the US. China gets some algorithms that it might have had trouble getting before. Insofar as you really do not like that, negotiate it. You can have a horse-trading arrangement where, when the US and China are making the deal, the US says, “Since we are giving you all this stuff with the transparency, why don’t you give us something else in return?” We can try to work that out.
Like a more favorable compute distribution.
Like a more favorable compute distribution, for example. There has to be some combination of carrots and sticks and trading going back and forth that we think would be in the interest of both sides. Another thing worth mentioning is that it is not as big a gift to China as you might think, because security is poor at these companies. Through their spy networks and leaks, they are probably getting most of the information anyway.
13. Why AI policy debates miss the technology
I do not know if you guys saw Sam’s tweet yesterday, basically saying that they are going to pause training for a while. It made me think: Why not just stop now? You guys are actually quite bullish about some of the positive things that AI can do. There were loads of examples in your article, but one example was that, in hospitals, we could have little devices that decontaminate the air and stop the transmission of diseases and stuff like that. So it is not like you guys actually think we should stop.
We do think we should do something like Plan A as soon as possible. First of all, I think that just stopping everything now—Plan S—would be better than the default. I would rather just stop everything now than continue on our current trajectory. Secondly, our actual recommendation would be to do Plan A.
You do a temporary inference-only pause so that you can set up all the transparency and verification infrastructure. Then you can continue in this more distributed, transparent, cautious way, as previously described. Having continued in that way, you basically go up to the level that you feel like you can control reliably, until you feel confident that you have solved alignment well enough that you can give up on control. That is what we depict happening over the course of the 2030s in our scenario.
Why do you guys think that AI discourse is so bad? Is it unique to AI, or is discourse just bad in general?
Discourse is bad in general, right?
Well, why?
I do not know. I think there is a lot to say. One thing is that there are massive incentives and motivated reasoning for people at AI companies to think that what they are doing is justified and good, and that they should not take costly actions that would make the situation better, because those costly actions would be bad for whatever reason.
So, there’s a bunch of rationalization. I think most of the effect is probably just that discourse is generically hard and bad. Twitter or whatever is very unnuanced and argumentative.
I think LessWrong in particular, which is where I try to do most of my discourse, actually has pretty good discourse quality on average. Generally, when I write a post and go into the comments, I think the comments are quite thoughtful and technically informed, especially relative to other places like Twitter.
And then maybe a final thing is just that, in D.C. in particular—which is an area I care a lot about—I really want D.C., the government of the United States and other governments, to react well to AI technology. I think there are 2 problems going on there. One is that there isn’t very much AI expertise, right? The government isn’t hiring really high-quality technical AI experts who really know what they’re doing.
The second is this generic issue: I think the conversation isn’t happening because the incentives for everyone in D.C. aren’t toward truthfully and accurately understanding the situation. The incentives are for every individual to say stuff that sounds good and looks good, and that’s within the D.C. Overton window, so they can make a lot of friends. Unfortunately, I think this comes apart from the actual reality.
I think the actual reality of the situation is that a bunch of controversial and niche views about AI turned out to be true, right? The whole AGI hypothesis is correct. I think D.C. basically just hasn’t come to grips with that. So almost all of the discourse happening there is fundamentally anchored on completely wrong assumptions about how the technology works, in particular the assumption that it’s mostly fake news and it’s all a bubble.
Or that it’ll be like the next internet.
Or that it’ll be like the next internet, which is maybe better than it was a few years ago, when it was even more bearish, but it’s still not nearly bullish enough on the technology, in my opinion.
14. Is AI normal technology? The remaining disagreement
I think it was the AI Snake Oil guys who had an article saying that AI is normal technology. Obviously, it’s been said that AI is not a normal technology. It’s really quite different, right? So he’s very, very skeptical about AI, but you had some great discussions with him, and none of you changed your mind as a result of that. I think quite an interesting thing is the skeptics’ perspective as well. A lot of skeptics think that folks in Silicon Valley just aren’t being sincere. What they’re saying isn’t sincere.
Sure. With respect to sincerity, I would say some people in Silicon Valley aren’t being sincere, but others are.
Such as us.
Such as us, but also some of the people at the AI companies are being sincere. Not all of them. I don’t think you should trust what the leadership of the company says in general.
15. AI as Normal Technology
Anyhow, we co-authored a blog post or an article with the “AI as Normal Technology” people. It was a post about what we agreed on. You can go read it; there are 10 points in there that we both agree on. One highlight from our perspective is that we had a bit of a truce where we said, “Yes, AI right now may be a normal technology, but in the future it will not be normal. In particular, in the future it’ll be more like humans in the cloud, and all this crazy stuff is going to start happening,” as we describe in AI 2027. They agreed that if you get humans in the cloud, that would not be a normal technology. They just think you’re not going to get that, at least not for many years.
AI—is it possible that in the next few years we will have AIs that are like humans in the cloud and that can do all sorts of knowledge work in a way that substitutes for humans, including AI research, for example? And then that, sometime after that, we will have robots, perhaps controlled by those AIs, that can do physical work in a way that broadly substitutes for humans? Our claim is that, yes, that sort of thing will be achievable in the next few years. Their claim is, “No, not in the next few years. That’s much farther away.” I think that is the main source of our disagreement.
We would agree with them that if that level of AI and robotics is still very, very far away, then maybe AI is more of a normal technology. It’ll be like the next internet or something. But we just think that it’s actually on a path to get to that level of capability soon.
Can you be more specific about what the core cruxes are and what would make either of you change your mind?
I don’t know if I have a useful answer to that. I think there are a lot of different things we argue back and forth about. One thing that would change my mind is that people keep talking about the limitations of the current paradigm, but then the limitations of the current paradigm keep getting overcome within the current paradigm.
One thing that would change my mind is if someone was actually right about one of these limitations. If someone right now is going around saying that data efficiency or non-verifiable tasks is the current limitation, and then several years go by and it becomes clear that there’s basically no progress on overcoming that limitation—that the AIs of 2029 are no better at fuzzy, non-verifiable tasks than the AIs of 2025—then I’d be like, “Okay, this feels like a real barrier. This feels like something where we’re starting to feel the elephant. We’re running up against some sort of real, actual barrier that was correctly predicted by theory by some people to be there, and now it’s actually there.”
By contrast, from my perspective, there are just loads of experts going around talking about all these barriers, and then we keep plowing through them as if they’re not there. So actually running up against some sort of barrier like that would lengthen my timelines quite a lot.
Another thing that would lengthen my timelines quite a lot is more political change. If there was a war with China and most of the chips got destroyed by missiles, that would lengthen my timelines.
Exactly, like a vintage, though.
Okay, that’s true. On the bright side, if there was more of an international deal to pace the frontier, that would lengthen my timelines, et cetera.
I feel like the fundamental disagreement is really just this view of whether AI is more like electricity or airplanes, or whether AI is more like humans in the cloud. All of the intuitions are downstream of this core reference class of what we’re thinking the future will be like in the near term.
I think that if we froze AI progress right now and didn’t train any new models, it would become more of a normal technology. There are just so many things that Claude 5 can’t do, and if we couldn’t get any new models beyond Claude 5, then it would be more like the internet. We would restructure a lot of our professions and workflows to incorporate copies of Claude 5 doing parts of them, but then the humans would shift to doing more of the things that Claude 5 can’t do.
There would be a big change in a lot of things, and it would be like the next internet in the sense that it would change everything in some sense, but it wouldn’t fundamentally change anything.
16. Validity of the single processor approach to achieving large scale computing capabilities
Yeah. Right now you get these Amdahl’s law-type effects, where the AI can do some fraction of the workflow, but then it gets bottlenecked on the parts of the workflow that the humans have to do.
I still feel like there’s this thing where, when I think of the future, I’m imagining AI doing 100% of the workflow of a bunch of important economic workflows. You don’t get this bottlenecking effect. That’s a pretty qualitative change from the situation today that causes pretty fundamentally different predictions of what the world looks like. Whether AI actually gets to literally 100% of a bunch of very, very important tasks is the core question underlying the difference in worldviews.
I think that if we do see recursive structural adaptation, which is coherent—and I think it’s likely to be quite divergent—then that’s it for me. That’s not a normal technology.
Can you flesh that out a bit more? Would this be something like the Hugging Face swarm? Tell us more about what you would see that would be the thing for you.
Yeah. For me, adaptivity is the word most synonymous with intelligence.
I think these new RL-trained models we have aren’t the same as what we had many years ago. We said, “Scale is all you need,” and the clue is in the name. Scaling means you take a scalar property of a system and scale it up.
These RL systems are still self-attention transformers, but they’re different. They’re actually structurally different. There are new types of training, new architectures used differently, and so on.
A bunch of humans did some experiments and adapted the structure to create a system that had different scaling properties. I can imagine a future where we actually have some kind of recursive loop in which the system is adapting itself, deciding what things are interesting, and evolving by itself. When that happens efficiently, I think that’s a different type of technology.
Yeah, that sounds kind of similar to what we would say about recursive self-improvement and automating the AI research process itself. Yeah.
Yeah. I guess we agree, then.
Amazing, guys. It’s been an honor and a pleasure having you both on MLST. Thank you so much for joining us today.
Thank you for having us.
Yep. Yeah, appreciate it.