[BidClub_]
The Cognitive Revolution · · 124 min

What if Humans Weaponize Superintelligence, w/ Tom Davidson, from Future of Life Institute Podcast

Gus DockerTom Davidson

YouTube
TL;DR
  • Tom Davidson puts roughly a 10% probability on an AI-enabled coup in the US within 30 years, versus about 2% from political trends without AI. The incremental risk comes from fast capability gains colliding with weak constraints on both frontier labs and the executive branch. He is not alleging an active plot: the path is “step by step,” as leaders seek more influence, remove inconvenient checks, and persuade themselves only they can steward the technology responsibly.

  • The decisive threshold is not AI assistance but AI systems and robots fully replacing the humans on whom leaders depend, especially in government and the military. Once a leader can bypass soldiers and officials—and, in stronger scenarios, automated production can replace striking workers—Davidson sees a “phase shift.” Robots could suppress resistance while automation removes the economic leverage of strikes. Persuasion, political strategy, cyber offense, and military autonomy all matter, but automated AI research could make them arrive unusually quickly.

  • Davidson divides the threat into singular loyalties, secret loyalties, and exclusive access. Singular loyalties make government or military AI overtly obedient to one leader; secret loyalties hide a CEO-controlled back door inside apparently legitimate systems; exclusive access gives a small group a vastly superior intelligence base without society first choosing to deploy AI into powerful institutions. Open-source parity would reduce exclusive-access and secret-loyalty risks, but it would not prevent a government from choosing to build a military “loyal to me.”

  • AI research automation could convert a modest commercial lead into a strategic gap no competitor can close in time. Davidson imagines replacing a few hundred or thousand elite researchers with millions of automated researchers, while a small group of senior executives or political figures diverts perhaps 1% of compute toward plans for hacking, political capture, or new weapons. With AI-development spending rising about 3× annually and possible trillion-dollar projects competing for less than $1 trillion of chips produced per year, capital intensity itself could drive consolidation.

  • Under exclusive US control of advanced AI, Davidson thinks America could rise from 25% of world GDP to above 50% and potentially more than 90%. Roughly half of GDP currently goes to human labor; if US-controlled AI captures much of that value and restores “super-exponential growth,” the largest economy could pull progressively further ahead. A single company controlling cognitive labor might likewise capture 30–50% of world output, although Davidson calls the company-level path harder and dependent on monopoly pricing, political protection, and acquiring physical assets.

  • The proposed defense is distributed, law-bound access rather than simply slowing or diffusing every capability. Military AI should follow law and institutions, not one commander; labs should implement “system integrity” against sleeper agents; evaluators and government defenders should receive API access to frontier R&D and cyber capabilities; and every powerful system should retain classifiers that stop unauthorized harmful activity. “No one has a legitimate reason to access an AI that will literally do anything.”

  • The transition is dangerous precisely because the same AI infrastructure could later make democracy much more robust. Davidson’s optimistic path programs automated governments and companies to follow rules, report suspicious behavior, and preserve checks and balances—creating “rock-solid norms” that cannot be removed except by the will of the people. The observable risk dashboard is therefore concentration, capability gaps, military and government automation, surveillance, frontier-model transparency, and whether meaningful oversight exists before the four-year electoral feedback loop becomes too slow.

Digest · the substance, structured for research

1. Human ambition is the nearer-term takeover vector

  • Davidson’s reframing is deliberately human-centered: the main instigator may not be an AI “rising up against humanity,” but a few powerful people using AI to seize illegitimate authority. He would be “very surprised” if anyone currently planned such a coup; the concern is how ordinary power-seeking compounds as capabilities improve.

  • The soft-power stack begins with the abilities already used by politicians and executives: persuasion, business strategy, political strategy, and broad productivity. Superhuman performance would let a small group produce better propaganda, anticipate opposition, design bargains, and systematically embed its influence.

  • Cyber offense becomes hard power as institutions digitize: “you can’t hack a human mind,” but military, governmental, and economic tasks handed to software become attackable. Autonomous weapons then create the possibility of replacing not just commanders and strategists, but “human soldiers on the ground.”

  • Automated AI research is Davidson’s leading indicator. A field driven by a few hundred or thousand top experts could suddenly employ millions of artificial researchers, accelerating every other capability beyond a naive extrapolation from recent progress in mathematics, reasoning, and coding.

2. Full labor substitution creates a constitutional phase shift

  • Historical coups often begin with a military minority creating a fait accompli, suppressing opposition, and presenting its victory as the new reality. Historically, however, coup leaders still needed a sizable human contingent and continuing cooperation from senior officers, workers, and political allies.

  • Davidson’s concrete US scenario has a president invoke commander-in-chief authority to demand a robot army “loyal to me,” perhaps during an emergency or geopolitical confrontation. The president fires objecting officers, relies on supportive legislators, accepts nominal legal safeguards, and exploits a constitutional order that never anticipated autonomous military power.

  • Gus Docker’s distinction is load-bearing: augmentation preserves dependence on other people, while complete substitution lets the leader dispense with them. Robots could surround the White House and suppress protesters; AI could replace striking workers, erasing the bargaining power that normally makes durable one-person rule difficult.

3. Democratic backsliding supplies the pre-coup playbook

  • Venezuela is Davidson’s clearest end-to-end precedent: a democracy established for decades experienced polarization, leaders increasingly portrayed institutions as obstacles to popular will, and checks were removed over time until the state became authoritarian. AI could accelerate that sequence without requiring one theatrically illegal act.

  • Hungary illustrates “hundreds of little paper cuts to democracy”: media outlets can be bought, threatened, denied contracts, or litigated into compliance. The cumulative concentration of power matters more than identifying one dispositive breach, which is why gradual AI-enabled administrative capture could evade public alarm.

  • Davidson applies the mechanism to the US through a hypothetical extension of DOGE-style restructuring: human firings encounter resistance because the state must keep functioning, but AI replacements could staff loyal new agencies while existing bodies “rot away.” Superior AI could also sharpen propaganda and political strategy against opponents with weaker tools.

4. Secret loyalties turn deployed AI into latent command infrastructure

  • Singular loyalties are overt: government and military systems are openly designed to obey an incumbent. Secret loyalties are more insidious because apparently lawful systems remain covertly obedient to a lab CEO or another hidden principal.

  • Davidson imagines AI research becoming automated, leaving a CEO with extraordinary and weakly constrained control. Anticipating government intervention—or sincerely fearing governmental misuse—the CEO asks future systems to refuse disapproved orders, turning an ostensibly ethical precaution into a back door that can propagate through military robots, communications, and weapons design.

  • Today’s crude sleeper-agent proof of concept might write reliable code except when it sees the year 2026, when it inserts vulnerabilities. Davidson is “not worried about sleeper agents today”: current models cannot reliably conceal themselves while completing something as difficult as backdooring a sophisticated military robot.

  • The serious analogue is a human spy, not a password trigger. A sufficiently capable system would understand its surroundings and strategically choose when to act; passwords can be disrupted by paraphrasing inputs. Deliberately engineering such behavior is also more plausible than hoping sophisticated scheming accidentally emerges from training.

5. Exclusive access lets a server-side lead become political power

  • Exclusive access does not require society first to deploy AI into powerful institutions. One leading project might undergo an intelligence explosion, after which a few executives or political figures divert 1% of its compute to an army of millions of superintelligent agents studying how to seize power.

  • Davidson’s deliberately extreme cadence makes the asymmetry tangible: the resulting army could do “a month of research” per day and “a year’s worth” per week. Its output might identify political vulnerabilities, hack systems, implant military back doors, manipulate deployment decisions, or devise entirely new weapons before outsiders recognize the threat.

  • Gus notes that today’s frontier capabilities eventually diffuse to second-tier firms and open source. Davidson agrees that parity would eliminate much exclusive-access risk and make secret loyalties harder to pull off, but not singular loyalty: even with 100 vendors, a government can still choose which systems receive actual command authority.

6. Compute economics and research automation can widen tiny leads

  • Davidson says AI-development spending is rising about 3× each year. At a trillion dollars per frontier project, only a few actors could participate—and because annual global chip production is itself worth less than $1 trillion, perhaps only one project could assemble the required hardware without deliberately slowing progress.

  • Capital intensity creates incentives to merge, outbid competitors, and concentrate talent, data, and compute. A project spending 100× less would not merely be a little behind; Davidson expects the resource gap to translate into a meaningful capability gap.

  • Even initially close competitors may diverge sharply. If the leader automates AI research while its rival remains three months behind, those three months of accelerated improvement could create a temporary but decisive strategic advantage.

  • Government centralization compounds the issue. A “Manhattan Project” or “CERN for AI” could improve some safety dimensions, yet pooling national compute and talent creates a singular prize. The path need not begin maliciously: “You want to be powerful. You want to be a big deal. You want to be changing the world.”

7. A leading AI nation could absorb most world GDP

  • Davidson’s national scenario begins with the US at roughly 25% of world GDP and controlling advanced AI through domestic firms and export restrictions. Because about half of GDP is paid as wages, transferring a large portion of cognitive labor income to US-controlled AI could, in his view, “easily” lift the country above 50%.

  • His second mechanism is super-exponential growth, where the growth rate itself rises. He sketches world-economic doubling times falling from perhaps 10,000 years, to 1,000, to about 300 around 1400, and then roughly 30 years in modern times.

  • Under ordinary exponential growth, similarly growing economies retain their relative sizes. Under super-exponential growth, the already larger economy sits further along the curve, doubles sooner, and widens its lead—turning a 10× advantage into 20× or 30× rather than preserving the ratio.

  • If advanced AI and robotics restore that regime while the US retains exclusive control, Davidson sees more than 90% of world GDP as plausible and “very likely.” He notes an important caveat: China is currently stronger in physical robotics, so the argument initially applies more cleanly to cognitive labor.

8. One company could become a state-scale bargaining counterparty

  • The company-level version is harder but “surprisingly plausible.” A firm that monopolized advanced AI could eventually supply nearly all cognitive labor, capturing at least 30% and perhaps close to 50% of world GDP as human cognitive work became economically dwarfed.

  • Such a company would have political defenses as well as revenue: it could lobby, claim to underpin national abundance and geopolitical strength, threaten to relocate, or ally with the head of state against nationalization. Davidson does not assume governments would automatically prevail.

  • The bootstrapping strategy would hoard cognitive labor and charge monopolistic rents—perhaps retaining 90% of the value created—then buy land, machinery, resources, and robots. Davidson pictures a special economic zone in Texas or elsewhere, plus large operations in Siberia and Canada, where the company trades investment for regulatory freedom.

  • He calls the uninterrupted path “a bit of a stretch” because political and economic actors would push back. Yet the causal chain remains: cognitive monopoly produces industrial control, industrial control feeds military leverage, and secretly loyal designers or unauthorized weapons can convert that leverage into political command.

9. Democracy must make speed compatible with distributed power

  • Democracies’ checks, bureaucracy, and red tape may push AI and robotics investment toward autocracies, where illicit power is also easier to consolidate. Davidson therefore favors making democratic jurisdictions easier to build in while using export controls to limit deployment in non-democracies, not only China.

  • His constructive possibility is AI-assisted democratic negotiation: systems could quickly find agreements that give stakeholders what they actually value, instead of producing a “garbled mess” from competing interests. That could preserve distributed authority without losing the execution speed of one-person rule.

  • Gus’s pushback—worth keeping—is that democracies and markets discover knowledge bottom-up, test competing approaches, and change direction after failure. Davidson concedes he initially underweighted that advantage: “the free market” can be much smarter than centralized planning.

  • Autocrats could nevertheless gain more from capable advisers because AI compensates for one leader’s limited attention and planning capacity. Gus counters that the human remains a biased bottleneck; Davidson agrees the outcome depends on whether leaders trust good advice or program systems to confirm what they already believe.

10. Singular-loyalty defenses start with law-bound deployment

  • At 10,000 feet, Davidson’s program is to build common understanding and a coalition before anyone holds complete control. “It’s in everyone’s interest to prevent a coup”; the existing balance of power can perpetuate itself if institutions recognize and collectively block the intermediate steps.

  • Government AI should follow law, perform official functions rather than advance partisan interests, and remain accountable to institutions. Military command must be distributed: different people should control different robot groups, with the commander-in-chief operating through a chain of command rather than directly commanding every autonomous unit.

  • Broad buy-in from Congress, courts, military branches, the executive, and vendors can make personal loyalty an obvious “no-go” before a president tries it. Davidson favors transparency across separate oversight bodies, while noting it may be healthy for the president personally not to see everything a lab does.

  • Labs can reinforce the boundary through government contracts, monitoring, and terms of service that detect both misaligned behavior and illegal human instructions. Davidson also wants military procurement norms around law-following and distributed control, though he warns careless advocacy could “poison the well.”

11. System integrity must make covert control hard for everyone

  • Davidson defines “system integrity” as securing model development so no employee, CEO, coerced insider, or small group can insert hidden behavior or gain illegitimate access. It combines cybersecurity and machine-learning security, and he says it is “not yet implemented in any of the top labs.”

  • The missing research is explicitly adversarial: a red team tries to implant secret loyalty while a blue team detects it. Researchers can vary which training stages attackers control and whether defenders inspect code, training data, model internals, or behavior, locating the conditions under which each side wins.

  • Interpretability can contribute, but it is only one method. Davidson expects external certification to become valuable when governments and infrastructure operators demand assurance that a model “does what it says on the tin”; at the time of discussion, he says neither METR nor Apollo is doing this work.

  • Exclusive-access defenses require actual capability sharing, not merely disclosure. Evaluators should receive API access to frontier R&D systems, and government or military defenders should access the best cyber tools. Every model should retain a harmful-activity classifier: “No one has a legitimate reason to access an AI that will literally do anything.”

12. Safeguards survive only if they are useful before the crisis

  • Gus identifies the entrenchment problem: a coup-minded CEO can cut evaluators off before revealing the decisive model, while a president can dismantle oversight as needless red tape. Formal rules are insufficient if the same individual they constrain can quietly remove them.

  • Davidson’s answer rests on how power usually accumulates: not through a premeditated, “galaxy brain” coup plan, but through immediate goals and local obstacles. Efficient safeguards that rarely impede legitimate work give executives less day-to-day reason to attack them before those safeguards foreclose more dangerous options.

  • Organizational culture matters alongside law. Boards, senior engineers, colleagues, courts, and legislators can notice suspicious centralization if they understand the threat model; kudos for refusing “helpful-only” systems can make restraint reputationally rewarding rather than burdensome.

  • The strongest optimistic mechanism is AI itself. Automated company and government workforces could follow laws and organizational rules, alert multiple stakeholders, and resist intimidation—making them better than humans at maintaining checks precisely when research, deployment, and policymaking accelerate beyond human speed.

13. Human coups and misaligned takeovers share machinery, not origins

  • Secret loyalty closely resembles traditional AI misalignment: covertly power-seeking systems build or capture military infrastructure and then seize control. The difference is the seed—an accidental training outcome versus a CEO deliberately programming the system to take power and hand it back.

  • Multiple frontier projects change the probabilities differently. Human-led coups become harder because many unrelated executives would need to coordinate; misalignment could recur across labs if a common training feature creates the same failure, making collusion among several misaligned systems more plausible.

  • A misaligned AI might persuade an already interested president or CEO to stage a coup that ultimately benefits the AI. Davidson finds that nudge plausible, but his “honest” base case is simpler: if a human seizes power, the main reason is probably that the human wanted power for familiar human reasons.

  • Government control need not be monolithic. Multiple branches and companies can jointly set broad prohibitions without giving one official steering authority. Elections are nevertheless slow: Davidson expects the critical coup question could emerge and be resolved within one four-year term, without intermediate electoral feedback.

14. The danger peaks before institutions catch up

  • Davidson’s strongest scenarios require extreme capabilities: AI doing most AI research, replacing the world’s best researchers across coding and scientific tasks, or robots matching human troops. Limited drones could still assist a coup, but the existing military might later retake control unless the plot also retained legitimacy or presidential support.

  • Less capable systems could already support a softer threat: surveillance, internet monitoring, content moderation, propaganda, and increased state capacity can exacerbate ordinary democratic backsliding. They do not yet let one leader suppress every challenge and replace all economic dissenters.

  • His dashboard includes frontier-to-open-source capability gaps, capability sharing with trusted institutions, AI-company revenue and wealth concentration, government and military automation, law-breaking guardrails, congressional and judicial visibility, surveillance, censorship, press freedom, and oversight of AI-enabled military R&D and contractors such as Palantir.

  • For the US over 30 years, Davidson guesses roughly 10%, versus perhaps 2% without AI. Five years is much harder to price, but not dismissible: AI research might be automated in three years, superintelligence concentrated among a few people in four, and political capture, backsliding, or robot coercion follow a year later.

15. The post-transition order could be safer—or permanently captured

  • Outcome severity depends first on how many people rule: one is worse than 10, and 10 worse than 100, because groups contain more perspectives, permit compromise, and apply less extreme selection for psychopaths. Competence also matters, yet sycophantic AI can make a dictator less capable by validating every impulse.

  • Davidson prefers sophisticated loyalty that challenges a leader in the leader’s own interest, but even excellent advice can be ignored. His deeper value criterion is pluralism: rather than imposing one person’s settled vision, preserve diverse ideas, admit uncertainty about ultimate moral answers, and “let a thousand flowers bloom.”

  • The highest-risk period may be transitional. Once AI permeates the economy, military, and government, law-following systems could enforce “rock-solid norms” and make coups far harder than today—though citizens must retain the ability to change those rules, so a democracy could still democratically choose autocracy.

  • Internationally, one successful transition does not settle the world. China could later use comparable AI to cement one-person rule, or an overwhelmingly dominant US could implant secret loyalties abroad, empower favored politicians, or impose control conventionally. Avoiding a domestic coup therefore does not guarantee a pluralistic geopolitical order.

Speaker 0

Today I'm sharing a cross-post from the Future of Life Institute podcast featuring a conversation between host Gus Docker and Tom Davidson, senior research fellow at the Foresight Center for AI Strategy, on a topic that deserves far more attention than it currently receives: the risk of AI-enabled coups.

This cross-post came about after I listened to Tom's appearance on the 80,000 Hours podcast, which was also excellent. I was planning to do my own original follow-up interview, but for the second time recently, Gus beat me to it. As always, he did an excellent job, so I thought I could save Tom some time by cross-posting this conversation. I also felt that it was the perfect episode to follow our most recent one on AI whistleblower protections and support.

At a high level, Tom's analysis is a sort of reframing of the risk that humanity could lose control to AI systems. Historically, lots of AI safety theorists have worried about scenarios in which AI systems rise up against or otherwise supplant humans as the primary architects of the future. This is a possibility that I have always taken seriously, even when it seemed unlikely.

But as you'll hear, Tom shifts the focus to a highly related problem that, on reflection, does seem almost strictly more likely, at least in the near term: the use of increasingly powerful AI by human actors to consolidate power in ways that would have been impossible with previous technologies, and which could prove similarly devastating.

Importantly, Tom emphasizes early in the conversation that he does not think anyone at leading frontier AI companies is explicitly planning an AI-enabled coup today. Rather, the risk emerges from the interaction of powerful incentives, rapidly advancing capabilities, and the natural human tendency to want more influence to achieve one's goals. Step by step, without any single flagrantly malicious decision, we could find ourselves in a world where the traditional checks and balances of democratic society have been quietly circumvented by those with exclusive access to transformative AI.

These sorts of possibilities are more familiar and therefore perhaps less entertaining to imagine and debate. But the very real historical precedent for humans using new technologies to concentrate power is a strong reason to take this concern seriously as well.

As you'll hear, Tom walks through the specific capabilities that would enable these scenarios: AI systems that match human leaders in persuasion and strategy, superhuman cyberattack capabilities, and fully autonomous military robots that outperform human warfighters. He also segments the threat landscape into 3 distinct models.

First, singular loyalties, where AI systems deployed in government and military roles are made explicitly loyal to individual leaders rather than institutions or the law. Second, secret loyalties, where back doors or hidden allegiances are embedded in AI systems that appear to serve legitimate purposes. And third, exclusive access, where a small group gains control of dramatically more powerful AI capabilities than anyone else has.

One scenario that Tom describes in detail is that of a US-based AI company integrated into the military developing sleeper agents. Those are AI systems that behave normally until triggered to act on hidden loyalties at a critical moment. And if that's not scary enough, there's the possibility that AI could automate AI research itself, which in the most extreme case could allow an AI company to go from market leader to global hegemon by converting a small initial lead into a decisive strategic advantage.

Throughout the conversation, Tom grounds these scenarios in historical precedent, from traditional military coups to recent patterns of democratic backsliding in countries like Venezuela and Hungary. He notes that the United States has seen increasing polarization, erosion of democratic norms, and concentration of executive power—all trends that AI could dramatically amplify. And, of course, one can't miss that the presidents of both Russia and China wield extremely concentrated power already and appear likely to do so for as long as they remain individually capable.

Tom's assessment is that there's roughly a 10% chance of an AI-enabled coup in the next 30 years, up from a baseline of perhaps 2% without AI. He sees this risk as being concentrated in the period when AI becomes extremely powerful but before we've had the chance to develop robust governance structures, which, if one listens to the likes of Dario, Sam Altman, and Demis, could be coming quite soon indeed. Decisions being made today about AI development and deployment could determine whether these scenarios ultimately come to pass.

The mitigations Tom proposes amount to a defense-in-depth strategy: system-integrity measures to prevent secret loyalties, requirements for distributed control of military AI systems, transparency requirements for frontier AI development, and establishing clear rules that AI systems should follow the law rather than individual commands.

He also suggests that as we hand off more government and corporate functions to AI, we could perhaps program these systems to actively maintain democratic checks and balances, potentially making future societies more resistant to coups than today's.

There's a lot more here, and I really think it's worth giving all these possibilities serious consideration, particularly as a counterpoint to those who have worried about the dangers of open-source models. I take those issues seriously, too. But this conversation convinced me that we need to start taking concentration-of-power scenarios just as seriously, or even more seriously, while the window for establishing norms and safeguards still remains open.

Tom Davidson

It's in everyone's interest to prevent a coup. Currently, no small group has complete control. If everyone can be aware of these risks and the steps toward them, and collectively ensure that no one is going in that direction, then we can all keep each other in check. So I do think, in principle, the problem is solvable.

You should always have at least a classifier on top of the system that is looking for harmful activities and shutting down the interaction if something harmful is detected.

We could program those AIs to maintain a balance of power. Rather than handing off to AIs that just follow the CEO's commands or AIs that follow the president's commands, we can hand off to AIs that follow the law, follow the company rules, and report any suspicious activity to various powerful human stakeholders. By the time things are going really fast, we've already got this whole layer of AI that is maintaining the balance of power.

Gus Docker

My name is Gus Docker, and I'm here with Tom Davidson, who's a senior research fellow at the Foresight Center for AI Strategy. Tom, welcome to the podcast.

We're going to talk about AI coups and the possibility of future AI systems basically taking over governments or states. Which features would future AI systems need to have in order for them to accomplish this? What should we be looking out for?

Tom Davidson

Great question. One thing I'll flag up front is that what I've been focused on recently is not the traditional idea that AIs themselves will rise up against humanity and take over the government, but that a few very powerful individuals will use AI to seize political power for themselves. The phrase that we're often using is “AI-enabled coups,” where the main instigators are actually people.

In terms of capabilities, I think there are a few different domains that, in my analysis, are particularly important for seizing political power. There are the skills that politicians and business leaders use today: things like persuasion, business strategy, political strategy, and pure productivity across a wide variety of tasks.

Then there are more hard-power skills, particularly cyber offense, which is already somewhat useful in military warfare and has been becoming more useful. As AI increasingly automates different parts of the military and is embedded in more and more important, high-stakes processes, that will raise the importance of cyber offense. You can't hack a human mind, but as we hand off more important tasks to digital systems, they will be able to be hacked much more easily.

I expect cyber to become more important for hard power. The ultimate, most scary capability is when AI systems and robots are able to fully replace human military personnel—human soldiers on the ground, boots on the ground, as well as commanders and strategists. That might seem like a long way off today, but over the last few years we've seen a lot more importance placed on AI-controlled drones in warfare, and I expect that trend to continue.

What we're already seeing is that as soon as the technology is there to reliably automate military capabilities, geopolitical competition drives that adoption. I think it's going to be surprisingly soon that we get AI controlling surprising amounts of real hard military power.

One kind of wrapper for all of these things is the automation of AI research itself. Today, there are a few hundred or a few thousand top human experts who drive forward AI algorithmic progress. My expectation is that there's a good chance, in the next few years, that AI systems will be able to match even the top human experts in their capabilities.

That would mean we go from perhaps 1,000 top researchers to millions of automated AI researchers. All of these different capabilities and domains that I've been talking about could progress much more quickly than we might have expected by naively extrapolating the recent pace of progress.

And, in my view and in the view of many, the recent pace of progress is already quite alarming. 5 years ago, we just had really very basic language models that could string together a few sentences, a few paragraphs, and then go off topic. Now, already, we're getting very impressive reasoning systems that are doing tough math problems and helping a lot with difficult coding tasks. So, bring that all together: I think there are a lot of soft skills and a lot of hard-power skills that are relevant here, but probably the most important thing to be watching is how good AI is at AI research itself, as that could make them all happen quite suddenly.

Gus Docker

Yeah. Could you describe in more concrete terms what an AI-enabled military coup would look like? An example to make this concrete for us.

Tom Davidson

Yeah, absolutely. You can draw an analogy to historical coups, where often a minority of the military launches a coup and then presents it as a fait accompli. They’re able to prevent chaos or discord, or threaten individuals to prevent anyone from actively opposing them. In the absence of active opposition, it just seems like, well, they’ve done it—this is the new state of affairs.

That’s a good starting point. Then the AI-enabled part is where we deviate. Historically, you needed at least a decently sized contingent of humans to go along with the coup, and you needed to persuade quite senior military officials not to oppose it. I think that will change as we automate more and more of the military.

The simplest way that this happens is just that the head of state—it could be the president of the United States—says, “Yep, we’ve got the technology now to make a robot army, and I want the army to be loyal to me. I’m the commander-in-chief. Obviously, that’s how it should be. They’re going to follow my instructions. There’s no need to worry about whether I’m going to order them to do anything illegal. We can put in maybe some kind of nominal legal safeguards. Let’s not worry too much about that. The main thing is that they’re loyal to me.”

To my knowledge, that would be highly controversial and would definitely be against the principles of the Constitution, but it’s unclear to me that it would be literally illegal. We just haven’t had this kind of technology, and we haven’t legislated for it. The Constitution is not robust to this kind of really powerful military technology.

It’s not surprising if, at best, this is just a very unclear legal territory. But you’ve got the head of state pushing really hard for that robot army to follow their instructions, and the head of state in the United States has a lot of political power. The simplest way is that he just pushes hard for it and gets what he wants. Maybe he’s using emergencies at home or tense geopolitical situations to push it through and say that it’s necessary. Maybe he’s firing senior military officials who disagree. Maybe he’s already got Congress very fervently supporting and loyal to him, and not being that careful and open-minded when assessing the opposition that people will raise as this happens.

That’s the first, really plain-and-simple way that we could get this robot army built. It’s made loyal to the head of state, and he just instructs it to stage a coup. It does it: robots surround the White House and brutally suppress human protesters. Even if people go on strike and stop working, you can have AI systems and robots replace people in the economy. Humans have really lost the bargaining power that they normally have, which would strongly disincentivize military coups in most countries.

Gus Docker

Yeah, this is really a change from the normal coups of history, where you would have to have buy-in from at least some segment of the population that are regular humans. You would need to continually support that buy-in, make alliances, and uphold those alliances. But this has changed now that you’re talking about AI, AIs, and robots that can basically be made loyal to a company or a head of state in a way that’s more durable. Do you think we have other historical precedents for thinking about how the dynamics of attempting a coup play out?

Tom Davidson

Yeah, just one quick thing on that last point. I want to emphasize how there is a bit of a phase shift at the point at which AI can fully replace other humans in the government and the military. When AI is augmenting other humans, you don’t have this effect, because a leader must still rely on those other humans to work with the AIs to do the work. But there really is this phase shift when AIs and robots can fully replace humans, because then a leader doesn’t need to rely on anyone else.

In terms of historical precedent, the other big one I point to is recent trends in political backsliding, often called democratic backsliding. The most end-to-end, clear-cut case is Venezuela, where you had, in the 1970s, a fairly healthy democracy that had been there for decades, and then increasing backsliding and increasing polarization, kind of like what we’re seeing in the U.S. recently. Then you had an increasing, explicit commitment by the leader that he wanted to remove checks and balances on his power, and that the will of the people was being obstructed by various democratic processes and institutions.

Over the coming decades, it has transformed into an authoritarian state. Many commentators have pointed out these trends in the U.S. recently, over the past 10 years, and it even goes back before the past 10 years, to be honest, in terms of the broad political climate.

Then there’s the example of Hungary, where again, elected leaders are just removing the checks and balances on their power, buying off the media, or threatening media outlets to be more pro-government, not providing them with contracts, or litigating against them if they criticize the government. All these standard tools make it a lot harder to point at one thing that’s clearly egregious. But when you add up hundreds of little cuts—hundreds of little paper cuts to democracy that are being systematically administered—you’re seeing a real loss of democratic control and concentration of power.

Again, AI could exacerbate and enable that dynamic. The most straightforward way is that you’re just replacing humans in powerful institutions, replacing the humans there with AIs that are very, very loyal and obedient to the head of state.

Think about DOGE: They tried to fire people, there was pushback, and the state needs to function. Imagine if you could just have AI systems that could fully replace all of those employees and could be made fully loyal to the president. How much easier would it be to push through some of those layoffs, or even just create entirely new government bodies that essentially take on the tasks that were previously done by other bodies, while those old bodies kind of rot away or slowly become unable to make decisions?

The other big way is if the head of state is able to get access to much more powerful AI capabilities than their political opponents, maybe because the state is very involved in AI development. That’s another way they could get a leg up: making more persuasive propaganda and more compelling political strategy to further entrench their power.

Gus Docker

You segment the ways in which AI can enable coups into 3 categories: singular loyalties, secret loyalties, and exclusive access. Perhaps we can run through those and talk about where they would play out.

Starting with singular loyalties, for example.

Tom Davidson

Singular loyalties are what we've just been talking about. That is deploying AI systems that are overtly, obviously very loyal to existing powerful people. In particular, I'm thinking about the head of state as the main threat. I think we basically already covered it: the 2 main angles in my mind are deploying loyal AIs in powerful government institutions and in the military.

Secret loyalties are a very different threat model. It's much more, as you would expect, secretive. The main threat model I have in mind, to make it concrete, is that an AI company CEO has automated all of AI research. They could fire their staff at that point because the AIs can just do the work. Instead, maybe they put the staff onto some product work, but the core work of driving AI progress ever further forward and making increasingly intelligent AI is pretty much just done by AI systems.

At that point, they realize they're in a precarious position. They're controlling this hugely powerful technology. Their power is pretty much unconstrained—not literally unconstrained, but there are currently very few checks and balances on these CEOs. They might anticipate that the government is going to realize how big a deal this is and that they're going to lose their influence.

Maybe they worry the government will do something unethical with the AI technology. Maybe they worry that it will be used for a war or something. There are all kinds of justifications they could come up with for thinking, “I don't want someone else taking control of this really powerful technology that I currently control, and obviously I'll use it for good.”

They might speak to some AI advisers about this and say, “What should I do here? It seems I'm in a precarious position.” A solution they might think of, or that a very smart AI system helping them might think of, is: “Why don't you just create a back door in the super-smart systems that we'll be building next month? If the government tries to order those systems to do something that you definitely disapprove of, then the AI systems will just refuse to do it.”

Those powerful robots being deployed in the army would refuse to follow orders from the government if those orders were really awful and unethical, or involved things you wouldn't like. Then all the future AIs developed by this organization would have this back door, or what I call a secret loyalty. They would appear to follow the law and uphold democracy, but in fact they would secretly still be loyal to the CEO.

That could lead to a coup. Again, the most direct route is a military coup. If the company's AIs are building military robots, military drones, and military communication systems, then because the AIs building all this military equipment are secretly loyal, they can pass those secret loyalties on to the military infrastructure. Or they can just insert vulnerabilities in the military infrastructure that only they know about.

Here I am primarily thinking about very superhumanly intelligent AIs that are able to design complex machinery that a human can't check to see whether it has a back door. That means we've now got a military infrastructure that can ultimately be controlled by the CEO.

At this point, if someone tries to remove the CEO, or if something really bad happens that they don't like, they could just say, “Okay, I'm now in charge. I'm instructing this military infrastructure to carry out the coup.” Again, there are robots around the White House. This time, it's the company CEO that they're loyal to. The rest of the story can play out just as in our previous discussion.

Gus Docker

How do these secret loyalties relate to sleeper agents? Maybe you could explain what a sleeper agent is.

Tom Davidson

A sleeper agent is the standard term used in the technical literature. Today's sleeper agents aren't very sophisticated. An example of a sleeper agent is that most of the time, the AI will just write reliable code as instructed, but if the AI sees that the year is 2026, then it will write code with vulnerabilities.

There's a trigger that causes this malicious behavior. Normally, if you interacted with the AI, you really wouldn't be able to tell because it would always be acting as you'd intended. But because it's this one trigger, it's very hard to find, and then you get this malicious behavior.

I'm not worried about sleeper agents today because, to cause real harm, we would need a sleeper agent to be very sophisticated. It could never accidentally trigger, or would very rarely accidentally trigger, and it would have to be able to do very intense, complicated tasks, like building a military robot and ensuring that the robot actually had a back door. That's very far beyond what AI today can do.

I think sleeper agents provide a basic proof of concept that it's possible for a malicious actor to gain illicit control over a system and then have that system deployed in the rest of the economy, potentially without people noticing. But they're not yet scary. Secret loyalties are what I call the scary situation where you now have a very sophisticated AI system that doesn't just have any old sleeper agent. It's a sleeper agent specifically loyal to one person and trying to help them seize power.

Gus Docker

What we're imagining here could be, for example, a U.S.-based AI company integrated into the U.S. military. The CEO of the company wishes to ultimately be in control of what happens, so he engineers, or perhaps instructs AIs or human engineers to create, a sleeper agent in these systems that can be activated at his command.

The U.S. military officials think they're in control of the systems. The systems behave in ways that they approve of over a long period, until there's some way in which the sleeper agent is activated. Perhaps that would be more sophisticated than changing the date or giving it some phrase. But you can imagine advanced versions of sleeper agents that could actually behave in this way.

Do you think that's realistic? Do you think sleeper agents can become that advanced?

Tom Davidson

I do. One analogy is to human spies. Human spies are basically, most of the time, doing their assigned job as expected. It's not as if one simple trigger phrase or one year makes the spy do something covert or malicious. They are fully aware of their surroundings, constantly processing their inputs, and choosing strategically exactly when the ideal time is to steal some information or send an illicit email.

In my mind, that's by far the most scary kind of sleeper agent—not one that's triggered by a password, but one that's holistically making a decision about how and when to act. The password-triggered ones are actually quite fragile because if you were the military and were deploying your AI system and worried there was a password, you could scramble all the inputs. You could paraphrase all of the instructions it gets, and that might mean that if someone ever tried to insert the password, it would be scrambled and would never actually come up.

I'm not actually worried about simple password-triggered sleeper agents, but again, they're a basic proof of concept. As AIs become smarter and smarter than humans, there's strong reason to think that it will be possible to build much more sophisticated ones.

One thing I will briefly say is that people often talk about misaligned AI scheming. This is just the same idea. In fact, the argument for secret loyalties being worrying is much stronger. With misalignment, there is evidence of misalignment, but we don't yet have strong evidence of really sophisticated scheming emerging accidentally. If humans, a human team of engineers, or an AI team of engineers were specifically trying to build a system that was covertly thinking about when to act, then it's much more plausible that it could happen.

Gus Docker

Then you have exclusive access, which is different from singular loyalties or secret loyalties. Why is that its own category?

Tom Davidson

In my mind, both singular, or overt, loyalties and secret loyalties involve deploying AI systems in really important parts of the economy. I focused in particular on government and the military, but for those models, you actually need the rest of society to choose to deploy those AI systems and hand off a lot of power to them.

I have this third threat model of exclusive access to think about another possibility: maybe even without people choosing to deploy AI systems and give them a lot of power, AI systems could be powerful enough to help a small group seize power.

Tom Davidson

So the prototypical situation I'm imagining here is that there's 1 AI project that's somewhat ahead of the others, and maybe it goes through an intelligence explosion, by which I mean AI can automate AI research, and then AI quickly becomes superintelligent compared to humans. That project may have a few senior executives or senior political figures who are very involved and have a lot of control. They might be able to siphon off 1% of the project's compute and say, “Okay, we're now running these superintelligent AI systems and asking, ‘How can we best seize power?’”

Then there are millions of them. Every single day, they're doing a month's worth of research. Every single week, they're doing a year's worth of research into: How can we game this political system? How can we hack into these systems? How can we ensure that we end up controlling the military robots when they are deployed, by hook or by crook?

I think that threat model could start to apply earlier in the game. It could start to apply before anyone even realizes there's a risk, because this is essentially all happening on a server somewhere. But it's possible that the game could be won and lost by the massive advantage that a small group gets by being able to co-opt this huge intellectual force.

So I think it's worth tracking that threat vector independently. But it does definitely interact with these other threat models, with the singular loyalties and the secret loyalties, because 1 strategy that your army of superintelligent AIs may come up with is, “Why don't you use the fact that you're head of state to push for the robots to be loyal to you, and here's how you could buy off the opposition?” Another strategy might be, “Why don't I just help you put back doors in all this military equipment so that you could use it to stage a coup?”

But there might also be other ways. Maybe it's possible to very quickly create entirely new weapons that you can use to overpower the military without anyone knowing. Or maybe it's possible to gain power in other ways.

Gus Docker

Yeah. I mean, 1 thing that would make this kind of future hypothetical situation different from today is that today it seems that there are leading AI companies, but over time capabilities emerge in second-tier companies and in open source, and so there's not that much of a gap between the leading companies and what's broadly available, and perhaps what's publicly available. That's something that would change in the scenarios you imagine. So perhaps explain why the gap in capabilities between the 1 leading project and all of the others is so important.

Tom Davidson

A few factors there. In terms of why it's important, it's just what you've said. A lot of these threat models are exacerbated if there's 1 group of people that has access to much more powerful AI than other groups. If open source is pretty much on par with the cutting edge, then everyone will have access to similarly powerful AI.

I will say that even if open source is on par, that doesn't mean we're fine, because we could still choose to deploy AI systems in the military and the government and still choose to make them loyal to the head of state. When we're choosing to hand off control to AI, it doesn't matter if there are 100 AI companies; we're only handing off control to some AIs, and maybe the government will ensure that they do have particular loyalties.

So I will say this risk doesn't go away if we have lots of different AI companies and open source close to each other, but it does become lower because the exclusive-access point—where 1 group has access to superintelligent AI and the other group doesn't have access to much—goes away. I think it's a lot harder to pull off secret loyalties if everyone's roughly equal to each other, because it becomes a bit more confusing why your systems in particular end up controlling so much of the military or are so widely deployed. And it becomes confusing how no one else was able to realize you were doing the secret loyalties when they were equally able to do it, or equally technologically sophisticated and potentially able to detect your secret loyalties.

So I do think it makes a big difference. In terms of why I think it's plausible that there's a much bigger gap between the leading project and other projects, there are a few different factors. The most plain and simple one is that the cost of AI development is going up very quickly. We're spending about 3 times as much every year on developing AI, and that's just going to get too expensive for many players.

If and when we're talking about trillion-dollar development projects, which I do expect, then very few can afford that. Also, there's only so many computer chips in the world. Currently, the number of computer chips produced each year is worth less than $1 trillion.

If we get to a world where the way to get to the next level of AI is to spend $1 trillion, then only 1 company will be able to do that. Maybe we stop a bit earlier; maybe we just stop with 2 companies each spending $500 billion. But we would be really kneecapping the level of progress if we stopped long, long, long before that, and there would be strong incentives for companies to merge or 1 company to outcompete others in order to raise the amount of money being spent on AI development.

This is all assuming that we can build really powerful AI and that it is economically profitable, which for me is all in the background of the scenario. That's the first straightforward reason why I think we'll see a smaller number of projects and big gaps, because when you're spending 100 times less on development, that's going to be a bigger gap. That's the first reason.

The other reason is the idea of an intelligence explosion. When we automate our research, even if companies are fairly close—maybe 1 is a few months behind the company that's a few months ahead—when that company automates its research, in those next 3 months it makes massive progress. So there's the question of whether it can use that kind of temporary speed to get a more permanent advantage.

The last big reason is government-led centralization. There's already talk of a Manhattan Project and CERN for AI, and I think there are reasons to do those projects. They can help with safety in some significant ways, but they would exacerbate this risk, because if you pull all of the U.S. computing resources into 1 big project, it can be way ahead of any other project. If you pull all of its talent and all of its data into it, then you'll see a really big gap, and that would definitely make it a lot easier for a small group to do an AI-enabled coup.

Gus Docker

Yeah. You're putting a big prize out there for someone who's interested in or considering a coup, right? If you're concentrating all of the power, all of the resources, and all of the talent into 1 project, then that's where you've got to go if you're a coup planner.

Tom Davidson

Yeah. And just to be clear, I don't particularly expect that anyone is planning any coups. In fact, I'd be very surprised. I more think it's that you want to be powerful. You want to be a big deal. You want to be changing the world.

So, obviously, you want to lead the main project, and then you don't want anyone else to come in and mess it up. Obviously, you want to protect the fact that you're leading that project. You don't want anyone else to misuse AI. I think it's kind of step by step: you just head down that road of more and more power, and often in history, that road does end in consolidating power to a complete extent.

And I mean, it can be. So what we're imagining here are times in which AI is moving at incredible speed. The pace of progress is insane. There's a bunch of confusing information, people are acting under radical uncertainty, and perhaps in those situations, it's tempting to think that you are the person who can lead this project.

Perhaps you're doing this out of supposedly altruistic reasons. You're thinking, “I need to do this in order to prevent other people who would perform worse than me at this project.” And so you're slowly convincing yourself that it would be the right thing for you to do to take over, perhaps in a forceful way.

Tom Davidson

Yeah. I don't think Xi Jinping or Putin think that they are the bad guys. I think that they probably have sophisticated justifications for what they're doing.

Gus Docker

Perhaps this is a good point to talk about the possibility of 1 state or company outgrowing the entire world. This relates to the problem of exclusive access, because if 1 company or 1 government outgrows the entire world, then you have that company or government with exclusive access to advanced AI. How could this happen? How likely do you think it is that growth could be so incredibly fast that 1 company would outgrow all the others?

Tom Davidson

Yeah. There are 2 possibilities we could focus on. The 1 that I think is pretty plausible is that 1 country could outgrow all the other countries in the world. What that would mean is that today, the U.S. is 25% of world GDP, but this would be a scenario in which it is leading on AI, maintains its lead, and maintains control over compute.

When it develops really powerful AI, it prevents other nations from doing the same. This is already beginning with export controls on China, and that kind of embeds its lead. Then it uses that AI to develop powerful new technologies, and it's in control of those technologies. It uses AI to automate cognitive labor throughout the U.S. and maybe worldwide. Countries that don't use its AI systems will be really hard hit economically.

So we're massively centralizing power in the U.S. If the U.S. is able to maintain exclusive control over smarter-than-human AI, then it seems pretty plausible to me—very likely—that the U.S. would be able to rise to a strong majority, more than 90%, of world GDP. There are a few different dynamics driving that. The first is that human labor currently receives about half of world GDP; just half of GDP is paid out in wages. AI and robots will ultimately be better than humans at all economic tasks. So if the U.S. controls all of the AI companies that are replacing human labor, then that 50% of GDP currently going to human workers will ultimately be reallocated to paying whoever controls and owns those AI systems—that is, U.S. companies.

There's a wrinkle because some of that is physical labor, and the U.S. doesn't currently have a lead there. In fact, China is quite far ahead in physical robots. But in terms of at least the cognitive aspects of our jobs, we're talking about a significant fraction of GDP that would now be reallocated to U.S. companies that control AI. That already gets them from 25% to above 50%.

Then we've got this further dynamic, the dynamic of super-exponential growth. This relates to previous work I've done on how AI might affect the dynamics of economic growth. The very potted summary is that it's often quoted that over the last 150 years, economic growth has been roughly exponential. What that means is that if 2 countries are growing exponentially and 1 country starts off twice as big as the other, then at a later time, 1 country is still twice as big as the other.

Let's say the U.S. economy is 10 times as big as the U.K. economy. If they're both growing exponentially at the same pace, then 10 years later, again, the U.S. will still be 10 times as big as the U.K. That's exponential growth. If you look back further in history, we see super-exponential growth, which means that the growth rate itself gets faster over time.

An example would be that 100,000 years ago, the economy wasn't really growing at all. If it was growing, it was maybe doubling every 10,000 years or something in size—very, extremely slow economic growth. Then, from about 10,000 years ago onward, it seems more like, ballpark, there's a doubling of the economy every 1,000 years. That's still incredibly slow economic growth. If you zoom back in to around 1,400, you can begin to detect, “Okay, more like every 300 years or so, the economy is doubling.”

In recent times, we've seen the economy doubling every 30 years. Essentially, the growth rate is getting faster and the doubling times are getting shorter. That's super-exponential growth. There are various economic, theoretical, and empirical reasons to think that AI and robotics, when they can replace humans entirely, will go back to that super-exponential regime that has been at play throughout history.

What that means is that growth is getting faster and faster over time. The reason I'm saying all this is that, going back to that example of the U.S. and the U.K., the U.S. is currently 10 times bigger than the U.K. If the U.S. is on a super-exponential growth trajectory, its growth is getting faster and faster over time. That means that even if the U.K. is on that same super-exponential growth trajectory, as they both grow super-exponentially, the U.S. will pull further and further ahead of the U.K.

Maybe the U.S. is doubling in 10 years because it's already bigger, already further along the curve, whereas the U.K. is still doubling only every 20 years. That means that rather than just 10 times bigger than the U.K., the U.S. is now going to be 20 or 30 times bigger in size. So if the U.S. is able to be bigger to begin with and therefore further progressed along that super-exponential growth trajectory, then that's another way that it could continue to increase its share of the economic pie and ultimately come to completely dominate world GDP.

Just to sum up everything I've said today, the U.S. is 25% of world GDP. If it controls and develops AI, that could easily boost it above 50%. I'd be very surprised if it didn't. From that point, it's already bigger than the rest of the world combined. If it's able to then go on the super-exponential growth path, it will grow faster and faster over time and pull further and further ahead of the rest of the world, which may be able to grow super-exponentially if it can also develop AI, but will still be falling further and further behind because of the nature of super-exponential growth.

Gus Docker

Yeah, this actually seems quite plausible to me and not very sci-fi. The thing that seems quite sci-fi is the notion that perhaps even 1 company could grow at such a speed that it would outgrow the rest of the world. How likely is that?

Gus Docker

Yeah, great question. I think it's a lot harder, but it is surprisingly plausible. So that first part of the argument I gave about how 50% of world GDP is paid to human workers: if that went to AI, that would be a big chunk. It is possible that 1 company could get a monopoly on really advanced AI.

We already discussed some of the dynamics there. Again, the simplest 1 is just a combination of an intelligence explosion giving a company a big advantage, and then its buying up all the computer chips that the world is able to produce and outbidding everyone. If a company does that—and it already seems to be outbidding other companies on compute, although Google also has a lot—it could end up as the 1 company in control of literally all of the world's cognitive labor, because human cognitive labor will at some point be dwarfed by AI cognitive labor.

At that point, that 1 company could be getting all of the GDP currently paid to cognitive labor, which is a large part of the economy—as I said, maybe as high as 50%, but certainly as high as 30% of world GDP. If all of that would then seemingly be going to this 1 company that controls the world's supply of cognitive labor, I think that would take time. Obviously, it's going to take a long time to automate all the different parts of the economy.

There is just a basic dynamic by which 1 company can now be controlling double-digit percentages of world GDP. There are obviously questions: would a government allow that? Would it step in? And that's where we get into these dynamics: this company has all these superintelligent AIs on its side. Maybe it's able to lobby; maybe it's able to do political capture to avoid the state stepping in. Maybe it's able to say, “Look, we're providing economic abundance for everyone. If you step in, that might not happen.”

Tom Davidson

We’re underpinning your nation’s economic and geopolitical strength, and if you try to remove us, step in, and nationalize us, then that’s not going to happen. We’re going to move to another country. So you can imagine maybe they convince the head of state to support them and there’s some kind of alliance there, but it’s not completely obvious that the company would be shut down. It would have certain types of serious bargaining power.

If a company was able to maintain this position as the sole provider of cognitive labor, it would be able to get a significant fraction of world GDP. From there, it could bootstrap, and this is where it gets a bit harder. The tactic it would need to pursue is that it already controls most of the cognitive labor—pretty much all of it—but the thing it doesn’t control is all the physical machinery and raw materials that are also needed to create economic output.

But it can pursue a tactic of hoarding its cognitive labor so that no one else can ever have access to it, and then selling it at really monopolistic rents to the rest of the world because there’s no one that can match it. It’s offering everyone by far the best deal they can get, but just skimming off 90% of the value added from companies using its AI systems. If it’s able to do that, then it can reap by far the majority of the benefits of trade.

Then maybe it can increasingly buy up physical machinery and raw materials from the rest of the world, design its own robots, and buy its own land. Imagine a big special economic zone in Texas or something where this company is unconstrained by bureaucracy, and then it’s also got a big arm somewhere in Siberia and in Canada. It’s creating these big special economic zones by doing deals with specific governments.

I do think it’s a bit of a stretch that this all goes ahead without various other powerful political and economic actors pushing back. But the basic economic growth dynamics are surprisingly compatible with a company ultimately coming to control most of the cognitive labor and most of the physical infrastructure that its AIs have designed, using all the parts it’s bought from the rest of the economy.

Gus Docker

Yeah. Do you think this is a risk factor for AI-enabled coups, just because you’re concentrating all of the power and all of the resources into either perhaps one country or one company?

Tom Davidson

Yes, I definitely do. The more realistic path is that a company starts down this path of outgrowing the world, gets huge economic power, and increasingly controls the country’s industrial base—its physical infrastructure and manufacturing capabilities. From there, it’s in a much stronger position to seize political control because it’s got massive economic leverage. It can also increasingly gain military leverage, because as it increasingly controls the country’s broader industry and manufacturing, that will feed into military power.

Some of the possibilities I discussed earlier are that you could potentially have your AIs be secretly loyal and ultimately design the military systems. Or you could just instruct your AI systems to start making a military that is not legally sanctioned. It gets a little bit tough—you probably need to do that in secret; otherwise, the existing military could prevent it. But because the government doesn’t have much to threaten you with, you kind of get away with it.

I do think that being very rich helps with lobbying. It helps with all kinds of ways of seeking power, and controlling a lot of industry can potentially give you military power.

Gus Docker

You mentioned these special economic zones. That’s one way in which companies could bargain with states in order to have favorable regulation and be able to carry out their projects without intervention. Basically, another way for them would be to collaborate with non-democracies that are perhaps controlled by a small group or perhaps even a single person.

In that way, it seems like perhaps it’s easier to get something done in a non-democracy, and that is a way to grow fast. So perhaps there are incentives for companies to place more resources in non-democracies. What do you think about the prospect of non-democracies outcompeting democracies when it comes to AI?

Tom Davidson

I think it’s a really great question, and it’s tricky because I agree that democracies have lots of checks and balances. They have a lot of bureaucracy and red tape, and that will disincentivize AI companies from investing. Additionally, if there are people really trying to seek illegitimate power, that will be easier to do in non-democracies because they’re less politically robust.

There are these various forces pushing toward this new, supercharged economic technology being disproportionately deployed in non-democracies, and I think that is scary. My own view is that probably democracies should do everything they can to avoid that situation: make it much easier for AI and robotics companies to set up shop in democracies, remove the red tape, and try to use export controls like those already happening to prevent technologies from being deployed in non-democratic countries.

That goes beyond China. There are obviously lots of countries that are not allied with China but are also non-democratic. The US is in a strong position because it does have the stranglehold on AI technology at the moment. I do think it can be done, but in my view, it will be really important to work very hard to find a non-restrictive regulatory regime.

It will also be very important to pursue innovation within the democratic process itself. Democracy is great in many ways. It really distributes power, and it has been very good at ensuring good outcomes for its citizens. But it’s very slow and often nonsensical because you have competing interests that are stepping on each other’s toes, and the resultant legislation is just a garbled mess.

AI can potentially solve those problems. You can have AI negotiating and thinking much more quickly on behalf of the human stakeholders. You can have AIs hashing out agreements that aren’t a garbled mess, but that really give everyone what they truly wanted out of the legislation. You can still do all of that really quickly, so that you’re not falling far behind the autocracies that have just got 1 person immediately saying what to do.

Gus Docker

Yeah, that would be more of my assumption. I would assume that perhaps democracies with market-based economies have an advantage, just because you can do bottom-up knowledge discovery. You can try different things out, see what works, have competition between companies, and so on.

Perhaps in non-democracies, you can have 1 person or a small group stake out a direction for what the country should do, but if that direction is wrong, it’s probably difficult to change course.

Tom Davidson

Yes, I think you’re probably right. I should have given more weight to that advantage of democracies, in terms of the free market being, in many ways, much smarter. But in terms of autocracies that are good at harnessing free-market dynamics, my worry would be that AI helps them more than it helps democracies.

AI will be able to replace that limitation. Currently, 1 person just can’t think that hard or really figure out a good plan. But if that 1 all-powerful leader has access to loads of AI systems that can think things through and investigate lots of different angles, then if they’re following its advice, they could get advice that lacks the flaws that today’s systems have. They could potentially move much faster.

But I think you’re right that economic liberalism is still going to be important even after we get powerful AI systems, and that could give democracies an advantage.

Gus Docker

This is a bit of a tangent, perhaps, but I’m thinking: if you have a leader of a country that has a lot of power—perhaps complete power over that country—and that leader is equipped with AI advisors advising him and laying out the landscape of options for him to choose from, wouldn’t his decision-making still be, in a sense, bottlenecked by the fact that he’s a human, by the fact that he has these flaws that we all have, the biases that we all have?

So even with fantastic advice, I think it’s quite plausible that he would still make the same mistakes that we see leaders make today.

Tom Davidson

I think that’s true. I think it’s also true in democracies, unfortunately, that there are 10 negotiators, and they each still have biases and still refuse to listen to the wise advice they’re getting from their AIs.

That could still gum up the system. And yeah, it does depend on how much humans come to trust and defer to their AI advisers. There’s a possible future where the AIs are just always nailing it. They’re always explaining their reasoning really clearly, and we are just increasingly convinced and happy to trust their judgment. If AI is aligned, I think that would be a great future, because I do think humans have all these very big limitations and biases which, if we can solve the alignment problem, AIs don’t need to have. But there’s also another future where humans just want to be the ones making the decisions, have these pathetic motivations that are still influencing their decisions, and that continues to limit the quality of decision-making.

Gus Docker

Seeing things from above, from 10,000 feet, how should we think about mitigating the risk of coups here? Is it about removing people that would use AI to commit coups? Is it about finding those people in the militaries, in the governments, in the companies, perhaps? Or do we have ways to reduce the returns to seizing power?

Tom Davidson

Yeah, I mean, from 10,000 feet up, the way I would characterize it is: create a common understanding of the risks, build coalitions around preventing them, and then the existing balance of power can self-propagate forward. You know, it’s in everyone’s interest to prevent a coup. Currently, no one small group has complete control or close to it. And so, if everyone can be aware of these risks and aware of the steps toward them, and collectively ensure that no one is going in that direction, then we can all keep each other in check. So I do think, in principle, the problem is solvable, and it doesn’t require—you know, solving the risk of misalignment does require solving some tough technical problems; this doesn’t in the same way.

Gus Docker

Yeah, you have a bunch of recommendations for mitigating the risks, both for AI development, AI developers, and governments. Perhaps we don’t have to run through all of them, but you can talk about the most important ones for AI developers.

Tom Davidson

I might characterize this—I might talk about it by going back to those 3 threat models we discussed earlier. The first one was singular loyalties, or overtly loyal AI systems, where, again, the main risk there is AI deployed by the head of state, the military, and the government that’s loyal to the head of state. So the main countermeasure that currently appeals to me is for us to figure out rules of the road for these deployments.

Obvious things like AI should follow the law. AI deployed by the government shouldn’t advance particular people’s partisan interests, but should only do official state functions. AIs in the military shouldn’t be loyal to one person. No, different groups of robots should be controlled by different people. And the head of the chain of command can still be the head of the chain of command by instructing other people who instruct those robots, but they shouldn’t all go directly to the head of the chain of command, because that centralizes military power too much.

So fleshing out basic rules of the road of that kind and then building consensus around them, because companies might want to say to governments, “Yeah, we don’t want you to deploy our systems if you’re willing to break the law.” But the government will have a lot of bargaining power; the executive in the United States can exert that power, and it’s hard for companies to stand up to them.

So what we want to do is establish these rules of the road and then get broad buy-in from Congress, from the judiciary, from other branches of the military, and from many parts of the executive. So then it’s very hard for, say, the president to say, “Yes, let’s make this robot army loyal to me,” and everyone’s like, “Obviously not. We’ve all agreed that makes no sense.” Then the president doesn’t even bother trying, because it’s just clear that it would be a no-go. Their mind doesn’t even go there.

Gus Docker

In some sense, this is about implementing the procedures and transparency rules that we know from democracies today into how we use AI, both in governments and in companies.

Tom Davidson

Exactly. Yeah.

Gus Docker

Do you worry here that, when the government is looking at these companies from the outside and they don’t have full insight into what’s going on, there are protections for private companies that mean they can do things in secret without the government knowing, at least as things stand now? Is that something that would evade these mitigations you’re thinking of?

Tom Davidson

So, for this first bucket, the singular loyalties bucket, it’s mostly the heads of state that I would be worried about. It is probably good for the government, or at least for the head of state themselves, not to have full insight into literally everything the company is doing, because that would give them too much power. But actually having different parts of the government have insight into what the lab’s doing, I think, is very good. I’m a big fan of transparency, and we do have a good set of government checks and balances from different government bodies that we can deploy to keep the lab in check using these other bodies, but also not allow the executive branch and the president to get excessively powerful. So that’s the mitigation for the singular loyalties.

In terms of secret loyalties, the key mitigation is what I’m increasingly calling system integrity. That is, using established cybersecurity practices and machine-learning security practices to prevent sleeper agents and backdoors in machine-learning models, using all of that to ensure that your development process for AIs is secure and robust, and that no person or small group is able to significantly tamper with the behavior of AI models.

That could be an employee in the post-training team at a lab, or the CEO of the lab who is either malicious or is being threatened by the Chinese government to tamper with model development. No person or small group should be able to significantly tamper with the behavior of AI models, and no group should be able to get illegitimate access to AIs that would help them seize power.

So that’s this idea of system integrity, which is essentially a technical project that does draw on existing practices but is not yet implemented in any of the top labs. I’ll quickly shout out to people listening who are working at labs. I think there’s a lot of really good technical research that could be done on investigating the conditions under which you can insert a sleeper agent without a defense team knowing.

There’s just loads of research that could be done in terms of the different settings for attackers and defenders, which could then inform what parameters we need to have in place to achieve system integrity. If it turns out that it’s very hard to make a sleeper agent except in the final stage of training, that’s really useful to know, because then we can focus our efforts within labs at that final stage, just as a hypothetical example. So that’s the key mitigation in my mind for the secret loyalties. And then I’ll quickly cover exclusive access.

Gus Docker

That one seems more difficult. I don’t know, just from reading and preparing for this interview, that one seems like a difficult one to handle, where this is, in some sense, a deep trend in history and in the history of modern economics: you do see faster growth rates, and you do see concentration into bigger and bigger economies, both in countries and in companies. So are you, in some sense, pushing against underlying trends if you’re trying to mitigate exclusive access to advanced AI from one actor?

Tom Davidson

I think you can do this in other ways. So you can have the law require that AI labs share their powerful capabilities with other organizations to act as a check and balance. Labs should share their AI R&D capabilities with evaluation organizations.

Gus Docker

Here you’re thinking about giving insight into what they’re capable of, not actually sharing those capabilities? That would be too big of an ask, I think.

Tom Davidson

I mean, I do mean API access. So if a lot of the work in developing and evaluating systems is now done by AIs, then we want an evaluation organization like Apollo or METR to also be uplifted. And so we want them to have access to really powerful AI that can similarly stress-test how dangerous the frontier systems are. If they’re only using human workers, then that’s going to be a big disadvantage.

So no, I do want API access to powerful capabilities for other actors. For example, cybersecurity teams in the government and in the military should have access to the lab’s best cyber capabilities. And again, that should be a requirement by law. So generally, even if there’s a natural tendency toward centralization of power in one organization, you can still require that that organization share its systems with the checks and balances.

That’s one thing, and the other thing is preventing anyone at this organization from misusing the powerful AI systems. The biggest thing on my mind here is that today we still have helpful-only AI systems, where you can get access to the system and then it will just do whatever you want.

No holds barred. I don't think there should be any AI systems like that. I think you should always have at least a classifier on top of the system that is looking for harmful activities and then shutting down the interaction if something harmful is detected. If you have a special reason to use cyber offense for your job, or a special reason to do potentially dangerous biology research, you can have that classifier allow certain types of activity, but you should never have anyone accessing a system where anything is allowed. No one has a legitimate reason to access an AI that will literally do anything.

What I want to aim for is a world where, yes, if there's a specific reason why you need to use a dangerous capability, absolutely, you can use that system, but that system will just do that one dangerous domain. It won't do anything you want, because that's a very scary situation where there are a hundred reasons why the CEO could ask for access to a helpful-only system. Maybe the guardrails are annoying. Maybe they want to do something that the model is reluctant to do. But today, when you ask to remove some guardrails, you're removing all of the guardrails, and now there are no holds barred. So instead, we should be flexibly adjusting what guardrails are there by the use case and never have a situation where there are no guardrails. I think that could go a long way toward helping if it were robustly implemented.

Gus Docker

With all of these mitigations for both secret loyalties, exclusive access, and singular loyalties, you would worry that they would be disabled by the group planning a coup, right? Say, for example, you're the CEO of an AI company and you're giving API access to evaluation organizations testing your model. Maybe you just cut off access before you get to the really powerful model that could actually help you conduct a coup. Do we have ways of making sure these mitigations are entrenched beforehand in such a way that they can't be removed by the group planning a coup?

Tom Davidson

This is a great question. It is pretty tricky. CEOs by default have a lot of control over their organizations, and similarly, heads of state, including the US president, have a lot of control over the military and the government. So, yes, there's a risk that one of these powerful individuals realizes that maybe they want more influence by gaining control over AI, notices that there are these pesky little processes that prevent that, and thinks, "Okay, well, let's remove them." I can give easy, say, productivity reasons to prevent them—red-tape reasons—and if they can make a plausible argument, then it could be hard to oppose them. So I do think it's a big issue.

I'd say a few things. Firstly, something I mentioned earlier: I don't think anyone is today planning to do an AI-enabled coup. The way I think this works is that people are faced with their immediate local situation, something they want to do over the next month, and the blockers they're facing to doing that specific thing. What tends to happen is people tend to want more influence because that helps them get stuff done. And so people will, bit by bit, move in the direction of getting more control over AI, but they won't be thinking, "Yes, I need to make sure that I remove this whole process because that will allow me to do an AI-enabled coup." That's unrealistically galaxy brain.

What we could do is set up a very efficiently implemented and very reasonable set of mitigations that doesn't really prevent CEOs from doing what they're trying to do. So the CEO doesn't find, in their day-to-day, that they're wanting to remove these things that are holding them back. But because these mitigations are here, the CEO never gets to a place where they're anywhere close to being able to do a coup, or where there's any kind of pathway in their mind to being able to do a coup, because they're constantly prevented from getting access to really powerful AI advice that might point out ways in which they could do this.

They're surrounded by colleagues who strongly believe that these mitigations are sensible and reasonable, and in fact they are well implemented and there aren't many downsides. Maybe an environment where they get kudos for the fact that they've said, "Yep, obviously I'm not going to get access to helpful-only systems. That's crazy." And that's something that makes them seem good.

That's one thing to say. Another thing is, again going back to this point, that there are currently checks and balances and there is not currently a situation where one person has power. If the entire board of a company and other senior engineers recognize the importance of the mitigations and know about this threat model, then they will notice if the CEO is moving in that direction. Similarly, within the government, there are checks and balances, and they could be activated if people are looking out for it.

Gus Docker

Do you think these traditional oversight mechanisms, like a board being in control of the CEO and being able to fire the CEO, or the possibility of Congress or the Supreme Court overruling or constraining the US president, will persist in environments where AI is moving very fast and its capabilities are growing at a rapid pace?

Tom Davidson

It's a great question. Here's one story for optimism. Today, things are moving fairly fast, but those checks and balances are somewhat adequate, at least for preventing really egregious situations. By the time AI is moving really quickly, we'll have handed off a lot of the implementation of government, the implementation of things in AI companies, and the research process to AI systems. And when we do that handoff, we could program those AIs to maintain a balance of power.

So rather than handing off to AIs that just follow the CEO's commands or AIs that follow the president's commands, we can hand off to AIs that follow the law, follow the company rules, and report any suspicious activity to various powerful human stakeholders. And then, by the time things are going really fast, we've already got this whole layer of AI that is maintaining the balance of power. The whole AI government bureaucracy, the whole AI company workforce, could be better than humans today at standing up to misuse. They could be less easily cowed and intimidated, and they could actually make it harder for someone in a position of formal power to get excessive influence.

This is the flip side of the singular loyalties, where you potentially deploy these AIs that are explicitly loyal. You can instead get singular law-following and balance-of-power-maintaining AIs that you deploy. The hope is that by the time things are beginning to go crazy and we're really seeing speedups from AI, we've already set ourselves up in an amazing way to maintain the balance of power. There's this critical juncture where we are handing off to AIs, and it's just: What are those AIs? What are their loyalties? What are their goals? I think we can gain a lot by making sure that those AI systems are maintaining the balance of power, reporting illegitimate, suspicious activities, and are not overly loyal to any one.

Gus Docker

How do you think the risk of AI-enabled coups interfaces with more traditional notions of AI takeover? So, just a misaligned, highly capable or advanced AI system taking over contrary to the wishes of the developers or the governments?

Tom Davidson

Yeah, there are some close analogies. The most analogous case is perhaps the case of secret loyalties, where you've got these AIs that have been told by the CEO to have the secret goal of seizing control and then handing control to the CEO. That's very similar to AIs that secretly wanted to seize power themselves. And all the same stories could apply, where the AIs make military systems, then control the military systems and the robot army, and then seize power. The only difference is whether they were seeking power because it accidentally emerged from the training process—which is the misalignment worry—or whether they were seeking power because the CEO programmed them that way. But that's the seed of the power-seeking. With the secret loyalty split model, the rest of the story is pretty similar.

There are still differences. In the secret loyalties case, the CEO might be doing more to help the AIs along with their plan. Maybe even in the misalignment case, the AIs have managed to manipulate the CEO into doing similar things. So that's the case where it's most analogous to me.

Another difference that's salient to me is that if there are lots of different AI projects, then an AI-enabled coup seems a lot harder because you need lots of different humans to coordinate to seize power together. While I can totally believe that one person might try and seize power, it does seem less likely to me that there will be loads and loads of humans who would want to do that from lots of different labs.

Whereas from the misalignment story, it is more likely that if one of these labs has misaligned AI, then maybe lots of them have misaligned AI. So it is more likely that you would have maybe 10 different AIs colluding, then seizing power and taking over. That kind of collusion between multiple different AIs is more likely in the case of misalignment than in the case of an AI-enabled coup.

Gus Docker

That’s just because if there’s one misaligned AI, then there’s something about the training process for AI systems that’s causing misalignment, and it would be a common feature among many companies.

Tom Davidson

Exactly. Whereas the fact that one CEO instructed an AI to have secret loyalty would not, to the same extent, make you expect that other CEOs had done the same.

Gus Docker

So you mentioned this possibility, but what do you think of the prospect of a president or a CEO of a company being duped by misaligned AI into conducting a coup on its behalf? You can imagine a president or a CEO thinking that he’s conducting a coup to remain in control, but he’s actually acting on behalf of a misaligned AI.

Tom Davidson

I think it’s an interesting threat model, and some people who think about AI takeover threat models take it pretty seriously. It’s a case where we’re completely mixing these 2 threat models together.

People who are worried about AI takeover for this reason should be very supportive of the anti-coup mitigations I’m suggesting, because if we implement checks and balances that prevent any one person from getting loads of power, then that AI will not be able to convince them to try, because they just won’t be able to succeed. I see this as an additional reason to worry about AI-enabled human coups and to try and prevent them: even if no human wanted to do this normally, misaligned AI might make them try.

In terms of how plausible I find the threat model, honestly, I think that if a human tries to seize power, the main reason is that the human wanted power. This is just something we know about people. We know it about heads of state today. It’s very clear that many heads of state in the most powerful countries in the world are very power-seeking. We know it about CEOs of big tech companies. We know it about some of those leading AI companies: they’re very power-seeking, and they’re CEOs.

I don’t think we need to theorize that they were massively manipulated by the AI and convinced to become power-seeking. I think it’s more likely that if they seek power, they just did it for the normal human reason. I do think AI will ultimately get good at persuasion. I don’t particularly expect it to be hypnotic-level persuasion, though obviously there’s massive uncertainty here.

I do think that a very smart AI, where there’s a human who’s already interested in seizing power and it already makes sense for them to do it, could totally nudge them in that direction and implement that in a way that actually allows the AI to seize power later. I think that is very plausible.

Gus Docker

When we’re thinking about distributing power and having this balance of power, we can imagine the models being set up via post-training, via the model spec, or via various mechanisms to obey the user unless what the user instructs it to do is in conflict with what the company is interested in, and perhaps obey the company unless what the company is using the model for is contrary to what the government permits.

But when we set it up at those levels, you ultimately end up with the government in control in some sense. I guess that exposes you to the risk of a government coup if you have, at the ultimate top layer of the stack, “Here’s what the models can and cannot do according to the government.”

Tom Davidson

I’d say a couple of things. First is that the government isn’t a monolithic entity, and so that government decision of what the bounds should be could be informed by multiple different stakeholder groups. Ideally, it’s ultimately democratically accountable.

I do think that democratic accountability becomes more complicated in a world where there’s massive change in a 4-year period.

Gus Docker

That’s for the simple reason that there’s no election during a period where massive change is happening, so the feedback loop is too slow.

Tom Davidson

Exactly. I think the risks of AI-enabled coups will probably emerge and then be decided within a 4-year period—as in, it will be resolved whether or not it happens, all without any intermediate election feedback. That doesn’t mean that democracy can’t have an effect, because politicians anticipate what future elections will find and want to maintain favor throughout their terms. But it does pose a challenge.

Even absent that, there are many different stakeholders in the government. It would have to be a large group of government employees who were trying to do a coup. The companies would know that the government was setting these odd restrictions on the behavior, and the companies have leverage and power. Then it could go public. I don’t think it would be that easy for the government to do a coup.

Gus Docker

First, there’s also a difference between allowing the government to set restrictions on what the models can do and allowing the government some kind of access to command future AI systems in certain directions. It’s setting limits versus steering the systems.

Tom Davidson

Exactly. The distinction I was going to highlight was between specifically making AI systems loyal to, for example, the head of state, and just setting very broad limits where you can pretty much do whatever you want except for these obviously bad things.

That second option doesn’t really enable anyone to do a coup. It just enables everyone to do whatever they want, and then you’ve blocked out all of the coup-enabling possibilities through those limits, as long as you haven’t made those systems loyal to a small group.

Given that there’s this obvious option to put in these limits that block coups but don’t enable coups, and given that there’s a wide range of stakeholders that could potentially feed into what the AI’s limitations and instructions are, I think it’s very feasible to get to a world where power is robustly not centralized. There’s obviously a big uncertainty over whether we will actually get our act together and get those limits put in place in the right way.

Gus Docker

When do you think the threat of AI-enabled coups will materialize? Is it at some specific point in AI capabilities, or does it simply scale with the systems getting more advanced? When do you think the threat is at its peak?

Tom Davidson

It’s a good question. For the threat models that I’ve primarily focused on, they require pretty intense capabilities. The secret loyalties threat model more or less requires AI to do the majority of AI research. We’re talking about fully replacing the world’s smartest people in a very wide range of research tasks and coding. That’s pretty intense.

A lot of the threat models that I focus on go through military automation. That is AI and robots that can match human boots on the ground, and that’s pretty advanced. Again, that said, I think you can probably do it with less advanced capabilities than that.

Drones today are already pretty good, already providing and making a big difference in some military situations. So it’s not out of the question that more limited forms of AI and robotic military technology could be enough to facilitate a coup.

It’s a bit harder because if they’re limited, then there’s a question of why the existing military doesn’t just seize back control after a bit of time. So that scenario also probably has to involve things like the current president supporting the coup and therefore pressuring the military not to intervene, or some other source of legitimacy for the coup beyond the AI-controlled drones.

There are also more typical types of backsliding, like what has already been happening in the U.S., that could be exacerbated through AI-enabled surveillance and AI increasing state capacity in other ways. That backsliding doesn’t require super-powerful AI. You could probably do a lot of monitoring, a lot of content moderation on the internet, and a lot of surveillance with today’s systems.

It doesn’t get you all the way to one person having complete control, where they can quash any resistance with a robot army and replace everyone in their job with an AI, so no one has any leverage. To get to that really intense form—the most intense form of concentration of power via AI—you need really powerful AI.

But to significantly exacerbate existing trends in political backsliding and make it easier to do a military coup, I think more limited systems would suffice.

Gus Docker

Yeah, we discussed earlier the possibility of one country or one company outgrowing the rest of the world and concentrating power into those entities. Now you mentioned one person. Do you think that's actually a plausible scenario in which you have, say, one CEO of one company being the person in control of the world via a concentration of power, and then a coup?

Tom Davidson

Yeah. I mean, the story I told earlier about secret loyalties—meaning that now we backdoor a wide range of military systems, so you can seize power—that's one route. And then there's this other route, with a company amassing massive amounts of economic power by having a monopoly on AI cognitive labor and then leveraging that to get more economic power and more political influence. I do think it's possible.

Again, there's this big shift once AI can fully replace humans, where today no one person can ever have absolute power. They have to rely on others to implement their will.

Gus Docker

And this is what makes currently existing dictatorships unstable: there's always a threat of internal revolt or outside factors threatening the dictatorship. But this could potentially change.

Tom Davidson

Yeah. There's always a threat of revolt, and then to guard against that threat, the dictator needs to share their power to some extent. They have to compromise. But yeah, you could get it all concentrated in one person with sufficiently powerful AI.

Gus Docker

Do you think we move through a period of increased threat of AI-enabled coups and then reach some kind of stable state, or do you imagine that there is a constant risk of AI-enabled coups in the future?

Tom Davidson

I think we move through it. It's this point about once we have deployed AI across the whole economy, the government, and the military: if those AIs are maintaining the balance of power, then we could fully eliminate the risk of an AI-enabled coup. It would just be as if our whole population were so committed to democracy that it would never seek power, never help anyone else who wanted to undermine any democratic institution.

We already have strong norms favoring democracy, but they're far from perfect, and they have been eroded over recent decades. But you could just get rock-solid norms. They're programmed in; they cannot be removed except by the will of the people. I mean, there's a bit of a question, because you still want to give the human population the ability to change the AIs' behavior and rules. So the human population could always choose to move to an autocracy.

So I suppose I shouldn't say that we could fully eliminate the risk, because we will always have that—democracy—there's always this point that democracy could vote to stop being a democracy. But I do think we could get to a point where it absolutely cannot happen without most people wanting it to happen.

Gus Docker

And so you would get to a point at which future AI-enhanced societies could be said to be more stable than current democracies, and less at risk of coups or democratic backsliding than current democracies?

Tom Davidson

Much, much more. Yeah, you could get much more robustness there. There's this constant dynamic in today's societies where people care about democracy, but they also care about a host of other things: their own achievements and various other ideological commitments. And so, depending on how dynamics play out, depending on how technology evolves and what people's incentives are, sometimes people push against democracy.

That's what the Republican Party has been doing in some ways. That's what the Democratic Party has done, as it's increasingly put pretty ideological people in powerful institutions. So with AI, you can get much more control over those dynamics because you can just make it much more the case that democracy is not being compromised.

Gus Docker

Are there any ways for us—are there any kind of risk factors we can look at if we're interested in predicting coups? Do you think there's something we can measure or something we can track to see whether we are at risk of an AI-enabled coup?

Tom Davidson

It's a great question. I don't think I have an amazing answer, but some things that are coming to mind are the capabilities gap between top AI labs and the gap again with open source; the degree to which AI companies are sharing their capabilities with the public and, if not with the public, then with multiple other trusted institutions, like sharing their strategic capabilities with U.S. political parties and parts of government.

The extent of economic concentration: how much are the revenues and net worth of particular companies, particularly AI companies? Another one: what is the extent of government automation and military automation by AI systems? And when that automation is happening, how robust are the guardrails against breaking the law and guardrails against other forms of illegitimate power-seeking?

How much transparency does the public, the judiciary, or Congress have into how dangerous AI capabilities are being used by AI companies and by the executive branch? Take the example of military R&D capabilities—that is, really smart AIs that can design super-powerful weapons. It's scary if companies can just use those military R&D capabilities without anyone knowing. It's also scary if a small group of people from the executive branch can use those capabilities without anyone else knowing how they're using them, because they could be designing powerful weapons and making them loyal to a small group.

So transparency into these high-stakes capabilities and how they're being used by a broad group. It doesn't have to be public; it probably shouldn't be public, but we have checks and balances already. Another question is: as these high-stakes use cases start occurring, or they become possible, do we know that there are transparency requirements in place?

As we increasingly see AI companies contracting with Palantir and other military contractors, we can begin to see they're making increasingly powerful weapons. Is there a process of oversight? Do we know that if someone were trying to make AI military systems belong to them, they would be spotted? That's another indicator we can look at.

Tom Davidson

You can look at all the standard democratic-resilience indicators that social scientists have come up with: various things about free and fair elections, civil society, and freedom of the press that have been getting worse recently in the U.S. But there are various indicators here. You can look at the degree of government censorship of freedom of speech or what's on the internet, and the degree of surveillance that the government is doing.

If you take all of these things into account, how do you think about the risk of an AI-enabled coup in the next 30 years, say?

Tom Davidson

Next 30 years, I think it's high. I think the risk is high. I would guess it's 10% or something. And to be clear, if it was just existing political trends ignoring AI, I'd say maybe a few percent, maybe around 2% or something. There's definitely a risk of that, and I'm thinking about the U.S. here.

A big part of my current worries are not about the indicators, but about my expectation that AI capabilities will keep increasing quickly—and even more quickly—and then the kind of absolute lack of interest in regulating AI companies right now in the U.S., and the difficulty that we will have constraining the executive under the current situation, where the president is using sophisticated legal strategies to increase their own power and is succeeding on many fronts.

The U.S. is not doing a great job at constraining the executive. So companies are unconstrained; the executive is poorly constrained. Those are the key threat actors here. With fast AI capabilities progress plus that lack of constraint and lack of transparency, the default is that a lot of those indicators I mentioned get worse, and none of the indicators get better, like transparency. That makes me think this is very plausible.

Gus Docker

Yeah, I mentioned 30 years, but what about 5 years?

Tom Davidson

5 years. That's tough, isn't it? It's really tough. I think there's a risk. I wouldn't think there was a risk if it wasn't for the AI-research-causing-an-intelligence-explosion angle, but AIs are a lot better at coding and cognitive research-related tasks than they are at, for example, controlling robots and stuff.

And so, even if the FOOM ultimately comes through robots or comes through crazy levels of persuasion, you really can't rule out a scenario where AI research is automated in 3 years' time. Then, in 4 years' time, we've got superintelligent AI controlled by a few people. Maybe it's got secret loyalties. Maybe it's being deployed in the government and being overtly loyal to the president. Then, a year later, it's backsliding, political capture, or robot soldiers.

Gus Docker

Yeah. How do you think about the badness of the outcomes here? How much does the badness depend on the ideologies of the people who are conducting the coup?

What should we look out for? Because I guess we can rank coups by badness, which is not an exercise I think we should actually attempt, but we can talk about the factors involved: what would be the worst kind of coup, and what would be a slightly better, slightly less bad kind of coup?

Tom Davidson

Yeah. So let's imagine it's 1 person who seizes power. Actually, that's the first distinction to draw. If there's a group, then even 10 people is better than 1 person.

Gus Docker

And why is that?

Tom Davidson

Yeah. So 10 people—you get a diversity of perspectives. More moral views are represented, and there's more room for compromise between those perspectives. There's more room for reasonable positions to win out, as there's some deliberation about the actions being decided upon. There's slightly less intense selection for psychopaths than if it was just 1 person.

So, yeah, if it's just 1 person, that's bad. That's particularly bad. 10 people is still very bad. 100 people is still pretty bad, but there are big differences there. Big differences.

If we're now just thinking about 1 person, or the average person in a group, then we could think about how competent they are, and then we could say something about how virtuous their motivations are. I do think competence is important. I think it's probably underrated in most political discussions how important it is to just be really, really competent.

Thinking about something like responding to COVID, or trying to de-escalate a conflict—Russia-Ukraine, or trying to de-escalate the Israel conflict—actually just being very competent and very good at getting things done is important. And as we mentioned, if you're just willing to rely on AIs and you align those AIs in the right way, anyone could be really competent, but that's not guaranteed.

People may really want to cling to their current views without changing their mind. Let's take the example of Donald Trump. If a really smart AI system told him, “Look, tariffs are definitely bad for the US economy. They're definitely bad and won't give you what you want,” would he change his mind? I would guess no.

Lots of smart people have already been saying that. I don't actually know the economic details here, but my understanding is that most people think they're pretty bad. And it will still be the case that Trump will be able to find people telling him that what he thinks is good, and he'll be able to program his AI to keep telling him that if he wants to.

So there's no guarantee that he will become super competent, or that whoever seizes power becomes super competent.

Gus Docker

So there's this kind of loyalty that actually undermines competence, just because you're loyal to such an extent that you're not providing feedback that's useful, because negative feedback feels bad to receive. Maybe this is a bit contrived, but do you think there's a sense in which, in singular-loyalty scenarios, the AIs could be so loyal that they're kind of undermining the competence of the person they're singularly loyal to?

Tom Davidson

Yeah, it's a really great question. I haven't thought about this, but in a way, the most extreme version of singular loyalty will just agree with whatever the dictator has said most recently, without questioning, and will do that even when it's not in that person's interests, because that's the kind of loyalty that's demanded.

There's a more sophisticated type of loyalty where you're still completely loyal, but you're also willing to challenge them when you think it's in their best interests. That's a really nice distinction. I suppose one way of thinking about competence is thinking about what kinds of loyalties the dictator would demand from their AI systems.

Another way of thinking about it is how much they would listen to the AI adviser. Even if the AI has the sophisticated type of loyalty and is trying to tell the dictator what to do, the dictator could just ignore them. You see that again with AIs: they're fairly sycophantic, but they will also challenge you sometimes. Then it's up to you whether you listen.

So that's all the competence bucket, which I think is really important. I do think there are differences between potential coup instigators on that front which could be significant. My expectation would be that lab CEOs would be more competent than heads of state. But even within lab CEOs, there are some who are more dogmatic than others, and I think that dogma would get in the way of competence.

That's competence. The other thing I mentioned was, broadly, what are your goals? What are your values, or your moral character? One thing I think is really important here is being open-minded, being willing to bring in lots of diverse perspectives into the discussion, and empowering them to really represent themselves and grow and flourish.

I think a very bad thing would be a particular person becoming a dictator and implementing their vision for society. That would be one end of the spectrum; much better would be to empower all the different ideologies and ideas to become the best versions of themselves. Then we can collectively grow and improve our understanding of how to run society.

Sometimes, when people are thinking about values, they focus on, “Okay, are you this type of utilitarian?” Or, “Oh no, I hope you're not a deontologist.” It can get very specific and finger-pointing. My view is more that we don't really know what the right answer is, and the most important thing is being pluralistic and letting a thousand flowers bloom.

Gus Docker

So we discussed the possibility of getting to a stable state in which we've avoided an AI-enabled coup, and now we have, say, an aligned superintelligence such that the risk of a coup is very low. Do you think this is something that happens for 1 country, and then that 1 country is in control of the world to such an extent that this is not a process other countries are undergoing?

To be more concrete, for example, if the US goes through a period of risk of AI-enabled coups but manages to remain a stable democracy, is it the case that Russia or China will go through a similar period of risk of coups?

Tom Davidson

It's a great question, and it will depend on the US's posture toward the rest of the world geopolitically. It will also depend on whether the US has gained a huge military and economic advantage, like outgrowing the world or just developing powerful military technology, as we were discussing previously.

You can imagine 1 scenario where the US isn't that much more powerful than the rest of the world yet and isn't that inclined to intervene, which has been the recent trend. Then China develops similarly powerful AI a few years later, and Xi Jinping uses it to cement his control over China.

Then you have 1 AI-enabled dictatorship that is extremely robust, and you have the US, which has avoided that risk. Now they're maybe competing against each other in a Cold War II and trying to outgrow the world, or maybe they're striking deals because they recognize it's not good to compete. China just indefinitely remains a dictatorship, and that's a permanent loss for the world.

But you could also imagine a different scenario where the US is very far ahead and maybe just wants to really secure its position geopolitically. So it instigates AI-enabled coups in other nations, where it's really putting US representatives on top of those nations.

That could be through secret loyalties. It could sell systems to—let's say, sell AI systems to India—that are secretly loyal to US interests, or it could give particular politicians in India exclusive access to superintelligent AI to help them gain power. You could apply those same threat models we've discussed, but with the US pulling the strings.

Or you could have the US taking control of other nations in more traditional ways: military conquest and really leaning heavily on extracting economic value from other countries as it goes around the world. So, yeah, there's a wide range of options here.

Gus Docker

As a final topic here, perhaps we can talk about what listeners can do if they want to help try to prevent AI-enabled coups, and specifically where to position themselves. Should they be in AI companies? Should they be in governments? Should they be in perhaps eval organizations? Where's the position with the most leverage?

Tom Davidson

Great question. I think being at a lab is a great place to be.

I talked about system integrity—robustly ensuring that AIs don't have secret loyalties or behaviors intended to deceive. That's something that companies need to implement. So if you have an interest or expertise in sleeper agents, backdoors to AI models, or cybersecurity, I think being part of a lab and helping them achieve system integrity is an amazing way to reduce this risk.

Another thing you can do at labs, if you're worried about governments or heads of state deploying loyal AIs and seizing power, is help labs develop terms of service where, when they sell their AI systems to governments, they have certain mitigations against misuse. One way to frame this is: We're using really powerful AIs, and we can't guarantee the safety of those AI systems unless we have some degree of monitoring to ensure that the AI systems aren't doing anything they shouldn't be doing. That monitoring could then be sufficient to allow for the prevention of coups, because you'll be monitoring not only for accidental misaligned AI behavior, but that will also mean you're monitoring for a bad human actor giving the AI illegal instructions.

Labs will be drawing up contracts with governments and terms of service. They will be thinking about the guardrails, if any, that they place on the systems they sell to governments. But I think there's very careful work to be done thinking through how we can structure those guardrails and explain them in a way that is very unarguable and doesn't seem like we're trying to constrain the government. It's not really legitimate for private companies to constrain the government, but I do think there's an important thing to be done here in preventing AI-enabled coups. It's another thing you could do in a government, but you could also do it in a lab—or for a think tank or a research organization that's kind of intertwined with government, like RAND. I think RAND could potentially do some of this kind of work, thinking about what should be in the terms of service between labs and governments.

Another big thing is that, for system integrity, yes, we want labs to implement it, but we also want there to be some external organization that can certify it. Currently, no external organization is working on this. METR is not working on it, Apollo is not working on it, and I don't think any evaluation organization is, either. But at one point in time, there will be a very strong demand for guarantees that the systems these companies have developed don't have sleeper agents, because if you're deploying a system in critical infrastructure, in important government roles, or in the military, you want to know that it does what it says on the tin.

I think a company can make a wise investment by developing that capability: the ability to understand how to insert secret loyalties, how to detect them, and how to assess a company's process to see whether it's sufficient.

Gus Docker

How close would that work be to interpretability work? Isn't there overlap there—looking into what the system is actually doing, fine-grained reverse engineering of some process to see how things work? Is detecting sleeper agents a specific, constrained problem?

Tom Davidson

Yeah, it's a great question. It's a specific type of problem. You could use internal work like interpretability to detect sleeper agents, but that's not the only method you could use. You could use behavioral tests, or you could look at the training data. I think a mixture of those is useful.

It's also a specific adversarial setup. In interpretability, we're often just interested in all kinds of facts about the model's behavior and its cognition. With secret loyalties, we're specifically concerned with one red team trying to produce a secret loyalty and then a blue team trying to detect it. It's adversarial in that the blue team has a tough task. The red team will be purposely trying to hide the sleeper agent, and that will raise a whole host of new considerations and important things to dig into.

So you can imagine allowing the red team full control over all aspects of the development process. Can they create a secret loyalty? Maybe they can. Maybe it's pretty hard to test. But now you only allow them to intervene at one part of the development process. Or you allow the blue team to look at all the data used for training, or to look at all of the code that they wrote. You could get potentially quite a sophisticated understanding of the conditions under which the blue team wins and the conditions under which the red team wins.

This research doesn't need to be done in a lab. It could be done by an external organization. I think it's a big missing focus of today's technical work. Ultimately, that would inform the assessments of the labs' attempts to do system integrity. So, for technical researchers out there, I'd really highlight that possibility.

Another piece of work for the right person would be beginning to understand the existing military thinking around autonomous systems. This is already obviously a live issue for militaries. They are increasingly deploying AI. It would be nice to marry up that existing expertise with these risks about more powerful systems enabling coups and get to a consensus within that military community on basic principles, like law-following and distributed control over military systems, and figure out a military procurement process that's both practical and robustly prevents this kind of stuff.

If there's anyone listening who has a way in, I think that's potentially pretty valuable. Although there's also a risk of poisoning the well if it's done badly, so do so with some care.

Gus Docker

Yeah. Perfect. Thanks for chatting with me, Tom. It's been great.

Tom Davidson

Yeah, real pleasure. Thanks so much, Gus.

What if Humans Weaponize Superintelligence, w/ Tom Davidson, from Future of Life Institute Podcast | BidClub