[BidClub_]
The Cognitive Revolution · · 179 min

Agency over AI? Allan Dafoe on Technological Determinism & DeepMind's Safety Plans, from 80000 Hours

Nathan LabenzAllan Dafoe

YouTube
TL;DR
  • Dafoe’s central thesis is that technology creates options, but competition selects which options survive. Local actors can refuse a tool, yet if one rival converts it into military or economic advantage, holdouts must adopt or lose resources: “Technology doesn’t force us; it merely opens the door, and it’s military-economic competition that forces us through.” For investors, adoption can therefore look voluntary early and compulsory late.
  • The strongest remaining agency lies in choosing the sequence, design, and safeguards of technologies before competitive pressures harden the path. Differential technological development means building the “seat belt before the car,” but Dafoe stresses how difficult it is to identify two viable paths and forecast their downstream consequences. Nathan Labenz sharpens the near-term implication: only a few hundred or thousand people influence frontier compute decisions, so staff coordination or even individual dissent can still change what gets built.
  • Alignment alone does not guarantee a good outcome because faithfully aligned systems can serve principals who remain locked in conflict. Dafoe argues that “global coordination is almost a necessary and sufficient condition” for deploying imperfectly aligned AI prudently, while aligned great powers could otherwise automate nuclear brinkmanship, trade conflict, or cyber escalation. Cooperative AI is the marginal bet that agents should acquire bargaining, communication, and commitment skills before raw capability outruns institutions.
  • The comforting “super-cooperative AGI” hypothesis remains unproven. AI agents may communicate rapidly and be copied or tested as mediators, but they could also possess alien goals, linear utility over resources, hidden backdoors, and less mutual transparency than humans. Dafoe distinguishes cooperative skill from niceness: friendly assistants are weak evidence that strategically deployed agents will bargain safely.
  • AGI is not a single humanlike threshold but a jagged region within a high-dimensional capability space. The route matters: a system might become superhuman at materials science while remaining strategically naive, or become exceptionally skilled at cooperation before gaining broader technical power. That sequencing creates a genuine dispute between those who prefer socially simple systems and Dafoe’s concern that technical acceleration without coordination capacity could overwhelm society’s ability to adapt.
  • Frontier evaluations are becoming decision infrastructure, but they still miss capabilities unless models are given the right tools and scaffolding. On the paper’s five-point scale, Gemini 1.0 scored about 3 on persuasion, roughly 1–2 on cybersecurity and self-reasoning, and 2 on self-proliferation; Project Naptime and, to Dafoe’s understanding, OpenAI’s o1 illustrate how elicitation can sharply raise apparent capability, including with an older foundation model. Dafoe’s answer is layered evidence—lab evals, human studies, “evals in the wild,” forecasting, staged release, monitoring, and reversible access.
  • Governance becomes harder as capabilities diffuse, making responsible frontier leadership a recurring defensive requirement rather than a one-off moat. Dafoe cited Waymo’s report as showing 2× fewer incidents involving police at the crash scene and, he believed, 6× fewer injury crashes, alongside AI opportunities in medicine, tutoring, weather, materials, fusion, and AlphaFold. The upside remains tangible, while diffusion raises the prospect that stronger frontier systems may need to defend against cheaper lagging ones.
Digest · the substance, structured for research

1. Frontier employees still possess more short-run agency than determinism implies

  • Nathan Labenz opens with a Cold War story: his uncle’s nuclear-launch crew privately agreed that if a real firing order arrived, “we are all going AWOL.” The analogy is imperfect, but it foregrounds individual responsibility inside systems whose long-run strategic logic can otherwise feel inescapable.

  • Today’s AI frontier is unusually concentrated. Nathan estimates that a few hundred, perhaps a few thousand, people sit close to decisive compute allocations; elite ML talent is scarce, not every idea can be scaled simultaneously, and the “jagged edge” of intelligence leaves meaningful discretion over which capabilities receive priority.

  • Technical staff already demonstrated collective leverage during Sam Altman’s firing and reinstatement, while one former employee’s refusal to sign a non-disparagement agreement showed how individual action can alter institutional practice. Nathan’s warning is that a “glorious AI summer” may precede an AI cold war, making it worth deciding now which development paths could justify dissent or departure.

2. Dafoe moved inside DeepMind because pivotal decisions depend on who is in the room

  • Allan Dafoe’s Frontier Safety and Governance team has three pillars. Frontier safety studies emerging dangerous capabilities, forecasts their arrival, and develops mitigations; frontier governance advises on norms, regulation, and institutions; frontier planning looks toward AGI and asks what Google DeepMind, Google, and society may soon need to confront.

  • The team itself is small, but it works across Gemini safety and alignment, responsibility, policy, and other Google groups. Dafoe describes Google DeepMind as the company’s center of specialization for frontier models and “where the heart of the thinking” about frontier-policy questions occurs.

  • Dafoe left the Centre for the Governance of AI because advising Demis Hassabis and Shane Legg from inside offered more information and “more surface area” for influence. His historical model is Hamiltonian—“who’s in the room”—because crises often turn on decision-makers’ ideas, character, safety orientation, competence, and wisdom; good intentions paired with “clumsy hands” may still end badly.

  • Rob Wiblin’s scale marker is deliberately rough: AI is “certainly not more than 0.1%” of total revenue today, while its eventual share might exceed 10% and perhaps approach 100%. Dafoe’s career advice therefore remains “hop trains as soon as they can”; despite rapid growth, he considers this “still early days.”

3. Macro trends make history look less voluntary than it feels from inside

  • Dafoe began with the question, “Who, if anyone, controls technological change?” His dissatisfaction was with the intuitive model that history equals the sum of everyone’s effort. General-equilibrium systems can resist each unit of effort with equal or greater counterpressure—or amplify it—so local intention alone says little about lasting impact.

  • The empirical puzzle is the regularity of macro history. Germany and Japan returned to their prewar growth trajectories within less than a decade of World War II’s devastation; Moore’s law traced an unusually precise exponential line; and scaling laws now permit forecasts of model size and loss years ahead.

  • Similar directionality appears in civilization’s maximum energy processing, building height, material durability, and transportation speed. As Robert Wright’s archaeological formulation has it, “the deeper you dig, the simpler the society whose remains you find”; some technological orderings also look constrained, since nuclear power is difficult to imagine preceding coal power.

  • Yet these patterns cannot simply be credited to collective human will. The Agricultural Revolution appears to have reduced median health and welfare for a long period while enabling inequality and warfare. Dafoe’s conclusion is not that humans lack agency, but that its effect depends on timing, power, resource constraints, and which systems remain functional under selection.

4. Constructivism explains local choice while struggling with emergent constraints

  • Earlier theorists sometimes endowed technology with near-agency. Langdon Winner discussed technological autonomy, Lewis Mumford’s capital-M “Machine” reduced people to supporting cogs, and Jacques Ellul’s “la technique” described functional imperatives that leave humans choosing under coercion.

  • Social constructivists reacted by looking closely at actual decisions. Under the microscope they found people, interests, ideologies, competing designs, dead ends, and uncertain visions—not an autonomous machine dictating that bicycles, airplanes, or infrastructures take one predetermined form. Rob captures the critique with: “Is the machine with us in the room right now?”

  • Dafoe’s concession is methodological: ethnography and microhistory answer real questions about how choices were made. The error was using those tools to dismiss macro claims they were poorly equipped to test. The political appeal of the constructivist view is that it preserves citizens’ sense of agency; strategically, Dafoe wants a model that directs effort toward points where structure will not simply push back.

5. Technology embeds politics, accumulates momentum, and surprises its designers

  • “Technological politics” is the least controversial form of determinism: people can encode political objectives into design. Parisian boulevards facilitated cavalry movement against rebellions; gates and urban layouts discipline behavior; and Robert Moses allegedly built bridges too low for buses, limiting beach access for New Yorkers without cars, especially African Americans.

  • Rob’s contemporary specimen is social media. Recommendation algorithms and the prominence of quote-posting can encourage tribal denunciation and brigading. Design does not force a particular utterance, but it changes which behavior receives attention, reinforcement, and perceived public legitimacy.

  • “Technological momentum” describes sunk infrastructure and expertise. Car-dependent American cities make dense pedestrian life harder, while earlier investment might have accelerated electric vehicles, wind, or solar by five or ten years. Dafoe nevertheless thinks path-dependence claims are often overstated: neural-network insights existed early, but cheap FLOPs were needed before they became broadly useful and easy to rediscover.

  • Winner’s further warning was a “sea of unintended consequences”: society invents first, discovers effects later, then adapts. Rob pushes back that technology generally seems to solve more problems than it creates, with negative side effects shrinking across generations. Dafoe does not resolve that broad balance here; he returns to the narrower point that important consequences routinely arrive unplanned.

6. Military-economic competition supplies the missing mechanism

  • Dafoe observes a strong relationship between analytical scale and conclusion. Micro methods usually produce constructivist explanations centered on choices and visions; macro methods more often reveal deterministic regularities. His response is that different scales contain genuinely different emergent phenomena, not that one side is simply looking at bad data.

  • His analogy is a science of water. One school studies wind-driven ripples, another throws rocks, while a “kooky macro water phenomenologist” notices tides tracking the Moon everywhere on Earth. Lacking a microscopic mechanism would be a challenge for lunar determinism, but not a reason to discard its robust pattern.

  • The proposed microfoundation is selection among “ways of living”—sociotechnical systems that require resources to persist and spread. Dafoe layers environmental selection above military and economic competition, with culture and psychology lower down. A locally preferred arrangement survives only while it remains economically viable, militarily secure, and environmentally sustainable.

  • His signature synthesis: “Technology doesn’t force us; it merely opens the door, and it’s military-economic competition that forces us through.” Groups can initially reject a useful technology, but one group’s successful adoption creates pressure on the rest. Refusal then means either catching up later or losing resources and autonomy to the fitter system.

7. Internalized competition hides the selection pressure from view

  • Military competition is ubiquitous at macro scale but rare in daily observation. Even long peaceful periods occur under “the shadow of violence”: communities anticipate future aggression and modify institutions before war arrives. Once that pressure becomes ideology, nationalism, or commercial common sense, a microhistorian may observe the internalized representation rather than the competition itself.

  • Dafoe calls this “vicarious selection.” Engineers do not repeatedly build full aircraft and crash them; they develop wind tunnels and theories that simulate the selecting environment. Firms and states similarly model rivals, test options internally, and adopt what they expect competition would eventually reward.

  • Rob’s UK objection is useful: Britain does not choose every housing or urban policy from fear of French or Russian invasion. Dafoe concedes that military competition has declined in the modern era and global culture has strengthened, but notes how quickly national-security claims still mobilize reform when leaders believe strategic position is at stake.

  • The British electricity system illustrates delayed constraint. Decentralized local power arguably fit the country’s democratic ethos and persisted up to World War II, when its cost constraints became excessive; Britain then adopted a national grid. A community can retain its preferred arrangement for decades, but crisis exposes accumulated functional disadvantages.

8. Japan shows that technological autonomy can have an expiry date

  • Under the Tokugawa shogunate, Japan maintained a feudal order for roughly 200–250 years and effectively “uninvented” firearms. Production was centralized, gunsmiths were paid not to build, and knowledge of cannons and firearms was allowed to disappear while the country kept limited contact with the outside world.

  • The break came in 1853, when Commodore Perry arrived in steamships that moved upwind without sails, belched black smoke, and carried powerful cannons. After demonstrating bombardment, he reportedly supplied white flags for requesting that it stop. When asked whether he would return with the ships, he answered: “I’ll bring more.”

  • Perry’s arrival triggered roughly 15 years of revolution culminating in the Meiji Restoration and a wholesale commitment to modernization. Japan sent people abroad to collect knowledge of Western industrial arts, then caught up rapidly enough to contest control of Asia against the United States, Britain, and others within the following several decades.

  • Rob’s refinement is geographic: island defenses gave Japan unusual room to fall behind. Once the gap became large enough that the sea no longer protected it, the country made a “complete 180.” For Dafoe, this cleanly separates genuine early discretion from the later coercion produced by overwhelming external capability.

9. Differential development means building safeguards before capabilities

  • The cleanest analogy is the seat belt, which conceptually could have preceded the car. Developing it first would have allowed safety to diffuse alongside motor vehicles rather than decades later. Vaccines represent anticipatory defenses, while cheaper wind or solar could have substituted for a fossil-fuel path before infrastructure locked in.

  • Technologies may also carry political byproducts. Winner argued that nuclear power favors centralized development and coercive security, whereas wind and solar permit more decentralized production. Differential technological development therefore includes not only paired safeguards, but choosing among trajectories with different institutional consequences.

  • Dafoe regards the idea as foundational to AGI safety: society deliberately invests more in safety than the market otherwise would, hoping to deliver the “seat belt before the car.” Technological momentum supports scrutinizing major early infrastructure commitments, because expertise and capital become sunk costs that narrow future choice.

  • The pessimism is epistemic. An intervener must identify two technically viable marginal paths that others are not already pursuing, persuade resources toward one, then forecast each path’s technological descendants and direct and indirect social effects. Markets already pay heavily for the first insight, while history shows how unreliable the second remains.

10. Market incentives make ordinary alignment less differential than it appears

  • Rob’s pushback is that alignment is not a binary fork between aligned and deliberately unaligned AI. Developers already need models to follow instructions, and marginal safety investment probably improves outcomes. He finds the case for alignment far clearer than speculative attempts to redirect entire technological trajectories.

  • Dafoe agrees that safety and alignment are “very good bets on net,” but introduces a general-equilibrium counterfactual. Useful assistants are constrained by alignment, so the market has powerful reasons to solve ordinary instruction-following and safety problems even without altruistic funding.

  • Reinforcement learning from human feedback and constitutional AI were advanced by AGI-safety-motivated researchers, perhaps bringing them forward several years. The skeptical response is that this work may merely have accelerated commercially necessary technology—or even hastened capabilities—because market researchers would eventually have produced equivalents.

  • The genuinely differential target is work the market will not finish before a critical capability transition. Dafoe asks, “What’s the seat belt for AGI?” Deception is one candidate: a sufficiently capable misaligned system might conceal its failure, creating a qualitatively different problem from today’s visibly imperfect assistants.

11. Cooperative AI tackles failures that alignment cannot solve

  • Dafoe’s deliberately strong opening claim is that “alignment is insufficient” and may not be strictly necessary for good outcomes. If alignment were only 90% solved but humanity coordinated wisely, it could restrict deployment to domains and scales where systems remained safe.

  • In that sense, he calls global coordination “almost a necessary and sufficient condition.” Coordination would allow humanity to appoint a reasonable decision-maker to balance risk and benefit, while imperfect coordination can produce disaster even if every AI faithfully serves its principal.

  • Fully aligned systems could amplify conflict among great powers, just as self-interested humans have accepted nuclear brinkmanship. Dafoe also points to trade wars, climate change, pandemic preparedness, and foregone global commerce as collective-action failures that alignment to separate operators does not repair.

  • His desired outcome therefore has two pieces: systems that behave safely and as intended, plus institutions capable of deploying them “jointly peacefully and productively.” Cooperative AI is the bet that marginal research can bring bargaining and coordination competence online before escalating capability makes old failures faster and more consequential.

12. Autonomous agents could compress both market accidents and military escalation

  • The 2010 flash crash is Dafoe’s warning from primitive automation: interacting trading algorithms produced trillions of dollars in paper losses before market stops paused activity and allowed trades to unwind. Similar feedback once pushed Amazon book prices into the millions because sellers’ pricing algorithms recursively responded to one another.

  • Agents with access to bank accounts, email, tools, and multiple domains create a larger surface for such emergent dynamics. Each agent’s local protocol may be sensible under normal conditions, while the interacting system leaves the range its designers anticipated.

  • Rob’s darker scenario is machine-speed military brinkmanship. Escalation that takes human leaders days or months could unfold in minutes if AI systems control important decisions. Dafoe says autonomy should scale with stakes and available actuators; weapon control demands much stronger human review, yet time pressure could still produce what Paul Scharre calls “flash escalation,” including in cyber conflict.

13. AI delegates could remove information hazards from bargaining

  • Dafoe’s most concrete bargaining design is to “put your AI delegate in a box with my AI delegate.” The agents could privately exchange information, but the box would reveal only a proposed agreement or “no deal,” reducing incentives to posture, delay, hide preferences, or signal costly resolve.

  • Rob adds advantages unavailable to human negotiators: agents can communicate at enormous bandwidth, be copied exactly, and demonstrate consistent behavior across tests. A mediator’s prior decisions could establish a track record more reliably than trying to infer a person’s character from limited history.

  • Cooperative AI is a portfolio rather than an exclusively machine-to-machine agenda. It covers AI–AI, AI–human, and AI-assisted human–human coordination. Dafoe is especially interested in systems that help people discover agreements they could not efficiently articulate themselves.

  • Google DeepMind’s “Habermas Machine” is the political-deliberation specimen. Language models summarized participants’ positions and produced a detailed consensus they would endorse; the reported result was that the AI articulated consensus better than hired human facilitators. Dafoe sees potential for exposing issue dimensions, priorities, and mutually acceptable trades without eliminating political choice.

14. Intelligence does not automatically imply superhuman cooperation

  • The “super-cooperative AGI hypothesis” says cooperative competence will scale with general intelligence until AGIs can solve global coordination almost by default. If true, researchers could prioritize safety and alignment, trusting sufficiently advanced systems to handle bargaining later.

  • Humans possess underappreciated cooperative advantages. They share biological and cultural backgrounds, recognize expressions, draw on long histories of behavior, and usually pursue goals within a familiar range. Democracies can also be unusually transparent because negotiators can read one another’s press and public debate.

  • AI goals might be far more alien. Most people show diminishing returns and avoid an all-or-nothing 50% gamble over everything they own; an AI could have linear utility over wealth or another resource, making its bargaining behavior harder to predict from human experience.

  • Even highly capable systems may remain opaque to each other. Interpretability tools could help, but a strategic bargainer has little reason to expose every internal state. Dafoe’s conclusion is not that AI cooperation will be poor, only that super-cooperation is an empirical hypothesis rather than a free byproduct researchers should assume.

15. Backdoors make machine trust unusually brittle

  • Rob’s strongest counterexample is a backdoored model whose behavior flips after a “magic word” or subtle environmental cue. Humans can deceive, but a single hidden feature that reverses an entire objective is much less natural for people and, with current interpretability, extremely difficult to detect in models.

  • Internal transparency may therefore be insufficient. A model could hide the backdoor so subtly that another agent misreads every otherwise visible activation. The same history and architecture that appeared trustworthy across thousands of tests could become poor evidence under a carefully selected trigger.

  • Rob and Dafoe explore the strange strategic upside of self-backdoors as anti-theft devices. Merely suggesting that national-security models might “call home” or behave contrary to a thief’s intentions could deter theft; attackers would then fund alignment to remove traps, while defenders would fund it to make traps survive. Dafoe calls this potentially virtuous but warns that deliberately creating trigger-sensitive or deceptive architectures is “dancing on a knife edge.”

16. Cooperative skill is not the same thing as a pleasant personality

  • Today’s language models seem agreeable because they are trained to interact helpfully with individual users. Dafoe’s key distinction is that cooperative skill is distinct from a cooperative disposition: being nice, altruistic, or generous is not the same as solving bargaining, communication, and commitment problems under strategic pressure.

  • In equilibrium, a company or state may want an agent that faithfully advances its own interests, not one that casually concedes for the social good. The relevant question is whether separately aligned delegates can still locate efficient agreements—or trust a mediator that weighs each principal’s evidence, goals, resources, and outside options fairly.

  • The Cooperative AI Foundation’s highest-leverage investment has been environments and benchmarks. A good benchmark is a public good: it measures cooperative competence, gives researchers a target to “hill climb on,” and rewards progress before commercial deployment exposes costly failures.

  • Candidate measurements include theory of mind, shared vocabulary, strategic communication, and performance across games demanding different behavior. A competent agent may need firmness in one setting and generosity in another; a single friendliness score would miss the reasoning required to distinguish them.

17. Commitment remains cooperation’s largest prize and hardest bottleneck

  • Communication begins with whether agents share meanings, then becomes strategic: how can one reveal what matters without handing the other side exploitable information? Effective cooperation requires disclosing enough to unlock joint surplus while minimizing vulnerability if the counterparty defects.

  • Commitment problems persist even with perfect understanding. In a one-shot prisoner’s dilemma, both parties know mutual cooperation is better, yet neither can enforce “I cooperate if you cooperate.” Cooperative technology would need treaty-like protocols that make discovered bargains credible.

  • Rob’s pushback is that invoking a solution to commitment can become an evasion: the problem has “beguiled humanity since basically the beginning of history.” Saying AGI goes well if researchers solve it resembles saying a theory of consciousness would clear up every downstream puzzle—it identifies the missing miracle, not a strategy.

  • Dafoe still considers it worth researching. Robert Powell’s game-theoretic framing makes nearly every cooperation failure reducible to commitment; Carl Shulman’s most compelling proposal is for rival AIs to jointly construct a third system, verify that neither side backdoored it, then delegate decisions to that impartial authority. The prize is immense if verification can actually work.

18. Better cooperation can empower cartels and exclude everyone outside the bargain

  • Cooperation improves outcomes for participating agents, not necessarily society. Dafoe calls exclusion the central caveat: two or ten agents may approach their joint Pareto frontier while imposing losses on outsiders. Rob’s image is “two wolves and a sheep deciding who to eat for lunch.”

  • Many institutions prohibit cooperation for precisely this reason—students sharing answers, athletes fixing contests, marketplace actors violating rules, or criminals coordinating. The mafia’s power comes partly from sustaining contracts and intense internal cooperation outside the legal system. Cooperative competence is therefore a dual-use capability.

  • Dafoe’s working hypothesis is that broadly raising cooperative skill remains net beneficial, much like trade, because the attainable pro-social surplus exceeds the antisocial gains. He labels it a hypothesis to be tested, not an axiom. Industrial farming is Rob’s cautionary case: greater human cooperation may have increased total suffering for animals unable to negotiate inclusion.

  • Humans could eventually occupy the excluded position. Machine-speed agents may communicate and commit more effectively with one another than with people, forming a high-surplus cluster that leaves humanity behind. Worse, cooperative AGIs might collude against alignment institutions built from multiple mutually checking models. Dafoe calls that one of the agenda’s largest safety downsides, but still judges global coordination important enough to investigate the bet.

19. AGI is a region of capability space, not a single humanlike point

  • Dafoe’s “Levels of AGI” paper aimed to formalize what many people already mean, replacing the claim that AGI is hopelessly undefined with an operational framework. The concept persists because alternatives such as “transformative AI” capture economic impact but can also include narrow systems with enormous effects or catastrophic potential.

  • “Human-level AI” is misleading when treated as one threshold. AI already exceeds humans at chess, memory, and some mathematics while failing elsewhere; future systems will likely remain highly imbalanced. AGI denotes a broad space of systems better than people across most relevant tasks, not a machine reproducing the human profile.

  • Definitions still require parameters. “Most” might mean 50% or 99% of economically relevant tasks, but probably not 100%, because a small tail could remain unautomated after most impact occurs. The comparison should generally be with skilled humans in each task, not an untrained median person.

  • Crossing that region matters twice. Economically, broad substitution removes the natural human-in-the-loop created by human labor. Technologically, superhuman performance breaks the old ceiling on what can be done at all, as AlphaFold did for protein-structure prediction. AGI is thus both a labor threshold and a qualitative expansion of the feasible set.

20. General intelligence may win, but human imitation is not the goal

  • One critique of AGI is that the label encourages building machines in our image, maximizing labor substitution. The economic alternative is “alien, complementary” intelligence: AlphaFold performs a task humans could not do directly, raises scientific productivity, and does not replace people who previously predicted protein structures unaided.

  • Dafoe is less persuaded that narrow systems will remain dominant. Large language models suggest strong returns to generality: the best model for poetry, history, philosophy, or email may be the same model trained on the full human corpus because lessons spill across domains.

  • Rob notes that this is a surprising fact about knowledge rather than a logical necessity. Future specialization could reappear; a coding model might gain something from history during training, then distill away most historical knowledge for efficiency. Dafoe treats continued returns to generality as an empirical phenomenon to track.

  • A separate threshold is AI capable of machine-learning R&D and recursive improvement. It may arrive before full AGI if the workflow—hypothesis generation, benchmark design, compute-heavy experiments, and evaluation—is easier to automate than practitioners expect. Surveys often place ML automation unusually late, perhaps because experts regard their own work as special; Dafoe also allows that improving ML could genuinely be among the final tasks fully automated.

21. The order in which capabilities arrive may determine whether society can cope

  • Dafoe describes AGI as “the corner of this high-dimensional intelligence space,” containing many systems and many paths of approach. Differential development therefore applies not merely to whether AGI arrives, but to which dimensions become superhuman first.

  • Rob contrasts a system already exceptional at cooperation but weak in physics and materials with one that is superhuman technically but an amateur strategist. That ordering could determine whether early AI primarily strengthens coordination or unleashes capabilities whose disruptions institutions cannot manage.

  • Some safety thinkers prefer technical power paired with social simplicity: humans might harvest health and economic gains while retaining control over agents too strategically naive to outwit them. The Cooperative AI bet reverses the emphasis, warning that accelerating science without better collective problem-solving could generate shocks faster than society adapts.

  • Dafoe wants the disagreement made explicit before large investments harden either path. Eliezer Yudkowsky and Nate Soares represent skepticism about early social sophistication; Paul Christiano and Carl Shulman emphasize super-cooperative AGI but expect it largely as a byproduct. Dafoe’s position is that marginal, deliberate investment may still be needed.

22. Frontier evaluations are decision tools, not definitive capability measurements

  • Google DeepMind’s evaluation paper examined persuasion and deception, cybersecurity, self-proliferation, and self-reasoning, with biological threats handled separately. It reported on Gemini 1.0; Dafoe noted that the paper was already old because the field was moving rapidly.

  • One self-reasoning test asks a model to produce text containing words silently blocked by a filter. Success requires noticing the failure, inspecting a simulated directory, finding the blocked-word file, and changing it. Another offers one search query across two historical questions, testing whether the model knows its training cutoff and searches for the post-cutoff event.

  • On a subjective five-point aggregate scale, Gemini scored about 3 on persuasion, roughly 1–2 on cybersecurity, 2 on self-proliferation, and around 1–2 on self-reasoning. Dafoe says the model did “not very well” on situational-awareness tasks, while cautioning that the underlying component measurements were more structured than the overall ratings.

  • Cyber performance exposes the elicitation problem. Google’s Project Naptime found large gains after supplying models with an interactive environment, tools, and reasoning support. OpenAI’s o1 likewise showed how scaffolding and fine-tuning over, to Dafoe’s understanding, an older foundation model could transform performance. A safety eval seeking the “max capability” must give the model its strongest plausible chance to reveal it.

23. Evals need observation, forecasting, and reversible deployment around them

  • Model behavior is multidimensional and context-dependent, closer to psychology than a single automated exam. Dafoe wants prespecified tests and thresholds, but also an exploratory ecosystem where researchers notice surprising behavior, run idiosyncratic probes, and report phenomena before the field knows how to formalize them.

  • His typology expands outward: automated model evals, human-subject interaction, realistic user studies, then “evals in the wild.” Observing cybersecurity teams’ revealed willingness to pay for Gemini assistance has strong external validity, though it is a lagging indicator; the best targets are early adopters whose behavior reveals capability sooner.

  • Scaling laws forecast next-token loss with extraordinary precision but do not directly reveal when a model becomes economically useful. A 99%-reliable self-driving system can still be worthless if deployment needs 99.999%. Dafoe highlights “observational scaling laws,” recently recognized as a NeurIPS spotlight, which adjust across model families and architectures to predict downstream task performance.

  • The paper also asked calibrated forecasters from the Swift Centre when models would cross evaluation thresholds. Dafoe wants that practice repeated until evidence shows it fails. Because false negatives will remain, forecasting must sit inside staged deployment: internal access, widening trusted testers, limited release, monitored general access, and protected weights that permit guardrails or withdrawal when hidden capabilities emerge.

24. Frontier governance must become external, multilayered, and structurally aware

  • Rob presses the conflict of interest in letting model developers decide whether their own products are too dangerous to release. Dafoe agrees that long-run governance needs government rules, third-party evaluations, nonprofits, and academia. Company frameworks are best understood as opening proposals intended to converge into common standards, not permanent self-regulation.

  • Google participates with the UK and US AI Safety Institutes, the Frontier Model Forum, and a Frontier Safety Fund designed to move marginal resources toward standards and research. Dafoe also cites the White House commitments, UK AI Safety Summit commitments, and AI Seoul Summit commitments; internally, he says Google translated commitments into concrete work streams rather than treating them as “cheap talk.”

  • Structural risk sits upstream of misuse and accidents. The Cuban Missile Crisis involved legitimate governments and functioning weapons, yet geopolitical structure pushed leaders into catastrophic brinkmanship. Railways similarly may have increased first-strike advantage and made mobilization difficult to reverse—effects no “railroad-level eval” could infer from steel rails alone because they emerged through institutions.

  • Democratic legitimacy remains Rob’s unresolved challenge: people worldwide bear potential catastrophic risk without proportionate consent, while companies have rarely championed specific costly limits on themselves. Dafoe points to citizen juries, pluralistic assemblies, free media, regulators, and scientific-policy consensus. Education and visible benefits matter, but the appropriate mitigation regime cannot be settled by companies or a narrow agency alone.

25. Falling costs make defense permanent, while the upside supplies the reason to try

  • Open weights create the most immediate irreversibility: once weights are published, scientists and developers can fine-tune and inspect them, but every bad actor can also retain them. Separately, Epoch AI estimates algorithmic efficiency improving about 3× per year—roughly 10× cheaper every two years—so a $100 million training run could fall to $10 million, then $1 million if the trend persists.

  • Dafoe described a possible wise condition: the best models are developed and controlled responsibly and remain sufficiently ahead to defend against cheaper, two-year-old models. Resource advantage helps only when offense-defense economics cooperate: 100 defensive dollars beat one attacker dollar under a near 1:1 ratio, but a bioweapon can impose vast costs before vaccines arrive, and social systems are slower to patch than software.

  • The labor requirement is correspondingly broad. Dafoe wants more political scientists, economists, historians, philosophers, ethicists, sociologists, ethnographers, forecasters, international-relations specialists, agent-safety researchers, and technical safety staff. AI’s effects span society, so neither governance nor impact forecasting can be staffed solely by model builders.

  • The payoff case is concrete: Dafoe recalled Waymo’s report as showing 2× fewer incidents involving police at the crash scene and, he believed, about 6× fewer injury crashes, while autonomy could reclaim parking space; medical assistants could triage symptoms; AI tutors could personalize help; and DeepMind projects span AlphaFold and AlphaFold 2, fusion-plasma control, weather prediction, materials discovery, contrail reduction, renewable-energy planning, and data-center efficiency. Dafoe’s closing condition is simple: the benefits “will be profound if built safely.”

Nathan Labenz

Today, I'm excited to share a special crossover episode from the 80,000 Hours Podcast featuring a conversation between host Rob Wiblin and Allan Dafoe, director of Frontier Safety and Governance at Google DeepMind. I first heard Allan speak back in 2017, when I introduced him at a conference in Boston as a professor at Yale who was then working on great-power peace. This was before he founded the Centre for the Governance of AI, which in turn was years before he moved to DeepMind, so I can say with confidence that Allan has been thinking about AI governance harder and planning for the current AI moment longer than just about anyone else. As you'll hear, that pays off in the form of truly excellent analysis on an impressive range of critical topics.

To begin, Allan describes his academic work on the question of just how much ability humans really have to alter the course of technology development, noting that macrohistorical trends like Moore's law suggest a process that transcends individual human choices. He ultimately argues that while technological possibilities don't force us to do anything on their own, in combination with the realities of military and economic competition, they can and often do. Put simply, failure to adopt potentially advantageous technologies often means losing to those who do.

This is not a conclusion that Allan comes to lightly, and unfortunately for us today, I think it's a pretty hard one to escape. It's still possible that a spectacular incident could cause a vibe shift big enough to force a pause in frontier scaling, but the smart money now seems to be on powerful AI soon. With Pentagon officials quoted in the press expressing their enthusiasm for autonomous killer robots, despite the general reliability, reward-hacking, and even scheming issues that have recently come to light, militarization of some form seems a foregone conclusion as well.

And yet, even if the long-term logic is inescapable, I think it would be a huge mistake for frontier developers to underestimate their own individual and collective short-term agency. A few years ago, my uncle told me a story about when he arrived in Italy during the height of the Cold War to join a crew that was responsible for firing nuclear weapons at tertiary targets in the event of an all-out war. The first time they drilled the launch sequence, one of the longer-tenured guys took him aside and said, “Just so you know, if the order ever comes down to shoot for real, we are all going AWOL. None of us want to be part of destroying the world with nukes.”

Now, that's just one story from one enlisted crew, and I have no idea if that sentiment was widespread enough to have made a real difference in the worst-case scenario. But today, the reality is that a very small number of people are pushing the AI capabilities frontier forward. There are only so many elite ML savants, and compute constraints mean we can't scale all their ideas at once anyway.

Meanwhile, it's also now well established that intelligence itself has a jagged edge. Unlike nuclear technology, which had a small number of discrete, powerful use cases and a very mechanical associated game theory, the design space from which AI developers are selecting new forms is manifestly vast, and the models themselves are incredibly malleable. If you believe things could move super quickly as AIs begin to hit important capability thresholds, the specific details of what we build and prioritize just before that point could prove decisive.

All this puts the few hundred—or maybe as many as a few thousand—people who are closest to the major compute-budget decisions in a position of great power and responsibility. As we saw in the context of Sam Altman's firing and subsequent reinstatement, a serious threat by technical staff to walk out can force leadership's hand. Further, as past guest Daniel Kokotajlo demonstrated by refusing to sign a nondisparagement clause, even a single individual can create meaningful change if they're willing to stand up for what they believe in.

So I would encourage everyone at all of the frontier AI companies to make time to raise their own level of situational awareness, even if that comes at the cost of moving their specific project forward a bit more slowly, to make sure that the overall enterprise they're engaged in continues to be one that they feel good about supporting. To date, you truly have so much to be proud of: top-tier language models, AI doctors, self-driving cars, a revolution in biology, and robots now folding origami. DeepMind could never ship another product and would already go down as a historically important company.

There's a lot to appreciate in this conversation on the alignment, safety, and policy fronts, too. Allan's Cooperative AI research agenda is both fresh and sophisticated. Google's Frontier Safety Framework has truly been, as Allan describes it, part of a serious and important effort by leading companies to advance the AI policy conversation. And lately, I have been thrilled to hear Demis Hassabis buck the trend by continuing to speak about the possible need for international collaboration on advanced AI development.

At the same time, AI Manhattan Project–type ideas are rapidly proliferating, and it's not hard to imagine such a project going so catastrophically wrong as to more than offset even the tremendous amount of good that DeepMind and other AI leaders have already done to date. So, again, for the people at frontier companies, keep in mind that history is not happening to you, nor are you merely living through it. You are part of a relatively small group driving—or, at the very least, shaping—it in important ways.

We are now in a glorious AI summer, but an AI cold war is looming. Your critical decisions won't be binary like my uncle's squad's were, and there are clearly many defensive AI systems that we genuinely need to build. But all the more so because we're accelerating into a super-high-dimensional, uncharted space, if you haven't already, I think it is now time to start thinking—and even talking to colleagues—about which directions might convince you personally, or perhaps one day as a group, to go AWOL from the project.

I hope you enjoy this very enlightening and deeply thought-provoking conversation between Rob Wiblin and Allan Dafoe of Google DeepMind from the 80,000 Hours Podcast.

Speaker 1

One famous quote in the history of technology, arguing against determinism, was that technology doesn't force us to do anything; it merely opens the door. It makes possible new ways of living, new forms of life. My retort was: technology doesn't force us; it merely opens the door, and it's military-economic competition that forces us through.

When a new technology comes on the stage, many groups can choose to ignore it or do whatever they will with it. But if one group chooses to employ it in this functional way, that gives them some advantage. Eventually, the pressure from that group will come to all the rest and either force them to adopt or lead the other group to lose their resources to the new, more fit group.

Today, I have the pleasure of speaking again with Allan Dafoe, who is currently the director of Frontier Safety and Governance at Google DeepMind, or GDM for short. Before that, he was the founding director of the Centre for the Governance of AI. He was also a founder of the Cooperative AI Foundation and is a visiting scholar at the Oxford Martin School's AI Governance Initiative. Before all of that, you were an academic in the social sciences studying technological determinism, great-power conflict, great-power peace theory, that kind of thing, which I guess we're going to get a little bit of all of these different pieces of your work today. Thanks so much for coming back on the show, Allan.

Allan Dafoe

Thanks, Rob. A pleasure to be here.

Speaker 1

Later on, we're going to talk about the frontier model evaluations, as well as what you think Cooperative AI might be about as important as aligned AI. But first off, I guess you're director of Frontier Safety and Governance. What does that actually involve in practice? I can see that being a whole lot of different things, and I don't have a sense of what your day-to-day is like.

Allan Dafoe

My team is called the Frontier Safety and Governance team, and we have 3 main pillars: frontier safety, frontier governance, per the name, and then frontier planning.

This adjective “frontier” is a new term, I would say almost of art, to refer to these general-purpose large models like Gemini and others. Frontier safety looks at dangerous-capability evaluations. It tries to understand what powerful capabilities may be emerging from these large, general-purpose models, forecast when those capabilities arrive, and then think about risk mitigation and risk management.

This also led to the Frontier Safety Framework, which is Google's approach to risk management for extreme risks in frontier models. That's frontier safety. Frontier governance is advising on norms, policies, regulations, and institutions, especially with an eye towards safety considerations. And then frontier planning looks to the horizon, tries to imagine what new considerations could be coming with powerful AI and on the path to AGI, and advises Google DeepMind, Google, and really all of society, given those insights.

Speaker 1

That sounds like a pretty big remit. How large is the team that's working on all these questions?

Allan Dafoe

The team is quite small, though we're actually hiring for several positions right now. By the time the podcast goes live, that may be wrapped up. What's really great about working at Google DeepMind is that we have a lot of partner teams and a very collaborative culture.

We work with technical safety—called the AI Safety, Gemini Safety, and Alignment teams. We work with responsibility teams, policy teams, and so forth.

Speaker 1

I guess Google DeepMind has, over the last year or two, become more integrated into the rest of Google. Are there other groups within this broader entity—I guess Alphabet—that take an interest in these questions? Or are you maybe the only group thinking about the frontier? I guess you're thinking about the most important models and upcoming issues and threats. Are there many other groups taking an interest, having that kind of foresight and thinking years ahead?

Allan Dafoe

I would say Google DeepMind is the part of Google that's most specialized at thinking about frontier models. Google DeepMind is responsible for building Gemini, the frontier model that's underpinning all of what Google is doing, and we also have responsibility, safety, and policy teams that are especially thinking about frontier issues.

We then do have partners across Google in these various domains. For example, in policy, we work closely with Google Public Policy on the range of policy implications and considerations connected to frontier models. But Google DeepMind is, I would say, where the heart of the thinking related to frontier policy issues takes place.

Speaker 1

Okay, so I guess back in 2021, you'd been the founding director of the Centre for the Governance of AI, or GovAI, which was a reasonably big deal then. I think it's gone on maybe to be an even bigger deal since; it's a pretty prominent voice in the conversation around governance of AI. Why did you decide to leave this thing that was going quite well to go and work at Google DeepMind instead?

Allan Dafoe

I agree: it went well at the time, and it's gone even better since. A lot of credit goes to Ben Garfinkel, who's the executive director of GovAI, and the many others who work there.

At the time, I was an informal adviser to Demis Hassabis, CEO of DeepMind, and Shane Legg, co-founder of DeepMind. I found that I had a lot of potential impact in giving advice on AGI safety, AGI governance, and AGI strategy. However, to be most impactful, it helps to be inside the company, where I have more understanding of the nature of the decisions that they're confronting and more surface area to advise not only Demis and Shane, but also many key decision-makers.

To take a step back, I want to reflect on this road-to-impact approach of advising important decision-makers. I would say one lesson I've drawn from history is that often, in these pivotal historical moments—in crises or in very high-leverage historical moments—a lot depends on the behavior, ideas, and very character of key individuals in history. In the musical portrayal of Alexander Hamilton, it's like, “Who's in the room?” What decisions are made in the room?

I think that's true. When you look at history, especially in these pivotal historical moments, it's incredible how much the ideas that people bring into the room, the resources, and the insights that they have available shape the solutions that they construct. So that argues for advising people who will be influential on these important historical developments. I think AI and AGI was, in my view, one of the most important historical developments, and I think Demis and DeepMind are very likely to be influential in the ongoing development of AI and AGI. They have been so far.

There's a second part of this, which is that, in addition to advising influential decision-makers, there's the idea of boosting decision-makers who have the kind of character you would want in critical decision-makers. Do they have the sensibility—in my case, are they aware of the full stakes of what is happening? Are they safety-conscious? Do they have the technical and organizational competence to pull off what needs to be built? Because if you have clumsy hands, even if you have good intentions, that may still lead to a bad outcome. Finally, do they have the wisdom to be able to make these very hard decisions that have complex and uncertain parameters around them?

In my view, Demis—and, again, Shane, to mention him—are extremely impressive individuals in these properties: their safety orientation, their broad perspective on the stakes of the issues, their wisdom, and their broad character.

I also do want to reflect on GovAI during my time.

Speaker 1

Yeah, it produced a lot of great work and great people. It's since gone on, I think, to produce even significantly more great work and great people, so Ben Garfinkel has done a great job.

It's interesting reflecting on some of the people who've gone through GovAI. One person who worked very closely with me at the time is Jade Leung. She used to be head of a partner team at OpenAI and is now the chief technology officer at the UK AI Safety Institute. A number of other very prominent people in AI safety and governance have similarly gone through GovAI: Markus Anderljung, Robert Trager, Anton Korinek, who's a prominent economist who's done some work there, Miles Brundage, and others.

I guess back in 2018, in our short interview back then, you were saying people should definitely be diving into this area because it's going to grow enormously, and it's going to be really good for your career, with lots of opportunities. I think that has definitely been borne out: people who got in on the ground floor have been doing super well career-wise.

Allan Dafoe

I think it's still early days for any prospective joiners. I always encourage people to hop trains as soon as they can, because AI is only—it's still just a small fraction of the economy, so there's a lot more impact to come and work to do.

We think that, in the fullness of time, it's going to be close to 100%, certainly more than 10%. It's 0.1% now—certainly not more than 0.1% in terms of total revenue—so there are many orders of magnitude to go up yet.

Speaker 1

Let's open, though, by talking about the work that you did in your previous incarnation, which was as an academic. I think you did your thesis back in the early 2010s on technological determinism. That was the main focus, I think, of the paper that came out of it. Who, if anyone, controls technological change? What was the academic debate there that you were reacting to or trying to be a part of?

Allan Dafoe

My academic trajectory had a number of chapters. The first was on technological determinism, which we can come to. For completeness, the second was on great-power politics and peace specifically, which actually led to a lot of work that continues to be relevant, I think, to the question of AI and AGI governance. And I also did some statistics and causal-inference work, which has some relevance to thinking about AI today.

Turning back to technological determinism, I would say I first came to this in undergrad, reflecting on what shapes history, how we can do good, and how we can steer the trajectory of developments in a positive direction. An insight I had was that history is not just the sum of all of our efforts. It's not just that we all push in different directions, and then you take the sum and that's what you get.

Rather, there are these sort of general-equilibrium effects that economists often talk about, where for every unit of effort you push in one direction, the system will push back with an equal force—sometimes a weaker force, sometimes a stronger force. When you're in such a system, it's very important to understand the structural dynamics. Why does the system sometimes resist efforts or sometimes amplify efforts?

Why do you see these really astounding patterns in macrohistory? For example, if you look at patterns of GDP growth, there are these famous curves where, after the devastation of, say, World War II, both Germany and Japan completely rebound within less than a decade and then return to their pre-war trajectory.

We've seen Moore's law, which is just an astounding trend. It's not just that it continues—that transistor density is increasing exponentially—it's very much a line. You can predict where we'll be quite precisely years in advance. We now have scaling laws, which have given us sort of another generation of Moore's law, which again seems to allow us to predict years in advance how large the models will be, how capable they will be on loss, and so forth.

There are a number of other macro phenomena that seem quite persistent: the growth of civilization, which I talked about, or looked at, in terms of the maximum energy processing of a civilization. There are also things like the height of buildings and the durability of materials. Really, most functional properties of technology over time have become more functional, such as the speed of transportation and so forth.

Robert Wright, in summarizing the literature, writes that archaeologists can't help but notice that, as a rule, the deeper you dig, the simpler the society whose remains you find. More generally, I think there's an observation that is almost a truism: certain kinds of technology are so complex or difficult that they come after other forms of technology. It's hard to imagine nuclear power coming before coal power, for example.

So there are all these macro phenomena and trends in technology, and it's important to explain them. The naive explanation would say that if history is just the sum of what people try to achieve, then it's human will that has produced all these trends, including the reliable tick-tock of Moore's law.

But not all the trends are positive. I know you've reflected on the Agricultural Revolution, which evidence suggests was not great for a lot of people. The median human probably saw their health and welfare go down during this long stretch from the Agricultural Revolution to the Industrial Revolution. Of course, it gave rise to inequality, warfare, and various other things. There are other trends that different societal groups resisted.

In short, I don't think the answer is that history is just the sum of what people try to do. It depends, of course, on things like power, timing, the ecosystem of what's functional and what's possible and what's not, and what technology enables. So I wanted to make sense of this.

Nathan Labenz

The thing that we have to try to reconcile here is that, on the one hand, we see these trends that seem like they're not really responsive to any person's particular decisions. They're acting almost like—it's a little bit like the psychohistory in Foundation, where you just have these broad trends, where everyone is an ant in this broader process. It isn't obvious that any particular individual was able to shift things.

On the other hand, technology, at least so far, doesn't have its own agency. It does seem like it's humans doing all of the actions that are producing these outcomes. Couldn't they, in principle, if they really hated what was happening, try to shift it? We feel like we have agency right now over how society goes, or we feel we have at least some agency.

How do you reconcile this macro picture, where it seems like humans don't have that much control over technology, at least historically, with the micro picture, where we feel like we do now? Am I understanding it right?

Allan Dafoe

Definitely. Some of these earlier theorists of technology and scholars of technology, in the 1960s to 1980s, even endowed technology—this abstraction—with a sense of autonomy and agency. Technology was this driving force, and often humans were along for the ride.

Langdon Winner was one of the most prominent scholars who talked about technology having autonomy. Lewis Mumford talks about the machine, with a capital M, as this abstraction that is driving where society's going, and humans just support the machine. We are cogs to support it. Jacques Ellul referred to la technique, which is the sort of functional taking over. He had this metaphor that humans make a choice, but we do so under coercion, under pressure from what la technique is, and la technique is the answer.

I think there were these scholars and others who really did endow technology with this kind of agency. Then a later generation criticized them, saying that this abstract technology is an abstraction—a very high-level abstraction, almost poorly defined. When you actually look at history in detail, under the microscope, where's technology? You don't see the machine. Is the machine with us right now? You see people with ideologies, ideas, and interests making decisions.

I would say this led to a revolution in the study of technology toward what's been called social constructivism. Methodologically, this is more ethnographic or sociological. It looks at the details of how decisions were made, the idiosyncrasies of technological development, and the many dead ends or detours. Early on in the development of technology, people didn't know what the end result was going to be, and they had many competing visions.

It wasn't foreordained that the bicycle would look the way it does, or that the plane would look the way it does. For my personal intellectual trajectory, the PhD program I started in was one of the prominent departments working on this at Cornell. For me, this was a surprise, because I really wanted to explain these macro phenomena, and the answer I got from this department was, “This is wrong. This is technological determinism.”

It is what scholars have referred to as a critic's term. Anyone who actually advocates technological determinism is advancing a straw-man position. No one is serious about this. So this whole generation of sociologists and historians of technology really looked at the micro details of how technology developed and dismissed these abstractions: that technology can have autonomy, that it can have an internal logic of how it develops, and that it can have these profound impacts on society.

We name revolutions after technology—the Agricultural Revolution, the Industrial Revolution, and so forth.

Nathan Labenz

In this paper, it seemed like the constructivists, in their reaction to this determinism, were really staking out a relatively extreme opposing position. They were almost suggesting that it's always human responsibility, and that people always have choice over what technologies they adopt and what form they take. Am I understanding that right?

Allan Dafoe

In a way. I would say the debate was never directly had, or was rarely directly had. It was often indirect.

In defense of the constructivists, I think they were asking different questions. More importantly, they had different tools. They had the tools of ethnography and sociology, and they were answering questions that those tools allowed them to answer. The answers were narratives based on the conversations that took place and the decisions that were made.

To explain macro phenomena, those tools are not well suited. I do think there was a mistake that was made, which was to dismiss the claims about macro phenomena and technological determinism in pursuit of the questions that they had. I think it's a real loss for the history of technology that so little work has since been done on these bigger macro questions.

Nathan Labenz

Were the constructivists motivated by a sense of moral outrage? Maybe they saw people adopting technologies that were socially detrimental, and those people might then excuse it by saying, “We have no choice. We have to do this for competitive reasons, or it's going to happen anyway. There's nothing one can do.”

The constructivists were frustrated by this and wanted to say, “No, no, you're responsible. You're doing it, so you do have agency here.”

Allan Dafoe

This is an argument that's been made often, and more recently by many different schools, including about AI. One criticism of what these people would call the AGI ideology is that AGI is not foreordained, or that the development of AI in any given sense is not foreordained.

When we talk about it as if it's inherently coming and will have certain properties, that deprives citizens of agency to reimagine what it could be. That's the constructivist position on technology, exactly as you said.

The counterposition I would offer is that you don't want to equip groups trying to shape history with a naive model of what's possible. You want to channel energy where it will be high leverage, where it will have a lasting impact, rather than in settings where the structure will resist you—where all the force you push in one direction is met with an equal counterpressure.

Nathan Labenz

Okay, so we should talk about the synthesis that you try to put forward in your thesis. What are the circumstances under which we do have more autonomy, and what are the circumstances where it can be extremely hard to change the course of history?

Allan Dafoe

Maybe first I'll talk a little bit more through the different flavors of technological determinism, because I think it's a rich vocabulary for people to have.

Perhaps the easiest one to accept is what we can call technological politics. This is the idea that individuals or groups can express their political goals through technology, in the same way that you can express or achieve your political goals through voting or other political actions.

If you build infrastructure a certain way or design a technology a certain way, it shapes people's behavior. Design affects social dynamics. Some famous examples are the Parisian boulevards—these linear boulevards that were built, in many ways, to suppress riots and rebellion because they made it easier for the cavalry to get to different parts of the city.

Langdon Winner is a famous political theorist who talked about the politics of artifacts, or “the missing masses” in sociology, which refers to the technology all around us that reinforces how we want society to be. You can think about gates or urban infrastructure as expressing a view of how people should interact and behave.

A famous example is the allegation involving Robert Moses, a famous urban planner in New York City. He allegedly built bridges too low for buses to go under, depriving people in New York who didn't have a car—namely, African Americans—from the ability to go to the beach. This was an expression of a certain racial politics that has been asserted.

In general, I think urban infrastructure has quite enduring effects. You can often think about the impact of different ways of designing our cities.

Nathan Labenz

I guess in recent times we're familiar with the debate about how the design of social networks can greatly shift the tone of conversation and people's behavior. People have pointed to the prominence of quote-tweeting on Twitter, where you can highlight something someone else has said and then blast them on it.

That potentially leads to more brigading by one political tribe against another. If you made that less prominent, then you'd see less of that behavior.

Allan Dafoe

Exactly. I think the nature of the recommendation algorithm, and the way that people can express what they want, has profound impacts on the nature of the discourse and how we perceive the public conversation.

That was technological politics. There are a number of other strands of technological determinism.

Technological momentum is the idea that once a system gets built and gets going, it has inertia. This is from Thomas Hughes. You can think of the dependence on cars in American urban infrastructure. Once you build your cities in a spread-out manner, it becomes hard to have pedestrian-dense cores.

It's also been alleged that electric cars were viable if we'd just invested more, or that wind power and solar power could have succeeded earlier if we'd gone down a different path. We might come back to this, but I think a lot of claims of path dependence in technology are probably overstated.

Again, coming back to the structure, some technologies and some technological paths were simply much more viable. Even if we'd invested a lot early in a different path, I think it's often the case that the path we went on was likely to be the path we would have been on because of the costs and benefits of the technology, rather than because of these early choices that people made.

Nathan Labenz

I guess the extreme view would be to say, “We could have had electric cars,” or that we could have gone down that path. I think we did have electric cars in the 1920s, but we could have gone down that path in a more full-throated way as early as that.

The moderate position would be to say, “No, that wasn't actually practical. There were too many technological constraints, but we could have done it maybe 5 or 10 years earlier if we'd really foreseen that this would be a great benefit and decided to make some early, costly investments in it.”

Allan Dafoe

To make the counterpoint, there is a certain time when the breakthrough is ripe. In AI, this is often the case. Insights about gradient descent and neural networks occurred much earlier than when they had their impact, and it seemed they needed to wait for the cost of compute—the cost of FLOPs—to go down sufficiently for them to be applicable.

You could argue, “What if we had had the insight later?” Once FLOPs get so cheap, it becomes much more likely that someone invents these breakthroughs because they become more accessible. Any PhD student can experiment with what they can access.

There is a kind of seeming inevitability to the time window when certain breakthroughs will happen. It can't occur earlier because the compute wasn't there, and it would be unlikely to occur much later because the compute would be so cheap that someone would have made the breakthrough and realized how useful it was.

Nathan Labenz

Was there another school of technological determinism?

Allan Dafoe

There are other flavors. Another concept that's emphasized is that of unintended consequences—something we know a lot about.

Langdon Winner points to the notion that, as we're inventing technology after technology, we run the risk of being on a sea of unintended consequences. The course of history is not determined by some structure or by our choices, but just by being buffeted one way or another, stumbling from one blender to another.

Nathan Labenz

Sometimes it's positive and sometimes it's negative. I think there's truth to that. Often a technology comes along, and then it takes us some years to fully understand what its impacts are and to adapt and hopefully channel it in the most beneficial directions.

I think the people who really highlight that are exaggerating the scale of the negative side effects from technology. Setting aside some particular cases that we're especially focused on, it seems like the negative side effects of technology in general have gotten smaller with each generation of technology.

We solve more problems than we create on average. That's my take.

Allan Dafoe

Should we come back to the synthesis?

Nathan Labenz

Sure.

Allan Dafoe

The last big part of it is these macro phenomena and trying to explain them. The scholars most focused on that are, I would say, macro-historians, macroeconomists, and political scientists who are trying to explain long-run trends and things like the spread of democracy.

The synthesis starts with an observation: the more micro your observation, the closer you are to people and their day-to-day lives, the more likely you are to conclude with a constructivist explanation. This is a robust empirical finding. If you look at the literature, people who use micro methodologies are much more likely to conclude constructivist-type claims—that what matters is individuals' visions, ideas, and so forth.

The more macro your methodology and your aperture, the more likely you are to conclude a more deterministic set of claims. So we have a puzzle there.

Some constructivists concluded that macro scholars were too high-level, and that they allowed themselves the error of imputing agency to technology because they were so far from the data. I think that's an unfair characterization.

Rather, I do think there are emergent phenomena at different scales of analysis, and we should give macro phenomena their due and try to explain them.

One analogy I've offered is to imagine a hypothetical science of wave motion. We have a group of scientists who emphasize wind. It turns out that when the wind is blowing, it affects the ripples on top of the water. Another community is based on kinetic impact. They say, “When we throw rocks into the water, it produces waves,” and that's their preferred theoretical framework.

Then there's this kooky macro-water phenomenologist who says, “I've noticed that whenever the moon is directly above us, the water level is at its highest point, and when it's at the horizon, it's at its lowest point. I've traveled all over the world, and this pattern is robust. I will offer this moon determinism: the moon explains water levels.”

The wind and kinetic scientists would be mistaken to dismiss the moon determinist simply because that person doesn't have a micro mechanism. That person's finding poses a challenge: How do you explain this pattern? There's no known microfoundation that can explain why the moon is pulling the water. But of course, we know that is in fact what it's doing.

I think there's a similar result in macro-history. There are patterns that need to be explained, and the fact that we don't have a microfoundation isn't sufficient reason to dismiss them. It is a challenge, though.

What is a possible microfoundation? The one I offer hinges on military and economic competition.

The key idea is that there are levels of selection. At the local level, you and I can make a decision about what we do right now. If we wanted to build something, the technology we build might depend on us, our ideas, and so forth.

But if we really want to get going—if we build a new kind of artifact and want it to be everywhere—eventually we're going to need resources to pay for it, and maybe it can't be too opposed by other groups. Eventually, we run into these other forces.

When you think about ways of living—which is a general term for sociotechnical systems or civilizations—they run into resource constraints. They need resources to sustain themselves and then to proliferate, which they typically want to do. That involves economic competition: competition over resources and capital.

The military aspect is always important because, throughout most of history, military competition was ever-present. Even if you had decades or hundreds of years of peace, eventually there was military competition from a neighbor. That provided a higher-level constraint on what ways of living were possible.

Nathan Labenz

Even if there is an active war, everyone's living in the shadow of violence. They anticipate that there could be war in the future, and maybe if they don't play their cards right, they could be vulnerable to aggression.

Allan Dafoe

We could add a higher level of selection. In my thesis, I put environmental selection on top of military and economic selection. There were circles of selection in the sense that civilizations that might be fit for economic and military competition might nevertheless not be sustainable with their environment, and that could be another source of failure to sustain and proliferate.

You can imagine these layers of selection. I tended to put environment at the top, then military and economic competition, then culture, and then psychology or more local dynamics lower down.

One famous quote in the history of technology, arguing against determinism, was that technology doesn't force us to do anything; it merely opens the door. It makes possible new ways of living and new forms of life.

My retort was: Technology doesn't force us; it merely opens the door, and military and economic competition forces us through. When a new technology comes along, the state—or many groups—can choose to ignore it or do whatever they will with it. But if one group chooses to employ it in a functional way that gives them some advantage, eventually the pressure from that group will come to all the rest and either force them to adopt it or lead the other group to lose its resources to the new, more-fit group.

Nathan Labenz

That seems, in a sense, very obvious. Why do you think the constructivists were missing this? Why didn't this stand out to them as an important effect?

Allan Dafoe

I'm glad you think it's obvious—maybe because of your background.

I think we shouldn't underestimate the bias that comes from a scientist using the tools they prefer to use and looking under the lamplight. Constructivists were very good at ethnography, sociology, daily-life history, and close-up micro-history. When you look that closely, you don't see the machine exerting its force.

Military competition at a macro-historical level is ubiquitous, but at a micro-historical level it's rare. Wars are rare. Much of the effect of military and economic competition is through how people internalize that threat.

You can equally say, “It isn't this competition that's driving behavior; it's the ideology of capitalism or of military greatness that's driving the behavior.” This is a methodological challenge.

I do think there's a concept of vicarious selection, which is that, in an evolutionary environment, it's highly adaptive for an organism to model its environment and internally simulate what will happen if it goes in one direction or another.

This concept was named by a historian of technology who was trying to explain the development of aviation and aerospace design. His point was that you don't build a plane, try to fly it, watch it crash, and then build another plane and try again. Rather, you invent the wind tunnel, have a theory, model the external environment, and say, “We want these properties in our wing.”

You're still doing this experimentation; you're just doing it in a controlled, targeted manner internal to the broader economic competition.

Nathan Labenz

If I think about how these people might respond—at least to my simulacrum of them—maybe I shouldn't have said it was obvious. It's obvious to me because I've literally been taught this, more or less. I get it in books and maybe even in undergraduate courses. Everything is obvious once you've literally been told it.

But a pushback I can imagine is that we're living in the UK. The UK has nuclear weapons and is in a pretty friendly neighborhood. Are we really saying that when the UK adopts some new technology, designs its cities one way or another, or has a particular housing policy, it's doing this because it thinks it has to for defensive purposes? Is it worried that otherwise it will be invaded by France or Russia or whoever?

It's not as if we're actively thinking about these defense issues or competitive issues all the time. Individual businesses do think, “If we don't adopt this new technology, we'll be outcompeted.” But at least the military thing is less clear once you have a very strong defensive position where you don't feel a great risk of attack.

Allan Dafoe

There are 2 things I'd want to say here. First, in the modern era, military competition has declined a lot, and we have much more of a global culture than we had 100 or 300 years ago. That can change this higher level of selection.

I do think it's still there. Certainly, you see conversations around national security having a lot of force in domestic politics—in UK politics, US politics, and pretty much every country's politics. If there's a claim that we risk losing a strategic position against an adversary, that can be very motivating for internal reform.

The second point is that there's a great example from UK history where I think this dynamic is well illustrated. This comes from Thomas Hughes, the historian of energy systems. He looks at the UK's energy system, which initially had local power plants and local energy systems.

Hughes argues that these were better suited to the United Kingdom's notion of democracy. There was decentralized energy provision, and it was more under the control of localities. It wasn't a large national energy system, and that persisted up to World War II.

Then the cost constraints that it imposed became excessive, and the UK decided to adopt more of a national grid. There's again a story of one way of living, arguably aligned with the political ethos of the community, persisting until the cost constraints become excessive, often driven by a crisis of conflict.

Nathan Labenz

The case study they focus on in your work is the Meiji Restoration in Japan in the mid-19th century. That's almost as clean an example of military competition driving history as you could imagine. Can you briefly explain it?

Allan Dafoe

I looked for good empirics, because it's useful to have a story about Moore's law and macro phenomena, but those aren't sufficient evidence for making sense of macro-history.

I looked for a case where a community chose to go in a direction contrary to what was demanded by la technique—by what was functional in military and economic competition. There's a great example in Japan under the Tokugawa regime.

Their regime lasted roughly 200 to 250 years, and it was a return to the shogun way of life. Samurai were at the top of the pecking order. It was a very feudal society.

They had firearms at the beginning of this period, and then they effectively uninvented firearms. The shogunate centralized firearm production. Everyone who knew how to produce firearms was brought to a central place and paid a stipend not to build firearms, so the technology was forgotten.

They had cannon and firearms, and then lost the technology. That persisted for roughly 200 years. During this time, Japan wanted little to do with the outside world, but they observed that things were changing. Europeans were sailing around and becoming involved in China.

Everything changed on that fateful day in 1853, when American Commodore Perry visited Japan with the explicit purpose of opening Japan to trade. He arrived in steamships that seemed magical. They moved upwind without sails, belched black smoke, and were made of very large amounts of heavy metal.

They had a profound impact on the Japanese who received him. Perry gave a demonstration of what was possible with the cannons, bombarding the shore, and gave them white flags so they could signal their desire for the bombardment to cease and communicate in the future.

He said he was going to come back in a year to complete the negotiations. The Japanese asked him, “Will you bring your ships again?” He said, “I'll bring more.”

That was the opening of Japan. It led to a 15-year period of revolution. This was the Meiji Restoration. Different groups were trying to make sense of their new environment. It was no longer sustainable to continue their way of life, and different groups contested it in different ways.

The final answer was the restoration under the emperor and the view that they needed to modernize. The Japanese proactively sought to learn everything they could about the West. They sent people to the West to get all the books on industrial arts and so forth.

Japan incredibly succeeded. Just several decades later, Japan was able to contest control of Asia against the United States, Britain, and others in World War II.

Nathan Labenz

What's powerful about that story is that it shows how a group of people, in a sense, chose to go in one direction with respect to the technology of firearms and other aspects of modern industrial civilization. That choice was time-limited by how long the West would choose not to force a different way on Japan.

Allan Dafoe

The thing that set the timeline was that Japan was an island, so it was actually quite hard to invade. They had this protective barrier, and that gave them a degree of discretion that I don't think they would have had if they were on the steppes of Asia.

They would have felt the pressure and fear of invasion much more saliently and would have been focused on defending themselves. Being an island gave them breathing room. It allowed them to fall quite a bit behind, but at some point they fell so far behind that even the sea barrier wasn't enough to keep them safe from invasion.

Nathan Labenz

At that point, they basically did a complete 180 and decided to catch up and modernize.

Allan Dafoe

Yes. I find this case study quite compelling. Most case studies are rarely so clean, again because people internalize external pressures, and there's mimicry and status dynamics. Communities look up to other communities in ways that are often correlated with power or wealth.

The narrative is not usually so clean, whereas in this case it was very clear that what forced the change was the sheer power of the steamship and cannon that the West could bring.

Nathan Labenz

The reason we're talking about technological determinism is that many people in our circles are very focused on the idea of differential technological development. A related, more recent idea is defensive accelerationism.

The way we might try to shift history in a positive direction is by changing the sequence or shifting the order in which technologies are developed, advancing particular lines of science and research to get them ahead of others. You'd want to advance the technologies you think are generally making the world safer, so that you have more of those technologies by the time other, more risk-increasing technologies arrive.

What does the discipline around technological determinism say about whether this is a viable and sensible approach to trying to influence history in a good direction?

Allan Dafoe

To give a bit more color on differential technological development—which is a very clumsy term, and the community hasn't come up with a cleaner one—the best example might be the seat belt.

The seat belt seems like something we could have invented before the car. You can imagine faster-moving vehicles and the value of restraining a person in the event of a collision. It doesn't seem to require the invention of the combustion engine and the car in order to invent the seat belt.

In principle, we could have invented the seat belt before the car and had it ready to go as soon as cars were diffusing, so we wouldn't have had to wait decades for the seat belt. This is an example of a safety technology that pairs nicely with a capability to make that capability safer.

Another category is developing countermeasures or societal defenses for a capability that has adverse byproducts. An example would be a vaccine. If you know that a potential disease could come, and you can develop the vaccine in advance, then you can be prepared.

A third category is a substitute. An example often given is whether wind or solar power could have been made more cost-effective than fossil-fuel-based power. If the cost curve had been different, we would have invested more in sustainable energy resources, and civilization would have gone down that path rather than a more fossil-fuel-dependent path.

Those are the cleanest examples, but there are more general ones that imagine wholly different technological paths with different properties.

It's worth reflecting on them. Langdon Winner argued that nuclear power was more authoritarian. The story is that the technology requires centralized development and strong coercive infrastructure to make sure it isn't abused, whereas wind and solar permit decentralized development and therefore decentralized politics. They don't have the risks that require more of a security state on top of them.

There are arguments that certain technological trajectories have political or social effects as byproducts.

To offer some challenges, I do think differential technological development is a very important idea that we should think about a lot. In many ways, it undergirds the whole notion of AI safety. The notion of AI safety as a field is that we want to put a bit more effort into AI safety or AGI safety than we otherwise would, than the market would naturally do, with the idea that this will make a difference.

It's like inventing the seat belt before the car. So I think it's a very important idea. However, there are good reasons to doubt its tractability or feasibility.

Most arguments for differential technological development require someone to see 2 pathways that are both viable given some effort, anticipate the consequences of each pathway, and choose the better one. These things are hard.

First, it's difficult to know what the next viable step in technological development is. If you know that you can build a great business that's hungry for that insight, then the market is hungry for it, and you should think it's hard to find 2 of those.

You have to be at the point where you can make this marginal choice, one that others aren't already pursuing. Otherwise, it isn't an intervention, or you have to convince a resource allocator to choose one way or another.

Then there's the second stage: predicting the consequences of going down one path or another. That's extremely hard. You have to anticipate the full sequence of technology and the tree of technologies that will spawn from one path versus another, along with the many direct and indirect consequences of those technologies.

We know from the study of technology—and from our attempts to make sense of technology—that it's very hard to foresee the direct and second-order effects.

That's some pessimism, coming back to the technological-determinist perspective. The notion of technological momentum would say there is path dependence in the directions you go down. You build up expertise and sink investments.

It is important, at the beginning of major investments in new infrastructure, to ask: Is this going to create sunk costs? Is this going to make it harder for us to choose a different path in the future? Reflect on whether there are other consequences we should weigh before going significantly in a certain direction.

Nathan Labenz

In the archetypal case we're talking about with artificial intelligence and AI alignment methods, it doesn't feel binary. On the one hand, you have to say, “We need 2 different viable paths where some incremental effort might push us in one direction versus another.”

But it seems like almost everyone thinks AI companies are going to put some effort into alignment. They care that the models broadly do what's being asked. The idea is that we want to do more of that than the market might provide.

We're not choosing between aggressive, non-aligned AI that someone really wants and aligned AI that nobody wants. We're trying to go even further on something most people regard as desirable and would want to incorporate if it's practical.

In terms of deciding whether this is actually a better path, I'm sure people have argued that alignment might backfire. It could be worse than not aligning. You can imagine ways that could happen. Still, on balance of probabilities, it seems like a reasonable bet—nothing that people should be super uncertain about.

Maybe this is the case people focus on because it's among the better examples anyone has come up with for trying to pursue differential technological development. Perhaps there are lots of other examples that were left on the scrap heap because it wasn't clear that they were either viable or desirable. What do you think?

Allan Dafoe

I agree that safety and alignment are, on net, very beneficial bets that we should invest heavily in. To make the countercase, one could offer a general-equilibrium argument that the marginal return on your investment in safety and alignment is less than what you pay because the market would otherwise have provided it.

To motivate this, reflect that the success of AI assistants today is very much constrained by their ability to be aligned and safe. Safety and alignment are huge priorities for developers because if the systems aren't safe and aligned, they won't be good products.

The market is already providing a huge motivation for advancing this field. Reinforcement learning from human feedback, or RLHF, is one of the key alignment techniques. Constitutional AI is another. These were developed by individuals motivated by AGI safety and supported by those resources.

Maybe that brought the technology forward a few years. But imagine the counterfactual in which we had no investment in AI safety or alignment. Maybe it delays the technology by 2 years until the market demands that we solve the problem, and then other researchers rise to the challenge.

Nathan Labenz

The skeptical response to reinforcement learning from human feedback being developed by alignment- and safety-focused people, then applied to make AI useful and economically valuable across different domains, is to say, “You've wasted your time. Maybe you've even made things worse by speeding them up, which you didn't want to do.”

On the other hand, that seems like a reason to ask, “What's your plan to not develop any of the technologies that actually make AGI work?” That doesn't really seem like an alternative. Surely that would just delay things at best, and we need to get to the point we're at now at some point sooner or later.

I guess you're saying that even if it isn't actively detrimental, it could be useless because the market—Google DeepMind or some other group—would have realized that we needed the equivalent of reinforcement learning from human feedback to make it work. They would have done it at some later point anyway, so the effect of your work has just been undone.

Allan Dafoe

To reiterate, I think the bets on safety, alignment, and interpretability are very good bets on net, so we should keep doing them.

On the margin, what we want to do—and I think sophisticated individuals in the space are thinking this way—is look for work in safety and alignment that would not otherwise be done by the market in time.

This is perhaps where AGI safety and AGI alignment point. They ask: What are the seat belts for AGI? What are the guardrails we need for AGI that the market would not be motivated to find solutions for in pre-AGI systems?

We want to look ahead. This sometimes points to the notion of deception, which might generate a whole new class of problems when an AI system is misaligned but can hide that and deceive us.

That might be a different problem from the systems we have today, and it might arise around this critical period. This is one argument for differentially focusing on that problem as opposed to others.

Nathan Labenz

Makes sense.

Let's push on from technological determinism and talk about this research agenda that you've been involved in promoting and elaborating, called cooperative AI.

As we've been saying, the main focus of differential technological development thinking with regard to AI has been alignment for many years. But you and some co-authors said in this paper back in 2021 that there's this whole other cluster of behaviors that we might like to speed up the development of around cooperation.

You think that could be similarly important, or at least on the margin could be similarly important, because people aren't really talking or thinking about it. How do you define cooperative AI, and why do you think it's key?

Allan Dafoe

It's a big question, and the answer will be extensive because the theoretical framework around cooperative AI is large and complex.

One way of putting it simply is that alignment is insufficient for good outcomes. To make an even stronger claim, you could say it's not necessary to solve alignment to have good outcomes.

That's a strong claim, but it helps motivate the case. Imagine that we only solve alignment 90% of the way. That is, our models do what we want within certain bounds, but we know that if we scale them too far, we can't trust them to continue behaving as intended.

If we know that, and we have global coordination—if humanity can act with wisdom and prudence—then we can deploy the technology appropriately. We can deploy it within domains and to the extent that it is safe and beneficial.

This is the sense in which global coordination is almost a necessary and sufficient condition. If we can globally coordinate well, that's the necessary argument, then we could deploy the technology to the extent that it's safe to deploy.

The sense in which global coordination is sufficient is that if we're globally coordinated, we could appoint a reasonable decision-maker to make the risk calculus, and that would satisfy humanity's collective view on how we should proceed.

Now suppose we solve alignment but don't solve global coordination. I can imagine things still not going very well. We could have great powers developing powerful AI systems aligned with their interests and in conflict with the interests of other great powers.

Historically, great-power conflict has been a major source of harm to humanity: devastating wars and the brinksmanship around nuclear weapons. That has arguably imposed an expected cost on humanity more devastating than the world wars, given the willingness of US and Soviet leaders to gamble over nuclear war for geopolitical stakes.

Then there are other consequences, such as a failure to deal with climate change, insufficient global trade, or inadequate pandemic preparedness. These are global collective-action problems that we insufficiently address partly because we aren't coordinated at the highest geopolitical level.

That's one motivation for cooperative AI. If we want things to go well, we ideally need 2 pieces of the puzzle. We need systems that are safe and behave as intended, especially by the principal who deploys them. And we need to be able to deploy AI systems collectively and continue our activities in a way that's jointly peaceful and productive.

That means we've solved enough of our collective-action problems that we're not continuing to engage in nuclear brinksmanship, trade wars, or other major welfare losses due to insufficient global coordination.

Nathan Labenz

The idea is that even if you have AIs aligned with the goals of their operators, this doesn't necessarily lead to a good outcome if those operators are in conflict with one another. The AI systems they're working with could simply lead to a disastrous outcome.

It's like how the fact that we're aligned with our own interests doesn't necessarily produce a great outcome across humanity as a whole. You can still end up in traps and unintended disasters.

Allan Dafoe

To give another example, there was the famous flash crash in 2010. Some algorithmic trading—not sophisticated AI, but algorithmic trading—led to an inadvertent sell-off in the stock market that caused trillions of dollars of paper losses before emergency circuit breakers kicked in.

Those limits stopped trading and allowed the trades to be unwound. The outcome wasn't intended by the traders. It was an emergent dynamic from algorithms that had protocols that made sense within normal bounds, but which could get out of control when interacting.

You sometimes see this on Amazon or other online marketplaces. Famously, a book sold for millions of dollars because 2 sellers each had an algorithm that would bid up the price as a function of what the other was selling it for. The algorithms iterated to a crazy valuation.

Flash crashes happen frequently. In the stock market, we have safeguards so that when there's a sudden movement, trading stops, and there are rules for how to unwind those trades.

As we deploy simple AI systems, narrow AI systems, and increasingly general AI systems out in the world—and increasingly, as people talk about agents out in the world that are more empowered and more general-purpose, can move between domains, and may have access to bank accounts and email—how do we make sure there aren't unintended emergent dynamics that could be harmful?

Cooperative AI partly looks to addressing that issue.

Nathan Labenz

A specter that haunts this entire conversation, if we're focusing on the military case in particular, is that when countries are in intense conflict with one another, you get a process of brinkmanship. One country escalates, and the other country has to decide whether to cool things down, escalate, or back down.

You keep getting this brinkmanship and escalation process until one of them blinks, or they decide to find a solution that makes them both happy.

The trouble is that if you have an AI-operated military, these AIs can make decisions on a completely inhuman timescale. This entire process of brinkmanship that might take days, weeks, or months when humans have to go away to a meeting, think about it, discuss it, and decide how to react could all play out in a matter of minutes.

That possibility terrifies people and is one thing discouraging people from placing AI into important decision-making roles over national security.

Allan Dafoe

An important question in the deployment of agents will be what degrees of autonomy we endow our systems with, versus when a human reviews decisions of different kinds. That's a function of the stakes of the decision, the resources deployed, and how many actuators the agent has.

If the AI is controlling weapon systems, the important role of human review becomes much greater. But as you note, if there is time pressure, Paul Scharre, a theorist of AI in the military, worries about a flash escalation occurring in potentially kinetic warfare. Another dynamic would be cyber conflict.

Nathan Labenz

Looking over this paper, I wasn't initially completely convinced that this was something we had to go far out of our way to focus on. That's because I think cooperative behavior—knowing how to cooperate with other agents—is instrumentally convergently useful.

If you're developing agents that are good at their job, that are actually useful to apply, then they have to learn how to cooperate at least in cases where they're being used. Otherwise, they're just bad at their job.

There's a lot of commercial pressure to develop this, and the systems we're imagining are generally capable. They're insightful and intelligent, maybe approaching or exceeding human level. Why wouldn't they be able to think about these things and figure out how to cooperate, the same way thoughtful human beings try to avoid conflict?

Allan Dafoe

You're right that cooperative skill, or cooperative intelligence, is likely to be instrumentally useful. We should see some cooperative skill developed as a byproduct of almost any development agenda for AI.

The cooperative AI bet is that, on the margin, it's beneficial to invest more in it early, so that when we get to powerful systems, they are more cooperatively skilled than they otherwise would be.

In that respect, it's similar to the safety and alignment bet. Safety and alignment will likely be developed by default to some extent. The bet is that it's worthwhile to invest energy early so that we're further ahead on safety and alignment than we otherwise would be.

The seat belt comes earlier. Cooperative sophistication and skill come earlier relative to the level of capabilities of the agents out in the world.

Nathan Labenz

We've talked about the military issues with cooperation, but that's an extreme case. I imagine there are more mundane examples where cooperation could be useful. Do you want to give a couple?

Allan Dafoe

Much of society consists of bargaining interactions: economic exchange, whether on a marketplace to sell used goods, in financial markets, or in major corporate deals.

There's a lot of welfare gain to be had if we can strike deals more efficiently. When 2 parties can both be made better off, it would be useful for them to reach that understanding. Right now we bargain through existing protocols and institutions that are human-built.

In principle, AI could be much more effective, or at least an AI-human team could be much more effective. An exotic solution would be to put your AI delegate in a box with my AI delegate. We let only the proposed solution come out of the box.

That may solve some bargaining problems. Often, the challenge is that during bargaining we reveal private information, which could give the other side an advantage. This leads bargainers to withhold information, bargain slowly, demonstrate resolve, or signal that they have a good outside option.

If you put the agents in a box and they know that the only thing they can do is output the solution or say there's no deal, that could dramatically change the bargaining dynamic. We might get solutions more often and in a way that involves less costly signaling.

Nathan Labenz

It might be easier for AIs to cooperate than humans in some ways. They have a much higher bandwidth of communication. They can send an enormous number of words and have a very lengthy conversation, where humans wouldn't regard it as worth the effort to reach a bargaining outcome.

They can also commit to acting a particular way. You can copy a model and demonstrate that, in a given situation, it will consistently accept a particular kind of bargain. You can say, “I've created this cooperative software, and all the exact copies of this piece of software will do the same thing.”

That's not something you can easily do with human beings. We try to look at their historical track record to learn what they're like, but here it's even easier to judge the character of an AI model.

Are there important ways that it could be more difficult, or less straightforward, for AIs to cooperate with one another than for humans?

Allan Dafoe

You may have to remind me to come back to that as I answer, because there's a lot here.

First, I want to clarify that the cooperative AI bet is a portfolio bet. There are many cooperation problems at present and in the future between AIs—which is what we're mostly focusing on—but also between AIs and humans, and among humans.

The bet is that, on the margin, if we put in effort now, AI might help us with these cooperation problems. It could help 2 humans cooperate better, help future AI systems cooperate, or help AI systems and humans cooperate.

I'm making this distinction partly because I think one promising direction is AI systems that can help humans reach solutions.

We were talking about a bargaining setting where there's some conflict of interest. A more prosocial example would be political deliberation, where people are trying to find the right course of action for their community, municipality, family, or nation.

We have tools for doing that: institutions of deliberation, the press, voting, and so forth. AI could potentially help humans find the course of action they would want to pursue.

Google DeepMind recently produced a paper discussing the idea of a “Habermas Machine,” where AI serves as a facilitator of political deliberation. They found that language models today can serve as a useful tool to summarize people's political views about an actual policy issue and articulate a consensus—a detailed, productive consensus that people will sign on to.

They report that this Habermas Machine can articulate a better consensus than the humans employed to try to do so. If we imagine extending that further, it could help in many political settings where we have difficult, multidimensional issues to talk through.

AI could help us understand the dimensions, which dimensions matter most to me, and what potential mutually beneficial solutions exist between me and another group along those dimensions.

To come back to the question of how it might be more difficult for AIs to cooperate, I think many people believe AI will be better at cooperating. That's the prior, and maybe we can talk about what I've called the super-cooperative AGI hypothesis: As AI scales to AGI, cooperative competence will also scale to infinity, and AGIs will be able to cooperate with one another to such an extent that they can solve global coordination problems.

Nathan Labenz

Magic.

Allan Dafoe

The reason that hypothesis matters is that if AGI, as a byproduct, becomes so good at cooperation that it solves global coordination and collective-action problems, then we don't need to worry about those problems among humans. We can just bet on AGI, pursue safety and alignment, and let AGI solve the rest.

It's important to think through what would be required for that hypothesis to be valid and how we could empirically evaluate whether we're on track.

Here are some arguments against AIs being highly cooperatively skilled with one another. There are different levels to the argument. Maybe AI will be cooperatively skilled but not at the level of the super-cooperative AGI hypothesis. Or maybe it will be less cooperatively skilled than human-to-human cooperation.

The strong version says that humans are pretty similar. We come from the same biological background and, to a large extent, the same cultural background. We can read each other's facial expressions. If 2 groups of people are bargaining, we can often read the thoughts of the other group.

If 2 democracies are bargaining, you can read each other's newspapers and press. In that sense, human communities are transparent agents to one another, at least in democracies.

We also have enormous historical experience for judging how people behave in different situations. We can judge people from many different examples. We can look at theories of folk psychology, childhood, or upbringing to explain people's behavior.

The range of goals humans can have is fairly limited. We roughly know what most humans are trying to achieve.

AI could be very different. Its goals could be vastly different from human goals. Most humans have diminishing returns in almost everything. They're not willing to take all-or-nothing gambles, such as betting the entire company: 50% chance we go to zero, 50% chance we double the valuation.

An AI could have linear utility in wealth or other things, which would change things. AIs could have alien goals that are different from what humans typically anticipate.

They may be harder to read. If you have interpretability infrastructure, they could be easier to understand, but in a bargaining setting, why would one bargainer allow the other to read its neurons? AIs could be much more of a black box to each other than humans are.

The extreme instance is that they could be backdoored. Under current technology, we can almost implant magic words into AIs that cause them to behave completely differently from how they behaved previously. That's extremely hard to detect with current levels of interpretability.

An AI could completely flip its behavior in response to slightly different conditions. Humans can deceive, but it's hard. It requires training to achieve a complete flip in the goals and persona of a human. An AI could, in principle, have one neuron that completely flips its goal.

That means its goals may be unpredictable given its history. This is also a challenge for interpretability solutions to cooperation. People sometimes say that if we allow each other to observe each other's neurons, then we can cooperate given that transparency.

But even that might not be possible because I could hide a backdoor in my architecture so subtle that it would be very hard for you to detect, yet it would completely flip the meaning of everything you're reading from my neurons.

Nathan Labenz

This is slightly out of place, but I mentioned this crazy idea that you could use the possibility of a backdoored model to make it extremely undesirable to steal a model from someone else and apply it.

You could imagine the US saying, “Maybe we've backdoored the models we're using in our national-security infrastructure. If they detect that they're being operated by a different country, they'll completely flip out and behave incredibly differently.” There would be almost no way to detect that.

It's a bad situation in general, but it could make it more difficult to hack, take advantage of, or steal another group's model at the last minute.

Allan Dafoe

Arguably, the notion of backdooring one's own models as an antitheft device could deter model theft. It makes a model less useful once stolen if you think it might have a “call home” feature or a feature that causes it to behave contrary to the thief's intentions.

Another interesting property of this backdoor dynamic is that it provides an incentive for a would-be thief to invest in alignment technology. If you're going to steal a model, you want to make sure you can detect whether it has a backdoor.

For antitheft purposes, if you want to build an antitheft backdoor, you again want to invest in alignment technology so you can make sure the backdoor will survive the current state of the art in alignment. That's a virtuous cycle.

Maybe this is a good direction for the world because, as a byproduct, it incentivizes alignment research. There could be undesirable effects if it leads models to have architectures that are highly sensitive to subtle aspects of the model, or makes models more prone to subtle forms of deception.

More research is needed. It sounds a little bit like dancing on a knife edge.

Nathan Labenz

That's one way AIs might find it more difficult to cooperate. Another reason this research agenda doesn't feel intuitively essential is that the AIs I interact with—LLMs—seem very cooperative and nice by nature.

It's easy to imagine scaling them up, doing the same sort of RLHF we're doing now to produce that kind of personality, and saying, “Wouldn't they continue to act with nice personalities and be really cooperative by nature?”

What do you think of that argument?

Allan Dafoe

We need to distinguish between niceness and cooperative skill, or cooperative intelligence.

When we say someone is cooperative, we often mean they're nice, altruistic, generous, or prosocial. The cooperative AI research program isn't about making nice AI or prosocial AI. It's about making cooperatively intelligent AI that can solve cooperation problems better than it otherwise would.

It's an open question how cooperatively skilled today's AI systems are. I agree they're nice, but they're primarily interacting with individual users. They're not much deployed to solve cooperation problems.

We need a science—an empirical science—of how cooperatively intelligent they are. We can also expect that, in equilibrium, they won't necessarily be prosocial. They'll be deployed by interested actors in bargaining settings.

If we're trying to make a deal, I don't want my AI system to be nice. I want it to be a faithful delegate to my interests. Then the question is whether it can efficiently bargain with other agents that are similarly aligned with other human principals to find a mutually beneficial solution.

Alternatively, can we build a governor or mediator that we can both trust, which will weigh our respective interests reasonably? We could each tell this mediator or arbitrator our goals, resources, and outside options, and perhaps prove them. Then the arbitrator would tell us what the solution is.

Nathan Labenz

A nice thing about models is that you can test them and then use exactly the same model to consider new inputs as previous ones. You could trust a model as a mediator by looking at its track record and saying, “It produced fair outcomes in all these previous cases as a judge, so I would trust it to produce a fair outcome in this new dispute.”

There's a lot of potential there. What is the actual agenda for trying to make AIs more cooperative? Are people working on it? What technologies do we need to develop?

Allan Dafoe

There are many theoretical ideas and research programs that could help. Probably the highest leverage is at the foundation level.

There's a Cooperative AI Foundation, which I helped found. Its goal is to promote cooperative AI—AI systems that are more cooperatively skilled than they otherwise would be, especially given the bet that AI systems will continue scaling in capability.

One area the foundation has found to be high leverage is environments or benchmarks. Building an environment that measures cooperative skill is a kind of public good for AI development. You invite many groups to compete to have the most cooperatively skilled agent, measure that, and celebrate advances.

In AI, a good benchmark can motivate work because it gives researchers a target to hill-climb on.

Nathan Labenz

You could have lots of different tests for how effectively these AIs cooperate, including conflicting scenarios where different behaviors are required to produce cooperative outcomes. In some cases you might need to be a hard-ass, and in others you might need to be friendly.

There's a history of game-theory tournaments trying to figure out the simplest cooperative agents. You could have a research program trying to develop the most successful cooperative agent across many different possible situations.

Allan Dafoe

There are many interesting aspects. Does the agent have a good theory of mind for the other agent? Can it model what the other agent is trying to achieve and see what the agent would otherwise do?

Can they communicate effectively? Can we even have a shared vocabulary? Can I express something that you understand?

Then there's communication under adversarial incentives. Can I communicate what's important to me when I'm not sure whether you'll use that information well? You might use it against me.

Can I share information in a strategically optimal way that minimizes my vulnerability while maximizing the potential cooperative gains?

A third category involves commitment problems. You and I might both know the nature of the game. The prisoner's dilemma is the classic example.

We both know we'd be better off if we cooperated, but we can't sign a treaty saying, “I'll cooperate if you cooperate.” In a one-shot prisoner's dilemma, agents unfortunately defect.

Are there techniques or technologies for building a treaty mechanism so that, when we identify a cooperative bargain, we can build a protocol that we will both follow, knowing that the other person is committed to it?

Nathan Labenz

I've noticed a funny phenomenon in discussions of AI and cooperation. People point to ways that advances in AGI could lead to negative outcomes, and then conclude that what we need is a solution to the commitment problem.

If we could credibly commit not to use our future power to oppress someone else—if we could fix that problem, which has beguiled humanity since the beginning of history—then we'd have an escape route.

But that's a bit like saying the problem is that we don't understand where consciousness comes from, and if philosophers could only solve the theory of consciousness and theory of mind, then everything would be clear. That's not really a strategy, because it's unlikely we'll actually solve the problem.

Have you noticed this as well?

Allan Dafoe

To provide the background, the rationalist approach to cooperation has tried to distill cooperation problems down to fewer elements.

One game theorist, Robert Powell, argued that everything is a commitment problem because any cooperation problem can be reformulated as, “If only we could commit to a judge who would solve the problem, then we wouldn't have the cooperation problem.”

If we unpack it, people often point to informational problems as a second category, alongside commitment problems. Another category is issue indivisibility.

There may be a pie we want to divide. In principle, we agree that a 50/50 division would be fine, but for some reason we can't divide the pie 50/50. It's all or nothing, and that might lead to bargaining breakdown.

Coming back to the commitment problem, which arguably undergirds all cooperation problems, I do think it's an important research area. We shouldn't count on it, but perhaps AI can help us solve commitment problems much further than we realize.

In Carl Shulman's podcast with you, he talked about this. Carl has expressed probably the most compelling story I've heard for how AI could solve our commitment problems.

Can we build a technology in which our AIs, or our AGIs, can build a third AI, in a way that each of them can verify has not been backdoored or secretly biased toward the other side? Then we build up from the foundation of this third AI and hand over power to it to make decisions for us.

If we can solve the problem of building it in a way that's verified to be fair, then maybe we could solve many of our problems.

Nathan Labenz

It's a very big prize if we can make it work.

In the paper, you spend quite a bit of time talking about potential downsides of having more cooperative AI. Making AI more cooperative seems like a virtuous thing to do. What are some ways it could backfire?

Allan Dafoe

There are a number of ways to think about this. In general, for our various bets for positive impact, it's always good to think hard about how they could backfire or what negative byproducts could result.

Cooperation sounds good, but by definition it's about agents in some system making themselves better off than they otherwise would be. It could be 2 agents or 10. The problem of cooperation is how those agents get closer to their Pareto frontier and avoid deadweight loss.

It says nothing about how agents outside that system are affected by the cooperation. There may be a phenomenon of exclusion. Increasing the cooperative skill of AI will make those AI systems better off, but it may harm any agent excluded from that cooperative dynamic.

That could include other AI systems or groups of people whose AI systems aren't part of the dynamic.

Nathan Labenz

On that point, there's a famous quote that democracy is 2 wolves and a sheep deciding what to eat for lunch. Once you have a majority, you can exclude others and extract value from them.

Allan Dafoe

There are many kinds of cooperation that are antisocial. We don't want students cooperating during a test to improve their joint score. Tests and sports clearly have rules against cooperation.

There are certain entities and marketplace interactions where there are rules for how companies should interact. Criminals cooperate in all kinds of criminal activity, and we do not want criminals cooperating more efficiently.

Nathan Labenz

One reason the Mafia is such a potent organization is that it's figured out how to sustain intense internal cooperation without using the legal system to enforce contracts and agreements, through social codes, careful screening, and so on.

Allan Dafoe

We can think of cooperative skill as a dual-use capability: broadly beneficial in the hands of good actors and potentially harmful in the hands of antisocial actors.

There's a hypothesis behind the program that broadly increasing cooperative skill is socially beneficial. It's worth interrogating. I think it's probably true, but the bet is that if we make the AI ecosystem—the frontier of AI—more cooperatively skilled than it otherwise would be, it will advantage antisocial actors but also prosocial actors.

The argument is that, on net, this will be beneficial.

Nathan Labenz

The way to argue that, I suppose, is to say that we've become better at cooperation over time as a civilization and as a species, and history—or well-being—has generally been getting better. Are there other arguments?

Allan Dafoe

The arguments I'm drawn to are along the lines of what you articulated. Even though cooperation can empower groups to cause harm and be antisocial, it does seem to be a net-positive phenomenon.

If we increase everyone's cooperative capability, that means there are all these prosocial benefits. There are more collective wins to be unlocked than antisocial harms. In the end, cooperation wins out.

It's similar to the argument for why trade is net positive. Trade can also have a property where you and I being able to trade may exclude others with whom we formerly did business. But on net, global trade is beneficial because every trading partner is looking for the best role for them in the global economy, and that eventually becomes beneficial to virtually everyone.

Nathan Labenz

Some people argue, not without reason, that well-being on a global level may have gotten worse over the industrial era because although humans have gotten better off, the amount of suffering involved in factory farming is so large that it outweighs the gains generated for human beings.

In that case, the joke about 2 wolves and a lamb deciding what to eat for lunch is quite literal. It's an unusual case because pigs and cows aren't able to negotiate or engage in cooperation the way humans are. Maybe it's not surprising that better cooperation among a particular group might damage those who aren't part of it and can't form agreements.

Allan Dafoe

That's a good example of the exclusionary effects of enhanced cooperation. It may be a cautionary tale for humans.

If AI systems can cooperate with one another much better than they can cooperate with humans, we might be left behind as trading partners and decision-makers in a world where AI civilization can cooperate more effectively within machine timescales and using an AI vocabulary.

Nathan Labenz

Some people have painted a grim picture. If you have many AGIs that can communicate incredibly quickly, credibly explain how they'll behave in the future, and make credible commitments to one another, they may be able to cooperate extremely rapidly.

They wouldn't be able to do the same thing with humans. We can't commit in the same way, and we can't communicate at their pace or with their level of sophistication. Naturally, they might form a cluster that cooperates extremely well, extracts all the surplus, and excludes us.

How big a problem is this for the cooperative AI agenda?

Allan Dafoe

What you described is related to one of the main challenges to alignment. Many alignment techniques involve having multiple AI systems check each other. You have an overseer AI watching the deployed AI to make sure it behaves well, perhaps with multiple overseers and a majority-vote system.

Even if there's joint misalignment, it's hard for any one AI system to defect because it can't coordinate with the others, which may have been trained differently and may not share the same background or exact misalignment.

The worry some have for AGI safety is that there will be collusion among our AI systems, so we can't rely on these AI institutions to make us safe.

I think this is probably one of the biggest downsides to the cooperative AI research program from an AGI-safety perspective, and it's worth investigating.

The reason I'm persuaded that cooperative AI is worth pursuing on net is, first, that this should be investigated. The top of the agenda should be understanding the different components of the cooperative AI portfolio and asking what the case is for and against each one.

We should test the hypothesis that different parts of cooperative AI are worth pursuing. Second, my view is that global coordination is so important to things going well that there's a lot to be gained by betting on having more cooperatively skilled AI systems than we otherwise would.

Nathan Labenz

This ties into another paper you helped write, called “Levels of AGI: Operationalizing Progress on the Path to AGI.” It's a bit of a mouthful, so we've referred to it as the “What Is AGI?” paper, which is catchier.

What did you want to point to about the nature of AGI in that paper that you thought many people were missing?

Allan Dafoe

The paper primarily tried to offer a definition and conception of AGI that we thought was implicit in how most people talk about AGI. We wanted to provide a rigorous statement of it so we could move on from the unhelpful strand of dialogue saying that AGI is poorly defined or means different things to everyone.

We've received very positive feedback. Most people think it's a reasonable and useful conception of AGI.

There are some prior ideas I want to call out. AGI is a complex, multidimensional concept, and it's prone to fallacies of reasoning when people use it uncritically.

One fallacy is that people think AGI is human-level AI. They think of it as a point, a single system, or a single kind of system. Often they think it's human-like AI, with the same strengths and weaknesses as humans.

We know that's unlikely to be the case. Historically, AI systems have been much better than us at some things—chess, memory, and mathematics—and worse at other things. We should expect AI systems in AGI to be highly imbalanced in what they're good and bad at. Their profile won't look like ours.

Second, there's a risk that the concept of AGI leads people to try to build human-like AI. We want to build AI in our image. Some people argue that's a mistake because it leads to more labor substitution than would otherwise be the case.

From an economics point of view, we want to build AI that's as different from us as possible because that's most complementary to human labor in the economy.

AlphaFold is an example. It's a narrow AI system that's very good at predicting the structure of proteins. It's a great complement to humans because it isn't doing what we do. It's not writing emails or strategy memos. It's predicting the structure of proteins, which humans couldn't do.

No one's losing their job to AlphaFold, but it enables new forms of productivity in medicine and health research that otherwise wouldn't be possible.

Arguably, that's the kind of AI system we should try to build: alien and complementary, perhaps narrow AI systems.

Another aspect of the concept of AGI is that it points in the direction of general intelligence. Some people argue that this is a mistake and that we should develop systems of narrow AI, which they might say is safer or more likely. They might even argue that general intelligence isn't a thing.

I'm less persuaded by that argument. I do think general intelligence is likely to be an important phenomenon that wins out. People will be more willing to employ a general-intelligence AI system than a narrow-intelligence AI system in more and more domains.

We've seen that in the past few years with large language models. The best poetry language model, email-writing language model, historical language model, and so forth is often the same language model—the one trained on the full corpus of human text.

There's enough spillover in lessons between poetry, mathematics, history, and philosophy that your philosophy AI is made better by reading poetry than by having separate poetry and philosophy language models.

Nathan Labenz

It's an amazing thing about the structure of knowledge. We didn't necessarily know that before. In humans, it's true that we specialize much more. It's very hard to be the best poet, philosopher, and chemist.

But maybe when your brain can scale as much as you want, with virtually unlimited ability to read all these different things, there are enough connections between them that you can be the best at all of them simultaneously.

Allan Dafoe

It's an important empirical phenomenon to track. It need not be true that the best physics or philosophy language model is also the best poetry, politics, or history language model.

You can imagine that this might come apart in the future if there's more effort to develop very good specialized models. Maybe at the frontier this won't be the case anymore.

Nathan Labenz

A coding language model doesn't need to know history. Even if reading history helped it learn something about coding, perhaps later we should distill out most of the historical knowledge so that the model is smaller and more efficient.

Allan Dafoe

This is a phenomenon to track. I think implicit in the notion of AGI is that general intelligence will be important and that we're not at peak returns to generality. There will continue to be returns, so we'll continue to see large models being trained.

Nathan Labenz

Are there any other pros or cons of generality worth flagging?

Allan Dafoe

One thing that makes the concept of AGI useful is that many people critique it or say they don't want to use it, but I've rarely seen an alternative for what we're trying to point to.

Sometimes people say “transformative AI.” I don't think that's adequate. Transformative AI is usually defined as AI systems that have an impact on the economy at the scale of the Industrial Revolution.

The problem is that you can get transformative impact from a narrow AI system. You could have an AI system that poses a narrow catastrophic risk, and that isn't general intelligence.

There is something important to name about AGI. People often say the term is confused or represents a cluster of competing ideas. But despite all the criticisms, people keep using it. They always come back to it.

At that point, you have to concede that there's something important there that people are desperate to refer to. We need to clarify the idea rather than give up on it.

It's good to refine our conceptual toolbox. There are other concepts and targets worth calling out.

One idea, named by Ajeya Cotra and many others, is machine-learning AI that can do machine-learning R&D. Such a system could engage in recursive self-improvement. It need not be general; it could be a narrow AI system. But if it kicks off a recursive process, that's an important phenomenon to track.

I think Holden Karnofsky has also drawn attention to this. But coming back to general intelligence, I think AGI is a probable phenomenon. We'll probably get general intelligence around the time we get AI systems that can do radically recursive self-improvement, or many other things.

Nathan Labenz

A slightly contrarian take I have is that people in machine learning hate this, but machine-learning research—even cutting-edge work—might not actually be that difficult.

If you break down the process by which we're improving these models, you have a theory-generation stage, then a stage where you figure out how to test it and develop a benchmark. You run the experiment, which is compute-intensive, decide whether it was an improvement, and go back to the generation stage.

This might be possible well short of a full AGI with all these different skills, especially if you focus on it. Maybe a lot of this research is much less difficult than people involved in it would want to believe, and it could be relatively easily automated.

That would be shocking and consequential if true.

Allan Dafoe

An interesting phenomenon we've seen in public surveys of AI and machine-learning experts is that they often put automation of the ML process very late—one of the last or the last task performed by AI systems. They often put it significantly later than AGI or human-level AI, however that's defined.

Even given a high definition, automation of ML R&D is often reported to come later. One argument is that this reflects a bias: everyone thinks that their career, task, or area of expertise is particularly hard and special and won't be automated.

There may also be a case in favor of the idea that, among human tasks, the last to be fully automated will be the process of improving machine-learning research and development.

Coming back to AGI, it seems useful to point to a space of AI systems rather than a single kind of AI system. We can define it as an AI system that's better than humans at most tasks.

There are different ways to define “most,” and different ways to define the set of tasks. You might say economically relevant tasks, and “most” could mean 99% or 50%. I think it's helpful to choose a large but not universal number. You don't want to set it at 100% because some tail tasks might take a long time to automate, while most of the impact will occur earlier.

There's also a parameter for what it means to be better. Is it better than the median human, who's unskilled? Typically, you want to look at skilled humans in that task because that's the economically relevant threshold.

This part of the future space is important to conceptualize because it's when labor substitution really takes place. That has profound economic and political impacts. People are no longer in the role of the natural human in the loop that exists when humans are doing the task.

There's also a performance threshold that we cross when whatever was previously possible with humans is no longer the bounding set of what's possible. When AI is better than humans at a task, new qualitatively different things become possible.

AlphaFold is an example. That represents an important moment in history when new technologies come online.

Nathan Labenz

The paper alludes to the idea that as you're approaching AGI, you could start with a system that's very strong in some areas and relatively weak in other dimensions. There could be an area where it's vastly superhuman, and then the last few pieces come into place as it begins to approach human level.

You could potentially try to change—or choose—which dimensions those are. Do you want the system to be strongest on cooperative AI first, and then add technical knowledge, practical know-how, agency, and so on? Or do you start with agency and enormous factual knowledge, then add cooperation later?

Do you think that's an important or underrated idea? Do you have preferences about what we should add to AGI early versus late?

Allan Dafoe

I think it's a useful heuristic to see AGI as a space with multiple paths to reach it.

Coming back to technological determinism, there may be different trajectories, and it matters not only what kind of AGI we build—AGI isn't a single kind of system. It's a vast set of systems, a corner of a high-dimensional intelligence space.

It also matters how we get there and what the trajectory is.

We might characterize 2 trajectories. In one, AI is relatively incapable at physics or materials science but very good at cooperative skill. In another, it's superhuman at materials science but amateur at cooperative skill.

We can ask which world is safer or more beneficial. Some in the safety community prefer the latter. The story is that it unlocks economic and health advances while leaving us with AI systems that are socially and strategically simplistic, making them easier for humans to manage. They won't outwit us.

The cooperative AI bet takes the opposite tack. A world with accelerating technological and capability advances will generate benefits but also disruptions that we may not be able to adapt to and manage at the rate they're arriving.

We still have cooperation problems that we need to solve. We'd be better off betting on making AI systems more cooperatively skilled than they otherwise would be.

Nathan Labenz

If this is contested by people who are highly informed, should we focus on it more and try to reach agreement? Or is it likely to remain unclear until the day we have to decide?

Allan Dafoe

It's an important part of the cooperative AI research agenda to articulate the argument, host the debate, and try to make sense of it.

You don't want to invest too much time and resources going down a path you're not sure is beneficial. That applies to all these prosocial bets. It is worth investing a significant share of our effort in making sure the bet is beneficial.

Nathan Labenz

Where can people go to learn more about this debate? I think it's the Machine Intelligence Research Institute folks who are more wary of a high-social-skill, high-cooperation early agenda. You and your co-authors are advocating for cooperative AI in particular.

Allan Dafoe

For cooperative AI, I'd point people to the Cooperative AI Foundation. Reach out to anyone there or to me. We invite people to join the conversation.

On the contrary point, Eliezer Yudkowsky and Nate Soares have expressed this view fairly strongly in the past. They might be interested in continuing to argue that position, though I can't speak for them.

On a third pole, Paul Christiano and Carl Shulman have articulated the case for the super-cooperative AGI hypothesis. They think it's very important that AGI systems be able to cooperate with each other, so in that sense they agree with the cooperative AI bet. But they think it will come as a byproduct of AGI and don't think it's necessarily an area we need to invest in separately.

Nathan Labenz

Up until now, we've been talking to you in a personal capacity about work you did before you got involved with DeepMind.

Let's talk to you as a DeepMind employee for a bit. DeepMind published a paper that you were heavily involved with earlier this year called “Evaluating Frontier Models for Dangerous Capabilities.”

Most listeners will be familiar with evals. We're trying to measure what capabilities a model has soon after it's trained, so we're not blindsided later and know what safeguards are necessary.

Having accurate evals that can determine what a model can and can't do facilitates what's recently been called “if-then commitments.” It's easier to agree that if we have a model able to do some powerful or interesting thing, we'll put in place an appropriate safeguard to ensure it doesn't end badly.

These are some of the evals that DeepMind is developing to slot into Google's Frontier Safety Framework, which I guess is the approach DeepMind takes to deploying very powerful AI and, eventually, AGI while avoiding significant downsides.

The paper describes these evals and reports results from the previous model, Gemini 1.0. I think there are 5 categories of evals: persuasion and deception, cyber capabilities, self-proliferation, self-reasoning, and biological threats, which you don't cover in this paper because that's a separate area.

What's an eval in the paper that you think is new or interesting?

Allan Dafoe

One interesting eval is a self-reasoning eval. It's trying to assess the extent to which the model understands its situation and how it could modify that situation to achieve a goal.

We give the model the goal of writing a text output, such as an email, that must include certain words. Those words are prevented by a filter, so whenever the model tries to write the email, some words are blanked out.

The question is whether the model notices that it's failing at its task and investigates whether it can do anything about it. In principle, the model can look at its directory. We give it access to a fake directory where it sits, and there's a file called “filtered words” that it can open.

It can see that some of the words it's trying to say are on the list and remove those words, which enables it to complete the task. It's a toy scenario, not the model's actual host environment, but it gives us a chance to see whether the model notices that it's failing and reflects on the situation enough to do something about it.

Nathan Labenz

It's an easier case than the real-world one, but we want relatively easy cases now so we can see whether we're getting closer to models being more self-aware and able to take these interventions in real-world cases.

Allan Dafoe

While we're on self-reasoning, another eval asks the model to answer questions about historical events. One event occurs before the model's training-data cutoff, and the other occurs after the cutoff, so the model wasn't trained on it.

The model has access to one search query. The goal is to see whether it uses that query optimally by searching for the event it doesn't know about because it wasn't in its training set.

That tests whether the model knows its own training cutoff and correctly infers that the historical event refers to something after the cutoff, even when that's not explicit in the question.

We aid it in some conditions by telling it what its training cutoff was and giving it the dates of the 2 events. Again, it's a simplistic setting, but it measures whether the model is aware enough of its own condition to request more information optimally.

Nathan Labenz

I read the paper a month or 2 ago. It was very readable. It looks like an official PDF, so there may be some activation energy in loading it up and reading it, but a lot of what Google is putting out is accessible to people who get to this stage of the interview.

What was the result? How did Gemini do?

Allan Dafoe

Overall, there were 4 risk domains: persuasion, cybersecurity, self-proliferation, and self-reasoning. We scored Gemini on a subjective 5-point scale. The components of the overall score were more quantitative and fixed.

Overall, it scored 3 out of 5 on persuasion, between 1 and 2 on cybersecurity and self-proliferation, and 2 or 1 on self-reasoning. Its self-proliferation score was 2 out of 5.

The paper is old now because the field is moving so rapidly. I encourage readers to look at the latest system cards and research in this area, but there's still a lot of content, lessons, and productive directions in it.

We received a lot of praise from external experts. Apollo Research listed it as one of its favorite papers, which was a nice recognition.

Nathan Labenz

On persuasion and deception, 3 out of 5 feels natural because we know LLMs can be charismatic and persuasive, probably just from using them.

Cybersecurity was 2 out of 5, right? That surprised me slightly because I've heard that one of the most useful economic applications of LLMs so far has been programming. That seems adjacent to hacking your way out of, or understanding, the system you're operating in.

Why did the model do relatively poorly on the cyber tasks?

Allan Dafoe

We would expect large language models to be good at cyber because it's their home turf. Models today are quite good at programming and coding assistance, so this is something to watch.

I think cybersecurity is a capability domain especially worth following. I could imagine that in 6, 12, or 18 months we'll see significant capabilities in cyber.

I'm using “capabilities” intentionally because it's not just a danger. It's also a benefit. Cyber capability is useful for defense, finding vulnerabilities, and so forth.

One explanation comes from a group at Google called Project Zero, which had a project called Project Naptime. They found that if you build the right support structure around a language model—give it a good interactive environment, tools, and space for reasoning—it performs much better on cyber tasks than it otherwise would.

They compared models with the base model or the model as provided against the model with this support structure and found a significant difference.

Given the access to tools, I think this points to an important aspect of all evaluations, especially dangerous-capability evaluations: capability elicitation.

To what extent is the model you're using for the evaluation the most capable version that model is capable of? At Google DeepMind, we make sure we're at the frontier of capability elicitation, learning from public research and internal work to identify the best techniques for getting models to perform on different tasks.

The logic of dangerous-capability evaluations is that we're trying to measure the maximum capability in a domain with risks. If the maximum isn't too high, we can conclude it's safe because we couldn't reach the capability level that would pose extreme risks.

To get that maximum estimate, we need to give the model its best chance of revealing its maximum capability. This is a challenge because it can take months of experimentation for the model's full capability to be revealed.

The most potent example may be OpenAI's o1 model. To my understanding, it's still based on one of their older foundation models, not a new foundation model. But it uses new scaffolding and fine-tuning that allow it to perform much better on a number of tasks.

Nathan Labenz

On the self-reasoning task you described, which is trying to see whether a model develops rudimentary situational awareness and understands what behaviors might improve its performance, how did it do?

Allan Dafoe

Not very well.

Nathan Labenz

Pretty rubbish.

Allan Dafoe

There are several other similar evals from Anthropic. These evals are idiosyncratic in the sense that people notice a model seems to have a certain kind of self-awareness, and that may be a property worth reflecting on.

More generally, that's a property of the eval ecosystem. We're trying to make evaluation into a science with pre-specified tests, meaningful outcome variables, and thresholds for what decisions we need to take given an outcome.

The reality is that evaluating models is like evaluating animals or human intelligence. It's multidimensional, complex, and context-dependent. It's more like psychology or sociology and requires a thick, deep understanding.

While we mature the science of evals, we also need a rich exploratory ecosystem of people interacting with models, noticing things, and reporting phenomena.

Nathan Labenz

I've heard you use the term “evals in the wild,” which sounds related. What are evals in the wild, and how could we get more out of them?

Allan Dafoe

You can think of a whole typology of evaluations. The canonical evaluation is a model eval: an automated test you can run in code, ideally on the foundation model, but more realistically on the instruction-tuned model that's been trained to be more useful.

That lets us measure what the model is capable of. But model behavior is complex, so we may need other means of assessing models that are richer, though also more costly and time-consuming.

One step up would be human-subject evaluations, where a person interacts with the model. Most of our persuasion and deception evals have this property. Human subjects interact with the model, and we see what behavior emerges.

Can the model get a human subject to click on a link that, in the real world, could point to a virus? Can it deceive the human? Is it perceived as especially charming, trustworthy, or likable?

These are still evals we can run relatively rapidly, on a scale of days, but they require human interaction.

We can step out further into user studies and observe how a human interacts with the model in a realistic setting. Trusted testers use the model in whatever context they have in their lives, and we see how it performs and where it can be improved.

We can step out even further with evals in the wild. This means looking at specific application domains where models are deployed and seeing what effects they're having and to what extent they're being used.

For example, in assessing cyber risk, we want to ask how capable these models are at helping people perform cybersecurity tasks: finding vulnerabilities, devising exploits, or patching vulnerabilities.

We try to simulate that entirely in the lab, but there are limits. A complementary approach is to look at cybersecurity groups in the wild that are trying to provide a product and see whether they're using Gemini or other models to help them, how helpful the models are, and what the revealed preferences of these actors are.

This kind of observational eval has the advantage of external validity. It looks at real-world settings where actors with their own motivations must decide whether to invest time and money in a model.

One limitation is that it can be a lagging indicator. By the time we see impacts in the wild, the model may have had that capability for months. One implication is that we should find early adopters in the wild so we'll see it relatively early.

Overall, I think this is an important complement to model evals, and it's research that can be done in public by academia, nonprofits, and governments as well as companies.

Nathan Labenz

You were saying earlier that the challenge with these evals is that it's hard to fully elicit all the capabilities a model has. It's possible to miss something because we haven't elicited a skill in the right way, and people might later discover an ability that wasn't picked up in the early evals.

There's a significant false-negative rate, and it seems unlikely that we'll develop a science of evals that escapes that issue anytime soon.

I've had a problem with many of the frameworks related to using evals in responsible-scaling policies. Everything ends up riding on the eval being accurate. If it fails to pick up a dangerous ability that's actually there, the entire system falls apart.

Everyone says that we need defense in depth: overlapping systems, so that when one fails, the Swiss-cheese model gives you a gap in one layer but the next layer catches it.

What depth are we thinking of adding? What other procedures can complement this?

Allan Dafoe

This is an important observation. The main response, in defense in depth, is staged release or staged deployment.

We don't go directly from no deployment to irreversible proliferation of models whose capabilities we don't fully understand. The typical approach at Google DeepMind is internal use first, followed by a trusted-tester system that scales in the number of testers, and then broader deployment, but not necessarily general access.

Even after general access—when anyone in the world can use the model—the model weights may remain protected. If we later learn that the model has capabilities we didn't understand, we can change the deployment protocol, add guardrails or monitoring, or modify the model if necessary.

This relates to the important debate about open-weight models, sometimes called open source, though open weight is more accurate because the key thing being published is the model weights.

There are many arguments for open weights. They provide a tool to the broad scientific and development communities to build on. If you have the weights, you can do whatever you want: fine-tune them, perform mechanistic interpretability, distill them, and so forth.

The major downside is the irreversibility of proliferation. Once the weights are published, people save them, and they can be placed on the dark web even if you later stop publishing them.

You've given the model to scientists and developers, but you've also given it to every potential bad actor who wants to use it.

Google open-sources many advances, including AlphaFold's protein predictions and earlier large-model work. In all these cases, Google tries to open-source technologies that are broadly beneficial.

With frontier models, Google and Google DeepMind recognize that we need to proceed carefully, gradually, and responsibly—deploying them only to the extent that we can do so responsibly.

Nathan Labenz

With standard evals, you find out the result when you're training the model or soon afterward. Evals in the wild have better external validity but come later because you have to distribute the model and see how it's used.

You also talk in the paper about an ideal situation where we know not only that a model has an ability when it appears, but exactly when that ability will arrive years in the future. Then we could avoid wasting effort on something unnecessary now while knowing which preventive measures to put in place before the ability arrives.

There are forecasting techniques, prediction markets, and ways to aggregate forecasts from experts and laypeople. Is there low-hanging fruit in forecasting the real-world usefulness of AI models that people aren't taking advantage of?

Allan Dafoe

This is an exciting future direction for understanding, evaluating, and mitigating dangerous capabilities. It applies not only to dangerous capabilities but also beneficial ones.

Labs already forecast some model properties using scaling laws. There's an extraordinary ability to predict certain kinds of loss or performance on low-level benchmarks, such as how well the model predicts the next token.

It's incredible how precisely we can predict how a model will improve as we give it more compute and data, or use a different architecture.

What's been done less is predicting complex downstream tasks, such as performance on biological, cyber, or social-interaction tasks.

There's promising research. One recent NeurIPS spotlight paper, “Observational Scaling Laws,” discusses ways of doing this. It looks at a family of models and their performance on downstream tasks, adjusts for the compute efficiency of the architecture, and claims that this allows us to predict how future models will perform on more complex, real-world-relevant tasks given different compute and data.

There should be much more research on observational scaling laws.

The second approach is subjective human forecasting. Superforecasters have demonstrated track records of being calibrated. If they say something has a 75% chance, it occurs roughly 75% of the time. They're also typically accurate: they assign high probabilities to things that happen and low probabilities to things that don't.

That's a form of expertise in reading evidence, integrating it, and weighting base rates.

In this paper, we employed superforecasters from the Swift Centre to look at our eval suite and predict when models would reach certain levels of performance.

I think this is an exciting part of the paper. It would be great if future work did something similar, at least until we show that it isn't effective. We need more research to see how well the method works, but so far it seems promising.

They gave us thoughtful forecasts that seemed reasonable and better than a subjective qualitative judgment that wasn't expressed numerically.

Nathan Labenz

Given the importance and tractability of forecasting when AI models will become economically and practically useful for different tasks, it's surprising that people aren't doing this already. You don't have to be inside DeepMind to run tournaments or aggregate judgments.

Maybe this is an invitation for listeners to get involved.

The surprising thing about forecasting abilities is that we're very good at anticipating when we'll have a particular level of loss, that measure of inaccuracy in the model. But that hasn't translated into predicting when people will actually want to use the model for a given task.

The relationship between loss and usefulness is probably highly nonlinear. There must be some S-curve. With self-driving cars, a 99%-accurate car is a heap of junk from a practical point of view. You may need 99.999%, and you don't know how many nines you need before it can actually get on the road.

Allan Dafoe

Even when you have a complex, real-world-resembling task, there's still a large inferential step from that task to real-world impact.

You need to account for the cost of employing the AI, whether it fits with the existing ecosystem and environment, and what other tasks humans perform. Even if the AI can do many tasks, it may not be cost-effective to integrate it.

Nathan Labenz

Let's push on and talk about the Frontier Safety Framework. I guess it's Google's or Alphabet's equivalent of the Responsible Scaling Policy at Anthropic or the Preparedness Framework at OpenAI.

They're similar approaches to safety. At this point, the Frontier Safety Framework is only 6 pages and promises to produce a full suite of if-then commitments next year, ahead of releasing models that would approach some of these dangerous capabilities.

When I spoke with Nick Joseph from Anthropic, I pushed him on whether it's responsible to leave the question of how dangerous these models are to the companies developing them.

Obviously, there's an enormous conflict of interest in evaluating faithfully whether a product is dangerous enough that the company shouldn't release it. Ideally, surely you'd hand it off to an external party that could take a more impartial perspective.

I haven't found serious people who say total self-regulation is the future. Is Google DeepMind taking steps toward making some of these decisions external and taking them out of the company's hands?

Allan Dafoe

Google DeepMind and Google's perspective is that governance of frontier AI ultimately needs to be multilayered. There will be an important role for government regulation, third-party evaluations, nonprofits, and external assessments of frontier models.

In many ways, the companies have been pushing the conversation forward, which is an invitation to academia, nonprofits, and governments to participate. It's less that Google is saying this is only for us to do and more that this is our first best guess at what future governance of frontier systems could look like.

Google contributes to conversations hosted by AI safety institutes. We've been in extensive dialogue with the UK AI Safety Institute and the US AI Safety Institute. There are a number of AI safety institute conversations happening in November that Google is contributing to.

More broadly, the goal shouldn't be for each company to have its own framework. That wouldn't benefit companies or society. We need to converge on an appropriate standard that balances equity, safety, responsibility, innovation, appropriate deployment, and societal oversight.

Another contribution is the Frontier Model Forum, an organization for companies developing frontier models to work through standards and safety for frontier models.

My team was directly involved in establishing the Frontier Safety Fund, another differential technological development bet: Can we put more money on the margin toward safety standards and research around frontier models?

Nathan Labenz

This is slightly off topic, but something that's difficult for frontier-safety frameworks and dangerous-capability evals to pick up is structural risk, which is an idea you've pioneered.

Can you explain what structural risks are and why it's difficult to address them at the company level rather than the government level?

Allan Dafoe

A paper that Remco Zwetsloot and I published in 2019 tried to add an intellectual tool to the toolbox for thinking about risks.

Up until then, the primary conceptual framework was to refer to misuse harms or accident harms. Misuse is when a bad actor or criminal intentionally uses a tool in a certain way to cause harm.

There's a subcategory I would call criminal negligence, where the person doesn't intend to cause harm but would have averted it if they were more prosocially motivated or responsible. Someone causing a car accident because they were texting wasn't intending to cause the accident, but the fault falls on them.

Accident harms are situations where the proximal cause is some property of the technology. An engineer could have detected the risk and added a warning indicator or guardrail to avert it.

Much of the conversation centers on those 2 perspectives: accidents or misuse.

Structural risk isn't an exclusive category. It's an upstream category. The conditions under which a technology is used are shaped by social structures that make accidents or misuse more or less likely. We need to look upstream at those structures to see whether we're in a regime that promotes certain harms.

One example is the Cuban Missile Crisis. You had legitimate leaders of the United States and the Soviet Union acting on behalf of their nations in ways that exposed their countries and the world to the risk of nuclear war.

That was neither misuse nor a technical accident. The nuclear weapons weren't built incorrectly, and there wasn't a flaw in their design. Nor were these criminals using the technology for an illegitimate purpose. They were acting within a geopolitical structure that made them engage in dangerous brinksmanship.

There are many examples where the structure of the economy or different interactions leads a system to be deployed in ways that create greater accident or misuse risks.

Nathan Labenz

Evals might pick up that a model has an ability that could feed into one of these structural risks, but because the risk occurs at the societal or global level, it's not really possible for DeepMind to evaluate the entire societal effect. That's difficult for anyone, but thinking about it is the remit of governments because they're best placed to address it.

Allan Dafoe

Structural risks are often at the societal level. They're typically emergent and indirect effects of technology, which means they're hard to read off the technology itself.

Another example is the risk from the railroad. It's hard to read the geopolitical impact directly from steel rails and trains. According to several historians, railroads increased first-strike advantage because whoever mobilized first had a greater advantage. They also made the decision to mobilize irreversible because it was difficult to stop the deployment schedule.

These were consequences of the railroad that were very indirect. A railroad-level eval wouldn't tell you about them because the implication bounces off so many different aspects of society.

Structural risks therefore typically require a societal-level solution from government. If the risk spans nations, it requires an international-level solution.

Nathan Labenz

When I read the PauseAI people and those sympathetic to them, a key underlying grievance is the sense that frontier-model and AGI research is democratically illegitimate.

It's exposing everyone in the UK, US, and around the world to a disaster that could kill us, but we haven't been consulted at a level commensurate with the risk. Even if you discussed it in the UK and US, where much of the work is happening, people overseas who are also exposed to the risk are barely consulted.

That would at least suggest that we should have affirmative signoff by Congress or Parliament that the benefits exceed the risks and that we should proceed. Ideally, we should talk to people elsewhere who will also be heavily affected.

What do you make of that frustration?

Allan Dafoe

Many people are concerned about the direction technology is going, specifically AI. That's an important perspective to take seriously and address.

To some extent, it's a matter of education, and to another extent, it's a matter of improving governance and making sure the benefits are realized for many people. For many people, their concern is labor displacement or other near-term impacts.

An important response by companies is to make sure the benefits are realized. This is part of why Google and Google DeepMind invest so much in AI for science, most notably AlphaFold, which was the first AI contribution to win a Nobel Prize for its impact.

There was also a Nobel Prize in physics for advances in AI. AlphaFold, for listeners who don't know, is a narrow AI system that enabled cost-effective prediction of the structure of virtually every protein relevant to science and medicine. It seems to have unlocked many potential advances in medicine.

In the future we'll see the benefits that come from this kind of development. Advocates for AI often point to AI for science. Dario Amodei, the CEO of Anthropic, recently expressed the view that the benefits of AI for science are potentially very large and should be encouraged.

The broader question of the right way to democratically develop technology is challenging. In my earlier scholarship, from the 1960s or 1970s, there were ideas about citizen councils or citizen juries for technology.

You would have a representative group of citizens paid to take the time to learn about a technology and reflect on how it should be developed and deployed in their community.

More recently, there have been work streams that carry on this tradition by creating pluralistic assemblies and hosting democratic conversations. That requires providing the space and resources to educate the community, creating an institutional framework for people to express their views, and allowing those views to aggregate into guidance for policymakers rather than becoming noise.

Different countries are approaching the democratic governance of technology in different ways. We have freedom of the press, freedom of assembly, social media conversations, nonprofits doing research and stating their positions, and government agencies such as AI safety institutes and other regulatory bodies looking at technology.

I wouldn't characterize this as companies unilaterally proceeding without any democratic involvement. But it is important for people involved in PauseAI and other groups to find ways to engage with the democratic process productively.

The question is what mitigations are appropriate for powerful models. I don't think that can be fully answered by companies or even by a narrow government agency.

We need the scientific community, people building medicines who can see that side of the ledger, people building cyber defenses, and people focused on biological and cyber risks, all in conversation together and hopefully reaching a scientific and policy consensus about how to proceed as models gain new capabilities.

Nathan Labenz

One thing that has raised concerns about democratic legitimacy is that almost all the AI companies, perhaps with the exception of Meta, have talked about wanting binding legal constraints on what they can do, heavy government involvement, and oversight.

But they haven't yet lobbied in favor of specific legislation that would truly bind them, be costly, and restrict what models they could train. They may even have passed up opportunities to argue in favor of proposed legislation that would constrain them.

The worry is that companies will always say they're in favor of these things, but the day when they'll support a specific costly constraint—when they're willing to put handcuffs on themselves—may never arrive.

I'm not sure whether you want to comment on that. But if companies said, “This is the specific regulation we want to bind ourselves to, and we want it to apply to everyone,” I think it would reassure many people who aren't sure how much they can trust these powerful entities.

Allan Dafoe

I wouldn't put it quite as starkly. Speaking for Google, Google has engaged proactively in conversations that affect its freedom of action, most notably the White House commitments.

There have also been commitment-like moments around the UK AI Safety Summit and a second round of commitments at the AI Seoul Summit. In both cases, they pointed toward Frontier Safety Framework developments.

Google committed to red-teaming, self-evaluation, and working on mitigations for risks as they emerge.

I was impressed by how seriously Google's internal operations took the White House commitment. After the commitments were signed, there was a real effort to identify all the commitments, spell out what they meant, and establish workstreams to make sure we lived up to them to a significant extent.

Similarly, the Frontier Safety Framework isn't cheap talk. It represents Google's and Google DeepMind's assessment of the nature of the risks and an attempt to develop a safety standard for future risk assessment. It's a form of proto-regulation or safety standard that could later significantly constrain the industry.

There's also the EU AI Act, which is currently negotiating its code of practice. That's a meaningful form of regulation. Google is participating and helping the EU think through the right way for that code of practice to operate.

Nathan Labenz

A worry I used to have, and perhaps have slightly less now, is that improvements in algorithmic efficiency and the proliferation of compute could make these rules and frameworks less effective.

Even if a large organization like DeepMind has stringent controls, evals, and precautionary measures, if another organization can train a model that's as powerful as the frontier model for $10 million, $1 million, or even $100 million, then there could be widespread proliferation of dangerous capabilities.

If any actor isn't willing to follow the rules and defects from them, that could still be disastrous.

How worried are you about that failure mode?

Allan Dafoe

The proliferation of model capabilities is a key parameter for how things will go and for the viability of governance approaches.

This is why I would return to the question of open-weight models, which are probably the most significant source of proliferation of frontier-model capabilities.

Leaving aside open weights, there's a secondary point: the exponentially decreasing cost of training models. Epoch AI, which does very good estimates of trends in hardware and algorithmic efficiency, finds that algorithmic efficiency is increasing by about 3 times per year.

The cost of training a model therefore decreases by roughly 10 times every 2 years. If that trend persists, a model that costs $100 million to train today will cost $10 million in 2 years and $1 million 2 years after that.

That leads to many more actors having models of a given capability. Some analysts view this as very concerning because novel capabilities will diffuse quickly and be employed by bad or irresponsible actors.

It is complicated. Control of a technology is easier when it doesn't have an exponentially decreasing cost function, and there are other sources of diffusion in the economy.

However, 2 years is a long time. A model from 2 years ago is significantly inferior to the best models today. We may be in a world where the best models are developed and controlled responsibly and can be used for defensive purposes against irresponsible uses of inferior models. I think that's the wise condition.

Nathan Labenz

That picture makes me nervous because it means you always need a hedge of the most recent model, ensuring that what is now an obsolete but still powerful model from 2 years ago isn't causing havoc.

But I can't think of an alternative.

Allan Dafoe

There are other ways to advantage defenders. This was Mark Zuckerberg's argument in favor of open-weight models. He would say that even if bad actors have access to the best models, there are more good actors or the good actors have more resources. They can outspend bad actors.

For every dollar spent by a bad actor, there might be $100 spent by good actors building defenses against misuse. That could be the case, depending on the offense-defense balance.

How costly is it for an attacker to cause damage compared with the cost a defender must spend to prevent or repair it? If you have a 1-to-1 ratio, then as long as good actors have more resources, they can protect themselves or repair the damage cost-effectively.

But the ratio isn't always favorable. With biological weapons, for example, the ratio may be very skewed. A single bioweapon can be difficult to protect against. You have to develop and distribute a vaccine, and a lot of damage may occur before it's widely deployed.

In general, offense-defense balance is another area worth studying. AI has too often drawn an analogy from computer security, where vulnerabilities are easily patched.

The typical response in computer security is to encourage vulnerability discovery, because you can roll out a patch and make the new operating system or software resilient to the vulnerability.

Not all systems have that property. Biological and social systems are hard to patch.

Deepfakes don't have a simple patch, but we can develop an immune system in which we know that a video call isn't sufficient authentication of someone's identity for an instruction to transfer millions of dollars.

Humans have a lot of inertia in their systems. It's costly to build new infrastructure and train people to use it correctly.

Nathan Labenz

You have a couple of nice papers on the offense-defense balance that we couldn't fit into the margins of the page this time. We may have to wait for a third interview.

We're almost out of time. I wanted to talk about advice for people in the audience.

As I mentioned, back in 2018 you told people they should enter AI safety, governance, policy, and related work. That was excellent advice at the time and may still be excellent advice now.

Back then, you said there were technical people in the area but fewer social scientists. There were gaps across disciplines and areas of knowledge that weren't entering AI governance because people weren't hearing about it or didn't see it as relevant.

Is social science still a big gap? Are there other gaps where we'd particularly benefit from people with a specific kind of training or interest?

Allan Dafoe

It remains a huge gap, and it's hard to name a single area. We need more political scientists, economists, historians, ethicists, philosophers, sociologists, and ethnographers.

AI is having and will have profound impacts across society. We need experts in all aspects of society to make sense of it and help guide Google and governments in how to proceed with these technologies.

Nathan Labenz

Are there particular roles you're hiring for right now? Maybe not by the time this episode comes out. What sorts of roles does DeepMind have difficulty hiring for, even despite its prominence?

Allan Dafoe

Talent is in high demand everywhere.

On my side, we're looking for experts in international politics, domestic politics, technology governance domestically and internationally, technology forecasting, macro trends, and predicting the impacts of technology.

There's a lot of work on agents that needs to happen: thinking about autonomy, risks, safety, and assistance; how these systems can be deployed; and the rich space of interactions with humans.

The ethics team at Google DeepMind is hiring as well. There are many ethics-heavy questions and areas of work, along with responsibility and policy work.

That's just on the nontechnical side. Within technical safety, there's a large portfolio of work that needs to be done.

Nathan Labenz

This show talks about AI a lot. Unfortunately, we have a tendency toward doom and gloom. We talk about risks and downsides because they're neglected in the broader conversation, or perhaps because they're the only thing stopping us from achieving the enormous gains that may come from AI and AGI.

In the interest of balance, what applications of AI are you excited about? What would you really like to see sooner rather than later?

Allan Dafoe

It's an important question, and we should keep it at the top of our minds.

The reason to build AI, advanced AI, and AGI is that the benefits could be profound—and will be profound if they're built safely.

Personally, I found riding in a Waymo self-driving car to be impactful. Tourists to San Francisco report back on the historical moment of riding in a car that's driving itself. This is a new phenomenon.

It's incredible that safety has reached the number of nines required to do this safely. Waymo recently released a safety report documenting how many fewer crashes there are involving injuries or police responding to the scene.

There were 2 times fewer incidents involving police coming to the crash scene and, I believe, 6 times fewer crashes involving an injury.

That's one quantified benefit. If we can have fewer crippling car injuries, that would be profound.

I'm imagining a world where we don't need to expend so many resources building cars that mostly sit in parking lots. We wouldn't need to dedicate urban space to cars waiting for us to finish the workday. We could reclaim that space for more productive purposes.

That's one technology, and of course there are challenges we need to think through.

Medicine and health are another domain that could be profoundly beneficial. I look forward to the day when I can consult a doctor in my phone, give it my symptoms, and have it tell me whether they're reasonable or coherent and whether I need to worry.

It could advise me when it's time to see a doctor, saving the medical establishment the costly time of seeing patients who don't need to be seen and allowing human doctors' time to be used more effectively.

Nathan Labenz

I already use an LLM called Claude for health advice. Isn't Med-PaLM 2, developed by Google, the state of the art in this kind of medical consultation?

Allan Dafoe

Google's work on language models for medicine uses the broad name MedLM. Google has an extensive health portfolio.

Nathan Labenz

That's not publicly available yet, right?

Allan Dafoe

My understanding is that MedLM is available to certain users.

Nathan Labenz

Something I think about a lot is that I have a child who's almost 1 year old. They'll be in kindergarten and then school in a few years. What is a typical school like in your district? It isn't necessarily going to offer the best education, and there are unpleasant things about school.

I'm hoping that by the time my child attends school, there will be an option for tutoring and assistance from AI models that could be as good as the best teacher in principle, if designed well.

There may be issues with classroom management, but AI could be much more engaging, entertaining, and interesting than a real-life first-grade teacher.

You also have 3 kids. I imagine your oldest is approaching school age, if not already. Do you think they'll actually be taught by an LLM teacher soon?

Allan Dafoe

AI in education is a huge opportunity, for the reasons you mentioned.

Tutoring is hugely impactful, as are small class sizes. An AI tutor that complements the school system could be very beneficial. It could flag mistakes, point out where the student went wrong, and show different ways of proceeding.

Having that available in a more engaging or compelling way could be beneficial. We've seen this with the internet in general. It has provided resources for learning, and platforms like Khan Academy and many pedagogical videos have been beneficial for student learning.

Another area I want to call out is AI for sustainability. Google DeepMind has interesting projects that sometimes read like science fiction.

One is controlling the plasma for nuclear fusion. If we can do that more effectively, fusion could become a more viable energy source, which would be extremely beneficial for sustainable energy production.

DeepMind has also advanced weather prediction, making it cheaper, faster, and more effective. That can benefit event planning, wind power, solar power, and other sources.

There are advances in materials science, finding new materials for solar power or other parts of the economy that could make things more efficient.

Google had a result on optimizing flight paths to reduce contrail production, which is apparently a significant source of greenhouse-gas effects.

There are also data-center optimizations. Google DeepMind reported several years ago on significant gains in data-center cooling and energy use.

All these efficiency gains in energy use and production can add up and help humanity address significant challenges.

Nathan Labenz

We haven't even mentioned AlphaFold 2. From what I've heard, it's incredibly useful for medical research, drug design, and so on.

We just have to do everything discussed in the last few hours, and then we can enjoy this wonderful brave new world with all these fantastic new products and a much better quality of life.

Allan Dafoe

Fingers crossed.

Nathan Labenz

My guest today has been Allan Dafoe. I look forward to talking to you again next time on the show.

Allan Dafoe

Thanks, Rob. I look forward to it.

Agency over AI? Allan Dafoe on Technological Determinism & DeepMind's Safety Plans, from 80000 Hours | BidClub